Training methods, devices, storage media and equipment for costume conversion models

By training an outfit transformation model and using a generator and discriminator to generate predicted outfit transformation image samples, the problems of outfit inability to be changed and poor real-time performance in existing technologies are solved. This achieves efficient outfit transformation effects with natural replacement and adaptation to complex scenes, and is suitable for real-time scenarios.

CN116486210BActive Publication Date: 2026-05-26GUANGZHOU BOGUAN TELECOMM TECH LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU BOGUAN TELECOMM TECH LTD
Filing Date
2023-04-27
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing clothing-changing methods based on 3D human body points cannot change the original attire, have poor real-time performance, cannot adapt to complex scenes, and have poor clothing-changing effects, making them unsuitable for real-time scenarios such as live streaming.

Method used

By acquiring image samples and key point maps, a costume transformation model is trained using a generator and a discriminator to generate predicted costume image samples. The model parameters are then updated using the discriminator's discrimination results, enabling natural costume replacement and adaptation to complex scenes.

Benefits of technology

It enables a natural replacement of the original attire, adapts to complex scenes, improves the dressing effect, and can be applied to real-time scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116486210B_ABST
    Figure CN116486210B_ABST
Patent Text Reader

Abstract

This disclosure relates to the field of computer technology, specifically to a method, apparatus, computer-readable storage medium, and electronic device for training an outfit conversion model. The method includes: acquiring an image sample to be processed; acquiring a key point bitmap; inputting the image sample to be processed and the key point bitmap into a generator in a model to be trained to obtain a predicted outfit conversion image sample that converts the original outfit corresponding to the portrait portion into the target outfit; using a discriminator to distinguish between the predicted outfit conversion image sample and a reference image sample to obtain a discrimination result; updating the neural network parameters of the model to be trained based on the predicted outfit conversion image sample, the reference image sample, and the discrimination result, so that the generator in the trained model is used as the outfit conversion model. The technical solution of this disclosure can solve the problem of poor outfit conversion effects in related technologies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and more specifically, to a method for training a costume conversion model, a costume conversion method, a costume conversion model training device, a costume conversion device, a computer-readable storage medium, and an electronic device. Background Technology

[0002] With the rapid development of software and hardware, clothing-changing technology is being applied more and more widely in various fields. For example, clothing-changing technology is widely used in areas such as smart virtual fitting, virtual avatar reconstruction, and smart clothing shopping.

[0003] In related technologies, clothing replacement based on 3D human body points is commonly used. Specifically, clothing replacement based on 3D human body points refers to the technique of replacing clothing on a human body model using the coordinate positions of 3D points.

[0004] However, the solutions in the relevant technologies cannot change the original attire, but can only cover it up. They cannot achieve the ideal dressing effect, have poor real-time performance, and cannot be applied to real-time scenarios such as live streaming. Furthermore, the dressing effect is poor in complex scenarios (such as when the attire is more complicated and there are more movements).

[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] The purpose of this disclosure is to provide a method and apparatus for training a costume conversion model, a computer-readable storage medium and an electronic device, which can solve the problem of low training efficiency of costume conversion models in related technologies.

[0007] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part from practice of this disclosure.

[0008] According to a first aspect of this disclosure, a method for training an outfit conversion model is provided, characterized in that the method includes: acquiring an image sample to be processed; wherein the image sample to be processed includes a human portrait portion and an original outfit corresponding to the human portrait portion; acquiring a key point map corresponding to the human portrait portion; wherein the key point map is used to indicate key points of the human portrait; inputting the image sample to be processed and the key point map into a generator in a model to be trained to obtain a predicted outfit conversion image sample that converts the original outfit corresponding to the human portrait portion into a target outfit; wherein the model to be trained has a reference image sample, the reference image sample including a reference human portrait portion and a target outfit corresponding to the reference human portrait portion; inputting the predicted outfit conversion image sample and the reference image sample into a discriminator in the model to be trained, and using the discriminator to discriminate the predicted outfit conversion image sample and the reference image sample to obtain a discrimination result; updating the neural network parameters of the model to be trained based on the predicted outfit conversion image sample, the reference image sample, and the discrimination result, so as to use the generator in the trained model to be trained as the outfit conversion model.

[0009] According to a second aspect of this disclosure, a costume conversion method is provided, characterized in that the method includes: acquiring an image; wherein the image includes a human portrait portion and an original costume corresponding to the human portrait portion; inputting the image into a costume conversion model to obtain a costume conversion image that converts the original costume corresponding to the human portrait portion into a target costume; wherein the costume conversion model corresponds to a target costume, and the costume conversion model is trained by a costume conversion model training method as described in any of the above.

[0010] According to a third aspect of this disclosure, a training apparatus for an outfit conversion model is provided, characterized in that the apparatus comprises: an image sample acquisition module for acquiring an image sample to be processed; wherein the image sample to be processed includes a human portrait portion and an original outfit corresponding to the human portrait portion; a key point map acquisition module for acquiring a key point map corresponding to the human portrait portion; wherein the key point map is used to indicate key points of the human portrait; and a prediction image acquisition module for inputting the image sample to be processed and the key point map into a generator in a model to be trained to obtain a predicted outfit conversion image sample that converts the original outfit corresponding to the human portrait portion into a target outfit. The training model includes a reference image sample, which includes a reference portrait portion and the target costume corresponding to the reference portrait portion. The discrimination result acquisition module is used to input the predicted costume image sample and the reference image sample into the discriminator in the training model, and the discriminator distinguishes the predicted costume image sample and the reference image sample to obtain the discrimination result. The training model update module is used to update the neural network parameters of the training model according to the predicted costume image sample, the reference image sample and the discrimination result, so as to use the generator in the trained training model as the costume conversion model.

[0011] According to a fourth aspect of this disclosure, a costume conversion model training apparatus is provided, characterized in that the apparatus comprises: an image acquisition module for acquiring an image; wherein the image includes a human portrait portion and an original costume corresponding to the human portrait portion; and an image conversion module for inputting the image into a costume conversion model to obtain a costume image that converts the original costume corresponding to the human portrait portion into a target costume; wherein the costume conversion model corresponds to a target costume, and the costume conversion model is trained by a costume conversion model training method as described in any of the above.

[0012] According to a fifth aspect of this disclosure, a computer-readable storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the costume conversion model training method of the first aspect of the above embodiments or the costume conversion method of the second aspect of the above embodiments.

[0013] According to a sixth aspect of this disclosure, an electronic device is provided, comprising:

[0014] One or more processors; and

[0015] A memory is used to store one or more programs that, when executed by one or more processors, cause the one or more processors to implement the costume conversion model training method of the first aspect of the above embodiments or the costume conversion method of the second aspect of the above embodiments.

[0016] The technical solutions provided in this disclosure may have the following beneficial effects:

[0017] In one embodiment of the present disclosure, a method for training a costume conversion model is provided. This method involves acquiring a sample image to be processed, obtaining a key point map corresponding to the portrait portion, inputting the sample image to be processed and the key point map into a generator in the model to be trained, obtaining a predicted costume conversion image sample that converts the original costume corresponding to the portrait portion into the target costume, inputting the predicted costume conversion image sample and a reference image sample into a discriminator in the model to be trained, obtaining a discrimination result by discriminating the predicted costume conversion image sample and the reference image sample, and updating the neural network parameters of the model to be trained based on the predicted costume conversion image sample, the reference image sample, and the discrimination result, so that the generator in the trained model can be used as the costume conversion model. The solution disclosed herein offers several advantages. First, it allows for the replacement of existing attire, resulting in a more natural look and adaptability to complex scenes, leading to better costume changes. Second, by considering key aspects of the human face during costume changes, the model learns about the human face structure, further enhancing the costume change effect. Third, the trained costume conversion model can convert the attire of a person in an image in real time, making it applicable to real-time scenarios and highly versatile.

[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0020] Figure 1 The illustration shows a schematic diagram of an exemplary system architecture for which an outfit conversion model training method can be applied according to embodiments of the present disclosure;

[0021] Figure 2 A flowchart illustrating an exemplary embodiment of the present disclosure of a method for training an outfit conversion model is shown schematically.

[0022] Figure 3 This schematically illustrates a flowchart of a process for determining a key point map corresponding to a human figure portion using multiple first key points, multiple second key points, and multiple intermediate key points in an exemplary embodiment of this disclosure.

[0023] Figure 4 This schematically illustrates a flowchart of an exemplary embodiment of the present disclosure, in which a predicted costume image sample is obtained based on a costume area mask and intermediate image samples to convert the original costume corresponding to the human figure portion into the target costume.

[0024] Figure 5 This schematically illustrates a diagram showing the output of a costume change area mask and an intermediate image sample via two branches in an exemplary embodiment of this disclosure.

[0025] Figure 6 This schematically illustrates a flowchart of an exemplary embodiment of the present disclosure, in which a target portrait segmentation box is determined based on a mask image of the previous frame's clothing change area, and the original image sample is segmented using the target portrait segmentation box to obtain an image sample to be processed.

[0026] Figure 7 This schematically illustrates a flowchart of a predicted costume change image sample that converts the original costume corresponding to a human figure portion into a target costume based on edge features in an exemplary embodiment of this disclosure.

[0027] Figure 8 This schematic diagram illustrates a reparameterized structure of a generator according to an exemplary embodiment of the present disclosure.

[0028] Figure 9This schematically illustrates a flowchart of an exemplary embodiment of the present disclosure, in which an image sample is obtained by segmenting an original image sample using a target human figure segmentation box to obtain an image sample to be processed.

[0029] Figure 10 This schematically illustrates a flowchart of an exemplary embodiment of the present disclosure in which an image is input into a costume conversion model to obtain a costume image that converts the original costume corresponding to the portrait portion into the target costume.

[0030] Figure 11 This schematic diagram illustrates the composition of a costume conversion model training device according to an exemplary embodiment of the present disclosure;

[0031] Figure 12 This schematic diagram illustrates the composition of a costume conversion device according to an exemplary embodiment of the present disclosure;

[0032] Figure 13 The schematic diagram illustrates a structural schematic of a computer system suitable for implementing an electronic device according to exemplary embodiments of the present disclosure. Detailed Implementation

[0033] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., may be employed. In other instances, well-known structures, methods, apparatuses, implementations, materials, or operations are not shown or described in detail to avoid obscuring various aspects of this disclosure.

[0034] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, or in one or more software-hardened modules, or in different network and / or processor devices and / or microcontroller devices.

[0035] With the development of terminal devices and the gaming industry, a large number of games of different themes have emerged to meet the needs of players. Typically, game logic can be controlled by a server. After receiving instructions uploaded by the client, the server can issue game protocols based on the game logic, and the client will then execute relevant actions according to the game protocols.

[0036] However, when the game logic is complex, multiple game protocols may be sent to the client at the same time. When the client receives multiple game protocols, it will execute these game protocols simultaneously, resulting in chaotic in-game behavior.

[0037] Figure 1 A schematic diagram of an exemplary system architecture for which the costume conversion model training method of embodiments of the present disclosure can be applied is shown.

[0038] like Figure 1 As shown, system architecture 1000 may include one or more of terminal devices 1001, 1002, and 1003, network 1004, and server 1005. Network 1004 is used as a medium to provide a communication link between terminal devices 1001, 1002, and 1003 and server 1005. Network 1004 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0039] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included. For example, server 1005 could be a server cluster composed of multiple servers.

[0040] Users can use terminal devices 1001, 1002, and 1003 to interact with server 1005 via network 1004 to receive or send messages, etc. Terminal devices 1001, 1002, and 1003 can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, and desktop computers. Additionally, server 1005 can be a server providing various services.

[0041] In one embodiment, the execution entity of the costume conversion model training method of this disclosure can be terminal devices 1001, 1002, and 1003. Server 1005 can acquire the image samples to be processed sent by terminal devices 1001, 1002, and 1003, and obtain the key point map corresponding to the portrait portion in server 1005. The image samples to be processed and the key point map are input into the generator in the model to be trained to obtain a predicted costume conversion image sample that converts the original costume corresponding to the portrait portion into the target costume. The predicted costume conversion image sample and the reference image sample are then input into the model to be trained. In the discriminator, the predicted costume change image sample and the reference image sample are discriminated to obtain the discrimination result. Then, the discrimination result, the predicted costume change image sample, and the reference image sample are sent to the terminal devices 1001, 1002, and 1003, so that the terminal devices 1001, 1002, and 1003 receive the discrimination result, the predicted costume change image sample, and the reference image sample sent by the server, and update the neural network parameters of the model to be trained according to the predicted costume change image sample, the reference image sample, and the discrimination result, so as to use the generator in the trained model as the costume conversion model.

[0042] In addition, the instant message processing status display method of this disclosure can be executed through terminal devices 1001, 1002, 1003, etc., to obtain image samples to be processed, obtain key point maps corresponding to the human portrait, input the image samples to be processed and the key point maps into the generator in the model to be trained, obtain predicted costume change image samples that convert the original costume corresponding to the human portrait into the target costume, input the predicted costume change image samples and reference image samples into the discriminator in the model to be trained, obtain the discrimination result by the discriminator to discriminate the predicted costume change image samples and reference image samples, and update the neural network parameters of the model to be trained according to the predicted costume change image samples, reference image samples and discrimination results, so as to use the generator in the trained model as the costume conversion model.

[0043] refer to Figure 2 The diagram illustrates a flowchart of a costume conversion model training method in this exemplary embodiment, which may include the following steps:

[0044] Step S210: Obtain an image sample to be processed; wherein the image sample to be processed includes a human figure and the original attire corresponding to the human figure.

[0045] Step S220: Obtain the key point map corresponding to the portrait portion; wherein, the key point map is used to indicate the key points of the portrait;

[0046] Step S230: Input the image sample to be processed and the key point map into the generator in the model to be trained to obtain the predicted dressing image sample that converts the original costume corresponding to the portrait part into the target costume; wherein, the model to be trained has a reference image sample, which includes a reference portrait part and the target costume corresponding to the reference portrait part.

[0047] Step S240: Input the predicted costume change image sample and the reference image sample into the discriminator in the model to be trained, and use the discriminator to distinguish the predicted costume change image sample and the reference image sample to obtain the discrimination result;

[0048] Step S250: Update the neural network parameters of the model to be trained based on the predicted costume change image samples, reference image samples, and discrimination results, so as to use the generator in the trained model as the costume conversion model.

[0049] In one embodiment of the present disclosure, a method for training a costume conversion model is provided. This method involves acquiring a sample image to be processed, obtaining a key point map corresponding to the portrait portion, inputting the sample image to be processed and the key point map into a generator in the model to be trained, obtaining a predicted costume conversion image sample that converts the original costume corresponding to the portrait portion into the target costume, inputting the predicted costume conversion image sample and a reference image sample into a discriminator in the model to be trained, obtaining a discrimination result by discriminating the predicted costume conversion image sample and the reference image sample, and updating the neural network parameters of the model to be trained based on the predicted costume conversion image sample, the reference image sample, and the discrimination result, so that the generator in the trained model can be used as the costume conversion model. The solution disclosed herein offers several advantages. First, it allows for the replacement of existing attire, resulting in a more natural look and adaptability to complex scenes, leading to better costume changes. Second, by considering key aspects of the human face during costume changes, the model learns about the human face structure, further enhancing the costume change effect. Third, the trained costume conversion model can convert the attire of a person in an image in real time, making it applicable to real-time scenarios and highly versatile.

[0050] Below, we will combine Figure 2 The embodiments will provide a more detailed description of steps S210 to S250 of the outfit conversion model training method in this exemplary embodiment.

[0051] Step S210: Obtain the image sample to be processed;

[0052] In one example embodiment of this disclosure, an image sample to be processed can be obtained. The image sample to be processed includes a human figure portion and the original attire corresponding to the human figure portion. Specifically, the image sample to be processed is an image sample containing a human figure portion and the original attire corresponding to the human figure portion. The human figure portion refers to the part of the image containing a person, which may include the person's facial features, outline, torso, etc. These people can be real people or fictional people.

[0053] Specifically, the original attire may include a top, trousers, shoes, a hat, a bag, glasses, accessories (such as earrings, necklaces, rings, bracelets, etc.), a skirt, gloves, etc. It should be noted that this disclosure does not specifically limit the specific type of attire.

[0054] In one example embodiment of this disclosure, the image sample to be processed can be a portrait image captured by a camera module or a portrait image synthesized artificially. It should be noted that this disclosure does not impose any special limitations on the source of the image sample to be processed.

[0055] Step S220: Obtain the key point map corresponding to the portrait portion;

[0056] In one example embodiment of this disclosure, after obtaining the key point map corresponding to the portrait portion through the above steps, the key point map corresponding to the portrait portion can be acquired. The key point map is used to indicate the key points of the portrait. Specifically, the key point map corresponding to the portrait portion can be acquired through a key point detection algorithm, and the key point map corresponding to the portrait portion includes multiple key points of the portrait.

[0057] For example, key points in a portrait can be key points of the corresponding limbs, such as elbow key points, knee key points, hand key points, head key points, etc.

[0058] It should be noted that this disclosure does not impose any special restrictions on the specific locations of key points in a portrait or the specific methods for obtaining the key point map corresponding to the portrait portion.

[0059] In one example embodiment of this disclosure, a key point map corresponding to the human portrait portion can be obtained through a key point detection model.

[0060] In one example embodiment of this disclosure, the keypoint map includes multiple keypoints, the positions of which are related to the target attire. Specifically, the multiple keypoints in the keypoint map may include human body keypoints and facial keypoints. Human body keypoints refer to multiple keypoints of the human limbs in the portrait portion, and facial keypoints refer to multiple keypoints of the face in the portrait portion.

[0061] For example, 21 key points can usually be collected for human body key points; 160 key points can usually be collected for facial key points.

[0062] Specifically, when obtaining the key point map corresponding to the portrait portion, the key point map corresponding to the portrait portion can be obtained based on the target outfit corresponding to the model to be trained.

[0063] For example, if the target outfit includes glasses, then it is necessary to obtain facial key points, that is, the key point map corresponding to the portrait part includes facial key points; if the target outfit does not include glasses, then it is not necessary to obtain facial key points; if the target outfit only includes pants, then it is not necessary to obtain key points of the upper limbs.

[0064] Through the embodiments of this disclosure, the key points to be acquired can be determined according to the type of the target attire. When the target attire is a partial attire, the acquisition of key points can be reduced, thereby improving the efficiency of attire conversion.

[0065] Step S230: Input the image sample to be processed and the key point map into the generator in the model to be trained to obtain the predicted dressing image sample that converts the original costume corresponding to the portrait part into the target costume.

[0066] In one example embodiment of this disclosure, after obtaining the image sample to be processed and the key point map corresponding to the portrait portion through the above steps, the image sample to be processed and the key point map can be input into the generator in the model to be trained to obtain a predicted costume-changing image sample that converts the original attire corresponding to the portrait portion into the target attire. The model to be trained corresponds to a reference image sample, which includes a reference portrait portion and the target attire corresponding to the reference portrait portion. Specifically, the model to be trained is the base model used to perform the costume conversion task. The costume conversion task refers to the task of converting the original attire corresponding to the portrait portion in an image into the target attire; the target object of the costume conversion task is the portrait portion in the image.

[0067] In one example embodiment of this disclosure, an outfit conversion model can be obtained by training the model to be trained in order to complete the outfit conversion task.

[0068] In one example embodiment of this disclosure, the predicted costume-changing image sample that converts the original costume corresponding to the portrait portion into the target costume means that the model to be trained corresponds to the target costume, and the image sample to be processed contains a portrait portion, which corresponds to an original costume. After the image sample to be processed is input into the generator of the model to be trained, the original costume of the portrait portion in the image sample to be processed can be converted into the target costume, thereby achieving the effect of changing the costume of the portrait portion in the image sample to be processed.

[0069] In one example embodiment of this disclosure, the image sample to be processed and the key point bitmap can be input into the generator in the model to be trained to obtain a predicted costume-changing image sample that converts the original attire corresponding to the portrait portion into the target attire. Specifically, the model to be trained is based on a generative adversarial network (GAN). The model includes a generator and a discriminator. The generator generates the predicted costume-changing image sample that converts the original attire corresponding to the portrait portion into the target attire, while the discriminator distinguishes between the predicted costume-changing image sample and the reference image sample.

[0070] In one exemplary embodiment of this disclosure, the generator can be a deep learning network. For example, the generator can be a residual neural network, which may include a convolutional network, a residual network, and a deconvolutional network cascaded in sequence. After the image sample to be processed and the key point bitmap are input into the generator, the image sample to be processed and the key point bitmap are processed by the convolutional network, the residual network, and the deconvolutional network in sequence to generate a predicted dressing image sample.

[0071] It should be noted that the generator can also be other processing models such as recurrent neural networks, and this disclosure does not impose any special restrictions on the specific structure of the generator.

[0072] In one example embodiment of this disclosure, before inputting the image sample to be processed and the key point map into the generator in the model to be trained, the image sample to be processed and the key point map can be converted into portrait image feature vector and key point feature vector, and then the portrait image feature vector and key point feature vector are input into the model to be trained to obtain a predicted dressing image sample that converts the original costume corresponding to the portrait part into the target costume.

[0073] It should be noted that this disclosure does not impose any special limitations on the specific methods for converting the image samples to be processed and the key point map into portrait image feature vectors and key point feature vectors.

[0074] Furthermore, when inputting the portrait image feature vector and keypoint feature vector into the model to be trained, the portrait image feature vector and keypoint feature vector can be concatenated, and the concatenated portrait image feature vector and keypoint feature vector can be input into the model to be trained. For example, the concat (concatenation) operation can be used to concatenate the portrait image feature vector and keypoint feature vector.

[0075] It should be noted that this disclosure does not impose any special limitations on the specific method of splicing the feature vectors of human portrait images and the feature vectors of key points.

[0076] Step S240: Input the predicted costume change image sample and the reference image sample into the discriminator in the model to be trained, and use the discriminator to distinguish the predicted costume change image sample and the reference image sample to obtain the discrimination result;

[0077] In one exemplary embodiment of this disclosure, after obtaining the predicted costume-changing image sample through the above steps, the predicted costume-changing image sample and the reference image sample can be input into the discriminator in the model to be trained. The discriminator then distinguishes between the predicted costume-changing image sample and the reference image sample to obtain a discrimination result. Specifically, the discrimination result can be used to distinguish whether the image input to the discriminator is a predicted costume-changing image sample or a reference image sample.

[0078] In one exemplary embodiment of this disclosure, the discriminator can be a convolutional neural network (CNN). For example, the CNN may include an input layer, convolutional layers, pooling layers, and fully connected layers. Additionally, a classifier (such as a Softmax classifier) ​​may be added for classification. After the predicted costume change image sample and the reference image sample are input into the CNN, the CNN can extract features from the predicted costume change image sample and the reference image sample, and then further determine whether these features belong to a specific category, thereby achieving discrimination between the predicted costume change image sample and the reference image sample.

[0079] Furthermore, a classifier (such as a Softmax classifier) ​​can be cascaded directly after the generator; or, the first discriminator DB can be other discriminative models such as a Support Vector Machine (SVM) or a Bayesian classifier.

[0080] It should be noted that this disclosure does not impose any special restrictions on the specific structure of the discriminator.

[0081] Step S250: Update the neural network parameters of the model to be trained based on the predicted costume change image samples, reference image samples, and discrimination results, so as to use the generator in the trained model as the costume conversion model.

[0082] In one exemplary embodiment of this disclosure, after obtaining the discrimination result through the above steps, the neural network parameters of the model to be trained can be updated based on the predicted costume-changing image samples, reference image samples, and the discrimination result, so that the generator in the trained model can be used as the costume conversion model. Specifically, the loss function of the model to be trained can be determined by the predicted costume-changing image samples, reference image samples, and discrimination result, and the neural network parameters of the model to be trained can be updated using the loss function. Specifically, the loss function of the model to be trained includes a generation loss function and a discrimination loss function. The generation loss function can be determined by the predicted costume-changing image samples and reference image samples, and the discrimination loss function can be determined using the discrimination result.

[0083] Specifically, the neural network parameters of the model to be trained may include the number of model layers, the number of feature vector channels, and the learning rate. When updating the neural network parameters of the model to be trained, the number of model layers, the number of feature vector channels, and the learning rate can be updated to train the item classification model.

[0084] In one example embodiment of this disclosure, the neural network parameters of the model to be trained can be updated using the backpropagation algorithm. After training, the generator in the trained model to be trained is used as the costume conversion model.

[0085] In one example embodiment of this disclosure, the neural network parameters of the model to be trained can be updated based on the predicted costume-changing image samples, reference image samples, and the discrimination results. When the model to be trained meets the convergence condition, the generator in the trained model is used as the costume conversion model. Specifically, the convergence condition of the model to be trained means that the loss function of the model to be trained reaches the target value. At this point, the model to be trained is marked as having met the convergence condition, and the generator in the trained model can be used as the costume conversion model.

[0086] Specifically, when the loss function of the model to be trained does not reach the target value, backpropagation can be performed, and optimization algorithms such as gradient descent can be used to update the neural network parameters of the generator and discriminator in the model to be trained. For example, when the generator and discriminator are convolutional neural network models, the convolution weights and bias parameters of the convolutional neural network model can be updated, and steps S210 to S250 above can be repeated until the loss function reaches the target value.

[0087] It should be noted that this disclosure does not impose any special limitations on the specific method by which the neural network parameters of the model to be trained are updated based on the predicted costume change image samples, the reference image samples, and the discrimination results.

[0088] In one example embodiment of this disclosure, a first image region corresponding to a first target sub-suit can be obtained in the portrait portion, and multiple first key points can be obtained in the first image region. In the portrait portion, a second image region corresponding to a second target sub-suit can be obtained, and multiple second key points can be obtained in the second image region. Multiple first key points and multiple intermediate key points between the multiple second key points can also be obtained. A key point bitmap corresponding to the portrait portion is determined using the multiple first key points, the multiple second key points, and the multiple intermediate key points. (Refer to...) Figure 3 As shown, determining the key point map corresponding to the portrait portion through multiple first key points, multiple second key points, and multiple intermediate key points may include the following steps S310 to S340:

[0089] Step S310: In the portrait part, obtain the first image region corresponding to the first target sub-dress, and obtain multiple first key points in the first image region;

[0090] Step S320: In the portrait part, obtain the second image region corresponding to the second target sub-dress, and obtain multiple second key points in the second image region;

[0091] In one example embodiment of this disclosure, in the portrait portion, a first image region corresponding to a first target sub-suite can be obtained, along with multiple first key points within that first image region. Similarly, a second image region corresponding to a second target sub-suite can be obtained, along with multiple second key points within that second image region. Specifically, the target suit includes a first target sub-suite and a second target sub-suite, which are not adjacent. The first image region corresponding to the first target sub-suite can be obtained; this first image region can be the area occupied by the first target sub-suite, or it can include the area occupied by the first target sub-suite, meaning the first image region corresponding to the first target sub-suite is larger than the area occupied by the first target sub-suite. Likewise, the second image region corresponding to the second target sub-suite can be obtained; this second image region can be the area occupied by the second target sub-suite, or it can include the area occupied by the second target sub-suite, meaning the second image region corresponding to the second target sub-suite is larger than the area occupied by the second target sub-suite. After obtaining the first image region and the second image region, the first image region and the second image region can be detected to obtain multiple first key points in the first image region and multiple second key points in the second image region.

[0092] It should be noted that this disclosure does not impose any special limitations on the specific methods for obtaining the first / second image regions corresponding to the first / second target sub-suits, or on obtaining multiple first / second key points in the first / second image regions.

[0093] Step S330: Obtain multiple first keypoints and multiple intermediate keypoints between multiple second keypoints; wherein, the first keypoints and the second keypoints are respectively related to the intermediate keypoints;

[0094] Step S340: Determine the key point map corresponding to the portrait part through multiple first key points, multiple second key points, and multiple intermediate key points.

[0095] In one example embodiment of this disclosure, after obtaining multiple first key points in the first image region and multiple second key points in the second image region through the above steps, multiple intermediate key points between the multiple first key points and the multiple second key points can be obtained. The first key points and the second key points are respectively related to the intermediate key points. Specifically, since the first image region and the second image region are not adjacent, there are other regions between them. Multiple intermediate key points between the multiple first key points corresponding to the first image region and the multiple second key points corresponding to the second image region can be obtained, and these intermediate key points are respectively related to the first key points and the second key points. At this point, the multiple first key points, the multiple second key points, and the multiple intermediate key points can be determined as key points of the portrait portion, and a key point bitmap corresponding to the portrait portion can be generated using these key points.

[0096] For example, the first image region is the head region of the portrait, and the second image region is the lower limb region of the portrait. Multiple first key points can be determined in the head region, multiple second key points can be determined in the lower limb region, and multiple intermediate key points can be determined between the multiple first key points and the multiple second key points (the intermediate key points are key points in the upper limb region of the portrait, and the key points in the upper limb region are related to the first key points in the head region and the second key points in the lower limb region, respectively). The key point map corresponding to the portrait is determined through the above three types of key points.

[0097] It should be noted that this disclosure does not impose any special limitations on the specific method of determining the key point map corresponding to the portrait portion through multiple first key points, multiple second key points, and multiple intermediate key points.

[0098] Through the above steps S310-S340, in the portrait portion, a first image region corresponding to the first target sub-dress is obtained, and multiple first key points are obtained in the first image region. In the portrait portion, a second image region corresponding to the second target sub-dress is obtained, and multiple second key points are obtained in the second image region. Multiple first key points and multiple intermediate key points between the multiple second key points are also obtained. The key point bitmap corresponding to the portrait portion is determined by the multiple first key points, multiple second key points, and multiple intermediate key points. Through the embodiments of this disclosure, when the areas occupied by the target dress are two adjacent areas, intermediate key points can be added to enhance the structural integrity between the key points, avoid structural errors, facilitate model training and inference, and improve the performance of the dress conversion task.

[0099] In one exemplary embodiment of this disclosure, the image sample to be processed and the key point bitmap can be input into the generator in the model to be trained to obtain a costume change region mask and intermediate image samples. Based on the costume change region mask and intermediate image samples, a predicted costume change image sample is obtained that converts the original costume corresponding to the portrait portion into the target costume. (Refer to...) Figure 4 As shown, obtaining predicted costume-changing image samples that convert the original costume corresponding to the portrait portion into the target costume based on the costume-changing area mask map and intermediate image samples can include the following steps S410 to S420:

[0100] Step S410: Input the image sample to be processed and the key point bitmap into the generator in the model to be trained to obtain the costume area mask map and intermediate image sample.

[0101] In one exemplary embodiment of this disclosure, the image sample to be processed and the key point bitmap can be input into the generator in the model to be trained to obtain a costume change region mask and intermediate image samples. The costume change region mask indicates the region where the costume change will take place, and the intermediate image samples indicate the temporary result of converting the original costume corresponding to the portrait portion into the target costume. Specifically, the output of the generator of the model to be trained can include two parts: a costume change region mask and intermediate image samples.

[0102] Specifically, a costume change area mask can be used to indicate the area where the costume change is to be performed. For example, if the target costume is gloves, the costume change area mask can be an image of the same size as the image sample to be processed. In this costume change area mask, the position corresponding to the hand area of ​​the human figure in the image sample to be processed can be marked (e.g., blacked out) so that the area where the costume change is to be performed can be represented by the costume change area mask.

[0103] Specifically, intermediate image samples can be used to indicate a temporary result of converting the original costume corresponding to the portrait portion into the target costume. In this temporary result, the costume-changing area is not specifically processed through the costume-changing area mask map. Alternatively, the temporary result still needs to be further processed to obtain a more accurate predicted costume-changing image sample.

[0104] For example, the mask image for the costume change area can be a mask image.

[0105] It should be noted that this disclosure does not impose any special limitations on the specific methods for obtaining the mask image of the costume change area and the intermediate image samples.

[0106] Specifically, after inputting the image samples to be processed and the key point bitmap into the generator of the model to be trained, two outputs are obtained, including the costume change region mask map and intermediate image samples. This can be implemented through two branches, and the specific output design is as follows: Figure 5 As shown, the convolutional layers can use 3*3 or 5*5 convolutional kernels, the linear layers can be LeakyReLU (Leaky Rectified Linear Unit), the batch normalization layers can be BN (Batch Normalization), and the activation layers can be the Tanh function (hyperbolic tangent function) or the sigmoid function (S-shaped function).

[0107] Step S420: Based on the mask map of the dressing area and the intermediate image samples, obtain the predicted dressing image samples that convert the original costume corresponding to the portrait part into the target costume.

[0108] In one exemplary embodiment of this disclosure, after obtaining the costume change region mask and intermediate image samples through the above steps, a predicted costume change image sample that converts the original outfit corresponding to the portrait portion into the target outfit can be obtained based on the costume change region mask and intermediate image samples. Specifically, the area to be costumed can be determined through the costume change region mask, and the intermediate image samples can be further processed using the costume change region mask.

[0109] For example, the area to be dressed up can be determined by the dressing area mask map, and the transformation weight of the area to be dressed up can be increased to regenerate the predicted dressing image sample.

[0110] Furthermore, a predicted costume-changing image sample can be obtained by converting the original outfit corresponding to the portrait portion into the target outfit based on the costume-changing region mask, intermediate image samples, and image samples to be processed. Specifically, the region to be costumed can be determined using the costume-changing region mask, and the costume-changing image of the region to be costumed can be obtained from the intermediate image samples based on the costume-changing region mask. Additionally, a non-costume-changing image (not requiring costume-changing) can be obtained from the image samples to be processed based on the costume-changing region mask. The costume-changing image and the non-costume-changing image are combined to obtain the predicted costume-changing image sample that converts the original outfit corresponding to the portrait portion into the target outfit. Through this embodiment, for regions requiring costume-changing, the image at the corresponding position in the intermediate image sample can be used; for regions not requiring costume-changing, the image in the image samples to be processed can be used. The predicted costume-changing image sample determined in this way has a better costume-changing effect.

[0111] It should be noted that this disclosure does not impose any special limitations on the specific method of obtaining the predicted costume image sample that converts the original costume corresponding to the human figure into the target costume based on the costume area mask map and intermediate image samples.

[0112] Through the above steps S410-S420, the image sample to be processed and the key point bitmap can be input into the generator in the model to be trained to obtain the costume change region mask and intermediate image samples. Based on the costume change region mask and intermediate image samples, a predicted costume change image sample is obtained to convert the original costume corresponding to the portrait portion into the target costume. Through the embodiments of this disclosure, the area to be costumed can be determined by the costume change region mask, which can improve the costume change effect.

[0113] In one example embodiment of this disclosure, an original image sample can be obtained, along with a previous frame image sample and a key point bitmap corresponding to the portrait portion of the previous frame image sample. These are input into the generator of the model to be trained to obtain a previous frame clothing change region mask. A target portrait segmentation box is determined based on the previous frame clothing change region mask. The original image sample is then segmented using the target portrait segmentation box to obtain the image sample to be processed. (Refer to...) Figure 6 As shown, the target portrait segmentation box is determined based on the mask image of the clothing change area in the previous frame. The original image sample is then segmented using the target portrait segmentation box to obtain the image sample to be processed. This process may include the following steps S610 to S620:

[0114] Step S610: Obtain the original image sample, obtain the previous frame image sample and the key point bitmap corresponding to the portrait part of the previous frame image sample, and input the previous frame dressing area mask map obtained by the generator in the model to be trained.

[0115] In one example embodiment of this disclosure, an original image sample, a previous frame image sample, and a key point bitmap corresponding to the portrait portion of the previous frame image sample can be obtained and input into the generator of the model to be trained to obtain a mask map of the previous frame's clothing change region. Specifically, a portion of the original image sample is the image sample to be processed. When converting the original attire corresponding to the portrait portion in the image sample to be processed into the target attire, the image sample to be processed can be segmented from the original image sample using the target portrait segmentation box, thereby reducing the amount of data the model processes for the image sample to be processed, and thus improving the training and inference speed of the model.

[0116] Specifically, the image samples to be processed are multiple consecutive image samples. When determining the target portrait segmentation box of the current frame's image sample, the previous frame's clothing change region mask can be obtained. Specifically, after inputting the previous frame's image sample and the key point bitmap corresponding to the portrait portion of the previous frame's image sample into the model to be trained, the above embodiment's scheme can obtain the previous frame's clothing change region mask and the previous frame's intermediate image sample.

[0117] Step S620: Determine the target portrait segmentation box based on the mask image of the clothing change area in the previous frame, and perform image segmentation on the original image sample using the target portrait segmentation box to obtain the image sample to be processed.

[0118] In one example embodiment of this disclosure, after obtaining the mask map of the previous frame's clothing change area through the above steps, the area to be changed in the previous frame's image sample can be determined. Since the interval between frames is short, the area to be changed in the previous frame can be determined as the area to be changed in the current frame, and the target portrait segmentation box corresponding to the image sample to be processed in the current frame can be determined accordingly. The image sample to be processed is obtained by image segmentation of the original image sample through the target portrait segmentation box.

[0119] Furthermore, after obtaining the previous frame's costume change area mask, the candidate image segmentation box can be determined using the previous frame's costume change area mask, and the candidate image segmentation box can be augmented. The augmented result is then determined as the target image segmentation box.

[0120] For example, the candidate image segmentation box is a segmentation box of the same size as the area to be dressed up as indicated by the dressing area mask in the previous frame. Since the position of the human image part will change between different frames, the candidate image segmentation box can be enlarged. For example, the length and width of the candidate image segmentation box can be increased, and the result after the enlargement process can be determined as the target human image segmentation box.

[0121] For example, the target portrait segmentation bounding box can be a ROI (Region of Interest).

[0122] Through the above steps S610-S620, the original image sample can be obtained, the previous frame image sample and the key point bitmap corresponding to the portrait part of the previous frame image sample can be obtained and input into the generator of the model to be trained to obtain the previous frame clothing change region mask map, the target portrait segmentation box can be determined according to the previous frame clothing change region mask map, and the original image sample can be segmented by the target portrait segmentation box to obtain the image sample to be processed. Through the embodiments of this disclosure, the target portrait segmentation box of the current frame can be determined by the previous frame clothing change region mask map corresponding to the previous frame image sample, eliminating the need to determine the portrait segmentation box every time the model is trained, thus improving the training efficiency of the model to be trained.

[0123] In one example embodiment of this disclosure, edge features corresponding to the portrait portion can be obtained, and a predicted costume-changing image sample can be generated by converting the original attire corresponding to the portrait portion into the target attire based on the edge features. (Refer to...) Figure 7 As shown, the predicted costume change image sample, which converts the original costume corresponding to the portrait portion into the target costume based on edge features, may include the following steps S710 to S720:

[0124] Step S710: Obtain the edge features corresponding to the portrait portion;

[0125] Step S720: Based on edge features, convert the original attire corresponding to the portrait portion into a predicted costume image sample of the target attire.

[0126] In one exemplary embodiment of this disclosure, after inputting the image sample to be processed into the generator of the model to be trained through the above steps, the edge features of the portrait portion of the image sample can be obtained, and the original attire corresponding to the portrait portion can be converted into a predicted costume image sample of the target attire based on the edge features. Specifically, the edge portion of the portrait portion can be processed using the Sobel algorithm (edge ​​detection algorithm).

[0127] Specifically, the edge features of the human figure can be used to indicate the structural features of the edge of the human figure, the clothing features of the edge of the human figure, the color features of the edge of the human figure, etc.

[0128] It should be noted that this disclosure does not specifically limit the meaning of the edge features corresponding to the human portrait portion.

[0129] Furthermore, to obtain the edge features corresponding to the human image, a reparameterized structure can be used for the generator of the model to be trained. Specifically, the role of the reparameterized structure is to optimize model training without increasing the model's inference time. In actual model applications, the reparameterized structure is equivalent to a regular convolutional layer.

[0130] For example, such as Figure 8 As shown, this is a reparameterized structure for a generator, in which edge detection is added to obtain the edge features corresponding to the human portrait. 1*1 and 3*3 represent convolution kernels.

[0131] Through the above steps S710-S720, edge features corresponding to the portrait portion can be obtained, and a predicted costume image sample corresponding to the original costume portion can be converted into the target costume based on the edge features. Through the embodiments of this disclosure, edge features can be considered when performing costume conversion tasks to adapt to the importance of edge information in costume conversion tasks.

[0132] In one exemplary embodiment of this disclosure, after acquiring the image sample to be processed, portrait segmentation can be performed on the image sample to retain the portrait portion and original attire. Specifically, the portrait portion in the image sample to be processed can be detected using a portrait detection algorithm, and the image sample to be processed can be segmented to retain the portrait portion and original attire; alternatively, the image sample to be processed can be segmented, the portrait portion and original attire in the image sample to be processed can be determined, the portrait portion and original attire in the image sample to be processed can be retained, and the remaining portion can be filled with a solid color.

[0133] The embodiments of this disclosure can reduce the amount of image samples processed by the model, thereby improving the model's processing speed and thus improving the efficiency of the costume conversion task, meeting the real-time requirements of the costume conversion task.

[0134] In one example embodiment of this disclosure, an original image sample can be acquired, a previous frame image sample can be acquired, a target image segmentation box can be obtained by augmenting the portrait segmentation box, and the original image sample can be segmented using the target image segmentation box to obtain the image sample to be processed. (Refer to...) Figure 9 As shown, the process of segmenting the original image sample using the target human image segmentation bounding box to obtain the image sample to be processed may include the following steps S910 to S920:

[0135] Step S910: Obtain the original image sample and the previous frame image sample;

[0136] In one example embodiment of this disclosure, an original image sample and a previous frame image sample can be obtained. The previous frame image sample is an image obtained after image segmentation using a portrait segmentation bounding box. Specifically, a portion of the original image sample is the image sample to be processed. When converting the original attire corresponding to the portrait portion in the image sample to be processed into the target attire, the image sample to be processed can first be segmented from the original image sample using the target portrait segmentation bounding box. This reduces the amount of data the model processes for the image sample to be processed, thereby improving the training and inference speed of the model.

[0137] Specifically, the image samples to be processed are multiple consecutive image samples. When determining the target human figure segmentation box of the current frame image sample to be processed, the human figure segmentation box corresponding to the previous frame image sample can be obtained. Specifically, the previous frame image sample is obtained by segmenting the previous frame original image sample through the human figure segmentation box.

[0138] For example, the target portrait segmentation bounding box can be a ROI (Region of Interest).

[0139] Step S920: The portrait segmentation box is augmented to obtain the target portrait segmentation box. The original image sample is then segmented using the target portrait segmentation box to obtain the image sample to be processed.

[0140] In one exemplary embodiment of this disclosure, after obtaining the original image sample and the portrait segmentation box corresponding to the previous frame image sample through the above steps, the portrait segmentation box can be augmented to obtain the target portrait segmentation box, and the original image sample can be segmented using the target portrait segmentation box to obtain the image sample to be processed. Specifically, since the position of the portrait portion changes between different frames, the portrait segmentation box can be augmented. For example, the length and width of the candidate portrait segmentation box can be increased, and the result after augmentation is determined as the target portrait segmentation box. The original image sample can then be segmented using the target portrait segmentation box to obtain the image sample to be processed.

[0141] Through the above steps S910-S920, the original image sample and the previous frame image sample can be obtained. The portrait segmentation box is augmented to obtain the target portrait segmentation box. The original image sample is then segmented using the target portrait segmentation box to obtain the image sample to be processed. Through the embodiments of this disclosure, the portrait segmentation box corresponding to the previous frame image sample can be augmented to obtain the target portrait segmentation box for the current frame original image sample. This eliminates the need to determine the portrait segmentation box every time the model is trained, thus improving the training efficiency of the model.

[0142] In one example embodiment of this disclosure, the image samples to be processed can be augmented. For instance, before augmentation, the data set is image sample A to be processed - reference image sample A; after augmentation, the data set is image sample A to be processed - reference image sample A, image sample B to be processed - reference image sample A, and image sample C to be processed - reference image sample A.

[0143] Specifically, the stable diffusion model can be used for data augmentation.

[0144] In one example embodiment of this disclosure, an image can be acquired and input into a costume conversion model to obtain a costume image that converts the original costume corresponding to the portrait portion into the target costume. (Refer to...) Figure 10 As shown, inputting the image into the costume conversion model to obtain a costume image that converts the original costume corresponding to the portrait portion into the target costume may include the following steps S1010~S1020:

[0145] Step S1010: Acquire the image;

[0146] Step S1020: Input the image into the costume conversion model to obtain a costume image that converts the original costume corresponding to the portrait part into the target costume.

[0147] In one example embodiment of this disclosure, an image can be acquired. The image includes a human portrait portion and the original attire corresponding to the portrait portion. An attire conversion model corresponds to a target attire, and the attire conversion model is trained using the attire conversion model training method described in any of the above embodiments. Specifically, this attire conversion model, corresponding to the target attire, can convert the original attire of the human portrait portion in the image into the target attire to achieve a costume change effect.

[0148] It should be noted that this disclosure does not specifically limit the types of the original attire and the target attire.

[0149] Through the above steps S1010 to S1020, an image can be acquired and input into the costume conversion model to obtain a costume image that converts the original costume corresponding to the portrait part into the target costume.

[0150] In one embodiment of the present disclosure, a method for training a costume conversion model is provided. This method involves acquiring a sample image to be processed, obtaining a key point map corresponding to the portrait portion, inputting the sample image to be processed and the key point map into a generator in the model to be trained, obtaining a predicted costume conversion image sample that converts the original costume corresponding to the portrait portion into the target costume, inputting the predicted costume conversion image sample and a reference image sample into a discriminator in the model to be trained, obtaining a discrimination result by discriminating the predicted costume conversion image sample and the reference image sample, and updating the neural network parameters of the model to be trained based on the predicted costume conversion image sample, the reference image sample, and the discrimination result, so that the generator in the trained model can be used as the costume conversion model. The solution disclosed herein offers several advantages. First, it allows for the replacement of existing attire, resulting in a more natural look and adaptability to complex scenes, leading to better costume changes. Second, by considering key aspects of the human face during costume changes, the model learns about the human face structure, further enhancing the costume change effect. Third, the trained costume conversion model can convert the attire of a person in an image in real time, making it applicable to real-time scenarios and highly versatile.

[0151] It should be noted that the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0152] Furthermore, in an exemplary embodiment of this disclosure, an apparatus for training an outfit conversion model is also provided. (Refer to...) Figure 11 As shown, a costume conversion model training device 1100 includes: an image sample acquisition module 1110, a point map acquisition module 1120, a prediction image acquisition module 1130, a discrimination result acquisition module 1140, and a model update module 1150.

[0153] The system includes the following modules: an image sample acquisition module for acquiring image samples to be processed, which include a human portrait and the original attire corresponding to the portrait; a key point map acquisition module for acquiring key point maps corresponding to the human portrait, which indicate the key points of the portrait; a prediction image acquisition module for inputting the image samples to be processed and the key point maps into the generator of the model to be trained to obtain a predicted costume-changing image sample that converts the original attire corresponding to the human portrait into the target attire, where the model to be trained has a corresponding reference image sample, which includes a reference human portrait and the target attire corresponding to the reference human portrait; a discrimination result acquisition module for inputting the predicted costume-changing image sample and the reference image sample into the discriminator of the model to be trained, and obtaining a discrimination result by discriminating between the predicted costume-changing image sample and the reference image sample; and a model to be trained update module for updating the neural network parameters of the model to be trained based on the predicted costume-changing image sample, the reference image sample, and the discrimination result, so that the generator in the trained model to be trained can be used as the costume conversion model.

[0154] In one exemplary embodiment of this disclosure, based on the aforementioned scheme, the key point map includes multiple key points, the positions of which are related to the target attire.

[0155] In an exemplary embodiment of this disclosure, based on the aforementioned scheme, the target outfit includes a first target sub-outfit and a second target sub-outfit, which are not adjacent to each other; to obtain a key point map corresponding to the portrait portion, the apparatus further includes: a first key point acquisition unit, configured to acquire a first image region corresponding to the first target sub-outfit in the portrait portion, and acquire a plurality of first key points in the first image region; a second key point acquisition unit, configured to acquire a second image region corresponding to the second target sub-outfit in the portrait portion, and acquire a plurality of second key points in the second image region; a third key point acquisition unit, configured to acquire a plurality of first key points and a plurality of intermediate key points between the plurality of second key points; wherein the first key points and the second key points are respectively related to the intermediate key points; and a point map generation unit, configured to determine a key point map corresponding to the portrait portion through the plurality of first key points, the plurality of second key points, and the plurality of intermediate key points.

[0156] In one exemplary embodiment of this disclosure, based on the aforementioned scheme, the image sample to be processed and the key point map are input into the generator in the model to be trained to obtain a predicted costume-changing image sample that converts the original outfit corresponding to the portrait portion into the target outfit. The apparatus further includes: a costume-changing region mask acquisition unit, used to input the image sample to be processed and the key point map into the generator in the model to be trained to obtain a costume-changing region mask and an intermediate image sample; wherein, the costume-changing region mask is used to indicate the area for costume changing, and the intermediate image sample is used to indicate the temporary result of converting the original outfit corresponding to the portrait portion into the target outfit; and a costume-changing region mask application unit, used to obtain a predicted costume-changing image sample that converts the original outfit corresponding to the portrait portion into the target outfit based on the costume-changing region mask and the intermediate image sample.

[0157] In an exemplary embodiment of this disclosure, based on the aforementioned scheme, the image sample to be processed has a previous frame image sample. To obtain the image sample to be processed, the apparatus further includes: a previous frame costume change region mask acquisition unit, used to acquire the original image sample, acquire the previous frame image sample, and input the key point bitmap corresponding to the portrait portion of the previous frame image sample into the generator of the model to be trained to obtain the previous frame costume change region mask; and a first target portrait segmentation box determination unit, used to determine the target portrait segmentation box according to the previous frame costume change region mask, and perform image segmentation on the original image sample through the target portrait segmentation box to obtain the image sample to be processed.

[0158] In an exemplary embodiment of this disclosure, based on the aforementioned scheme, a predicted costume change image sample is obtained that converts the original costume corresponding to the portrait portion into the target costume. The apparatus further includes: an edge feature acquisition unit, used to acquire edge features corresponding to the portrait portion; and an edge feature application unit, used to convert the original costume corresponding to the portrait portion into the target costume based on the edge features.

[0159] In one exemplary embodiment of this disclosure, based on the foregoing scheme, the apparatus further includes: a portrait segmentation unit, used to perform portrait segmentation on the image sample to be processed, retaining the portrait portion and the original attire in the image sample to be processed.

[0160] In an exemplary embodiment of this disclosure, based on the aforementioned scheme, the image sample to be processed has a previous frame image sample. To obtain the image sample, the apparatus further includes: an original image sample acquisition unit, used to acquire the original image sample and the previous frame image sample; wherein, the previous frame image sample is an image obtained after image segmentation by a portrait segmentation box; and a second target portrait segmentation box acquisition unit, used to amplify the portrait segmentation box to obtain a target portrait segmentation box, and to perform image segmentation on the original image sample using the target portrait segmentation box to obtain the image sample to be processed.

[0161] An embodiment of this disclosure provides a costume conversion model training device that can acquire image samples to be processed, acquire key point maps corresponding to the human portrait portion, input the image samples to be processed and the key point maps into the generator in the model to be trained, obtain predicted costume conversion image samples that convert the original costume corresponding to the human portrait portion into the target costume, input the predicted costume conversion image samples and reference image samples into the discriminator in the model to be trained, and obtain a discrimination result by discriminating the predicted costume conversion image samples and reference image samples through the discriminator, and update the neural network parameters of the model to be trained based on the predicted costume conversion image samples, reference image samples and discrimination results, so as to use the generator in the trained model as the costume conversion model. The solution disclosed herein offers several advantages. First, it allows for the replacement of existing attire, resulting in a more natural look and adaptability to complex scenes, leading to better costume changes. Second, by considering key aspects of the human face during costume changes, the model learns about the human face structure, further enhancing the costume change effect. Third, the trained costume conversion model can convert the attire of a person in an image in real time, making it applicable to real-time scenarios and highly versatile.

[0162] Since the functional modules of the costume conversion model training device in the example embodiments of this disclosure correspond to the steps of the example embodiments of the costume conversion model training method described above, for details not disclosed in the device embodiments of this disclosure, please refer to the embodiments of the costume conversion model training method described above.

[0163] Furthermore, in an exemplary embodiment of this disclosure, an outfit conversion device is also provided. (Refer to...) Figure 12 As shown, an outfit conversion device 1200 includes: a human image acquisition module 1210 and a human image conversion module 1220.

[0164] The image acquisition module is used to acquire images, including a human figure and the original attire corresponding to the human figure. The image conversion module is used to input the image into the attire conversion model to obtain a costume image that converts the original attire corresponding to the human figure into the target attire. The attire conversion model corresponds to the target attire and is trained using any of the above-mentioned attire conversion model training methods.

[0165] Since the functional modules of the costume conversion device in the example embodiments of this disclosure correspond to the steps of the example embodiments of the costume conversion method described above, for details not disclosed in the device embodiments of this disclosure, please refer to the embodiments of the costume conversion method described above.

[0166] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0167] Furthermore, in an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described outfit conversion model training method and the above-described outfit conversion method is also provided.

[0168] Those skilled in the art will understand that various aspects of this disclosure can be implemented as a system, method, or program product. Therefore, various aspects of this disclosure can be embodied in the following forms: a completely hardware embodiment, a completely software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."

[0169] The following reference Figure 13 To describe an electronic device 1300 according to such an embodiment of the present disclosure. Figure 13 The electronic device 1300 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.

[0170] like Figure 13 As shown, the electronic device 1300 is manifested in the form of a general-purpose computing device. The components of the electronic device 1300 may include, but are not limited to: at least one processing unit 1310, at least one storage unit 1320, a bus 1330 connecting different system components (including storage unit 1320 and processing unit 1310), and a display unit 1340.

[0171] The storage unit stores program code, which can be executed by the processing unit 1310 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure. For example, the processing unit 1310 can perform actions such as... Figure 2The steps shown are as follows: Step S210: Obtain an image sample to be processed; wherein the image sample to be processed includes a human figure and the original costume corresponding to the human figure; Step S220: Obtain a key point map corresponding to the human figure; wherein the key point map is used to indicate the key points of the human figure; Step S230: Input the image sample to be processed and the key point map into the generator in the model to be trained to obtain a predicted costume-changing image sample that converts the original costume corresponding to the human figure into the target costume; wherein the model to be trained has a reference image sample, which includes a reference human figure and the target costume corresponding to the reference human figure; Step S240: Input the predicted costume-changing image sample and the reference image sample into the discriminator in the model to be trained, and use the discriminator to discriminate the predicted costume-changing image sample and the reference image sample to obtain a discrimination result; Step S250: Update the neural network parameters of the model to be trained according to the predicted costume-changing image sample, the reference image sample and the discrimination result, so as to use the generator in the trained model as the costume conversion model.

[0172] In one exemplary embodiment of this disclosure, based on the aforementioned scheme, the key point map includes multiple key points, the positions of which are related to the target attire.

[0173] In an exemplary embodiment of this disclosure, based on the aforementioned scheme, the target outfit includes a first target sub-outfit and a second target sub-outfit, wherein the first target sub-outfit and the second target sub-outfit are not adjacent to each other; obtaining the key point map corresponding to the portrait portion includes: in the portrait portion, obtaining a first image region corresponding to the first target sub-outfit, and obtaining a plurality of first key points in the first image region; in the portrait portion, obtaining a second image region corresponding to the second target sub-outfit, and obtaining a plurality of second key points in the second image region; obtaining a plurality of first key points and a plurality of intermediate key points between the plurality of second key points; wherein the first key points and the second key points are respectively related to the intermediate key points; and determining the key point map corresponding to the portrait portion through the plurality of first key points, the plurality of second key points, and the plurality of intermediate key points.

[0174] In one exemplary embodiment of this disclosure, based on the aforementioned scheme, the image sample to be processed and the key point map are input into the generator in the model to be trained to obtain a predicted costume-changing image sample that converts the original outfit corresponding to the portrait portion into the target outfit. This includes: inputting the image sample to be processed and the key point map into the generator in the model to be trained to obtain a costume-changing region mask and an intermediate image sample; wherein, the costume-changing region mask is used to indicate the region where the costume-changing is performed, and the intermediate image sample is used to indicate the temporary result of converting the original outfit corresponding to the portrait portion into the target outfit; and obtaining the predicted costume-changing image sample that converts the original outfit corresponding to the portrait portion into the target outfit based on the costume-changing region mask and the intermediate image sample.

[0175] In an exemplary embodiment of this disclosure, based on the aforementioned scheme, the image sample to be processed has a previous frame image sample. Obtaining the image sample to be processed includes: obtaining an original image sample, obtaining the previous frame image sample and the key point bitmap corresponding to the portrait part of the previous frame image sample, inputting the previous frame clothing change region mask map obtained by the generator in the model to be trained; determining the target portrait segmentation box according to the previous frame clothing change region mask map, and performing image segmentation on the original image sample through the target portrait segmentation box to obtain the image sample to be processed.

[0176] In one exemplary embodiment of this disclosure, based on the aforementioned scheme, obtaining a predicted costume change image sample that converts the original costume corresponding to the portrait portion into the target costume includes: acquiring edge features corresponding to the portrait portion; and converting the original costume corresponding to the portrait portion into the target costume based on the edge features.

[0177] In one exemplary embodiment of this disclosure, based on the aforementioned scheme, human portrait segmentation is performed on the image sample to be processed, preserving the human portrait portion and original attire in the image sample to be processed.

[0178] In an exemplary embodiment of this disclosure, based on the aforementioned scheme, the image sample to be processed has a previous frame image sample. Obtaining the image sample includes: obtaining an original image sample and obtaining a previous frame image sample; wherein, the previous frame image sample is an image obtained after image segmentation by a portrait segmentation box; the portrait segmentation box is augmented to obtain a target portrait segmentation box, and the original image sample is segmented using the target portrait segmentation box to obtain the image sample to be processed.

[0179] An embodiment of this disclosure provides an electronic device that can acquire an image sample to be processed, acquire a key point map corresponding to the human portrait portion, input the image sample to be processed and the key point map into a generator in a model to be trained, obtain a predicted costume-changing image sample that converts the original costume corresponding to the human portrait portion into a target costume, input the predicted costume-changing image sample and a reference image sample into a discriminator in the model to be trained, and obtain a discrimination result by discriminating the predicted costume-changing image sample and the reference image sample through the discriminator. The neural network parameters of the model to be trained are updated according to the predicted costume-changing image sample, the reference image sample and the discrimination result, so as to use the generator in the trained model as a costume conversion model. The solution disclosed herein offers several advantages. First, it allows for the replacement of existing attire, resulting in a more natural look and adaptability to complex scenes, leading to better costume changes. Second, by considering key aspects of the human face during costume changes, the model learns about the human face structure, further enhancing the costume change effect. Third, the trained costume conversion model can convert the attire of a person in an image in real time, making it applicable to real-time scenarios and highly versatile.

[0180] Storage unit 1320 may include readable media in the form of volatile storage units, such as random access memory (RAM) 1321 and / or cache memory 1322, and may further include read-only memory (ROM) 1323.

[0181] Storage unit 1320 may also include a program / utility 1324 having a set (at least one) program module 1325, such program module 1325 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.

[0182] Bus 1330 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0183] Electronic device 1300 can also communicate with one or more external devices 1370 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 1300, and / or with any device that enables electronic device 1300 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 1350. Furthermore, electronic device 1300 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 1360. As shown, network adapter 1360 communicates with other modules of electronic device 1300 via bus 1330. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 1300, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0184] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0185] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible embodiments, various aspects of this disclosure may also be implemented as a program product including program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure.

[0186] In one example embodiment of this disclosure, a computer-readable signal medium is provided, which can acquire an image sample to be processed; wherein the image sample to be processed includes a human portrait portion and the original attire corresponding to the human portrait portion; a key point map corresponding to the human portrait portion is acquired; wherein the key point map is used to indicate the key points of the human portrait; the image sample to be processed and the key point map are input into a generator in a model to be trained to obtain a predicted costume-changing image sample that converts the original attire corresponding to the human portrait portion into a target attire; wherein the model to be trained has a reference image sample, which includes a reference human portrait portion and the target attire corresponding to the reference human portrait portion; the predicted costume-changing image sample and the reference image sample are input into a discriminator in the model to be trained, and the discriminator distinguishes the predicted costume-changing image sample and the reference image sample to obtain a discrimination result; the neural network parameters of the model to be trained are updated according to the predicted costume-changing image sample, the reference image sample, and the discrimination result, so that the generator in the trained model to be trained is used as the costume conversion model.

[0187] In one exemplary embodiment of this disclosure, based on the aforementioned scheme, the key point map includes multiple key points, the positions of which are related to the target attire.

[0188] In an exemplary embodiment of this disclosure, based on the aforementioned scheme, the target outfit includes a first target sub-outfit and a second target sub-outfit, wherein the first target sub-outfit and the second target sub-outfit are not adjacent to each other; obtaining the key point map corresponding to the portrait portion includes: in the portrait portion, obtaining a first image region corresponding to the first target sub-outfit, and obtaining a plurality of first key points in the first image region; in the portrait portion, obtaining a second image region corresponding to the second target sub-outfit, and obtaining a plurality of second key points in the second image region; obtaining a plurality of first key points and a plurality of intermediate key points between the plurality of second key points; wherein the first key points and the second key points are respectively related to the intermediate key points; and determining the key point map corresponding to the portrait portion through the plurality of first key points, the plurality of second key points, and the plurality of intermediate key points.

[0189] In one exemplary embodiment of this disclosure, based on the aforementioned scheme, the image sample to be processed and the key point map are input into the generator in the model to be trained to obtain a predicted costume-changing image sample that converts the original outfit corresponding to the portrait portion into the target outfit. This includes: inputting the image sample to be processed and the key point map into the generator in the model to be trained to obtain a costume-changing region mask and an intermediate image sample; wherein, the costume-changing region mask is used to indicate the region where the costume-changing is performed, and the intermediate image sample is used to indicate the temporary result of converting the original outfit corresponding to the portrait portion into the target outfit; and obtaining the predicted costume-changing image sample that converts the original outfit corresponding to the portrait portion into the target outfit based on the costume-changing region mask and the intermediate image sample.

[0190] In an exemplary embodiment of this disclosure, based on the aforementioned scheme, the image sample to be processed has a previous frame image sample. Obtaining the image sample to be processed includes: obtaining an original image sample, obtaining the previous frame image sample and the key point bitmap corresponding to the portrait part of the previous frame image sample, inputting the previous frame clothing change region mask map obtained by the generator in the model to be trained; determining the target portrait segmentation box according to the previous frame clothing change region mask map, and performing image segmentation on the original image sample through the target portrait segmentation box to obtain the image sample to be processed.

[0191] In one exemplary embodiment of this disclosure, based on the aforementioned scheme, obtaining a predicted costume change image sample that converts the original costume corresponding to the portrait portion into the target costume includes: acquiring edge features corresponding to the portrait portion; and converting the original costume corresponding to the portrait portion into the target costume based on the edge features.

[0192] In one exemplary embodiment of this disclosure, based on the aforementioned scheme, human portrait segmentation is performed on the image sample to be processed, preserving the human portrait portion and original attire in the image sample to be processed.

[0193] In an exemplary embodiment of this disclosure, based on the aforementioned scheme, the image sample to be processed has a previous frame image sample. Obtaining the image sample includes: obtaining an original image sample and obtaining a previous frame image sample; wherein, the previous frame image sample is an image obtained after image segmentation by a portrait segmentation box; the portrait segmentation box is augmented to obtain a target portrait segmentation box, and the original image sample is segmented using the target portrait segmentation box to obtain the image sample to be processed.

[0194] One embodiment of this disclosure provides a computer-readable signal medium that can acquire image samples to be processed, acquire key point maps corresponding to the human portrait portion, input the image samples to be processed and the key point maps into a generator in a model to be trained, obtain a predicted costume-changing image sample that converts the original costume corresponding to the human portrait portion into the target costume, input the predicted costume-changing image sample and a reference image sample into a discriminator in the model to be trained, and obtain a discrimination result by discriminating the predicted costume-changing image sample and the reference image sample through the discriminator. The neural network parameters of the model to be trained are updated according to the predicted costume-changing image sample, the reference image sample and the discrimination result, so that the generator in the trained model to be trained can be used as the costume conversion model. The solution disclosed herein offers several advantages. First, it allows for the replacement of existing attire, resulting in a more natural look and adaptability to complex scenes, leading to better costume changes. Second, by considering key aspects of the human face during costume changes, the model learns about the human face structure, further enhancing the costume change effect. Third, the trained costume conversion model can convert the attire of a person in an image in real time, making it applicable to real-time scenarios and highly versatile.

[0195] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0196] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0197] Program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing devices can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0198] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0199] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

Claims

1. A method for training an outfit conversion model, characterized in that, The method includes: Obtain an image sample to be processed; wherein the image sample to be processed includes a human figure portion and the original attire corresponding to the human figure portion; A key point map corresponding to the portrait portion is obtained based on the target costume corresponding to the model to be trained; wherein, the key point map is used to indicate the key points of the portrait; the target costume includes a first target sub-costume and a second target sub-costume, and the first target sub-costume and the second target sub-costume are not adjacent to each other; obtaining the key point map corresponding to the portrait portion includes: in the portrait portion, obtaining a first image region corresponding to the first target sub-costume, and obtaining a plurality of first key points in the first image region; in the portrait portion, obtaining a second image region corresponding to the second target sub-costume, and obtaining a plurality of second key points in the second image region; obtaining the plurality of first key points and a plurality of intermediate key points between the plurality of second key points; wherein, the first key points and the second key points are respectively related to the intermediate key points; and determining the key point map corresponding to the portrait portion through the plurality of first key points, the plurality of second key points, and the plurality of intermediate key points. The image sample to be processed and the key point map are input into the generator in the model to be trained to obtain a predicted costume change image sample that converts the original costume corresponding to the portrait part into the target costume; wherein, the model to be trained corresponds to a reference image sample, which includes a reference portrait part and the target costume corresponding to the reference portrait part. The predicted costume change image sample and the reference image sample are input into the discriminator in the model to be trained, and the discriminator judges the predicted costume change image sample and the reference image sample to obtain the discrimination result; The neural network parameters of the model to be trained are updated based on the predicted costume change image samples, the reference image samples, and the discrimination results, so that the generator in the trained model to be trained can be used as the costume conversion model.

2. The method according to claim 1, characterized in that, The key point map includes multiple key points, the positions of which are related to the target outfit.

3. The method according to claim 1, characterized in that, The step of inputting the image sample to be processed and the key point map into the generator in the model to be trained to obtain a predicted costume-changing image sample that converts the original costume corresponding to the portrait portion into the target costume includes: The image sample to be processed and the key point map are input into the generator in the model to be trained to obtain the costume change area mask map and intermediate image sample; wherein, the costume change area mask map is used to indicate the area to be costumed, and the intermediate image sample is used to indicate the temporary result of converting the original costume corresponding to the portrait part into the target costume; Based on the mask map of the changing area and the intermediate image samples, a predicted changing image sample is obtained to convert the original outfit corresponding to the portrait part into the target outfit.

4. The method according to claim 3, characterized in that, The image sample to be processed has a previous frame image sample, and the process of obtaining the image sample to be processed includes: Obtain the original image sample, obtain the previous frame image sample and the key point bitmap corresponding to the portrait part of the previous frame image sample, and input the previous frame dressing area mask map into the generator in the model to be trained. The target portrait segmentation box is determined based on the previous frame's clothing change area mask, and the original image sample is segmented using the target portrait segmentation box to obtain the image sample to be processed.

5. The method according to claim 1, characterized in that, The process of obtaining a predicted costume-changing image sample that converts the original attire corresponding to the portrait portion into the target attire includes: Obtain the edge features corresponding to the portrait portion; Based on the edge features, a predicted costume change image sample is generated by converting the original costume corresponding to the portrait portion into the target costume.

6. The method according to claim 1, characterized in that, The method further includes: The image sample to be processed is segmented into human figures, and the human figure portion and the original attire in the image sample to be processed are preserved.

7. The method according to claim 1, characterized in that, The image sample to be processed has a previous frame image sample, and the process of obtaining the image sample to be processed includes: Obtain the original image sample and the previous frame image sample; wherein, the previous frame image sample is the image obtained after image segmentation using a portrait segmentation bounding box; The target image segmentation box is obtained by amplifying the image segmentation box, and the original image sample is segmented using the target image segmentation box to obtain the image sample to be processed.

8. A method for changing attire, characterized in that, The method includes: Acquire an image; wherein the image includes a human figure and the original attire corresponding to the human figure; The image is input into the costume conversion model to obtain a costume image that converts the original costume corresponding to the portrait into the target costume. The costume conversion model corresponds to a target costume, and the costume conversion model is trained by the costume conversion model training method as described in any one of claims 1-7.

9. A training device for a costume conversion model, characterized in that, The device includes: An image sample acquisition module is used to acquire an image sample to be processed; wherein, the image sample to be processed includes a human figure portion and the original attire corresponding to the human figure portion; A key point map acquisition module is used to acquire a key point map corresponding to the portrait portion based on the target attire corresponding to the model to be trained; wherein, the key point map is used to indicate the key points of the portrait; the target attire includes a first target sub-attire and a second target sub-attire, and the first target sub-attire and the second target sub-attire are not adjacent; acquiring the key point map corresponding to the portrait portion includes: in the portrait portion, acquiring a first image region corresponding to the first target sub-attire, and acquiring a plurality of first key points in the first image region; in the portrait portion, acquiring a second image region corresponding to the second target sub-attire, and acquiring a plurality of second key points in the second image region; acquiring the plurality of first key points and a plurality of intermediate key points between the plurality of second key points; wherein, the first key points and the second key points are respectively related to the intermediate key points; and determining the key point map corresponding to the portrait portion through the plurality of first key points, the plurality of second key points, and the plurality of intermediate key points. The prediction image acquisition module is used to input the image sample to be processed and the key point map into the generator in the model to be trained to obtain a prediction dressing image sample that converts the original outfit corresponding to the portrait part into the target outfit; wherein, the model to be trained corresponds to a reference image sample, the reference image sample includes a reference portrait part and the target outfit corresponding to the reference portrait part; The discrimination result acquisition module is used to input the predicted costume change image sample and the reference image sample into the discriminator in the model to be trained, and to obtain the discrimination result by the discriminator discriminating the predicted costume change image sample and the reference image sample; The training model update module is used to update the neural network parameters of the training model based on the predicted costume change image samples, the reference image samples, and the discrimination results, so as to use the generator in the trained training model as the costume conversion model.

10. A costume conversion device, characterized in that, The device includes: An image acquisition module is used to acquire an image; wherein the image includes a human portrait portion and the original attire corresponding to the human portrait portion; An image conversion module is used to input the image into an outfit conversion model to obtain a costume image that converts the original outfit corresponding to the portrait portion into the target outfit; wherein, the outfit conversion model corresponds to the target outfit, and the outfit conversion model is trained by the outfit conversion model training method as described in any one of claims 1-7.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-8.

12. An electronic device, characterized in that, include: One or more processors; as well as A memory for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1-8.