Multi-angle clothes changing method and system, electronic equipment and storage medium

By generating an attention weight matrix and a target latent space map, combined with multi-view analysis and a cacheable diffusion model, the problem of high requirements for model and clothing poses is solved, achieving efficient, multi-angle, and high-quality dress-up effects, and supporting real-time dress-up from multiple outfit images.

CN121505085APending Publication Date: 2026-02-10HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511806082.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing technologies have high requirements for model and clothing poses, resulting in cumbersome and inefficient image acquisition, making it impossible to generate high-quality dressing effects in a timely manner. Furthermore, they lack the ability to coordinate the replacement of supporting elements. Existing diffusion models have a large computational load and cannot achieve timely dressing changes.

Method used

By acquiring model and outfit images, an attention weight matrix and target latent space map are generated. Multi-view analysis and noise addition are performed. Combined with a cacheable multi-reference diffusion model, the computational load is reduced, achieving consistency between the angles of the model and outfit images and high-quality outfit changes.

Benefits of technology

It enables efficient generation of high-quality multi-angle costume images without requiring high pose, reducing computational load, improving costume efficiency and effect, and supporting real-time costume changing from multiple costume images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121505085A_ABST
    Figure CN121505085A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-angle dress changing method and system, electronic equipment and a storage medium, and the method comprises the steps: obtaining a model drawing and at least one dress-up drawing, and determining the type of each dress-up drawing; generating an attention weight matrix corresponding to each dress-up picture based on the type of each dress-up picture and the model picture; performing multi-view analysis on each dress-up picture to generate a target potential space picture of each dress-up picture consistent with the angle of the model picture; performing primary noise addition on the set of the target potential space diagrams of the dress-up diagrams to obtain final dress-up space features; performing multiple times of noise addition on a weighted potential space diagram of the model diagram in the overall space features to obtain final model space features; wherein the overall spatial feature comprises a weighted potential spatial map of the model chart obtained by combining an attention weight matrix corresponding to each dress-up chart with a potential spatial map of the model chart, and a final dress-up spatial feature; and redrawing based on the final spatial features of the model to obtain a reloading model drawing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a multi-angle dressing method and system, electronic device, and storage medium. Background Technology

[0002] With the rapid development of generative artificial intelligence technology, its application scenarios in image generation and editing are constantly expanding. Therefore, current technologies can be used to generate images of a specified clothing model wearing a specified outfit from a given clothing image, thus providing a more cost-effective and efficient way to obtain images of a specified model wearing a specified outfit.

[0003] Specifically, the process involves capturing images of the model in a standardized pose, standing facing forward with their clothing fully visible, to obtain a model image. Similarly, the clothing image also needs to be a standard frontal view. Then, a diffusion model is used to iteratively process all features of both the model and clothing images, adding noise and other techniques to complete the clothing change, resulting in a model image after the change.

[0004] However, existing methods have high requirements for the model's and clothing's poses, making the image acquisition process overly cumbersome and inefficient. Furthermore, discrepancies in angles between the model and clothing can easily lead to poor generation results. Moreover, these methods primarily replace the main clothing, lacking the ability to coordinate the replacement of accessories such as headwear, shoes, and socks, thus failing to meet diverse current needs. Finally, existing diffusion models require significant time to process all features during inference, preventing timely clothing changes. Summary of the Invention

[0005] In view of the shortcomings of the prior art, this application provides a multi-angle dress-up method and system, electronic device and storage medium to solve the problem that the prior art cannot guarantee high-quality dress-up for multiple outfits in a timely manner.

[0006] To achieve the above objectives, this application provides the following technical solution:

[0007] The first aspect of this application provides a multi-angle costume changing method, including:

[0008] Obtain a model image and at least one outfit image, and determine the type of each outfit image;

[0009] Based on the type of each outfit image and the model image, generate an attention weight matrix corresponding to each outfit image;

[0010] By performing multi-view analysis on each of the aforementioned outfit images, a target latent space map of each outfit image with the same angle as the model image is generated;

[0011] A noise addition is performed on the set of target latent space maps of each of the aforementioned costume images to obtain the final spatial features of the costume.

[0012] Multiple noise additions are made to the weighted latent space map of the model image in the overall spatial features to obtain the final spatial features of the model; wherein, the overall spatial features are obtained by concatenating the weighted latent space map of the model image and the final spatial features of the outfit; the weighted latent space map of the model image is obtained by combining the attention weight matrix corresponding to each outfit image with the latent space map of the model image;

[0013] Based on the final spatial features of the model, a redrawing is performed to obtain a model image in different outfits.

[0014] Optionally, in the above-described multi-angle dress-up method, generating an attention weight matrix corresponding to each of the dress-up images based on the type of each dress-up image and the model image includes:

[0015] The type of each outfit image and the model image are encoded to obtain the text features of the type of each outfit image and the latent space map of the model image.

[0016] The text features of each type of costume image and the latent space map of the model image are analyzed using a cross-attention mechanism to obtain the attention weight matrix corresponding to each costume image.

[0017] Optionally, in the above-described multi-angle costume change method, the step of generating target latent space maps of each costume image with the same angle as the model image by performing multi-view analysis on each of the costume images includes:

[0018] For each of the costume images, a multi-branch encoder is used to process the costume image to obtain the mean and logarithmic variance of the latent space map of the costume image in each target direction;

[0019] The mean and logarithmic variance of the latent space map in the target direction corresponding to the angle range of the model image are selected; wherein, the angle of the model image is obtained by processing the model image through a residual network and then performing angle transformation on the processing result;

[0020] The target latent space map of the costume image is generated using the mean and logarithmic variance of the selected latent space map.

[0021] Optionally, in the above-described multi-angle costume change method, the step of processing the costume image using a multi-branch encoder to obtain the mean and logarithmic variance of the latent space map of the costume image in each target direction includes:

[0022] The costume image is feature extracted by five convolutional layers and one residual layer in a multi-branch encoder to obtain the latent space map of the costume image in each target direction.

[0023] The latent space graphs of the costume image in each target direction are processed by the fully connected layer in the multi-branch encoder to obtain the mean and logarithmic variance of the latent space graphs of the costume image in each target direction.

[0024] Optionally, in the above-described multi-angle costume change method, generating the target latent space map of the costume image using the mean and logarithmic variance of the selected latent space map includes:

[0025] The standard deviation of the latent space map is calculated using the log-variance of the selected latent space map;

[0026] Gaussian distribution sampling is performed based on the standard deviation and mean of the latent space map to obtain latent variables;

[0027] The latent variables are mapped to high-dimensional vectors to obtain latent high-dimensional vectors;

[0028] The potential high-dimensional vector is adjusted into a three-dimensional feature map to obtain the target potential space map of the costume image.

[0029] Optionally, in the above-described multi-angle dress-up method, the step of adding noise to the set of target latent space maps of each of the dress-up images to obtain the final spatial features of the dress-up includes:

[0030] The set of target latent space maps of each of the aforementioned costume images is noise-added using a cacheable multi-reference diffusion model to obtain the final spatial features of the costume.

[0031] The training method for the cacheable multi-reference diffusion model includes:

[0032] The acquired training data is processed using the cacheable multi-reference diffusion model to obtain the current costume change image;

[0033] Based on the deviation between the current outfit image and the outfit images in the training data, a current first loss and various current second losses are calculated; wherein, the current first loss is the loss when all the outfit sample images in the training data are used; and the current second loss is the loss when only a single outfit sample image in the training data is used.

[0034] The parameters of the cacheable multi-reference diffusion model are updated based on the current second loss and the current first loss, and the process of processing the acquired training data through the cacheable multi-reference diffusion model is returned to obtain the current costume image until the total loss converges.

[0035] The second aspect of this application provides a multi-angle dressing system, including:

[0036] An acquisition unit is used to acquire a model image and at least one outfit image, and to determine the type of each outfit image;

[0037] The weight determination unit is used to generate an attention weight matrix corresponding to each of the costume images based on the type of each costume image and the model image, respectively.

[0038] The feature generation unit is used to generate target latent space maps of each of the costume images by performing multi-view analysis on each of the costume images, with the same angle as the model image.

[0039] The decoration noise addition unit is used to add noise to the set of target latent space maps of each decoration image to obtain the final decoration space features;

[0040] The cyclic addition unit adds noise multiple times to the weighted latent space map of the model image in the overall spatial features to obtain the final spatial features of the model; wherein, the overall spatial features are obtained by concatenating the weighted latent space map of the model image and the final spatial features of the outfit; the weighted latent space map of the model image is obtained by combining the attention weight matrix corresponding to each outfit image with the latent space map of the model image;

[0041] An image generation unit is used to redraw the model based on the model's final spatial features to obtain a model image in different outfits.

[0042] Optionally, in the above-described multi-angle costume changing system, the weight determination unit includes:

[0043] An encoding unit is used to encode the type of each of the costume images and the model image respectively, to obtain the text features of the type of each costume image and the latent space map of the model image;

[0044] The weight matrix generation unit is used to analyze the combination of the text features of the type of each costume image and the latent space map of the model image through a cross-attention mechanism to obtain the attention weight matrix corresponding to each costume image.

[0045] Optionally, in the above-described multi-angle costume changing system, the feature generation unit includes:

[0046] A multi-branch processing unit is used to process each of the costume images separately using a multi-branch encoder to obtain the mean and logarithmic variance of the latent space map of the costume image in each target direction.

[0047] A selection unit is used to select the mean and logarithmic variance of the latent space map in the target direction corresponding to the angle range of the model image; wherein the angle of the model image is obtained by processing the model image through a residual network and then performing angle transformation on the processing result;

[0048] The feature map generation unit is used to generate a target latent space map of the costume image using the mean and logarithmic variance of the selected latent space map.

[0049] Optionally, in the above-described multi-angle changing system, the multi-branch processing unit includes:

[0050] The feature extraction unit is used to extract features from the costume image through five convolutional layers and one residual layer in the multi-branch encoder to obtain the latent space map of the costume image in each target direction.

[0051] The parameter calculation unit is used to process the latent space map of the costume image in each target direction through the fully connected layer in the multi-branch encoder to obtain the mean and logarithmic variance of the latent space map of the costume image in each target direction.

[0052] Optionally, in the above-described multi-angle changing system, the selection unit includes:

[0053] The standard deviation calculation unit is used to calculate the standard deviation of the latent space map using the logarithmic variance of the selected latent space map;

[0054] A sampling unit is used to perform Gaussian distribution sampling based on the standard deviation and mean of the latent space map to obtain latent variables;

[0055] A mapping unit is used to map the latent variables into high-dimensional vectors to obtain latent high-dimensional vectors;

[0056] The dimension adjustment unit is used to adjust the potential high-dimensional vector into a three-dimensional feature map to obtain the target potential space map of the costume image.

[0057] Optionally, the multi-angle changing system described above also includes:

[0058] The training unit is used to process the acquired training data through the cacheable multi-reference diffusion model to obtain the current costume image;

[0059] The loss calculation unit is used to calculate a current first loss and various current second losses based on the deviation between the current costume image and the costume images in the training data; wherein, the current first loss is the loss when all the costume sample images in the training data are used; and the current second loss is the loss when only a single costume sample image in the training data is used.

[0060] The parameter update unit is used to update the parameters of the cacheable multi-reference diffusion model based on each of the current second loss and the current first loss, and return to the training unit until the total loss converges.

[0061] A third aspect of this application provides an electronic device, comprising:

[0062] Memory and processor;

[0063] The memory is used to store programs;

[0064] The processor is used to execute the program, which, when executed, is specifically used to implement the multi-angle dressing method as described in any of the above.

[0065] The fourth aspect of this application provides a computer storage medium for storing a computer program, which, when executed by a processor, is used to implement the multi-angle dressing method as described in any of the preceding claims.

[0066] This application provides a multi-angle dress-up method, which acquires a model image and at least one outfit image, and determines the type of each outfit image, thus supporting the uploading of multiple types of outfit images for simultaneous dress-up. Based on the type of each outfit image and the model image, an attention weight matrix is ​​generated for each outfit image. This matrix is ​​used during subsequent redrawing to ensure different regions of the model image have different weights, better delineating the redrawing area, reducing the abruptness of the dress-up edges, coordinating the dress-up process across different outfit images, and ensuring the dress-up effect of each image. By performing multi-view analysis on each outfit image, target latent space maps of each outfit image with the same angle as the model image are generated. This multi-view analysis ensures that the angles of the model image and the outfit image are consistent, eliminating the need for poses with high modal requirements. Then, noise is added to the set of target latent space maps of each outfit image to obtain the final spatial features of the outfit. Multiple noise additions are then applied to the weighted latent space map of the model image within the overall spatial features to obtain the final spatial features of the model. The overall spatial features are obtained by concatenating the weighted latent space map of the model images with the final spatial features of the outfit. The weighted latent space map of the model images is obtained by combining the attention weight matrix corresponding to each outfit image with the latent space map of the model images. Finally, the model images are redrawn based on the final spatial features of the model to obtain the outfit-changing model images. This freezes the features of multiple outfit images after adding noise once. Subsequently, noise is only added to the features of the weighted model images in a cyclical manner, reducing the computational load during cyclic noise addition, accelerating the inference process, and realizing real-time outfit changing. Therefore, a method for high-quality and timely outfit changing of multiple outfit images simultaneously is realized. Attached Figure Description

[0067] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0068] Figure 1 A flowchart illustrating a multi-angle garment changing method provided in this application embodiment;

[0069] Figure 2 A flowchart illustrating a method for generating an attention weight matrix corresponding to a costume image, provided in an embodiment of this application;

[0070] Figure 3 A flowchart illustrating a method for multi-view analysis of costume images provided in this application embodiment;

[0071] Figure 4 A schematic diagram of the architecture of a multi-view module provided in an embodiment of this application;

[0072] Figure 5 A flowchart illustrating a method for processing a costume image using a multi-branch encoder, provided in an embodiment of this application;

[0073] Figure 6 A schematic diagram of the architecture of a multi-branch encoder provided for an embodiment of this application;

[0074] Figure 7 A flowchart illustrating a method for generating a target potential space map of a costume image, as provided in an embodiment of this application;

[0075] Figure 8 This is a schematic diagram illustrating the logic of adding noise to feature data, as provided in an embodiment of this application.

[0076] Figure 9 A flowchart illustrating a training method for a cacheable multi-reference diffusion model provided in this application embodiment;

[0077] Figure 10 This application provides a schematic diagram of the architecture of a multi-angle dressing system.

[0078] Figure 11 This is a schematic diagram of the architecture of an electronic device provided in an embodiment of this application. Detailed Implementation

[0079] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0080] In this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0081] This application provides a multi-angle clothing changing method, such as... Figure 1 As shown, it includes the following steps:

[0082] S101. Obtain a model image and at least one outfit image, and determine the type of each outfit image.

[0083] It should be noted that, since multi-view analysis is subsequently introduced in this application embodiment, the acquired model image and multiple outfit images do not need to conform to high-standard poses. The acquired multiple outfit images can include various types of outfit images, such as images of tops, bottoms, shoes, etc. Furthermore, this application embodiment can coordinate the changing of various types of outfits, so multiple outfit images of different types can be acquired simultaneously for outfit changing. Similarly, there can be one or multiple model images. When there are multiple model images, the method provided in this application embodiment can be executed on each model image separately, thereby applying the outfit to each model image.

[0084] Since different types of outfit images require changing clothes to different areas of the model image, in order to achieve accurate clothing changes for each type of outfit image and reduce the abruptness of the clothing change edges, thus obtaining high-quality outfit change images, it is necessary to determine the type of each outfit image so that the clothing change area can be processed accordingly based on the type of outfit image.

[0085] Alternatively, each costume image can be input into the reverse model, which outputs the type of each costume image in text form.

[0086] S102. Generate the attention weight matrix corresponding to each outfit image based on the type of each outfit image and the model image.

[0087] To ensure high-quality replacement of each outfit image with the model image and reduce the abruptness of the outfit change edges, a cross-attention mechanism is introduced. This mechanism generates an attention weight matrix for each outfit image based on its type and the model image. As a result, for different outfit image types, the redrawing weight of the model image is higher in more relevant positions when replacing the top view of that type.

[0088] Optionally, in another embodiment of this application, one specific implementation of step S102 is as follows: Figure 2 As shown, it includes:

[0089] S201. Encode the type of each outfit image and the model image respectively to obtain the text features of the type of each outfit image and the latent space map of the model image.

[0090] Specifically, since the dress-up image is text-based, its type can be encoded using a multimodal model to achieve cross-modal semantic understanding and obtain the textual features of the dress-up image's type. Therefore, optionally, the CLIP model can be used to encode the type of each costume image to obtain the textual features of the type of each costume image.

[0091] The model image can be encoded using a variational autoencoder (VAE) to obtain its latent space map. .

[0092] S202. The text features of each outfit image type and the latent space map of the model image are analyzed by using the cross-attention mechanism to obtain the attention weight matrix corresponding to each outfit image.

[0093] Specifically, for each costume image, a cross-attention mechanism is used to calculate the text features of the costume image type. Potential space map of model images Attention weight matrix between:

[0094]

[0095] in, ; ; and represents the learnable parameter matrix; n represents the total number of reference graphs. Optionally, the Q and K values ​​are shared with the cross-attention module in the original model.

[0096] S103. By performing multi-view analysis on each outfit image, a target potential space map of each outfit image with the same angle as the model image is generated.

[0097] Since there are no high standards for the poses of the model and outfit images, their angles can vary. Therefore, to ensure that the angles of the changed clothing and the model images are compatible, a multi-view analysis is performed on the outfit images to obtain their latent space maps from various perspectives. Then, the angles of the model images are analyzed, and latent space maps of outfit images with the same angles are selected. Finally, the latent space maps of the selected outfit images are used to generate target latent space maps for each outfit image with the same angle as the model image.

[0098] Optionally, in another embodiment of this application, one specific implementation of step S103 is as follows: Figure 3 As shown, it includes:

[0099] S301. For each costume image, process the costume image using a multi-branch encoder to obtain the mean and logarithmic variance of the latent space map of the costume image in each target direction.

[0100] The target direction is the set direction, that is, the set view direction. There are usually 6 target directions: forward, right-forward, right, backward, left, and left-forward.

[0101] Specifically, the costume image is encoded using a multi-branch encoder to obtain a latent space map of the costume image in each target direction. The features in the latent space map of the costume image in each target direction are then calculated to obtain the mean and log-variance of the latent space map of the costume image in each target direction. For example, ... Figure 4 As shown, the costume image is input into the multi-branch encoder, which outputs 6 sets of mean and logarithmic variance.

[0102] Optionally, in another embodiment of this application, one specific implementation of step S301 is as follows: Figure 5 As shown, it includes:

[0103] S501. The costume image is feature extracted by five convolutional layers and one residual layer in the multi-branch encoder to obtain the latent space map of the costume image in each target direction.

[0104] like Figure 6 As shown, the multi-branch encoder consists of five convolutional layers, one residual layer, and two fully connected layers. Specifically, the residual layer can be composed of two convolutional layers. Therefore, by using the five convolutional layers and one residual layer in the multi-branch encoder to extract features from the costume image, the latent space map of the costume image in each target direction is obtained.

[0105] S502. The latent space map of the costume image in each target direction is processed by the fully connected layer in the multi-branch encoder to obtain the mean and logarithmic variance of the latent space map of the costume image in each target direction.

[0106] The features of the costume image in the latent space map of each target direction are calculated by using two fully connected layers in the multi-branch encoder, thereby obtaining the mean of the latent space map in each target direction. Sum of logarithmic variance .like Figure 6 As shown, six means were calculated using two fully connected layers respectively. and 6 log-variance , .

[0107] S302. Select the mean and logarithmic variance of the potential space map in the target direction corresponding to the angle range of the model diagram.

[0108] Specifically, such as Figure 4As shown in the figure, an angle calculator is used to calculate the angle of the model image, and the calculated angle, the mean value and logarithmic variance of the potential space map of the model image in the target direction are input into the multi-view diffusion model. The multi-view diffusion model selects the mean value and logarithmic variance of the potential space map in the target direction with an angle close to that of the model image, so as to analyze a potential space map by using the selected mean value and logarithmic variance of the potential space map.

[0109] Specifically, to obtain the angle of the model image by using an angle calculator, the model image can be input into a residual network, and the residual network processes the model image and inputs a processing result. , and the processing result is subjected to angle conversion angle = x * 180, so as to obtain the angle of the model image. .

[0110] Since each target direction corresponds to an angle range, for example: forward corresponds to -22.5° ≤ angle ≤ 22.5°, right forward corresponds to 22.5° < angle ≤ 67.5°, right corresponds to 67.5° < angle ≤ 112.5°, backward corresponds to |angle| > 157.5°, left corresponds to -112.5° ≤ angle < -67.5°, left forward corresponds to -67.5° ≤ angle < -22.5°. Therefore, the target direction where the angle of the model image is located can be determined according to the angle range where the angle of the model image is located. And the potential space map in this target direction is the closest to the angle of the model image, so the potential space map in this target direction is selected.

[0111] S303. Generate the target potential space map of the dressing image by using the mean value and logarithmic variance of the selected potential space map.

[0112] After the mean value and logarithmic variance of the potential space map with the same target direction are selected, the parameters of the feature distribution in this target direction can be determined. Therefore, a potential space map for generating a dressing image consistent with the angle of the model image can be constructed based on this, as the target potential space map of the dressing image.

[0113] Optionally, in another embodiment of the present application, a specific implementation manner of step S303 is as Figure 7 shown, including:

[0114] S701. Calculate the standard deviation of the potential space map by using the logarithmic variance of the selected potential space map.

[0115] Specifically, the calculation formula for the standard deviation σ of the potential space map is: σ = exp(0.5 * log(σ^2)).

[0116] S702. Perform Gaussian distribution sampling based on the standard deviation and mean value of the potential space map to obtain latent variables.

[0117] Specifically, the Gaussian distribution of the latent space map can be obtained through its standard deviation and mean. Therefore, sampling from its Gaussian distribution yields the latent variable z: ;in, From the standard

[0118] S703. Map the latent variables to high-dimensional vectors to obtain the latent high-dimensional vectors.

[0119] Specifically, the latent variable z is mapped to a high-dimensional vector through a fully connected layer, representing a flattened version of the latent space graph, as follows:

[0120]

[0121] in, This is the weight matrix of the fully connected layer, where, ; It is the bias of the fully connected layer; It is the output of the fully connected layer.

[0122] S704. Adjust the potential high-dimensional vector into a three-dimensional feature map to obtain the target potential space map of the costume image.

[0123] In order to process it into a latent space map of the subsequent input format, the output of the fully connected layer is reshaped into a 3D feature map, resulting in a latent space map that serves as the target latent space map for the costume image.

[0124]

[0125] in, .

[0126] S104. Add noise to the set of target latent space maps of each decoration image to obtain the final space features of the decoration.

[0127] It should be noted that the costume image does not need to be changed during the dress-up process; only the model image needs to be changed, replacing the costume image with the model image. Therefore, to accelerate inference and improve dress-up efficiency, this application only adds noise once to the target latent space map of the costume image and freezes its features. Subsequent sampling cycles only change the features of the model image. In other words, this application proposes a cacheable multi-reference diffusion model that saves the feature map of the costume image after one noise addition. When adding noise N times subsequently, only the latent space map of the model image needs to be calculated, thereby effectively reducing computational load and improving processing efficiency.

[0128] Specifically, the target potential space maps of each costume image are combined to obtain a set of potential space maps. Then, noise is added to it to obtain the added features. ,Right now:

[0129]

[0130] Where ε represents a single noise addition process.

[0131] S105. Add noise multiple times to the weighted latent space map of the model image in the overall spatial features to obtain the final spatial features of the model.

[0132] The overall spatial features are obtained by concatenating the weighted latent space map of the model images with the final spatial features of the outfit. The weighted latent space map of the model images is obtained by combining the attention weight matrix corresponding to each outfit image with the latent space map of the model images.

[0133] Specifically, the outfit change is achieved using the latent space maps of the model image and each outfit image. Therefore, the final spatial features of the outfit are concatenated with the latent space map of the model image. To focus on important areas, the attention weight matrix corresponding to each outfit image needs to be combined with the latent space map of the model image; that is, the latent space map of the model image is weighted to obtain a weighted latent space map of the model image, which is then concatenated with the final spatial features of the outfit to obtain the overall spatial features. Then, during the iterative addition of noise, the final spatial features of the outfit are frozen, and noise is added only to the weighted latent space map of the model image within the overall spatial features each time.

[0134] Therefore, for a model image, after encoding by an encoder, the latent space map of the model image is obtained. Then, after adding noise N times, the final spatial features of the model are obtained:

[0135]

[0136] That is, such as Figure 8 As shown, first, noise is added to the set of target latent space maps (I1, I2, I3, I4) of the costume image to obtain (E1, E2, E3, E4). Subsequently, noise is added to the latent space maps of (E1, E2, E3, E4) and the model image. Figure X When adding noise using a combination of methods, only the latent space of the model image is considered. Figure X Add noise to change it from X to The outputs (E1, E2, E3, E4) remain unchanged, thus eliminating the need to add noise to the features of the weights each time.

[0137] S106. Redraw the model based on its final spatial features to obtain a model image in different outfits.

[0138] Specifically, the image is drawn according to the final spatial characteristics of the model, thus obtaining the image of the model after the change of clothes.

[0139] Optionally, in another embodiment of this application, one specific implementation of step S104 includes:

[0140] By adding noise to the set of target latent space maps of each costume image using a cacheable multi-reference diffusion model, the final spatial features of the costume are obtained.

[0141] In other words, this application provides a cacheable multi-reference diffusion model that freezes the features of multiple outfit images after adding noise once, reducing the amount of computation when adding noise repeatedly, accelerating the inference process, and enabling real-time outfit changing.

[0142] Accordingly, embodiments of this application provide a training method for a cacheable multi-reference diffusion model, such as... Figure 9 As shown, it includes:

[0143] S901. The acquired training data is processed using a cacheable multi-reference diffusion model to obtain the current costume image.

[0144] The training data can include multiple sets. Each set of training data includes multiple outfit sample images and one model sample image as samples. It also includes images of the outfit after changing only to each outfit sample image and images of the outfit after changing to all outfit sample images, as sample labels. This is to facilitate subsequent analysis of the loss of the outfit images after changing to each outfit sample image and the outfit images after changing to all outfit sample images, thereby improving the outfit changing effect of each outfit sample and the overall coordinated outfit changing effect.

[0145] Specifically, the specific implementation of step S901 is the same as that of steps S104 to S106, so it will not be described again.

[0146] S902. Based on the deviation between the current costume image and the costume images in the training data, calculate the current first loss and each current second loss.

[0147] Here, the current first loss L is the loss when replacing all the outfit sample images in a set of training data. Each current second loss Li is the loss when replacing only a single outfit sample image in a set of training data.

[0148] Specifically, the loss can be calculated by sampling the following loss function L:

[0149]

[0150] Where E is the expectation; The image is the original image; t is the time step. is the noise added to the original image during forward diffusion; c is the condition, i.e., the camera parameters; h(c) is the embedding feature of the reference image; The noise predicted by the diffusion model; This is the diffused image.

[0151] S903. Determine whether the current total loss has converged.

[0152] The current total loss is calculated based on each current second loss and the current first loss.

[0153] Optionally, the current total loss is the sum of the mean of all current second losses and the current first loss. Therefore, the current total loss is specifically:

[0154]

[0155] If it is determined that the current total loss has not converged, it means that the model can be further optimized, so proceed to step S904. If it is determined that the current total loss has converged, then proceed to step S905.

[0156] S904. Update the parameters of the cacheable multi-reference diffusion model based on each current second loss and the current first loss.

[0157] It should be noted that after updating the parameters of the cacheable multi-reference diffusion model, the process returns to step S901 to perform the next round of training.

[0158] S905, The cacheable multi-reference diffusion model has been determined and the process is complete.

[0159] This application provides a multi-angle costume-changing method. It acquires a model image and at least one costume image, and determines the type of each costume image, thus supporting the uploading of multiple types of costume images for simultaneous costume-changing. Based on the type of each costume image and the model image, an attention weight matrix is ​​generated for each costume image. This matrix is ​​used during subsequent redrawing to ensure different weights for different areas of the model image, better dividing the redrawing area, reducing the abruptness of the costume-changing edges, coordinating the costume-changing process across different images, and ensuring the costume-changing effect for each image. By performing multi-view analysis on each costume image, target latent space maps of each costume image with angles consistent with the model image are generated. This multi-view analysis ensures that the model image and costume image have consistent angles, eliminating the need for poses with high modal requirements. Then, noise is added to the set of target latent space maps of each costume image to obtain the final costume spatial features. Multiple noise additions are then performed on the weighted latent space map of the model image within the overall spatial features to obtain the final model spatial features. The overall spatial features are obtained by concatenating the weighted latent space map of the model images with the final spatial features of the outfit. The weighted latent space map of the model images is obtained by combining the attention weight matrix corresponding to each outfit image with the latent space map of the model images. Finally, the model images are redrawn based on the final spatial features of the model to obtain the outfit-changing model images. This freezes the features of multiple outfit images after adding noise once. Subsequently, noise is only added to the features of the weighted model images in a cyclical manner, reducing the computational load during cyclic noise addition, accelerating the inference process, and realizing real-time outfit changing. Therefore, a method for high-quality and timely outfit changing of multiple outfit images simultaneously is realized.

[0160] Another embodiment of this application provides a multi-angle changing system, such as Figure 10 As shown, it includes:

[0161] The acquisition unit 1001 is used to acquire model images and at least one outfit image, and to determine the type of each outfit image.

[0162] The weight determination unit 1002 is used to generate the attention weight matrix corresponding to each outfit image based on the type of each outfit image and the model image.

[0163] The feature generation unit 1003 is used to generate target latent space maps of each outfit image with the same angle as the model image by performing multi-view analysis on each outfit image.

[0164] The decoration noise addition unit 1004 is used to add noise to the set of target latent space maps of each decoration image to obtain the final decoration space features.

[0165] Unit 1005 iteratively adds noise multiple times to the weighted latent space map of the model image in the overall spatial features to obtain the final spatial features of the model. The overall spatial features are obtained by concatenating the weighted latent space map of the model image with the final spatial features of the outfit. The weighted latent space map of the model image is obtained by combining the attention weight matrix corresponding to each outfit image with the latent space map of the model image.

[0166] The image generation unit 1006 is used to redraw the model based on the final spatial features of the model to obtain the model image after changing clothes.

[0167] Optionally, in another embodiment of the multi-angle dressing system provided in this application, the weight determination unit includes:

[0168] The encoding unit is used to encode the type of each outfit image and the model image respectively, so as to obtain the text features of the type of each outfit image and the latent space map of the model image.

[0169] The weight matrix generation unit is used to analyze the combination of text features of each outfit image type and latent space map of the model image through a cross-attention mechanism to obtain the attention weight matrix corresponding to each outfit image.

[0170] Optionally, in another embodiment of the multi-angle dressing system provided in this application, the feature generation unit includes:

[0171] The multi-branch processing unit is used to process each costume image separately using a multi-branch encoder to obtain the mean and logarithmic variance of the latent space map of the costume image in each target direction.

[0172] The selection unit is used to select the mean and logarithmic variance of the latent space map corresponding to the angular range of the model image's angle in the target direction. The angle of the model image is obtained by processing the model image through a residual network and then performing angle transformation on the processing result.

[0173] The feature map generation unit is used to generate a target latent space map of the costume map using the mean and logarithmic variance of the selected latent space map.

[0174] Optionally, in another embodiment of the multi-angle changing system provided in this application, the multi-branch processing unit includes:

[0175] The feature extraction unit is used to extract features from the costume image through five convolutional layers and one residual layer in the multi-branch encoder to obtain the latent space map of the costume image in each target direction.

[0176] The parameter calculation unit is used to process the latent space map of the costume image in each target direction through the fully connected layer in the multi-branch encoder to obtain the mean and logarithmic variance of the latent space map of the costume image in each target direction.

[0177] Optionally, in another embodiment of the multi-angle changing system provided in this application, the selected unit includes:

[0178] The standard deviation calculation unit is used to calculate the standard deviation of the latent space map using the logarithmic variance of the selected latent space map.

[0179] The sampling unit is used to perform Gaussian distribution sampling based on the standard deviation and mean of the latent space map to obtain latent variables.

[0180] The mapping unit is used to map latent variables to high-dimensional vectors, resulting in latent high-dimensional vectors.

[0181] The dimension adjustment unit is used to adjust the potential high-dimensional vector into a three-dimensional feature map to obtain the target latent space map of the costume image.

[0182] Optionally, in another embodiment of the multi-angle dressing system provided in this application, the system further includes:

[0183] The training unit is used to process the acquired training data through a cacheable multi-reference diffusion model to obtain the current costume image.

[0184] The loss calculation unit is used to calculate the current first loss and each of the current second losses based on the deviation between the current costume image and the costume images in the training data. The current first loss is the loss when all costume sample images from the training data are used. The current second loss is the loss when only a single costume sample image from the training data is used.

[0185] The parameter update unit is used to update the parameters of the cacheable multi-reference diffusion model based on each current second loss and the current first loss, and then return to the training unit until the total loss converges.

[0186] It should be noted that the specific working process of each unit provided in the above embodiments of this application can be referred to the implementation process of the corresponding steps in the above method embodiments, and will not be repeated here.

[0187] Another embodiment of this application provides an electronic device, such as... Figure 11 As shown, it includes:

[0188] Memory 1101 and processor 1102.

[0189] The memory 1101 is used to store the program.

[0190] The processor 1102 is used to execute the program stored in the memory 1101. When the program is executed, it is specifically used to implement the multi-angle dressing method provided in any of the above embodiments.

[0191] Another embodiment of this application provides a computer storage medium for storing a computer program, which, when executed by a processor, is used to implement the multi-angle dressing method provided in any of the above embodiments.

[0192] Computer storage media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0193] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0194] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multi-angle costume changing method, characterized in that, include: Obtain a model image and at least one outfit image, and determine the type of each outfit image; Based on the type of each outfit image and the model image, generate an attention weight matrix corresponding to each outfit image; By performing multi-view analysis on each of the aforementioned outfit images, a target latent space map of each outfit image with the same angle as the model image is generated; A noise addition is performed on the set of target latent space maps of each of the aforementioned costume images to obtain the final spatial features of the costume. Multiple noise additions are made to the weighted latent space map of the model image in the overall spatial features to obtain the final spatial features of the model; wherein, the overall spatial features are obtained by concatenating the weighted latent space map of the model image and the final spatial features of the outfit; the weighted latent space map of the model image is obtained by combining the attention weight matrix corresponding to each outfit image with the latent space map of the model image; Based on the final spatial features of the model, a redrawing is performed to obtain a model image in different outfits.

2. The method according to claim 1, characterized in that, The step of generating an attention weight matrix corresponding to each of the outfit images based on the type of each outfit image and the model image includes: The type of each outfit image and the model image are encoded to obtain the text features of the type of each outfit image and the latent space map of the model image. The text features of each type of costume image and the latent space map of the model image are analyzed using a cross-attention mechanism to obtain the attention weight matrix corresponding to each costume image.

3. The method according to claim 1, characterized in that, The step of generating target latent space maps of each of the costume images by performing multi-view analysis on each of the costume images, with the same angle as the model image, includes: For each of the costume images, a multi-branch encoder is used to process the costume image to obtain the mean and logarithmic variance of the latent space map of the costume image in each target direction; The mean and logarithmic variance of the latent space map in the target direction corresponding to the angle range of the model image are selected; wherein, the angle of the model image is obtained by processing the model image through a residual network and then performing angle transformation on the processing result; The target latent space map of the costume image is generated using the mean and logarithmic variance of the selected latent space map.

4. The method according to claim 3, characterized in that, The process of processing the costume image using a multi-branch encoder to obtain the mean and logarithmic variance of the latent space map of the costume image in each target direction includes: The costume image is feature extracted by five convolutional layers and one residual layer in a multi-branch encoder to obtain the latent space map of the costume image in each target direction. The latent space graphs of the costume image in each target direction are processed by the fully connected layer in the multi-branch encoder to obtain the mean and logarithmic variance of the latent space graphs of the costume image in each target direction.

5. The method according to claim 3, characterized in that, The step of generating the target latent space map of the costume image using the mean and logarithmic variance of the selected latent space map includes: The standard deviation of the latent space map is calculated using the log-variance of the selected latent space map; Gaussian distribution sampling is performed based on the standard deviation and mean of the latent space map to obtain latent variables; The latent variables are mapped to high-dimensional vectors to obtain latent high-dimensional vectors; The potential high-dimensional vector is adjusted into a three-dimensional feature map to obtain the target potential space map of the costume image.

6. The method according to claim 1, characterized in that, The step of adding noise to the set of target latent space maps of each of the aforementioned costume images to obtain the final spatial features of the costume includes: The set of target latent space maps of each of the aforementioned costume images is noise-added using a cacheable multi-reference diffusion model to obtain the final spatial features of the costume. The training method for the cacheable multi-reference diffusion model includes: The acquired training data is processed using the cacheable multi-reference diffusion model to obtain the current costume change image; Based on the deviation between the current outfit image and the outfit images in the training data, a current first loss and various current second losses are calculated; wherein, the current first loss is the loss when all the outfit sample images in the training data are used; and the current second loss is the loss when only a single outfit sample image in the training data is used. The parameters of the cacheable multi-reference diffusion model are updated based on the current second loss and the current first loss, and the process of processing the acquired training data through the cacheable multi-reference diffusion model is returned to obtain the current costume image until the total loss converges.

7. A multi-angle changing system, characterized in that, include: An acquisition unit is used to acquire a model image and at least one outfit image, and to determine the type of each outfit image; The weight determination unit is used to generate an attention weight matrix corresponding to each of the costume images based on the type of each costume image and the model image, respectively. The feature generation unit is used to generate target latent space maps of each of the costume images by performing multi-view analysis on each of the costume images, with the same angle as the model image. The decoration noise addition unit is used to add noise to the set of target latent space maps of each decoration image to obtain the final decoration space features; The cyclic addition unit adds noise multiple times to the weighted latent space map of the model image in the overall spatial features to obtain the final spatial features of the model; wherein, the overall spatial features are obtained by concatenating the weighted latent space map of the model image and the final spatial features of the outfit; the weighted latent space map of the model image is obtained by combining the attention weight matrix corresponding to each outfit image with the latent space map of the model image; An image generation unit is used to redraw the model based on the model's final spatial features to obtain a model image in different outfits.

8. The system according to claim 7, characterized in that, The weight determination unit includes: An encoding unit is used to encode the type of each of the costume images and the model image respectively, to obtain the text features of the type of each costume image and the latent space map of the model image; The weight matrix generation unit is used to analyze the combination of the text features of the type of each costume image and the latent space map of the model image through a cross-attention mechanism to obtain the attention weight matrix corresponding to each costume image.

9. An electronic device, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is used to execute the program, which, when executed, is specifically used to implement the multi-angle dressing method as described in any one of claims 1 to 6.

10. A computer storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, is used to implement the multi-angle dressing method as described in any one of claims 1 to 6.