Image processing method, apparatus and device

By using the clothing and posture feature extraction module in the intelligent drawing model, a dress-up effect image of a portrait photo is generated, which solves the problems of long time consumption and high cost in the existing technology and realizes fast and low-cost dress-up effect generation.

CN117252777BActive Publication Date: 2026-05-29HANGZHOU QUWEI SCI & TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU QUWEI SCI & TECH
Filing Date
2023-09-27
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies lack simple and effective methods for changing outfits in a single photo in the field of portrait photo enhancement. They cannot quickly generate various different outfit effects, and existing methods are time-consuming, costly, and cannot effectively preserve the user's original pose.

Method used

By acquiring the image to be processed and the descriptive text, the clothing feature extraction module and posture feature extraction module in the intelligent drawing model are used to generate an initial dressing effect image, preserving the consistency of human posture and facial area, thus enabling rapid dressing.

Benefits of technology

It enables the rapid generation of costume renderings without constructing a virtual human body model, improving the efficiency of costume generation and reducing application costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117252777B_ABST
    Figure CN117252777B_ABST
Patent Text Reader

Abstract

The application provides an image processing method, device and equipment, and relates to the technical field of image processing. The method comprises the following steps: acquiring a to-be-processed image and a description text; determining a first human body mask, a face mask and a human body posture graph according to the to-be-processed image; then processing the description text by using a clothing feature extraction module in a preset intelligent drawing model to obtain target clothing features corresponding to the description text; processing the human body posture graph by using a posture feature extraction module in the intelligent drawing model to obtain human body posture features; and processing the target clothing features, the human body posture features and the face mask by using an image generation module in the intelligent drawing model to generate an initial dressing effect graph. Through the determination of the target clothing features and the human body posture features, the image generation module generates the initial dressing effect graph corresponding to the description text, and the human body posture in the initial dressing effect graph is consistent with the human body posture in the to-be-processed image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and more specifically, to an image processing method, apparatus, and device. Background Technology

[0002] Currently, there are few simple and effective methods for changing outfits in a single photo in the field of portrait photo beautification that can meet users' needs to quickly generate various different outfit effects after taking a photo. Taking scenic spot photos as an example, if users want to see various ancient style outfit effects through just a natural photo without changing clothes and makeup, there is currently no effective solution.

[0003] Current technologies can achieve the effect of changing clothes by constructing a complete virtual human body model through 3D human body modeling, or by constructing a realistic human body model through multiple photos taken of the user in special environments, or by constructing a model training for a specific look through a massive dataset of paired clothing photos. However, existing methods are time-consuming, costly, have poor changing effects, and cannot effectively preserve the user's original posture. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of the prior art by providing an image processing method, apparatus, and device. This allows for the rapid alteration of the image to be processed by determining the characteristics of the target clothing and the human posture characteristics of the image to be processed. The initial alteration effect image generated by the image generation module is consistent with the human posture in the initial alteration effect image.

[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows:

[0006] In a first aspect, embodiments of this application provide an image processing method, including:

[0007] Obtain the image to be processed and the descriptive text;

[0008] Based on the image to be processed, determine the first human body mask, the face mask, and the human body pose map;

[0009] Based on the description text, the clothing feature extraction module in the preset intelligent drawing model is used for processing to obtain the target clothing features corresponding to the description text;

[0010] Based on the human posture diagram, the posture feature extraction module in the intelligent drawing model is used for processing to obtain human posture features;

[0011] Based on the target clothing features, the human posture features, and the facial mask, the image generation module in the intelligent drawing model is used to process the data and generate an initial dress-up effect image.

[0012] In an optional implementation, the method further includes:

[0013] Based on the initial costume change effect image, the first human body mask, and the facial mask, a target human body mask is generated;

[0014] Based on the target clothing features, the human body posture features, and the target human body mask, the image generation module in the intelligent drawing model is used to process the data and generate a target clothing effect image.

[0015] In an optional implementation, generating the target human body mask based on the initial costume change effect image, the first human body mask, and the facial mask includes:

[0016] Based on the initial costume change effect diagram, determine the second human body mask;

[0017] A third human body mask is generated by mixing the second human body mask and the first human body mask.

[0018] The target human body mask is generated by mixing the reverse images of the third human body mask and the face mask.

[0019] In an optional implementation, the step of mixing the second human body mask and the first human body mask to generate a third human body mask includes:

[0020] The third human body mask is generated by blending the second human body mask and the first human body mask with a brightening layer.

[0021] In an optional implementation, the step of generating the third human body mask by blending the brightening layers based on the second human body mask and the first human body mask includes:

[0022] Determine the first brightness value of each first pixel position in the first human body mask in multiple color channels;

[0023] Determine the position of each first pixel in the second human body mask at the second brightness value of the plurality of color channels;

[0024] Based on the maximum brightness value of each first pixel position in the first brightness value and the second brightness value of each color channel, the first target pixel value of each first pixel position in each color channel is determined respectively;

[0025] The third human body mask is generated based on the first target pixel value of each first pixel position in the plurality of color channels.

[0026] In an optional implementation, the step of generating the target human body mask by mixing the inverse images of the third human body mask and the facial mask includes:

[0027] The target human body mask is generated by blending darkening layers based on the inverse images of the third human body mask and the face mask.

[0028] In an optional implementation, the step of generating the target human body mask by performing darkening layer blending based on the inverse image of the third human body mask and the face mask includes:

[0029] Determine the third brightness value of each second pixel position in the third human body mask in multiple color channels;

[0030] The position of each second pixel in the facial mask is determined by the fourth luminance value of the plurality of color channels;

[0031] Based on the minimum brightness value among the inverse brightness values ​​of the third and fourth brightness values ​​of each second pixel position in each color channel, the second target pixel value of each second pixel position in each color channel is determined respectively.

[0032] The target human body mask is generated based on the second target pixel value of each second pixel position in the plurality of color channels.

[0033] In an optional implementation, before processing the description text using a clothing feature extraction module in a preset intelligent drawing model to obtain the target clothing features corresponding to the description text, the method further includes:

[0034] Acquire multiple clothing sample images and a tagging file for each clothing sample image, wherein the tagging file for each clothing sample image contains sample description text;

[0035] Based on the multiple clothing sample images and their corresponding labeling files, a preset image style adapter model is trained to generate the clothing feature extraction module.

[0036] Secondly, embodiments of this application also provide an image processing apparatus, the apparatus comprising:

[0037] The acquisition module is used to acquire the image to be processed and the descriptive text.

[0038] The determination module is used to determine a first human body mask, a face mask, and a human body pose map based on the image to be processed.

[0039] The processing module is used to process the description text using the clothing feature extraction module in the preset intelligent drawing model to obtain the target clothing features corresponding to the description text.

[0040] The processing module is further configured to process the human posture diagram using the posture feature extraction module in the intelligent drawing model to obtain human posture features.

[0041] The processing module is further configured to process the target clothing features, the human body posture features, and the facial mask using the image generation module in the intelligent drawing model to generate a preliminary dress-up effect image.

[0042] Thirdly, embodiments of this application also provide a computer device, including: a processor, a storage medium, and a bus, wherein the storage medium stores program instructions executable by the processor, and when the computer device is running, the processor communicates with the storage medium via the bus, and the processor executes the program instructions to perform the steps of the image processing method as described in any of the first aspects.

[0043] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the image processing method as described in any of the first aspects.

[0044] The beneficial effects of this application are:

[0045] This application provides an image processing method, apparatus, and device, comprising: acquiring an image to be processed and descriptive text; determining a first human body mask, a face mask, and a human body pose diagram based on the image to be processed; then processing the image based on the descriptive text using a clothing feature extraction module in a preset intelligent drawing model to obtain target clothing features corresponding to the descriptive text; processing the image based on the human body pose diagram using a pose feature extraction module in the intelligent drawing model to obtain human body pose features; and finally processing the image based on the target clothing features, human body pose features, and face mask using an image generation module in the intelligent drawing model to generate an initial dress-up effect image. The method of this application determines the target clothing features and human posture features through the clothing feature extraction module and posture feature extraction module in the preset intelligent drawing model. This enables the image generation module to generate an initial dress-up effect image corresponding to the descriptive text based on the target clothing features. Furthermore, the method controls the human posture in the generated initial dress-up effect image to be consistent with the human posture in the image to be processed based on the human posture features. In addition, since the facial mask is input to the image generation module for processing, the facial area in the initial dress-up effect image is consistent with the facial area in the image to be processed. Therefore, it is not necessary to construct a virtual human body model through 3D modeling to achieve dress-up. Only the image to be processed needs to be input and processed by the intelligent drawing model to achieve dress-up, thereby improving dress-up generation efficiency and reducing application costs. Attached Figure Description

[0046] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 A structural diagram of an intelligent drawing model provided in an embodiment of this application;

[0048] Figure 2 This is one of the schematic flowcharts of an image processing method provided in an embodiment of this application;

[0049] Figure 3 This is a second schematic flowchart of an image processing method provided in an embodiment of this application;

[0050] Figure 4 The third schematic flowchart of an image processing method provided in this application embodiment;

[0051] Figure 5 The fourth schematic flowchart of an image processing method provided in this application embodiment;

[0052] Figure 6 Fifth schematic flowchart of an image processing method provided in this application embodiment;

[0053] Figure 7 A schematic flowchart of an image processing method provided in this application embodiment is shown in Figure 6.

[0054] Figure 8 This is a schematic diagram of the functional modules of an image processing apparatus provided in an embodiment of this application;

[0055] Figure 9 This is a schematic diagram of a computer device provided in an embodiment of this application. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.

[0057] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0058] In the description of this application, it should be noted that if the terms "upper", "lower", etc. appear to indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship that the product of this application is usually placed in, it is only for the convenience of describing this application and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.

[0059] Furthermore, the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Additionally, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0060] It should be noted that, where there is no conflict, the features in the embodiments of this application can be combined with each other.

[0061] To achieve rapid clothing replacement in images to be processed, this application provides an image processing method. First, the image to be processed and descriptive text are acquired. The image to be processed is then processed to determine a first human body mask, a facial mask, and a human pose diagram. Next, the clothing feature extraction module in a preset intelligent drawing model is used to process the descriptive text to obtain the target clothing features corresponding to the descriptive text. Then, the pose feature extraction module in the intelligent drawing model is used to process the human pose diagram to obtain human pose features. Finally, based on the target clothing features, human pose features, and facial mask, the image generation module in the intelligent drawing model is used to generate an initial clothing replacement effect image. This initial clothing replacement effect image can replace the clothing in the image to be processed, achieving rapid clothing replacement.

[0062] Figure 1 A structural diagram of an intelligent drawing model provided in an embodiment of this application is shown below. Figure 1As shown, the intelligent drawing model 100 is used to generate an initial dress-up effect image based on the image to be processed. The intelligent drawing model 100 includes: a clothing feature extraction module 110, a posture feature extraction module 120, and an image generation module 130. The clothing feature extraction module 110 is used to process the descriptive text to obtain the target clothing features corresponding to the descriptive text; the posture feature extraction module 120 is used to process the human posture image to obtain the human posture features; and the image generation module 130 is used to generate an initial dress-up effect image based on the target clothing features, human posture features, and a facial mask.

[0063] The image processing method provided in this application embodiment will be explained in detail below with reference to the accompanying drawings and specific examples. The image processing method provided in this application embodiment can be implemented by a computer device pre-installed with a preset image generation algorithm or detection software, by running the algorithm or software. The computer device can be, for example, a server or a terminal, and the terminal can be a user computer. Figure 2 This is one of the schematic flowcharts of an image processing method provided in an embodiment of this application. Figure 2 As shown, the method includes:

[0064] S101. Obtain the image to be processed and the description text.

[0065] In this embodiment, the image to be processed is selected and input by the user, and the user inputs descriptive text, which is the clothing change requirement of the image to be processed. The descriptive text is, for example, descriptions such as long hair, red clothing, etc., so that the computer device can obtain the image to be processed and the descriptive text input by the user, and perform subsequent processing based on the image to be processed and the descriptive text. The image to be processed is a portrait image, including a portrait area and a background area.

[0066] It should be noted that if the user does not enter any descriptive text, the descriptive text will be the default text, and subsequent processing will be performed based on the acquired image to be processed and the default text.

[0067] S102. Based on the image to be processed, determine the first human body mask, the face mask, and the human body pose map.

[0068] Specifically, a preset segmentation algorithm is used to segment the human body region and facial region of the image to be processed, thereby obtaining the first human body mask and facial mask. The preset segmentation algorithm can be a semantic segmentation network such as U-net, SegNet, Deeplab, or SegmentAnything.

[0069] It also employs a pre-defined human pose estimation algorithm to identify key human points in the image to be processed, thereby obtaining a human pose map. The pre-defined human pose estimation algorithm can be the deep learning-based human pose estimation framework OpenPose.

[0070] The image to be processed is then processed according to the preset segmentation algorithm and the preset human pose estimation algorithm to determine the first human mask, the face mask, and the human pose map.

[0071] S103. Based on the description text, the clothing feature extraction module in the preset intelligent drawing model is used for processing to obtain the target clothing features corresponding to the description text.

[0072] S104. Based on the human posture diagram, the posture feature extraction module in the intelligent drawing model is used for processing to obtain the human posture features.

[0073] The preset intelligent drawing model is used to process the image to be processed and finally generate the initial dress-up effect image. The preset intelligent drawing model can adopt the intelligent painting generation model Stable Diffusion.

[0074] The pre-defined intelligent drawing model includes a clothing feature extraction module, which converts descriptive text into target clothing features corresponding to the descriptive text. Specifically, the clothing feature extraction module first performs text reversal on the descriptive text to obtain word vectors, and then determines the target clothing features corresponding to the descriptive text based on the word vectors.

[0075] The pose feature extraction module is a neural network model called Controlnet, which can be used as an extension model of the preset intelligent drawing model. The pose feature extraction module controls the preset intelligent drawing model, which is a diffusion model used to generate images based on text or images. During the image generation process, the pose feature extraction module can introduce human pose images to intervene in the image generation process. Specifically, the pose feature extraction module processes the human pose images to obtain human pose features, thereby controlling the generation of the final costume effect image through human pose features.

[0076] S105. Based on the characteristics of the target clothing, the characteristics of the human posture, and the facial mask, the image generation module in the intelligent drawing model is used to process the image and generate an initial dress-up effect image.

[0077] Specifically, the preset intelligent drawing model also includes an image generation module. The image generation module processes the target clothing features, human posture features, and facial mask to obtain an initial dress-up effect image. The clothing feature extraction module and posture feature extraction module determine the target clothing features and human posture features, so that the initial dress-up effect image can retain the human posture and facial area in the image to be processed, thereby realizing the replacement of clothing in the human portrait area of ​​the image to be processed.

[0078] In summary, this application provides an image processing method, including: acquiring an image to be processed and descriptive text; determining a first human body mask, a face mask, and a human body pose map based on the image to be processed; then processing the descriptive text using a clothing feature extraction module in a preset intelligent drawing model to obtain target clothing features corresponding to the descriptive text; processing the human body pose map using a pose feature extraction module in the intelligent drawing model to obtain human body pose features; and finally processing the target clothing features, human body pose features, and face mask using an image generation module in the intelligent drawing model to generate an initial dress-up effect image. The method of this application determines the target clothing features and human posture features through the clothing feature extraction module and posture feature extraction module in the preset intelligent drawing model. This enables the image generation module to generate an initial dress-up effect image corresponding to the descriptive text based on the target clothing features. Furthermore, the method controls the human posture in the generated initial dress-up effect image to be consistent with the human posture in the image to be processed based on the human posture features. In addition, since the facial mask is input to the image generation module for processing, the facial area in the initial dress-up effect image is consistent with the facial area in the image to be processed. Therefore, it is not necessary to construct a virtual human body model through 3D modeling to achieve dress-up. Only the image to be processed needs to be input and processed by the intelligent drawing model to achieve dress-up, thereby improving dress-up generation efficiency and reducing application costs.

[0079] This application also provides another possible implementation of the image processing method. Figure 3 This is a second schematic flowchart illustrating an image processing method provided in an embodiment of this application. Figure 3 As shown, the method also includes:

[0080] S201. Generate the target human body mask based on the initial costume effect image, the first human body mask, and the facial mask.

[0081] In this embodiment, the human body mask, the first human body mask, and the face mask corresponding to the initial costume change effect image are processed to obtain the target human body mask.

[0082] S202. Based on the characteristics of the target clothing, the characteristics of the human body posture, and the target human body mask, the image generation module in the intelligent drawing model is used to process the data and generate a target clothing effect image.

[0083] Specifically, the image generation module processes the target clothing features, human posture features, and target human mask to obtain the target clothing effect image. The clothing feature extraction module and posture feature extraction module determine the target clothing features and human posture features, so that the initial clothing effect image can retain the human posture and target human region in the image to be processed, thereby realizing the replacement of clothing in the target human region of the image to be processed.

[0084] In the method provided in this application embodiment, a target human body mask is generated based on an initial costume change effect image, a first human body mask, and a facial mask. Then, based on the target clothing features, human posture features, and the target human body mask, an image generation module in an intelligent drawing model is used to process the image to generate a target costume change effect image. Since the initial costume change effect image only retains the facial region and human posture, the human body region in the initial costume change effect image is inconsistent with the human body region of the image to be processed, and the background region is not retained. In order to obtain an accurate costume change effect image that retains the human body region and background region of the image to be processed, the initial costume change effect image, the first human body mask, and the facial mask are processed to obtain the target human body mask. Finally, based on the target clothing features, human posture features, and the target human body mask, an image generation module in an intelligent drawing model is used to process the image to generate the target costume change effect image, so that the target costume change effect image retains most of the background region and facial region, and only the clothing and hairstyle of the human body region are replaced.

[0085] This application also provides another possible implementation of the image processing method. Figure 4 This is a third schematic flowchart illustrating an image processing method provided in an embodiment of this application. Figure 4 As shown, based on the initial costume change effect image, the first human body mask, and the facial mask, a target human body mask is generated, including:

[0086] S301. Determine the second human body mask based on the initial costume change effect diagram.

[0087] In this embodiment, a preset segmentation algorithm is used to segment the human body region of the initial costume effect image to obtain a second human body mask. The preset segmentation algorithm can be a semantic segmentation network such as U-net, SegNet, Deeplab, or Segment Anything.

[0088] S302. The second human body mask and the first human body mask are mixed to generate a third human body mask.

[0089] Optionally, a third human body mask can be generated by blending the second and first human body masks with a brightening layer.

[0090] Specifically, the brightness values ​​of multiple color channels at each pixel position of the second and first human body masks are brightened and mixed to obtain the third human body mask.

[0091] S303. Blend the reverse images of the third human body mask and the face mask to generate the target human body mask.

[0092] Optionally, a darkening layer blending is performed based on the inverse image of the third human body mask and the face mask to generate the target human body mask.

[0093] Specifically, the brightness values ​​of multiple color channels at each pixel position of the reverse image of the third human body mask and the face mask are darkened and mixed to obtain the target human body mask.

[0094] In the method provided in this application embodiment, a second human body mask is determined based on an initial costume change effect image. Then, the second human body mask and the first human body mask are mixed to generate a third human body mask. Finally, the third human body mask and the reverse image of the facial mask are mixed to generate a target human body mask. To obtain an accurate costume change effect image, the human body area and background area of ​​the image to be processed are preserved. Therefore, the second human body mask and the first human body mask are mixed to obtain a third human body mask. The human body area of ​​the final costume change effect image, i.e., the area to be costumed, is determined. Then, the third human body mask and the reverse image of the facial mask are mixed to generate a target human body mask. The human body area of ​​the final costume change effect image, i.e., the area to be costumed, and the background area are determined. Thus, an accurate target costume change effect image can be generated through the image generation module, so that most of the background area and facial area are preserved in the target costume change effect image, and only the clothing and hairstyle of the human body area are replaced.

[0095] This application also provides another possible implementation of the image processing method. Figure 5 This is a fourth schematic flowchart illustrating an image processing method provided in an embodiment of this application. Figure 5 As shown, a third human body mask is generated by blending the second and first human body masks using a brightening layer, including:

[0096] S401. Determine the first brightness value of each first pixel position in the first human body mask in multiple color channels.

[0097] S402. Determine the second brightness value of each first pixel position in the second human body mask in multiple color channels.

[0098] In this embodiment, if the first human body mask is MaskA and the first pixel position is (i, j), then the first brightness values ​​of the three color channels of the first human body mask MaskA at position (i, j) are determined to be (R_MaskA, G_MaskA, B_MaskA), where R_MaskA is the brightness value of the red channel of the first human body mask MaskA at position (i, j), G_MaskA is the brightness value of the green channel of the first human body mask MaskA at position (i, j), and B_MaskA is the brightness value of the blue channel of the first human body mask MaskA at position (i, j).

[0099] Similarly, if the second human body mask is MaskB and the first pixel position is (i, j), then the second brightness values ​​of the three color channels of the second human body mask MaskB at position (i, j) are determined to be (R_MaskB, G_MaskB, B_MaskB), where R_MaskB is the brightness value of the red channel of the second human body mask MaskB at position (i, j), G_MaskB is the brightness value of the green channel of the second human body mask MaskB at position (i, j), and B_MaskB is the brightness value of the blue channel of the second human body mask MaskB at position (i, j).

[0100] S403. Based on the maximum brightness value of the first brightness value and the second brightness value of each first pixel position in each color channel, determine the first target pixel value of each first pixel position in each color channel.

[0101] S404. Generate a third human body mask based on the first target pixel values ​​of each first pixel position in multiple color channels.

[0102] Specifically, the first target pixel value at the first pixel position (i, j) in each color channel is represented as follows:

[0103] R_MaskC=max(R_MaskA, R_MaskB);

[0104] G_MaskC=max(G_MaskA, G_MaskB);

[0105] B_MaskC=max(B_MaskA, B_MaskB);

[0106] Where R_MaskC is the first target pixel value of the red channel at position (i, j), G_MaskC is the first target pixel value of the green channel at position (i, j), and B_MaskC is the first target pixel value of the blue channel at position (i, j).

[0107] Then the first target pixel values ​​of the three color channels of the third human body mask MaskC at position (i, j) are determined to be (R_MaskC, G_MaskC, B_MaskC).

[0108] In the method provided in this application embodiment, the first brightness value of each first pixel position in a first human body mask in multiple color channels is determined; the second brightness value of each first pixel position in a second human body mask in multiple color channels is determined; based on the maximum brightness value among the first and second brightness values ​​of each first pixel position in each color channel, the first target pixel value of each first pixel position in each color channel is determined; finally, a third human body mask is generated based on the first target pixel values ​​of each first pixel position in multiple color channels. This allows the third human body mask to be used to determine the human body area, i.e., the area requiring costume change, in the final costume change effect image.

[0109] This application also provides another possible implementation of the image processing method. Figure 6 This is the fifth schematic flowchart illustrating an image processing method provided in an embodiment of this application. Figure 6 As shown, based on the inverse images of the third human body mask and the face mask, a darkening layer blending is performed to generate the target human body mask, including:

[0110] S501. Determine the third brightness value of each second pixel position in multiple color channels in the third human body mask.

[0111] S502, Determine the fourth brightness value of each second pixel position in the face mask in multiple color channels.

[0112] In this embodiment, if the third human body mask is MaskC and the second pixel position is (a, b), then the third brightness values ​​of the three color channels of the third human body mask MaskC at position (a, b) are determined to be (R_MaskC, G_MaskC, B_MaskC), where R_MaskC is the brightness value of the red channel of the third human body mask MaskC at position (a, b), G_MaskC is the brightness value of the green channel of the third human body mask MaskC at position (a, b), and B_MaskC is the brightness value of the blue channel of the third human body mask MaskC at position (a, b).

[0113] Similarly, if the facial mask is MaskD and the second pixel position is (a, b), then the fourth brightness values ​​of the three color channels of the facial mask MaskD at position (a, b) are determined to be (R_MaskD, G_MaskD, B_MaskD), where R_MaskD is the brightness value of the red channel of the facial mask MaskD at position (a, b), G_MaskD is the brightness value of the green channel of the facial mask MaskD at position (a, b), and B_MaskB is the brightness value of the blue channel of the facial mask MaskB at position (a, b).

[0114] S503. Based on the minimum brightness value among the third brightness value and the inverse brightness value of the fourth brightness value of each second pixel position in each color channel, determine the second target pixel value of each second pixel position in each color channel.

[0115] S504. Generate a target human body mask based on the second target pixel values ​​of each second pixel position in multiple color channels.

[0116] Specifically, the second target pixel value at the second pixel position (a, b) in each color channel is represented as follows:

[0117] R_MaskE=min(R_MaskC, 255-R_MaskD);

[0118] G_MaskE=min(G_MaskC, 255-G_MaskD);

[0119] B_MaskE=min(B_MaskC, 255-B_MaskD);

[0120] Where R_MaskE is the second target pixel value of the red channel at position (a, b), G_MaskE is the second target pixel value of the green channel at position (a, b), and B_MaskE is the second target pixel value of the blue channel at position (a, b).

[0121] Then the second target pixel values ​​of the three color channels of the target human body mask Mask E at position (a, b) are determined to be (R_MaskE, G_MaskE, B_MaskE).

[0122] The method provided in this application embodiment determines the third brightness value of each second pixel position in a third human body mask in multiple color channels, determines the fourth brightness value of each second pixel position in a face mask in multiple color channels, and determines the second target pixel value of each second pixel position in each color channel based on the minimum brightness value among the inverse brightness values ​​of the third and fourth brightness values ​​of each second pixel position in each color channel. A target human body mask is then generated based on the second target pixel values ​​of each second pixel position in multiple color channels. This allows the human body area (i.e., the area requiring costume change) and the background area of ​​the final costume change effect image to be determined using the target human body mask.

[0123] This application also provides another possible implementation of the image processing method. Figure 7 This is a sixth schematic flowchart illustrating an image processing method provided in an embodiment of this application. Figure 7 As shown, before processing the clothing feature extraction module in the preset intelligent drawing model according to the description text to obtain the target clothing features corresponding to the description text, the method also includes:

[0124] S601. Obtain multiple clothing sample images and a labeling file for each clothing sample image. The labeling file for each clothing sample image contains sample description text.

[0125] In this embodiment, multiple clothing sample images can be obtained from clothing photos taken in scenic spots or photo studios. The multiple clothing sample images are preprocessed to crop and scale their size to 512×768, and the human figure area accounts for more than 40%.

[0126] Each clothing sample image is labeled with a description of the corresponding clothing sample image in English, such as: long hair, white shirt, blur background, etc., so that each clothing sample image corresponds to a label file.

[0127] S602. Based on multiple clothing sample images and their corresponding labeling files, train the preset image style adapter model to generate a clothing feature extraction module.

[0128] The preset image style adapter model adopts the Low-Rank Adaptation of Large Language Models (LoRA) model to fine-tune the intelligent drawing model. The LoRA model is used to train multiple clothing sample images and corresponding labeled files to obtain the clothing feature extraction module, which can determine the target clothing features corresponding to the description text based on the description text.

[0129] The image processing apparatus and computer equipment provided in any of the above embodiments of this application will be explained below. The specific implementation process and the resulting technical effects are the same as those in the corresponding method embodiments. For the sake of brevity, parts not mentioned in this embodiment can be referred to the corresponding content in the method embodiments.

[0130] Figure 8 This is a schematic diagram of the functional modules of an image processing apparatus provided in an embodiment of this application. Figure 8 As shown, the image processing apparatus 200 includes:

[0131] The acquisition module 210 is used to acquire the image to be processed and the descriptive text.

[0132] The determination module 220 is used to determine the first human body mask, the face mask, and the human body pose map based on the image to be processed.

[0133] The processing module 230 is used to process the description text using the clothing feature extraction module in the preset intelligent drawing model to obtain the target clothing features corresponding to the description text.

[0134] The processing module 230 is also used to process the human posture diagram using the posture feature extraction module in the intelligent drawing model to obtain human posture features.

[0135] The processing module 230 is also used to process the target clothing features, human posture features and facial mask using the image generation module in the intelligent drawing model to generate a preliminary dressing effect image.

[0136] Optionally, the image processing apparatus 200 further includes:

[0137] The generation module is used to generate the target human body mask based on the initial costume effect image, the first human body mask, and the facial mask.

[0138] The processing module 230 is also used to process the target clothing features, human posture features and target human mask using the image generation module in the intelligent drawing model to generate a target dressing effect image.

[0139] Optionally, the generation module is also used to determine a second human body mask based on the initial costume effect image; to generate a third human body mask by mixing the second human body mask and the first human body mask; and to generate a target human body mask by mixing the third human body mask and the reverse image of the facial mask.

[0140] Optionally, the generation module is also used to perform a brightening layer blending based on the second human body mask and the first human body mask to generate a third human body mask.

[0141] Optionally, the generation module is further configured to: determine the first brightness value of each first pixel position in the first human body mask in multiple color channels; determine the second brightness value of each first pixel position in the second human body mask in multiple color channels; determine the first target pixel value of each first pixel position in each color channel based on the maximum brightness value among the first and second brightness values ​​of each first pixel position in each color channel; and generate a third human body mask based on the first target pixel value of each first pixel position in multiple color channels.

[0142] Optionally, the generation module is also used to generate a target human body mask by performing darkening layer blending based on the inverse image of the third human body mask and the face mask.

[0143] Optionally, the generation module is further configured to: determine the third brightness value of each second pixel position in the third human body mask in multiple color channels; determine the fourth brightness value of each second pixel position in the face mask in multiple color channels; determine the second target pixel value of each second pixel position in each color channel based on the minimum brightness value among the inverse brightness values ​​of the third and fourth brightness values ​​of each second pixel position in each color channel; and generate a target human body mask based on the second target pixel values ​​of each second pixel position in multiple color channels.

[0144] Optionally, the image processing apparatus 100 further includes:

[0145] The training module is used to acquire multiple clothing sample images and the labeling file of each clothing sample image. Each clothing sample image's labeling file contains sample description text. Based on the multiple clothing sample images and their corresponding labeling files, a preset image style adapter model is trained to generate a clothing feature extraction module.

[0146] The above-described device is used to execute the method provided in the foregoing embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.

[0147] These modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more microprocessors, or one or more Field Programmable Gate Arrays (FPGAs). Alternatively, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a system-on-a-chip (SOC).

[0148] Figure 9 This is a schematic diagram of a computer device provided in an embodiment of this application, which can be used for image processing. Figure 9 As shown, the computer device 300 includes: a processor 310, a storage medium 320, and a bus 330.

[0149] Storage medium 320 stores machine-readable instructions executable by processor 310. When the computer device is running, processor 310 communicates with storage medium 320 via bus 330, and processor 310 executes the machine-readable instructions to perform the steps of the above method embodiment. The specific implementation and technical effects are similar, and will not be described again here.

[0150] Optionally, this application also provides a storage medium 320, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the above-described method embodiments. The specific implementation and technical effects are similar, and will not be repeated here.

[0151] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0152] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0153] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0154] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0155] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An image processing method, characterized in that, include: Obtain the image to be processed and the descriptive text; Based on the image to be processed, determine the first human body mask, the face mask, and the human body pose map; Based on the description text, the clothing feature extraction module in the preset intelligent drawing model is used for processing to obtain the target clothing features corresponding to the description text; Based on the human posture diagram, the posture feature extraction module in the intelligent drawing model is used for processing to obtain human posture features; Based on the target clothing features, the human body posture features, and the facial mask, the image generation module in the intelligent drawing model is used to process the data and generate an initial dress-up effect image. The method further includes: Based on the initial costume change effect image, the first human body mask, and the facial mask, a target human body mask is generated; Based on the target clothing features, the human body posture features, and the target human body mask, the image generation module in the intelligent drawing model is used to process the image and generate a target costume change effect image; so that the target costume change effect image retains most of the background area and facial area of ​​the image to be processed. The step of generating a target human body mask based on the initial costume change effect image, the first human body mask, and the facial mask includes: Based on the initial costume change effect diagram, determine the second human body mask; A third human body mask is generated by mixing the second human body mask and the first human body mask. The target human body mask is generated by mixing the reverse images of the third human body mask and the face mask.

2. The method according to claim 1, characterized in that, The step of mixing the second human body mask and the first human body mask to generate a third human body mask includes: The third human body mask is generated by blending the second human body mask and the first human body mask with a brightening layer.

3. The method according to claim 2, characterized in that, The step of generating the third human body mask by blending the brightening layers based on the second human body mask and the first human body mask includes: Determine the first brightness value of each first pixel position in the first human body mask in multiple color channels; Determine the position of each first pixel in the second human body mask at the second brightness value of the plurality of color channels; Based on the maximum brightness value of each first pixel position in the first brightness value and the second brightness value of each color channel, the first target pixel value of each first pixel position in each color channel is determined respectively; The third human body mask is generated based on the first target pixel value of each first pixel position in the plurality of color channels.

4. The method according to claim 1, characterized in that, The step of generating the target human body mask by mixing the reverse images of the third human body mask and the facial mask includes: The target human body mask is generated by blending darkening layers based on the inverse images of the third human body mask and the face mask.

5. The method according to claim 4, characterized in that, The step of generating the target human body mask by performing darkening layer blending based on the inverse image of the third human body mask and the face mask includes: Determine the third brightness value of each second pixel position in the third human body mask in multiple color channels; The position of each second pixel in the facial mask is determined by the fourth luminance value of the plurality of color channels; Based on the minimum brightness value among the inverse brightness values ​​of the third and fourth brightness values ​​of each second pixel position in each color channel, the second target pixel value of each second pixel position in each color channel is determined respectively. The target human body mask is generated based on the second target pixel value of each second pixel position in the plurality of color channels.

6. The method according to claim 1, characterized in that, Before processing the description text using the clothing feature extraction module in a preset intelligent drawing model to obtain the target clothing features corresponding to the description text, the method further includes: Acquire multiple clothing sample images and a tagging file for each clothing sample image, wherein the tagging file for each clothing sample image contains sample description text; Based on the multiple clothing sample images and their corresponding labeling files, a preset image style adapter model is trained to generate the clothing feature extraction module.

7. An image processing apparatus, characterized in that, The device includes: The acquisition module is used to acquire the image to be processed and the descriptive text. The determination module is used to determine a first human body mask, a face mask, and a human body pose map based on the image to be processed. The processing module is used to process the description text using the clothing feature extraction module in the preset intelligent drawing model to obtain the target clothing features corresponding to the description text. The processing module is further configured to process the human posture diagram using the posture feature extraction module in the intelligent drawing model to obtain human posture features. The processing module is further configured to process the target clothing features, the human body posture features, and the facial mask using the image generation module in the intelligent drawing model to generate an initial dress-up effect image. The device further includes: The generation module is used to generate a target human body mask based on the initial costume effect image, the first human body mask, and the facial mask; The processing module is further configured to process the target clothing features, the human body posture features, and the target human body mask using the image generation module in the intelligent drawing model to generate a target dressing effect image; so that the target dressing effect image retains most of the background area and facial area of ​​the image to be processed; The generation module is further configured to determine a second human body mask based on the initial costume effect image; generate a third human body mask by mixing the second human body mask and the first human body mask; and generate the target human body mask by mixing the third human body mask and the reverse image of the face mask.

8. A computer device, characterized in that, include: The computer device includes a processor, a storage medium, and a bus. The storage medium stores program instructions executable by the processor. When the computer device is running, the processor communicates with the storage medium via the bus, and the processor executes the program instructions to perform the steps of the image processing method as described in any one of claims 1 to 6.