Image processing method and device, model training method and device, equipment and storage medium

By dividing images into regions and applying region-specific color mapping, the method enhances the precision and harmony of image style transfer, addressing the limitations of traditional methods.

CN120318062APending Publication Date: 2025-07-15SHUXING TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510536779.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In traditional image processing technology, the restoration and accuracy of image style conversion are low, making it difficult to achieve high-quality target style conversion.

Method used

By acquiring multiple area masks and color mapping relationships of the original image, using the weight prediction network and the mask prediction network, the area mask and color mapping relationships are generated respectively, and refined mapping and fusion processing are performed to generate the target image.

Benefits of technology

Improves the style conversion accuracy and harmony of the target image, ensuring natural transition and high-precision conversion of image color styles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318062A_ABST
    Figure CN120318062A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing method and device, a model training method and device, equipment and a storage medium. The image processing method comprises the following steps: acquiring an original image, wherein the original image comprises a plurality of areas; based on the image features of the original image, obtaining a plurality of area masks of the original image, the plurality of area masks respectively corresponding to a plurality of areas of the original image; based on the original image, obtaining a color mapping relation corresponding to the plurality of areas by using a weight prediction network; and obtaining a target image based on the original image, the region masks corresponding to the plurality of regions and the color mapping relationship. By adopting the method, the style conversion precision of the target image obtained by conversion can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technologies, and particularly to an image processing method, a model training method, a device, a device and a storage medium. Background Art

[0002] With the development of image processing technologies, various images with different styles can be generated through color transformation of images.

[0003] In traditional technologies, a neural network model is mainly used to predict the weights of several three-dimensional look-up tables (3D LUTs) based on an original image, and after weighted fusion of multiple preset 3D LUTs, the color of the original image is transformed to achieve the target style.

[0004] However, in traditional technologies, there is a problem that the restored degree and accuracy of the converted image to the target style are relatively low. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide an image processing method, a model training method, a device, a device and a storage medium that can improve the style conversion accuracy of a target image obtained by conversion.

[0006] In a first aspect, this application provides an image processing method, and the method includes:

[0007] Obtain an original image, where the original image includes multiple regions;

[0008] Based on the image features of the original image, obtain multiple region masks of the original image, and the multiple region masks respectively correspond to the multiple regions of the original image;

[0009] Based on the original image, use a weight prediction network to obtain the color mapping relationships corresponding to the multiple regions;

[0010] Based on the original image, the region masks corresponding to the multiple regions, and the color mapping relationships, obtain a target image.

[0011] In one of the embodiments, the step of using a weight prediction network to obtain the color mapping relationships corresponding to the multiple regions based on the original image includes:

[0012] Based on the original image, use the weight prediction network to obtain the weights corresponding to the multiple regions;

[0013] For each region, perform a fusion process on the three-dimensional look-up table corresponding to the region according to the weight corresponding to the region, so as to obtain the color mapping relationships corresponding to the multiple regions.

[0014] In one embodiment, obtaining the target image based on the original image, the region masks corresponding to the multiple regions, and the color mapping relationship includes:

[0015] For each region, performing a rendering process on the original image according to the color mapping relationship corresponding to the region to obtain the rendering maps corresponding to the multiple regions;

[0016] Based on the multiple rendering maps and the region masks corresponding to the multiple regions, obtaining the region images corresponding to the multiple regions;

[0017] Fusing the multiple region images to obtain the target image.

[0018] In one embodiment, the multiple regions include multiple target regions, and obtaining the region images corresponding to the multiple regions based on the multiple rendering maps and the region masks corresponding to the multiple regions includes:

[0019] For each target region, performing a fusion process on the rendering map of the target region and the region mask corresponding to the target region to obtain the target rendering maps corresponding to the multiple target regions;

[0020] For each target region, performing a fusion process on the target rendering map of the target region and the inverse image of the mask of the original image to obtain the region images corresponding to the multiple target regions.

[0021] In one embodiment, the multiple regions further include a full - map region, and the method further includes:

[0022] Performing a fusion process on the rendering map of the full - map region and the region mask of the full - map region to obtain the target rendering map corresponding to the full - map region;

[0023] Performing a fusion process on the target rendering map corresponding to the full - map region and the mask of the original image to obtain the region image corresponding to the full - map region.

[0024] In one embodiment, fusing the multiple region images to obtain the target image includes:

[0025] Performing a fusion process on the region images corresponding to the multiple target regions and the region image corresponding to the full - map region to obtain the target image.

[0026] In one embodiment, the multiple target regions include at least one of the following:

[0027] Multiple brightness regions with different brightness levels for each brightness region;

[0028] Multiple color regions, with different colors for each color region;

[0029] Multiple main body regions, with different main bodies corresponding to each main body region;

[0030] Multiple local regions, with different importance levels of each local region in the original image.

[0031] In one embodiment, the image features include at least one of luminance feature, color feature, shape feature, texture feature, semantic feature, embedding feature, and local feature.

[0032] In one embodiment, the target image is a style image with a target style corresponding to the original image,

[0033] At least one of the weight prediction network and the mask prediction network is trained based on a sample image and a sample style image with a target style corresponding to the sample image, and the mask prediction network is used to generate a region mask corresponding to at least one region for the original image based on the image features of the original image.

[0034] In a second aspect, the present application provides a model training method, and the method includes:

[0035] Obtain a sample image and a sample style image with a target style corresponding to the sample image, where the sample image includes multiple regions;

[0036] Based on the image features of the sample image, obtain multiple region masks of the sample image, and the multiple region masks respectively correspond to the multiple regions of the sample image;

[0037] Based on the sample image, use an initial weight prediction network to obtain a color mapping relationship corresponding to the multiple regions;

[0038] Based on the sample image, the region masks corresponding to the multiple regions, and the color mapping relationship, obtain a transformed image;

[0039] Train at least one of the initial weight prediction network and the initial mask prediction network according to the transformed image and the sample style image; the initial mask prediction network is used to generate a region mask corresponding to at least one region for the sample image based on the image features of the sample image.

[0040] In a third aspect, the present application provides an image processing device, and the device includes:

[0041] A first acquisition module, configured to acquire an original image, where the original image includes multiple regions;

[0042] A second acquisition module, configured to acquire a plurality of region masks of the original image based on the image features of the original image, where the plurality of region masks respectively correspond to a plurality of regions of the original image;

[0043] A third acquisition module, configured to obtain a color mapping relationship corresponding to the plurality of regions based on the original image by using a weight prediction network;

[0044] A processing module, configured to obtain a target image based on the original image, the region masks corresponding to the plurality of regions, and the color mapping relationship; the target image is a style conversion image of the original image.

[0045] In a fourth aspect, the present application provides a model training device, where the device includes:

[0046] A first acquisition module, configured to acquire a sample image and a sample style image with a target style corresponding to the sample image, where the sample image includes a plurality of regions;

[0047] A second acquisition module, configured to acquire a plurality of region masks of the sample image based on the image features of the sample image, where the plurality of region masks respectively correspond to a plurality of regions of the sample image;

[0048] A third acquisition module, configured to obtain a color mapping relationship corresponding to the plurality of regions based on the sample image by using an initial weight prediction network;

[0049] A fourth acquisition module, configured to obtain a conversion image based on the sample image, the region masks corresponding to the plurality of regions, and the color mapping relationship;

[0050] A training module, configured to train at least one of the initial weight prediction network and the initial mask prediction network according to the conversion image and the sample style image; the initial mask prediction network is configured to generate a region mask corresponding to at least one region for the sample image based on the image features of the sample image.

[0051] In a fifth aspect, the present application provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the methods described in the first aspect and the second aspect are implemented.

[0052] In a sixth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the methods described in the first aspect and the second aspect are implemented.

[0053] In a seventh aspect, the present application provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the methods described in the first aspect and the second aspect.

[0054] For the above image processing method, model training method, device, equipment, and storage medium, the original image includes multiple regions. Based on the image features of the original image, region masks corresponding to the multiple regions can be obtained. Based on the original image, a color mapping relationship corresponding to the multiple regions can be obtained by using a weight prediction network. Thus, based on the original image, the region masks corresponding to the multiple regions, and the color mapping relationship, a target image can be obtained. In this way, the refined mapping of the colors of different regions of the original image can be achieved through the region masks corresponding to the multiple regions and the color mapping relationship, ensuring the natural transition of the image color style of the original image to the target color style, and improving the accuracy of the target color style and the harmony of the target color style in the converted target image. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.

[0056] Figure 1 It is an application environment diagram of the image processing method in an embodiment;

[0057] Figure 2 It is a schematic flowchart of the image processing method in an embodiment;

[0058] Figure 3 It is a schematic flowchart of the image processing method in another embodiment;

[0059] Figure 4 It is a schematic flowchart of the image processing method in another embodiment;

[0060] Figure 5 It is a schematic flowchart of the model training method in an embodiment;

[0061] Figure 6 It is a schematic diagram of the process of image processing in an embodiment;

[0062] Figure 7 It is a schematic flowchart of the complete image processing method in an embodiment;

[0063] Figure 8 It is a schematic flowchart of the complete model training method in an embodiment;

[0064] Figure 9 is a structural block diagram of an image processing apparatus in an embodiment;

[0065] Figure 10 is a structural block diagram of a model training apparatus in an embodiment;

[0066] Figure 11 is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0067] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0068] With the development of computer technology, a variety of social application programs (Application, abbreviated as APP) have emerged. Publishers can record their lives or share the usage experiences of items in images or videos by publishing images or videos and attaching text information in the APP. Usually, the publisher can first take a picture or video and then enter the publishing page to publish the taken picture or video, or select the already taken content and enter the publishing page to publish the content. However, the style of the image can be different in different scenarios. When publishing an image or video, in order to make the style of the image or video more close to the shooting scene, the publisher may need to perform a conversion process on the style of the image or video to convert the style of the image or video into the desired target image style.

[0069] For the first content publishing method, that is, the publishing method of taking a picture first and then entering the publishing page. Usually, after the shooting is completed, the preview editing page can be entered. The taken picture can be displayed on the preview editing page, and various editing operations can be performed on the taken picture. For example, image style conversion and other processes can be performed on the taken picture, and then enter the publishing page to publish the video or image after shooting and editing, generating a published content.

[0070] For the second content publishing method, that is, the publishing method of selecting the already taken content and entering the publishing page, one or more of the already taken videos and pictures can be selected from the picture library. After the selection is completed, the preview editing page can be entered first, and various editing operations can be performed on the preview editing page. For example, image style conversion and other processes can be performed. After the editing is completed, enter the publishing page to publish the content, generating a published content.

[0071] It should be noted that the image processing method provided by the embodiments of the present application can be applied to the scenario of image style conversion processing of the image to be published in the above two different content publishing methods, or can also be applied to other image processing scenarios. The application scenarios of the embodiments of the present application are not limited herein, and the image processing method provided by the embodiments of the present application can be applied to any image processing scenario.

[0072] The image processing method provided by the embodiments of the present application can be applied to, for example Figure 1 the terminal 102 as shown. The terminal 102 can respond to an image processing request and convert the image style of the image to be processed into a target image style through an image processing model. Among them, the image processing model is a pre-trained neural network model. Optionally, the image processing model can be trained in the terminal 102, or the image processing model can also be trained in the server. After the image processing model is trained, the trained image processing model is sent to the terminal 102 so that the terminal 102 can use the trained image processing model to perform image style conversion processing. Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart TVs, smart in-vehicle devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc.

[0073] In an exemplary embodiment, as Figure 2 shown, an image processing method is provided. In this embodiment, the method is exemplified by being applied to a terminal. It can be understood that the method can also be applied to a server and can also be applied to a system including a terminal and a server and implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0074] S201, obtain an original image, where the original image includes multiple regions.

[0075] Among them, the original image can be an image to be published in a social APP, or can be an image saved in a gallery, etc. The original image can be the captured image in the first content publishing method described above. That is to say, the application scenario in this embodiment can be a scenario of performing various editing operations on the image captured in real time. Or, the original image can be the captured image in the second content publishing method described above. That is to say, the application scenario in this embodiment can be a scenario of performing various editing operations on the already captured image. Exemplarily, the original image can be any one of landscape images, portrait images, architectural images, food images.

[0076] Exemplarily, multiple regions of the original image can include at least one of the full-image region, multiple brightness regions, multiple color regions, multiple subject regions, and multiple local regions. Among them, the brightness of each brightness region in the multiple brightness regions is different; the color of each color region in the multiple color regions is different; the subject corresponding to each subject region in the multiple subject regions is different; the importance of each local region in the multiple local regions in the original image is different. Exemplarily, the original image can be divided according to the image features of the original image to determine multiple regions of the original image. For example, multiple brightness regions of the original image can be determined according to the brightness features of the original image; or multiple color regions of the original image can be determined according to the color features of the original image; or multiple brightness regions and multiple color regions of the original image can be determined according to the brightness features and color features of the original image.

[0077] S202. Based on the image features of the original image, obtain multiple region masks of the original image, and the multiple region masks respectively correspond to multiple regions of the original image.

[0078] Optionally, the image features of the original image can include at least one of brightness features, color features, shape features, texture features, semantic features, embedding features, local features, and global features. Correspondingly, the multiple region masks of the original image can include at least one of multiple brightness masks, multiple color masks, multiple shape masks, skin-color masks, non-skin-color masks, multiple subject masks, and region masks of the full-image region.

[0079] Exemplarily, based on the luminance feature of the original image, a first luminance mask and a second luminance mask of the original image can be obtained; the first luminance mask and the second luminance mask respectively correspond to a first luminance region and a second luminance region in the original image, wherein the luminance of the first luminance region is different from that of the second luminance region. Based on the color feature of the original image, a first color mask and a second color mask of the original image can be obtained, and the first color mask and the second color mask respectively correspond to a first color region and a second color region in the original image, wherein the color of the first color region is different from that of the second color region. Based on the semantic feature of the original image, a first subject mask and a second subject mask of the original image can be obtained, and the first subject mask and the second subject mask respectively correspond to a first subject region and a second subject region in the original image, and the subject in the first subject region is different from the subject in the second subject region.

[0080] It should be noted that the mask can be a binary image or matrix for specifying the region of interest in the image. In this embodiment, based on the image feature of the original image, a binary image or matrix of different regions in the original image can be generated according to the pixel values of the original image, or the original image and the image feature of the original image can be input into a mask prediction network to obtain multiple region masks of the original image.

[0081] Exemplarily, taking the feature of the original image as the luminance feature as an example, based on the luminance values of each pixel of the original image in the luminance dimension, the highlight region and the shadow region of the original image can be determined, and then the mask images of the highlight region and the shadow region are determined as multiple region masks of the original image. It can be understood that the highlight region of the original image refers to the region with high luminance in the original image, and the higher the luminance, the larger the luminance value of the pixel. Assuming that in this scenario, the multiple region masks of the original image include a highlight mask and a shadow mask, then the mask corresponding to the region with the luminance value of L in the original image can be determined as the highlight mask, and the mask corresponding to the region with the luminance value of 1 - L in the original image can be determined as the shadow mask.

[0082] As an example, the detailed process of obtaining the highlight mask and the shadow mask of the original image with the image feature being the luminance feature is explained below, as Figure 3 shown, the above S202 includes:

[0083] S301, converting the original image to the Lab color space.

[0084] Among them, the Lab color space is a color space. The L component in the Lab color space is used to represent the brightness of pixels, and its value range is [0, 100], indicating from pure black to pure white; the a component represents the range from red to green, and its value range is [127, -128]; the b component represents the range from yellow to blue, and its value range is [127, -128].

[0085] It can be understood that the original color space of the original image in this embodiment is the RGB color space. In this embodiment, the RGB color values of the original image can be converted into XYZ color values, and then the XYZ color values can be converted into the values of the three channels of the Lab color space, so as to realize the conversion of the original image to the Lab color space.

[0086] S302: Based on the brightness values of each pixel of the original image in the Lab color space, perform masking processing on the original image to obtain a highlight mask and a shadow mask.

[0087] As mentioned in the above embodiment, the L component in the Lab color space is used to represent the brightness of pixels. In this embodiment, based on the brightness values of each pixel of the original image in the Lab color space, the highlight area and the shadow area of the original image can be determined, and then a highlight mask and a shadow mask can be generated. Optionally, in this embodiment, a highlight mask and a shadow mask of the original image can be generated through a brightness mask generator.

[0088] In this embodiment, the process of converting the original image to the Lab color space is relatively simple, and the original image can be quickly converted to the Lab color space, so that based on the brightness values of each pixel of the original image in the Lab color space, masking processing can be quickly performed on the original image, and the efficiency of obtaining the highlight mask and the shadow mask of the original image can be improved.

[0089] S203: Based on the original image, use a weight prediction network to obtain the color mapping relationships corresponding to multiple regions.

[0090] First of all, it should be noted that the color mapping relationship in this embodiment can be a Three-Dimensional Look-Up Table (3D LUT). The 3D LUT is a three-dimensional matrix data structure used to store and quickly find the color mapping relationship, realize the accurate conversion from the input color space to the output color space, so as to achieve the purpose of color adjustment and correction. Exemplarily, the weight prediction network in this embodiment can be used to predict the weights corresponding to multiple regions of the original image, and then based on the weights corresponding to multiple regions and the 3D LUT, the color mapping relationships corresponding to multiple regions can be obtained.

[0091] Optionally, in this embodiment, based on the original image, a weight prediction network can be used to obtain the weights corresponding to multiple regions of the original image. Then, for each region of the original image, the 3D LUT corresponding to that region is fused according to the weight corresponding to each region to obtain the color mapping relationship corresponding to each region.

[0092] Exemplarily, taking the multiple regions of the original image including a highlight region and a shadow region as an example, in this embodiment, the original image can be input into the weight prediction network, so that the weight prediction network obtains the weight corresponding to the highlight region and the weight corresponding to the shadow region based on the image features of the original image; then, combining the weight corresponding to the highlight region and the 3D LUT, the color mapping relationship of the highlight region is obtained, and combining the weight corresponding to the shadow region and the 3D LUT, the color mapping relationship of the shadow region is obtained.

[0093] S204. Based on the original image, the region masks corresponding to multiple regions, and the color mapping relationship, obtain the target image.

[0094] Optionally, in this embodiment, the original image, the region masks corresponding to multiple regions of the original image, and the color mapping relationship can be input into an image processing model. Through the image processing model, combining the region masks corresponding to multiple regions of the original image and the color mapping relationship, multiple regions of the original image are rendered, realizing the conversion of the image style of the original image, and converting the original image into a target image with a target style.

[0095] In the above image processing method, the original image includes multiple regions. Based on the image features of the original image, region masks corresponding to multiple regions can be obtained. Based on the original image, a weight prediction network can be used to obtain the color mapping relationships corresponding to multiple regions. Thus, based on the original image, the region masks corresponding to multiple regions, and the color mapping relationship, the target image can be obtained. In this way, the fine mapping of the colors of different regions of the original image can be realized through the region masks corresponding to multiple regions and the color mapping relationship, ensuring the natural transition of the image color style of the original image to the target color style, and improving the accuracy of the target color style and the harmony of the target color style in the converted target image.

[0096] In some scenarios, for each region of the original image, the image style can be converted, and then the images after style conversion of each region are subjected to image fusion processing to obtain the target image after style conversion of the original image. In one embodiment, as Figure 4 shown, the above S204 includes:

[0097] S401. For each region, render the original image according to the color mapping relationship corresponding to the region to obtain the renderings corresponding to multiple regions.

[0098] Optionally, in this embodiment, for each region of the original image, the original image can be rendered according to the color mapping relationship and color mapping relationship renderer corresponding to each region to obtain a rendered image corresponding to each region.

[0099] Exemplarily, the multiple regions of the original image can include multiple target regions, and the multiple target regions can include at least one of multiple brightness regions, multiple color regions, multiple main body regions, and multiple local regions.

[0100] In this embodiment, if the multiple regions of the original image include the above multiple target regions, the original image can be rendered according to the color mapping relationship corresponding to each target region to obtain a rendered image corresponding to each target region. Exemplarily, taking the multiple target regions of the original image as the highlight region and the shadow region as an example, the original image can be rendered according to the color mapping relationship (3D LUT) corresponding to the shadow region to obtain a rendered image corresponding to the shadow region; the original image can be rendered according to the color mapping relationship (3D LUT) corresponding to the highlight region to obtain a rendered image corresponding to the highlight region.

[0101] As another example, in addition to the above multiple target regions, the multiple regions of the original image can also include the full-image region of the original image. Then, in this embodiment, the original image can also be rendered according to the color mapping relationship corresponding to the full-image region to obtain a rendered image corresponding to the full-image region.

[0102] S402. Based on the multiple rendered images and the region masks corresponding to the multiple regions, obtain the region images corresponding to the multiple regions.

[0103] Exemplarily, in this embodiment, if the multiple regions of the original image include the above multiple target regions, S402 can include: for each target region, perform a fusion process on the rendered image of the target region and the region mask corresponding to the target region to obtain a target rendered image corresponding to each target region. Then, for each target region, perform a fusion process on the target rendered image corresponding to each target region and the inverse image of the mask of the original image to obtain a region image corresponding to each target region. It should be noted that if the pixel value of each pixel of the mask of the original image is a, then the pixel value of each pixel of the inverse image of the mask of the original image is 1 - a.

[0104] Further, as another example, if multiple regions of the original image further include a full-image region, where the full-image region of the original image can be understood as the original image itself, then in this embodiment, S402 further includes: fusing the rendering image of the full-image region and the region mask of the full-image region to obtain a target rendering image corresponding to the full-image region, and then fusing the target rendering image corresponding to the full-image region and the mask of the original image to obtain a region image corresponding to the full-image region. It should be noted that the mask of the original image mentioned in this embodiment can also be regarded as a kind of region mask, that is, the mask corresponding to the full-image region.

[0105] S403, fuse multiple region images to obtain a target image.

[0106] In this embodiment, if multiple regions of the original image include the above-mentioned multiple target regions, then S403 may include: fusing the region images corresponding to the multiple target regions to obtain a target image of the original image. If multiple regions of the original image further include a full-image region in addition to the above-mentioned multiple target regions, then S403 may include: fusing the region images corresponding to the multiple target regions and the region image corresponding to the full-image region to obtain a target image of the original image.

[0107] In this embodiment, since the rendering process is performed on each region of the original image according to the color mapping relationship corresponding to each region, for each region of the original image, a rendering process with higher precision can be performed, ensuring the precision of the multiple rendering images corresponding to the multiple regions of the original image, and thus ensuring the precision of the region images corresponding to the multiple regions of the original image based on the rendering images and region masks corresponding to the multiple regions of the original image. In this way, after fusing the multiple region images, the precision of the target image corresponding to the original image is ensured.

[0108] In some scenarios, based on the image features of the original image, a region mask corresponding to at least one region can be generated for the original image by using a mask prediction network. The above target image can be a style image with a target style corresponding to the original image. As an optional implementation manner, at least one of the above weight prediction network and mask prediction network is trained according to a sample image and a sample style image with a target style corresponding to the sample image.

[0109] Optionally, in this embodiment, the original image can be input into an image processing model. The image processing model can include at least one of a weight prediction network and a mask prediction network, enabling the image processing model to use at least one of the weight prediction network and the mask prediction network to execute the image processing method described in the above embodiment to obtain a corresponding output image. Then, according to the sample style image with the target style corresponding to the sample image and the output image, calculate the value of the loss function of the image processing model, and perform gradient backpropagation through the value of the loss function to adjust the adjustable parameters in the image processing model, and train the image processing model, that is, train at least one of the weight prediction network and the mask prediction network.

[0110] In addition, in this embodiment, the image processing model further includes a three-dimensional lookup table corresponding to multiple regions of the original image randomly initialized in advance. During the training process of the image processing model, when performing gradient backpropagation based on the value of the loss function of the image processing model, the trainable parameters of the three-dimensional lookup table corresponding to multiple regions of the original image can be updated and adjusted, so that the image style of the output image is closer to the target style.

[0111] In this embodiment, the region masks corresponding to multiple regions of the original image can be obtained by using the mask prediction network based on the image features of the original image. At least one of the weight prediction network and the mask prediction network can be trained according to the sample image and the sample style image with the target style corresponding to the sample image. In this way, the color mapping relationships corresponding to multiple regions of the original image and the region masks corresponding to multiple regions can be quickly obtained by using the trained weight prediction network and mask prediction network. Furthermore, based on the original image, the region masks corresponding to multiple regions of the original image, and the color mapping relationships, the target image can be quickly obtained, ensuring the efficiency of obtaining the target image after style conversion corresponding to the original image.

[0112] In this embodiment, the training processes of at least the weight prediction network and the mask prediction network used in the process of obtaining the style image with the target style corresponding to the original image will be explained. In an exemplary embodiment, as Figure 5 shown, a model training method is provided. In this embodiment, this method is exemplified by being applied to a terminal. In this embodiment, the method includes the following steps:

[0113] S501, obtain a sample image and a sample style image with the target style corresponding to the sample image, where the sample image includes multiple regions.

[0114] Among them, the sample image can be an image to be published in a social APP, or can be an image saved in a gallery. Exemplarily, the sample image can be any one of landscape images, portrait images, architectural images, and food images. The sample style image can be an image obtained by performing target image style conversion on the sample image, where the target image style can be the image style corresponding to a pre-determined filter. For example, the target image style can be a dim style, a high-brightness style, a soft-light style, a vivid color style, an oil painting style, a retro style, etc.

[0115] Optionally, in this embodiment, after determining the target style corresponding to the sample image, the image style of the sample image can be adjusted by at least one of the color, saturation, and brightness of the sample image through an image processor to obtain the sample style image corresponding to the sample image. Alternatively, an image pair pre-processed by an image processor can be obtained, the original image in the image pair can be used as the sample image, and the processed image in the image pair can be used as the sample style image.

[0116] Exemplarily, the multiple regions included in the sample image can include at least one of a full-image region, multiple brightness regions, multiple color regions, multiple subject regions, and multiple local regions. Among them, the brightnesses of the multiple brightness regions are different; the colors of the multiple color regions are different; the subjects corresponding to the multiple subject regions are different; the importance degrees of the multiple local regions in the original image are different.

[0117] S502, based on the image features of the sample image, obtain multiple region masks of the sample image, and the multiple region masks respectively correspond to the multiple regions of the sample image.

[0118] Optionally, the image features of the sample image can include at least one of brightness features, color features, shape features, texture features, semantic features, embedding features, local features, and global features. Correspondingly, the multiple region masks of the sample image can include at least one of multiple brightness masks, multiple color masks, multiple shape masks, skin color masks, non-skin color masks, multiple subject masks, and region masks of the full-image region.

[0119] Exemplarily, based on the luminance features of the sample image, a first luminance mask and a second luminance mask of the sample image can be obtained; the first luminance mask and the second luminance mask respectively correspond to a first luminance region and a second luminance region in the sample image, wherein the luminance of the first luminance region is different from that of the second luminance region. Based on the color features of the sample image, a first color mask and a second color mask of the sample image can be obtained, the first color mask and the second color mask respectively correspond to a first color region and a second color region in the sample image, wherein the color of the first color region is different from that of the second color region. Based on the semantic features of the sample image, a first subject mask and a second subject mask of the sample image can be obtained, the first subject mask and the second subject mask respectively correspond to a first subject region and a second subject region in the sample image, and the subject of the first subject region is different from that of the second subject region.

[0120] Optionally, in this embodiment, based on the image features of the sample image, an initial mask prediction network can be used to generate region masks corresponding to at least one region of the sample image.

[0121] S503, based on the sample image, use the initial weight prediction network to obtain color mapping relationships corresponding to multiple regions.

[0122] First of all, it should be noted that the color mapping relationship in this embodiment can be a three-dimensional look-up table (3D LUT). The 3D LUT is a three-dimensional matrix data structure used to store and quickly look up the color mapping relationship, realize the accurate conversion from the input color space to the output color space, so as to achieve the purpose of color adjustment and correction. Exemplarily, the initial weight prediction network in this embodiment can be used to predict the weights corresponding to multiple regions of the sample image, and then based on the weights corresponding to multiple regions of the sample image and the 3D LUT, color mapping relationships corresponding to multiple regions can be obtained.

[0123] Optionally, in this embodiment, based on the sample image, the initial weight prediction network can be used to obtain the weights corresponding to multiple regions of the original image, and then for each region of the sample image, the 3D LUT corresponding to each region is fused according to the weight corresponding to each region to obtain the color mapping relationship corresponding to each region.

[0124] Exemplarily, taking the example that multiple regions of the sample image include a highlight region and a shadow region, in this embodiment, the sample image can be input into the initial weight prediction network, so that the initial weight prediction network obtains the weight corresponding to the highlight region and the weight corresponding to the shadow region based on the image features of the sample image; then, combining the weight corresponding to the highlight region and the 3DLUT, the color mapping relationship of the highlight region is obtained, and combining the weight corresponding to the shadow region and the 3D LUT, the color mapping relationship of the shadow region is obtained.

[0125] S504. Obtain a transformed image based on the sample image, the region masks corresponding to multiple regions, and the color mapping relationship.

[0126] Optionally, in this embodiment, for each region of the sample image, the original image can be rendered according to the color mapping relationship corresponding to each region to obtain the rendered images corresponding to multiple regions, and for the full-image region of the sample image, the sample image can be rendered according to the color mapping relationship corresponding to the full-image region to obtain the rendered image corresponding to the full-image region of the sample image.

[0127] Then, the rendered image of each region and the corresponding region mask are fused to obtain the target rendered image of each region, and then the target rendered image of each region and the inverse image of the mask of the sample image are fused to obtain the sample region image corresponding to each region; the rendered image corresponding to the full-image region of the sample image and the region mask of the full-image region are fused to obtain the target rendered image corresponding to the full-image region, and then the target rendered image corresponding to the full-image region and the mask of the sample image are fused to obtain the sample region image corresponding to the full-image region. After that, the sample region images corresponding to each region and the sample region image corresponding to the full-image region are fused to obtain the transformed image of the sample image.

[0128] S505. Train at least one of the initial weight prediction network and the initial mask prediction network according to the transformed image and the sample style image; the initial mask prediction network is used to generate region masks corresponding to at least one region for the sample image based on the image features of the sample image.

[0129] Optionally, in this embodiment, the value of the loss function can be calculated according to the sample style image with the target style corresponding to the sample image and the transformed image, and the adjustable parameters of at least one of the initial weight prediction network and the initial mask prediction network can be adjusted by backpropagating the gradient through the value of the loss function to train at least one of the initial weight prediction network and the initial mask prediction network.

[0130] In the above model training method, the sample image includes multiple regions. Based on the image features of the sample image, region masks corresponding to the multiple regions of the sample image can be obtained. Then, based on the sample image, a color mapping relationship corresponding to the multiple regions is obtained using the initial weight prediction network. Further, based on the sample image, the region masks corresponding to the multiple regions of the sample image, and the color mapping relationship, the fine-grained mapping of the colors of different regions of the sample image is realized, ensuring the natural transition of the image color style of the sample image to the target style, improving the accuracy of the target style and the harmony of the target style in the obtained converted image. Thus, at least one of the initial weight prediction network and the initial mask prediction network can be accurately trained according to the converted image and the sample style image, so that the at least one trained initial weight prediction network and initial mask prediction network can perform fine-grained mapping of the colors of different regions of the original image, ensuring the natural transition of the style of the original image, and improving the accuracy of the target style and the harmony of the target style of the obtained target image.

[0131] Exemplarily, taking the multiple regions of the original image including a shadow region, a highlight region, and a full-image region as an example, the image processing method provided by the present disclosure will be described in detail. Figure 6 A schematic diagram showing its processing process is as follows Figure 6As shown, the method includes: based on the image features and mask prediction network of the original image, obtaining the highlight mask, shadow mask and full-image mask of the original image; based on the original image, using the weight prediction network to obtain the shadow area weight, highlight area weight and full-image area weight, fusing the weight corresponding to the shadow area with the 3D LUT to obtain the 3D LUT of the shadow area, fusing the weight corresponding to the highlight area with the 3D LUT to obtain the 3D LUT of the highlight area, and fusing the weight corresponding to the full-image area with the 3D LUT to obtain the 3D LUT of the full-image area; using the 3D LUT renderer to render the original image and the 3D LUT of the shadow area respectively to obtain the shadow rendering image, rendering the original image and the 3D LUT of the highlight area to obtain the highlight rendering image, and rendering the original image and the 3D LUT of the full-image area to obtain the rendering image of the full-image area; then, fusing the shadow rendering image and the shadow mask to obtain the target rendering image of the shadow area, fusing the highlight rendering image and the highlight mask to obtain the target rendering image of the highlight area, and fusing the rendering image of the full-image area and the full-image mask to obtain the target rendering image of the full-image area; then fusing the target rendering image of the shadow area and the inverse image of the mask of the original image to obtain the shadow area image, fusing the target rendering image of the highlight area and the inverse image of the mask of the original image to obtain the highlight area image, and fusing the target rendering image of the full-image area and the mask of the original image to obtain the global area image; afterwards, fusing the shadow area image, the highlight area image and the global area image to obtain the target image.

[0132] For the convenience of those skilled in the art, the following provides a detailed introduction to the image processing method provided by the present disclosure, as Figure 7 shown, the method may include:

[0133] S1, obtaining an original image, where the original image includes multiple regions.

[0134] S2, based on the image features of the original image, using a mask prediction network to obtain multiple region masks of the original image and the region mask of the full-image region, where the multiple region masks respectively correspond to multiple regions of the original image; wherein, the image features include at least one of brightness feature, color feature, shape feature, texture feature, semantic feature, embedding feature, and local feature.

[0135] S3, based on the original image, using a weight prediction network to obtain the weights corresponding to the above multiple regions and the weight corresponding to the full-image region.

[0136] S4, for each region, fusing the three-dimensional lookup table corresponding to the region according to the weight corresponding to the region to obtain the color mapping relationships corresponding to the multiple regions.

[0137] S5. Perform a fusion process on the three-dimensional lookup table corresponding to the full map region according to the weight corresponding to the full map region to obtain the color mapping relationship corresponding to the full map region.

[0138] S6. For each region, perform a rendering process on the original image according to the color mapping relationship corresponding to each region to obtain the rendered images corresponding to multiple regions.

[0139] S7. Perform a fusion process on the rendered image of each region and the corresponding region mask to obtain the target rendered image corresponding to each region.

[0140] S8. Perform a fusion process on the target rendered image of each region and the inverse image of the mask of the original image to obtain the region image corresponding to each region.

[0141] S9. For the full map region, perform a rendering process on the original image according to the color mapping relationship corresponding to the full map region to obtain the rendered image corresponding to the full map region.

[0142] S10. Perform a fusion process on the rendered image corresponding to the full map region and the region mask of the full map region to obtain the target rendered image corresponding to the full map region.

[0143] S11. Perform a fusion process on the target rendered image corresponding to the full map region and the mask image of the original image to obtain the region image corresponding to the full map region.

[0144] S12. Perform a fusion process on the region image corresponding to each region and the region image corresponding to the full map region to obtain the target image with the target style corresponding to the original image.

[0145] For the convenience of those skilled in the art to understand, the following details the training process of at least one of the mask prediction network, weight prediction network, three-dimensional lookup tables corresponding to multiple regions and / or the full map region in the image processing method provided by the present disclosure, as Figure 8 shown, the method may include:

[0146] T1. Obtain a sample image and a sample style image with the target style corresponding to the sample image, and the sample image includes multiple regions.

[0147] T2. Based on the image features of the sample image, use the initial mask prediction network to obtain multiple region masks of the sample image and the region mask of the full map region, and the multiple region masks respectively correspond to the multiple regions of the sample image; wherein, the image features include at least one of brightness feature, color feature, shape feature, texture feature, semantic feature, embedding feature, and local feature.

[0148] T3. Based on the sample image, using the weight prediction network, obtain the weights corresponding to multiple regions of the sample image and the weight corresponding to the full-image region.

[0149] T4. For each region of the sample image, perform a fusion process on the three-dimensional lookup table corresponding to the region according to the weight corresponding to the region to obtain the color mapping relationships corresponding to multiple regions.

[0150] T5. Perform a fusion process on the three-dimensional lookup table corresponding to the full-image region according to the weight corresponding to the full-image region to obtain the color mapping relationship corresponding to the full-image region.

[0151] T6. For each region, perform a rendering process on the sample image according to the color mapping relationship corresponding to each region to obtain the sample renderings corresponding to multiple regions.

[0152] T7. Perform a fusion process on the sample rendering of each region and the corresponding region mask to obtain the final rendering corresponding to each region.

[0153] T8. Perform a fusion process on the final rendering of each region and the inverse image of the mask of the sample image to obtain the sample region image corresponding to each region.

[0154] T9. For the full-image region, perform a rendering process on the sample image according to the color mapping relationship corresponding to the full-image region to obtain the sample rendering corresponding to the full-image region.

[0155] T10. Perform a fusion process on the sample rendering corresponding to the full-image region and the region mask of the full-image region to obtain the final rendering corresponding to the full-image region.

[0156] T11. Perform a fusion process on the final rendering corresponding to the full-image region and the mask image of the sample image to obtain the sample region image corresponding to the full-image region.

[0157] T12. Perform a fusion process on the sample region image corresponding to each region and the sample region image corresponding to the full-image region to obtain the converted image with the target style corresponding to the sample image.

[0158] T13. According to the sample style image and the above-mentioned converted image, obtain the value of the loss function.

[0159] T14. According to the value of the loss function, adjust the trainable parameters of at least one of the initial mask prediction network, the initial weight prediction network, and the three-dimensional lookup tables corresponding to multiple regions and / or the full-image region, and train at least one of the initial mask prediction network, the initial weight prediction network, and the three-dimensional lookup tables corresponding to multiple regions and / or the full-image region.

[0160] It should be noted that for the descriptions in the above steps, reference can be made to the relevant descriptions in the above embodiments, and the effects are similar. Therefore, this embodiment will not be elaborated herein.

[0161] It should be understood that although the steps in the flowcharts involved in the above embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0162] Based on the same inventive concept, an embodiment of the present application also provides an image processing apparatus for implementing the above-mentioned image processing method. The solution provided by this apparatus to solve problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the following image processing apparatus can refer to the limitations on the image processing method in the above text, and will not be elaborated herein.

[0163] In an exemplary embodiment, as Figure 9 shown, an image processing apparatus is provided, including: a first acquisition module 10, a second acquisition module 11, a third acquisition module 12, and a processing module 13, where:

[0164] The first acquisition module 10 is configured to acquire an original image, and the original image includes multiple regions.

[0165] The second acquisition module 11 is configured to acquire multiple region masks of the original image based on the image features of the original image, and the multiple region masks respectively correspond to the multiple regions of the original image.

[0166] The third acquisition module 12 is configured to obtain a color mapping relationship corresponding to multiple regions based on the original image by using a weight prediction network.

[0167] The processing module 13 is configured to obtain a target image based on the original image, the region masks corresponding to the multiple regions, and the color mapping relationship.

[0168] The image processing apparatus provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effect are similar, and will not be elaborated herein.

[0169] Based on the above embodiments, optionally, the above third acquisition module 12 includes: a first acquisition unit and a second acquisition unit, where:

[0170] The first acquisition unit is configured to obtain weights corresponding to multiple regions based on the original image by using a weight prediction network.

[0171] The second acquisition unit is configured to, for each region, perform a fusion process on the three-dimensional lookup table corresponding to the region according to the weight corresponding to the region to obtain a color mapping relationship corresponding to multiple regions.

[0172] The image processing device provided in this embodiment can execute the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0173] Based on the above embodiments, optionally, the above processing module includes: a processing unit, a third acquisition unit, and a fusion unit, where:

[0174] The processing unit is configured to, for each region, perform a rendering process on the original image according to the color mapping relationship corresponding to the region to obtain a rendered image corresponding to multiple regions.

[0175] The third acquisition unit is configured to obtain region images corresponding to multiple regions based on multiple rendered images and region masks corresponding to multiple regions.

[0176] The fusion unit is configured to fuse multiple region images to obtain a target image.

[0177] The image processing device provided in this embodiment can execute the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0178] Based on the above embodiments, the above multiple regions include multiple target regions. The third acquisition unit is specifically configured to, for each target region, perform a fusion process on the rendered image of the target region and the region mask corresponding to the target region to obtain a target rendered image corresponding to multiple target regions; for each target region, perform a fusion process on the target rendered image of the target region and the inverse image of the mask of the original image to obtain region images corresponding to multiple target regions.

[0179] Optionally, the multiple target regions include at least one of the following:

[0180] Multiple brightness regions with different brightness levels for each brightness region;

[0181] Multiple color regions with different colors for each color region;

[0182] Multiple main body regions with different main bodies corresponding to each main body region;

[0183] Multiple local regions, with different degrees of importance of each local region in the original image.

[0184] Optionally, the image features include at least one of brightness feature, color feature, shape feature, texture feature, semantic feature, embedding feature, and local feature.

[0185] The image processing apparatus provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0186] Based on the above embodiment, the multiple regions further include a full-image region, and the apparatus further includes: a first fusion module and a second fusion module, where:

[0187] The first fusion module is configured to perform a fusion process on the rendering image of the full-image region and the region mask of the full-image region to obtain a target rendering image corresponding to the full-image region.

[0188] The second fusion module is configured to perform a fusion process on the target rendering image corresponding to the full-image region and the mask of the original image to obtain a region image corresponding to the full-image region.

[0189] The image processing apparatus provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0190] Based on the above embodiment, the above fusion unit is configured to perform a fusion process on the region images corresponding to the multiple target regions and the region image corresponding to the full-image region to obtain a target image.

[0191] The image processing apparatus provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0192] Based on the above embodiment, the above target image is a style image with a target style corresponding to the original image, and the apparatus further includes: a training module, where:

[0193] The training module is configured to train at least one of a weight prediction network and a mask prediction network according to a sample image and a sample style image corresponding to the sample image with a target style, and the mask prediction network is configured to generate a region mask corresponding to at least one region for the original image based on the image features of the original image.

[0194] The image processing apparatus provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0195] Each module in the above image processing device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0196] Based on the same inventive concept, an embodiment of the present application further provides a model training device for implementing the above-mentioned model training method. The implementation solutions provided by this device to solve problems are similar to those described in the above method. Therefore, the specific limitations in one or more embodiments of the model training device provided below can refer to the limitations on the model training method in the above text, and will not be repeated here.

[0197] In an exemplary embodiment, as Figure 10 shown, a model training device is provided, including: a first acquisition module 20, a second acquisition module 21, a third acquisition module 22, a fourth acquisition module 23, and a training module 24, where:

[0198] The first acquisition module 20 is configured to acquire a sample image and a sample style image with a target style corresponding to the sample image, and the sample image includes multiple regions.

[0199] The second acquisition module 21 is configured to acquire multiple region masks of the sample image based on the image features of the sample image, and the multiple region masks respectively correspond to the multiple regions of the sample image.

[0200] The third acquisition module 22 is configured to obtain a color mapping relationship corresponding to multiple regions based on the sample image by using an initial weight prediction network.

[0201] The fourth acquisition module 23 is configured to obtain a converted image based on the sample image, the region masks corresponding to the multiple regions, and the color mapping relationship;

[0202] The training module 24 is configured to train at least one of the initial weight prediction network and the initial mask prediction network according to the converted image and the sample style image; the initial mask prediction network is configured to generate region masks corresponding to at least one region for the sample image based on the image features of the sample image.

[0203] The model training device provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effect are similar, and will not be repeated here.

[0204] Each module in the above model training device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0205] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 11 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a model training method or an image processing method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the shell of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0206] Those skilled in the art can understand that Figure 11 the structure shown in

[0207] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0208] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0209] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0210] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., and are not limited thereto.

[0211] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.

[0212] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.

Claims

1. An image processing method, characterized in that, The method includes: Obtaining an original image, where the original image includes multiple regions; Based on the image features of the original image, obtaining multiple region masks of the original image, where the multiple region masks respectively correspond to the multiple regions of the original image; Based on the original image, using a weight prediction network to obtain the color mapping relationships corresponding to the multiple regions; Based on the original image, the region masks corresponding to the multiple regions, and the color mapping relationships, obtaining a target image.

2. The method according to claim 1, wherein The step of based on the original image, using a weight prediction network to obtain the color mapping relationships corresponding to the multiple regions includes: Based on the original image, using the weight prediction network to obtain the weights corresponding to the multiple regions; For each region, performing a fusion process on the three-dimensional lookup table corresponding to the region according to the weight corresponding to the region, to obtain the color mapping relationships corresponding to the multiple regions.

3. The method according to claim 2, characterized in that, The step of based on the original image, the region masks corresponding to the multiple regions, and the color mapping relationships, obtaining a target image includes: For each region, performing a rendering process on the original image according to the color mapping relationship corresponding to the region, to obtain the rendered images corresponding to the multiple regions; Based on the multiple rendered images and the region masks corresponding to the multiple regions, obtaining the region images corresponding to the multiple regions; Fusing the multiple region images to obtain the target image.

4. The method according to claim 3, characterized in that, The multiple regions include multiple target regions, and the step of based on the multiple rendered images and the region masks corresponding to the multiple regions, obtaining the region images corresponding to the multiple regions includes: For each target region, performing a fusion process on the rendered image of the target region and the region mask corresponding to the target region, to obtain the target rendered images corresponding to the multiple target regions; For each target region, performing a fusion process on the target rendered image of the target region and the inverse image of the mask of the original image, to obtain the region images corresponding to the multiple target regions.

5. The method according to claim 4, wherein The multiple regions further include a full-image region, and the method further includes: Performing a fusion process on the rendered image of the full-image region and the region mask of the full-image region, to obtain the target rendered image corresponding to the full-image region; Performing a fusion process on the target rendered image corresponding to the full-image region and the mask of the original image, to obtain the region image corresponding to the full-image region.

6. The method according to claim 5, wherein The step of fusing the multiple region images to obtain the target image includes: Performing a fusion process on the region images corresponding to the multiple target regions and the region image corresponding to the full-image region, to obtain the target image.

7. The method according to claim 4, wherein The multiple target regions include at least one of the following: Multiple brightness regions, where the brightness of each brightness region is different; Multiple color regions, where the color of each color region is different; Multiple subject regions, where the subject corresponding to each subject region is different; Multiple local regions, where the importance of each local region in the original image is different.

8. The method according to claim 7, wherein The image features include at least one of brightness feature, color feature, shape feature, texture feature, semantic feature, embedding feature, and local feature.

9. The method according to any one of claims 1 to 8, characterized in that, The target image is a style image with a target style corresponding to the original image, At least one of the weight prediction network and the mask prediction network is trained according to a sample image and a sample style image with a target style corresponding to the sample image. The mask prediction network is used to generate a region mask corresponding to at least one region for the original image based on the image features of the original image.

10. A model training method, characterized in that, The method includes: Obtaining a sample image and a sample style image with a target style corresponding to the sample image, where the sample image includes multiple regions; Based on the image features of the sample image, obtaining multiple region masks of the sample image, where the multiple region masks respectively correspond to the multiple regions of the sample image; Based on the sample image, obtaining the color mapping relationships corresponding to the multiple regions using an initial weight prediction network; Based on the sample image, the region masks corresponding to the multiple regions, and the color mapping relationships, obtaining a transformed image; Training at least one of the initial weight prediction network and the initial mask prediction network according to the transformed image and the sample style image; the initial mask prediction network is used to generate a region mask corresponding to at least one region for the sample image based on the image features of the sample image.

11. An image processing apparatus, characterized in that, The device includes: A first acquisition module, configured to acquire an original image, where the original image includes multiple regions; A second acquisition module, configured to acquire multiple region masks of the original image based on the image features of the original image, where the multiple region masks respectively correspond to the multiple regions of the original image; A third acquisition module, configured to obtain the color mapping relationships corresponding to the multiple regions using a weight prediction network based on the original image; A processing module, configured to obtain a target image based on the original image, the region masks corresponding to the multiple regions, and the color mapping relationships; the target image is a style-transformed image of the original image.

12. A model training device, characterized in that, The device includes: A first acquisition module, configured to acquire a sample image and a sample style image with a target style corresponding to the sample image, where the sample image includes multiple regions; A second acquisition module, configured to acquire multiple region masks of the sample image based on the image features of the sample image, where the multiple region masks respectively correspond to the multiple regions of the sample image; A third acquisition module, configured to obtain the color mapping relationships corresponding to the multiple regions using an initial weight prediction network based on the sample image; A fourth acquisition module, configured to obtain a transformed image based on the sample image, the region masks corresponding to the multiple regions, and the color mapping relationships; A training module, configured to train at least one of the initial weight prediction network and the initial mask prediction network according to the transformed image and the sample style image; the initial mask prediction network is used to generate a region mask corresponding to at least one region for the sample image based on the image features of the sample image.

13. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 10 are implemented.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 10 are implemented.

15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.