Image style migration method, and training method and device of image style migration model

By generating an image style transfer model trained by adversarial networks, using image content loss and discriminant loss, the artifact problem caused by style differences between image sensors is solved, and high-quality image style transfer is achieved, which is suitable for intelligent driving scenarios.

CN120339045APending Publication Date: 2025-07-18HORIZON JOURNEY (SHANGHAI) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510512200.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing image style migration method in the YUV domain leads to artifacts in the migrated image, and fails to effectively reduce the style differences between different image sensors, affecting the data acquisition cost and perception tasks of the autonomous driving system.

Method used

By using the image style transfer model based on the generative adversarial network, the sample image is enhanced by using the first generation network, and training is combined with image content loss and discriminative loss, the trained generation network is obtained to realize image style transfer.

Benefits of technology

It effectively reduces style differences between image sensors, improves the image quality after migration, maintains the details of the original image, reduces the amount of calculation and reduces the appearance of artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339045A_ABST
    Figure CN120339045A_ABST
Patent Text Reader

Abstract

The invention discloses an image style migration method and a training method and device of an image style migration model, and the method comprises the steps: carrying out the enhancement processing of a first sample image based on a first generation network, and obtaining a first enhanced sample image; determining an image content loss based on the first sample image and the first enhanced sample image; performing classification processing on the first enhanced sample image by using a discrimination network to obtain a first discrimination loss; iteratively training the first generation network based on the image content loss and the first discrimination loss to obtain a trained first generation network; and obtaining an image style migration model based on the trained first generation network. According to the technical scheme of the invention, the first generative network is trained through the image content loss and the first discrimination loss, so that the first generative network can be trained directly based on the generative adversarial loss of comparative learning, thereby achieving the purpose of image style migration. Meanwhile, due to the fact that pixel-by-pixel regression is not involved, the migrated image has good image quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing, and in particular, to an image style transfer method, a training method and device for an image style transfer model. Background Art

[0002] With the rapid development of deep learning technology, intelligent driving perception based on image sensors has gradually emerged. However, due to differences in photosensitive devices, noise differences in amplifier circuits, gain differences in circuits, etc. among image sensors produced by different manufacturers or different processes, differences in chromaticity and brightness will be caused at the image level. Therefore, there will be significant style differences in image data between different modules, that is, the perception model trained on the current module is difficult to generalize to another module. Therefore, there is an urgent need for a method to reduce the data style differences caused by different modules, and further reduce the data acquisition cost in the autonomous driving system. Summary of the Invention

[0003] Currently, most existing image style transfer solutions are carried out in the YUV domain (a color space) after ISP (Image Signal Processing) processing, and this method is pixel-by-pixel regression, which easily causes artifacts in the transferred image. To solve the above technical problems, the present disclosure provides an image style transfer method, a training method and device for an image style transfer model, which can solve the problem that the existing pixel-by-pixel regression method easily causes artifacts in the transferred image, thereby affecting downstream perception tasks.

[0004] In the first aspect of the present disclosure, a training method for an image style transfer model is provided, including: determining a first sample image collected by a first camera; performing enhancement processing on the first sample image based on a first generation network to obtain a first enhanced sample image corresponding to the first sample image; determining an image content loss based on the first sample image and the first enhanced sample image; performing classification processing on the first enhanced sample image by using a discriminant network to obtain a first discriminant loss; iteratively training the first generation network based on the image content loss and the first discriminant loss to obtain a trained first generation network; and obtaining the image style transfer model based on the trained first generation network.

[0005] In the second aspect of the present disclosure, an image style transfer method is provided, including: determining a to-be-processed image collected by a first camera; processing the to-be-processed image based on an image style transfer model to obtain a stylized image corresponding to the to-be-processed image; wherein, the image style transfer model is trained by using the training method for the image style transfer model provided in the first aspect above.

[0006] The third aspect of the present disclosure provides a training device for an image style transfer model, including: a first sample image determination module for determining a first sample image collected by a first camera; a first enhanced image generation module for performing enhancement processing on the first sample image based on a first generation network to obtain a first enhanced sample image corresponding to the first sample image; a content loss determination module for determining an image content loss based on the first sample image and the first enhanced sample image; a first discriminant loss determination module for classifying the first enhanced sample image by using a discriminant network to obtain a first discriminant loss; a first generation network training module for iteratively training the first generation network based on the image content loss and the first discriminant loss to obtain a trained first generation network; and a transfer model determination module for obtaining the image style transfer model based on the trained first generation network.

[0007] The fourth aspect of the present disclosure provides an image style transfer device, including: an image determination module for determining a to-be-processed image collected by a first camera; an image generation module for processing the to-be-processed image based on an image style transfer model to obtain a stylized image corresponding to the to-be-processed image; wherein, the image style transfer model is trained by using the training method for an image style transfer model provided in the first aspect above.

[0008] The fifth aspect of the present disclosure provides a computer-readable storage medium storing a computer program for executing the training method for an image style transfer model provided in the first aspect above, or the image style transfer method provided in the second aspect above.

[0009] The sixth aspect of the present disclosure provides an electronic device, including: a processor; a memory for storing executable instructions of the processor; and the processor for reading the executable instructions from the memory and executing the instructions to implement the training method for an image style transfer model provided in the first aspect above, or the image style transfer method provided in the second aspect above.

[0010] The seventh aspect of the present disclosure provides a computer program product, when the instructions in the computer program product are executed by a processor, implementing the training method for an image style transfer model provided in the first aspect above, or the image style transfer method provided in the second aspect above.

[0011] Based on the training method of the image style transfer model provided by the present disclosure, the first enhanced sample image after style transformation can be obtained by processing the first sample image through the first generation network; furthermore, the first generation network is trained by using the image content loss obtained based on the first sample image and the first enhanced sample image, and the first discrimination loss obtained by classifying the first enhanced sample image through the discrimination network, and the trained first generation network is obtained, and then the trained image style transfer model is obtained; in this way, the first generation network can be directly trained based on the generation adversarial loss of contrastive learning to achieve the purpose of image style transfer; at the same time, since the first generation network is directly trained based on the generation adversarial loss of contrastive learning to achieve the purpose of image style transfer in the embodiments of the present disclosure and does not involve pixel-by-pixel regression, the transferred image has better image quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 FIG. is a schematic diagram of a scenario applicable to the training method of the image style transfer model provided by an exemplary embodiment of the present disclosure.

[0013] Figure 2A FIG. is a schematic flowchart of the training method of the image style transfer model provided by an exemplary embodiment of the present disclosure.

[0014] Figure 2B FIG. is a schematic structural diagram of the first generation network provided by an exemplary embodiment of the present disclosure.

[0015] Figure 3 FIG. is a schematic flowchart of the training method of the image style transfer model provided by another exemplary embodiment of the present disclosure.

[0016] Figure 4 FIG. is a schematic flowchart of the training method of the image style transfer model provided by still another exemplary embodiment of the present disclosure.

[0017] Figure 5A FIG. is a schematic flowchart of the training method of the image style transfer model provided by still another exemplary embodiment of the present disclosure.

[0018] Figure 5B FIG. is a schematic structural diagram of the training network of the image style transfer model provided by an exemplary embodiment of the present disclosure.

[0019] Figure 6 FIG. is a schematic flowchart of the training method of the image style transfer model provided by still another exemplary embodiment of the present disclosure.

[0020] Figure 7 FIG. is a schematic flowchart of the training method of the image style transfer model provided by still another exemplary embodiment of the present disclosure.

[0021] Figure 8It is a schematic flowchart of an image style transfer method provided by an exemplary embodiment of the present disclosure.

[0022] Figure 9 It is a schematic structural diagram of a training device for an image style transfer model provided by an exemplary embodiment of the present disclosure.

[0023] Figure 10 It is a schematic structural diagram of a training device for an image style transfer model provided by another exemplary embodiment of the present disclosure.

[0024] Figure 11 It is a schematic structural diagram of a training device for an image style transfer model provided by still another exemplary embodiment of the present disclosure.

[0025] Figure 12 It is a schematic structural diagram of an image style transfer device provided by an exemplary embodiment of the present disclosure.

[0026] Figure 13 It is a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. Detailed implementation manners

[0027] To explain the present disclosure, exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. It should be understood that the present disclosure is not limited by the exemplary embodiments.

[0028] It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present disclosure.

[0029] Application Overview

[0030] The image style transfer model obtained by using the training method for the image style transfer model provided by the embodiments of the present disclosure can be applied to, for example, the intelligent driving scenario of a vehicle and any other applicable scenarios.

[0031] As Figure 1 shown, at least a plurality of cameras 11 are provided on the vehicle 10, and the images output by the cameras 11 are raw images without any processing, that is, RAW images. The training method for the image style transfer model provided by the embodiments of the present disclosure performs image style transfer and training of the image style transfer model based on the RAW image. After the training of the image style transfer model is completed, the trained image style transfer model can be used to perform image style transfer on the image to be processed to obtain a stylized image. Further, the stylized image can be used as input data for vehicle environment perception and input into the vehicle environment perception model to obtain a corresponding environment perception result. It should be noted that Figure 1For illustrative purposes only, the embodiments of the present disclosure do not limit Figure 1 the number and positions of multiple cameras 11 in

[0032] Generally, there are two methods for image style transfer: The first is image style transfer performed in the YUV domain after ISP processing, and this transfer method is usually pixel-by-pixel regression. For example, if the YUV domain image input to the neural network is an image of 5*5*3, then the neural network needs to regress the pixel values of each pixel point in this image, that is, 75 pixel values need to be regressed. Therefore, this method has a large computational amount and pixel-by-pixel regression easily causes artifacts in the transferred image, and style transfer in the YUV domain loses the advantage of the high bit width of the original RAW image. The second method is to debug a CCM matrix for image style transfer through a large number of manual experiments, and then use this CCM matrix for image style transfer. However, this method not only consumes manpower, but also does not consider the differences in underlying signals, signal-to-noise ratios, color differences, etc. caused by RAW images collected based on different sensors, different devices, or different scenarios. At this time, the same CCM matrix is used for image processing, resulting in poor image style transfer effects.

[0033] To solve the above problems, the embodiments of the present disclosure provide a training method for an image style transfer model. By processing a first sample image through a first generation network, a first enhanced sample image after style transformation can be obtained; then, using the image content loss obtained based on the first sample image and the first enhanced sample image, and the first discrimination loss obtained by classifying the first enhanced sample image through a discrimination network to train the first generation network, a trained first generation network is obtained, and then a trained image style transfer model is obtained; in this way, the first generation network can be directly trained based on the generative adversarial loss of contrastive learning to achieve the purpose of image style transfer; at the same time, since the embodiments of the present disclosure directly train the first generation network based on the generative adversarial loss of contrastive learning to achieve the purpose of image style transfer and do not involve pixel-by-pixel regression, the transferred image has better image quality.

[0034] Exemplary Method

[0035] Figure 2A FIG. is a schematic flowchart of a training method for an image style transfer model provided by the embodiments of the present disclosure. This embodiment can be applied to an electronic device, such as a server, as Figure 2A shown, and this method includes the following steps S201 - step S206.

[0036] Step S201: Determine the first sample image collected by the first camera.

[0037] Exemplarily, as Figure 1 shown, the RAW image collected by the first camera in the camera 11 can be obtained first as the first sample image. This first sample image can also be referred to as the target domain RAW map. The image style transfer model in the embodiments of the present disclosure needs to process this target domain RAW map, so as to, while retaining the image content of the target domain RAW map, transfer and transform its image style to be the same as that of the source domain image. Among them, the source domain image can be Figure 1 the original RAW map collected by other cameras (such as the second camera, and the first camera and the second camera are different cameras) in the camera 11 except this first camera; or, the source domain image can be Figure 1 the image obtained after processing the original RAW map collected by other cameras in the camera 11 except this first camera, for example, the SRGB image corresponding to the original RAW map after processing.

[0038] It should be noted that the image style transfer in the embodiments of the present disclosure refers to applying the style features (such as color, texture, shape, etc.) of the style image (source domain) to the content image (target domain), so as to generate a new image. Therefore, in the embodiments of the present disclosure, the target domain refers to the domain corresponding to the style features of the original image, which is the starting point of the style transfer. The source domain refers to the domain corresponding to the style features to be transferred, which is the end point of the style transfer.

[0039] In some examples, as Figure 1 shown, the sample image collected by the first camera in the camera 11 in the first acquisition scenario can be obtained as the target domain image; and, the sample image collected by the second camera in the camera 11 in the second scenario can be used as the source domain image; where the first camera and the second camera are the same camera. That is to say, in the embodiments of the present disclosure, the source domain image and the target domain image can also be images collected by the same camera in different acquisition scenarios. For example, the source domain image and the target domain image are images collected by the same camera during the day and at night respectively.

[0040] Moreover, in the embodiments of the present disclosure, the first camera can be any camera in the camera 11, and the embodiments of the present disclosure do not limit this.

[0041] Step S202: Perform enhancement processing on the first sample image based on the first generation network to obtain a first enhanced sample image corresponding to the first sample image.

[0042] Exemplarily, during the training process of the embodiments of the present disclosure, a first generation network is included. The function of this first generation network is to receive a first sample image as input and generate a first enhanced sample image similar to the source domain style. Among them, the network structure of the first generation network can refer to the examples described below Figure 2B below.

[0043] In some examples, in the embodiments of the present disclosure, the first sample image can be enhanced based on the first generation network to obtain a first enhanced sample image corresponding to the first sample image; this first enhanced sample image is a sample image after image style transfer of the first sample image. This first generation network (the first CCM Module) is a network based on parameter estimation. Without involving transformations on the local content of the image, it can achieve a global style transformation, so it can better preserve the detail information of the original image. In some examples, the first sample image can be processed based on the first generation network first to obtain a color transformation matrix corresponding to the first sample image; then the first sample image is multiplied by the color transformation matrix corresponding to the first sample image to obtain the first enhanced sample image. Among them, the color transformation matrix corresponding to the first sample image is composed of a 3x3 matrix, and each element in the color transformation matrix represents the weight of the corresponding color channel. When the first sample image is transformed and enhanced through the corresponding color transformation matrix, the value of each color channel will be weighted and adjusted according to the weights in the matrix, and then the first enhanced sample image is obtained.

[0044] It should be noted that the CCM color transformation matrix in the embodiments of the present disclosure can be 3-channel, for example, 3x3 in size. In addition, the CCM matrix can also be other variants. For example, the CCM matrix can be 4x3 in size, and the embodiments of the present disclosure do not limit this. And, in the embodiments of the present disclosure, the 3x3 size is taken as an example for illustrative purposes, not a limitation thereof.

[0045] In some embodiments, such as Figure 2BAs shown in the figure, the network structure of the first generation network at least includes a residual module 21, a pooling module 22, and a data reorganization module 23. Among them, taking the first sample image as the input, after multiple residual processes (such as three residual processes) in the residual module 21 (ResConvBlock), features with the same spatial dimension as the image and 9 channels are output. Furthermore, the pooling module 22 (GlobalPooling) is used to perform global pooling on the 9-channel features to obtain the pooled features. Then, the data reorganization module 23 (Reshape) is used to reorganize the pooled features to obtain 9 parameters of the CCM (Color Correction Matrix) matrix. For example, if the pooled feature is a vector, the data reorganization module 23 is required to reshape the vector into a 3*3 CCM matrix. Finally, the CCM matrix including 9 parameters is multiplied by the first sample image to obtain the first enhanced sample image.

[0046] It should be noted that the network structure of the first generation network described above is only an exemplary illustration, and the embodiments of the present disclosure do not limit this. For example, in actual applications, the residual module 21 in the first generation network can be replaced with a downsampling module. That is, the downsampling module is used to perform downsampling on the first sample image to obtain features with the same spatial dimension as the image and 9 channels.

[0047] In the embodiments of the present disclosure, the dimension (size) of the CCM color transformation matrix depends on the channel dimension of the input and output images of the application scenario. For example, if the input image is a four-channel image of RGBW and the output channel is a three-channel image of RGB, the output of the first generation network is features with 12 channels. On this basis, global pooling is performed on the 12-channel features to obtain a 4x3 CCM color check matrix.

[0048] Step S203: Determine the image content loss based on the first sample image and the first enhanced sample image.

[0049] Exemplarily, the image content loss Loss can be determined based on the first sample image and the first enhanced sample image. cl This image content loss is used to ensure the content consistency between the first sample image and the first enhanced sample image during the conversion process, so it can also be called the contrast learning loss; among them, content consistency means that when the style features of the style image are transferred to the content image, the original semantics, structure, and main features of the content image are not damaged or distorted; therefore, the image content loss in the embodiments of the present disclosure is used to ensure that the original semantics, structure, and main features of the image remain unchanged during the process of converting the first sample image into the first enhanced sample image.

[0050] In some examples, the first sample image and the first enhanced sample image can be input into the MLP netF network, and then the image content loss Loss is obtained based on the output of the MLP netF network cl Calculate the difference between the image features of the first sample image and the image features of the first enhanced sample image, so as to ensure that the image content of the first enhanced sample image remains unchanged

[0051] In some examples, the training process of the first sample image at least includes a generation stage. In the generation stage, the main objective is to optimize the first generation network, and this stage requires at least the use of the image content loss Loss cl For example, the first enhanced sample image and the first sample image can be input into a feature extraction network (such as the MLP netF network) at the same time to calculate the difference between their image features, and ensure that the image content of the first enhanced sample image remains unchanged

[0052] Step S204: Use the discriminative network to classify the first enhanced sample image to obtain the first discriminative loss

[0053] Exemplarily, the training process of the embodiments of the present disclosure further includes a discriminative network, and the objective of this discriminative network is to classify samples and determine whether the samples are real (samples corresponding to the source domain) or generated (samples corresponding to the target domain). Furthermore, the discriminative network can be used to classify the first enhanced sample image to obtain the first discriminative loss Loss adv1 Among them, the discriminative network includes but is not limited to: discriminator, UNET RESNET, etc. And, the specific network structure of the discriminator in the embodiments of the present disclosure is not limited

[0054] In some examples, when optimizing the first generation network in the generation stage, it is necessary to keep the weights of the discriminative network (D) unchanged. In addition to using the image content loss Loss in this stage cl it is also necessary to use the first discriminative loss Loss adv1 For example, the first enhanced sample image can be input into the discriminative network (D) with unchanged weights to obtain the first discriminative loss Loss adv1 where This first discriminative loss is used to make the discriminative network classify the first enhanced image sample as real as possible. Among them, D is the discriminative network, and Target.Enhance is the first enhanced sample image

[0055] Step S205: Iteratively train the first generation network based on the image content loss and the first discriminative loss to obtain the trained first generation network

[0056] Exemplarily, when iteratively training the first generation network based on the image content loss and the first discriminant loss, the training can be terminated to obtain the trained first generation network in response to the first generation network satisfying the convergence condition. Among them, the first generation network satisfying the convergence condition can be achieved in the following three ways: The first is that the image content loss and the first discriminant loss are respectively less than a first preset value and a second preset value, where the first preset value and the second preset value can be the same or different. The second is that the change in the weights of the first generation network between two iterations is less than a third preset value. The third is that the number of iterations reaches a preset number. It should be noted that in actual use, those skilled in the art can select an appropriate convergence condition to train the first generation network according to actual needs.

[0057] Step S206: Based on the trained first generation network, obtain an image style transfer model.

[0058] Exemplarily, an image style transfer model can be obtained based on the trained first generation network. For example, the image style transfer model includes the trained first generation network and an image processing module. The first generation network is used to process the first sample image to obtain a color correction matrix corresponding to the first sample image; the image processing module is used to perform convolution on the first sample image and the color correction matrix corresponding to the first sample image to obtain a first sample enhanced image. Furthermore, the image to be processed can be input into the trained image style transfer model to obtain a stylized image with the same style as the source domain image.

[0059] The training method of the image style transfer model provided by the embodiments of the present disclosure can obtain a first enhanced sample image after style transformation by processing the first sample image through the first generation network; furthermore, the first generation network is trained using the image content loss obtained based on the first sample image and the first enhanced sample image, and the first discriminant loss obtained by classifying and processing the first enhanced sample image using the discriminant network, to obtain the trained first generation network and then obtain the trained image style transfer model; in this way, the first generation network can be directly trained based on the generation adversarial loss of contrast learning to achieve the purpose of image style transfer; at the same time, since the embodiments of the present disclosure directly train the first generation network based on the generation adversarial loss of contrast learning to achieve the purpose of image style transfer and do not involve pixel-by-pixel regression, the transferred image has better image quality.

[0060] As Figure 3 shown, on the basis of the above Figure 2A shown embodiment, Figure 2A the discriminant network in

[0061] can be trained through the following steps S301 - S304. Step S301: Determine the second sample image collected by the second camera.

[0062] Exemplarily, as Figure 1 shown, the RAW image captured by the second camera in the camera 11 can be obtained first as the second sample image. That is, the second sample image can be the source domain RAW image. Here, the second camera can be any camera in the camera 11, and the embodiments of the present disclosure do not limit this. And the second camera and the first camera are different cameras.

[0063] Step S302: Perform enhancement processing on the second sample image based on the second generation network to obtain a second enhanced sample image corresponding to the second sample image.

[0064] Exemplarily, during the training process of the embodiments of the present disclosure, there is also a second generation network. The function of the second generation network is to receive the second sample image as input and generate a second enhanced sample image similar to the label ground truth (target style image). Among them, the network structure of the second generation network can be implemented with reference to the Figure 2B shown network structure, which will not be elaborated here.

[0065] In some examples, in the embodiments of the present disclosure, the second sample image can be enhanced based on the second generation network to obtain a second enhanced sample image corresponding to the second sample image; the second enhanced sample image is a sample image after image style transfer of the second sample image. The second generation network (second CCM Module) is a network based on parameter estimation, which can achieve a global style transformation without involving transformations on the local content of the image, so it can better preserve the detail information of the original image. In some examples, the second sample image can be processed based on the second generation network first to obtain a color transformation matrix corresponding to the second sample image; then the second sample image is multiplied by the color transformation matrix corresponding to the second sample image to obtain the second enhanced sample image. Among them, the color transformation matrix corresponding to the second sample image is composed of a 3x3 matrix, and each element in the color transformation matrix represents the weight of the corresponding color channel. When the second sample image is transformed and enhanced through the corresponding color transformation matrix, the value of each color channel will be weighted and adjusted according to the weights in the matrix, and then the second enhanced sample image is obtained.

[0066] Step S303: Use the discriminant network to classify the first enhanced sample image and the second enhanced sample image to obtain a second discriminant loss.

[0067] Exemplarily, in the embodiments of the present disclosure, the foregoing discriminant network can also be used to classify the first enhanced sample image and the second enhanced sample image to obtain a second discriminant loss Loss adv2 .

[0068] In some examples, for the first sample image, the generative adversarial loss based on contrastive learning is mainly adopted. In addition to the above-mentioned generative stage, the training process of the first sample image at least further includes a discriminative stage. Furthermore, in the discriminative stage, it is necessary to keep the weights of the first generative network and the second generative network unchanged and optimize the weights of the discriminative network (D). At this time, it is necessary to input the first enhanced sample image and the second enhanced sample image into the discriminative network simultaneously to obtain the second discriminative loss Loss. adv2 . The network weights of the discriminative network are optimized only by the second discriminative loss Loss. adv2 , that is, the second discriminative loss is used to enable the discriminative network to have a real discrimination ability, that is, to be able to discriminate the second enhanced sample image as real and the first enhanced sample image as fake. At this time, the second discriminative loss Loss adv2 can be obtained through the following formula (1):

[0069]

[0070] where Target.Enhance is the first enhanced sample image, Source.Enhance is the second enhanced sample image, is the expected value.

[0071] Step S304: Iteratively train the discriminative network based on the second discriminative loss to obtain the trained discriminative network.

[0072] Exemplarily, the discriminative network can be iteratively trained based on the second discriminative loss to obtain the trained discriminative network. The input of this discriminative network is the first enhanced sample image and / or the second enhanced sample image, and the output of this discriminative network is 0 or 1; where an output of 1 indicates that the discriminative network's discrimination result for the input is real, and an output of 0 indicates that the discriminative network's discrimination result for the input is fake. The purpose of training the discriminative network is to enable the discriminative network to finally discriminate the second enhanced sample image as real and the first enhanced sample image as fake. At the same time, the purpose of training the first generative network is to generate a realistic first enhanced sample so that the discriminative network discriminates the first enhanced sample image as real. In this way, during the training process, the first generative network and the second generative network perform adversarial training in an alternating training manner, and finally enable the first generative network to generate a first enhanced sample image with the same style as the second enhanced sample image.

[0073] In some examples, in response to the discriminative network satisfying the convergence condition, the training is terminated to obtain the trained discriminative network. Among them, the discriminative network satisfying the convergence condition can be achieved in the following three ways: The first is that the loss of the discriminative network is less than a certain preset value. The second is that the change in the weights of the discriminative network between two iterations is less than a certain preset value. The third is that the number of iterations reaches the preset number.

[0074] It should be noted that in the actual use process, those skilled in the art can select appropriate convergence conditions to train the discrimination network according to actual needs. However, since the first generation network, the second generation network, and the discrimination network all need to update parameters during one iteration of training, the third method can be preferentially selected to determine the termination condition of the training process.

[0075] The training method of the image style transfer model provided by the embodiments of the present disclosure first determines a second sample image collected by a second camera; enhances the second sample image based on a second generation network to obtain a second enhanced sample image corresponding to the second sample image; then classifies the first enhanced sample image and the second enhanced sample image by using a discrimination network to obtain a second discrimination loss; and finally iteratively trains the discrimination network based on the second discrimination loss to obtain a trained discrimination network. In this way, since the discrimination network in the embodiments of the present disclosure is trained by the second discrimination loss obtained from the first enhanced sample image in the target domain and the second enhanced sample image in the source domain, and the first enhanced sample image is determined by the first sample image, the first sample image in the target domain can have the same color and brightness space as the source domain (the second enhanced sample image) through this discrimination network.

[0076] As Figure 4 shown, on the basis of the above Figure 3 shown embodiments, Figure 3 the second generation network in

[0077] Step S401: Determine the target style image corresponding to the second sample image.

[0078] Exemplarily, the target style image can be the SRGB image corresponding to the second sample image. In the embodiments of the present disclosure, this target style image can be used as the label (true value) of the second sample image to supervise the training of the first generation network. Among them, the SRGB (standard Red Green Blue, a color space standard) image is based on the RGB principle, but the color values are processed by a specific conversion formula, and its color values are more in line with the perceptual characteristics of the human eye.

[0079] Step S402: Determine the image error loss based on the second enhanced sample image and the target style image.

[0080] Exemplarily, the image error loss Loss mse can be determined based on the second enhanced sample image and the target style image; among them, this image error loss Loss mse is pixel-level and is mainly used to measure the deviation between the second enhanced sample image and the target style image.

[0081] In some examples, since each row of the color correction matrix (CCM) represents the transformation coefficients of a primary color channel (red, green, blue, or the base colors in other color spaces), in the prior art, if the CCM color correction matrix is used to process an image, it is necessary to set the constraint that the sum of the channel gains is 1 to balance the weights of each channel, so as to ensure that each color channel will not shift during the conversion process, resulting in color distortion and the conservation of color energy during the color conversion process. However, setting the constraint that the sum of the channel gains is 1 will lose the high-bit information of the original RAW image, and the constraint that the sum of the channel gains is 1 can only adjust the color of the target domain. In the embodiments of the present disclosure, the image error loss Loss is used in the source domain mse to constrain the brightness and white balance of each channel, so that there is no need to set the constraint that the sum of the channel gains is 1. Therefore, in the embodiments of the present disclosure, while adjusting the color mapping, it is also possible to achieve the adaptive alignment of the brightness between different camera modules without losing the high-bit information of the original RAW (the second sample image).

[0082] In some embodiments, based on the second enhanced sample image and the target style image, determining the image error loss includes: determining the difference between the second enhanced sample image and the target style image; and determining the image error loss based on the difference.

[0083] Exemplarily, the image error loss Loss can be determined by the following formula (2) mse :

[0084] Loss mse = ||Source.RGB - Source.Enhance|| Formula (2);

[0085] where Source.RGB is the second sample image and Source.Enhance is the second enhanced sample image.

[0086] In some examples, by calculating the image error loss Loss between the second enhanced sample image and the target style image mse , it is possible to constrain the second generation network to output an image in the high-bit sRGB color space when inputting the second sample image. Here, since the image error loss Loss mse can constrain the color space of the second enhanced sample image, and the second generation network outputs a global color correction matrix, it is possible to ensure the high-bit characteristics of the second enhanced sample image.

[0087] Step S403: Iteratively train the second generation network based on the image error loss to obtain the trained second generation network.

[0088] Exemplarily, when iteratively training the second generation network based on the image error loss to obtain the trained second generation network, the training can be terminated to obtain the trained second generation network in response to the second generation network satisfying the convergence condition. Among them, there are three ways to implement the second generation network satisfying the convergence condition: the first is that the loss of the second generation network is less than a certain preset value. The second is that the change in the discriminator network weights between two iterations is less than a certain preset value. The third is that the number of iterations reaches a preset number.

[0089] It should be noted that in the actual use process, those skilled in the art can select appropriate convergence conditions to train the second generation network according to actual needs. However, since the first generation network, the second generation network, and the discriminator network all need to update parameters during one iteration training, the third method can be preferentially selected to determine the termination condition of the training process.

[0090] The training method of the image style transfer model provided by the embodiments of the present disclosure determines the target style image corresponding to the second sample image; determines the image error loss based on the second enhanced sample image and the target style image; iteratively trains the second generation network based on the image error loss to obtain the trained second generation network; thus, since the embodiments of the present disclosure perform semi-supervised training based on the target style image and do not need to constrain the white balance at the parameter level, the embodiments of the present disclosure do not involve the constraint that the sum of the gains of each color channel is 1. Furthermore, while adjusting the color mapping, the embodiments of the present disclosure can also achieve brightness adaptive alignment between different camera modules without losing the high-bit information of the original RAW image.

[0091] As Figure 5A shown, on the basis of the above Figure 2A shown embodiment, step S203 may include the following steps S2031 - step S2033.

[0092] Step S2031, based on a preset partitioning rule, respectively perform image block partitioning on the first enhanced sample image and the first sample image to obtain a plurality of first image blocks corresponding to the first enhanced sample image, and a plurality of second image blocks corresponding to the first sample image.

[0093] Exemplarily, based on a preset partitioning rule, the first enhanced sample image can be partitioned into image blocks (patch partitioning) to obtain a plurality of first image blocks corresponding to the first enhanced sample image. At the same time, based on this preset partitioning rule, the first sample image is partitioned into image blocks to obtain a plurality of second image blocks corresponding to the first sample image. For example, the preset partitioning rule can be to divide the image into multiple 3*3 image blocks. Further, the first enhanced sample image can be divided into multiple 3*3 first image blocks, and the first sample image can be divided into multiple 3*3 second image blocks.

[0094] It should be noted that no specific division rule is limited in the embodiments of the present disclosure. Those skilled in the art can determine the specific division rule according to factors such as the size of the first sample image or the first enhanced sample image.

[0095] Step S2032: Respectively perform feature extraction on the multiple first image patches and the multiple second image patches based on the feature extraction network to obtain the first image patch features corresponding to the respective first image patches and the second image patch features corresponding to the respective second image patches.

[0096] Exemplarily, in the training process of the embodiments of the present disclosure, there is also a feature extraction network, which is used to perform feature extraction on the first sample image and the first sample enhanced image. For example, this feature extraction network can be the MLPnetF network. No specific network structure of the feature extraction network is limited in the embodiments of the present disclosure.

[0097] In some examples, the multiple first image patches and the multiple second image patches can be respectively subjected to feature extraction based on a multi-layer perceptron network (MLP netF network) to obtain the first image patch features corresponding to the respective first image patches and the second image patch features corresponding to the respective second image patches. Of course, the multiple first image patches and the multiple second image patches can also be respectively subjected to feature extraction based on other feature extraction networks other than the MLP netF network, and the embodiments of the present disclosure do not limit this.

[0098] Step S2033: Determine the image content loss based on the respective first image patch features and the respective second image patch features.

[0099] Exemplarily, the principle that the similarity between the first image patch features and the second image patch features at the same position should be large and the similarity between the first image patch features and the second image patch features at different positions should be small can be utilized to determine the image content loss; wherein, the same position means that the position of the first image patch feature in the first enhanced sample image is the same as the position of the second image patch feature in the first sample image; similarly, different positions mean that the position of the first image patch feature in the first enhanced sample image is different from the position of the second image patch feature in the first sample image. In some examples, for a certain first image patch, the first similarity between the first image patch feature corresponding to the first image patch and the second image patch feature corresponding to the second image patch at the same position can be determined, as well as the second similarity between the first image patch feature corresponding to the first image patch and the second image patch feature corresponding to the second image patch at different positions; furthermore, based on the ratio between the first similarity and the second similarity, the image content loss corresponding to the first image patch can be determined.

[0100] In some examples, the image content loss Loss in the embodiments of the present disclosure can be determined by the following formula (3). cl :

[0101]

[0102] Wherein, v is the first image patch feature corresponding to the first image patch; v + is the second image patch feature corresponding to the second image patch (positive sample) having the same position as the first image patch; v - is the second image patch feature corresponding to the second image patch (negative sample) having a different position from the first image patch; N is the number of image patches of the second image patches having different positions from the first image patch, which is a hyperparameter; τ is used to control the discrimination degree of the first generation network for negative samples. The smaller the τ value, the more the first generation network focuses on difficult negative samples, and τ is also a hyperparameter.

[0103] It should be noted that the above N is not the number of all negative samples, and negative samples can be randomly selected. For example, if there are 100 negative samples, 50 negative samples can be randomly selected, then the number of N is 50.

[0104] In some embodiments, based on each first image patch feature and each second image patch feature, determining the image content loss includes: determining the position information of the first image patch feature in the first enhanced sample image; based on the first sample image and the position information, determining multiple second target image patch features from multiple second image patch features; based on the first image patch feature and the multiple second target image patch features, determining the image patch content loss; based on the image patch content losses corresponding to each first image patch, determining the image content loss.

[0105] Exemplarily, for a certain first image patch feature, the position information of the first image patch feature in the first enhanced sample image can be determined first, and then based on the position information and the first sample image, multiple second target image patch features can be determined from multiple second image patch features; for example, the second target image patch feature is the image patch feature at this position information in the first sample image, or the second target image patch feature is a partial image patch feature at other positions in the first sample image except this position information. Finally, based on the first image patch feature and the multiple second target image patch features, determining the image patch content loss corresponding to the first image patch feature; furthermore, based on the image patch content losses corresponding to each first image patch, determining the image content loss.

[0106] The training method of the image style transfer model provided by the embodiments of the present disclosure divides the first enhanced sample image and the first sample image into image patches respectively based on a preset division rule, to obtain a plurality of first image patches corresponding to the first enhanced sample image and a plurality of second image patches corresponding to the first sample image; performs feature extraction on the plurality of first image patches and the plurality of second image patches respectively based on a feature extraction network, to obtain first image patch features corresponding to the respective first image patches and second image patch features corresponding to the respective second image patches; determines an image content loss based on the respective first image patch features and the respective second image patch features; thus, it is possible to ensure the consistency of the image details of the image before style transfer and the image after style transfer by calculating the difference between the first image patch features and the second image patch features.

[0107] In some embodiments, as Figure 5B shown, the network structure corresponding to the training method of the image style transfer model provided by the embodiments of the present disclosure at least includes a first generation network 51, a second generation network 52, a discriminant network 53, and a feature extraction network 54. Among them, the network structure can be divided into two branches, namely branch 1 and branch 2. Branch 1 is mainly used to process the transformation of the source domain RAW image (the second sample image), and branch 2 is mainly used to process the conversion of the luminance and color spaces of the target domain RAW image (the first sample image).

[0108] First, in the data acquisition stage, it is necessary to obtain the source domain RAW image, the SRGB image (the target style image) of the source domain, and the target domain RAW image. There are significant differences in luminance and color between the source domain RAW image and the target domain RAW image due to differences in the camera modules.

[0109] Secondly, in the network training stage, for branch 1, it is mainly to enhance the luminance space and color space of the source domain to the SRGB space corresponding to the SRGB image based on the source domain RAW image and the SRGB image (used as the ground truth) of the source domain; among them, the second generation network 52 takes the source domain RAW image as the input and outputs the source domain stylized image (the second enhanced sample image) corresponding to the source domain RAW image. For branch 2, it is mainly to map the target domain RAW image to the SRGB space in an unsupervised manner based on the target domain RAW image and the discriminant network 53; among them, the first generation network 51 takes the target domain RAW image as the input and outputs the target domain stylized image (the first enhanced sample image) corresponding to the target domain RAW image.

[0110] That is to say, in the embodiments of the present disclosure, the second color correction matrix corresponding to the source-domain RAW image is regressed by using the source-domain RAW image through the second generation network 52. The product of the second color correction matrix corresponding to the source-domain RAW image and the source-domain RAW image can adjust the brightness and color, and then align with the SRGB image in the source domain, while the bit depth remains unchanged. Since the purpose of image style transfer is to transfer the style of the source domain to the target domain, it is necessary to restore the target-domain RAW image to a color and brightness space similar to that of the source domain. Therefore, the discrimination loss of the discrimination network 53 is used to constrain the consistency of color and brightness. Similarly, in the embodiments of the present disclosure, the first color correction matrix corresponding to the target-domain RAW image is regressed by using the target-domain RAW image through the first generation network 51. The product of the first color correction matrix corresponding to the target-domain RAW image and the target-domain RAW image can obtain the target-domain stylized image. At the same time, the feature extraction network 54 can extract features from the target-domain RAW image and the target-domain stylized image, and determine the image content loss based on the extracted features to ensure that the content of the transferred image remains unchanged. Thus, it can be seen that whether it is the source-domain RAW image or the target-domain RAW image, after being processed by the first generation network 51 or the second generation network 52, a high-bit SRGB space image can be obtained (because both the source-domain stylized image and the target-domain stylized image are SRGB space images).

[0111] It should be noted that although the SRGB image in the source domain is used as the strong supervision for the enhanced image in the source domain in the embodiments of the present disclosure, since the network essentially regresses a global brightness and color space conversion, the strong supervision in the source domain will not destroy the high-bit characteristics of the enhanced image, enabling the source-domain RAW image to have the advantage of high bit depth after passing through the second generation network 52. At the same time, the target RAW image and the source-domain RAW image with different chromaticities and brightnesses can be mapped to the same SRGB space after passing through the first generation network 51 and the second generation network 52 respectively while keeping the original RAW information unchanged, so as to reduce the style difference between the two and effectively improve the perception performance of the downstream perception model on cross-camera module data.

[0112] As Figure 6 shown, on the basis of the above Figure 2A shown embodiment, step S202 may include the following steps S2021 - step S2022.

[0113] Step S2021: Perform transformation processing on the first sample image based on the first generation network to obtain the first color correction matrix.

[0114] Exemplarily, the input of the first generation network is the first sample image, and the output is the color correction matrix corresponding to the first sample image, that is, the first color correction matrix. In the embodiments of the present disclosure, when the input of the first generation network is different, the CCM color correction matrix output by the network is different, and the parameter values of the color correction matrix are adaptively output according to the input.

[0115] In some embodiments, performing a transformation process on the first sample image based on the first generation network to obtain the first color correction matrix includes: performing a residual process on the first sample image based on at least one residual module in the first generation network to obtain residual features corresponding to the first sample image; performing global pooling processing on the residual features to obtain the pooled residual features; and reorganizing the pooled residual features to obtain the first color correction matrix.

[0116] Exemplarily, as Figure 2B shown, the first generation network includes at least one residual module 21, a pooling module 22, and a data reorganization module 23. Therefore, the at least one residual module 21 can be used to perform a residual process on the first sample image to obtain residual features corresponding to the first sample image; further, the pooling module 22 is used to perform global pooling processing on the residual features to obtain the pooled residual features; finally, the data reorganization module 23 is used to reorganize the pooled residual features to obtain a 3×3 first color correction matrix.

[0117] It should be noted that the network structure of the above first generation network is only for exemplary illustration, and the embodiments of the present disclosure do not limit this. In actual applications, the functional structure of the first generation network can also be implemented in other ways. For example, the first sample image can be downsampled to obtain downsampled features corresponding to the first sample image; further, global pooling processing is performed on the downsampled features corresponding to the first sample image to obtain the pooled downsampled features; finally, the data reorganization module is used to reorganize the pooled downsampled features to obtain a 3×3 first color correction matrix.

[0118] Step S2022: Perform enhancement processing on the first sample image based on the first color correction matrix to obtain a first enhanced sample image corresponding to the first sample image.

[0119] Exemplarily, the first enhanced sample image corresponding to the first sample image can be determined based on the product between the first color correction matrix and the first sample image.

[0120] In some embodiments, enhancing the first sample image based on the first color correction matrix to obtain a first enhanced sample image corresponding to the first sample image includes: performing linear interpolation processing on the first sample image to obtain an interpolated first sample image; multiplying the interpolated first sample image by the first color correction matrix to obtain a first enhanced sample image corresponding to the first sample image.

[0121] Exemplarily, in the actual processing, the first sample image can be processed first, such as performing image interpolation processing on the first sample image. Then, based on the product between the processed first sample image and the first color correction matrix, the first enhanced sample image corresponding to the first sample image is determined.

[0122] The training method of the image style transfer model provided by the embodiments of the present disclosure obtains a first color correction matrix by performing transformation processing on the first sample image based on the first generation network; enhancing the first sample image based on the first color correction matrix to obtain a first enhanced sample image corresponding to the first sample image; in this way, by regressing a CCM matrix, a global transformation can be performed on the first sample image directly collected by the camera to achieve the purpose of image style transfer, thereby better preserving the detail information of the original image.

[0123] As Figure 7 shown, based on the above Figure 2A shown embodiment, step S205 may include the following steps S2051 - step S2054.

[0124] Step S2051: Determine a first gradient corresponding to the image content loss based on the image content loss.

[0125] Exemplarily, a first gradient corresponding to the image content loss can be determined based on the image content loss. Wherein, the first gradient is the change rate of the image content loss with respect to the parameters of the first generation network.

[0126] Step S2052: Determine a second gradient corresponding to the first discriminant loss based on the first discriminant loss.

[0127] Exemplarily, a second gradient corresponding to the first discriminant loss can be determined based on the first discriminant loss. Wherein, the second gradient is the change rate of the first discriminant loss with respect to the parameters of the first generation network.

[0128] Step S2053: Superimpose the first gradient and the second gradient, and use the superimposed gradient to update the network parameters of the first generation network to iteratively train the first generation network.

[0129] Step S2054: Determine the trained first generation network based on the first generation network with updated parameters.

[0130] Exemplarily, the first gradient and the second gradient can be superimposed, and the superimposed gradient is used to update the network parameters of the first generation network, thereby iteratively training the first generation network. That is, the first generation network is jointly trained by the image content loss and the first discriminant loss.

[0131] The training method of the image style transfer model provided by the embodiments of the present disclosure determines the first gradient corresponding to the image content loss based on the image content loss; determines the second gradient corresponding to the first discriminant loss based on the first discriminant loss; superimposes the first gradient and the second gradient, and uses the superimposed gradient to update the network parameters of the first generation network to iteratively train the first generation network; based on the first generation network with updated parameters, determines the trained first generation network; thus, it is possible to train the first generation network using the generation adversarial loss based on contrast learning, so as to achieve the purpose of image style transfer.

[0132] Figure 8 It is a schematic flowchart of an image style transfer method provided by the embodiments of the present disclosure. This embodiment can be applied to an electronic device, such as an in-vehicle information system device, such as Figure 8 shown, the method includes the following steps S801 - step S802.

[0133] Step S801, determine the image to be processed collected by the first camera.

[0134] Exemplarily, as Figure 1 shown, the RAW image collected by the first camera in the camera 11 can be obtained first as the image to be processed.

[0135] Step S802, process the image to be processed based on the image style transfer model to obtain a stylized image corresponding to the image to be processed;

[0136] wherein, the image style transfer model is trained by the training method of the above-mentioned image style transfer model.

[0137] Exemplarily, after training to obtain a trained image style transfer model, the image to be processed can be processed based on the trained image style transfer model to obtain a stylized image corresponding to the image to be processed; wherein, the stylized image can have the image style of the image collected by the second camera on the basis of maintaining the image content of the image to be processed.

[0138] The training method of the image style transfer model provided by the embodiments of the present disclosure includes determining a to-be-processed image collected by a first camera; processing the to-be-processed image based on the image style transfer model to obtain a stylized image corresponding to the to-be-processed image; wherein, the image style transfer model is trained by the training method of the above-mentioned image style transfer model; thus, since the image style transfer model is trained based on the image content loss and the first discriminant loss, in the embodiments of the present disclosure, the image style transfer model can be directly trained based on the generative adversarial loss of contrastive learning so as to achieve the purpose of image style transfer.

[0139] In some embodiments, processing the to-be-processed image based on the image style transfer model to obtain a stylized image corresponding to the to-be-processed image includes: performing a transformation process on the to-be-processed image based on a first generation network in the image style transfer model to obtain a second color correction matrix; performing an enhancement process on the to-be-processed image based on the second color correction matrix to obtain a stylized image corresponding to the to-be-processed image.

[0140] Exemplarily, a trained first generation network can be used to process the to-be-processed image to obtain a color correction matrix corresponding to the to-be-processed image; further, convolving the to-be-processed image with the color correction matrix corresponding to the to-be-processed image to obtain a stylized image corresponding to the to-be-processed image.

[0141] Exemplary device

[0142] Figure 9 A training device for an image style transfer model provided by the embodiments of the present disclosure, as Figure 9 shown, the training device 90 of the image style transfer model includes a first sample image determination module 91, a first enhanced image generation module 92, a content loss determination module 93, a first discriminant loss determination module 94, a first generation network training module 95, and a transfer model determination module 96.

[0143] The first sample image determination module 91 is configured to determine a first sample image collected by a first camera;

[0144] The first enhanced image generation module 92 is configured to perform an enhancement process on the first sample image based on the first generation network to obtain a first enhanced sample image corresponding to the first sample image;

[0145] The content loss determination module 93 is configured to determine an image content loss based on the first sample image and the first enhanced sample image;

[0146] The first discriminant loss determination module 94 is configured to perform a classification process on the first enhanced sample image by using a discriminant network to obtain a first discriminant loss;

[0147] The first generation network training module 95 is configured to iteratively train the first generation network based on the image content loss and the first discriminant loss, so as to obtain the trained first generation network;

[0148] The migration model determination module 96 is configured to obtain an image style migration model based on the trained first generation network.

[0149] In some embodiments, the training device 90 further includes a discriminant network training module, configured to determine a second sample image collected by a second camera; perform enhancement processing on the second sample image based on a second generation network to obtain a second enhanced sample image corresponding to the second sample image; use the discriminant network to perform classification processing on the first enhanced sample image and the second enhanced sample image to obtain a second discriminant loss; iteratively train the discriminant network based on the second discriminant loss to obtain the trained discriminant network.

[0150] In some embodiments, the training device 90 further includes a second generation network training module, configured to determine a target style image corresponding to the second sample image; determine an image error loss based on the second enhanced sample image and the target style image; iteratively train the second generation network based on the image error loss to obtain the trained second generation network.

[0151] In some embodiments, as Figure 10 shown, the content loss determination module 93 includes a block division unit 931, a feature extraction unit 932, and a content loss determination unit 933.

[0152] The block division unit 931 is configured to respectively perform image block division on the first enhanced sample image and the first sample image based on a preset division rule, so as to obtain a plurality of first image blocks corresponding to the first enhanced sample image and a plurality of second image blocks corresponding to the first sample image;

[0153] The feature extraction unit 932 is configured to respectively perform feature extraction on the plurality of first image blocks and the plurality of second image blocks based on a feature extraction network, so as to obtain first image block features corresponding to the respective first image blocks and second image block features corresponding to the respective second image blocks;

[0154] The content loss determination unit 933 is configured to determine the image content loss based on the respective first image block features and the respective second image block features.

[0155] In some embodiments, the content loss determination unit 933 is specifically configured to determine the position information of the first image block feature in the first enhanced sample image; based on the first sample image and the position information, determine a plurality of second target image block features from the plurality of second image block features; determine an image block content loss based on the first image block feature and the plurality of second target image block features; determine the image content loss based on the image block content losses corresponding to the respective first image blocks.

[0156] In some embodiments, the second generation network training module is specifically configured to determine the target style image corresponding to the second sample image; determine the difference between the second enhanced sample image and the target style image; determine the image error loss based on the difference; and iteratively train the second generation network based on the image error loss to obtain the trained second generation network.

[0157] In some embodiments, as Figure 11 shown, the first enhanced image generation module 92 includes a correction matrix generation unit 921 and a first enhanced image generation unit 922.

[0158] The correction matrix generation unit 921 is configured to perform transformation processing on the first sample image based on the first generation network to obtain a first color correction matrix.

[0159] The first enhanced image generation unit 922 is configured to perform enhancement processing on the first sample image based on the first color correction matrix to obtain the first enhanced sample image corresponding to the first sample image.

[0160] In some embodiments, the correction matrix generation unit 921 is specifically configured to perform residual processing on the first sample image based on at least one residual module in the first generation network to obtain the residual features corresponding to the first sample image; perform global pooling processing on the residual features to obtain the pooled residual features; and perform recombination on the pooled residual features to obtain the first color correction matrix.

[0161] In some embodiments, the first generation network training module 95 is specifically configured to determine the first gradient corresponding to the image content loss based on the image content loss; determine the second gradient corresponding to the first discriminant loss based on the first discriminant loss; superimpose the first gradient and the second gradient, and use the superimposed gradient to update the network parameters of the first generation network to iteratively train the first generation network; and determine the trained first generation network based on the first generation network with updated parameters.

[0162] For the beneficial technical effects corresponding to the exemplary embodiments of the above image style transfer model training device 90, reference can be made to the corresponding beneficial technical effects in the above exemplary method section, which will not be elaborated here.

[0163] Figure 12 For an image style transfer device provided in an embodiment of the present disclosure, as Figure 12 shown, the image style transfer device 120 includes an image determination module 121 and an image generation module 122.

[0164] The image determination module 121 is configured to determine the image to be processed collected by the first camera.

[0165] An image generation module 122 is configured to process a to-be-processed image based on an image style transfer model to obtain a stylized image corresponding to the to-be-processed image;

[0166] The image style transfer model is trained by the training method of the above image style transfer model.

[0167] In some embodiments, the image generation module 122 is specifically configured to perform transformation processing on the to-be-processed image based on a first generation network in the image style transfer model to obtain a second color correction matrix; and perform enhancement processing on the to-be-processed image based on the second color correction matrix to obtain a stylized image corresponding to the to-be-processed image.

[0168] For the beneficial technical effects corresponding to the exemplary embodiments of the above image style transfer device 120, reference may be made to the corresponding beneficial technical effects in the above exemplary method section, which will not be elaborated here.

[0169] Exemplary electronic device

[0170] Figure 13 The following is a structural diagram of an electronic device 130 provided by an embodiment of the present disclosure, including at least one processor 131 and a memory 132.

[0171] The processor 131 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 130 to perform desired functions.

[0172] The memory 132 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 131 may run one or more computer program instructions to implement the training method and / or image style transfer method of the image style transfer model in various embodiments of the present disclosure above and / or other desired functions.

[0173] In one example, the electronic device 130 may further include: an input device 133 and an output device 134, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0174] The input device 133 may further include, for example, a keyboard, a mouse, etc.

[0175] The output device 134 can output various information to the outside, and it can include, for example, a display, a speaker, a printer, a communication network, and remote output devices connected thereto, etc.

[0176] Of course, for simplicity, Figure 13 only some of the components related to the present disclosure in the electronic device 130 are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application scenarios, the electronic device 130 may further include any other appropriate components.

[0177] Exemplary computer program products and computer-readable storage media

[0178] In addition to the above methods and devices, embodiments of the present disclosure can also provide a computer program product, including computer program instructions, which, when run by a processor, cause the processor to execute the steps in the training method of the image style transfer model and / or the steps in the image style transfer method of various embodiments of the present disclosure described in the above "Exemplary Method" section.

[0179] The computer program product can be written in any combination of one or more programming languages to write program code for performing the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0180] Furthermore, embodiments of the present disclosure can also be a computer-readable storage medium, on which computer program instructions are stored, which, when run by a processor, cause the processor to execute the steps in the training method of the image style transfer model and / or the steps in the image style transfer method of various embodiments of the present disclosure described in the above "Exemplary Method" section.

[0181] A computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium includes, for example but not limited to, a system, device or component of electricity, magnetism, optics, electromagnetic, infrared ray, or semiconductor, or any combination of the above. More specific examples of the readable storage medium (an exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0182] The basic principles of the present disclosure have been described in conjunction with specific embodiments. However, the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that they are essential for each embodiment of the present disclosure. In addition, the above-described specific details are only for the purposes of illustration and easy understanding, rather than limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0183] Those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure also intends to include these changes and modifications.

Claims

1. A training method for an image style transfer model, comprising: Determine a first sample image collected by a first camera; Perform enhancement processing on the first sample image based on a first generation network to obtain a first enhanced sample image corresponding to the first sample image; Determine an image content loss based on the first sample image and the first enhanced sample image; Use a discriminative network to perform classification processing on the first enhanced sample image to obtain a first discriminative loss; Iteratively train the first generation network based on the image content loss and the first discriminative loss to obtain a trained first generation network; Obtain the image style transfer model based on the trained first generation network.

2. The method according to claim 1, wherein the discriminative network is trained through the following steps: Determine a second sample image collected by a second camera; Perform enhancement processing on the second sample image based on a second generation network to obtain a second enhanced sample image corresponding to the second sample image; Use the discriminative network to perform classification processing on the first enhanced sample image and the second enhanced sample image to obtain a second discriminative loss; Iteratively train the discriminative network based on the second discriminative loss to obtain a trained discriminative network.

3. The method according to claim 2, wherein the second generation network is trained through the following steps: Determine a target style image corresponding to the second sample image; Determine an image error loss based on the second enhanced sample image and the target style image; Iteratively train the second generation network based on the image error loss to obtain a trained second generation network.

4. The method according to claim 1, wherein The determining the image content loss based on the first sample image and the first enhanced sample image includes: Perform image block partitioning on the first enhanced sample image and the first sample image respectively based on a preset partitioning rule to obtain a plurality of first image blocks corresponding to the first enhanced sample image and a plurality of second image blocks corresponding to the first sample image; Perform feature extraction on the plurality of first image blocks and the plurality of second image blocks respectively based on a feature extraction network to obtain first image block features corresponding to each of the first image blocks and second image block features corresponding to each of the second image blocks; Determine the image content loss based on each of the first image block features and each of the second image block features.

5. The method according to claim 4, wherein, The determining the image content loss based on each of the first image block features and each of the second image block features includes: Determine the position information of the first image block feature in the first enhanced sample image; Based on the first sample image and the position information, determine a plurality of second target image block features from the plurality of second image block features; Determine an image block content loss based on the first image block feature and the plurality of second target image block features; Determine the image content loss based on the image block content loss corresponding to each of the first image blocks.

6. The method according to claim 3, wherein, The determining the image error loss based on the second enhanced sample image and the target style image includes: Determine the difference between the second enhanced sample image and the target style image; Based on the difference, determine the image error loss.

7. The method according to any one of claims 1-6, wherein, The enhancing the first sample image based on the first generation network to obtain a first enhanced sample image corresponding to the first sample image includes: Performing a transformation process on the first sample image based on the first generation network to obtain a first color correction matrix; Enhancing the first sample image based on the first color correction matrix to obtain a first enhanced sample image corresponding to the first sample image.

8. The method according to claim 7, wherein The performing a transformation process on the first sample image based on the first generation network to obtain a first color correction matrix includes: Performing a residual process on the first sample image based on at least one residual module in the first generation network to obtain a residual feature corresponding to the first sample image; Performing global pooling on the residual feature to obtain a pooled residual feature; Recombining the pooled residual feature to obtain the first color correction matrix.

9. The method according to any one of claims 1-6, wherein The iteratively training the first generation network based on the image content loss and the first discriminant loss to obtain a trained first generation network includes: Determining a first gradient corresponding to the image content loss based on the image content loss; Determining a second gradient corresponding to the first discriminant loss based on the first discriminant loss; Superposing the first gradient and the second gradient, and using the superposed gradient to update network parameters of the first generation network to iteratively train the first generation network; Determining a trained first generation network based on the first generation network with updated parameters.

10. An image style transfer method, including: Determining a to-be-processed image collected by a first camera; Processing the to-be-processed image based on an image style transfer model to obtain a stylized image corresponding to the to-be-processed image; wherein, the image style transfer model is trained by the method according to any one of claims 1-9.

11. The method according to claim 10, wherein, The processing the to-be-processed image based on the image style transfer model to obtain a stylized image corresponding to the to-be-processed image includes: Performing a transformation process on the to-be-processed image based on a first generation network in the image style transfer model to obtain a second color correction matrix; Enhancing the to-be-processed image based on the second color correction matrix to obtain a stylized image corresponding to the to-be-processed image.

12. A training device for an image style transfer model, including: A first sample image determination module, configured to determine a first sample image collected by a first camera; A first enhanced image generation module, configured to enhance the first sample image based on a first generation network to obtain a first enhanced sample image corresponding to the first sample image; A content loss determination module, configured to determine an image content loss based on the first sample image and the first enhanced sample image; A first discriminant loss determination module, configured to perform a classification process on the first enhanced sample image by using a discriminant network to obtain a first discriminant loss; The first generation network training module is configured to iteratively train the first generation network based on the image content loss and the first discriminant loss to obtain the trained first generation network; The transfer model determination module is configured to obtain the image style transfer model based on the trained first generation network.

13. An image style transfer device, comprising: An image determination module, configured to determine a to-be-processed image collected by a first camera; An image generation module, configured to process the to-be-processed image based on the image style transfer model to obtain a stylized image corresponding to the to-be-processed image; Wherein, the image style transfer model is trained based on the method according to any one of claims 1-9.

14. A computer-readable storage medium storing a computer program for executing the training method of the image style transfer model according to any one of claims 1-9 above, or the image style transfer method according to any one of claims 10-11 above.

15. An electronic device, the electronic device comprising: A processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the training method of the image style transfer model according to any one of claims 1-9 above, or the image style transfer method according to any one of claims 10-11 above.