Contrast enhancement method and device for shadow region, model training method and device and medium
Through the pre-trained contrast enhancement model and secondary supervision mechanism, the shadowed areas are accurately positioned and local contrast enhancement is solved, and the problem of insufficient recognition of shadowed areas in the prior art is improved, and the image display effect and model training efficiency are improved.
Patent Information
- Application Number
- CN202510837808.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-20
AI Technical Summary
Existing local contrast enhancement technologies cannot accurately identify and enhance shadowed areas, resulting in unnecessary adjustment of the contrast in non-shaded areas, affecting the naturalness of the image and the overall visual effect, and are prone to excessive enhancement or insufficient enhancement when dealing with complex scenes.
The pre-trained contrast enhancement model is adopted to accurately locate the shadowed area by segmenting mask generation model and image segmentation model, and use the image enhancement model to perform local contrast enhancement, while retaining the original contrast information of the non-shaded area, and optimizing model parameters in combination with the secondary supervision mechanism.
It improves the contrast of shadowed areas, enhances the image display effect and detail performance, reduces the calculation amount, solves the shortcomings of traditional global adaptive enhancement algorithms, and improves the efficiency and accuracy of model training.
Smart Images

Figure CN120374479A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of image processing, and particularly relates to a method for enhancing the contrast of shadow regions, a model training method, a device, and a medium. Background Art
[0002] With the continuous development of display technology, people's requirements for visual effects are constantly increasing. Whether watching videos, playing games, or engaging in other visual activities, people hope to obtain a more realistic, clear, and smooth picture effect. In many practical applications, images are often affected by uneven illumination, resulting in low contrast in shadow regions, loss of details, and poor display effects. However, most of the existing local contrast enhancement techniques adopt global adaptive enhancement methods. Although they can improve the overall contrast of the image to a certain extent, their effects are limited when processing specific local regions. These methods often cannot accurately identify and enhance shadow regions, resulting in unnecessary adjustment of the contrast in non-shadow regions, thereby affecting the naturalness and overall visual effect of the image. In addition, due to the lack of precise control over local regions, these techniques are prone to problems of over-enhancement or under-enhancement when processing complex scenes, and cannot meet users' requirements for high-quality images. Summary of the Invention
[0003] In view of the above problems, the present disclosure provides a method for enhancing the contrast of shadow regions, a model training method, a device, and a medium, aiming to specifically enhance the contrast of shadow regions in an image, thereby improving the image display effect and detail performance.
[0004] According to a first aspect of the present disclosure, there is provided a method for enhancing the contrast of shadow regions, including: Inputting a first image into a pre-trained contrast enhancement model, where the first image includes a first shadow region; Outputting a second image by the contrast enhancement model, where the second image includes a second shadow region corresponding to the first shadow region, and the contrast of the second shadow region is higher than that of the first shadow region. Among them, the contrast enhancement model generates a segmentation mask for the first shadow region, and performs local contrast enhancement on the first shadow region based on the segmentation mask while retaining the original contrast information of the non-shadow region.
[0005] Optionally, the contrast enhancement model includes a segmentation mask generation model, an image segmentation model, and an image enhancement model. The outputting of the second image by the contrast enhancement model includes: Performing a binarization operation on the first image by the segmentation mask generation model to generate a segmentation mask for the first shadow region; The dilation convolution operation is performed on the segmentation mask of the first shadow region by the segmentation mask generation model to obtain the dilated segmentation mask of the first shadow region; The first image is segmented by the image segmentation model based on the segmentation mask of the first shadow region to obtain the first segmentation map of the first shadow region; The first image is segmented by the image segmentation model based on the dilated segmentation mask of the first shadow region to obtain the second segmentation map of the first shadow region; The second image is obtained by the image enhancement model by mosaicking the first segmentation map, the second segmentation map, and the first image.
[0006] Optionally, the step of the image enhancement model mosaicking the first segmentation map, the second segmentation map, and the first image to obtain the second image includes: The image enhancement model performs image enhancement operations on the first segmentation map, the second segmentation map, and the first image respectively, and mosaics the first segmentation map, the second segmentation map, and the first image after the image enhancement operations; An image enhancement operation is performed on the image obtained by image mosaicking to obtain the second image.
[0007] According to a second aspect of the present disclosure, a method for training a contrast enhancement model for a shadow region is provided, including: Obtain a training data set, where the training data set includes a plurality of training samples and shadow region labels corresponding to the training samples; Input the training samples into a target contrast enhancement model, where the target contrast enhancement model includes a target segmentation mask generation model, a target image segmentation model, and a target image enhancement model; Adopt a secondary supervision mechanism to calculate the loss functions of the target segmentation mask generation model and the target image enhancement model respectively; Use the loss functions for backpropagation learning to optimize the model parameters of the target segmentation mask generation model, the target image segmentation model, and the target image enhancement model.
[0008] Optionally, the step of adopting a secondary supervision mechanism to calculate the loss functions of the target segmentation mask generation model and the target image enhancement model respectively includes: Input the training samples into the target segmentation mask generation model, and the target segmentation mask generation model performs a binarization operation on the training samples to generate a segmentation mask of the shadow region in the training samples; Based on the segmentation mask of the shadow region in the training sample and the corresponding shadow region label of the training sample, the loss function of the target segmentation mask generation model is calculated.
[0009] Optionally, the adopting the secondary supervision mechanism to calculate the loss functions of the target segmentation mask generation model and the target image enhancement model respectively further includes: The target segmentation mask generation model performs a dilation convolution operation on the segmentation mask of the shadow region in the training sample to obtain a dilated segmentation mask of the shadow region in the training sample; The target image segmentation model segments the training sample based on the segmentation mask of the shadow region in the training sample to obtain a first segmentation map of the shadow region in the training sample; The target image segmentation model segments the training sample based on the dilated segmentation mask of the shadow region in the training sample to obtain a second segmentation map of the shadow region in the training sample; The target image enhancement model splices the first segmentation map, the second segmentation map and the training sample to obtain a target image of the training sample; Based on the target image of the training sample and the corresponding shadow region label of the training sample, the loss function of the target image enhancement model is calculated.
[0010] Optionally, the step that the target image enhancement model splices the first segmentation map, the second segmentation map and the training sample to obtain a target image of the training sample includes: The target image enhancement model respectively performs image enhancement operations on the first segmentation map, the second segmentation map and the training sample, and splices the first segmentation map, the second segmentation map and the training sample after the image enhancement operations; An image enhancement operation is performed on the image obtained by image splicing to obtain a target image of the training sample.
[0011] According to the third aspect of the present disclosure, there is provided a contrast enhancement device for a shadow region, including: An image input unit, configured to input a first image into a pre-trained contrast enhancement model, where the first image includes a first shadow region; A contrast enhancement unit, which outputs a second image by the contrast enhancement model, where the second image includes a second shadow region corresponding to the first shadow region, and the contrast of the second shadow region is higher than the contrast of the first shadow region. Wherein, the contrast enhancement model generates a segmentation mask of the first shadow region, and performs local contrast enhancement on the first shadow region based on the segmentation mask, while retaining the original contrast information of the non-shadow region.
[0012] According to a fourth aspect of the present disclosure, there is provided a training device for a contrast enhancement model of a shadow region, including: A training data acquisition unit configured to acquire a training data set, where the training data set includes a plurality of training samples and shadow region labels corresponding to the training samples; A training sample input unit configured to input the training samples into a target contrast enhancement model, where the target contrast enhancement model includes a target segmentation mask generation model, a target image segmentation model, and a target image enhancement model; A loss function calculation unit configured to calculate loss functions of the target segmentation mask generation model and the target image enhancement model respectively by adopting a secondary supervision mechanism; A model parameter tuning unit configured to perform backpropagation learning by using the loss functions to tune model parameters of the target segmentation mask generation model, the target image segmentation model, and the target image enhancement model.
[0013] According to a fifth aspect of the present disclosure, there is provided an electronic device, including: a processor, a memory, and a program stored on the memory and executable on the processor, where when the program is executed by the processor, the steps of the method described above are implemented.
[0014] According to a sixth aspect of the present disclosure, there is provided a storage medium having a computer program or instruction stored thereon, where when the computer program or instruction is executed by a processor, the steps of the method described above are implemented.
[0015] According to a seventh aspect of the present disclosure, there is provided a chip, including: The contrast enhancement device described above to implement the method.
[0016] According to a seventh aspect of the present disclosure, there is provided a chip, including: The training device described above to implement the method.
[0017] The present disclosure brings the following beneficial effects: The contrast enhancement method for the shadow area provided by the present disclosure inputs the first image into a pre-trained contrast enhancement model. The contrast enhancement model generates a segmentation mask for the first shadow area of the first image, performs local contrast enhancement on the first shadow area based on the segmentation mask, and at the same time preserves the original contrast information of the non-shadow area. The second image output by the contrast enhancement model includes a second shadow area corresponding to the first shadow area, and the contrast of the second shadow area is higher than that of the first shadow area. The non-shadow area of the second image preserves the contrast information of the corresponding non-shadow area of the first image. In this way, by accurately locating the shadow area and specifically enhancing the contrast of the shadow area, the problem that the contrast of the non-shadow area is unnecessarily adjusted by the traditional global adaptive enhancement algorithm is effectively solved, the image display effect and detail performance are improved, and the calculation amount is reduced.
[0018] The training method for the contrast enhancement model of the shadow area provided by the present disclosure obtains a training data set. The training data set includes multiple training samples and shadow area labels corresponding to the training samples. The training samples are input into a target contrast enhancement model. The target contrast enhancement model includes a target segmentation mask generation model, a target image segmentation model, and a target image enhancement model. A secondary supervision mechanism is adopted to calculate the loss functions of the target segmentation mask generation model and the target image enhancement model respectively, and use the loss functions for backpropagation learning to optimize the model parameters of the target segmentation mask generation model, the target image segmentation model, and the target image enhancement model. In this way, by introducing the secondary supervision mechanism, the collaborative optimization of the segmentation mask generation and image enhancement processes is realized, and the model training efficiency and accuracy are improved.
[0019] Other features and advantages of the present disclosure will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure are achieved and obtained by the structures specifically pointed out in the specification and the drawings.
[0020] To make the above objectives, features, and advantages of the present disclosure more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings
[0021] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above and other objectives, features, and advantages of the present disclosure will become clearer. In the drawings: Figure 1 It is a schematic flowchart of the contrast enhancement method for the shadow area provided by an embodiment of the present disclosure; Figure 2A It is a block diagram of the algorithm of the contrast enhancement model provided by an embodiment of the present disclosure; Figure 2BSchematic diagram of the loss function of the target segmentation mask generation model provided according to an embodiment of the present disclosure; Figure 2C Schematic diagram of the loss function of the target image enhancement model provided according to an embodiment of the present disclosure; Figure 3 Flow schematic diagram of the training method of the contrast enhancement model for the shadow area provided according to an embodiment of the present disclosure; Figure 4 Structural schematic diagram of the contrast enhancement device for the shadow area provided according to an embodiment of the present disclosure; Figure 5 Structural schematic diagram of the training device of the contrast enhancement model for the shadow area provided according to an embodiment of the present disclosure; Figure 6 Structural schematic diagram of the electronic device provided according to an embodiment of the present disclosure. Detailed implementation manners
[0022] Various embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. In each of the drawings, the same elements are denoted by the same or similar reference numerals. For clarity, the various parts in the drawings are not drawn to scale.
[0023] The following terms are used herein: The shadow area refers to, for the application scenario displayed by the display panel, the shadow area is, for example, an area in the image where the local brightness is significantly lower than the surrounding area due to an object blocking the light source (such as the sun, light, etc.). The shadow area may include self-shadows (shadow areas formed by the object's own structure blocking light) and projection shadows (shadows projected by the object onto other surfaces), etc.
[0024] Contrast refers to the degree of brightness difference between adjacent regions or objects in an image, and is usually quantified and characterized by the ratio or difference between the maximum brightness value and the minimum brightness value. In the field of digital image processing, contrast reflects the visibility and layering of image details. A higher contrast can make the edges, textures, and structural features in the image more clearly distinguishable, while low-contrast areas often appear gray and blurred. Especially in the shadow area, due to the brightness compression effect caused by light attenuation, the contrast value of this area will be significantly reduced, making it difficult to identify the detailed information.
[0025] Figure 1 Flow schematic diagram of the contrast enhancement method for the shadow area provided according to an embodiment of the present disclosure. As Figure 1 shown, the contrast enhancement method of the embodiment of the present disclosure includes: In step S110, the first image is input into the pre-trained contrast enhancement model, and the first image includes a first shadow area.
[0026] In some embodiments, the first image may be a photographic picture of the real world or a two-dimensional image of a three-dimensional model in a virtual environment. It should be noted that the virtual environment is the virtual environment displayed (or provided) when the application runs on the terminal. The virtual environment may be a three-dimensional virtual environment or a two-dimensional virtual environment. The three-dimensional virtual environment may be a simulation environment of the real world, a semi-simulated and semi-fictional environment, or a purely fictional environment. The first image includes a first shadow area. In real life, an object will form a shadow under sunlight (the shadow area and the shaded area in this disclosure are equivalent concepts).
[0027] In some embodiments, the contrast enhancement model is a pre-trained algorithm model for locally enhancing the contrast of the shadow area in the first image while retaining the original contrast information of the non-shadow area. In some embodiments, the first image is input into the pre-trained contrast enhancement model to facilitate the use of the contrast enhancement model to perform the contrast enhancement method of the embodiments of the present disclosure.
[0028] In step S120, a second image is output by the contrast enhancement model. The second image includes a second shadow area corresponding to the first shadow area, and the contrast of the second shadow area is higher than that of the first shadow area. Wherein, the contrast enhancement model generates a segmentation mask of the first shadow area, and locally enhances the contrast of the first shadow area based on the segmentation mask while retaining the original contrast information of the non-shadow area.
[0029] Figure 2A The algorithm block diagram of the contrast enhancement model according to an embodiment of the present disclosure is shown below. Figure 2A The contrast enhancement method of the embodiments of the present disclosure will be described in detail below. As Figure 2AAs shown, in some embodiments, the contrast enhancement model includes a segmentation mask generation model 210, an image segmentation model 220, and an image enhancement model 230. In some embodiments, the segmentation mask generation model 210 generates a segmentation mask for the first shadow region. The image segmentation model 220 cuts out the first shadow region from the first image based on the segmentation mask. The image enhancement model 230 performs local contrast enhancement on the first shadow region specifically while preserving the original contrast information of the non-shadow region. In some embodiments, the segmentation mask generation model 210 includes a UNet deep neural network 211. The UNet deep neural network 211 constructs a feature extraction layer by stacking multiple residual blocks (ResBlock) to perform feature extraction on the input image. These residual blocks adopt a skip connections structure, enhancing the network's learning ability while alleviating the vanishing gradient problem in deep networks. In some embodiments, the first image is input into the UNet deep neural network 211, and the UNet deep neural network 211 performs a binarization operation on the first image to generate a segmentation mask (mask) for the first shadow region. The first image can be converted into a grayscale image for binarization processing, and the grayscale image is converted into a binary image (i.e., the segmentation mask for the first shadow region) using the threshold binarization method. The segmentation mask for the first shadow region can achieve the division of the shadow region and the non-shadow region in the first image, thereby clearly separating the shadow region. For example, the grayscale value of the pixels in the shadow region of the first image is 1, and the grayscale value of the pixels in the non-shadow region is 0. In some embodiments, the segmentation mask generation model 210 further includes a dilated convolution block (Dilated Block) 212. The segmentation mask for the first shadow region is input into the dilated convolution block 212, and the dilated convolution block 212 performs a dilated convolution operation on the segmentation mask for the first shadow region to obtain a dilated segmentation mask for the first shadow region. Through dilated convolution, the receptive field can be expanded to capture more extensive context information. The boundary range of the segmentation mask for the first shadow region is smaller than the boundary range of the dilated segmentation mask for the first shadow region, and the dilated segmentation mask for the first shadow region includes the surrounding associated regions outside the segmentation mask for the first shadow region. In some embodiments, the image segmentation model 220 includes a first segmentation layer 221 and a second segmentation layer 222 constructed by, for example, a multiplier. In some embodiments, the segmentation mask for the first shadow region and the first image are input into the first segmentation layer 221, and the first segmentation layer 221 segments the first image based on the segmentation mask for the first shadow region to obtain a first segmentation map of the first shadow region. The dilated segmentation mask for the first shadow region and the first image are input into the second segmentation layer 222, and the second segmentation layer 222 segments the first image based on the dilated segmentation mask for the first shadow region to obtain a second segmentation map of the first shadow region.In some embodiments, the first segmentation map, the second segmentation map, and the first image are stitched by the image enhancement model 230 to obtain a second image. The second image includes a second shadow region corresponding to the first shadow region, and the contrast of the second shadow region is higher than that of the first shadow region. The image enhancement model 230 includes an original image enhancement layer 231, a segmentation mask image enhancement layer 232, an extended segmentation mask image enhancement layer 233, an image stitching layer 234, and a stitched image enhancement layer 235. The original image enhancement layer 231 is constructed by, for example, residual blocks and is used to enhance the first image. The segmentation mask image enhancement layer 232 is constructed by, for example, residual blocks and is used to perform an image enhancement operation on the first segmentation map of the first shadow region of the first image. The extended segmentation mask image enhancement layer 233 is constructed by, for example, residual blocks and is used to perform an image enhancement operation on the second segmentation map of the first shadow region of the first image. The first segmentation map, the second segmentation map, and the first image after the image enhancement operation are stitched by the image stitching layer 234. The stitched image enhancement layer 235 performs an image enhancement operation on the image obtained by the image stitching to obtain the second image.
[0030] Figure 3 FIG. is a schematic flowchart of a method for training a contrast enhancement model for a shadow region according to an embodiment of the present disclosure. As Figure 3 shown, the training method of the embodiment of the present disclosure includes: In step S310, a training data set is obtained, and the training data set includes a plurality of training samples and shadow region labels corresponding to the training samples.
[0031] In some embodiments, the training data set can be obtained from image samples collected from a public data set or an actual application scenario. The training samples contain shadow region bounding boxes and pixel-level shadow region labels marked by a professional annotation tool. The label data can adopt a dual-channel annotation method of RGB color space and grayscale mask, where the first channel records the color information of the original image, and the second channel marks the exact position of the shadow region through binarization processing.
[0032] In step S320, the training samples are input into a target contrast enhancement model, and the target contrast enhancement model includes a target segmentation mask generation model, a target image segmentation model, and a target image enhancement model.
[0033] In some embodiments, the target contrast enhancement model can refer to Figure 2A the contrast enhancement model. When the target contrast enhancement model is trained to meet the training completion condition, the target contrast enhancement model is the Figure 2A contrast enhancement model.
[0034] In step S330, a secondary supervision mechanism is adopted to calculate the loss functions of the target segmentation mask generation model and the target image enhancement model respectively.
[0035] In some embodiments, the training samples are input into the target segmentation mask generation model, and the target segmentation mask generation model performs a binarization operation on the training samples to generate a segmentation mask of the shadow region in the training samples. Figure 2B It is a schematic diagram of the loss function of the target segmentation mask generation model provided according to an embodiment of the present disclosure. As Figure 2B shown, based on the segmentation mask of the shadow region in the training samples and the corresponding shadow region labels of the training samples, the loss function Lmask of the target segmentation mask generation model is calculated. In some embodiments, the target segmentation mask generation model performs a dilation convolution operation on the segmentation mask of the shadow region in the training samples to obtain a dilated segmentation mask of the shadow region in the training samples. The target image segmentation model segments the training samples based on the segmentation mask of the shadow region in the training samples to obtain a first segmentation map of the shadow region in the training samples. The target image segmentation model segments the training samples based on the dilated segmentation mask of the shadow region in the training samples to obtain a second segmentation map of the shadow region in the training samples. The target image enhancement model stitches together the first segmentation map and the second segmentation map of the shadow region in the training samples, and the training samples to obtain the target image of the training samples. In some embodiments, the target image enhancement model performs image enhancement operations on the first segmentation map and the second segmentation map of the shadow region in the training samples, and the training samples respectively, stitches together the first segmentation map and the second segmentation map of the shadow region in the training samples after the image enhancement operation, and the training samples, and performs an image enhancement operation on the image obtained by the image stitching to obtain the target image of the training samples. Figure 2C It is a schematic diagram of the loss function of the target image enhancement model provided according to an embodiment of the present disclosure. As Figure 2C shown, in some embodiments, based on the target image of the training samples and the corresponding shadow region labels of the training samples, the loss function Limage of the target image enhancement model is calculated.
[0036] In step S340, using the loss function for backpropagation learning to optimize the model parameters of the target segmentation mask generation model, the target image segmentation model, and the target image enhancement model.
[0037] In some embodiments, the loss function Lmask of the target segmentation mask generation model and the loss function Limage of the target image enhancement model are jointly optimized. For example, specifically, the total loss function Ltotal = α·Lmask + β·Limage can be constructed by weighted summation, where α and β are hyperparameters set according to task requirements. The gradient of the total loss function is transmitted to the entire network architecture through the backpropagation algorithm, and the Adam optimizer is used to synchronously update the UNet network parameters in the target segmentation mask generation model, the convolution kernel parameters of the dilated convolution block, and the parameters of each residual block in the target image enhancement model. During the training process, a dynamic learning rate adjustment strategy is set. When the validation set loss does not show a significant decrease for three consecutive training epochs, the learning rate is automatically multiplied by the decay factor 0.8 until the preset upper limit of the training rounds is reached or the loss function converges within the threshold range.
[0038] It can be understood that after the model training is completed, the contrast enhancement method of the embodiments of the present disclosure realizes fine processing through a multi-scale feature fusion strategy to improve the image enhancement effect. Specifically, the original segmentation mask and the dilated segmentation mask output by the segmentation mask generation model correspond to regional features in different receptive field ranges: the original segmentation mask retains accurate shadow boundary information, while the dilated segmentation mask captures the light transition zone features around the shadow area by adjusting the dilation rate parameter (usually set to 2-4 times). The image segmentation model performs a per-pixel multiplication operation on these two masks and the original image to obtain a two-scale segmentation map including the core shadow area and the edge transition area. The image enhancement model integrates the feature vectors of the two-scale segmentation map and the original image through channel concatenation at the splicing layer, enabling the enhancement operation to more comprehensively consider the local and global information of the image, thereby improving the local contrast while maintaining the overall naturalness of the image.
[0039] Figure 4 The structural schematic diagram of a contrast enhancement device for a shadow area provided according to an embodiment of the present disclosure is shown. As Figure 4 The contrast enhancement device 400 for the shadow area shown includes an image input unit 410 and a contrast enhancement unit 420.
[0040] The image input unit 410 is configured to input a first image into a pre-trained contrast enhancement model, and the first image includes a first shadow area.
[0041] The contrast enhancement unit 420 outputs a second image from the contrast enhancement model. The second image includes a second shadow region corresponding to the first shadow region, and the contrast of the second shadow region is higher than that of the first shadow region. Wherein, the contrast enhancement model generates a segmentation mask of the first shadow region, and performs local contrast enhancement on the first shadow region based on the segmentation mask while preserving the original contrast information of the non-shadow region.
[0042] Since the specific process of enhancing the contrast of the shadow region has been described in detail above, it will not be elaborated here.
[0043] Figure 5 The structural schematic diagram of a training device for a contrast enhancement model of a shadow region provided according to an embodiment of the present disclosure is shown. As Figure 5 shown, the training device 500 for the contrast enhancement model of the shadow region includes a training data acquisition unit 510, a training sample input unit 520, a loss function calculation unit 530, and a model parameter tuning unit 540.
[0044] The training data acquisition unit 510 is configured to acquire a training data set, and the training data set includes a plurality of training samples and shadow region labels corresponding to the training samples.
[0045] The training sample input unit 520 is configured to input the training samples into a target contrast enhancement model, and the target contrast enhancement model includes a target segmentation mask generation model, a target image segmentation model, and a target image enhancement model.
[0046] The loss function calculation unit 530 is configured to calculate the loss functions of the target segmentation mask generation model and the target image enhancement model respectively by adopting a quadratic supervision mechanism.
[0047] The model parameter tuning unit 540 is configured to perform backpropagation learning by using the loss function to tune the model parameters of the target segmentation mask generation model, the target image segmentation model, and the target image enhancement model.
[0048] Since the specific process of training the contrast enhancement model has been described in detail above, it will not be elaborated here.
[0049] An embodiment of the present disclosure further provides an electronic device 600, as Figure 6 shown, including a memory 620, a processor 610, a power supply component 630, a network interface 640, an input / output interface 650, and a program stored in the memory 620 and executable on the processor 610. When the program is executed by the processor 610, it can implement the various processes of the above embodiments of the method and achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0050] The embodiments of the present disclosure further provide a chip, including Figure 4 the contrast enhancement device 400 shown in the figure, which implements the steps of the method as described above. The chip here includes general-purpose processors (such as CPU, GPU), mobile device main processors (AP), programmable logic chips (such as FPGA), and application-specific integrated circuits (such as ASIC), etc. The beneficial effects that can be achieved by the method provided by the embodiments of the present disclosure can be realized. For details, please refer to the previous embodiments and will not be elaborated here.
[0051] The embodiments of the present disclosure further provide a chip, including Figure 5 the training device 500 shown in the figure, which implements the steps of the method as described above. The chip here includes general-purpose processors (such as CPU, GPU), mobile device main processors (AP), programmable logic chips (such as FPGA), and application-specific integrated circuits (such as ASIC), etc. The beneficial effects that can be achieved by the method provided by the embodiments of the present disclosure can be realized. For details, please refer to the previous embodiments and will not be elaborated here.
[0052] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling related hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. For this reason, the embodiments of the present disclosure further provide a storage medium, on which a computer program or instructions are stored. When the computer program or instructions are executed by a processor, the various processes of the above embodiments in the method can be realized.
[0053] Since the instructions stored in the storage medium can execute the steps in the method provided by the embodiments of the present disclosure, the beneficial effects that can be achieved by the method provided by the embodiments of the present disclosure can be realized. For details, please refer to the previous embodiments and will not be elaborated here. The specific implementation of each of the above operations can be referred to the previous embodiments and will not be elaborated here.
[0054] In summary, according to the embodiments of the present disclosure, the first image is input into the pre-trained contrast enhancement model. The contrast enhancement model generates a segmentation mask of the first shadow region of the first image. Based on the segmentation mask, local contrast enhancement is performed on the first shadow region while retaining the original contrast information of the non-shadow region. The second image output by the contrast enhancement model includes a second shadow region corresponding to the first shadow region. The contrast of the second shadow region is higher than that of the first shadow region. The non-shadow region of the second image retains the contrast information of the corresponding non-shadow region of the first image. In this way, by accurately locating the shadow region and specifically enhancing the contrast of the shadow region, the problem that the contrast of the non-shadow region is unnecessarily adjusted by the traditional global adaptive enhancement algorithm is effectively solved, the image display effect and detail performance are improved, and the calculation amount is reduced.
[0055] According to an embodiment of the present disclosure, a training data set is obtained. The training data set includes a plurality of training samples and the shadow region labels corresponding to the training samples. The training samples are input into a target contrast enhancement model. The target contrast enhancement model includes a target segmentation mask generation model, a target image segmentation model, and a target image enhancement model. A secondary supervision mechanism is adopted to calculate the loss functions of the target segmentation mask generation model and the target image enhancement model respectively, and reverse learning is performed using the loss functions to optimize the model parameters of the target segmentation mask generation model, the target image segmentation model, and the target image enhancement model. In this way, by introducing the secondary supervision mechanism, the collaborative optimization of the segmentation mask generation and image enhancement processes is realized, and the model training efficiency and accuracy are improved.
[0056] Finally, it should be noted that: Obviously, the above embodiments are merely examples for clearly illustrating the present disclosure, rather than limiting the implementation manners. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or variations derived therefrom still fall within the protection scope of the present disclosure.
Claims
1. A method for enhancing the contrast of a shadow region, comprising: Inputting a first image into a pre-trained contrast enhancement model, wherein the first image includes a first shadow region; Outputting a second image by the contrast enhancement model, the second image including a second shadow region corresponding to the first shadow region, and the contrast of the second shadow region being higher than that of the first shadow region. Wherein, the contrast enhancement model generates a segmentation mask of the first shadow region, performs local contrast enhancement on the first shadow region based on the segmentation mask, and simultaneously retains the original contrast information of the non-shadow region.
2. The contrast enhancement method according to claim 1, wherein, The contrast enhancement model includes a segmentation mask generation model, an image segmentation model, and an image enhancement model. The outputting of the second image by the contrast enhancement model includes: Performing a binarization operation on the first image by the segmentation mask generation model to generate a segmentation mask of the first shadow region; Performing a dilated convolution operation on the segmentation mask of the first shadow region by the segmentation mask generation model to obtain a dilated segmentation mask of the first shadow region; Segmenting the first image by the image segmentation model based on the segmentation mask of the first shadow region to obtain a first segmentation map of the first shadow region; Segmenting the first image by the image segmentation model based on the dilated segmentation mask of the first shadow region to obtain a second segmentation map of the first shadow region; Performing image stitching on the first segmentation map, the second segmentation map, and the first image by the image enhancement model to obtain the second image.
3. The contrast enhancement method according to claim 2, wherein The performing image stitching on the first segmentation map, the second segmentation map, and the first image by the image enhancement model to obtain the second image includes: Performing image enhancement operations on the first segmentation map, the second segmentation map, and the first image respectively by the image enhancement model, and performing image stitching on the first segmentation map, the second segmentation map, and the first image after the image enhancement operations; Performing an image enhancement operation on the image obtained by image stitching to obtain the second image.
4. A method for training a contrast enhancement model for a shadow region, comprising: Obtaining a training data set, the training data set including a plurality of training samples and shadow region labels corresponding to the training samples; Inputting the training samples into a target contrast enhancement model, the target contrast enhancement model including a target segmentation mask generation model, a target image segmentation model, and a target image enhancement model; Adopting a secondary supervision mechanism to calculate the loss functions of the target segmentation mask generation model and the target image enhancement model respectively; Performing backpropagation learning using the loss functions to optimize the model parameters of the target segmentation mask generation model, the target image segmentation model, and the target image enhancement model.
5. The training method according to claim 4, wherein The adopting a secondary supervision mechanism to calculate the loss functions of the target segmentation mask generation model and the target image enhancement model respectively includes: Input the training sample into the target segmentation mask generation model, and the target segmentation mask generation model performs a binarization operation on the training sample to generate a segmentation mask of the shadow area in the training sample; Based on the segmentation mask of the shadow area in the training sample and the corresponding shadow area label of the training sample, calculate the loss function of the target segmentation mask generation model.
6. The training method according to claim 5, wherein, The step of using the secondary supervision mechanism to calculate the loss functions of the target segmentation mask generation model and the target image enhancement model respectively further includes: The target segmentation mask generation model performs a dilated convolution operation on the segmentation mask of the shadow area in the training sample to obtain a dilated segmentation mask of the shadow area in the training sample; The target image segmentation model segments the training sample based on the segmentation mask of the shadow area in the training sample to obtain a first segmentation map of the shadow area in the training sample; The target image segmentation model segments the training sample based on the dilated segmentation mask of the shadow area in the training sample to obtain a second segmentation map of the shadow area in the training sample; The target image enhancement model splices the first segmentation map, the second segmentation map and the training sample to obtain a target image of the training sample; Based on the target image of the training sample and the corresponding shadow area label of the training sample, calculate the loss function of the target image enhancement model.
7. The training method according to claim 6, wherein, The step that the target image enhancement model splices the first segmentation map, the second segmentation map and the training sample to obtain a target image of the training sample includes: The target image enhancement model performs image enhancement operations on the first segmentation map, the second segmentation map and the training sample respectively, and splices the first segmentation map, the second segmentation map and the training sample after the image enhancement operations; Perform an image enhancement operation on the image obtained by image splicing to obtain a target image of the training sample.
8. A contrast enhancement device for a shadow area, comprising: An image input unit for inputting a first image into a pre-trained contrast enhancement model, where the first image includes a first shadow area; A contrast enhancement unit, where the contrast enhancement model outputs a second image, the second image includes a second shadow area corresponding to the first shadow area, and the contrast of the second shadow area is higher than that of the first shadow area. Among them, the contrast enhancement model generates a segmentation mask of the first shadow area, performs local contrast enhancement on the first shadow area based on the segmentation mask, and at the same time retains the original contrast information of the non-shadow area.
9. A training device for a contrast enhancement model of a shadow area, comprising: A training data acquisition unit for acquiring a training data set, where the training data set includes a plurality of training samples and shadow area labels corresponding to the training samples; A training sample input unit for inputting the training samples into a target contrast enhancement model, where the target contrast enhancement model includes a target segmentation mask generation model, a target image segmentation model, and a target image enhancement model; A loss function calculation unit for calculating the loss functions of the target segmentation mask generation model and the target image enhancement model respectively by adopting a quadratic supervision mechanism; A model parameter tuning unit for performing backpropagation learning by using the loss functions to tune the model parameters of the target segmentation mask generation model, the target image segmentation model, and the target image enhancement model.
10. An electronic device, comprising: A processor, a memory, and a program stored on the memory and executable on the processor, where when the program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A storage medium having a computer program or instruction stored thereon, where when the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
12. A chip, comprising: The contrast enhancement device according to claim 8, implementing the method according to any one of claims 1 to 3.
13. A chip, comprising: The training device according to claim 9, implementing the method according to any one of claims 4 to 7.
Citation Information
Patent Citations
Floor tile image shadow removal method and device, computer equipment and storage medium
CN115546073A
Semantic segmentation method and device and storage medium
CN117523560A
Image shadow removing method and device based on diffusion model, equipment and medium
CN118691497A
Shadow removing method and system based on diffusion, segmentation and super-resolution model
CN119048357A
Shadow recovery method, model training method, system, medium, product and equipment
CN119399035A