A method, device, equipment and medium for determining a shadow model

Through the design of directional encoding and decoder, combined with occlusion patches and shadowless images, the preset shadow model is corrected, which solves the problem of inefficient shadow removal in the existing technology and achieves efficient shadow removal effect.

CN117830161BActive Publication Date: 2025-05-30SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410010903.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-04
Publication Date
2025-05-30
Estimated Expiration
2044-01-04

AI Technical Summary

Technical Problem

In the prior art, when shadow removal, the calculation cost of non-shaded areas is high, resulting in inefficiency.

Method used

By obtaining the triple set of images to be trained and the preset shadow model, the occlusion patch and non-occlusion patch in the shadow image are determined, and the occlusion patch is encoded using the encoder to directionally, and the preset shadow model is corrected to improve the shadow removal effect.

Benefits of technology

By directed encoding of occlusion patches, the efficiency of shadow area processing is improved, the calculation cost of non-shaded areas is reduced, and the performance of shadow removal is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117830161B_ABST
    Figure CN117830161B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, device and medium for determining a shadow model. The method includes: obtaining a set of training image triples and a preset shadow model, where each training image triple in the set of training image triples includes a shadow image, a shadow mask and a shadowless image; for the training image triples in the set of training image triples, determining the occluded patches and non-occluded patches in the shadow image according to the shadow image and the shadow mask; processing the occluded patches by an encoder in the preset shadow model to obtain an encoding result; determining a shadow removal result and a loss function corresponding to the shadow removal result according to the encoding result, the non-occluded patches, the shadowless image and a decoder in the preset shadow model; and correcting the preset shadow model according to the loss function to determine the final shadow model. By performing directional encoding on the occluded patches, the efficiency of processing the shadow area is ensured, and the performance of shadow removal is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a method, apparatus, device and medium for determining a shadow model. Background Art

[0002] Shadow removal is an important issue in the field of computer vision because shadows can lead to the loss of image details and scene inconsistencies. Traditional shadow removal methods mainly rely on manually designed priors and image attributes such as gradients, regions, and illumination. In recent years, deep learning methods such as convolutional neural networks (CNNs) and Transformer models have made significant progress in the field of shadow removal. However, existing methods usually involve a large amount of computation when dealing with non-shadow regions outside the shadow area, resulting in low efficiency.

[0003] In recent research, many methods use CNNs to achieve shadow removal, such as ST-GAN, BMNet, etc. These methods improve the performance of shadow removal by learning data representations, but usually need to process the entire input image, resulting in a high computational cost. Recently, some methods have tried to use Transformer models to aggregate global context information to assist shadow removal. However, similar to CNN methods, their computational cost in non-shadow regions is still high. Summary of the Invention

[0004] The present invention provides a method, apparatus, device and medium for determining a shadow model to improve the effect of the shadow model on shadow removal.

[0005] According to a first aspect of the present invention, there is provided a method for determining a shadow model, including:

[0006] Obtaining a set of training image triples and a preset shadow model, where each training image triple in the set of training image triples includes a shadow image, a shadow mask, and a shadowless image;

[0007] For the training image triples in the set of training image triples, determining occluded patches and non-occluded patches in the shadow image according to the shadow image and the shadow mask;

[0008] Processing the occluded patches according to an encoder in the preset shadow model to obtain an encoding result;

[0009] Determining a shadow removal result and a loss function corresponding to the shadow removal result according to the encoding result, the non-occluded patches, the shadowless image, and a decoder in the preset shadow model;

[0010] Correcting the preset shadow model according to the loss function to determine a final shadow model.

[0011] According to a second aspect of the present invention, there is provided a shadow model determination device, including:

[0012] A model acquisition module, configured to acquire a set of training image triples and a preset shadow model, where each training image triple in the set of training image triples includes a shadow image, a shadow mask, and a shadowless image;

[0013] A patch determination module, configured to, for the training image triples in the set of training image triples, determine occluded patches and non-occluded patches in the shadow image according to the shadow image and the shadow mask;

[0014] A result determination module, configured to process the occluded patches according to an encoder in the preset shadow model to obtain an encoding result;

[0015] A function determination module, configured to determine a shadow removal result and a loss function corresponding to the shadow removal result according to the encoding result, the non-occluded patches, the shadowless image, and a decoder in the preset shadow model;

[0016] A model determination module, configured to correct the preset shadow model according to the loss function to determine a final shadow model.

[0017] According to a third aspect of the present invention, there is provided an electronic device, where the electronic device includes:

[0018] At least one processor; and

[0019] A memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the shadow model determination method according to any embodiment of the present invention.

[0021] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for enabling a processor to execute the shadow model determination method according to any embodiment of the present invention when executed.

[0022] In the technical solution of the embodiment of the present invention, by obtaining a set of training image triples and a preset shadow model, each training image triple in the set of training image triples includes a shadow image, a shadow mask, and a shadowless image; for the training image triples in the set of training image triples, according to the shadow image and the shadow mask, the occluded patches and non-occluded patches in the shadow image are determined; the occluded patches are processed by the encoder in the preset shadow model to obtain an encoding result; according to the encoding result, the non-occluded patches, the shadowless image, and the decoder in the preset shadow model, a shadow removal result and a loss function corresponding to the shadow removal result are determined; the preset shadow model is corrected according to the loss function to determine the final shadow model. By the directional encoding of the occluded patches, the efficiency of processing the shadow area is ensured, and the performance of shadow removal is improved.

[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0025] Figure 1 is a flowchart of a method for determining a shadow model according to Embodiment 1 of the present invention;

[0026] Figure 2 is an example diagram of a transformer block in a method for determining a shadow model according to Embodiment 1 of the present invention;

[0027] Figure 3 is an example diagram of a framework in a method for determining a shadow model according to Embodiment 1 of the present invention;

[0028] Figure 4 is a result comparison diagram in a method for determining a shadow model according to Embodiment 1 of the present invention;

[0029] Figure 5 is a schematic structural diagram of a device for determining a shadow model according to Embodiment 2 of the present invention;

[0030] Figure 6 is a schematic structural diagram of an electronic device for implementing the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0033] Embodiment 1

[0034] Figure 1 A flowchart of a method for determining a shadow model is provided for Embodiment 1 of the present invention. This embodiment is applicable to the determination of a shadow model. This method can be executed by a shadow model determination device, which can be implemented in the form of hardware and / or software, and the shadow model determination device can be configured in an electronic device. As Figure 1 shown, the method includes:

[0035] S110. Obtain a set of training image triples and a preset shadow model. Each training image triple in the set of training image triples includes a shadow image, a shadow mask, and a shadowless image.

[0036] In this embodiment, the set of training image triplets can be understood as a collection of image triplets for training. The shadow image can be understood as an image including a shadow part. For example, when a vehicle is irradiated by sunlight, there will be a body shadow and other images including shadow parts. The shadow mask can be understood as a mask that only includes the shadow part in the image, which is an image represented by binary. The shadowless image can be understood as an image without a shadow, such as the image without a body shadow in the above example. The preset shadow model can be understood as a model for shadow removal. The preset shadow model is designed with an asymmetric architecture of an encoder and a decoder. In this preset shadow model, the network parameters and computational complexity are imposed on the encoder's processing of the shadow image, while the encoder only uses fewer parameters and less computational effort to process the non-shadow area.

[0037] Specifically, the processor can obtain the set of training image triplets and the preset shadow model from the corresponding storage medium. The set of training image triplets includes multiple training image triplets, and each training image triplet consists of a shadow image, a shadow mask corresponding to the shadow part in the shadow image, and a shadowless image without a shadow.

[0038] S120. For the training image triplets in the set of training image triplets, determine the occluded patches and non-occluded patches in the shadow image according to the shadow image and the shadow mask.

[0039] In this embodiment, the occluded patch can be understood as a patch including shadow occlusion. Since the shadow shape is usually irregular, the occluded patch usually includes a completely occluded patch and a partially occluded patch. The non-occluded patch can be understood as a patch without shadow occlusion.

[0040] Specifically, for each training image triplet included in the set of training image triplets, the processor can distinguish the shadow area and the shadowless area according to the values in the shadow mask. For example, the area with a value of 1 (white) in the shadow mask represents the shadow area, and the area with a value of 0 (black) in the shadow mask represents the shadowless area. The processor can divide the shadow image into a series of image blocks, that is, form a series of regular and non-overlapping patches, and divide the shadow mask in the same way. Since the shadow mask corresponds to the shadow image, these image blocks are distinguished according to the values in the shadow mask. If an image block has a value representing a shadow, then this image block is the shadow area, that is, corresponding to the occluded patch. If the image block does not contain a value representing a shadow, that is, all values represent non-shadow, then this image block is the non-shadow area, that is, corresponding to the non-occluded patch.

[0041] Exemplarily, the training image triplet is (I sf ,I s ,M), where I sfDenote the non - shadow image as I s Denote the shadow image as M, where M represents the shadow mask. Given a shadow image (where H and W are the height and width of the image respectively), the processor can divide it into a series of regular and non - overlapping patches and P seq ={p k} with k = 1, 2, …, M (M = HW / p 2 ), where p is the size of the image patch. The processor divides the corresponding shadow mask in the same way into patches By utilizing the numerical information in the shadow mask, we collect the occluded patches p s and the non - occluded patches p ns . Since the shadow shape is usually irregular, the occluded patches usually include fully occluded patches and partially occluded patches.

[0042] S130. Process the occluded patches according to the encoder in the preset shadow model to obtain an encoding result.

[0043] In this embodiment, the encoder can be understood as an encoder for directionally encoding the shadow region. The encoding result can be understood as the result obtained after being processed by the encoder.

[0044] Among them, the encoder includes at least two first Transformer blocks. The first Transformer block includes a multi - head self - attention layer, a multi - layer perceptron layer, and layer normalization. Exemplarily, the encoder is composed of N Transformer blocks. Each Transformer block includes multi - head self - attention (MSA), multi - layer perceptron (MLP), and layer normalization (LN). The hyperparameter of the embedding dimension in each Transformer block is 128.

[0045] Specifically, the processor can process the extracted occluded patches through the first Transformer block in the encoder in the preset shadow model to obtain an encoding result.

[0046] S140. Determine the shadow - removal result and the loss function corresponding to the shadow - removal result according to the encoding result, the non - occluded patches, the non - shadow image, and the decoder in the preset shadow model.

[0047] In this embodiment, the decoder can be understood as being used to determine the result without shadow. The shadow - removal result can be understood as the image after removing the shadow. The loss function can be understood as a function used to characterize the deviation between the model result and the actual non - shadow image.

[0048] Among them, the decoder in the preset shadow model includes a second Transformer block, and the embedding dimension of the second Transformer block is smaller than that of the first Transformer block. The embedding dimension hyperparameter of the Transformer block in the decoder part is 64.

[0049] Exemplarily, Figure 2 FIG. is an example diagram of a Transformer block in a shadow model determination method provided in Embodiment 1 of the present invention. From Figure 2 it can be seen that the Transformer block may include a linear mapping layer, layer normalization, and multi-head self-attention.

[0050] Specifically, in order to minimize the computational cost, the encoder only operates on the occluded patches. However, since shadow removal requires reconstructing the entire image, the decoder will receive as input the complete features of the latent representations from the encoder and the non-occluded patches. Since the encoder and the decoder have an asymmetric structure, before inputting to the decoder, the encoded result and the non-occluded patches need to be first mapped to the same embedding dimension as the decoder, and the mapped result is input to the decoder of the preset shadow model. To further save computational resources, only a single second Transformer block is set in the decoder of the preset shadow model, and the embedding dimension of the second Transformer block is smaller than that of the first Transformer block. The mapped result is decoded through the second Transformer, pyramid feature module, convolutional module, etc. in the decoder to determine the shadow removal result, and the loss function is determined by calculating the loss between the shadow removal result and the shadow-free image in the ideal case.

[0051] S150. Modify the preset shadow model according to the loss function to determine the final shadow model.

[0052] In this embodiment, the final shadow model can be understood as the finally trained model that meets the accuracy condition.

[0053] Specifically, the processor can modify and iterate the parameters in the preset shadow model through the loss function until the accuracy condition is met, and then obtain the final shadow model.

[0054] The technical solution of the embodiment of the present invention is to obtain a set of training image triples and a preset shadow model. Each training image triple in the set of training image triples includes a shadow image, a shadow mask, and a shadowless image. For the training image triples in the set of training image triples, according to the shadow image and the shadow mask, the occluded patches and non-occluded patches in the shadow image are determined. The occluded patches are processed by the encoder in the preset shadow model to obtain an encoding result. According to the encoding result, the non-occluded patches, the shadowless image, and the decoder in the preset shadow model, a shadow removal result and a loss function corresponding to the shadow removal result are determined. The preset shadow model is corrected according to the loss function to determine the final shadow model. By directionally encoding the occluded patches, the efficiency of processing the shadow area is ensured, and the performance of shadow removal is improved.

[0055] Exemplarily, Figure 3 FIG. 1 is a schematic diagram of a framework in a method for determining a shadow model provided in Embodiment 1 of the present invention. From Figure 3 it can be seen that the occluded patches and non-occluded patches can be distinguished from the shadow image and the shadow mask. To save the computational burden, the encoder only operates on the occluded patches (including fully occluded and partially occluded patches), and the lightweight decoder constructs the shadow removal result through the encoding result of the encoder and the latent representation of the non-occluded patches, that is, the output image is obtained.

[0056] Further, on the basis of the above embodiment, the step of determining the occluded patches and non-occluded patches in the shadow image according to the shadow image and the shadow mask can be optimized as:

[0057] The shadow image is patch-divided in a set manner to obtain image patches; each image patch is classified according to the values of each region in the shadow mask; if all the values included in the image patch are the first preset value, the image patch is used as a non-occluded patch; otherwise, the image patch is used as an occluded patch.

[0058] In this embodiment, the set manner can be understood as a pre-set manner for image division, such as according to pixel points or size ratios, etc., which can be set according to requirements. The image patch can be understood as each result after the image is divided. The region can be understood as the pixel region corresponding to each value. The value can be understood as a value for representing shadow and non-shadow, usually 1 or 0. The first preset value can be understood as a value for representing a non-occluded region. For example, if 0 represents a shadowless region, the corresponding first preset value is 0.

[0059] Specifically, the processor can divide the shadow image into patches in a set manner, which is consistent with the division method of the shadow image, and the formed patches also correspond one by one. At this time, a shadow mask can be used to determine the value corresponding to each image patch. For example, the area with a value of 1 (white) in the shadow mask represents the shadow area, and the area with a value of 0 (black) in the shadow mask represents the non-shadow area. The processor can classify each image patch according to the values of each area in the shadow mask. If all the values included in the image patch are the first preset value, the image patch is regarded as a non-occluded patch. For example, if all the values are 0, it corresponds to a non-occluded patch; otherwise, the image patch is regarded as an occluded patch. For example, if all the values are 1, the image patch is an occluded patch and is a fully occluded patch. If some of the values are 1, it is a partially occluded patch.

[0060] Further, on the basis of the above embodiment, the steps of determining the shadow removal result and the loss function corresponding to the shadow removal result according to the encoding result, non-occluded patches, non-shadow image, and decoder in the preset shadow model can be optimized as follows:

[0061] Map the encoding result and non-occluded patches through the linear layer in the preset shadow model to obtain a target encoding result with the same dimension as the decoder in the preset shadow model, and the latent representation of the non-occluded patches; decode the target encoding result and the latent representation through the second transformer block in the decoder to obtain a transformed result; process the transformed result according to the pyramid feature module in the decoder to obtain pyramid features; perform convolution and upsampling on the pyramid features to obtain the shadow removal result; determine the corresponding loss function according to the shadow removal result and the non-shadow image.

[0062] In this embodiment, the target encoding result can be understood as the result of converting the embedding dimension of the encoding result to be the same as the embedding dimension of the decoder. The latent representation can be understood as the complete feature of the latent representation of the non-occluded patches.

[0063] Specifically, map the encoding result and non-occluded patches through the linear layer in the preset shadow model to obtain a target encoding result with the same embedding dimension as the decoder in the preset shadow model, and the latent representation of the non-occluded patches. Decode the target encoding result and the latent representation through the second transformer block in the decoder to obtain a transformed result. After the second transformer block, a lightweight pyramid feature module (PFM) is introduced. First, perform pooling operations on the transformed result at sizes 4, 8, and 16 according to the pyramid feature module in the decoder to obtain pyramid features. Perform convolution on the pyramid features to reduce their dimensions and perform upsampling operations. Features of various scales are cascaded to generate the shadow removal result, and determine the corresponding loss function according to the shadow removal result and the non-shadow image.

[0064] Based on the above embodiments, the corresponding loss function can be further optimized according to the shadow removal result and the shadow-free image, including:

[0065] Determine the first loss sub-function according to the shadow removal result and the shadow-free image; determine the second loss sub-function according to the feature map extracted from the target layer in the preset shadow model; determine the corresponding loss function according to the first loss sub-function, the second loss sub-function and the preset balance weight.

[0066] In this embodiment, the first loss sub-function can be understood as being used to measure the difference between the shadow removal result and the ideal shadow-free image. The target layer can be understood as the layer number in the feature space for determining the second loss sub-function. The second loss sub-function can be understood as measuring the difference between the shadow removal result and the ideal shadow-free image in the feature space. The preset balance weight can be understood as the set weight parameter for balancing the second loss sub-function.

[0067] Specifically, it can be trained in a supervised manner, and the first loss sub-function can be determined by the following formula in combination with the shadow removal result ( ) and the corresponding shadow-free image I sf to determine the first loss sub-function

[0068]

[0069] To further improve the quality of the result, perceptual loss is introduced, that is, the second loss sub-function, which measures the difference between the shadow removal result and the shadow-free image in the feature space. Exemplary target layers can be the 3rd, 8th, and 15th layers. Specifically, the second loss sub-function is defined as:

[0070]

[0071] where VGG 3,8,15 (·) is the feature map extracted from the 3rd, 8th, and 15th layers of the pre-trained VGG16 model.

[0072] The loss function is determined through the calculation relationship among the first loss sub-function, the second loss sub-function and the preset balance weight, then the loss function can be:

[0073]

[0074] where λ percep. is the balance weight.

[0075] It should be noted that pre-training a deep neural network usually requires a large amount of training data. However, manually collecting a sufficient number of image pairs and creating realistic shadow effects can be a daunting task.

[0076] As a first alternative embodiment of Embodiment 1 of the present invention, based on the above embodiment, the step of constructing the image triple to be trained includes:

[0077] Obtain a preset image dataset and a preset shadow detection dataset; perform non-linear adjustment on each input image included in the preset image dataset through a preset gamma correction algorithm to obtain an image to be synthesized; determine a target shadow image according to the target shadow mask, the image to be synthesized, and the input image in the preset shadow detection dataset; determine the image triple to be trained according to the target shadow image, the shadowless image corresponding to the input image, and the target shadow mask.

[0078] In this embodiment, the preset image dataset can be understood as a set of some pre-set images. The preset gamma correction algorithm can be understood as a method of performing non-linear tone editing on an image by editing the gamma curve of the image. The image to be synthesized can be understood as the result of gamma curve correction. The preset shadow detection dataset can be understood as a set including shadow masks of various shapes, sizes, and positions. The target shadow mask can be understood as a mask of a randomly selected shadow part. The input image can be understood as the original image in the preset image dataset. The shadowless image corresponding to the input image can be understood as the image without a shadow corresponding to the input image. The target shadow image can be understood as the image obtained by adding the target shadow mask to the input image.

[0079] Specifically, shadows are usually characterized by features such as reduced contrast, weakened illumination, and color deviation. To solve this problem, first, an input image can be randomly selected from the preset image dataset and denoted as I in . Then, the processor can perform non-linear adjustment on the input image using a gamma correction (GC) operation to obtain the image to be synthesized I GC , which is defined as:

[0080]

[0081] where the constant weight a is set to 1, and the value of γ is randomly selected from the range [1, 3].

[0082] Furthermore, the processor can use an image decomposition model to synthesize the target shadow image I s in combination with the target shadow mask, which is denoted as:

[0083] I s = I GC * M+(1 - M)*I in

[0084] Where M represents a binarized shadow mask.

[0085] Due to the high computational efficiency of generating the synthetic shadow image, the processor can quickly generate a large number of image triples (I sf , I s , M) during the training phase, where I sf is the shadowless image corresponding to the target shadow image. Due to the asymmetric architecture of the shadow model framework, it is also efficient during the initial pre-training phase.

[0086] The technical solution of the embodiment of the present invention determines the occluded patches and non-occluded patches through the shadow mask and the shadow image, processes the occluded patches through the encoder, and through the design of the directional encoding and the lightweight decoder, based on the transformer model, can better aggregate global context information, improve the quality of image generation. It can process the shadow area more efficiently, improve the performance of shadow removal, adopt a shadow-customized pre-training strategy, and use a preset gamma correction algorithm to synthesize the target shadow image according to the target shadow mask, the image to be synthesized, and the input image of the shadow data set, improve the model's understanding of shadow features, and thus obtain higher performance at low cost. Compared with other methods based on CNN or Transformer, the present invention is more efficient and can complete image shadow removal in a short time, improving the processing efficiency.

[0087] Exemplarily, in order to verify the shadow removal effect in the present invention, a comparison is made with the latest technology. Figure 4 This is a result comparison diagram in a shadow model determination method provided in the first embodiment of the present invention. The experimental data sets used in the comparison: The present invention has conducted experiments on two widely used real-world data sets (ISTD+ and SRD), where ISTD+ includes 1330 training triple images and 540 test triple images, and SRD includes 2680 training pairs and 408 test pairs. Evaluation metrics: In order to ensure a fair comparison of the shadow removal performance, three commonly used evaluation metrics are adopted, including root mean square error (RMSE), peak signal-to-noise ratio (PSNR), and structural similarity (SSIM). Comparison with the latest technology: Comparisons are made with a variety of other latest shadow removal methods, including DeShadowNet, DSC, DHAN, BMNet, SG-ShadowNet, Inp-Shadow, and DMTN, etc. Through quantitative comparison, the present invention demonstrates excellent performance under multiple metrics. Inference speed comparison: The present invention has compared the inference speed with other shadow removal methods on the ISTD+ test data set. The results show that compared with methods such as ShadowFormer, the processing speed of the present invention has increased by nearly 21 times. Visual effect comparison: Visual effect comparisons are made on the ISTD+ and SRD data sets. ThroughFigure 4 It can clearly show the superiority of the shadow removal effect of the present invention compared with other methods. Figure 4 On the far left is the original input image. It can be seen that the image includes the shadows of the person with an umbrella and the tree. After removing the shadows by the SP+M+I-Net, SG-ShadowNet, and ShadowFormer methods, it can be seen that the shadows at the floor tiles are significantly inconsistent with those in the shadowless image, and even deformities and other phenomena occur. However, the result processed by the method of the present invention is closer to the shadowless image and has a better effect.

[0088] Embodiment 2

[0089] Figure 5 It is a schematic structural diagram of a shadow model determination device provided in Embodiment 2 of the present invention.

[0090] As Figure 5 shown, the device includes:

[0091] A model acquisition module 41, configured to acquire a set of training image triples and a preset shadow model, and each training image triple in the set of training image triples includes a shadow image, a shadow mask, and a shadowless image;

[0092] A patch determination module 42, configured to determine an occluded patch and a non-occluded patch in the shadow image according to the shadow image and the shadow mask for the training image triples in the set of training image triples;

[0093] A result determination module 43, configured to process the occluded patch according to the encoder in the preset shadow model to obtain an encoding result;

[0094] A function determination module 44, configured to determine a shadow removal result and a loss function corresponding to the shadow removal result according to the encoding result, the non-occluded patch, the shadowless image, and the decoder in the preset shadow model;

[0095] A model determination module 45, configured to correct the preset shadow model according to the loss function to determine a final shadow model.

[0096] In the technical solution of the embodiment of the present invention, by obtaining a set of training image triples and a preset shadow model, each training image triple in the set of training image triples includes a shadow image, a shadow mask, and a shadowless image; for the training image triples in the set of training image triples, according to the shadow image and the shadow mask, the occluded patches and non-occluded patches in the shadow image are determined; the occluded patches are processed by the encoder in the preset shadow model to obtain an encoding result; according to the encoding result, the non-occluded patches, the shadowless image, and the decoder in the preset shadow model, a shadow removal result and a loss function corresponding to the shadow removal result are determined; the preset shadow model is corrected according to the loss function to determine the final shadow model. By the directional encoding of the occluded patches, the efficiency of shadow area processing is ensured, and the performance of shadow removal is improved.

[0097] Wherein, the encoder includes at least two first transformer blocks, and each first transformer block includes a multi-head self-attention layer, a multi-layer perceptron layer, and layer normalization.

[0098] Wherein, the decoder in the preset shadow model includes a second transformer block, and the embedding dimension of the second transformer block is smaller than the embedding dimension of the first transformer block.

[0099] Further, the patch determination module 42 is specifically configured to:

[0100] Divide the shadow image into patches in a set manner to obtain image patches;

[0101] Classify each of the image patches according to the values of the regions in the shadow mask;

[0102] If all the values included in the image patch are the first preset value, the image patch is used as a non-occluded patch;

[0103] Otherwise, the image patch is used as an occluded patch.

[0104] Further, the function determination module 44 includes:

[0105] A characterization determination unit, configured to map the encoding result and the non-occluded patches through a linear layer in the preset shadow model to obtain a target encoding result having the same dimension as the decoder in the preset shadow model, and a latent characterization of the non-occluded patches;

[0106] A result determination unit, configured to decode the target encoding result and the latent characterization through the second transformer block in the decoder to obtain a transformed result;

[0107] A feature determination unit, configured to process the transformation result according to a pyramid feature module in the decoder to obtain pyramid features;

[0108] A result determination unit, configured to perform convolution and upsampling on the pyramid features to obtain a shadow removal result;

[0109] A function determination unit, configured to determine a corresponding loss function according to the shadow removal result and the shadowless image.

[0110] Among them, the function determination unit is specifically configured to:

[0111] Determine a first loss sub-function according to the shadow removal result and the shadowless image;

[0112] Determine a second loss sub-function according to the feature map extracted from the target layer in the preset shadow model;

[0113] Determine a corresponding loss function according to the first loss sub-function, the second loss sub-function and a preset balance weight.

[0114] Optionally, the device further includes a construction module:

[0115] The construction module is specifically configured to:

[0116] Obtain a preset image data set and a preset shadow detection data set;

[0117] Perform non-linear adjustment on each input image included in the preset image data set through a preset gamma correction algorithm to obtain a to-be-synthesized image;

[0118] Determine a target shadow image according to the target shadow mask, the to-be-synthesized image and the input image in the preset shadow detection data set;

[0119] Determine the to-be-trained image triple according to the target shadowless image, the to-be-synthesized image and the target shadow mask.

[0120] The shadow model determination device provided by the embodiments of the present invention can execute the shadow model determination method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0121] Embodiment III

[0122] Figure 6FIG. 0 shows a schematic structural diagram of an electronic device 60 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0123] As Figure 6 shown, the electronic device 60 includes at least one processor 61, and a memory communicatively connected to the at least one processor 61, such as a read-only memory (ROM) 62, a random access memory (RAM) 63, etc., wherein the memory stores a computer program executable by the at least one processor. The processor 61 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 62 or the computer program loaded from the storage unit 68 into the random access memory (RAM) 63. In the RAM 63, various programs and data required for the operation of the electronic device 60 can also be stored. The processor 61, the ROM 62, and the RAM 63 are connected to each other via a bus 64. An input / output (I / O) interface 65 is also connected to the bus 64.

[0124] A plurality of components in the electronic device 60 are connected to the I / O interface 65, including: an input unit 66, such as a keyboard, a mouse, etc.; an output unit 67, such as various types of displays, speakers, etc.; a storage unit 68, such as a magnetic disk, an optical disk, etc.; and a communication unit 69, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 69 allows the electronic device 60 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0125] The processor 61 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 61 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 61 executes the various methods and processes described above, such as the shadow model determination method.

[0126] In some embodiments, the shadow model determination method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 68. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 60 via the ROM 62 and / or the communication unit 69. When the computer program is loaded into the RAM 63 and executed by the processor 61, one or more steps of the shadow model determination method described above may be performed. Alternatively, in other embodiments, the processor 61 may be configured to perform the shadow model determination method by any other suitable means (e.g., by means of firmware).

[0127] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0128] The computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0129] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0130] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0131] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0132] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0133] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0134] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A shadow model determination method, characterized in that: include: Acquire a set of image triples to be trained and a preset shadow model, wherein each image triple to be trained in the set of image triples to be trained includes a shadow image, a shadow mask, and a shadow-free image; For the to-be-trained image triple in the to-be-trained image triple set, determining an occluded patch and a non-occluded patch in the shadow image according to the shadow image and the shadow mask; Processing the occlusion patch according to an encoder in a preset shadow model to obtain an encoding result; Determine a shadow removal result and a loss function corresponding to the shadow removal result according to the encoding result, the non-occluded patch, the shadow-free image and a decoder in the preset shadow model; Modifying the preset shadow model according to the loss function to determine a final shadow model; The encoder includes at least two first transformer blocks, wherein the first transformer block includes a multi-head self-attention layer, a multi-layer perceptron layer and a layer normalization; the decoder in the preset shadow model includes a second transformer block, and the embedding dimension of the second transformer block is smaller than the embedding dimension of the first transformer block; The step of determining the shadow removal result and the loss function corresponding to the shadow removal result according to the encoding result, the non-occluded patch, the shadow-free image and the decoder in the preset shadow model includes: Mapping the encoding result and the non-occluded patch through a linear layer in the preset shadow model to obtain a target encoding result having the same decoder dimension as that in the preset shadow model and a potential representation of the non-occluded patch; Decoding the target encoding result and the potential representation by a second transformer block in the decoder to obtain a transformation result; Processing the transformation result according to the pyramid feature module in the decoder to obtain a pyramid feature; Convolving and upsampling the pyramid features to obtain a shadow removal result; A corresponding loss function is determined according to the shadow removal result and the shadow-free image.

2. The method according to claim 1, characterized in that The determining, according to the shadow image and the shadow mask, the occluded patches and the non-occluded patches in the shadow image comprises: Divide the shadow image into patches according to a set method to obtain image patches; Classifying each of the image patches according to the value of each region in the shadow mask; If all values ​​included in the image patch are first preset values, the image patch is used as a non-occluded patch; Otherwise, the image patch is taken as an occlusion patch.

3. The method according to claim 1, characterized in that The determining a corresponding loss function according to the shadow removal result and the shadow-free image includes: Determining a first loss sub-function according to the shadow removal result and the shadow-free image; Determining a second loss sub-function according to a feature map extracted from a target layer in the preset shadow model; A corresponding loss function is determined according to the first loss sub-function, the second loss sub-function and a preset balance weight.

4. The method according to claim 1, characterized in that: The step of constructing the image triples to be trained includes: Obtain a preset image data set and a preset shadow detection data set; Performing nonlinear adjustment on each input image included in the preset image data set by using a preset gamma correction algorithm to obtain an image to be synthesized; Determine a target shadow image according to the target shadow mask in the preset shadow detection data set, the image to be synthesized and the input image; The image triplet to be trained is determined according to the target shadow image, the image to be synthesized and the target shadow mask.

5. A shadow model determination device, characterized in that: include: A model acquisition module, used to acquire a set of image triples to be trained and a preset shadow model, wherein each image triple to be trained in the set of image triples to be trained includes a shadow image, a shadow mask and a shadow-free image; a patch determination module, configured to determine, for the to-be-trained image triple in the to-be-trained image triple set, an occluded patch and a non-occluded patch in the shadow image according to the shadow image and the shadow mask; A result determination module, used for processing the occlusion patch according to an encoder in a preset shadow model to obtain an encoding result; A function determination module, configured to determine a shadow removal result and a loss function corresponding to the shadow removal result according to the encoding result, the non-occluded patch, the shadow-free image and a decoder in the preset shadow model; A model determination module, used to modify the preset shadow model according to the loss function to determine a final shadow model; The encoder includes at least two first transformer blocks, wherein the first transformer block includes a multi-head self-attention layer, a multi-layer perceptron layer and a layer normalization; the decoder in the preset shadow model includes a second transformer block, and the embedding dimension of the second transformer block is smaller than the embedding dimension of the first transformer block; Wherein, the function determination module includes: A representation determination unit, configured to map the encoding result and the non-occluded patch through a linear layer in the preset shadow model to obtain a target encoding result having the same decoder dimension as that in the preset shadow model and a potential representation of the non-occluded patch; A result determination unit, configured to decode the target encoding result and the potential representation through a second transformer block in the decoder to obtain a transformation result; A feature determination unit, configured to process the transformation result according to a pyramid feature module in the decoder to obtain a pyramid feature; A result determination unit, used for performing convolution and up-sampling on the pyramid features to obtain a shadow removal result; A function determination unit is used to determine a corresponding loss function according to the shadow removal result and the shadow-free image.

6. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so as to enable the at least one processor to perform the shadow model determination method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the shadow model determination method according to any one of claims 1 to 4 when executed.

Citation Information

Patent Citations

  • Image shadow removal model and construction method, device and application thereof

    CN115375589A

  • Detecting conflicts between multiple different encoded signals within imagery, using only a subset of available image data, and robustness checks

    US20220343454A1