Instance shadow detection method
By constructing data sets under different light source scenarios and using pairing query mechanisms and contrasting morphological alignment modules, the problem of inconsistent matching relationship between objects and shadows under complex lighting conditions is solved, and efficient instance shadow detection is achieved.
Patent Information
- Application Number
- CN202510321517.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, under complex lighting conditions, the offset-based bidirectional relationship learning method leads to inconsistent matching relationships between objects and shadows, which reduces detection accuracy and efficiency.
By constructing data sets in different light source scenarios, using deep learning models to extract advanced features, and using paired query mechanisms and contrasted morphological alignment modules for morphological alignment, we can achieve segmented instance detection of goals and shadows.
This improves the efficiency of instance shadow detection under complex lighting conditions, avoids the accumulation of complexity and errors caused by traditional bidirectional matching, and improves detection accuracy.
Smart Images

Figure CN120298567A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a method for detecting instance shadows. Background Art
[0002] In the fields of image processing and computer vision, instance shadow detection is a key task, aiming to detect shadow instances in images and associate them with the object instances that project the shadows. This technology has wide value in application scenarios such as image editing, augmented reality, autonomous driving, and light direction estimation.
[0003] In the prior art, most of them focus on single light source and natural light scenarios, usually assuming that each object projects only one shadow. The offset-based bidirectional relationship learning method realizes association matching through bidirectional offset prediction from object to shadow and shadow to object.
[0004] However, under complex lighting conditions, such as artificial multi-light source scenarios, a single object may project multiple shadows, and the offset-based bidirectional relationship learning method leads to inconsistent bidirectional matching relationships, reducing the detection accuracy and efficiency of targets and shadows. Summary of the Invention
[0005] In view of this, it is necessary to provide a method for constructing an instance shadow detection model and a method for detecting instance shadows to solve the problem of inconsistent matching relationships between objects and shadows in the prior art.
[0006] To solve the above problems, in a first aspect, the present invention provides a method for detecting instance shadows, including: Group and label the objects and shadows in the target scene image to obtain the labeled image; Input the labeled image into a trained instance shadow detection model to obtain the instance shadows of the target scene image; Wherein, the trained instance shadow detection model includes: Construct a data set under different light source scenarios; Input the data set into a preset deep learning model to obtain high-level features; Input the high-level features into a pixel decoder to obtain pixel embeddings and multi-scale features; Input the multi-scale features into a pairing grouping layer, and the pairing grouping layer extracts the association features between the target and the shadow from the multi-scale features based on the pairing query mechanism to obtain paired output embeddings; Input the pixel embeddings and the paired output embeddings into a contrast morphological alignment module for morphological alignment to obtain the segmentation instances of the object and the shadow.
[0007] In a possible implementation, the preset deep learning model includes one or more of ResNet, MobileNet, and VGG.
[0008] In a possible implementation, after constructing the dataset under different light source scenarios, it further includes: Performing data augmentation on the dataset; Grouping and annotating the objects and shadows in the dataset to obtain a preprocessed dataset.
[0009] In a possible implementation, the pairing query mechanism is to modify the query vector of the decoder in the preset neural network framework into a paired query vector.
[0010] In a possible implementation, the pairing grouping layer includes a pairwise mask attention mechanism and a self-attention mask mechanism. The pairwise mask attention mechanism includes a self-mask attention mechanism and a cross-mask attention mechanism. The pairing grouping layer extracts the correlation features between the target and the shadow from the multi-scale features based on the pairing query mechanism to obtain a paired output embedding, including: The cross-mask attention mechanism performs cross-scale feature interaction and fusion on the multi-scale features based on the paired query vector to obtain a first correlation feature; The self-mask attention mechanism further extracts and correlates the first correlation feature to obtain a second correlation feature; Inputting the second correlation feature into the self-attention mask mechanism to obtain a paired output embedding.
[0011] In a possible implementation, inputting the pixel embedding and the paired output embedding into the contrast morphological alignment module to obtain the segmentation instances of the object and the shadow, including: The contrast morphological alignment module performs a multiplication operation on the pixel embedding and the paired output embedding to obtain the separation instances of the object and the shadow.
[0012] In a possible implementation, performing edge optimization on the separation instances of the object and the shadow to obtain the aligned separation instances of the object and the shadow; The performing edge optimization on the separation instances of the object and the shadow includes: Obtaining a detail mask based on the separation instances of the object and the shadow; The detail mask undergoes a two-dimensional discrete convolution transform to obtain the frequency domain representation of the separation instance; Performing morphological alignment on the separation instance based on the frequency domain representation to obtain the aligned separation instances of the object and the shadow.
[0013] In a second aspect, the present invention further provides a device for detecting instance shadows, including: An image annotation module, which is used to group and annotate the objects and shadows in the target scene image to obtain the annotated image; An instance shadow segmentation module, which is used to input the annotated image into a trained instance shadow detection model to obtain the instance shadows of the target scene image; Among them, the trained instance shadow detection model includes: Construct a data set under different light source scenarios; Input the data set into a preset deep learning model to obtain high-level features; Input the high-level features into a pixel decoder to obtain pixel embeddings and multi-scale features; Input the multi-scale features into a pairing grouping layer, and the pairing grouping layer extracts the association features between the target and the shadow from the multi-scale features based on the pairing query mechanism to obtain paired output embeddings; Input the pixel embeddings and the paired output embeddings into a contrast morphological alignment module for morphological alignment to obtain the segmentation instances of the object and the shadow.
[0014] In a third aspect, the present invention further provides an electronic device, including a memory and a processor, wherein, The memory is used to store a program; The processor is coupled to the memory and is used to execute the program stored in the memory to implement the steps in a method for detecting an instance shadow described in any one of the above implementation manners.
[0015] In a fourth aspect, the present invention further provides a computer-readable storage medium, which is used to store computer-readable programs or instructions, and when the programs or instructions are executed by a processor, the steps in a method for detecting an instance shadow described in any one of the above implementation manners can be implemented.
[0016] The beneficial effects of the present invention are as follows: A method for detecting instance shadows provided by the present invention first groups and labels the objects and shadows in the target scene image to obtain the labeled image, and then inputs the labeled image into a trained instance shadow detection model to obtain the instance shadows of the target scene image. Among them, the trained instance shadow detection model includes constructing a dataset under different light source scenarios, inputting the dataset into a preset deep learning model to obtain high-level features, inputting the high-level features into a pixel decoder to obtain pixel embeddings and multi-scale features, inputting the multi-scale features into a pairing grouping layer, and the pairing grouping layer extracts the correlation features between the target and the shadow from the multi-scale features based on the pairing query mechanism to obtain paired output embeddings. On the basis of the Mask2Former structure, the query vectors are paired to realize the paired relationship between the target and the shadow. The pixel embeddings and the paired output embeddings are input into a contrast morphological alignment module for morphological alignment to obtain the segmentation instances of the object and the shadow. The present invention matches the target and the shadow in the dataset through the pairing query mechanism, thereby establishing a corresponding relationship between the target and the shadow in the dataset without bidirectional matching. Under the action of the pairing query mechanism, the pairing grouping layer further extracts the correlation features between the target and the shadow, avoiding the complexity and error accumulation caused by bidirectional matching in the traditional method and improving the detection efficiency. Description of the Drawings Figure 1 It is a flowchart of a method of an embodiment of a method for detecting instance shadows provided by the present invention; Figure 2 It is a schematic structural diagram of an instance shadow detection model provided by an embodiment of the present invention; Figure 3 For the present invention Figure 1 It is a schematic flowchart of an embodiment of S105 in the present invention; Figure 4 It is a schematic structural diagram of a pairing grouping layer in an embodiment of the present invention; Figure 5 It is a schematic structural diagram of a contrast morphological alignment module in an embodiment of the present invention; Figure 6 It is a comparison diagram of visual prediction results between the method of the present invention and the existing method in instance shadow detection; Figure 7 It is a schematic flowchart of an embodiment of a construction device of an instance shadow detection model provided by the present invention; Figure 8 It is a schematic structural diagram of an embodiment of an electronic device provided by the present invention. Detailed Embodiments
[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0018] In the description of the embodiments of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0019] In the embodiments of the present invention, the descriptions such as "first" and "second" are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Therefore, the technical features defined with "first" and "second" may explicitly or implicitly include at least one such feature.
[0020] Referring to "embodiment" herein means that a specific feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0021] Before presenting the embodiments, the following terms will be explained first.
[0022] First, introduce Mask2Former. The overall framework of Mask2Former is divided into a pixel-level model, a Transformer model, and a segmentation model. First, image features are extracted through the backbone, and then they are sent to the decoder to generate pixel embedding features. In the Transformer model, low-resolution image features are used as K and V, and N pre-segmented embedding vectors Q are generated through the Transformer decoder in combination with the initialized query vector query. One branch of Q is sent for classification through the MLP, and the other is mapped to the space of pixel embeddings for mask prediction. Finally, the mask prediction and the class prediction are multiplied matrix-wise and sent for segmentation.
[0023] Embodiment 1: The present invention provides a method for constructing an instance shadow detection model and a method for detecting instance shadows, which will be described separately below.
[0024] Figure 1 A schematic flowchart of an embodiment of the method for constructing an instance shadow detection model provided by the present invention is shown as Figure 1 shown. The method for constructing an instance shadow detection model includes: S101. Group and label the objects and shadows in the target scene image to obtain the labeled image; S102. Input the labeled image into the trained instance shadow detection model to obtain the instance shadow of the target scene image; Among them, the trained instance shadow detection model includes: Construct a dataset under different light source scenarios; Input the dataset into a preset deep learning model to obtain high-level features; Input the high-level features into a pixel decoder to obtain pixel embeddings and multi-scale features; Input the multi-scale features into a pairing grouping layer, and the pairing grouping layer extracts the correlation features between the target and the shadow from the multi-scale features based on the pairing query mechanism to obtain paired output embeddings; Input the pixel embeddings and the paired output embeddings into a contrast morphological alignment module for morphological alignment to obtain the segmentation instances of the object and the shadow.
[0025] Compared with the prior art, a method for detecting an instance shadow provided in this embodiment first groups and labels the objects and shadows in the target scene image to obtain the labeled image, and then inputs the labeled image into the trained instance shadow detection model to obtain the instance shadow of the target scene image. Among them, the trained instance shadow detection model includes constructing a dataset under different light source scenarios, inputting the dataset into a preset deep learning model to obtain high-level features, inputting the high-level features into a pixel decoder to obtain pixel embeddings and multi-scale features, inputting the multi-scale features into a pairing grouping layer, and the pairing grouping layer extracts the correlation features between the target and the shadow from the multi-scale features based on the pairing query mechanism to obtain paired output embeddings. On the basis of the Mask2Former structure, the query vectors are paired to realize the paired relationship between the target and the shadow. Input the pixel embeddings and the paired output embeddings into a contrast morphological alignment module for morphological alignment to obtain the segmentation instances of the object and the shadow. The present invention matches the target and the shadow in the dataset through the pairing query mechanism, thereby establishing a corresponding relationship between the target and the shadow in the dataset, without two-way matching. Under the action of the pairing query mechanism, the pairing grouping layer further extracts the correlation features between the target and the shadow, avoiding the complexity and error accumulation caused by two-way matching in the traditional method and improving the detection efficiency. In a specific embodiment of the present invention, as Figure 2As shown in the figure, the structure of the model corresponding to the method of this embodiment includes: a backbone network, a pixel decoder, a pairing grouping layer, and a contrastive morphological alignment module. The input image is processed by a backbone network for multi-scale feature extraction to generate pixel embeddings and multi-scale features, which are input into the pairing grouping layer. Then, through a pairing query mechanism, the target object and the shadow are grouped into pairing queries. Relevant features are extracted in the pairing grouping layer, and the pairing relationship between the target and the shadow is further optimized through the contrastive morphological alignment module. Finally, the model outputs the segmentation masks and the association relationships of the target and the shadow.
[0026] In some embodiments of the present invention, after constructing the dataset under different light source scenarios, it further includes: Performing data augmentation on the dataset; Grouping and annotating the objects and shadows in the dataset to obtain a preprocessed dataset.
[0027] In a specific embodiment of the present invention, the data required for training the model is collected, including USSO and SOBAv2. The USSO dataset covers multiple scenarios (natural light, multiple light sources, etc.), and the pairing relationship between the target and the shadow in the dataset is annotated in detail, so that the annotated dataset is paired with the query mechanism of this application, that is, the query vector in the form of pairing.
[0028] In some embodiments of the present invention, the pairing query mechanism is to modify the query vector of the decoder in a preset neural network framework into a query vector in the form of pairing.
[0029] In a specific embodiment of the present invention, the preset neural network framework is the Transformer framework.
[0030] In a specific embodiment of the present invention, the query vector in Mask2Former is modified into the form of target pairing, that is, the target query and the shadow query are combined to form a pairing query, and the target and the corresponding shadow are detected respectively. For example, an object query and a shadow query are combined together as a pair of tasks, which are responsible for detecting the object and its corresponding shadow respectively. These queries are called "pairing queries", denoted as , where represents the object query, and represents the shadow query. The total number of object and shadow queries is N, so N / 2 pairs are formed.
[0031] Group the object query and the shadow query to form a pairing query as the input of the decoder. Through this pairing method, the preliminary association detection between the object and the shadow can be efficiently completed, the detection process can be significantly simplified, and the matching efficiency can be improved.
[0032] In some embodiments of the present invention, such asFigure 3 As shown, the pairing grouping layer includes a pairwise masked attention mechanism and a self-attention masking mechanism. The pairwise masked attention mechanism includes a self-masking attention mechanism and a cross-masking attention mechanism. The pairing grouping layer extracts the correlation features between the target and the shadow from the multi-scale features based on the pairing query mechanism to obtain a paired output embedding, including: S301. The cross-masking attention mechanism performs cross-scale feature interaction and fusion on the multi-scale features based on the query vector in the pairing form to obtain the first correlation feature; S302. The self-masking attention mechanism further extracts and correlates the first correlation feature to obtain the second correlation feature; S303. Input the second correlation feature into the self-attention masking mechanism to obtain the paired output embedding.
[0033] In a specific embodiment of the present invention, as Figure 4 shown, the pairing grouping layer adopts a pairing query mechanism. The target query and the shadow query form a pairing query to detect the target and the corresponding shadow respectively. The pairing query serves as the Query, combines the multi-scale features as the Key and Value, and participates in the calculation of the cross-masking attention. In this way, the pairing query can extract the correlation features between the target and the shadow from the multi-scale features to generate the paired output embedding and . The pairing grouping layer strengthens the interaction relationship between the target and the shadow through cross-masking attention and self-masking attention, effectively improving the accuracy of the target and shadow pairing relationship. Among them, and respectively represent the attention masks derived from the object and shadow instances of the th layer of the Decoder Layer.
[0034] The cross-masking attention mechanism is used to extract the relevant features between the object and the shadow, and at the same time, the self-masking attention mechanism enhances the capture of the object and shadow's own features. The two attention mechanisms complement each other, further improving the accuracy and robustness of the object and shadow matching.
[0035] In some embodiments of the present invention, inputting the pixel embedding and the paired output embedding into the contrast morphological alignment module for morphological alignment to obtain the segmentation instances of the object and the shadow includes: The contrast morphological alignment module performs a multiplication operation on the pixel embedding and the paired output embedding to obtain the separation instances of the object and the shadow.
[0036] In some embodiments of the present invention, it further includes: Edge optimization is performed on the separation instances of the object and the shadow to obtain the separated instances of the aligned object and shadow; The edge optimization of the separation instances of the object and the shadow includes: A detail mask is obtained based on the separation instances of the object and the shadow; The detail mask undergoes a two-dimensional discrete convolution transform to obtain the frequency domain representation of the separation instance; Based on the frequency domain representation, morphological alignment is performed on the separation instance to obtain the separated instances of the aligned object and shadow.
[0037] In some embodiments of the present invention, as Figure 5 shown, the paired output embeddings and calculate the confidence levels of the object and the shadow, and the matching score (Matching Score, an evaluation of the matching degree between the object and the shadow) through a feed-forward network (FFN). At the same time, the paired output embeddings and respectively perform multiplication operations with the high-resolution pixel embeddings to obtain the final segmentation instances of the object and the shadow and .
[0038] Figure 5 The dashed box part of Figure 5 shows the structure for optimizing the morphological alignment of the target and the shadow. Among them, the detail masks and are generated from the segmentation instances and through a label decoupling method, so as to obtain the boundary regions and morphological features of the target and the shadow. To reduce the computational complexity, the detail masks and are encoded into the low-dimensional frequency domain representations of the target and the shadow and through the discrete cosine transform (DCT). Subsequently, the morphological alignment matrix between and is calculated through a dot product operation, and the morphological alignment is optimized using a contrastive learning strategy. Finally, the contrastive morphological alignment loss is calculated to optimize the model. The optimization objectives of the model include the classification loss , the mask loss , the matching score loss , and the contrastive morphological alignment loss . The final total loss function is: . Among them, , , and are loss weight hyperparameters.
[0039] Example 2: In this example, two datasets are used to evaluate the model performance. The SOBAv2 dataset contains a standard test set (SOBAv2-T) and a more challenging test set with a high number of instances (SOBAv2-C). In addition, the Universal Scene Shadow-Object dataset (USSO) is used, which covers complex scenarios such as natural light and multiple light sources, and a subset USSO-A that only contains artificial light source scenarios is designed.
[0040] This example is implemented based on the Detectron2 framework and uses ResNet-101 as the backbone network of the instance shadow detection model. The AdamW optimizer is used during the optimization process, and the initial learning rate is , for the SOBAv2 dataset, the learning rate is decayed at the 20,000th and 40,000th iterations, and the training stops at the 45,000th iteration; for the USSO dataset, the learning rate is decayed at the 80,000th iteration, and the training stops at the 90,000th iteration.
[0041] As Figure 6 shown, the comparative visualization effects of various methods applied to outdoor and indoor scenes in the USSO dataset are presented. This example can always accurately identify all target-shadow instances, while other methods have missed detections.
[0042] The workflow of a method for constructing an instance shadow detection model provided in this example includes: Step 1: Data collection and import. Collect the data required for training the model, including USSO and SOBAv2. The USSO dataset covers multiple scenarios (natural light, multiple light sources, etc.), and the pairing relationship between the target and the shadow is detailedly annotated. These data are introduced into the training process through the data import module.
[0043] Step 2: Data preprocessing. In the data preprocessing module, data augmentation is performed on the original images by applying random cropping, horizontal flipping, vertical flipping, and random rotation.
[0044] Step 3: Training of the instance shadow detection model. The warm-up strategy is used during training, and the learning rate is linearly increased within the first 10 iterations. At the same time, the weight decay is set to 0.05 to improve the regularization ability of the model. The long side of the input image is limited to 1,333 pixels while maintaining the aspect ratio, and the short side ranges from 640 to 800 pixels. The data augmentation strategy includes random horizontal flipping operations. Each loss hyperparameter is set to , , and .
[0045] Step 4: Model verification. Use the USSO and SOBAv2 verification datasets to evaluate the performance of the instance shadow detection model, mainly detecting segmentation accuracy, matching efficiency, and detection robustness. The verification results are used to adjust hyperparameters and optimize the model performance.
[0046] To better implement a method for constructing an instance shadow detection model in an embodiment of the present invention, correspondingly, based on a method for constructing an instance shadow detection model, as Figure 7 shown, an embodiment of the present invention also provides a device for constructing an instance shadow detection model. A device 700 for constructing an instance shadow detection model includes: An image annotation module 701 for grouping and annotating objects and shadows in a target scene image to obtain an annotated image; An instance shadow segmentation module 702 for inputting the annotated image into a trained instance shadow detection model to obtain instance shadows in the target scene image; Among them, the trained instance shadow detection model includes: Constructing datasets under different light source scenarios; Inputting the dataset into a preset deep learning model to obtain high-level features; Inputting the high-level features into a pixel decoder to obtain pixel embeddings and multi-scale features; Inputting the multi-scale features into a pairing grouping layer, and the pairing grouping layer extracts the association features between the target and the shadow from the multi-scale features based on the pairing query mechanism to obtain paired output embeddings; Inputting the pixel embeddings and the paired output embeddings into a contrast morphological alignment module for morphological alignment to obtain segmentation instances of the object and the shadow.
[0047] The device 700 for constructing an instance shadow detection model provided in the above embodiment can implement the technical solutions described in the above embodiment of the method for detecting an instance shadow. For the specific implementation principles of the above modules or units, reference can be made to the corresponding content in the above embodiment of the method for detecting an instance shadow, which will not be elaborated here.
[0048] As Figure 8 shown, the present invention also correspondingly provides an electronic device 800. The electronic device 800 includes a processor 801, a memory 802, and a display 803. Figure 8 Only some components of the electronic device 800 are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0049] In some embodiments, the processor 801 may be a central processing unit (CPU), a microprocessor, or other data processing chips, which are used to run the program code stored in the memory 802 or process data, such as a method for constructing an instance shadow detection model in the present invention.
[0050] In some embodiments, the processor 801 may be a single server or a server group. The server group can be centralized or distributed. In some embodiments, the processor 801 can be local or remote. In some embodiments, the processor 801 can be implemented on a cloud platform. In some embodiments, the cloud platform may include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an internal cloud, a multi-cloud, etc., or any combination of the above.
[0051] In some embodiments, the memory 802 may be an internal storage unit of the electronic device 800, such as a hard disk or a memory of the electronic device 800. In some other embodiments, the memory 802 may also be an external storage device of the electronic device 800, such as a plug-in hard disk equipped on the electronic device 800, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.
[0052] Furthermore, the memory 802 may also include both the internal storage unit and the external storage device of the electronic device 800. The memory 802 is used to store the application software installed on the electronic device 800 and various types of data.
[0053] In some embodiments, the display 803 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. The display 803 is used to display the information in the electronic device 800 and to display a visual user interface. The components 801-803 of the electronic device 800 communicate with each other through a system bus.
[0054] In one embodiment, when the processor 801 executes a program for constructing an instance shadow detection model in the memory 802, the following steps may be implemented: Group and label the objects and shadows in the target scene image to obtain the labeled image; Input the labeled image into the trained instance shadow detection model to obtain the instance shadow of the target scene image; Among them, the trained instance shadow detection model includes: Construct a data set under different light source scenarios; Input the data set into a preset deep learning model to obtain high-level features; Input the high-level features into a pixel decoder to obtain pixel embeddings and multi-scale features; Input the multi-scale features into a pairing grouping layer, and the pairing grouping layer extracts the correlation features between the target and the shadow from the multi-scale features based on a pairing query mechanism to obtain paired output embeddings; Input the pixel embeddings and the paired output embeddings into a contrast morphological alignment module for morphological alignment to obtain segmentation instances of the object and the shadow.
[0055] It should be understood that when the processor 801 executes a program for constructing an exemplary shadow detection model in the memory 802, in addition to the above functions, other functions can also be implemented. For specific details, reference can be made to the descriptions of the corresponding method embodiments above.
[0056] Furthermore, the embodiments of the present invention do not specifically limit the type of the mentioned electronic device 800. The electronic device 800 can be a mobile phone, a tablet computer, a personal digital assistant (PDA), a wearable device, a laptop computer, or other portable electronic devices. Exemplary embodiments of the portable electronic device include, but are not limited to, portable electronic devices running IOS, android, microsoft, or other operating systems. The above portable electronic devices can also be other portable electronic devices, such as a laptop computer with a touch-sensitive surface (such as a touch panel). It should also be understood that in some other embodiments of the present invention, the electronic device 800 may not be a portable electronic device, but a desktop computer with a touch-sensitive surface (such as a touch panel).
[0057] Correspondingly, an embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium is used to store computer-readable programs or instructions. When the programs or instructions are executed by a processor, the steps or functions in the construction of the exemplary shadow detection model provided by the above method embodiments can be implemented.
[0058] Those skilled in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The computer program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory, or a random access memory, etc.
[0059] The above has introduced in detail the method for constructing an example shadow detection model provided by the present invention and the method for detecting an example shadow. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for detecting instance shadows, characterized in that, Including: Group and label the objects and shadows in the target scene image to obtain the labeled image; Input the labeled image into the trained instance shadow detection model to obtain the instance shadow of the target scene image; Among them, the trained instance shadow detection model includes: Construct a dataset under different light source scenarios; Input the dataset into a preset deep learning model to obtain high-level features; Input the high-level features into a pixel decoder to obtain pixel embeddings and multi-scale features; Input the multi-scale features into a pairing grouping layer, and the pairing grouping layer extracts the correlation features between the target and the shadow from the multi-scale features based on the pairing query mechanism to obtain paired output embeddings; Input the pixel embeddings and the paired output embeddings into a contrast morphological alignment module for morphological alignment to obtain the segmentation instances of the object and the shadow.
2. The method for detecting an example shadow according to claim 1, wherein The preset deep learning model includes one or more of ResNet, MobileNet, and VGG.
3. The method for detecting an example shadow according to claim 1, wherein, After constructing the dataset under different light source scenarios, it further includes: Perform data augmentation on the dataset; Group and label the objects and shadows in the dataset to obtain the preprocessed dataset.
4. The method for detecting an example shadow according to claim 1, wherein, The pairing query mechanism is to modify the query vector of the decoder in a preset neural network framework into a paired query vector.
5. The method for detecting an example shadow according to claim 4, wherein, The pairing grouping layer includes a pairwise mask attention mechanism and a self-attention mask mechanism. The pairwise mask attention mechanism includes a self-mask attention mechanism and a cross-mask attention mechanism. The pairing grouping layer extracts the correlation features between the target and the shadow from the multi-scale features based on the pairing query mechanism to obtain paired output embeddings, including: The cross-mask attention mechanism performs cross-scale feature interaction and fusion on the multi-scale features based on the paired query vector to obtain the first correlation feature; The self-mask attention mechanism performs further feature extraction and correlation on the first correlation feature to obtain the second correlation feature; Input the second correlation feature into the self-attention mask mechanism to obtain the paired output embeddings.
6. The method for detecting an example shadow according to claim 1, wherein, The step of inputting the pixel embeddings and the paired output embeddings into a contrast morphological alignment module for morphological alignment to obtain the segmentation instances of the object and the shadow includes: The contrast morphological alignment module performs a multiplication operation on the pixel embeddings and the paired output embeddings to obtain the separation instances of the object and the shadow.
7. The method for detecting an instance shadow according to claim 6, wherein It further includes: Optimize the edges of the separation instances of the object and the shadow to obtain the aligned separation instances of the object and the shadow; The step of optimizing the edges of the separation instances of the object and the shadow includes: Obtain a detail mask based on the separation instances of the object and the shadow; The detail mask undergoes a two-dimensional discrete convolution transform to obtain the frequency domain representation of the separation instances; Perform morphological alignment on the separation instances based on the frequency domain representation to obtain the aligned separation instances of the object and the shadow.
8. A detection device for instance shadows, characterized in that, Including: An image annotation module for grouping and labeling the objects and shadows in the target scene image to obtain the labeled image; An instance shadow segmentation module for inputting the labeled image into the trained instance shadow detection model to obtain the instance shadow of the target scene image; Among them, the trained instance shadow detection model includes: Construct a dataset under different light source scenarios; Input the dataset into a preset deep learning model to obtain high-level features; Input the high-level features into a pixel decoder to obtain pixel embeddings and multi-scale features; Input the multi-scale features into a paired grouping layer, and the paired grouping layer extracts the association features between the target and the shadow from the multi-scale features based on the paired query mechanism to obtain paired output embeddings; Input the pixel embeddings and the paired output embeddings into a contrast morphological alignment module for morphological alignment to obtain the segmentation instances of the object and the shadow.
9. An electronic device, characterized in that, It includes a memory and a processor, wherein, The memory is used to store programs; The processor is coupled to the memory and is used to execute the program stored in the memory to implement the steps in the method for detecting an instance shadow described in any one of claims 1 to 7 above.
10. A computer-readable storage medium, characterized in that, For storing computer-readable programs or instructions, when the programs or instructions are executed by a processor, they can implement the steps in the method for detecting an instance shadow described in any one of claims 1 to 7 above.