Stain detection method and device based on spectrum abnormal region perception

Through multi-spectral image processing and attention mechanism identification, the problem that intelligent sweeping robots cannot identify key contaminated areas is solved, and efficient stain detection and cleaning are achieved.

CN120375331AActive Publication Date: 2025-07-25QINGDAO TAPER ROBOTICS CO LTD +1
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510863934.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-25
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

Existing intelligent sweeping robots cannot effectively identify key contaminated areas, resulting in low cleaning efficiency and the inability to implement differentiated treatment for different stain types.

Method used

The stain detection method based on spectral abnormal area perception is adopted, and the feature map of the target area is obtained through multi-spectral images. The stain area is identified by the attention mechanism and the mask decoder, and the feature fusion is combined with the channel and spatial attention feature map to generate the stain mask.

Benefits of technology

It realizes accurate identification of stained areas, improves the cleaning efficiency of the sweeping robot, reduces repetitive work, and improves the cleaning effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375331A_ABST
    Figure CN120375331A_ABST
Patent Text Reader

Abstract

The invention provides a stain detection method and device based on spectrum abnormal region perception, and relates to the technical field of computer vision, and the method comprises the steps: inputting a second multispectral image of a region of interest into an image encoder, and obtaining a feature map of the region of interest, obtaining a channel attention feature map and a space attention feature map corresponding to the feature map by using an attention mechanism, and fusing the channel attention feature map and the space attention feature map to obtain a to-be-spliced feature map; splicing and fusing the principal component feature map and the to-be-spliced feature map through channel dimensions to obtain a combined feature map; and inputting the joint feature map into a mask decoder to obtain a stain mask, and determining a stain area in the target area based on the stain mask. According to the stain detection method and device based on spectrum abnormal area perception, the stains on the ground are accurately recognized through the multispectral image, so that the sweeping robot can carry out key cleaning on the stain area, and the cleaning efficiency of the sweeping robot is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular, to a stain detection method and device based on spectral anomaly region perception. Background Art

[0002] With the rapid development of artificial intelligence and Internet of Things technologies, smart home cleaning devices have gradually achieved basic environmental perception and path planning capabilities. However, there are still significant limitations in the existing technologies. For example, although intelligent floor sweeping robots already have basic spatial navigation capabilities, their cleaning logic still remains at the mechanized stage of "indiscriminate coverage". The device often performs repetitive cleaning on the entire house floor according to a preset program, which can neither actively identify key pollution areas nor implement differential treatment for different stain types, resulting in low cleaning efficiency.

[0003] Based on this, there is an urgent need for a method that can identify dirty areas to reduce the repetitive work of the floor sweeping robot and improve the cleaning efficiency. Summary of the Invention

[0004] The purpose of this application is to provide a stain detection method and device based on spectral anomaly region perception, which can accurately identify stains on the ground through multi-spectral images, enabling the floor sweeping robot to focus on cleaning the stain areas and improving the cleaning efficiency of the floor sweeping robot.

[0005] This application provides a stain detection method based on spectral anomaly region perception, including: Obtain a first multi-spectral image of a target area, and after masking non-concerned areas in the first multi-spectral image, obtain a second multi-spectral image; the first multi-spectral image includes spectral images corresponding to multiple different bands, and one band corresponds to one spectral image; input the second multi-spectral image into a first image encoder to obtain a feature map of the concerned area, and use an attention mechanism to obtain a channel attention feature map and a spatial attention feature map corresponding to the feature map of the concerned area; the number of channels of the first image encoder is the same as the number of bands of the second multi-spectral image; after fusing the channel attention feature map and the spatial attention feature map, obtain a feature map to be stitched, and perform stitching fusion on the principal component feature map of the second multi-spectral image and the feature map to be stitched through the channel dimension to obtain a joint feature map; input the joint feature map into a first mask decoder to obtain a stain mask, and determine the stain area in the target area based on the stain mask.

[0006] Optionally, after masking the non - concerned regions in the first multispectral image to obtain a second multispectral image, it includes: inputting the multispectral image into a pre - trained second image encoder to obtain a multispectral feature map, and inputting the multispectral feature map into a text encoder to obtain a semantic feature vector; inputting the semantic feature vector as a query, the multispectral feature map as keys and values into a cross - attention network to obtain a fusion feature corresponding to the multispectral image; inputting the fusion feature into a second mask decoder to obtain a non - concerned region image mask, and multiplying the non - concerned region image mask with the spectral images in the first multispectral image respectively to obtain the second multispectral image of the target region; wherein, the non - concerned region image mask is used to mask the non - concerned regions in the first multispectral image.

[0007] Optionally, inputting the second multispectral image into a first image encoder to obtain a feature map of the concerned region includes: inputting the second multispectral image into the first image encoder, using the shallower layers in the first image encoder to capture local texture features and using the deeper layers in the first image encoder to expand the receptive field, finally obtaining the feature map of the concerned region; wherein, the shallower layers include: convolutional kernels of a preset size; the deeper layers include: dilated convolutions with a preset dilation rate.

[0008] Optionally, using the attention mechanism to obtain a channel attention feature map corresponding to the feature map of the concerned region includes: calculating the global statistic of each channel based on the feature values of each channel at any pixel position in the feature map of the concerned region; calculating a channel attention weight vector used to characterize the importance of each channel based on the global statistic of each channel; performing an exclusive - OR operation between the channel attention weight vector and the feature map of the concerned region to obtain the channel attention feature map.

[0009] Optionally, using the attention mechanism to obtain a spatial attention feature map corresponding to the feature map of the concerned region includes: calculating a first feature map obtained by average pooling the feature map of the concerned region along the spatial dimension and a second feature map obtained by max - pooling the feature map of the concerned region along the spatial dimension; performing a concatenation operation on the first feature map and the second feature map in the spatial dimension to generate a spatial attention weight; performing an exclusive - OR operation between the spatial attention weight vector and the feature map of the concerned region to obtain the spatial attention feature map.

[0010] Optionally, inputting the joint feature map into a first mask decoder to obtain a stain mask, and determining a stain area within the target area based on the stain mask includes: Upsampling the joint feature map and then inputting it into the first mask decoder to obtain a preliminary stain mask; Using morphological post-processing to process the preliminary stain mask and perform refined adjustment on the preliminary stain mask to obtain a final stain mask; The morphological post-processing includes: Performing an erosion-dilation sequence operation using an elliptical kernel with a size adapted to the stain area; After multiplying the final stain mask by the second multispectral image, obtaining a stain multispectral image, and determining the stain area within the target area based on the stain multispectral image.

[0011] This application also provides a stain detection device based on spectral anomaly area perception, including: An acquisition module, configured to acquire a first multispectral image of a target area; An image processing module, configured to mask non-concerned areas in the first multispectral image to obtain a second multispectral image; The first multispectral image includes: spectral images corresponding to multiple different bands, and one band corresponds to one spectral image; A feature extraction module, configured to input the second multispectral image into a first image encoder to obtain a feature map of the concerned area, and use an attention mechanism to obtain a channel attention feature map and a spatial attention feature map corresponding to the feature map of the concerned area; The number of channels of the first image encoder is the same as the number of bands of the second multispectral image; The feature extraction module is further configured to fuse the channel attention feature map and the spatial attention feature map to obtain a feature map to be stitched, and perform stitching fusion in the channel dimension on the principal component feature map of the second multispectral image and the feature map to be stitched to obtain a joint feature map; The image processing module is further configured to input the joint feature map into a first mask decoder to obtain a stain mask; A stain detection module, configured to determine a stain area within the target area based on the stain mask.

[0012] Optionally, the image processing module is specifically configured to input the multispectral image into a pre-trained second image encoder to obtain a multispectral feature map, and input the multispectral feature map into a text encoder to obtain a semantic feature vector; The image processing module is specifically further configured to use the semantic feature vector as a query, the multispectral feature map as keys and values, and input them into a cross-attention network to obtain a fusion feature corresponding to the multispectral image; The image processing module is specifically further configured to input the fusion feature into a second mask decoder to obtain an image mask of the non-concerned area, and multiply the image mask of the non-concerned area by the spectral images in the first multispectral image respectively to obtain the second multispectral image of the target area; Wherein, the image mask of the non-concerned area is used to mask the non-concerned areas in the first multispectral image.

[0013] Optionally, the feature extraction module is specifically configured to input the second multi-spectral image into the first image encoder, use the shallow layer in the first image encoder to capture local texture features, and use the deep layer in the first image encoder to expand the receptive field, and finally obtain the feature map of the region of interest; wherein, the shallow layer includes: a convolutional kernel with a preset size; the deep layer includes: atrous convolution with a preset dilation rate.

[0014] Optionally, the feature extraction module is specifically configured to calculate the global statistic of each channel based on the feature value of each channel at any pixel position in the feature map of the region of interest; the feature extraction module is further specifically configured to calculate the channel attention weight vector used to characterize the importance of each channel based on the global statistic of each channel; the feature extraction module is further specifically configured to perform an exclusive NOR operation on the channel attention weight vector and the feature map of the region of interest to obtain the channel attention feature map.

[0015] Optionally, the feature extraction module is specifically configured to calculate the first feature map after average pooling of the feature map of the region of interest along the spatial dimension and the second feature map after max pooling along the spatial dimension; the feature extraction module is further specifically configured to perform a splicing operation on the first feature map and the second feature map in the spatial dimension to generate a spatial attention weight; the feature extraction module is further specifically configured to perform an exclusive NOR operation on the spatial attention weight vector and the feature map of the region of interest to obtain the spatial attention feature map.

[0016] Optionally, the image processing module is specifically configured to perform upsampling on the joint feature map and then input it into the first mask decoder to obtain a preliminary stain mask; the image processing module is specifically configured to use morphological post-processing to process the preliminary stain mask and perform fine-tuning on the preliminary stain mask to obtain a final stain mask; the morphological post-processing includes: performing an erosion-dilation sequence operation using an elliptical kernel with an area adaptive to the stain size; the image processing module is specifically configured to multiply the final stain mask by the second multi-spectral image to obtain a stain multi-spectral image; the stain detection module is specifically configured to determine the stain region in the target region based on the stain multi-spectral image.

[0017] The present application also provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the stain detection method based on spectral anomaly region perception as described in any one of the above are implemented.

[0018] The present application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the stain detection method based on spectral anomaly region perception as described in any one of the above are implemented.

[0019] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the stain detection method based on spectral anomaly region perception as described in any one of the above are implemented.

[0020] The stain detection method and device based on spectral anomaly region perception provided by the present application obtain a first multispectral image of a target region, and after masking the non-concerned regions in the first multispectral image, a second multispectral image is obtained; the first multispectral image includes spectral images corresponding to multiple different bands, and one band corresponds to one spectral image; the second multispectral image is input into a first image encoder to obtain a feature map of the concerned region, and a channel attention feature map and a spatial attention feature map corresponding to the feature map of the concerned region are obtained by using an attention mechanism; the number of channels of the first image encoder is the same as the number of bands of the second multispectral image; after fusing the channel attention feature map and the spatial attention feature map, a feature map to be stitched is obtained, and the principal component feature map of the second multispectral image and the feature map to be stitched are fused by stitching in the channel dimension to obtain a joint feature map; the joint feature map is input into a first mask decoder to obtain a stain mask, and based on the stain mask, the stain region in the target region is determined. In this way, the stains on the ground are accurately identified through the multispectral image, enabling the sweeping robot to focus on cleaning the stain region and improving the cleaning efficiency of the sweeping robot. Description of the Drawings

[0021] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 It is a schematic diagram of the hardware environment of an interaction method of an intelligent device according to an embodiment of the present application; Figure 2 It is a flowchart of the stain detection method based on spectral anomaly region perception provided by the present application; Figure 3 It is a schematic diagram of the structure of the stain detection device based on spectral anomaly region perception provided by the present application; Figure 4It is a schematic structural diagram of the electronic device provided by this application. Specific embodiments

[0023] In order to enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0025] According to one aspect of the embodiments of this application, a stain detection method based on spectral anomaly region perception is provided. This stain detection method based on spectral anomaly region perception is widely applied to whole-house intelligent digital control application scenarios such as Smart Home, smart home, smart home appliance ecosystem, Intelligence House ecosystem, etc. Optionally, in this embodiment, the above-mentioned stain detection method based on spectral anomaly region perception can be applied to, for example Figure 2 the hardware environment composed of the terminal device 102 and the server 104 as shown. As Figure 2 shown, the server 104 is connected to the terminal device 102 through the network, and can be used to provide services (such as application services, etc.) for the terminal or the client installed on the terminal. A database can be set on the server or independently of the server, and is used to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server, and are used to provide data operation services for the server 104.

[0026] The above network may include, but is not limited to, at least one of the following: a wired network, a wireless network. The above wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, a local area network. The above wireless network may include, but is not limited to, at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 is not limited to a PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projection device, smart TV, smart clothes hanger, smart curtain, smart audio and video, smart socket, smart speaker, smart sound box, smart fresh air device, smart kitchen and bathroom equipment, smart bathroom equipment, smart floor sweeping robot, smart window cleaning robot, smart mopping robot, smart air purification device, smart steam box, smart microwave oven, smart kitchen water heater, smart purifier, smart water dispenser, smart door lock, etc.

[0027] There are several existing technical routes and limitations in the stain detection methods in the related art: Image segmentation - feature extraction architecture: It improves the multi - target parallel processing efficiency through regional detection strategies, but fails to effectively distinguish stain features from complex backgrounds; RGB channel selective detection method: It only performs stain detection on the blue and green channels of the image, but there is still insufficient spectral resolution and it cannot capture the features of special stains in the near - infrared and short - wave infrared. Then, these methods are difficult to detect stains with similar reflection characteristics to the ground substrate, and are also prone to missing and misdetecting stains in a cluttered background, resulting in redundant cleaning path planning. The deficiencies of these existing technical solutions will bring an unpleasant experience to users.

[0028] In summary, only regional stain detection cannot solve the interference of the background, and only detecting stains on the blue and green channels cannot improve the detection of difficult stains. In order to enable the floor sweeping robot to accurately identify ground stains, reduce missing and misdetection, and thus achieve an accurate cleaning route of the floor sweeping robot guided by stain detection.

[0029] Based on this, the embodiment of the present application provides a stain detection method that fuses multi - spectral images. This method excludes the interference of cluttered backgrounds by finding the ground mask, strengthens the spatial - spectral information of the obtained image by fusing multi - spectral features and image enhancement algorithms, and constructs a stain detection system with strong anti - interference ability.

[0030] The following will combine the accompanying drawings and, through specific embodiments and their application scenarios, elaborate in detail on the stain detection method based on spectral anomaly region perception provided by the embodiment of the present application.

[0031] As Figure 2As shown, a stain detection method based on spectral anomaly region perception provided by an embodiment of the present application may include the following steps 201 to 204: Step 201: Obtain a first multi-spectral image of a target region, and after masking the non-concerned regions in the first multi-spectral image, obtain a second multi-spectral image.

[0032] Among them, the first multi-spectral image includes spectral images corresponding to multiple different bands, and one band corresponds to one spectral image.

[0033] Exemplarily, the above target region is any region on the cleaning path of a floor sweeping robot or a window cleaning robot. The floor sweeping robot or the window cleaning robot can collect the multi-spectral image of the target region through a sensor.

[0034] It should be noted that the multi-spectral image collected by the sensor contains spectral images corresponding to different bands. In the embodiment of the present application, the spectral images of 10 bands (that is, 10 spectral images of the target region) are taken as an example for description.

[0035] Exemplarily, after obtaining the above first multi-spectral image, in order to eliminate some interference and narrow the processing range, the non-concerned regions can be masked by a mask. Taking the floor sweeping robot as an example, the non-ground regions in the image can be masked.

[0036] Specifically, the step of masking the non-concerned regions in the first multi-spectral image in step 201 to obtain the second multi-spectral image may further include the following steps 201a1 to 201a3: Step 201a1: Input the multi-spectral image into a pre-trained second image encoder to obtain a multi-spectral feature map, and input the multi-spectral feature map into a text encoder to obtain a semantic feature vector.

[0037] Step 201a2: Input the semantic feature vector as a query, the multi-spectral feature map as keys and values into a cross-attention network to obtain the fusion feature corresponding to the multi-spectral image.

[0038] Step 201a3: Input the fusion feature into a second mask decoder to obtain the non-concerned region image mask, and multiply the non-concerned region image mask by the spectral images in the first multi-spectral image respectively to obtain the second multi-spectral image of the target region.

[0039] Among them, the non-concerned region image mask is used to mask the non-concerned regions in the first multi-spectral image.

[0040] Exemplarily, in the embodiments of the present application, a binary mask of the ground area is generated by using the cross-attention mechanism of semantics and image features. Through semantic guidance and multi-spectral feature interaction, non-ground interference targets are stripped, and the processing range is limited to the key ground area.

[0041] Exemplarily, in order to obtain the multi-spectral image of the region of interest, the first spectral image can be passed through the image encoder obtained by contrastive learning pre-training to obtain a multi-spectral feature map , where H is the width of the multi-spectral feature map, W is the height of the multi-spectral feature map, is the number of channels of the extracted multi-spectral feature map, as specifically shown in the following formula (1): (Formula 1) where, is the multi-spectral image feature extractor, that is, the above-mentioned second image encoder.

[0042] Exemplarily, after that, it is also necessary to input the obtained above and the static semantic feature vector generated by the pre-trained CLIP text encoder into the cross-attention module to obtain , as specifically shown in the following formula (2): (Formula 2) where, is the cross-attention module, the text feature is used as the query Query, and the image feature is used as the key Key and value Value, and the ground-related region is focused through the attention weight.

[0043] Exemplarily, input the feature obtained by the cross-attention module into the mask decoder (i.e., the above-mentioned second decoder) to obtain a binary mask (i.e., the above-mentioned non-region-of-interest image mask), as specifically shown below:

[0044]

[0045]

[0046] where, is the deconvolution upsampling network, , is the mask decoder, is the obtained weight.

[0047] Finally, input the binary mask Multiply them with the first multispectral image respectively to obtain the required ground multispectral image , that is, the above-mentioned second multispectral image.

[0048] It should be noted that in the training stage, the designed loss functions are Dice Loss (i.e., ) and Focal Loss (i.e., ), and the formulas are as follows:

[0049]

[0050]

[0051] Among them, is the probability that the pixel i is predicted as a stain, is the true label (1 for stain and 0 for background), =0.25, =2, which is the loss weight for suppressing simple samples (background), , with Dice Loss dominating the global structure and Focal Loss refining local details.

[0052] Step 202: Input the second multispectral image into the first image encoder to obtain a feature map of the region of interest, and use the attention mechanism to obtain a channel attention feature map and a spatial attention feature map corresponding to the feature map of the region of interest.

[0053] Among them, the number of channels of the first image encoder is the same as the number of bands of the second multispectral image.

[0054] Exemplarily, after obtaining the multispectral image of the region of interest, perform orthogonal decomposition on the multi-channel multispectral image by principal component analysis to extract low-dimensional principal component features, and introduce a channel and spatial double-branch attention mechanism to extract key features from the spectral sensitivity and spatial distribution characteristics respectively. Then, the channel branch adaptively enhances the stain-discriminative bands, and the spatial branch focuses on local abnormal textures. The principal component features and the double-attention features jointly capture the global spectral anomaly features, and the two are fused to obtain a joint feature map.

[0055] Specifically, the step of inputting the second multispectral image into the first image encoder in the above step 202 to obtain a feature map of the region of interest may further include the following step 202a: Step 202a: Input the second multispectral image into the first image encoder. Use the shallow layer in the first image encoder to capture local texture features and the deep layer in the first image encoder to expand the receptive field, and finally obtain the feature map of the region of interest.

[0056] Among them, the shallow layer includes: a convolution kernel with a preset size; the deep layer includes: a dilated convolution with a preset dilation rate.

[0057] Exemplarily, the obtained second multispectral image is input into the first image encoder to obtain a multispectral ground feature map , specifically as shown in the following formula three: (Formula Three) Among them, is an image encoder based on UNet. The input channels of the extended UNet are extended to ten-channel segments, and the pre-trained weights are initialized by copying the RGB channels. Design a shallow layer (3×3 convolution) to capture local textures (such as spots of oil stains and infiltration edges of water stains); deep layer (dilated convolution): use a dilation rate r = 2 to expand the receptive field to 15×15 to identify large-area stains.

[0058] Specifically, in the above step 202, the step of obtaining the channel attention feature map corresponding to the feature map of the region of interest by using the attention mechanism may further include the following steps 202b1 to 202b3: Step 202b1: Calculate the global statistic of each channel based on the feature value of each channel in the feature map of the region of interest at any pixel position.

[0059] Step 202b2: Calculate the channel attention weight vector used to represent the importance of each channel based on the global statistic of each channel.

[0060] Step 202b3: After performing an exclusive OR operation on the channel attention weight vector and the feature map of the region of interest, obtain the channel attention feature map.

[0061] Exemplarily, the obtained multispectral ground feature map is input into the dual attention module to obtain a channel attention feature map and a spatial attention feature map using the dual attention mechanism.

[0062] Exemplarily, the channel attention formula is as follows:

[0063] Among them, is the feature value of the c-th channel at position The eigenvalue is the global statistic of the c-th channel. is the weight matrix of the fully connected layer. is the ReLU activation function. is the Sigmoid function, and s is the channel attention weight vector, representing the importance of each channel.

[0064] Specifically, in the above step 202, the steps of obtaining the spatial attention feature map corresponding to the feature map of the region of interest by using the attention mechanism may further include the following steps 202c1 to 202c3: Step 202c1: Calculate the first feature map after average pooling of the feature map of the region of interest along the spatial dimension and the second feature map after max pooling along the spatial dimension.

[0065] Step 202c2: After performing a concatenation operation on the first feature map and the second feature map in the spatial dimension, generate a spatial attention weight.

[0066] Step 202c3: After performing an exclusive OR operation on the spatial attention weight vector and the feature map of the region of interest, obtain the spatial attention feature map.

[0067] Exemplarily, the spatial attention formula is as follows:

[0068] where is the feature map after average pooling along the spatial dimension, is the feature map after max pooling along the spatial dimension. The features obtained by the dual attention mechanism .

[0069] Step 203: After fusing the channel attention feature map and the spatial attention feature map, obtain a feature map to be concatenated, and concatenate and fuse the principal component feature map of the second multispectral image and the feature map to be concatenated through the channel dimension to obtain a joint feature map.

[0070] Exemplarily, perform principal component analysis (PCA) on the obtained second multispectral image to reduce the dimension and extract the global spectral principal components , specifically as shown in the following formula:

[0071] where , is the eigenvector matrix corresponding to the first 3 largest eigenvalues.

[0072] Exemplarily, the principal component feature map obtained by principal component analysis and the features obtained by the dual attention mechanism module are spliced and fused in the channel dimension to form a joint feature map .

[0073]

[0074] Among them, .

[0075] Step 204: Input the joint feature map into the first mask decoder to obtain a stain mask, and determine the stain area within the target area based on the stain mask.

[0076] Exemplarily, after inputting the joint feature map into the decoding network, a preliminary high-resolution stain mask is generated to distinguish the real stain area from the ground area. Morphological post-processing is performed on the obtained preliminary mask to eliminate isolated noise points by first eroding and then dilating, and then fill internal holes by first dilating and then eroding to optimize the smoothness and continuity of the mask boundary, and finally output the stain localization result.

[0077] Specifically, the above step 204 may further include the following steps 204a1 to 204a3: Step 204a1: Upsample the joint feature map and then input it into the first mask decoder to obtain a preliminary stain mask.

[0078] Step 204a2: Process the preliminary stain mask using morphological post-processing to perform refined adjustment on the preliminary stain mask to obtain a final stain mask.

[0079] Among them, the morphological post-processing includes: performing an erosion-dilation sequence operation using an elliptical kernel with a stain area-adaptive size.

[0080] Step 204a3: Multiply the final stain mask by the second multispectral image to obtain a stain multispectral image, and determine the stain area within the target area based on the stain multispectral image.

[0081] Exemplarily, input the joint feature map into the mask decoder to obtain a high-resolution preliminary stain mask , as follows:

[0082] Among them, is a transposed convolution upsampling network , is the mask decoder is the obtained weight.

[0083] It should be noted that in the training stage, the loss functions are also designed as Dice Loss and Focal Loss, and the specific formulas are as follows:

[0084] Exemplarily, after initially obtaining the stain area mask, the mask is refined by a morphological post-processing optimization module to obtain the final high-resolution stain mask, and the specific process is as follows: An erosion-dilation sequence is performed using an N×N elliptical kernel (N is adaptively adjusted according to the stain area) to eliminate isolated noise points, fill holes (such as hair occlusion areas), and at the same time retain the smoothness of the stain edges. The obtained is multiplied by to obtain the final multi-spectral image of the stain, that is, the stain area is found.

[0085] In a possible implementation, in the training stage, a CM020 series hyperspectral imaging device is used for data acquisition. This camera has a wide spectral response range of 350 - 950 nm, and the imaging sensor is configured with a 1600×1200 pixel array. To overcome the interference of specular reflected light on the liquid pollutant surface to the detection result, a polarization filter component is placed at the front end of the optical lens, and a depolarization optical path is formed by adjusting the angle of two combined polarizers to suppress the specular reflection interference caused by environmental light on the liquid surface and smooth ground.

[0086] Exemplarily, the multi-spectral camera captures color images, gray images, and qs format files with multi-spectral information. A bmp format multi-spectral image is obtained through spectral inversion software. When inverting, the band interval and the number of required bands can be selected. Here, we select the band interval of 410 - 950, and a total of ten bands are inverted to finally obtain the multi-spectral image we captured.

[0087] After that, the obtained multi-spectral image is preprocessed, including: geometric correction and spatial transformation to denoise the multi-spectral image; scaling the multi-spectral image to 640×640 pixels, which is the multi-spectral image obtained through preprocessing. The obtained multi-spectral image is used to train the above-mentioned encoders and decoders.

[0088] Compared with the stain detection solutions in related technologies, the stain detection method based on spectral anomaly region perception provided by the embodiments of this application has the following advantages: ① By introducing ten-band (410-950nm) multispectral data, the specific reflection characteristics of stains under different spectra are captured, solving the problem of misdetection of similar textures by traditional RGB images and improving the cross-material generalization ability of the model. ② A combination of polarizers is used to suppress specular reflection noise, effectively eliminating the highlight interference on the surfaces of smooth floors and liquid pollutants. ③ A cross-attention based on semantic features and multispectral images is constructed to generate an accurate ground mask, excluding non-ground interference regions such as furniture and shadows, which helps to find high-precision stain regions subsequently. ④ A principal component analysis combined with a dual-attention mechanism is designed to focus on spectral anomaly regions and improve the feature decoding efficiency.

[0089] The stain detection method based on spectral anomaly region perception provided by the embodiments of this application acquires a first multispectral image of a target region, and after masking the non-concerned regions in the first multispectral image, a second multispectral image is obtained; the first multispectral image includes: spectral images corresponding to multiple different bands, and one band corresponds to one spectral image; the second multispectral image is input into a first image encoder to obtain a feature map of the concerned region, and a channel attention feature map and a spatial attention feature map corresponding to the feature map of the concerned region are obtained by using an attention mechanism; the number of channels of the first image encoder is the same as the number of bands of the second multispectral image; after fusing the channel attention feature map and the spatial attention feature map, a feature map to be stitched is obtained, and the principal component feature map of the second multispectral image and the feature map to be stitched are fused through stitching in the channel dimension to obtain a joint feature map; the joint feature map is input into a first mask decoder to obtain a stain mask, and based on the stain mask, the stain region in the target region is determined. In this way, the stains on the ground are accurately identified through multispectral images, enabling the sweeping robot to focus on cleaning the stain regions and improving the cleaning efficiency of the sweeping robot.

[0090] It should be noted that for the stain detection method based on spectral anomaly region perception provided by the embodiments of this application, the execution subject can be a stain detection device based on spectral anomaly region perception, or a control module in the stain detection device based on spectral anomaly region perception for executing the stain detection method based on spectral anomaly region perception. In the embodiments of this application, taking the stain detection device based on spectral anomaly region perception executing the stain detection method based on spectral anomaly region perception as an example, the stain detection device based on spectral anomaly region perception provided by the embodiments of this application is described.

[0091] It should be noted that in the embodiments of the present application, the stain detection methods shown in the above-mentioned various method drawings based on spectral anomaly region perception are all exemplified by taking one of the drawings in the embodiments of the present application as an example. Specifically, when implemented, the stain detection methods based on spectral anomaly region perception shown in the above-mentioned various method drawings can also be implemented in combination with any other combinable drawings schematically shown in the above embodiments, which will not be elaborated here.

[0092] The stain detection device based on spectral anomaly region perception provided by the present application will be described below, and the following description can be mutually referred to corresponding to the stain detection method based on spectral anomaly region perception described above.

[0093] Figure 3 It is a schematic structural diagram of the stain detection device based on spectral anomaly region perception provided by the embodiments of the present application, as Figure 3 shown, specifically including: An acquisition module 301, configured to acquire a first multi-spectral image of a target area; an image processing module 302, configured to mask non-concerned areas in the first multi-spectral image to obtain a second multi-spectral image; the first multi-spectral image includes: spectral images corresponding to multiple different bands, and one band corresponds to one spectral image; a feature extraction module 303, configured to input the second multi-spectral image into a first image encoder to obtain a feature map of the concerned area, and use an attention mechanism to obtain a channel attention feature map and a spatial attention feature map corresponding to the feature map of the concerned area; the number of channels of the first image encoder is the same as the number of bands of the second multi-spectral image; the feature extraction module 303 is further configured to fuse the channel attention feature map and the spatial attention feature map to obtain a feature map to be stitched, and perform stitching fusion on the principal component feature map of the second multi-spectral image and the feature map to be stitched through the channel dimension to obtain a joint feature map; the image processing module 302 is further configured to input the joint feature map into a first mask decoder to obtain a stain mask; a stain detection module 304, configured to determine a stain area in the target area based on the stain mask.

[0094] Optionally, the image processing module 302 is specifically configured to input the multi-spectral image into a pre-trained second image encoder to obtain a multi-spectral feature map, and input the multi-spectral feature map into a text encoder to obtain a semantic feature vector; the image processing module 302 is further specifically configured to use the semantic feature vector as a query, the multi-spectral feature map as keys and values and input them into a cross-attention network to obtain a fusion feature corresponding to the multi-spectral image; the image processing module 302 is further specifically configured to input the fusion feature into a second mask decoder to obtain an image mask of the non-attended area, and multiply the image mask of the non-attended area by the spectral image in the first multi-spectral image respectively to obtain a second multi-spectral image of the target area; wherein, the image mask of the non-attended area is used to mask the non-attended area in the first multi-spectral image.

[0095] Optionally, the feature extraction module 303 is specifically configured to input the second multi-spectral image into the first image encoder, use the shallower layers in the first image encoder to capture local texture features, and use the deeper layers in the first image encoder to expand the receptive field, and finally obtain a feature map of the attended area; wherein, the shallower layers include: convolutional kernels of a preset size; the deeper layers include: dilated convolutions with a preset dilation rate.

[0096] Optionally, the feature extraction module 303 is specifically configured to calculate global statistics for each channel based on the feature values of each channel at any pixel position in the feature map of the attended area; the feature extraction module 303 is further specifically configured to calculate a channel attention weight vector for characterizing the importance of each channel based on the global statistics of each channel; the feature extraction module 303 is further specifically configured to perform an exclusive NOR operation on the channel attention weight vector and the feature map of the attended area to obtain a channel attention feature map.

[0097] Optionally, the feature extraction module 303 is specifically configured to calculate a first feature map obtained by average pooling the feature map of the attended area along the spatial dimension and a second feature map obtained by max pooling the feature map of the attended area along the spatial dimension; the feature extraction module 303 is further specifically configured to perform a concatenation operation on the first feature map and the second feature map in the spatial dimension to generate a spatial attention weight; the feature extraction module 303 is further specifically configured to perform an exclusive NOR operation on the spatial attention weight vector and the feature map of the attended area to obtain a spatial attention feature map.

[0098] Optionally, the image processing module 302 is specifically configured to upsample the joint feature map and input it into the first mask decoder to obtain a preliminary stain mask; the image processing module 302 is specifically configured to use post-morphological processing to process the preliminary stain mask and perform fine-tuning on the preliminary stain mask to obtain a final stain mask; the post-morphological processing includes: performing an erosion-dilation sequence operation using an elliptical kernel with a stain area-adaptive size; the image processing module 302 is specifically configured to multiply the final stain mask by the second multispectral image to obtain a stain multispectral image; the stain detection module 304 is specifically configured to determine a stain area within the target area based on the stain multispectral image.

[0099] The stain detection device based on spectral anomaly region perception provided in this application acquires a first multispectral image of a target area, and after masking non-concerned areas in the first multispectral image, obtains a second multispectral image; the first multispectral image includes: spectral images corresponding to multiple different bands, and one band corresponds to one spectral image; inputting the second multispectral image into a first image encoder to obtain a feature map of the concerned area, and using an attention mechanism to obtain a channel attention feature map and a spatial attention feature map corresponding to the feature map of the concerned area; the number of channels of the first image encoder is the same as the number of bands of the second multispectral image; fusing the channel attention feature map and the spatial attention feature map to obtain a feature map to be stitched, and performing stitching fusion on the principal component feature map of the second multispectral image and the feature map to be stitched through the channel dimension to obtain a joint feature map; inputting the joint feature map into a first mask decoder to obtain a stain mask, and determining a stain area within the target area based on the stain mask. In this way, the stains on the ground are accurately identified through the multispectral image, enabling the sweeping robot to focus on cleaning the stain area and improving the cleaning efficiency of the sweeping robot.

[0100] Figure 4 An entity structure diagram of an electronic device is illustrated, such as Figure 4As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communications interface 420, and the memory 430 complete communication with each other through the communication bus 440. The processor 410 may call logical instructions in the memory 430 to execute a stain detection method based on the perception of spectral abnormal regions. The method includes: obtaining a first multispectral image of a target region, and after masking the non-concerned regions in the first multispectral image, obtaining a second multispectral image; the first multispectral image includes spectral images corresponding to multiple different bands, and one band corresponds to one spectral image; inputting the second multispectral image into a first image encoder to obtain a feature map of the concerned region, and using an attention mechanism to obtain a channel attention feature map and a spatial attention feature map corresponding to the feature map of the concerned region; the number of channels of the first image encoder is the same as the number of bands of the second multispectral image; after fusing the channel attention feature map and the spatial attention feature map, obtaining a feature map to be stitched, and stitchedly fusing the principal component feature map of the second multispectral image and the feature map to be stitched through the channel dimension to obtain a joint feature map; inputting the joint feature map into a first mask decoder to obtain a stain mask, and determining the stain region in the target region based on the stain mask. In this way, the stains on the ground are accurately identified through the multispectral image, enabling the sweeping robot to focus on cleaning the stain region and improving the cleaning efficiency of the sweeping robot.

[0101] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0102] On the other hand, the present application also provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the stain detection method based on spectral anomaly region perception provided by the above-mentioned various methods. The method includes: obtaining a first multi-spectral image of a target area, and masking non-concerned areas in the first multi-spectral image to obtain a second multi-spectral image; the first multi-spectral image includes spectral images corresponding to multiple different bands, and one band corresponds to one spectral image; inputting the second multi-spectral image into a first image encoder to obtain a feature map of the concerned area, and using an attention mechanism to obtain a channel attention feature map and a spatial attention feature map corresponding to the feature map of the concerned area; the number of channels of the first image encoder is the same as the number of bands of the second multi-spectral image; fusing the channel attention feature map and the spatial attention feature map to obtain a feature map to be stitched, and splicing and fusing the principal component feature map of the second multi-spectral image and the feature map to be stitched through the channel dimension to obtain a joint feature map; inputting the joint feature map into a first mask decoder to obtain a stain mask, and determining the stain area in the target area based on the stain mask. In this way, the stains on the ground are accurately identified through the multi-spectral image, enabling the sweeping robot to focus on cleaning the stain area and improving the cleaning efficiency of the sweeping robot.

[0103] On another aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the stain detection method based on spectral anomaly region perception provided by the above-mentioned various methods. The method includes: obtaining a first multi-spectral image of a target area, and masking non-concerned areas in the first multi-spectral image to obtain a second multi-spectral image; the first multi-spectral image includes spectral images corresponding to multiple different bands, and one band corresponds to one spectral image; inputting the second multi-spectral image into a first image encoder to obtain a feature map of the concerned area, and using an attention mechanism to obtain a channel attention feature map and a spatial attention feature map corresponding to the feature map of the concerned area; the number of channels of the first image encoder is the same as the number of bands of the second multi-spectral image; fusing the channel attention feature map and the spatial attention feature map to obtain a feature map to be stitched, and splicing and fusing the principal component feature map of the second multi-spectral image and the feature map to be stitched through the channel dimension to obtain a joint feature map; inputting the joint feature map into a first mask decoder to obtain a stain mask, and determining the stain area in the target area based on the stain mask. In this way, the stains on the ground are accurately identified through the multi-spectral image, enabling the sweeping robot to focus on cleaning the stain area and improving the cleaning efficiency of the sweeping robot.

[0104] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0105] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A stain detection method based on the perception of spectral abnormal regions, characterized in that, Including: Obtain a first multispectral image of a target area, and after masking non - concerned areas in the first multispectral image, obtain a second multispectral image; The first multispectral image includes spectral images corresponding to multiple different bands, and one band corresponds to one spectral image; Input the second multispectral image into a first image encoder to obtain a feature map of the concerned area, and use an attention mechanism to obtain a channel attention feature map and a spatial attention feature map corresponding to the feature map of the concerned area; the number of channels of the first image encoder is the same as the number of bands of the second multispectral image; After fusing the channel attention feature map and the spatial attention feature map, obtain a feature map to be stitched, and fuse the principal component feature map of the second multispectral image and the feature map to be stitched through stitching in the channel dimension to obtain a joint feature map; Input the joint feature map into a first mask decoder to obtain a stain mask, and determine the stain area in the target area based on the stain mask.

2. The stain detection method based on spectral anomaly region perception according to claim 1, characterized in that The step of masking non - concerned areas in the first multispectral image to obtain a second multispectral image includes: Input the multispectral image into a pre - trained second image encoder to obtain a multispectral feature map, and input the multispectral feature map into a text encoder to obtain a semantic feature vector; Use the semantic feature vector as a query, the multispectral feature map as keys and values, and input them into a cross - attention network to obtain a fusion feature corresponding to the multispectral image; Input the fusion feature into a second mask decoder to obtain a non - concerned area image mask, and multiply the non - concerned area image mask with the spectral images in the first multispectral image respectively to obtain the second multispectral image of the target area; Wherein, the non - concerned area image mask is used to mask non - concerned areas in the first multispectral image.

3. The stain detection method based on spectral anomaly region perception according to claim 2, wherein The step of inputting the second multispectral image into a first image encoder to obtain a feature map of the concerned area includes: Input the second multispectral image into the first image encoder, use the shallower layer in the first image encoder to capture local texture features, and use the deeper layer in the first image encoder to expand the receptive field, and finally obtain the feature map of the concerned area; Wherein, the shallower layer includes a convolutional kernel with a preset size; the deeper layer includes atrous convolution with a preset dilation rate.

4. The stain detection method based on spectral anomaly region perception according to any one of claims 1 to 3, characterized in that, The step of using the attention mechanism to obtain a channel attention feature map corresponding to the feature map of the concerned area includes: Based on the feature values of each channel at any pixel position in the feature map of the concerned area, calculate the global statistics of each channel; Based on the global statistics of each channel, calculate a channel attention weight vector used to characterize the importance of each channel; After performing an exclusive NOR operation on the channel attention weight vector and the feature map of the concerned area, obtain the channel attention feature map.

5. The stain detection method based on spectral anomaly region perception according to any one of claims 1 to 3, characterized in that The step of using the attention mechanism to obtain a spatial attention feature map corresponding to the feature map of the concerned area includes: Calculate a first feature map after average pooling of the feature map of the region of interest along the spatial dimension and a second feature map after max pooling along the spatial dimension; After performing a concatenation operation on the spatial dimension of the first feature map and the second feature map, generate a spatial attention weight; After performing an exclusive NOR operation on the spatial attention weight vector and the feature map of the region of interest, obtain the spatial attention feature map.

6. The stain detection method based on spectral anomaly region perception according to claim 1, wherein The inputting the joint feature map into a first mask decoder to obtain a stain mask, and determining the stain region within the target region based on the stain mask includes: Upsample the joint feature map and input it into the first mask decoder to obtain a preliminary stain mask; Use morphological post-processing to process the preliminary stain mask and perform fine-tuning on the preliminary stain mask to obtain a final stain mask; the morphological post-processing includes: performing an erosion-dilation sequence operation using an elliptical kernel with a size adapted to the stain area; Multiply the final stain mask by the second multispectral image to obtain a stain multispectral image, and determine the stain region within the target region based on the stain multispectral image.

7. A stain detection device based on the perception of spectral abnormal regions, characterized in that, The apparatus includes: An acquisition module, configured to acquire a first multispectral image of a target region; An image processing module, configured to mask a non-region of interest in the first multispectral image to obtain a second multispectral image; the first multispectral image includes: spectral images corresponding to multiple different bands, and one spectral image corresponds to one band; A feature extraction module, configured to input the second multispectral image into a first image encoder to obtain a feature map of the region of interest, and use an attention mechanism to obtain a channel attention feature map and a spatial attention feature map corresponding to the feature map of the region of interest; the number of channels of the first image encoder is the same as the number of bands of the second multispectral image; The feature extraction module is further configured to fuse the channel attention feature map and the spatial attention feature map to obtain a feature map to be concatenated, and fuse the principal component feature map of the second multispectral image and the feature map to be concatenated through concatenation in the channel dimension to obtain a joint feature map; The image processing module is further configured to input the joint feature map into a first mask decoder to obtain a stain mask; A stain detection module, configured to determine the stain region within the target region based on the stain mask.

8. The stain detection apparatus based on spectral anomaly region perception according to claim 7, wherein The image processing module is specifically configured to input the multispectral image into a pre-trained second image encoder to obtain a multispectral feature map, and input the multispectral feature map into a text encoder to obtain a semantic feature vector; The image processing module is specifically further configured to input the semantic feature vector as a query, the multispectral feature map as keys and values into a cross-attention network to obtain a fusion feature corresponding to the multispectral image; Specifically, the image processing module is further configured to input the fusion feature into a second mask decoder to obtain a non-attended region image mask, and multiply the non-attended region image mask by the spectral images in the first multispectral image respectively to obtain a second multispectral image of the target region; Wherein, the non-attended region image mask is used to mask the non-attended regions in the first multispectral image.

9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the stain detection method based on spectral anomaly region perception according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the stain detection method based on spectral anomaly region perception according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Cleaning robot and cleaning robot control method

    CN116687273A

  • Unmanned aerial vehicle target detection method and system based on multispectral information fusion

    CN117789062A

  • Rice seed vigor nondestructive testing method based on multispectral discriminant feature fusion

    CN118097277A

  • Image detection method and device based on multispectral image fusion, medium and equipment

    CN118155036A

  • Small sample hyperspectral remote sensing image change detection method based on graph convolution

    CN118447395A