A dust detection method, device and storage medium

The method improves dust detection accuracy and reduces resource consumption by employing image preprocessing and neural network analysis, effectively monitoring dust concentration and predicting violations in construction environments.

CN116309437BActive Publication Date: 2025-07-15CHINATOWER CO LTD HEBEI BRANCH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310251021.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-15
Publication Date
2025-07-15
Estimated Expiration
2043-03-15

AI Technical Summary

Technical Problem

The existing dust detection technology has low resource utilization, low accuracy, and strong environmental dependence, which is easy to trigger by mistake, resulting in high detection costs and poor accuracy.

Method used

Image processing and deep learning methods are used to generate synthetic images through diffusion models, image enhancement and dust density density recognition, combined with a scalable masking automatic encoder for dust detection, and a detection model is built to improve accuracy.

Benefits of technology

Reduces resource consumption, improves the accuracy and robustness of dust detection, reduces false triggering, and reduces detection costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116309437B_ABST
    Figure CN116309437B_ABST
Patent Text Reader

Abstract

The present application discloses a dust detection method, device and storage medium for dust detection. The dust detection method disclosed in the present application generates a synthetic image through a diffusion model, edits and enhances the generation result to create a clean and accurate generation result, then performs dust concentration density recognition, and then detects the dust in the construction site through the recognition result and image data. Finally, it analyzes the dust emission and predicts the illegal dust emission. The present application also provides a dust detection device and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computing technologies, and in particular, to a dust detection method, device, and storage medium. Background Art

[0002] In the prior art, most dust detection units detect the reflected light intensity in the air through photoelectric sensors. The photoelectric sensors have high requirements for the environment and low accuracy, and during the construction process, the dust concentration at each time period cannot be accurately analyzed. For the neural network method, the transformer time series prediction model based on the encoder-decoder has been widely used. However, during the training process of this model, with the increase in the model depth and data length, a large amount of memory space and computing resources are required, and the resource utilization rate is not high. The current detection technology not only has high detection costs, but is also not easy to implement, has a strong dependence on the detection environment, poor accuracy, and poor robustness. The sensor technology is also limited by various objective factors at the construction site and is extremely prone to false triggering of sensor signals in actual use, resulting in incorrect judgments by the sensors. Therefore, how to reduce the resource utilization rate and inaccurate recognition are technical problems that need to be solved. Summary of the Invention

[0003] In view of the above technical problems, embodiments of this application provide a dust detection method, device, and storage medium to reduce resource consumption and improve the accuracy of dust detection.

[0004] In a first aspect, a dust detection method provided by an embodiment of this application includes:

[0005] Collect an image of a detection area and preprocess the image to obtain a first image;

[0006] Perform dust concentration density recognition based on the first image to obtain a first recognition result;

[0007] Perform dust detection based on the first image and the first recognition result.

[0008] The preprocessing the image to obtain a first image includes:

[0009] Synthesize an image and perform enhancement.

[0010] Preferably, the synthesizing an image includes:

[0011] Given a conditional image, serialize the conditional image;

[0012] Use positional embedding and token embedding as inputs to a temporal encoder, calculate self-attention between all tokens in the same spatial index in time, and obtain dimension-reduced tokens htmp ;

[0013] Reshape the dimension-reduced token h tmp into a spatial token h spt , as the input to the spatial decoder;

[0014] The spatial decoder calculates the self-attention of the spatial dimensions between all tokens at the same time index, samples each spatial dimension, and finally obtains the spatial token h′ spt ;

[0015] Perform upsampling and 3D convolution operations to obtain the adjustment signal c.

[0016] Input the adjustment signal c into the preset diffusion model to synthesize an image.

[0017] Preferably, the enhancement includes:

[0018] Align and crop all the generated video frame images into regions of interest;

[0019] Encode the frames of the cropped regions of interest into latent features;

[0020] Adjust the image features with the DDIM forward process, then generate the edited frame images with the reverse process, and paste them into the corresponding regions of the original frames;

[0021] Convert the obtained RGB image into the HSV space;

[0022] Extract the luminance component V and perform histogram equalization on it;

[0023] Among them, the histogram equalization is carried out according to the following formula:

[0024]

[0025] Among them, MN is the total number of pixels in the image, L represents the number of gray levels of the image, n k is the number of pixels with gray level rk, n j is the number of pixels with gray level r j of pixels, j is an integer, k is an integer less than L, S k =T(r k ) is the histogram equalization result of the k-th gray level.

[0026] Preferably, the obtaining the first recognition result by identifying the dust concentration density according to the first image includes:

[0027] Cut the image into multiple blocks and set thresholds for each block;

[0028] For each of the blocks, adopt a smoothing method to eliminate the details of the image;

[0029] Convert each block into the HSV color space, extract the luminance component V, and map the component V into two independent one-dimensional waves;

[0030] Search for local minima in the two independent one-dimensional waves;

[0031] Calculate the density of different dust concentration regions for the entire image.

[0032] Among them, the smoothing method for eliminating image details includes:

[0033] Perform smoothing according to the following formula:

[0034]

[0035] Among them, S p is the smoothed image, I p is the image before smoothing, p is the pixel index, D x (p) and D y (p) are the changes of pixel p in the x and y directions, L x (p) is the total change capturing the overall x direction, L y (p) is the total change capturing the overall y direction, ε is a small positive number to avoid division by zero, λ is the weight controlling smoothness, and S is all the smoothed images.

[0036] Among them, the calculation of the density of different dust concentration regions includes:

[0037] Calculate the density of different dust concentration regions according to the following formula:

[0038]

[0039] Among them, D represents the density of different dust concentration regions in lines per inch, M represents the number of pixels between different concentration regions, N is the number of regions in the image, and S dpi is the resolution.

[0040] Preferably, the dust detection based on the first image and the first recognition result includes:

[0041] Generate an indicator request matrix of size T×N c , where T represents the total number of time frames, and N c represents the total number of different contents;

[0042] Define a request matrix R w ×N c based on an N (w) window, where represents the number of time windows of length w, where w is the time interval between two consecutive update times;

[0043] The request matrix R is split by overlapping sliding windows of length L (w) ;

[0044] Given the request pattern Xu of image data as the input of the ViT CAT architecture, the obtained multi-scale features are combined with the probability of the request content and the deviation of the request pattern to mark the image data as exceeding the standard or not exceeding the standard;

[0045] Build a detection model;

[0046] Dust detection is performed based on the constructed detection model.

[0047] Preferably, the construction of the detection model includes:

[0048] The detection model consists of a patch layer and a transformer encoder;

[0049] The inspection patch layer is composed of a first visual transformer ViT network and a second ViT network;

[0050] The transformer encoder consists of a multi-head attention module and a multi-layer perceptron module;

[0051] The first ViT network and the second ViT network are connected in parallel to a multi-head attention module, and the multi-head attention module is linked to the multi-layer perceptron module;

[0052] The first ViT network is used to collect temporal correlation, and the second ViT network is used to capture the correlation between different picture contents.

[0053] Wherein, the first ViT network and the second ViT network both include:

[0054] The first ViT network and the second ViT network are both patch networks;

[0055] The first ViT network is a time-based patch network, and the second ViT network is a content- and time-based patch network.

[0056] Preferably, the dust detection method of the present invention further includes:

[0057] A scalable masked autoencoder is used to predict whether dust emission is illegal;

[0058] The scalable masked autoencoder comprises:

[0059] Tokenizer, Adaptive Token Sampler, Encoder, and Decoder.

[0060] Preferably, the tokenizer includes:

[0061] The tokenizer marks the input image v of a given size T×C×H×W as tokens X of N dimensions d;

[0062] where T represents the number of time instances, C represents the input RGB channels, H×W represents the spatial resolution, H represents the height, W represents the width, N is the number of dimensions, and d represents the dimension.

[0063] Preferably, the encoder includes:

[0064] An analysis transformer and a quantizer.

[0065] Preferably, the decoder includes:

[0066] A synthesis transformer.

[0067] Preferably, the adaptive token sampler includes:

[0068] The adaptive token sampler includes a multi-head attention (MHA) network, a linear layer, and a Softmax activation layer;

[0069] The adaptive token sampler is expressed as:

[0070] Z = MHA(X);

[0071] P = Softmax(Linear(Z));

[0072] where X is the token generated by the tokenizer, MHA represents the MHA network, Linear represents the linear transformation, and Softmax represents the Softmax activation.

[0073] In a second aspect, an embodiment of the present application further provides a dust detection device, including:

[0074] An acquisition module, configured to acquire an image of a detection area and preprocess the image to obtain a first image;

[0075] A concentration density recognition module, configured to perform dust concentration density recognition based on the first image to obtain a first recognition result;

[0076] A detection module, configured to perform dust detection based on the first image and the first recognition result.

[0077] In a third aspect, an embodiment of the present application further provides a dust detection device, including a memory, a processor, and a user interface;

[0078] The memory is used to store a computer program;

[0079] The user interface is used to interact with the user;

[0080] The processor is configured to read the computer program in the memory. When the processor executes the computer program, the dust detection method provided by the present invention is implemented.

[0081] In a fourth aspect, an embodiment of the present application further provides a processor-readable storage medium storing a computer program, and when the processor executes the computer program, the dust detection method provided by the present invention is implemented.

[0082] Using the dust detection method of the present invention, a synthetic image is generated through a diffusion model, the generation result is edited and image enhanced to create a clean and accurate generation result, then the dust concentration density is identified, and then the dust at the construction site is detected based on the identification result and image data. Finally, the dust emission is analyzed to predict dust illegal emissions. A dust sensor is deployed at the construction site to collect the content of PM2.5 or PM10 in the air. When the content reaches the threshold, relevant data will be sent to the platform. The camera bound to the sensor is scheduled by the platform to turn to the orientation where the sensor device is located for capturing images and analysis, and the corresponding images are obtained as a data set to detect whether there are any illegal behaviors. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0084] Figure 1 It is a schematic flow chart of the dust detection method provided by an embodiment of the present application;

[0085] Figure 2 It is a schematic flow chart of the synthetic image provided by an embodiment of the present application;

[0086] Figure 3 It is a schematic flow chart of the dust concentration identification provided by an embodiment of the present application;

[0087] Figure 4 It is a schematic structural diagram of the detection model provided by an embodiment of the present application;

[0088] Figure 5 It is a schematic diagram of the dust detection device provided by an embodiment of the present application;

[0089] Figure 6 It is another schematic structural diagram of the dust detection device provided by an embodiment of the present application. Detailed Implementation Modes

[0090] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Apparently, the described embodiments are only some of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0091] Some terms appearing in the text are explained below:

[0092] 1. In the embodiments of the present invention, the term "and / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0093] 2. In the embodiments of the present application, the term "plurality" refers to two or more, and other quantifiers are similar.

[0094] 3. DDPM, the Denoising Diffusion Probabilistic Model, is a generative model that can perform high-quality image synthesis.

[0095] 4. RGB, namely the three-color space of red (R), green (G), and blue (B).

[0096] 5. HSV, namely the space of hue H, saturation S, and value V.

[0097] 6. RTV, the abbreviation of Relative Total Variation.

[0098] 7. ViT, Vision Transformer.

[0099] 8. ViT CAT, the Parallel Vision Transformer with Cross-Attention.

[0100] 9. CA, the abbreviation of Cross-Attention, Cross Attention, also known as Cross-Attention.

[0101] 10. MC, Multi-Image Content.

[0102] 11. MLP, Multilayer Perceptron.

[0103] 12. GELU, Gaussian Error Linear Unit.

[0104] 13. MSA, the abbreviation of Multihead Self-Attention, Multi-Head Self-Attention.

[0105] 14. Transformer, Transformer.

[0106] 15. SA, an abbreviation for Self - Attention, self - attention processing.

[0107] 16. AE, arithmetic encoder.

[0108] 17. AD, arithmetic decoder.

[0109] 18. DDIM, denoising diffusion implicit model.

[0110] 19. TS, that is, time series.

[0111] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0112] It should be noted that the display order of the embodiments of the present application only represents the sequence of the embodiments, and does not represent the superiority or inferiority of the technical solutions provided by the embodiments.

[0113] See Figure 1 , a schematic diagram of a dust detection method provided by an embodiment of the present application. As Figure 1 shown, the method includes steps S101 to S103:

[0114] S101. Collect an image of the detection area, and pre - process the image to obtain a first image;

[0115] S102. Perform dust concentration density recognition according to the first image to obtain a first recognition result;

[0116] S103. Perform dust detection according to the first image and the first recognition result.

[0117] Preferably, the pre - processing includes synthesizing an image and enhancing it.

[0118] The dust detection method of the present invention can be used for dust-raising machine identification at the construction site, and a dust density identification and dust illegal emission prediction method are proposed. A synthetic image is generated through a diffusion model, and the synthetic image is edited and enhanced to create a cleaner and more accurate generation result, and then the dust concentration density is identified. Finally, the dust at the construction site is detected through the identification results and image data, and its dust emission is analyzed to predict illegal dust emission. For example, a dust sensor is deployed at the construction site to collect the content of PM2.5 or PM10 in the air. When the content reaches a threshold, relevant data is sent to the platform. The platform dispatches the camera bound to the sensor to turn to the position where the sensor device is located, captures the picture, and obtains relevant image information as a data set to determine whether the dust data is normal.

[0119] It should be noted that, in the present invention, a cleaner and more accurate generation result refers to pasting the dust area of the edited frame image to the corresponding area of the original frame, without any inconsistency with the scene, redundant parts, unreasonable places, and no sense of incongruity.

[0120] As an optional example, in the above S101, preprocessing the image to obtain the first image includes synthesizing the image and enhancing the image, which are described below respectively:

[0121] 1. Synthesize images.

[0122] As an alternative example, a composite image such as Figure 2 As shown, including:

[0123] S201, given a conditional image, serializing the conditional image;

[0124] S202, taking the position embedding and token embedding as the input of the temporal encoder, calculating the self-attention between all tokens in the same spatial index in time, and obtaining the reduced-dimensional token h tmp ;

[0125] S203, the dimension-reduced token h tmp Reshape into space marker h spt , as the input of the spatial decoder;

[0126] S204, the spatial decoder calculates the self-attention of the spatial dimension between all tags in the same time index, samples each spatial dimension, and finally obtains the spatial tag h′ spt ;

[0127] S205, performing upsampling and 3D convolution operations to obtain a regulated signal c;

[0128] S206: Input the adjustment signal c into the preset diffusion model synthetic image.

[0129] The above steps S201 to S206 are described below with examples. Considering the problems of insufficiently collected data, limited data samples, and high annotation costs. If the datasets used in the training process and the testing process meet the preset specific requirements, the performance of any classification model is effective. That is to say, the larger, more balanced, and more representative the dataset is, the more trustworthy the obtained results are. Therefore, before detection, image synthesis is performed to fill the problem of insufficient datasets. Here, a sequence-aware diffusion model is cited to synthesize images given a series of conditional images. It is a scalar index sequence for tensor indexing. L represents the time length. The present invention adopts a transformer-based attention module for generating an adjustment signal of the diffusion model. This attention module will guide which frames of the adjustment image X will be beneficial to generate the missing frames.

[0130] Given the guiding adjustment image X, the shape of the unattenuated token h obtained from the non-overlapping linear projection window on the image.

[0131] The attention module is factorized into a time encoder and a spatial decoder with high computational efficiency. The time encoder calculates the self-attention between all tokens in the same spatial index in time. Then, the self-attention is calculated along the time dimension. For the spatial decoder, the output of the time encoder is reshaped into spatial tokens, and the spatial decoder calculates the self-attention of the spatial dimension between all tokens in the same time index. Then, sampling is performed on each spatial dimension. Finally, the output of a block is unfolded, and upsampling and 3D convolution operations are performed to obtain the adjustment signal c.

[0132] Since the transformer can mask specific indices of tokens, it can be trained and inferred even when frames are missing in longitudinal images. However, using a zero tensor for missing frames with non-zero positional encoding works better than masking the missing frames. Then, the diffusion model is extended by using a sequence-aware adjustment signal. In addition, an autoregressive sampling scheme is also used, which can effectively capture long-range temporal correlations during the inference process.

[0133] During the training process of the sequence-aware diffusion model, the input of the diffusion model is a target image randomly selected from the never-observed exponent and an adjustment signal from the previous exponent. The attention module and the conditional diffusion model can be pre-trained separately and fine-tuned together. The attention module can be pre-trained by minimizing the loss between the target image and the adjustment signal c, and the diffusion model can be pre-trained using a zero-value or random-value tensor as the adjustment signal. Then, using the autoregressive sampling scheme, the missing images are autoregressively interpolated with the synthetic images.

[0134] Image generation is performed through the above-mentioned generative adversarial network. For example, various background images collected from construction sites are combined with simulated dust effects to automatically generate synthetic dust emission scenarios. Then, these synthetic scenarios are used as data samples for training the detection model. The simulated dust effects are placed into background images with random position, rotation, and scaling parameters to automatically generate sufficient synthetic dust images for training the object detection model.

[0135] 2. Enhancement is performed.

[0136] As an optional example, the performing of enhancement includes:

[0137] Align and crop all the generated video frame images into regions of interest;

[0138] Encode the frames of the cropped regions of interest into latent features;

[0139] Adjust the image features with the DDIM forward process, then use the reverse process to generate edited frame images and paste them into the corresponding regions of the original frames;

[0140] Convert the obtained RGB images to the HSV space;

[0141] Extract the luminance component V and perform histogram equalization on it;

[0142] Among them, the histogram equalization is performed according to the following formula:

[0143]

[0144] Among them, MN is the total number of pixels in the image, L represents the number of gray levels of the image, n k is the number of pixels with gray level r k , n j is the number of pixels with gray level r j , j is an integer, k is an integer less than L,

[0145] S k = T(r k ) is the histogram equalization result of the k-th gray level.

[0146] The above-mentioned performing of enhancement is described below with examples. To maintain temporal consistency, a diffusion video frame autoencoder is used to edit the above-generated images to create clean and accurate results. First, all the generated video frame images are aligned and cropped into regions of interest, where N is the number of images. Then, the cropped frames are processed using the diffusion video frame autoencoder Encoded as latent features (latent features refer to information that can be extracted from the original image and can be used to generate or restore the original image). To encode an image with N video frames the latent features are encoded in a high-level semantic space . The diffusion video frame autoencoder consists of two separate semantic encoders (ID encoder E id and landmark encoder E lnd ) and a conditional noise estimator θ for diffusion modeling. The features encoded by the two semantic encoders are cascaded through an MLP to finally obtain the relevant features To encode the generated graph the noise estimator θ is used to run the deterministic forward process of DDIM conditional on and run the generative inverse process of conditional DDIM in a deterministic manner to reconstruct the encoded features

[0147] To extract the representative dust area image features therein, the ID features of each frame are averaged to:

[0148]

[0149] where represents the ID feature. Correspondingly, the landmark feature of each frame is calculated as to finally obtain the dust area image feature of each frame After that, the DDIM forward process is adjusted to calculate Thereafter, the edited frame image is generated through the conditional reverse process with After that, the dust area part of the edited frame image is pasted into the corresponding area of the original frame to create a cleaner and more accurate final result.

[0150] Specifically, first convert the obtained RGB image to the HSV (hue, saturation, value) space; then extract the luminance component V and perform histogram equalization on it. The histogram equalization used can be expressed as:

[0151]

[0152] where MN is the total number of pixels in the image, L represents the number of gray levels in the image, and n k is the number of pixels with gray level r k . Enhancing the image makes the recognition more accurate.

[0153] As an optional example, in S102 of the present invention, the dust concentration density recognition is performed according to the first image to obtain a first recognition result, as Figure 3 shown, including:​

[0154] S301. Cut the image into multiple blocks and set thresholds for each block;

[0155] Due to the complex color of dust, directly performing threshold segmentation on the entire image will result in many dust regions being unrecognizable. Therefore, first divide the image into multiple blocks, and then set thresholds for each block. It should be noted that the threshold is a pre-set value, which is set according to needs. It should be noted that the number of grids is not fixed, but depends on the image size. As an optional example, each block contains at least 250 pixels (50 pixels × 50 pixels).

[0156] S302. For each of the above-mentioned blocks, use a smoothing method to eliminate the details of the image;

[0157] As a preferred example, perform smoothing according to the following formula:

[0158]

[0159] where S p is the smoothed image, I p is the image before smoothing, p is the pixel index, D x (p) and D y (p) are the changes of pixel p in the x and y directions, L x (p) is the total change capturing the overall x direction, L y (p) is the total change capturing the overall y direction, ε is a small positive number to avoid division by zero, λ is the weight controlling smoothness, and S is all the smoothed images.

[0160] The effect of removing texture from the image is introduced by a new regularizer . λ is the weight controlling smoothness, and ε is a small positive number to avoid division by zero. The uneven brightness distribution in the image is weakened or eliminated. This result can be used for color segmentation. By dividing the elements of the original image I by the smoothed image I s , a gap-enhanced image can be obtained.

[0161] S303. Convert each block into the HSV color space, extract the brightness component V, and map the component V into two independent one-dimensional waves;

[0162] By using the brightness feature information of the image, the identification of dust regions in the image can be achieved. Intensity mutations can be observed in the transition regions. Therefore, the minimum intensity values can be found in the vertical and horizontal directions by using the projection algorithm. Assume that B(i, j) is the brightness value of the pixel in the i-th row and j-th column of the image. Then, the image can be mapped into two independent one-dimensional waves by using the following equation:

[0163] B(i) = ∑ i B(i, j);

[0164] B(j) = ∑ j B(i, j);

[0165] Where B(i) is the cumulative luminance value of the pixels in the i-th row, and B(j) is the cumulative luminance value of the pixels in the j-th column of the image. For the obtained image with enhanced transition regions, it is first converted to the HSV color space, and its luminance component V is extracted. Then the component V is mapped into two independent one-dimensional waves.

[0166] S304. Search for local minima in the two independent one-dimensional waves;

[0167] If B(i) > B(i - 1) and B(i) > B(i + 1) in an ideal smooth curve, then B(i) can be regarded as a peak; if B(i) < B(i - 1) and B(i) < B(i + 1), then B(i) can be regarded as a valley. However, the intensity of the pixel points in the image does not completely follow a decreasing distribution. Therefore, a 30-pixel spacing slider is used to search for local minima in the latitude and longitude directions respectively, and the distance between two adjacent local minima is greater than 30. The searched local minima are considered as the gaps between different dust concentration regions in the image.

[0168] S305. Calculate the density of different dust concentration regions for the entire image.

[0169] As an alternative example, according to the following formula, calculate the density of different dust concentration regions:

[0170]

[0171] Where D represents the density of different dust concentration regions in lines per inch, M represents the number of pixels between different concentration regions, N is the number of regions in the image, and S dpi is the resolution.

[0172] As an alternative example, in S103 of the present invention, the dust detection according to the first image and the first recognition result includes the following steps A1 to A6:

[0173] A1: Generate an indicator request matrix of size T × Nc, where T represents the total number of time frames, and Nc represents the total number of different contents;

[0174] Sort the image data in ascending order of time frames, and all the image data is c1. Therefore, if the image data c is requested at time T l , then generate an (T × N c ) indicator request matrix, where T represents the total number of time frames, and Nc Indicates the total number of different contents.

[0175] A2: Definition based on N w ×N c The window request matrix R (w) ,in represents the number of time windows of length w, where w is the time interval between two consecutive update times;

[0176] Define a (N w ×N c ) window, denoted by R(w), where Represents the number of time windows of length w, where w is the time interval between two consecutive update times. For example, Represents the time between two updates t u and t u -1 Image data c l The total number of requests, when the image data c l is requested within time t, then τ t,l =1.

[0177] A3: Split the request matrix R by overlapping sliding windows of length L (w) ;

[0178] To generate two-dimensional (2D) input samples, the window-based request matrix R is split by overlapping sliding windows of length L (w) Therefore, preparation by Represents the modified image data, where M is the total number of input images. u ∈R L×Nc is a 2D input sample, representing the update length L at time t u The request mode for all previous image data. Finally, the term y u ∈R 1×Nc Represents the corresponding label, where B represents storage capacity. u(l) Represents image data c l In t u+1 Sometimes the emissions may exceed the target, otherwise they will be zero.

[0179] A4: Given the request pattern Xu of image data as the input of the ViT CAT architecture, the obtained multi-scale features are combined with the probability of the request content and the deviation of the request pattern to mark the image data as exceeding the standard or not exceeding the standard;

[0180] Given the request pattern Xu of image data as the input of the ViT CAT architecture, the image data at this time is marked as exceeding the standard or not exceeding the standard in combination with the following criteria:

[0181] i) Probability of the requested content: Probability of requesting the image data cl 1 ≤ l ≤ Nc), update time t u :

[0182]

[0183] ii) Deviation of the request pattern: Deviation of the request pattern of the image data c l is represented as ζ for 1 ≤ l ≤ Nc), where a negative deviation indicates an increasing request pattern of the image data c1 over time. l

[0184] A5: Construct a detection model;

[0185] As an optional example, the detection model is as Figure 4 shown:

[0186] The described detection model consists of a patch layer (i.e., Figure 4 the Patching layer in Figure 4 ) and a Transformer encoder (i.e.,

[0187] the Transformer Encoder in Figure 4 ); Figure 4 The patch layer consists of a first Vision Transformer (ViT) network (i.e.,

[0188] the time-based patch in Figure 4 ) and a second ViT network (i.e., Figure 4 the content- and time-based patch in

[0189] The described first ViT network and the second ViT network are in parallel. The first ViT network and the second ViT network are connected to the multi-head attention module, and the multi-head attention module is linked to the multi-layer perceptron module;

[0190] The first ViT network is used to collect temporal correlations, and the second ViT network is used to capture correlations between different picture contents.

[0191] Among them, both the first ViT network and the second ViT network are patch networks;

[0192] The first ViT network is a time-based patch network, and the second ViT network is a content- and time-based patch network.

[0193] The time-based patch network uses time-based patches for the TS path to capture the temporal correlation of the content, where the size of each patch is S = L×1. The time-based patches respectively focus on the patterns of each content for a time series of length L for the request, where the total number of patches is N = (L×Nc) / (L×1) = Nc, which is the total number of contents.

[0194] In the content- and time-based patch network, in the MC path, the main goal is to capture the dependencies between all Nc contents within a short time range Ts. Therefore, the size of each patch is set to S = Ts×Nc, where the number of patches is N = (L×Nc) / (Ts×Nc) = L / Ts, where Ts << L.

[0195] The patch network, MSA module, and MLP module are described separately below.

[0196] 1. Patch Network

[0197] Generally, the input to the Transformer encoder in the ViT network is a sequence of embedded patches, which consists of patch embedding and positional embedding. In this regard, the two-dimensional input sample Xu is divided into N non-overlapping blocks, denoted by The following two patching methods are applied to the TS path and the MC path:

[0198] i) Time-based patch: To capture the temporal correlation of the content, time-based patches are used for the TS path, where the size of each patch is S = L×1. More precisely, the time-based patches respectively focus on the patterns of each content for a time series of length L for the request, where the total number of patches is N = (L×Nc) / (L×1) = Nc, which is the total number of contents.

[0199] ii) Content- and time-based patch: In the MC path, the main goal is to capture the dependencies between all Nc contents within a short time range Ts. Therefore, we set the size of each patch to S = Ts×Nc, where the number of patches is N = (L×Nc) / (Ts×Nc) = L / Ts, where Ts << L.

[0200] Each patch is flattened into a vector where (1 ≤ j ≤ N), called patch embedding, and the vector is projected linearly by E ∈ RS×d Embed it into the dimension \(d\) of the model, and then add a learnable embedding token \(x_{cls}\) at the beginning of the embedded patches. Finally, to encode the order of the input sequence, add a positional embedding \(E_{pos}\in\mathbb{R}^{(N + 1)\times d}\) after the patch embedding, where the final output of the patch and positional embeddings is denoted as \(Z_0\) and is given by:

[0201]

[0202] 2. MLP Module

[0203] The Multi-Layer Perceptron (MLP) module consists of two linear layers (LL) with Gaussian Error Linear Unit (GELU) activation functions, and together with the multi-head self-attention (MSA) block forms the Transformer encoder, where the total number of layers is denoted as \(NL\).

[0204] 3. MSA Module

[0205] The main goal of the Multihead Self-Attention (MSA) module is to attend to input samples from various representation subspaces at multiple points. More precisely, the MSA module consists of \(h\) heads with different trainable weight matrices and is executed \(h\) times in parallel. Finally, the outputs of the \(h\) heads are concatenated into a matrix and multiplied by where \(d_h\) is set to \(d / h\). Thus, the output of the MSA module is:

[0206] MSA(Z)=[SA1(Z);SA2(Z);...;SA h (Z)]W MSA (4)

[0207] where,

[0208]

[0209] Softmax represents the scaled similarity transformation probability;

[0210] represents the scaled dot product of \(Q\) and \(K\) with ;

[0211] Given the vector sequence \(Z_0\) as the input to the Transformer encoder, for \((1\leq L\leq NL)\), the outputs of the MSA and MLP modules in layer \(L\) are denoted as:

[0212] Z′1 = MSA(LayerNorm(Z l-1 )) + Z l-1 (6)

[0213] ZL = MLP(LayerNorm(Z′ L )) + Z′ L (7)

[0214] Solve the degradation problem through layer normalization. Finally, the output of the Transformer represented by is given by:

[0215]

[0216] where the provided to the LL module is used for classification tasks as follows:

[0217]

[0218] A6: Perform dust detection according to the constructed detection model.

[0219] As a preferred example, the dust detection method provided by the present invention may further include: Figure 1 S104: Use a scalable masked autoencoder to predict whether dust emissions are illegal.

[0220] wherein, the scalable masked autoencoder includes:

[0221] a tokenizer, an adaptive token sampler, an encoder, and a decoder.

[0222] As an alternative example, the tokenizer includes:

[0223] The tokenizer tokenizes the input image v of size T×C×H×W into N tokens X of dimension d;

[0224] where T represents the number of time instances, C represents the input RGB channels, H×W represents the spatial resolution, H represents the height, W represents the width, N is the number of dimensions, and d represents the dimension.

[0225] That is, for an input image v of size T×C×H×W, where T represents the number of time instances, C represents the input (RGB) channels, and H×W represents the spatial resolution, it is first tokenized by the tokenizer into N tokens X of dimension d, denoted as X.

[0226] As an alternative example, the encoder includes:

[0227] an analysis transformer, a quantizer.

[0228] In the encoder of the present invention, the visible token X sampled by encoding is passed

[0229] to generate the latent representation F v byv 。Then, given a sample observation of a random variable image V, along with the generative model p(v|y), the posterior distribution p(y|v) is sought. The posterior distribution generally cannot be expressed in a closed form. Therefore, its variational density q(y|v) can be used to represent it. Then, the posterior can be parameterized as The distribution parameters of the generative model are parameterized as p θ (v|y; θ), and the minimization of the KL divergence is continuously sought. Where y is the latent representation θ is the parameter. By feeding the input image v into the encoder, a feature representation s is generated, and the transformation g is analyzed s , and using the channel concatenation of s and Resize(V) as the input, a basic latent representation y1 is generated, where a spatial bicubic interpolation filter is selected for Resize. Analyze the transformation g v , and only use v as the input to generate an enhanced latent representation y2. Next, y1 and y2 are quantized (Q) and fed into the encoder (AE), which respectively produce a basic bitstream and an enhanced bitstream. The basic latent representation y1 is designated to capture the common information between v and s, while the enhanced latent representation y2 captures the information related only to v.

[0230] As an alternative example, the decoder includes:

[0231] A synthesis transformer.

[0232] In the decoder, on the decoder side, the corresponding bitstreams are then fed into the decoder (AD) to reconstruct the basic and enhanced latent representations and Using The synthesis transformation To reconstruct the feature representation Finally, an inference result is generated to analyze the dust emission. Using and The channel concatenation of, the synthesis transformation To reconstruct the input Then it is connected with the fixed learnable representation form of the mask token through the visible token representation Fv. Next, positional information is added to these two representations instead of rearranging them back to the original order. Finally, a prediction value is obtained through a lightweight transformer decoder It is analyzed to analyze the dust emission.

[0233] As an alternative example, the adaptive token sampler includes:

[0234] The adaptive token sampler includes a multi-head attention MHA network, a linear layer, and a normalized exponential function Softmax activation layer;

[0235] The adaptive token sampler is represented as:

[0236] Z = MHA(X);

[0237] P = Softmax(Linear(Z));

[0238] Where X is the token generated by the tokenizer, MHA represents the MHA network, Linear represents the linear transformation, and Softmax represents the softmax activation.

[0239] Given the token X generated by the tokenizer, it is passed through a lightweight multi-head attention (MHA) network, followed by a linear layer and Softmax activation, to obtain the probability scores P ∈ R of all tokens N ; Then, an N-dimensional categorical distribution P(p ∼ Categorical(N, P)) is performed, and a set of visible token indices I is drawn v (The set of masked token indices is given by I m = U - I v where U = {1, 2, 3,..., N} is the set of all indices, i.e., I m is the set complement of I v ). The number of sampled visible tokens N is calculated based on a predefined masking ratio ρ ∈ (0, 1) v and is equal to N × (1 - ρ).

[0240] In an embodiment of the present invention, the scalable masked autoencoder includes a tokenizer, an adaptive token sampler, an encoder, and a decoder. The masked reconstruction loss uses the mean squared error (MSE) loss between the prediction of the masked token and the patch-normalized true RGB value to optimize the scalable masked autoencoder (parameterized by ), as follows:

[0241]

[0242] where represents the predicted token, represents the local patch-normalized true RGB value.

[0243] The adaptive sampling loss uses the sampling loss to optimize the adaptive token sampling network (parameterized by θ), which enables the scalable masked autoencoder (parameterized by ) to perform gradient updates. The formula of is driven by the REAR algorithm in RL, where the visible token sampling process is regarded as an action, the scalable masked autoencoder is regarded as an environment, and the masked reconstruction loss Considered as a reward. Following the maximization of the expected reward in the REAR algorithm, it is proposed to optimize the sampling network by maximizing the expected reconstruction error to optimize the sampling network.

[0244] When a video with high-activity information and / or low-activity information regions is given, a high reconstruction error around the foreground can be observed. Since the goal is to sample more visible tokens from the high-activity regions and fewer tokens from the background, the sampling network is optimized by maximizing the expected reconstruction error on the masked tokens. When optimizing using the above rules, the adaptive token sampling network predicts high probability scores for tokens from high-activity regions compared to tokens from the background. This adaptive token sampling method of the scalable masked autoencoder is closely consistent with non-uniform sampling in compressive sensing, where more samples are assigned to high-activity regions and fewer samples are assigned to low-activity regions. Since the sampling assigns samples based on the level of spatio-temporal information, fewer tokens are required to achieve the same reconstruction error compared to random sampling. This also enables the scalable masked autoencoder to use a higher masking rate, which further reduces the computational burden and thus accelerates the pre-training process.

[0245] The objective function for optimizing the adaptive token sampling network can be expressed as:

[0246]

[0247] where is the probability of the masked token at index i inferred from the adaptive token sampling network parameterized by θ, is the reconstruction error of the i-th masked token. is the reconstruction error caused by the scalable masked autoencoder with parameters . Additionally, the gradient update of is prevented from propagating through the scalable masked autoencoder. Additionally, the logarithm of the probability is used to avoid precision errors caused by small probability values.

[0248] Using the dust detection method of the present invention, a synthetic image is generated through a diffusion model, the generated result is edited and image enhanced to create a clean and accurate generated result, and then the dust concentration density is identified to identify the dust in different concentration regions for subsequent preparation to improve the detection accuracy. Then, the dust at the construction site is detected based on the identification result and the image data. Finally, the dust emission is further analyzed through the dust detection, and finally, whether the dust emission is illegal is predicted.

[0249] Based on the same inventive concept, an embodiment of the present invention also provides a dust detection device, as Figure 5 shown, the device includes:

[0250] The acquisition module 401 is configured to acquire an image of the detection area and preprocess the image to obtain a first image;

[0251] The concentration density recognition module 402 is configured to perform dust concentration density recognition based on the first image to obtain a first recognition result;

[0252] The detection module 403 is configured to perform dust detection based on the first image and the first recognition result.

[0253] As a preferred example, the acquisition module 401 is further configured to preprocess the image to obtain a first image, including:

[0254] Synthesize an image and perform enhancement.

[0255] The synthesized image includes:

[0256] A given conditional image, and serialize the conditional image;

[0257] Use the positional embedding and token embedding as the input of the time encoder, calculate the self-attention between all tokens in the same spatial index in time, and obtain the dimension-reduced token h tmp ;

[0258] Reshape the dimension-reduced token h tmp into a spatial token h spt , as the input of the spatial decoder;

[0259] The spatial decoder calculates the self-attention of the spatial dimension between all tokens in the same time index, samples each spatial dimension, and finally obtains h′ spt ;

[0260] Perform upsampling and 3D convolution operations to obtain an adjustment signal c;

[0261] Input the adjustment signal c into the preset diffusion model to synthesize an image.

[0262] The performing enhancement includes:

[0263] Align and crop all the generated video frame images into regions of interest;

[0264] Encode the frames of the cropped regions of interest into latent features;

[0265] Adjust the image features with the DDIM forward process, then generate an edited frame image using the reverse process, and paste it into the corresponding region of the original frame;

[0266] Convert the acquired RGB image to the HSV space;

[0267] Extract the luminance component V and perform histogram equalization on it;

[0268] Among them, the histogram equalization is carried out according to the following formula:

[0269]

[0270] Among them, MN is the total number of pixels in the image, L represents the number of gray levels of the image, n k is the number of pixels with gray level rk, n j is the number of pixels with gray level r j , j is an integer, and k is an integer less than L.

[0271] As an alternative example, the concentration density recognition module 402 is further configured to perform dust concentration density recognition based on the first image to obtain a first recognition result, including:

[0272] Cut the image into multiple blocks and set thresholds for each block;

[0273] For each of the blocks, use a smoothing method to eliminate the details of the image;

[0274] Convert each block into the HSV color space, extract the luminance component V, and map the component V into two independent one-dimensional waves;

[0275] Search for local minima in the two independent one-dimensional waves;

[0276] Calculate the density of different dust concentration regions for the entire image.

[0277] The smoothing method to eliminate the details of the image includes:

[0278] Perform smoothing processing according to the following formula:

[0279]

[0280] Among them, S p is the smoothed image, I p is the image before smoothing, p is the pixel index, Dx(p) and D y (p) are the changes of pixel p in the x and y directions, L x (p) is the total change capturing the overall x direction, L y (p) is the total change capturing the overall y direction, ε is a small positive number to avoid division by zero, λ is the weight controlling the smoothness, and S is all the smoothed images.

[0281] The calculation of the density of different dust concentration regions includes:

[0282] Calculate the density of different dust concentration regions according to the following formula:

[0283]

[0284] Among them, D represents the density of different dust concentration regions in lines per inch, M represents the number of pixels between different concentration regions, N is the number of regions in the image, and S dpi is the resolution.

[0285] As an alternative example, the detection module 403 is further configured to perform dust detection according to the first image and the first recognition result, including:

[0286] Generating an indicator request matrix of size T×N c where T represents the total number of time frames and N c represents the total number of different contents;

[0287] Defining a request matrix R based on an N w ×N c window, where (w) represents the number of time windows of length w, and w is the time interval between two consecutive update times;

[0288] Dividing the request matrix R by an overlapping sliding window of length L (w) ;

[0289] Given the request pattern Xu of the image data as the input of the ViT CAT architecture, the obtained multi-scale features, combined with the probability of the request content and the deviation of the request pattern, label the image data as exceeding the standard and not exceeding the standard;

[0290] Constructing a detection model;

[0291] Performing dust detection according to the constructed detection model.

[0292] The constructing the detection model includes:

[0293] The detection model consists of a patch layer and a transformer encoder;

[0294] The patch layer consists of a first Vision Transformer (ViT) network and a second ViT network;

[0295] The transformer encoder consists of a multi-head attention module and a multi-layer perceptron module;

[0296] The first ViT network and the second ViT network are in parallel. The first ViT network and the second ViT network are connected to the multi-head attention module, and the multi-head attention module is linked to the multi-layer perceptron module;

[0297] The first ViT network is used to collect temporal correlations, and the second ViT network is used to capture correlations between different picture contents.

[0298] Both the first ViT network and the second ViT network include:

[0299] Both the first ViT network and the second ViT network are patch networks;

[0300] The first ViT network is a time-based patch network, and the second ViT network is a content- and time-based patch network.

[0301] As an alternative example, the dust detection device provided by the present invention further includes an analysis module 404, which is configured to predict whether dust emissions violate regulations by using a scalable masked autoencoder.

[0302] The scalable masked autoencoder includes:

[0303] A tokenizer, an adaptive token sampler, an encoder, and a decoder.

[0304] The tokenizer includes:

[0305] The tokenizer tokenizes an input image v of size T×C×H×W into N tokens X of dimension d;

[0306] where T represents the number of time instances, C represents the input RGB channels, H×W represents the spatial resolution, H represents the height, W represents the width, N is the number of dimensions, and d represents the dimension.

[0307] The encoder includes:

[0308] An analysis transformer and a quantizer.

[0309] The decoder includes:

[0310] A synthesis transformer.

[0311] The adaptive token sampler includes:

[0312] The adaptive token sampler includes a multi-head attention (MHA) network, a linear layer, and a Softmax activation layer;

[0313] The adaptive token sampler is expressed as:

[0314] Z = MHA(X);

[0315] P = Softmax(Linear(Z));

[0316] Wherein, X is a token generated by a tokenizer, MHA represents an MHA network, Linear represents a linear transformation, and Softmax represents a softmax activation.

[0317] It should be noted that the acquisition module 401 provided in this embodiment can implement all the functions included in step S101 of the above method embodiment, solve the same technical problem, and achieve the same technical effect, which will not be elaborated herein.

[0318] It should be noted that the concentration density recognition module 402 provided in this embodiment can implement all the functions included in step S102 of the above method embodiment, solve the same technical problem, and achieve the same technical effect, which will not be elaborated herein.

[0319] It should be noted that the detection module 403 provided in this embodiment can implement all the functions included in step S103 of the above method embodiment, solve the same technical problem, and achieve the same technical effect, which will not be elaborated herein.

[0320] It should be noted that the analysis module 404 provided in this embodiment can implement all the functions included in step S104 of the above method embodiment, solve the same technical problem, and achieve the same technical effect, which will not be elaborated herein.

[0321] It should be noted that the device and method of the present invention belong to the same inventive concept, solve the same technical problem, and achieve the same technical effect. The device embodiment can implement all the methods of the method embodiment, and the same parts will not be elaborated herein.

[0322] Based on the same inventive concept, the embodiment of the present invention further provides a dust detection device, as Figure 6 shown. The device includes:

[0323] including a memory 502, a processor 501, and a user interface 503;

[0324] The memory 502 is used to store a computer program;

[0325] The user interface 503 is used to interact with a user;

[0326] The processor 501 is used to read the computer program in the memory 502. When the processor 501 executes the computer program, it realizes:

[0327] acquiring an image of a detection area, and preprocessing the image to obtain a first image;

[0328] performing dust concentration density recognition on the first image to obtain a first recognition result;

[0329] Perform dust detection based on the first image and the first recognition result.

[0330] Among them, in Figure 6 , the bus architecture may include any number of interconnected buses and bridges, specifically various circuits of one or more processors represented by the processor 501 and the memory represented by the memory 502 are linked together. The bus architecture can also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface. The processor 501 is responsible for managing the bus architecture and general processing, and the memory 502 can store the data used by the processor 501 when performing operations.

[0331] The processor 501 can be a CPU, ASIC, FPGA, or CPLD, and the processor 501 can also adopt a multi-core architecture.

[0332] When the processor 501 executes the computer program stored in the memory 502, it implements any one of the dust detection methods in Embodiment 1.

[0333] It should be noted that the device provided in Embodiment 3 and the method provided in Embodiment 1 belong to the same inventive concept, solve the same technical problems, and achieve the same technical effects. The device provided in Embodiment 3 can implement all the methods in Embodiment 1, and the same parts will not be described in detail.

[0334] The present application also proposes a processor-readable storage medium. Among them, the processor-readable storage medium stores a computer program, and when the processor executes the computer program, it implements any one of the dust detection methods in Embodiment 1.

[0335] It should be noted that the division of units in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation. In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0336] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) containing computer-usable program codes.

[0337] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a means for realizing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0338] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that realizes the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0339] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.

Claims

1. A dust detection method, characterized in that, Including: Collect an image of the detection area and preprocess the image to obtain a first image; Perform dust concentration density recognition based on the first image to obtain a first recognition result, including: Cut the first image into multiple blocks and set a threshold for each block; For each block, use a smoothing method to eliminate the details of the image; Convert each block into the HSV color space, extract the brightness component V, and map the component V into two independent one-dimensional waves; Search for local minima in the two independent one-dimensional waves; Calculate the density of different dust concentration regions for the entire image to obtain the first recognition result; Perform dust detection based on the first image and the first recognition result, including: Generate an indicator request matrix of size T×Nc, where T represents the total number of time frames and Nc represents the total number of different contents; Define based on N w ×N c Request matrix of the window , where represents the number of time windows with a time interval of w between two consecutive update times; Partition the request matrix by an overlapping sliding window of length d Generate a two-dimensional input sample Xu, Xu ∈ R d×Nc ; The two-dimensional input sample Xu represents the request pattern of all image data before time t when updating the overlapping sliding window of length d u ; Given image data The requested pattern is used as the input of the ViTCAT architecture, and the obtained multi-scale features, combined with the probability of the requested content and the deviation of the requested pattern, mark the image data as over-standard and non-over-standard; where 1 ≤ l ≤ Nc, and Nc represents the total number of different contents; Construct a detection model; Perform dust detection according to the constructed detection model; The preprocessing the image to obtain a first image includes: Synthesize an image and perform enhancement; The synthesized image includes: Given an adjustment image, serialize the adjustment image; Take positional embeddings and token embeddings as the inputs of the temporal encoder, compute self-attention among all tokens in the same spatial index over time, and obtain tokens with reduced dimensions ; Reduce the dimension of the token Reshape it into a spatial token as the input of the spatial decoder; The spatial decoder computes self-attention over the spatial dimensions between all tokens at the same time index, samples each spatial dimension, and finally obtains spatial tokens ; Spatial markers token Perform upsampling and 3D convolution operations to obtain the adjustment signal c; Input the adjustment signal c and a target image randomly selected from the exponent into a preset diffusion model to synthesize an image, where , is a scalar index sequence of tensor indices, is the union with , , represents the time length; Use a scalable masked autoencoder to predict whether dust emissions are illegal; The scalable masked autoencoder includes: A tokenizer, an adaptive token sampler, an encoder, and a decoder; Probability of the requested content: The probability of the requested content represents the probability of the requested image data of, representing the probability of the requested image data of, representing the update time: ; among them, , represents the total number of requests for image data and during the time between two update times when requesting image data within time t ; Deviation of the request pattern: The deviation of the request pattern represents the image data of the deviation of the request pattern, representing the image data of the deviation of the request pattern, where a negative deviation represents the image data with an increasing request pattern over time.

2. The method according to claim 1, wherein The performing enhancement includes: Align and crop all generated synthesized images into regions of interest; Encode the frame images of the cropped regions of interest into latent features; Use the forward process of the denoising diffusion implicit model DDIM to adjust the image features of the region of interest, and then use the reverse process to generate an edited frame image and paste it into the corresponding region of the synthesized image; Convert the acquired RGB image into the HSV space; Extract the brightness component V and perform histogram equalization on the brightness component V; Wherein, the histogram equalization is performed according to the following formula: ; where, MN is the total number of pixels in the image, and L represents the number of gray levels of the image, is the number of pixels with gray level , is the number of pixels with gray level , j is an integer, k is an integer less than L, is the histogram equalization result of the k-th gray level.

3. The method according to claim 2, wherein The constructing a detection model includes: The detection model consists of a patch layer and a transformer encoder; The patch layer consists of a first ViT network and a second ViT network; The transformer encoder consists of a multi-head attention module and a multi-layer perceptron module; The first ViT network and the second ViT network are in parallel. After parallelization, the first ViT network and the second ViT network are connected to the multi-head attention module, and the multi-head attention module is linked to the multi-layer perceptron module; The first ViT network is used to collect temporal correlations, and the second ViT network is used to capture the correlations between different picture contents.

4. A dust detection device, characterized in that, Including: An acquisition module configured to collect an image of the detection area and preprocess the image to obtain a first image; A concentration density recognition module configured to perform dust concentration density recognition based on the first image to obtain a first recognition result, including: Cut the first image into multiple blocks and set a threshold for each block; For each block, use a smoothing method to eliminate the details of the image; Convert each block into the HSV color space, extract the brightness component V, and map the component V into two independent one-dimensional waves; Search for local minima in the two independent one-dimensional waves; Calculate the density of different dust concentration regions for the entire image to obtain the first recognition result; A detection module, configured to perform dust detection based on the first image and the first recognition result, including: Generate an indicator request matrix of size T×Nc, where T represents the total number of time frames and Nc represents the total number of different contents; Define based on N w ×N c Request matrix of the window , where Indicates the number of time windows with a time interval of w between two consecutive update times; Partition the request matrix by an overlapping sliding window of length d Generate a two-dimensional input sample Xu, Xu ∈ R d×Nc ; The two-dimensional input sample Xu represents the request pattern of all image data before time t when updating the overlapping sliding window of length d u ; Given image data The request pattern of is used as the input of the ViTCAT architecture, and the obtained multi-scale features, combined with the probability of the request content and the deviation of the request pattern, label the image data as over-standard and non-over-standard; where 1 ≤ l ≤ Nc, and Nc represents the total number of different contents; Construct a detection model; Perform dust detection according to the constructed detection model; The preprocessing of the image to obtain the first image includes: Synthesize the image and perform enhancement; The synthesized image includes: Given an adjustment image, serialize the adjustment image; Take positional embeddings and token embeddings as the inputs of the temporal encoder, calculate self-attention among all tokens in the same spatial index over time, and obtain tokens with reduced dimensions ; Reduce the dimension of the token Reshape it into a spatial token as the input to the spatial decoder; The spatial decoder computes self-attention over the spatial dimensions between all tokens at the same time index, samples each spatial dimension, and finally obtains spatial tokens ; Spatial marking token Perform upsampling and 3D convolution operations to obtain the adjustment signal c; Input the adjustment signal c and a target image randomly selected from the exponential into a preset diffusion model to synthesize an image, where , is a scalar index sequence of tensor indices, is the union with , , represent the time length; Use a scalable masked autoencoder to predict whether dust emissions are illegal; The scalable masked autoencoder includes: A tokenizer, an adaptive token sampler, an encoder, and a decoder; The probability of the requested content: The probability of the requested content represents the probability of the requested image data of which represents the probability of the requested image data of represents the update time ; among them, , represents the time between two updates and for the total number of requests for image data during the time t when requesting image data ; The deviation of the request pattern: The deviation of the request pattern represents the image data of the deviation of the request pattern, representing the image data of the deviation of the request pattern where a negative deviation represents the image data of an increasing request pattern over time.

5. A computer device, characterized in that, Includes a memory, a processor, and a user interface; The memory is used to store computer programs; The user interface is used to interact with the user; The processor is used to read the computer program in the memory. When the processor executes the computer program, it implements the dust detection method according to any one of claims 1 to 3.

6. A processor-readable storage medium, characterized in that, The processor-readable storage medium stores a computer program. When the processor executes the computer program, it implements the dust detection method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Intelligent analysis and treatment system and method for mine dust

    CN114511991A