Haze visibility detection method and equipment based on depth and brightness information and medium

By introducing depth and brightness information into visibility detection, a multimodal detection model is constructed, which solves the problem of insufficient visibility detection accuracy in haze scenarios, and achieves higher detection accuracy and reliability.

CN120067992AActive Publication Date: 2025-05-30NANJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510225733.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-09-24
Filing Date
2025-02-27
Publication Date
2025-05-30
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The existing visibility detection methods are insufficient for detection in haze scenarios, and it is necessary to improve the accuracy of visibility detection in haze scenarios.

Method used

A multimodal visibility detection model is constructed based on depth and brightness information, combining visual information with environmental parameters, and optimizing model design by introducing depth information and brightness information to extract visibility-related features.

Benefits of technology

It improves the accuracy and reliability of visibility detection in haze scenarios, effectively utilizes depth and brightness information, and enhances the detection accuracy of the model at night and haze conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067992A_ABST
    Figure CN120067992A_ABST
Patent Text Reader

Abstract

A haze visibility detection method and device based on depth and brightness information and a medium are used for foggy weather visibility detection, a multi-mode visibility detection model comprehensively utilizing image data and meteorological data is provided, the model takes the image data and the meteorological data as inputs at the same time, and the visibility detection accuracy is improved through a parallel network structure. The data of the two modes are respectively sent to different channels to extract respective features, and then the features of the two modes are fused to jointly invert the visibility value. According to the method, the visual information and the environmental parameters are combined by constructing the multi-mode visibility detection model, and the depth information and the brightness information are introduced to optimize the model design, so that the accuracy and the robustness of the model are improved, and the visibility can be detected more accurately and reliably.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - field of meteorology and computer vision, involves image processing and deep learning technologies, and is a haze visibility detection method, device and medium based on depth and brightness information. Background Art

[0002] In the field of meteorology, visibility is defined as the maximum visible distance that a normal - sighted observer can see a black target during the day or identify a luminous object at night, and haze is the main factor affecting visibility. In the aviation field, visibility is further defined as the runway visual range, that is, when the aircraft is located on the runway center line, the maximum distance at which the pilot can clearly see the runway surface markings, runway edge lights or center line lights. Visibility is a key criterion for determining whether an aircraft can take off and land safely. Therefore, accurate visibility detection is of crucial significance for flight planning and operation.

[0003] Currently, most airports rely on professional visibility detection equipment in visibility measurement, such as scattering visibility meters, transmissive visibility meters, and lidar visibility meters, etc. These devices are based on optical principles and calculate visibility by measuring optical parameters. However, these instruments are usually expensive, have high detection costs, and have strict requirements for the installation environment.

[0004] With the development of computer vision and deep learning technologies, intelligent analysis methods based on video images have been increasingly widely used in visibility detection. However, existing visibility detection models based on video images have many problems. For example: 1) Only use general network models to extract image features. Although these network models have certain feature extraction capabilities and can be well applied to general object recognition or detection tasks, visibility is different from object detection, and such general network models lack targeted design in capturing feature information related to visibility; 2) Only perform well in daytime scenes. However, in nighttime scenes, due to poor lighting conditions and scarce image information, the detection accuracy of the model drops significantly. Therefore, there is still much room for improvement in existing visibility detection technologies. Summary of the Invention

[0005] The problem to be solved by the present invention is that the existing visibility detection methods have poor detection effects under haze, and it is necessary to improve the detection accuracy of visibility in haze scenes.

[0006] The technical solution of the present invention is: a haze visibility detection method based on depth and brightness information, used for visibility detection in haze situations, including the following steps:

[0007] S1: Collect image data under different visibility conditions for the scene with the visibility to be measured, and simultaneously collect meteorological data and visibility true values corresponding to the image time in parallel to construct a data set;

[0008] S2: Build a visibility detection model, including an image data source channel, a meteorological data source channel, and a feature fusion module. The image data source channel and the meteorological data source channel are in a parallel structure. At the same time, image data and meteorological data are used as inputs. Through the parallel network structure, data of the two modalities are respectively sent to different channels to extract their respective features, and then the features of the two are fused through the feature fusion module to jointly invert the visibility value;

[0009] Among them, when the image data source channel extracts features, the ViT model is used as the backbone network, and depth information is introduced into the self-attention mechanism of the ViT model to obtain the depth information-driven layer-by-layer self-attention mechanism DDL Self-Attention. The specific process of the image data source channel extracting features is as follows:

[0010] Obtain the depth map of the image. The depth map is divided into P image patches according to the processing of the ViT model. For the depth value d of each image patch i Perform exponential transformation to obtain the depth weighting factor where i is the image patch index and β is a hyperparameter that controls the weight amplification;

[0011] Construct the depth weighting factors w of all image patches depth ∈R P into a depth weighting matrix W depth ∈R P×P , and each row of the matrix is w depth :

[0012]

[0013] w depth =[w depth,1 ,w dept h,2,……,w depth,P

[0014] In the stacking of the layer-by-layer encoder of the ViT model, introduce a dynamic adjustment factor α l to simulate the change process of attention when human vision perceives visibility, so that the image patches of the distant view gradually obtain more attention in the deep layer. The dynamic adjustment factor α of the l-th layer encoder l is:

[0015]

[0016] where L is the total number of encoder layers, l represents the current encoding layer index, and γ is a hyperparameter that controls the dynamic adjustment factor;

[0017] The depth weighting matrix of the l-th layer encoder is obtained as:

[0018] W​depth,l = α l ·W depth

[0019] Adjust the self-attention calculation result based on the depth weighting matrix in each layer of the encoder. The self-attention mechanism DDL Self-Attention of each layer of the encoder is as follows:

[0020]

[0021] Among them, Q, K, and V represent the query, key, and value vectors respectively, T represents the transpose, and d k is the embedding vector dimension, is the attention score matrix, and W depth,l is the depth weighting matrix;

[0022] Based on DDL Self-Attention, the calculation process of extracting image features through the encoder of the ViT model is as follows:

[0023] Z′ l = DDL MSA(LN(Z l-1 )) + Z l-1 l = 1, 2, …… L

[0024] Z l = MLP(LN(Z′ l )) + Z′ l l = 1, 2, …… L

[0025] Among them, LN represents layer normalization, Z l-1 represents the input of the l-th layer of the encoder, DDL MSA represents the DDL Self-Attention mechanism, Z′ l represents the output of the middle layer of the encoder, MLP represents the feed-forward layer, and Z l represents the output of the l-th layer of the encoder;

[0026] After stacking multiple layers of ViT encoders, the output of the last layer of the ViT encoder is passed through global average pooling to obtain the final global feature representation:

[0027] F image = GlobalAveragePooling(Z L )

[0028] Among them, Z L represents the feature vector output by the last layer of the ViT encoder, and F image is the extracted image feature;

[0029] S3: Train, validate, and test the constructed visibility detection model based on the dataset;

[0030] S4: Model deployment, input the real-time image for visibility detection.

[0031] Furthermore, in step S2, when fusing the features extracted from the image data and meteorological data through the feature fusion module, an adaptive weighted feature fusion strategy based on brightness information is adopted, and the features of the image and meteorological factors are dynamically weighted by considering the brightness information. Specifically:

[0032] First, the brightness value of the nighttime image is calculated to determine the illumination intensity of the current image. The image brightness B is obtained by calculating the average brightness after grayscale processing of the image:

[0033]

[0034] where n represents the nth pixel in the image, Gray(n) represents the grayscale value of the pixel, N is the total number of pixels in the image, and the brightness value B reflects the overall illumination intensity of the image, with a value range of 0 - 255;

[0035] Based on the brightness information B, the weighted coefficient of the image features is generated through the Sigmoid mapping function:

[0036]

[0037] where W image is the weighted coefficient of the image features, with a value range of 0 - 1, U is the set brightness threshold hyperparameter. When the brightness value B is higher than the brightness threshold U, the weighted coefficient W image is close to 1. When the brightness value B is lower than the brightness threshold U, the weighted coefficient W image is close to 0;

[0038] According to the calculated weighted coefficient, the image features and meteorological features are weighted and fused:

[0039] F fusion = W image ·F image +(1 - W image )·F weather

[0040] where F image is the image feature, F weather is the meteorological factor feature, F fusion is the fused feature after weighting. Input F fusion into the regression head for visibility prediction.

[0041] The present invention also provides an electronic device, which includes a processor and a memory. At least one instruction or at least one program segment is stored in the memory. The at least one instruction or at least one program segment is loaded and executed by the processor to implement the visibility detection model of the haze visibility detection method based on depth and brightness information, so as to perform real-time detection of haze visibility.

[0042] The present invention also provides a computer-readable storage medium, in which at least one instruction or at least one program segment is stored. When the at least one instruction or one program segment is executed, the visibility detection model of the haze visibility detection method based on depth and brightness information described above is implemented.

[0043] The characteristics of haze images and visibility detection tasks are that they need to capture the distribution of fog rather than common object recognition and detection. Therefore, the general vision recognition models in the prior art cannot well adapt to this task. After fully considering the characteristics of haze images, the present invention proposes the design of DDL Attention.

[0044] The present invention has the following beneficial effects:

[0045] (1) A method for detecting airport haze visibility based on depth and brightness information is proposed. By constructing a multi-modal visibility detection model, visual information is combined with environmental parameters. In particular, depth information and brightness information are introduced to optimize the model design, which can detect visibility more accurately and reliably. Existing detection methods do not fully utilize these information. The present invention comprehensively applies these two kinds of information to visual detection and the fusion of vision and meteorological detection, effectively improving the detection accuracy.

[0046] (2) When observing visibility with the human eye, first, a rough estimate will be made on whether there is haze in the scene and the size of the haze. Subsequently, the attention will gradually shift from the foreground in the scene to the background, and more attention will be given to the background area, trying to find the farthest area that can be recognized in the scene to further judge the size of visibility. Therefore, the background area in the haze image can better reflect the actual visibility size. Starting from the perspective of human eye observing haze visibility, on the basis of the baseline model ViT, the present invention designs a depth information-driven layer-by-layer self-attention mechanism DDL Self-Attention to better extract visibility-related features. DDL Self-Attention effectively utilizes depth information and gradually increases the attention weight of the background area in the image layer by layer. This way simulates the "from near to far" transfer process of human attention when perceiving haze visibility, gradually enhancing the attention to the background area, enabling the model to extract more effective feature information, better extract visibility features, rather than paying too much attention to invalid information such as the foreground area.

[0047] (3) Considering that the night visibility detection is directly related to the light source, and the worse the night illuminance is, the less effective feature information the image can provide. The present invention proposes an adaptive weighted feature fusion strategy based on brightness information, which uses the image brightness information as prior knowledge to dynamically weight the image and meteorological factor features, optimizes the fusion of the image and meteorological factor features under different night lighting conditions, and improves the accuracy and robustness of the model in the night environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a flowchart of the haze visibility detection method in the embodiment of the present invention.

[0049] Figure 2 It is a schematic structural diagram of the visibility detection model in the embodiment of the present invention.

[0050] Figure 3 It is a schematic structural diagram of the DDL Attention in the embodiment of the present invention.

[0051] Figure 4 It is a schematic structural diagram of the vision transformer encoder in the embodiment of the present invention.

[0052] Figure 5 It is a schematic diagram of the feature importance analysis in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0054] As Figure 1 shown, the present invention provides an airport haze visibility detection method based on depth and brightness information, including the following steps:

[0055] Step S1: Collect historical data and construct a data set.

[0056] Collect image data under different visibility conditions at the airport, and simultaneously collect meteorological data and visibility ground truth corresponding to the image time in parallel. Among them, the image data is obtained from the surveillance video captured by the high-definition cameras deployed at the airport, and the meteorological data is obtained from the automatic weather observation (AWOS) system deployed at the airport. This system records a large number of meteorological elements including temperature, wind speed, precipitation, etc., and records once every minute. In order to keep the image data and meteorological data aligned and synchronized, one frame is extracted from the video every minute for the image, and the visibility ground truth is obtained by the visibility detector deployed at the airport. Finally, a multi-modal data set is made, and the data set is divided into a training set, a validation set and a test set according to the ratio of 7:2:1.

[0057] Step 2: Construct a visibility detection model

[0058] As Figure 2 shown, the visibility detection model constructed by the present invention consists of four parts: an image data source channel, a meteorological data source channel, feature fusion, and a regression head. The image data source channel and the meteorological data source channel are in a parallel structure, and data of two modalities, namely images and meteorological factors, are respectively input into parallel backbone networks to extract their respective features. Then, the features of the two modalities are fused, and finally, the fused features are input into the regression head to jointly invert the visibility value.

[0059] In the image data source channel, first, the input original image is sharpened, and then the image features are extracted through the backbone network.

[0060] Among them, unsharp masking is used for image sharpening. First, the original image is smoothed to obtain a blurred image, then the difference between the original image and the blurred image is calculated to obtain a detail image, and finally, the detail image is added back to the original image according to a certain ratio to enhance the edges and details in the image. Through the sharpening process, the edge information in the image can be significantly enhanced, making the contours and details of the objects in the image clearer, highlighting the foreground and background in the image as well as the boundaries between the foggy and non-foggy areas, and improving the overall visual quality of the image.

[0061] Then, the sharpened image is input into the backbone network to extract image features. The present invention uses the Vision Transformer model architecture as the backbone network of the image channel and designs a depth-driven layerwise self-attention mechanism (Depth-Driven Layerwise Self-Attention, DDL Self-Attention) to better capture the visibility features.

[0062] Specifically, when humans perceive the visibility of a certain scene, they will first make a rough estimate of whether there is haze in the scene and the size of the haze. Subsequently, the attention will gradually shift from the foreground in the scene to the background, and more attention will be given to the background area, trying to find the farthest recognizable area in the scene to further judge the size of the visibility. Therefore, the clarity of the background area in the image is the key factor determining the size of the visibility. The information in the background area can better reflect the actual farthest visible distance, while the clarity of the foreground area is less affected by the change of visibility. Therefore, in the visibility detection task, more attention should be paid to the features of the background area in the image.

[0063] The ViT model divides the input image into n fixed-size image patches, captures the dependencies between each image patch through the global self-attention mechanism in the encoder, and then generates the final image features through the stacking of multiple layers of encoders. However, in the calculation process of the global self-attention mechanism, it may pay more attention to the foreground area where the features are clearer and more obvious. However, in a certain scene with different haze concentrations, the changes in the foreground area are not obvious, resulting in the image features finally extracted by the network being unable to fully represent different haze concentrations.

[0064] Therefore, the present invention designs a depth information-driven layer-by-layer self-attention mechanism (DDL Self-Attention) to simulate the attention transfer process of humans when perceiving haze visibility, as Figure 3 shown. By introducing depth information and dynamically adjusting the depth weighting factor layer by layer, the attention to the distant view area is gradually enhanced, realizing the attention transfer that conforms to human visual perception, so as to capture richer and more effective visibility features. The calculation steps of DDL Self-Attention are as follows:

[0065] For an original image I, first perform sharpening processing on it, and then obtain the corresponding depth map by passing the sharpened image through the monocular depth estimation model Depth Anything V2:

[0066] D = Depth Anything V2(S)

[0067] where S is the sharpened image, D is the obtained depth map, and the value range of the depth map is 0-1, representing the relative depth. Close to 0 represents the near view area, and close to 1 represents the distant view area.

[0068] The depth map D is also divided into P fixed-size image patches, and the depth value of each image patch is represented by the mean value of the depth values of all pixels within the image patch:

[0069]

[0070] where i is the image patch index, M represents the total number of pixels within the image patch, and D i,m represents the depth value of the m-th pixel in the i-th image patch.

[0071] Perform an exponential transformation on the depth value d i of each image patch to obtain the depth weighting factor:

[0072]

[0073] where d i is the depth value of the i-th image patch, and β is a hyperparameter that controls the weight amplification.

[0074] There are a total of P image patches, and the depth weighting factor w of the entire image depth ∈R P is constructed as a depth weighting matrix W depth ∈R P×P , and each row of it is w depth :

[0075]

[0076] w depth =[w depth,1 ,w dept h,2,……,w depth,P

[0077] Here, the depth weighting factor is extended into matrix form to match the dimension of the original attention matrix, so that during the calculation of attention, the depth weighting information can be applied to all image patches in parallel.

[0078] In the stacking of layer-by-layer encoders, a dynamic adjustment factor α is introduced l to simulate the change process of attention when human visual perception visibility, so that the image patches in the distance gradually obtain more attention in the deeper layers of the stacked encoders. The dynamic adjustment factor of the l-th layer encoder is:[[]]

[0079]

[0080] where L is the total number of encoder layers, l represents the current encoding layer index, and γ is a hyperparameter that controls the dynamic adjustment factor.

[0081] The depth weighting matrix of the l-th layer encoder can be obtained as:[[]]

[0082] W depth,l =α l ·W depth

[0083] In each layer encoder, the original self-attention calculation result is adjusted based on the depth weighting matrix. The calculation formula of DDL Self-Attention in each layer encoder is as follows:[[]]

[0084]

[0085] where Q, K, and V represent query, key, and value vectors respectively, T represents transpose, d k is the embedding vector dimension, is the attention score matrix, and W depth,l is the depth weighting matrix.

[0086] The corresponding ViT encoder structure is as Figure 4 ​As shown in the figure, the calculation process of extracting image features by the ViT encoder is as follows:

[0087] Z′ l = DDL MSA(LN(Z l-1 )) + Z l-1 l = 1, 2, …… L

[0088] Z l = MLP(LN(Z′ l )) + Z′ l l = 1, 2, …… L

[0089] Among them, LN represents layer normalization, Z l-1 represents the input of each layer of the encoder, DDL MSA represents the DDL multi-head self-attention mechanism of the present invention, Z′ l represents the output of the middle layer of the encoder, MLP represents the feed-forward layer, and Z l represents the output of the encoder.

[0090] After stacking multiple layers of ViT encoders, the output of the last layer of the ViT encoder is passed through global average pooling to obtain the final global feature representation:

[0091] F image = GlobalAveragePooling(Z L )

[0092] Among them, Z L represents the feature vector output by the last layer of the ViT encoder.

[0093] In this example, the input image size is fixed at 224×224×3, the image patch size is 16×16, the embedding dimension is 768, then the dimension of the encoder input vector is 196×768. After being processed by multiple layers of ViT encoders, the vector dimension remains unchanged. The feature vector output by the last layer of the ViT encoder passes through the global average pooling layer to obtain a 768-dimensional feature, which is the feature representation of the image mapped to the visibility value.

[0094] In the meteorological data source channel, first, the feature importance analysis of diverse meteorological factors is carried out to extract key meteorological factors, and then the meteorological features are extracted through the backbone network.

[0095] The meteorological instruments at the airport record a variety of meteorological factors. The meteorological factors that can significantly affect visibility are called key meteorological factors. The random forest is used for feature importance analysis to extract key meteorological factors, and the calculation method is as follows:

[0096]

[0097] Among them, node represents a node in the decision tree, nodes represents the set of all nodes in the decision tree, Δd(node) represents the reduction in impurity brought by feature f at node node, and S f,r represents the importance score of feature f on the r-th decision tree, R is the number of decision trees, and S f is the average score of feature f on all trees and serves as the feature importance score. Sort the feature importance scores of each meteorological factor, and select several meteorological factors with high scores as the key meteorological factors.

[0098] The results of feature importance analysis are as Figure 5 shown. The present invention selects several meteorological factors that have the most obvious influence on the visibility value, namely: horizontal wind speed (WS2A), temperature (TEMP), wind direction (WD2A), humidity (RH), vertical wind speed (CW2A), and air pressure (PAINS).

[0099] The extraction of meteorological features uses a fully connected neural network. Before inputting the key meteorological factors into the backbone network, the z-score normalization method is used to normalize the data to eliminate the difference in dimensions between different meteorological factors:

[0100]

[0101] Among them, x j is the j-th meteorological factor feature, μ j is the mean of the sample data of this meteorological factor, σ j is the standard deviation of the sample data of this meteorological factor, and x j,norm is the value of this meteorological factor after normalization.

[0102] Subsequently, input it into the fully connected neural network to extract features. Assume that the input data contains J meteorological factors, expressed as a vector x ∈ R J , and the neurons in each layer first perform a linear transformation on the input. For the h-th layer, the linear transformation formula is:

[0103] z (h) = W (h) a (h-1) + b (h)

[0104] Among them, W (h) is the weight matrix of the h-th layer, a (h-1) is the input of the (h - 1)-th layer, b (h) is the bias of the h-th layer, and z (h) is the result of the linear transformation of the h-th layer. Then, the result of the linear transformation is converted into a non-linear output through the activation function, where act() is the activation function:

[0105] a(h) = act(z (h) )

[0106] Assume that the last hidden layer is the H-th layer, and its output a (H) is the feature representation of mapping meteorological data to visibility values:

[0107] F weather = a (H)

[0108] In the example of the present invention, the input is 6 types of meteorological factors. Therefore, the number of neurons in the input layer of the fully connected neural network is 6, and the hidden layer is set to two layers. At the same time, in order to balance the roles of the two modal data and make the feature dimensions of the two modalities consistent, the number of neurons in the first layer is set to 512, and the number of neurons in the second layer is set to 768. The 768-dimensional features output by the second hidden layer are used as the feature representation of mapping meteorological factors to visibility values.

[0109] Furthermore, the image features and meteorological features are fused to jointly invert the visibility value by integrating the two modal features.

[0110] Considering that the visibility detection at night is greatly affected by environmental illumination. When the illumination intensity at night is high, the image can provide clear feature information for effective visibility detection. However, when the illumination is weak or there is no light source, the night image often lacks sufficient contrast and details and cannot provide effective visibility information, especially in the distant view area of the image, which makes it difficult for the model to extract key features related to visibility from the image. At this time, it is more accurate and effective to detect visibility through meteorological factors. For this reason, the present invention adopts an adaptive weighted feature fusion strategy based on brightness information, dynamically weights the image and meteorological factor features by considering the brightness information, so as to optimize feature fusion under different environmental illuminations and improve the accuracy of visibility detection. The steps are as follows:

[0111] First, calculate the brightness value of the night image to judge the illumination intensity of the current image. The image brightness B can be obtained by calculating the average brightness after graying the image:

[0112]

[0113] where n represents the n-th pixel in the image, Gray(n) represents the gray value of the pixel, N is the total number of image pixels, and the brightness value B reflects the overall illumination intensity of the image, with a value range of 0 - 255.

[0114] Based on the brightness information B, generate the weighted coefficient of the image features through the Sigmoid mapping function:

[0115]

[0116] Among them, W image is the weight coefficient of image features, with a value range of 0 - 1. U is the set brightness threshold hyperparameter. When the brightness value B is higher than the brightness threshold U, the weight coefficient W image is close to 1. When the brightness value B is lower than the brightness threshold U, the weight coefficient W image is close to 0.

[0117] According to the calculated weight coefficient, the image features and meteorological features are weighted and fused:

[0118] F fusion = W image ·F image +(1 - W image )·F weather

[0119] Among them, F image is the image feature, F weather is the meteorological factor feature, and F fusion is the weighted fusion feature.

[0120] Through the feature fusion strategy based on brightness information, the brightness information is used as prior knowledge to guide feature fusion and help the model learn. This strategy can better handle the changes in environmental lighting conditions in night visibility detection and improve the visibility detection accuracy of night scenes.

[0121] Finally, the attention-weighted feature representation is input into the regression head for the final visibility prediction. In this example, the regression head consists of two fully connected layers and the final output layer. The number of neurons in the fully connected layers is 1024 and 512 respectively. Since the present invention ultimately needs to predict the visibility value, the output layer has 1 neuron.

[0122] In step S3, the data set constructed in step S1 is input into the visibility detection model constructed in step S2 for model training, validation, and testing. Among them, the mean squared error function is used as the loss function during training and validation, and its calculation method is as follows:

[0123]

[0124] Among them, y t is the true visibility value of the t-th sample, is the predicted visibility value of the t-th sample, and Y is the number of samples.

[0125] The optimizer uses the AdamW optimizer with a learning rate set to 0.0001. The gradient of the loss with respect to all model parameters is calculated through backpropagation, and then the model parameters are updated to minimize the loss until the model converges. After each training epoch, the model is applied to the validation set, and the model hyperparameters are adjusted to optimize the training effect. Finally, the model with the best performance on the validation set is saved, and the model is finally evaluated using the test set data. By continuously optimizing the model parameters, the accuracy and robustness of the model are improved to ensure that the model can accurately detect visibility under various conditions.

[0126] In step S4, the visibility detection model trained and tested in step S3 is deployed to the visibility scenario to be measured, such as an airport monitoring system. The video data of the monitoring camera is obtained in real time and input into the model for visibility detection. The system can feedback the current visibility situation in real time according to the output result of the model to assist the operation decision-making of the airport.

[0127] The present invention can be implemented based on a computer program. Based on this, the present invention also provides an electronic device, which includes a processor and a memory. At least one instruction or at least one program segment is stored in the memory. The at least one instruction or at least one program segment is loaded and executed by the processor to implement the visibility detection model of the haze visibility detection method based on depth and brightness information, for real-time detection of haze visibility. A computer-readable storage medium is also provided. At least one instruction or at least one program segment is stored in the computer-readable storage medium. When the at least one instruction or one program segment is executed, the visibility detection model of the haze visibility detection method based on depth and brightness information is implemented. The instructions or program codes for implementing the method of the present invention can be written in any combination of one or more programming languages. The instructions or program codes can be executed entirely on the data processor, partially on the processor, executed partially on the processor and partially on a remote device as an independent software package, or executed entirely on a remote device or server.

Claims

1. A haze visibility detection method based on depth and brightness information is used for visibility detection in haze conditions, and its characteristics are: The following steps are involved: S1: Collect image data under different visibility conditions for the visibility scene to be tested, and collect meteorological data corresponding to the image time and the true value of visibility in parallel to construct a data set; S2: Construct a visibility detection model, including an image data source channel, a meteorological data source channel and a feature fusion module. The image data source channel and the meteorological data source channel are parallel structures. The image data and meteorological data are used as inputs. Through the parallel network structure, the data of the two modes are sent to different channels to extract their respective features. Then, the features of the two modes are fused through the feature fusion module to jointly invert the visibility value. Among them, when extracting features from the image data source channel, the ViT model is used as the backbone network, and depth information is introduced into the self-attention mechanism of the ViT model to obtain the layer-by-layer self-attention mechanism DDL Self-Attention driven by depth information. The features extracted from the image data source channel are as follows: Get the depth map of the image. The depth map is divided into P image blocks according to the processing of the ViT model. The depth value d of each image block is i Perform exponential transformation to obtain the depth weighting factor i is the image block index, β is the hyperparameter controlling weight amplification; The depth weighting factor w of all image blocks depth ∈R P Constructed as the depth weighted matrix W depth ∈R P×P , each row of the matrix is ​​w depth : w depth =[w depth,1 ,w dept h,2,……,w depth,P ] In the stacking of layer-by-layer encoders of the ViT model, a dynamic adjustment factor α is introduced l Simulating the change of attention when human visual perception visibility, the distant image blocks gradually gain more attention in the deep layer. The dynamic adjustment factor α of the encoder in the lth layer l for: Where L is the total number of encoder layers, l represents the current encoding layer index, and γ is a hyperparameter that controls the dynamic adjustment factor; The depth weighted matrix of the l-th layer encoder is obtained as: W depth,l =a l ·W depth In each layer of encoder, the self-attention calculation result is adjusted based on the depth weighted matrix. The self-attention mechanism DDL Self-Attention of each layer of encoder is as follows: Where Q, K, V represent query, key, and value vectors respectively, T represents transpose, and d k is the embedding vector dimension, is the attention score matrix, W depth,l is the depth weighting matrix; Based on DDL Self-Attention, the calculation process of extracting image features through the encoder of the ViT model is as follows: Z ′ l =DDL MSA(LN(Z l-1 ))+Z l-1 l=1,2,……L WITH l =MLP(LN(Z ′ l ))+Z ′ l l=1,2,……L Among them, LN represents layer normalization, Z l-1 represents the input of the encoder at layer l, DDL MSA represents the DDL Self-Attention mechanism, and Z ′ l represents the output of the middle layer of the encoder, MLP represents the feedforward layer, and Z l Represents the output of the l-th layer encoder; After stacking multiple layers of ViT encoders, the output of the last layer of ViT encoder is pooled globally to obtain the final global feature representation: F image =GlobalAveragePooling(Z L ) Where Z L represents the feature vector output by the last layer of ViT encoder, F image That is the extracted image feature; S3: Train, verify and test the constructed visibility detection model based on the dataset; S4: Model deployment, input real-time images for visibility detection.

2. The haze visibility detection method based on depth and brightness information according to claim 1 is characterized in that in step S1, image data is collected through monitoring video of the visibility scene to be tested, and one frame of the image is extracted from the video every one minute to keep the image data and meteorological data aligned and synchronized, the true value of visibility is obtained by deploying a visibility detector, and the constructed data set is divided into a training set, a validation set and a test set in a ratio of 7:2:

1.

3. The method for detecting haze visibility based on depth and brightness information according to claim 1, characterized in that Image sharpening is also performed before feature extraction from the image data source channel, and the image sharpening adopts the unsharp mask method.

4. The method for detecting haze visibility based on depth and brightness information according to claim 1, characterized in that The meteorological data source channel first performs feature importance analysis on various meteorological factors to extract key meteorological factors, and then extracts meteorological features through the backbone network, specifically: For the meteorological factors recorded by the weather station, the meteorological factors that can significantly affect visibility are called key meteorological factors. Random forest is used to perform feature importance analysis. The calculation method is as follows: Among them, node represents the node in the decision tree, nodes represents the set of all nodes in the decision tree, Δd(node) represents the reduction of impurity brought by feature f on node node, and S f,r represents the importance score of feature f on the rth tree, R is the number of decision trees, S f is the average score of feature f on all trees, which is used as the feature importance score. Sort the feature importance scores of each meteorological factor and select several meteorological factors with high scores as key meteorological factors; The meteorological features are extracted using a fully connected neural network. Before the key meteorological factors are input into the backbone network, the data are standardized using the z-score standardization method to eliminate the dimensional differences between different meteorological factors: Among them, x j is the jth meteorological factor characteristic, μ j is the mean of the sample data of the meteorological factor, σ j is the standard deviation of the meteorological factor sample data, x j,norm is the standardized value of the meteorological factor. Then, the normalized key meteorological factors are input into the fully connected neural network to extract features. Assuming that the input contains J meteorological factors, it is represented as a vector x∈R J , the neurons in each layer first perform a linear transformation on the input. For the hth layer, the linear transformation formula is: z (h) =W (h) a (h-1) +b (h) Among them, W (h) is the weight matrix of the hth layer, a (h-1) is the input of the h-1th layer, b (h) is the bias of the hth layer, z (h) is the linear transformation result of the hth layer, and then the linear transformation result is converted into a nonlinear output through the activation function, where act() is the activation function: a (h) =act(z (h) ) Assume that the last hidden layer is the Hth layer, and its output is a (H) Characterization for mapping meteorological data to visibility values: F weather =a (H) 。 5. The haze visibility detection method based on depth and brightness information according to claim 1, The feature is that in step S2, when the features extracted from the image data and the meteorological data are fused through the feature fusion module, an adaptive weighted feature fusion strategy based on brightness information is adopted to dynamically weight the image and meteorological factor features by considering the brightness information, specifically: First, the illumination intensity of the current image is determined by calculating the brightness value of the night image. The image brightness B is obtained by calculating the average brightness after graying the image: Where n represents the nth pixel in the image, Gray(n) represents the gray value of the pixel, N is the total number of pixels in the image, and the brightness value B reflects the overall light intensity of the image, with a value range of 0-255; Based on the brightness information B, the weighted coefficients of the image features are generated through the Sigmoid mapping function: Among them, W image is the image feature weight coefficient, with a value range of 0-1, and U is the set brightness threshold hyperparameter. When the brightness value B is higher than the brightness threshold U, the weight coefficient W image Close to 1, when the brightness value B is lower than the brightness threshold U, the weight coefficient W image Close to 0; According to the calculated weight coefficient, the image features and meteorological features are weighted fused: F fusion =W image ·F image +(1-W image )·F weather Among them, F image is the image feature, F weather is the meteorological factor characteristic, F fusion is the weighted fusion feature, and F fusion Input regression head to make visibility prediction.

6. An electronic device, characterized in that The electronic device includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the visibility detection model of the haze visibility detection method based on depth and brightness information as described in any one of claims 1 to 5, which is used for real-time detection of haze visibility.

7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one instruction or at least one program. When the at least one instruction or the program is executed, the visibility detection model of the haze visibility detection method based on depth and brightness information as described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • A traffic haze visibility detection method based on an improved Inception V4 network

    CN109948471A

  • Foggy day image visibility estimation method based on single image depth estimation

    CN112365467A

  • Foggy day visibility detection method based on two-channel deep network

    CN112365476A

  • Haze prediction method based on feature enhancement ConvLSTM

    CN112990531A

  • Visibility prediction method and device, storage medium and computer equipment

    CN115169592A

Cited By

  • Runway visual range prediction method based on multi-modal fusion

    CN120850226A

  • Self-correcting visibility measurement method based on vision-scattering spectrum bimodal fusion

    CN121830655A

  • Haze visibility detection method and system under night low-light condition

    CN121884068A

  • A method and system for detecting haze visibility under low light conditions at night

    CN121884068B