Fog concentration estimation method, readable storage medium and vehicle

By aligning and fusing features of visible light and infrared images, the problem of low fog concentration estimation accuracy was solved, achieving more accurate fog concentration estimation and improving the reliability of autonomous driving strategies.

CN121746362APending Publication Date: 2026-03-27GREAT WALL MOTOR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In existing technologies, the estimation of fog concentration through visible light image analysis suffers from low accuracy, especially in foggy weather, at night, and in backlit environments, leading to inaccurate fog concentration and affecting the reliability of autonomous driving strategies.

Method used

A feature-level alignment and fusion method is adopted. Feature maps are extracted and aligned and fused using visible light and infrared images along the vehicle's driving direction. Fog concentration is determined using a fog concentration estimation model. Infrared images are used to compensate for the blurring, uneven brightness, and distortion of visible light images, thereby improving the estimation accuracy.

Benefits of technology

By aligning and fusing features at the feature level, the amount of data and interference in the estimation process are reduced, feature differences are eliminated, the accuracy of fog concentration estimation is improved, and the reliability of autonomous driving strategies is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746362A_ABST
    Figure CN121746362A_ABST
Patent Text Reader

Abstract

The invention provides a fog concentration estimation method, a readable storage medium and a vehicle, and is applied to the technical field of vehicle control. The method comprises the following steps: acquiring a visible light feature map of a visible light image and an infrared feature map of an infrared image in the driving direction of a vehicle, aligning the visible light feature map and the infrared feature map, and fusing the aligned visible light feature map and infrared feature map to obtain an aligned feature map, and determining the fog concentration in the driving direction through a fog concentration estimation model based on the alignment feature map. According to the method, interference in the fog concentration estimation process can be reduced, the more accurate fog concentration can be obtained, and the accuracy in the fog concentration estimation process can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of vehicle control, and more particularly, to a fog concentration estimation method, a readable storage medium and a vehicle in the technical field of vehicle control. BACKGROUND

[0002] In the process of braking driving of a vehicle, it is necessary to estimate the fog concentration in the driving direction of the vehicle, and to formulate an automatic driving strategy according to the fog concentration.

[0003] A common practice is to obtain a visible light image in the driving direction of the vehicle, and to determine the fog concentration according to the analysis of the visible light image. However, this method has the problem of low accuracy. Therefore, there is an urgent need for a fog concentration estimation method to obtain accurate fog concentration during the driving of the vehicle. SUMMARY

[0004] The present application provides a fog concentration estimation method, a readable storage medium and a vehicle, which can improve the accuracy in the process of fog concentration estimation.

[0005] In a first aspect, a fog concentration estimation method is provided, the method comprising: obtaining a visible light feature map of a visible light image and an infrared feature map of an infrared image, the visible light image and the infrared image being images in the driving direction of the vehicle taken at the same time; aligning the visible light feature map and the infrared feature map to obtain the aligned visible light feature map and the infrared feature map; fusing the aligned visible light feature map and the infrared feature map to obtain an aligned feature map; determining the fog concentration in the driving direction based on the aligned feature map through a fog concentration estimation model.

[0006] In the embodiments of the present application, by using the alignment and fusion idea at the feature level, after the visible light feature map and the infrared feature map are aligned and fused, the fog concentration is estimated based on the aligned feature map obtained by fusion. The visible light feature map can filter out non-key information in the visible light image while including key information in the visible light image, and the infrared feature map can filter out non-key information in the infrared image while including key information in the infrared image. Therefore, when the fog concentration is estimated based on the aligned feature map, not only the amount of data in the estimation process can be reduced, but also the interference of non-key information on the estimation process can be avoided, thereby the accuracy of the estimation process can be improved. At the same time, after the infrared feature map and the visible light feature map are aligned, the feature difference and the spatial mismatch between the visible light image and the infrared image can be eliminated, and the key information in the visible light image and the key information in the infrared image can be effectively superimposed, so that the infrared image can make up for the problems of blur, uneven brightness and distortion in the visible light image. In simple terms, the feature-level alignment and fusion of the visible light feature map and the infrared feature map can not only provide more key information for the estimation of the fog concentration, but also reduce the interference in the estimation process, so that a more accurate fog concentration can be obtained, and the accuracy in the estimation process of the fog concentration can be improved.

[0007] Optionally, the obtaining the visible light feature map of the visible light image and the infrared feature map of the infrared image comprises: performing feature extraction on the visible light image to obtain a first feature map, performing feature extraction on the infrared image to obtain a second feature map; fusing the first feature map and the second feature map to obtain a fused feature map; and performing feature extraction on the fused feature map to obtain the visible light feature map and the infrared feature map.

[0008] In the embodiments of the present application, in the process of obtaining the visible light feature map and the infrared feature map, the first feature map and the second feature map are obtained based on the infrared image and the visible light image respectively, the two feature maps are fused to obtain a fused feature map, and then the fused feature map is feature-extracted to obtain the visible light feature map and the infrared feature map. In this way, while ensuring that the visible light feature map and the infrared feature map have key information in the infrared image and the visible light image, the amount of data in the visible light feature map and the infrared feature map can be reduced, thereby the efficiency of the entire prediction process can be improved.

[0009] Optionally, the fusing the first feature map and the second feature map to obtain a fused feature map comprises: generating a first spatial weight map of the first feature map based on the infrared image, in the first spatial weight map, a weight value at each spatial position is negatively correlated with a pixel value at a corresponding position in the infrared image; fusing the first feature map and the second feature map based on the first spatial weight map to obtain the fused feature map; or generating the first spatial weight map based on the visible light image, in the first spatial weight map, a weight value at each spatial position is positively correlated with a pixel value at a corresponding position in the visible light image; fusing the first feature map and the second feature map based on the first spatial weight map to obtain the fused feature map.

[0010] In the embodiments of the present application, in the process of fusing the first feature map and the second feature map, the first spatial weight map is generated based on the infrared image, and the weight of the first feature map and the second feature map in the fusion process is controlled according to the first spatial weight map, so that the weight of the second feature map can be increased in the area where the quality of the infrared image is high, and the weight of the first feature map can be increased in the area where the quality of the infrared image is low, thereby the problems existing in the visible light image can be better compensated based on the infrared image.

[0011] Similarly, in the process of fusing the first feature map and the second feature map, the first spatial weight map is generated based on the visible light image, and the weight of the first feature map and the second feature map in the fusion process is controlled according to the first spatial weight map, so that the weight of the second feature map can be increased in the blurred area, and the weight of the first feature map can be increased in the clear area, thereby the problems existing in the visible light image can be better compensated based on the infrared image.

[0012] Optionally, the fusing the first feature map and the second feature map to obtain a fused feature map comprises: determining a first weight of the first feature map and a second weight of the second feature map according to the light intensity in the driving direction, the first weight being positively correlated with the light intensity, and the second weight being negatively correlated with the light intensity; and fusing the first feature map and the second feature map based on the first weight and the second weight to obtain the fused feature map.

[0013] In the embodiments of the present application, in the process of fusing the first feature map and the second feature map, the first weight and the second weight are determined according to the light intensity in the driving direction, so that a greater weight can be configured for the visible light image when the light intensity is greater and the visible light image is clearer, so that the finally obtained fused feature image can better reflect the features of the driving direction.

[0014] Optionally, the feature extraction on the fusion feature map to obtain the infrared feature map and the visible light feature map comprises: extracting high-dimensional features in the fusion feature map to obtain an intermediate feature map; and performing feature extraction on the intermediate feature map to obtain the infrared feature map and the visible light feature map.

[0015] In the embodiments of the present application, high-dimensional feature extraction is performed on the fusion feature map obtained by fusing the first feature map and the second feature map to obtain an intermediate feature map, and the infrared feature map and the visible light feature map are extracted based on the intermediate feature map, so that the feature dimensions of the infrared feature map and the visible light feature map can be reduced.

[0016] Optionally, the fusing the aligned visible light feature map and the infrared feature map to obtain an aligned feature map comprises: generating a second spatial weight map of the visible light feature map based on the infrared image, wherein the weight values at each spatial position in the second spatial weight map are negatively correlated with the pixel values at the corresponding positions in the infrared image; performing weighted fusion on the visible light feature map and the infrared feature map based on the second spatial weight map to obtain the aligned feature map; or generating the second spatial weight map based on the visible light image, wherein the weight values at each spatial position in the second spatial weight map are positively correlated with the pixel values at the corresponding positions in the visible light image; and performing weighted fusion on the visible light feature map and the infrared feature map based on the second spatial weight map to obtain the aligned feature map.

[0017] In the embodiments of the present application, in the process of fusing the infrared light feature map and the visible light feature map, the second spatial weight map is generated based on the infrared image, and the weight of the visible light feature map and the infrared feature map in the fusion process is controlled according to the second spatial weight map, so that the weight of the infrared feature map can be increased in the area where the quality of the infrared image is high, and the weight of the visible light feature map can be increased in the area where the quality of the infrared image is low, thereby better compensating for various problems existing in the visible light image based on the infrared image.

[0018] Similarly, in the process of fusing the visible light feature map and the infrared feature map, the second spatial weight map is generated based on the visible light image, and the weight of the visible light feature map and the infrared feature map in the fusion process is controlled according to the second spatial weight map, so that the weight of the infrared feature map can be increased in the blurred area, and the weight of the visible light feature map can be increased in the clear area, thereby better compensating for various problems existing in the visible light image based on the infrared image.

[0019] Optionally, the fusion of the aligned visible light feature map and the infrared feature map comprises: determining a third weight of the visible light feature map and a fourth weight of the infrared feature map according to the light intensity in the driving direction, the third weight being positively correlated with the light intensity, and the fourth weight being negatively correlated with the light intensity; and fusing the aligned visible light feature map and the infrared feature map into the aligned feature map based on the third weight and the fourth weight.

[0020] In the embodiments of the present application, in the process of fusing the visible light feature map and the infrared feature map, the third weight and the fourth weight are determined according to the light intensity in the driving direction, so that a greater weight can be configured for the visible light image when the light intensity is greater and the visible light image is clearer, so that the finally obtained fused feature image can better reflect the features of the driving direction.

[0021] Optionally, the aligning of the visible light feature map and the infrared feature map to obtain the aligned visible light feature map and the infrared feature map comprises: determining an offset field of an offset feature map relative to a reference feature map based on the visible light feature map and the infrared feature map by using an offset field prediction network, the offset feature map being one of the visible light feature map and the infrared feature map, and the reference feature map being the other one of the visible light feature map and the infrared feature map; and aligning the offset feature map with the reference feature map according to the offset field.

[0022] In the embodiments of the present application, the offset field between the visible light feature map and the infrared feature map is predicted by using the offset field prediction network, and the visible light feature map and the infrared feature map are aligned based on the offset field, so that the fast alignment between the visible light feature map and the infrared feature map can be realized.

[0023] In a second aspect, a device for estimating fog concentration is provided, and the device comprises: a obtaining module configured to obtain a visible light feature map of a visible light image and an infrared feature map of an infrared image, the visible light image and the infrared image being images of a driving direction of a vehicle taken at the same time; an aligning module configured to align the visible light feature map and the infrared feature map to obtain the aligned visible light feature map and the infrared feature map; a fusing module configured to fuse the aligned visible light feature map and the infrared feature map to obtain an aligned feature map; a determining module configured to determine a fog concentration of the driving direction by using a fog concentration estimation model based on the aligned feature map.

[0024] In a third aspect, a vehicle is provided, and the vehicle comprises: a memory for storing executable program code; a processor for invoking and running the executable program code from the memory, so that the vehicle executes the method in any possible implementation manner of the first aspect.

[0025] In a fourth aspect, a program product is provided, which comprises executable program code, when the executable program code is run on a vehicle, so that the vehicle executes the method in any possible implementation manner of the first aspect.

[0026] In a fifth aspect, a readable storage medium is provided, which stores executable program code, when the executable program code is run on a vehicle, so that the vehicle executes the method in any possible implementation manner of the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0027] Figure 1 is a step flow chart of a fog concentration estimation method provided by an embodiment of the present application; Figure 2 is a determination process schematic diagram of a comprehensive evaluation value provided by an embodiment of the present application; Figure 3 is a schematic diagram of another fog concentration prediction process provided by an embodiment of the present application; Figure 4 is a structural schematic diagram of a fog concentration estimation device provided by an embodiment of the present application; Figure 5 is a structural schematic diagram of a vehicle provided by an embodiment of the present application. DETAILED DESCRIPTION

[0028] The technical solutions in the present application will be described clearly and exhaustively in combination with the drawings. In the description of the embodiments of the present application, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B: "and / or" in the text is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.

[0029] Hereinafter, the terms "first" and "second" are only used for description purposes, and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more features.

[0030] At present, in order to obtain the fog concentration in the driving direction of the vehicle, the visible light image in the driving direction is generally shot by the camera, and the fog concentration in the driving direction is determined through visible light image analysis. However, in foggy, night and backlight environments, the visible light image is prone to problems such as blurring, uneven brightness and distortion, and when the visible light image has the above problems, the final obtained fog concentration is inaccurate. Further, when the fog concentration is inaccurate, the reliability of the automatic driving strategy formulated according to the fog concentration is low, which is easy to cause safety accidents.

[0031] In order to solve the above problems, the embodiment of the present application provides a fog concentration estimation method, which firstly extracts the infrared feature map and the visible light feature map from the infrared image and the visible light image in the driving direction of the vehicle in the fog concentration estimation process, then aligns and fuses the infrared feature map and the visible light feature map, and finally determines the fog concentration in the driving direction based on the aligned feature map obtained by fusion by the fog concentration estimation model.

[0032] The method uses the alignment and fusion idea at the feature level, and after aligning and fusing the visible light feature map and the infrared feature map, the fog concentration is estimated based on the aligned feature map obtained by fusion. The visible light feature map can filter out the non-key information in the visible light image while including the key information in the visible light image, and the infrared feature map can filter out the non-key information in the infrared image while including the key information in the infrared image. Therefore, when the fog concentration is estimated based on the aligned feature map, not only the amount of data in the estimation process can be reduced, but also the interference of non-key information to the estimation process can be avoided, thereby the accuracy of the estimation process can be improved. At the same time, after aligning the infrared feature map and the visible light feature map, the feature difference and spatial mismatch between the visible light image and the infrared image can be eliminated, and the key information in the visible light image and the key information in the infrared image can be effectively superimposed, so that the infrared image can make up for the problems such as blurring, uneven brightness and distortion in the visible light image. In simple terms, the alignment and fusion of the visible light feature map and the infrared feature map at the feature level not only provides more key information for the estimation of the fog concentration, but also reduces the interference in the estimation process, so that a more accurate fog concentration can be obtained, and the accuracy in the fog concentration estimation process can be improved.

[0033] Referring to Figure 1 , Figure 1 is a step flowchart of a fog concentration estimation method provided by the embodiment of the present application. The execution subject of the method can be a vehicle control unit (VCU) in a vehicle, and the method can include the following steps: Step 101, obtaining a visible light feature map of a visible light image and an infrared feature map of an infrared image.

[0034] The visible light image and the infrared image are images of the driving direction of the vehicle at the same time, and the driving direction is the forward direction or the backward direction of the vehicle.

[0035] Exemplarily, the visible light camera and the infrared camera are installed on the vehicle, and the visible light camera and the infrared camera are respectively connected with the vehicle controller. During the driving of the vehicle, the vehicle controller captures the area where the road in the driving direction is located in real time through the visible light camera and the infrared camera to obtain a visible light image and an infrared image in the driving direction, and then extracts features of the visible light image and the infrared image to obtain a visible light feature map and an infrared feature map.

[0036] The sampling frequencies of the visible light camera and the infrared camera can be the same or different. During the driving of the vehicle, the vehicle controller periodically captures the area where the road in the driving direction is located through the visible light camera to obtain continuous visible light images, and periodically captures the area where the road in the driving direction is located through the infrared camera to obtain continuous infrared images. During the process of estimating the fog concentration, the vehicle controller can time-align the visible light images and the infrared images, select a visible light image with high definition from the continuous visible light images, and select an infrared image captured at the same time from the continuous infrared images, and extract features of the selected visible light image and infrared image to obtain a visible light feature map and an infrared feature map.

[0037] Optionally, the visible light feature map of the visible light image and the infrared feature map of the infrared image are obtained by: extracting features of the visible light image to obtain a first feature map, and extracting features of the infrared image to obtain a second feature map; fusing the first feature map and the second feature map to obtain a fused feature map; extracting features of the fused feature map to obtain the infrared feature map and the visible light feature map.

[0038] In an embodiment, during the process of obtaining the visible light feature map and the infrared feature map, first, features of the visible light image are extracted to obtain a first feature map, and features of the infrared image are extracted to obtain a second feature map, then the first feature map and the second feature map are fused to obtain a fused feature map, and finally the infrared feature map and the visible light feature map are extracted from the fused feature map.

[0039] Referring to Figure 2 , Figure 2 is a schematic diagram of a fog concentration prediction process provided by an embodiment of the present application. As shown in Figure 2 , a pre-trained feature extraction model, an alignment model and a fog concentration estimation model can be deployed in the vehicle controller.

[0040] The feature extraction model can be constructed based on a multi-scale fusion network (MSFNet), used to extract a first feature map from the visible light image and a second feature map from the infrared image, and fuse the first feature map and the second feature map to obtain a fused feature map. The feature extraction model includes Figure 2 The first convolution channel, the second convolution channel, the fusion module, and the first encoder are shown.

[0041] The first convolution channel corresponds to the visible light image, and the second convolution channel corresponds to the infrared image. The first convolution channel and the second convolution channel can be parallel shallow convolution channels, each including N convolution layers, each of which can be a separable convolution layer. N is an integer greater than 1, such as 3, 4, and 5. When the first convolution channel and the second convolution channel are shallow convolution channels, the parameter amount of the feature extraction model can be reduced, and the edge deployment performance of the feature extraction model can be improved.

[0042] Optionally, a channel attention mechanism can be introduced into the first convolution channel and the second convolution channel to improve the feature extraction capability of the visible light image and the infrared image. An adaptive residual fusion module can be introduced into the second convolution channel to capture the hot signal concentration area in the infrared image.

[0043] After obtaining the infrared image and the visible light image at the same time, the vehicle control unit can input the visible light image into the first convolution channel, and extract features of the visible light image through the first convolution channel to obtain a first feature map. At the same time, the vehicle control unit can input the infrared image into the second convolution channel, and extract features of the infrared image through the second convolution channel to obtain a second feature map. After obtaining the first feature map and the second feature map, the first feature map and the second feature map can be fused to obtain a fused feature map.

[0044] Optionally, the first feature map and the second feature map are fused to obtain a fused feature map, including: Generating a first spatial weight map of the first feature map based on the infrared image, in which the weight value at each spatial position is negatively correlated with the pixel value at the corresponding position in the infrared image; Based on the first spatial weight map, the first feature map and the second feature map are weighted and fused to obtain a fused feature map.

[0045] In an implementation, in the process of fusing the first feature map and the second feature map, a guided gated convolution mechanism can be adopted for fusion. First, the first spatial weight map corresponding to the first feature map can be determined based on the infrared image, and the first feature map and the second feature map are fused through the first spatial weight map to obtain the fused feature map.

[0046] As shown in FIG. 1, the fusion module includes a convolution subnetwork (also referred to as a gated generator) and a fusion operator, which is a gated fusion operator. Figure 2 When the first feature map and the second feature map are obtained, the vehicle controller can input the infrared image into the convolution subnetwork to obtain a spatial weight map (also referred to as a gating map or a heat map) output by the convolution subnetwork based on the infrared image, i.e., the first spatial weight map. The first spatial weight map is a spatial weight map for the first feature map. In the first spatial weight map, the weight value at each spatial position is negatively correlated with the pixel value at the corresponding position in the infrared image, i.e., the higher the pixel value at the corresponding position in the infrared image, the lower the weight value.

[0047] The fusion operator is as follows:

[0048] wherein, is the feature vector at the spatial position in the fused feature map, is the weight value at the spatial position in the first spatial weight map, is the feature vector at the spatial position in the first feature map, i.e. the weight in the fusion process. is the feature vector at the spatial position in the second feature map, is the weight in the fusion process, + =1. It can be understood that is the spatial weight map of the first feature map (i.e., the first spatial weight map), is the spatial weight map of the second feature map.

[0049] After obtaining the first feature map, the second feature map, and the first spatial weight map, the vehicle controller can fuse the first feature map and the second feature map by weighting based on the corresponding and to obtain the fused feature map.

[0050] It can be understood that is negatively correlated with the pixel value at the corresponding position of the infrared image, is positively correlated with the pixel value at the corresponding position of the infrared image. That is, in the feature fusion process, when the effect of the infrared image at a certain position is clear, the infrared image plays a leading role, and when the effect of the infrared image at a certain position is not clear, the visible light image plays a leading role.

[0051] Optionally, the first feature map and the second feature map are fused to obtain a fused feature map, including: generating a first spatial weight map based on the visible light image, in which the weight value at each spatial position is positively correlated with the pixel value at the corresponding position of the visible light image; weighting and fusing the first feature map and the second feature map based on the first spatial weight map to obtain the fused feature map.

[0052] In another embodiment, in the process of fusing the first feature map and the second feature map, a guided gated convolution mechanism can be used for fusion, a first spatial weight map corresponding to the first feature map can be determined based on the visible light image, and the first feature map and the second feature map are fused through the first spatial weight map to obtain the fused feature map.

[0053] is positively correlated with the pixel value at the corresponding position of the visible light image, Figure 2 Similarly, when determining the first weight map based on the visible light image, the vehicle controller can input the visible light image into the convolution subnetwork to obtain the first spatial weight map output by the convolution subnetwork based on the visible light image. The first spatial weight map is a spatial weight map for the first feature map, in which the weight value at each spatial position is positively correlated with the pixel value at the corresponding position of the visible light image, is negatively correlated with the pixel value at the corresponding position of the visible light image. That is, in the feature fusion process, when the visible light image at a certain position is relatively clear, the visible light image plays a leading role, and when the visible light image at a certain position is not clear, the infrared image plays a leading role.

[0054] The method of generating a first spatial weight map based on a visible light image and weighting and fusing a first feature map and a second feature map based on the first spatial weight map to obtain a fused feature map is similar to the method of generating a first spatial weight map based on an infrared image and weighting and fusing a first feature map and a second feature map based on the first spatial weight map to obtain a fused feature map, and details are not repeated herein.

[0055] Optionally, feature extraction is performed on the fused feature map to obtain an infrared feature map and a visible light feature map, including: extracting high-dimensional features in the fused feature map to obtain an intermediate feature map; The intermediate feature map is subjected to feature extraction to obtain the infrared feature map and the visible light feature map.

[0056] In an implementation, after the first feature map and the second feature map are fused to obtain the fused feature map, high-dimensional features in the fused feature map can be extracted to obtain the intermediate feature map, and then the infrared feature map and the visible light feature map are extracted from the intermediate feature map. In this way, the data amount in the infrared feature map and the visible light feature map can be reduced.

[0057] As shown in FIG. 1, the feature extraction model further includes a first encoder. After the first feature map and the second feature map are fused into the fused feature map by the fusion module, the fused feature map can be input into the first encoder. The first encoder can extract high-dimensional features in the fused feature map to obtain the intermediate feature map. Then, the intermediate feature map can be subjected to feature extraction to obtain the infrared feature map and the visible light feature map. Figure 2

[0058] The first encoder can convert the high-dimensional fused feature map into the low-dimensional intermediate feature map. While extracting high-level semantic features in the fused feature map, the data amount in the fused feature map can be reduced.

[0059] As shown in FIG. 1, the alignment model includes a second encoder and a third encoder. The second encoder corresponds to the visible light feature map, and the third encoder corresponds to the infrared feature map. After the intermediate feature map is obtained, the vehicle control unit first performs local window segmentation on the intermediate feature map, and then inputs the segmentation result into the second encoder and the third encoder, respectively. The second encoder can output the visible light feature map based on the intermediate feature map, and the third encoder can output the infrared feature map based on the fused feature map. Figure 2

[0060] In actual applications, the visible light feature map and the infrared feature map can also be obtained based on the fused feature map, that is, the first encoder is not arranged in the feature fusion model. After the fused feature map is obtained by the fusion module, the vehicle control unit first performs local window segmentation on the fused feature map, and then inputs the segmentation result into the second encoder and the third encoder, respectively. The second encoder can output the visible light feature map based on the fused feature map, and the third encoder can output the infrared feature map based on the fused feature map.

[0061] ​​In the embodiments of the present application, in the process of acquiring the visible light feature map and the infrared feature map, the first feature map and the second feature map are respectively acquired based on the infrared image and the visible light image, the two feature maps are fused to obtain a fused feature map, and then feature extraction is performed on the fused feature map to obtain the visible light feature map and the infrared feature map. In this way, while ensuring that the visible light feature map and the infrared feature map have key information in the infrared image and the visible light image, the data amount in the visible light feature map and the infrared feature map can be reduced, so that the efficiency of the entire prediction process can be improved.

[0062] In the embodiments of the present application, in the process of fusing the first feature map and the second feature map, the first spatial weight map is generated based on the infrared image, and the weight of the first feature map and the second feature map in the fusion process is controlled according to the first spatial weight map. In this way, the weight of the second feature map can be increased in the area where the quality of the infrared image is high, and the weight of the first feature map can be increased in the area where the quality of the infrared image is low, so that the problems existing in the visible light image can be better compensated based on the infrared image.

[0063] Similarly, in the process of fusing the first feature map and the second feature map, the first spatial weight map is generated based on the visible light image, and the weight of the first feature map and the second feature map in the fusion process is controlled according to the first spatial weight map. In this way, the weight of the second feature map can be increased in the blurred area, and the weight of the first feature map can be increased in the clear area, so that the problems existing in the visible light image can be better compensated based on the infrared image.

[0064] In the embodiments of the present application, high-dimensional feature extraction is performed on the fused feature map obtained by fusing the first feature map and the second feature map to obtain an intermediate feature map, and the infrared feature map and the visible light feature map are extracted based on the intermediate feature map. In this way, the feature dimension of the infrared feature map and the visible light feature map can be reduced.

[0065] Optionally, the first feature map and the second feature map are fused to obtain a fused feature map, including: According to the light intensity of the driving direction, a first weight of the first feature map and a second weight of the second feature map are determined, the first weight is positively correlated with the light intensity, and the second weight is negatively correlated with the light intensity; The first feature map and the second feature map are fused based on the first weight and the second weight to obtain a fused feature map.

[0066] In another embodiment, in the process of fusing the first feature map and the second feature map, the first weight of the first feature map and the second weight of the second feature map can be determined according to the light intensity of the driving direction of the vehicle.

[0067] Exemplarily, the vehicle is integrated with a light intensity sensor. During driving of the vehicle, the light intensity sensor can be used to collect light intensity in the driving direction of the vehicle in real time. Then, the first weight can be determined according to the light intensity, and the second weight can be determined according to the first weight, and the sum of the first weight and the second weight is 1. For example, a first light intensity threshold and a second light intensity threshold can be set in advance, the first light intensity threshold is less than the second light intensity threshold, and preset weight values 1, 2 and 3 from small to large are set, for example, the preset weight values 1, 2 and 3 are 0.3, 0.5 and 0.7 respectively. After detecting the light intensity by the light intensity sensor, when the light intensity is less than the first light intensity threshold, the first weight is determined to be 0.3, and the second weight can be determined to be 1-0.3=0.7; when the light intensity is greater than or equal to the first light intensity threshold and less than the second light intensity threshold, the first weight is determined to be 0.5, and the second weight can be determined to be 1-0.5=0.5; when the light intensity is greater than or equal to the second light intensity threshold, the first weight is determined to be 0.7, and the second weight can be determined to be 1-0.7=0.3.

[0068] After obtaining the first weight and the second weight, in the process of fusing the first feature map and the second feature map, the first feature map and the second feature map can be fused by using the fusion operator as described above, in the above fusion operator, the first weight is the second weight is.

[0069] In the embodiment of the application, in the process of fusing the first feature map and the second feature map, the first weight and the second weight are determined according to the light intensity in the driving direction, which can configure a larger weight for the visible light image when the light intensity is larger and the visible light image is clearer, so that the finally obtained fused feature image can better reflect the characteristics of the driving direction.

[0070] In actual application, before feature extraction of the visible light image and the infrared image, the infrared image and the visible light image can also be preprocessed, and the preprocessing can include brightness normalization and image enhancement and other processing steps, which are not limited in the embodiment.

[0071] It should be understood that the above is only an example, and the specific method of obtaining the visible light image and the infrared image, and the method of obtaining the visible light feature map and the infrared feature map from the visible light image and the infrared image respectively can include but not limited to the above examples.

[0072] Step 102, aligning the visible light feature map and the infrared feature map to obtain the aligned visible light feature map and the infrared feature map.

[0073] Step 103, fuse the aligned visible light feature map and the infrared feature map to obtain an aligned feature map.

[0074] In this embodiment, after obtaining the visible light feature map and the infrared feature map, the visible light feature map and the infrared feature map can be aligned to eliminate structural and semantic inconsistencies between the infrared image and the visible light image.

[0075] Optionally, step 102 can include: Based on the visible light feature map and the infrared feature map, determining, by an offset field prediction network, an offset field of an offset feature map relative to a reference feature map, the offset feature map being one of the visible light feature map and the infrared feature map, and the reference feature map being the other of the visible light feature map and the infrared feature map; Aligning the offset feature map with the reference feature map according to the offset field.

[0076] In an implementation, after obtaining the visible light feature map and the infrared feature map, one of the visible light feature map and the infrared feature map can be taken as an offset feature map, and the other can be taken as a reference feature map, an offset field of the offset feature map relative to the reference feature map is determined by an offset field prediction network, and then the offset feature map and the reference feature map are aligned.

[0077] As shown in FIG. 2, the alignment model can be constructed based on a deformable alignment network (Deformable Alignment Network), in addition to including the second encoder and the third encoder, the alignment model further includes an offset field prediction network (Offset Field Prediction Network), a deformable convolution layer, and a fusion module. The offset field prediction network can be formed by cascading several convolution layers. Figure 2

[0078] After obtaining the visible light feature map output by the second encoder and the infrared feature map output by the third encoder, the vehicle control unit can take the visible light feature map as a reference feature map, take the infrared feature map as an offset feature map, input the visible light feature map and the infrared feature map into the offset field prediction network at the same time, and obtain an offset field output by the offset field prediction network based on the visible light feature map and the infrared feature map. The offset field records how much distance (x direction and y direction) each feature vector on the infrared feature map needs to move to coincide with the feature vector at the corresponding position on the visible light feature map.

[0079] ​After obtaining the offset field, the offset field and the infrared feature map are input into a deformable convolution network, and the deformable convolution network performs deformable convolution on the infrared feature map based on the offset field to obtain an offset infrared feature map. The offset infrared feature map is aligned with the visible light feature map, and an aligned infrared feature map and a visible light feature map can be obtained. The offset infrared feature map and the visible light feature map are the aligned visible light feature map and infrared feature map.

[0080] After obtaining the aligned infrared feature map and the visible light feature map, the aligned infrared feature map and the visible light feature map are input into a fusion module. The fusion module fuses the aligned infrared feature map and the visible light feature map into one feature map, so as to obtain an aligned feature map. The understanding of the fusion module included in the alignment model can refer to the fusion module in the feature extraction model, and the present embodiment will not be repeated here.

[0081] In actual application, the infrared feature map can also be used as a reference feature map, and the visible light feature map can be used as an offset feature map. The offset field prediction network is used to predict the offset field for the visible light feature map. Then, the deformable convolution layer is used to offset the visible light feature map based on the offset field to obtain an offset visible light feature map. The offset visible light feature map and the infrared feature map are the aligned visible light feature map and infrared feature map.

[0082] In the embodiment of the present application, the offset field prediction network is used to predict the offset field between the visible light feature map and the infrared feature map, and the visible light feature map and the infrared feature map are aligned based on the offset field, so as to realize the rapid alignment between the visible light feature map and the infrared feature map.

[0083] In step 104, the fog concentration of the driving direction is determined by the fog concentration estimation model based on the aligned feature map.

[0084] In the embodiment, after obtaining the aligned feature map, the fog concentration of the driving direction can be estimated by the fog concentration estimation model based on the aligned feature map. For example, the fog concentration estimation model can be an estimation model constructed based on a regression network, and the input is the aligned feature map and the output is the fog concentration level or the specific fog concentration value. Figure 2 As shown in FIG. 8, the fog concentration estimation model can be pre-trained and deployed in the vehicle controller. After obtaining the aligned feature map output by the alignment model, the vehicle controller can input the aligned feature map into the fog concentration estimation model to obtain the fog concentration output by the fog concentration estimation model based on the aligned feature map.

[0085] In the fog concentration estimation process in the embodiments of this application, the visible light feature map of the visible light image and the infrared feature map of the infrared image in the driving direction of the vehicle are obtained, the visible light feature map and the infrared feature map are aligned, the aligned visible light feature map and infrared feature map are fused to obtain an aligned feature map, and the fog concentration in the driving direction is determined based on the aligned feature map through a fog concentration estimation model.

[0086] The method uses the alignment and fusion idea at the feature level, after aligning and fusing the visible light feature map and the infrared feature map, the fog concentration is estimated based on the aligned feature map obtained by fusion. The visible light feature map can filter out non-key information in the visible light image while including key information in the visible light image, and the infrared feature map can filter out non-key information in the infrared image while including key information in the infrared image. Therefore, when estimating the fog concentration based on the aligned feature map, not only the amount of data in the estimation process can be reduced, but also the interference of non-key information to the estimation process can be avoided, thereby improving the accuracy of the estimation process. At the same time, after aligning the infrared feature map and the visible light feature map, the feature difference and spatial mismatch between the visible light image and the infrared image can be eliminated, and the key information in the visible light image and the key information in the infrared image can be effectively superimposed, so that the infrared image can compensate for the problems of blur, uneven brightness and distortion in the visible light image. In simple terms, the feature-level alignment and fusion of the visible light feature map and the infrared feature map not only provides more key information for the estimation of the fog concentration, but also reduces the interference in the estimation process, so that a more accurate fog concentration can be obtained, and the accuracy in the fog concentration estimation process can be improved.

[0087] Referring to Figure 3 , Figure 3 is a schematic diagram of another fog concentration prediction process provided by the embodiments of this application. As shown in Figure 3 , a pre-trained feature extraction model, an alignment model and a fog concentration estimation model can be deployed in the vehicle controller. The feature extraction model can be constructed based on a multi-scale fusion network, including a first convolution channel and a second convolution channel, and the alignment model can be constructed based on a deformable alignment network, including an offset field prediction network, a deformable convolution layer and a fusion module. After obtaining the visible light image and the infrared image at the same time, the visible light image and the infrared image are input into the first convolution channel and the second convolution channel in the feature extraction model, respectively, to obtain a feature map output by the first convolution channel based on the visible light image (i.e. a visible light feature map), and to obtain a feature map output by the second convolution channel based on the infrared image (i.e. an infrared feature map).

[0088] Then, the visible light feature map and the infrared feature map are input into an offset field prediction network in the alignment model to obtain an offset field output by the offset field prediction network based on the visible light feature map and the infrared feature map. Next, a deformable convolution layer performs deformable convolution on the infrared feature map based on the offset field to align the infrared feature map and the visible light feature map. Finally, a fusion module can perform weighted fusion on the aligned visible light feature map and the infrared feature map to obtain an aligned feature map. The aligned feature map is input into the fog concentration estimation model to obtain a fog concentration output by the fog concentration estimation model based on the aligned feature map.

[0089] Optionally, the fusion of the aligned visible light feature map and the infrared feature map to obtain the aligned feature map comprises: generating a second spatial weight map of the visible light feature map based on the infrared image, wherein a weight value at each spatial position in the second spatial weight map is negatively correlated with a pixel value at a corresponding position in the infrared image; performing weighted fusion on the visible light feature map and the infrared feature map based on the second spatial weight map to obtain the aligned feature map.

[0090] In an implementation, in the process of fusing the aligned visible light feature map and the infrared feature map, a guided gated convolution mechanism can be used for alignment, a second spatial weight map corresponding to the visible light feature map is determined based on the infrared image, and the visible light feature map and the infrared feature map are fused through the second spatial weight map to obtain the aligned feature map.

[0091] In combination Figure 2 As shown in FIG. 6, similar to the fusion module in the feature extraction model, the fusion module in the alignment model can also include a fusion operator and a convolution subnetwork. After obtaining the aligned visible light feature map and the infrared feature map, the vehicle control unit can input the infrared image into the convolution subnetwork to obtain a spatial weight map output by the convolution subnetwork based on the infrared image, i.e., a second spatial weight map. The second spatial weight map is a spatial weight map for the visible light feature map, wherein a weight value at each spatial position in the second spatial weight map is negatively correlated with a pixel value at a corresponding position in the infrared image, i.e., the higher the pixel value at the corresponding position in the infrared image, the lower the weight value.

[0092] After obtaining the visible light feature map, the infrared feature map, and the second spatial weight map, the vehicle control unit can fuse the visible light feature map and the infrared feature map based on each weight value in the second spatial weight map using the fusion operator to obtain the aligned feature map.

[0093] Optionally, the fusion of the aligned visible light feature map and the infrared feature map to obtain the aligned feature map comprises: generate a second spatial weight map based on the visible light image, in which the weight value at each spatial position is positively correlated with the pixel value at the corresponding position in the visible light image; perform weighted fusion on the visible light feature map and the infrared feature map based on the second spatial weight map to obtain the aligned feature map.

[0094] In an implementation, in the process of fusing the aligned visible light feature map and the infrared feature map, a guided gated convolution mechanism can be used for alignment, a second spatial weight map corresponding to the visible light feature map is determined based on the visible light image, and the visible light feature map and the infrared feature map are fused through the second spatial weight map to obtain the aligned feature map.

[0095] In the process of determining the second weight map based on the visible light image, the vehicle controller can input the visible light image into a convolution subnetwork to obtain the second spatial weight map output by the convolution subnetwork based on the visible light image. The second spatial weight map is a spatial weight map for the visible light feature map, in which the weight value at each spatial position is positively correlated with the pixel value at the corresponding position in the visible light image. That is, in the feature fusion process, when the visible light image at a certain position is clear, the visible light image plays a leading role, and when the visible light image at a certain position is not clear, the infrared image plays a leading role.

[0096] For understanding of the fusion process of the infrared light feature map and the visible light feature map, reference can be made to the fusion process of the first feature map and the second feature map in the foregoing example, which will not be described herein.

[0097] In the process of fusing the infrared light feature map and the visible light feature map, a second spatial weight map is generated based on the infrared image, and the weight of the visible light feature map and the infrared feature map in the fusion process is controlled according to the second spatial weight map. The weight of the infrared feature map can be increased in the area where the quality of the infrared image is high, and the weight of the visible light feature map can be increased in the area where the quality of the infrared image is low, so that various problems existing in the visible light image can be better compensated based on the infrared image.

[0098] Similarly, in the process of fusing the visible light feature map and the infrared feature map, a second spatial weight map is generated based on the visible light image, and the weight of the visible light feature map and the infrared feature map in the fusion process is controlled according to the second spatial weight map. The weight of the infrared feature map can be increased in the blurred area, and the weight of the visible light feature map can be increased in the clear area, so that various problems existing in the visible light image can be better compensated based on the infrared image.

[0099] Optionally, fusing the aligned visible light feature map and the infrared feature map to obtain the aligned feature map comprises: The third weight of the visible light feature map and the fourth weight of the infrared feature map are determined according to the light intensity in the driving direction, the third weight is positively correlated with the light intensity, and the fourth weight is negatively correlated with the light intensity. The aligned visible light feature map and the infrared feature map are fused into an aligned feature map based on the third weight and the fourth weight.

[0100] In another implementation, in the process of fusing the visible light feature map and the infrared feature map, the third weight of the visible light feature map and the fourth weight of the infrared feature map can be determined according to the light intensity in the driving direction of the vehicle.

[0101] Exemplarily, a third light intensity threshold and a fourth light intensity threshold can be set in advance, and preset weight values 4, 5 and 6 can be set from small to large, for example, 0.2, 0.6 and 0.8 respectively. After detecting the light intensity by the light intensity sensor, when the light intensity is less than the third light intensity threshold, the third weight is determined to be 0.2, and the fourth weight can be determined to be 1-0.2=0.8; when the light intensity is greater than or equal to the third light intensity threshold and less than the fourth light intensity threshold, the third weight is determined to be 0.6, and the fourth weight can be determined to be 1-0.6=0.4; when the light intensity is greater than or equal to the fourth light intensity threshold, the third weight is determined to be 0.8, and the fourth weight can be determined to be 1-0.8=0.2.

[0102] After obtaining the third weight and the fourth weight, in the process of fusing the visible light feature map and the infrared feature map, the visible light feature map and the infrared feature map are fused according to the third weight and the fourth weight to obtain an aligned feature map.

[0103] In the embodiments of the present application, in the process of fusing the visible light feature map and the infrared feature map, the third weight and the fourth weight are determined according to the light intensity in the driving direction, and the greater the light intensity, the clearer the visible light image, so that the visible light image can be configured with a larger weight, so that the final fused feature image can better reflect the features in the driving direction.

[0104] In order to facilitate understanding of the present application, the training process of the feature extraction model, the alignment model and the fog concentration estimation model in the above examples is briefly introduced.

[0105] In the process of training the feature extraction model, the alignment model and the fog concentration estimation model, a training scheme combining piecewise pre-training and global fine-tuning can be used to ensure that the feature extraction model, the alignment model and the fog concentration estimation model have stable performance while realizing the collaborative optimization of the overall network on the target task.

[0106] The training process is divided into three stages. The first stage is the pre-training stage of the feature extraction model. The feature extraction model is trained alone to reconstruct the edge structure and texture information of the original image. The training loss is a weighted combination of the perceptual loss and the edge preservation loss. The goal is to enable the feature extraction model to perceive the structure information of the fog area.

[0107] The second stage is the pre-training stage of the alignment model. The model parameters of the feature extraction model are frozen, and only the alignment model is trained. The objective function includes the structural similarity reconstruction error and the multispectral boundary consistency indicator, which ensures that the spatial positions of the infrared image and the visible light image in the feature map achieve precise overlap in important semantic areas. In this stage, the alignment model is trained using sample pairs from the sample set. Each sample includes an infrared image and a visible light image at the same time, as well as the corresponding label. In this stage, the spatial adaptive ability of the alignment model is enhanced.

[0108] The third stage is the pre-training stage of the fog concentration estimation model. The fog concentration estimation model is trained alone to predict the fog concentration. The input is the fusion feature map, and the output is the fog concentration. Optimization is performed by minimizing the combined loss function of the mean square error and the grade classification cross-entropy.

[0109] After the independent training of the three feature extraction models, alignment models, and fog concentration estimation models is completed, the global fine-tuning stage is entered. The frozen state between models is released, and the complete feature extraction model, alignment model, and fog concentration estimation model are trained as a whole. Gradient propagation is performed using a unified end-to-end backpropagation mechanism.

[0110] In this stage, the input is a pair of time-aligned infrared images and visible light images, and the output is the fog concentration. The loss function is a weighted sum of the regression main loss, the structure reconstruction auxiliary loss, the boundary consistency loss, and the modal attention guidance loss. Through joint optimization, the optimal collaboration of parameters and information flow between the feature extraction, alignment, fusion, and evaluation stages is achieved.

[0111] In the global fine-tuning stage, the training samples can come from vehicle enterprise real scene collection data, covering various typical scenes such as cities, highways, and mountainous areas. The data is sent to the training task in batches according to weather intensity and visibility level. Multi-GPU parallel acceleration is used for long-period training to ensure the generalization ability of the models in real deployment environments. To avoid overfitting and modal bias, a modal occlusion enhancement mechanism is introduced during training. That is, the infrared image and the visible light image are randomly occluded in some training samples to guide the model to learn robust feature expression ability in the case of incomplete single-modal information.

[0112] After training is completed, the trained feature extraction model, alignment model, and fog concentration estimation model can be deployed in the vehicle controller.

[0113] It should be understood that the above is only an example, and the specific training process of the feature extraction model, the alignment model and the fog concentration estimation model can be set according to actual needs, and the present embodiment does not limit this.

[0114] The above is described in detail Figures 1 to 2 The fog concentration estimation method provided by the embodiment of the application is described in detail; the following will be combined Figure 4 and Figure 4 The device embodiment of the present application is described in detail. It should be understood that the device in the embodiment of the present application can execute the various methods of the foregoing embodiments of the present application, that is, the specific working processes of the following various products can be referred to the corresponding processes in the foregoing method embodiments.

[0115] Referring to Figure 4 , Figure 4 is a structural schematic diagram of a fog concentration estimation device provided by the embodiment of the present application. As Figure 4 indicated, the fog concentration estimation device 400 can include: The acquisition module 401 is configured to acquire a visible light feature map of a visible light image and an infrared feature map of an infrared image, the visible light image and the infrared image being images in a driving direction of a vehicle taken at the same time; The alignment module 402 is configured to align the visible light feature map and the infrared feature map to obtain the aligned visible light feature map and the infrared feature map; The fusion module 403 is configured to fuse the aligned visible light feature map and the infrared feature map to obtain an aligned feature map; The determination module 404 is configured to determine a fog concentration of the driving direction based on the aligned feature map by a fog concentration estimation model.

[0116] Optionally, the acquisition module 401 is specifically configured to perform feature extraction on the visible light image to obtain a first feature map, perform feature extraction on the infrared image to obtain a second feature map, fuse the first feature map and the second feature map to obtain a fused feature map, and perform feature extraction on the fused feature map to obtain the infrared feature map and the visible light feature map.

[0117] Optionally, the acquisition module 401 is specifically configured to generate a first spatial weight map of the first feature map based on the infrared image, in which the weight value at each spatial position is negatively correlated with the pixel value at the corresponding position in the infrared image; perform weighted fusion on the first feature map and the second feature map based on the first spatial weight map to obtain the fusion feature map; or generate the first spatial weight map based on the visible light image, in which the weight value at each spatial position is positively correlated with the pixel value at the corresponding position in the visible light image; perform weighted fusion on the first feature map and the second feature map based on the first spatial weight map to obtain the fusion feature map.

[0118] Optionally, the acquisition module 401 is specifically configured to determine a first weight of the first feature map and a second weight of the second feature map according to the light intensity of the driving direction, the first weight being positively correlated with the light intensity, and the second weight being negatively correlated with the light intensity; and perform weighted fusion on the first feature map and the second feature map based on the first weight and the second weight to obtain the fusion feature map.

[0119] Optionally, the acquisition module 401 is specifically configured to extract high-dimensional features in the fusion feature map to obtain an intermediate feature map; and perform feature extraction on the intermediate feature map to obtain the infrared feature map and the visible light feature map.

[0120] Optionally, the fusion module 403 is specifically configured to generate a second spatial weight map of the visible light feature map based on the infrared image, in which the weight value at each spatial position is negatively correlated with the pixel value at the corresponding position in the infrared image; perform weighted fusion on the visible light feature map and the infrared feature map based on the second spatial weight map to obtain the aligned feature map; or generate the second spatial weight map based on the visible light image, in which the weight value at each spatial position is positively correlated with the pixel value at the corresponding position in the visible light image; perform weighted fusion on the visible light feature map and the infrared feature map based on the second spatial weight map to obtain the aligned feature map.

[0121] Optionally, the fusion module 403 is specifically configured to determine a third weight of the visible light feature map and a fourth weight of the infrared feature map according to the light intensity of the driving direction, the third weight being positively correlated with the light intensity, and the fourth weight being negatively correlated with the light intensity; and fuse the aligned visible light feature map and the infrared feature map into the aligned feature map based on the third weight and the fourth weight.

[0122] Optionally, the alignment module 402 is specifically used to determine the offset field of the offset feature map relative to the reference feature map through an offset field prediction network based on the visible light feature map and the infrared feature map, wherein the offset feature map is one of the visible light feature map and the infrared feature map, and the reference feature map is the other of the visible light feature map and the infrared feature map; and to align the offset feature map with the reference feature map according to the offset field.

[0123] See Figure 5 , Figure 5 This is a structural schematic diagram of a vehicle provided in an embodiment of this application. For example... Figure 5 As shown, the vehicle 500 is, for example, a server, including a memory 501 and a processor 502. The memory 501 stores executable program code 5011, and the processor 502 is used to call and execute the executable program code 5011 to perform a fog concentration estimation method. Furthermore, embodiments of this application also protect a fog concentration estimation device, which may include a memory and a processor. The memory stores executable program code, and the processor is used to call and execute the executable program code to perform a fog concentration estimation method provided in embodiments of this application.

[0124] This embodiment can divide the device into functional modules according to the above method example. For example, each module can correspond to a separate functional module, or two or more functions can be integrated into one output module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0125] When each functional module is divided according to its corresponding function, the device may further include a determining module, a replacing module, and a controlling module. It should be noted that all relevant content of each step involved in the above method embodiments can be referenced to the functional description of the corresponding functional module, and will not be repeated here. It should be understood that the device provided in this embodiment is used to execute the above-described fog concentration estimation method, and therefore can achieve the same effect as the above-described implementation method.

[0126] When using an integrated unit, the device may include a determination module and a control module. Specifically, when the device is applied to a vehicle, the output module can be used to control and manage the vehicle's movements. The storage module can be used to support the vehicle in executing relevant program code, etc.

[0127] The output module can be a processor or a vehicle body setting module, which can implement or execute various exemplary logical blocks, modules and circuits shown in combination with the disclosure. The processor can also be a combination of computing functions, such as one or more microprocessor combinations, digital signal processing (DSP) and microprocessor combinations, etc., and the storage module can be a memory.

[0128] The embodiment also provides a readable storage medium, which stores executable program code, and when the executable program code is executed on the vehicle, the vehicle executes the related method steps to realize the fog concentration estimation method provided by the above embodiment.

[0129] The embodiment also provides a program product, which, when executed on the vehicle, causes the vehicle to execute the related steps to realize the fog concentration estimation method provided by the above embodiment.

[0130] The device, readable storage medium, program product or chip provided by the embodiment are used to execute the corresponding method provided above, and thus the beneficial effects achieved thereby can refer to the beneficial effects of the corresponding method provided above, which will not be described here.

[0131] Through the above description of the embodiments, those skilled in the art can understand that, for the convenience and brevity of description, only the above division of functional modules is taken as an example, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0132] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed mutual ones can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0133] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any skilled person in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for estimating fog concentration, characterized in that, The method includes: The visible light feature map of the visible light image and the infrared feature map of the infrared image are obtained, wherein the visible light image and the infrared image are images of the vehicle's driving direction captured at the same time. Align the visible light feature map and the infrared feature map to obtain the aligned visible light feature map and the infrared feature map; The aligned visible light feature map and the infrared feature map are fused together to obtain an aligned feature map; Based on the aligned feature map, the fog concentration in the driving direction is determined by a fog concentration estimation model.

2. The method as described in claim 1, characterized in that, The acquisition of the visible light feature map of the visible light image and the infrared feature map of the infrared image includes: Feature extraction is performed on the visible light image to obtain a first feature map, and feature extraction is performed on the infrared image to obtain a second feature map; The first feature map and the second feature map are fused to obtain a fused feature map; Feature extraction is performed on the fused feature map to obtain the infrared feature map and the visible light feature map.

3. The method as described in claim 2, characterized in that, The process of fusing the first feature map and the second feature map to obtain a fused feature map includes: A first spatial weight map is generated based on the infrared image to represent the first feature map. In the first spatial weight map, the weight value at each spatial location is negatively correlated with the pixel value at the corresponding location in the infrared image. Based on the first spatial weight map, the first feature map and the second feature map are weighted and fused to obtain the fused feature map. Alternatively, a first spatial weight map is generated based on the visible light image, wherein the weight value at each spatial location in the first spatial weight map is positively correlated with the pixel value at the corresponding location in the visible light image; based on the first spatial weight map, the first feature map and the second feature map are weighted and fused to obtain the fused feature map.

4. The method as described in claim 2, characterized in that, The process of fusing the first feature map and the second feature map to obtain a fused feature map includes: Based on the light intensity in the direction of travel, a first weight of the first feature map and a second weight of the second feature map are determined, wherein the first weight is positively correlated with the light intensity and the second weight is negatively correlated with the light intensity; Based on the first weight and the second weight, the first feature map and the second feature map are weighted and fused to obtain the fused feature map.

5. The method as described in claim 2, characterized in that, The step of extracting features from the fused feature map to obtain the infrared feature map and the visible light feature map includes: High-dimensional features are extracted from the fused feature map to obtain an intermediate feature map; features are extracted from the intermediate feature map to obtain the infrared feature map and the visible light feature map.

6. The method as described in claim 1, characterized in that, The fusion and alignment of the visible light feature map and the infrared feature map to obtain an aligned feature map includes: A second spatial weight map of the visible light feature map is generated based on the infrared image. In the second spatial weight map, the weight value at each spatial position is negatively correlated with the pixel value at the corresponding position in the infrared image. Based on the second spatial weight map, the visible light feature map and the infrared feature map are weighted and fused to obtain the alignment feature map. Alternatively, a second spatial weight map is generated based on the visible light image, in which the weight value at each spatial location is positively correlated with the pixel value at the corresponding location in the visible light image; based on the second spatial weight map, the visible light feature map and the infrared feature map are weighted and fused to obtain the alignment feature map.

7. The method as described in claim 1, characterized in that, The fusion and alignment of the visible light feature map and the infrared feature map to obtain an aligned feature map includes: Based on the light intensity in the direction of travel, a third weight of the visible light feature map and a fourth weight of the infrared feature map are determined. The third weight is positively correlated with the light intensity, and the fourth weight is negatively correlated with the light intensity. Based on the third weight and the fourth weight, the aligned visible light feature map and the infrared feature map are fused into the aligned feature map.

8. The method according to any one of claims 1-7, characterized in that, Aligning the visible light feature map and the infrared feature map to obtain aligned visible light feature map and infrared feature map includes: Based on the visible light feature map and the infrared feature map, the offset field of the offset feature map relative to the reference feature map is determined by an offset field prediction network. The offset feature map is one of the visible light feature map and the infrared feature map, and the reference feature map is the other of the visible light feature map and the infrared feature map. Align the offset feature map with the reference feature map based on the offset field.

9. A readable storage medium, characterized in that, The readable storage medium stores executable program code that, when run on a vehicle, causes the vehicle to perform the method as described in any one of claims 1 to 8.

10. A vehicle, characterized in that, The vehicle includes: a memory and a processor, the memory being used to store executable program code; the processor being used to call and run the executable program code from the memory, causing the vehicle to perform the method as described in any one of claims 1 to 8.