Vehicle control method, electronic equipment, vehicle, medium and product

Through multi-camera stitching and fusion image processing technology, we can identify bad weather and control the operation of the vehicle, which solves the problem of lidar failure in bad weather, and achieves the stable operation and safety of the vehicle in bad weather.

CN120589006APending Publication Date: 2025-09-05BYD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510773793.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

In severe weather conditions, the vehicle's lidar is susceptible to interference and leads to failure of perception functions, affecting functions such as active lane change and active follow-up, and endangering driving safety.

Method used

Surround-view images are captured through multiple vehicle surround-view cameras. After the stitching process, a pre-trained weather recognition model is used, combined with visible light and infrared camera images for fusion processing to identify bad weather and control vehicle operation.

Benefits of technology

In severe weather conditions, the vehicle's robust operating capabilities are enhanced, ensuring the safety of drivers and passengers and improving the accuracy and robustness of target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120589006A_ABST
    Figure CN120589006A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle control method, electronic equipment, a vehicle, a computer readable storage medium and a computer program product. The method comprises the steps of determining an identification result of current weather according to first driving environment image information; under the condition that the current weather is target weather, the vehicle is controlled according to second driving environment image information, and visibility corresponding to the target weather is smaller than or equal to a preset threshold value. Therefore, under the condition that the system recognizes that the weather is severe through the images of the surrounding environment of the vehicle, the system is automatically switched to the severe weather mode, the feature images in the severe weather are fused, a fusion image which is suitable, accurate and easy to recognize is generated, target detection is carried out on the fused image, and the vehicle is controlled to actively change lanes, actively follow the vehicle and the like according to the target detection result. Therefore, stable operation of the vehicle under the target weather is guaranteed, and the safety of the vehicle and drivers and passengers in the vehicle is guaranteed to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of vehicle technology, and in particular to a vehicle control method, electronic equipment, a vehicle, a computer-readable storage medium, and a computer program product. Background Art

[0002] In related technologies, vehicles often use lidar to sense the outside world and perform corresponding functions such as active lane change and active vehicle following. However, when vehicles encounter severe weather such as heavy rain or snow, lidar is easily affected by weather interference and may fail. This prevents the vehicle from correctly perceiving the outside world, causing functions such as active lane change and active vehicle following to malfunction or even fail, thus affecting the user's driving experience and even the safety of the user and the vehicle. Summary of the Invention

[0003] The present application provides a vehicle control method, an electronic device, a vehicle, a computer-readable storage medium, and a computer program product.

[0004] An embodiment of the present application provides a vehicle control method, the method comprising:

[0005] determining a recognition result of the current weather according to the first driving environment image information;

[0006] When the current weather is target weather, the vehicle is controlled according to the second driving environment image information, wherein the visibility corresponding to the target weather is less than or equal to a preset threshold.

[0007] Thus, in the embodiments of the present application, when the current weather recognition result is determined based on the first driving environment image information and the current weather is determined to be the target weather, the vehicle can be controlled based on the second driving environment image information, such as to control the vehicle to actively change lanes or actively follow other vehicles, thereby ensuring the vehicle's stable operation in the target weather and, to a certain extent, protecting the safety of the vehicle and its passengers. Furthermore, when equipment such as lidar fails due to weather interference, the vehicle can determine the current weather recognition result based on the first driving environment image information and control it based on the second driving environment image information, further ensuring the vehicle's stable operation in the target weather.

[0008] In some embodiments, the first driving environment image information includes a plurality of surround view images captured by a plurality of vehicle surround view cameras having different shooting directions.

[0009] In this way, the first driving environment image information includes multiple surround-view images captured by multiple vehicle surround-view cameras with different shooting directions. This multi-directional surround-view image provides a wider recognition range for inclement weather, avoiding misidentification in a single direction, thereby improving the accuracy of weather judgment.

[0010] In some embodiments, determining the recognition result of the current weather according to the first driving environment image information includes:

[0011] performing stitching processing on the plurality of surround view images to determine a stitched image;

[0012] The recognition result is determined according to the spliced ​​image and the plurality of surround-view images.

[0013] In this way, multiple surround view images are stitched together to create a stitched image. Recognition results are determined based on this stitched image and the multiple surround view images. This stitching of surround view images along the path integrates multi-dimensional environmental information, providing a more comprehensive input for subsequent weather recognition. Weather recognition based on this stitched image and the multiple surround view images can determine whether the current weather is severe.

[0014] In some embodiments, determining the recognition result based on the stitched image and the plurality of surround view images includes:

[0015] The recognition result is determined based on the spliced ​​image, the plurality of surround view images, and a pre-trained weather recognition model.

[0016] In this way, the recognition result is determined based on the stitched image and multiple surround view images, as well as the pre-trained weather recognition model. By analyzing the stitched image and multiple surround view images using the pre-trained weather recognition model, feature extraction, feature fusion, and weather recognition can be quickly performed, resulting in accurate weather recognition results.

[0017] In some embodiments, the weather recognition model includes a first feature extraction module, a first feature splicing module, and a weather recognition module connected in sequence;

[0018] The first feature extraction module is configured to perform feature extraction processing on the stitched image and the surround view image respectively to determine a first feature map of the stitched image and a second feature map of the surround view image;

[0019] The first feature splicing module is configured to splice the first feature map and a plurality of the second feature maps to determine a first spliced ​​feature map;

[0020] The weather recognition module is configured to determine the recognition result according to the first splicing feature map.

[0021] In this way, the weather recognition model includes a first feature extraction module, a first feature stitching module, and a weather recognition module, which are connected in sequence. The first feature extraction module is configured to perform feature extraction processing on the stitched image and the surround view image, respectively, to determine a first feature map for the stitched image and a second feature map for the surround view image; the first feature stitching module is configured to stitch the first feature map and multiple second feature maps to determine a first stitched feature map; and the weather recognition module is configured to determine a recognition result based on the first stitched feature map. In this way, the weather recognition module can perform feature extraction, feature stitching, and weather recognition processing on the stitched image and the surround view image, thereby accurately determining whether the vehicle is experiencing inclement weather.

[0022] In some embodiments, the first feature extraction module includes a first convolution unit and a first pooling unit;

[0023] The first convolution unit is configured to perform feature extraction processing on the stitched image and the surround view image respectively to determine a third feature map of the stitched image and a fourth feature map of the surround view image;

[0024] The first pooling unit is configured to perform pooling processing on the third feature map of the stitched image and the fourth feature map of the surround view image, respectively, to determine the first feature map of the stitched image and the second feature map of the surround view image.

[0025] In this way, the first feature extraction module includes a first convolution unit and a first pooling unit; the first convolution unit is configured to perform feature extraction processing on the stitched image and the surround view image, respectively, to determine a third feature map for the stitched image and a fourth feature map for the surround view image; the first pooling unit is configured to perform pooling processing on the third feature map of the stitched image and the fourth feature map of the surround view image, respectively, to determine a first feature map for the stitched image and a second feature map for the surround view image. In this way, by extracting features from the stitched image, it is possible to integrate all-round environmental information, extract global weather features, and avoid local interference from a single direction. By extracting features from the surround view image, it is possible to capture local detailed features in all directions. The first feature extraction module can provide global and local data basis for subsequent weather classification, improving environmental perception capabilities in severe weather conditions, thereby more accurately judging severe weather conditions.

[0026] In some embodiments, the first feature splicing module includes a splicing unit and a second convolution unit;

[0027] The splicing unit is configured to splice the first feature map and a plurality of the second feature maps to determine a second spliced ​​feature map;

[0028] The second convolution unit is configured to perform convolution processing on the second splicing feature map to determine the first splicing feature map.

[0029] In this way, the first feature splicing module includes a splicing unit and a second convolution unit. The splicing unit is configured to splice the first feature map and multiple second feature maps to determine a second spliced ​​feature map. The second convolution unit is configured to convolve the second spliced ​​feature map to determine the first spliced ​​feature map. Thus, after splicing by the splicing unit, the second spliced ​​feature map contains the overall weather conditions. The second convolution unit deeply fuses global and local features, reducing feature dimensionality and computational complexity, while enhancing the expression of key features and suppressing irrelevant noise, providing easier-to-judge data input for subsequent severe weather identification.

[0030] In some embodiments, the weather recognition module includes a third convolution unit, a fully connected layer unit, and a classification unit connected in sequence;

[0031] The third convolution unit is configured to perform convolution processing on the first splicing feature map to determine a third splicing feature map;

[0032] The fully connected layer unit is configured to determine feature data according to the third spliced ​​feature map;

[0033] The classification unit is configured to determine the recognition result according to the feature data.

[0034] In this way, the weather recognition module includes a third convolutional unit, a fully connected layer unit, and a classification unit connected in sequence. The third convolutional unit is configured to perform convolution processing on the first spliced ​​feature map to determine a third spliced ​​feature map; the fully connected layer unit is configured to determine feature data based on the third spliced ​​feature map; and the classification unit is configured to determine the recognition result based on the feature data. In this way, the fully connected layer unit maps the abstract feature space into a two-dimensional classification space, achieving a nonlinear mapping between features and weather categories, and obtaining a feature score. The classification unit can convert the score into a probability value, and the probability output provides a basis for determining inclement weather.

[0035] In some embodiments, the recognition result includes a first probability and a second probability, wherein the first probability is used to indicate the probability that the current weather is the target weather, and the second probability is used to indicate the probability that the current weather is not the target weather.

[0036] The recognition result includes a first probability and a second probability. The first probability indicates the probability that the current weather is the target weather, while the second probability indicates the probability that the current weather is not the target weather. This allows the recognition result to represent the probability of the current weather being inclement or non-inclement. A weighted calculation based on these probabilities can yield accurate weather recognition results.

[0037] In certain embodiments, the method further comprises:

[0038] Determining a first probability weighted result and a second probability weighted result at a current moment based on a first preset weight corresponding to the first probability, a second preset weight corresponding to the second probability, and the first probability and the second probability respectively corresponding to each piece of first driving environment image information acquired within a preset time period;

[0039] When the first probability weighted results and the second probability weighted results at multiple consecutive moments both meet a preset condition, the current weather is identified as the target weather.

[0040] In this way, the first and second probability weighting results for the current moment are determined based on the first preset weight corresponding to the first probability, the second preset weight corresponding to the second probability, and the first and second probabilities corresponding to each piece of first driving environment image information acquired within a preset time period. If the first and second probability weighting results for multiple consecutive moments meet the preset conditions, the current weather is identified as the target weather. In this way, the probability weighting of multiple moments can more accurately determine the current weather, improving the robustness of the entire system and user experience in complex weather conditions.

[0041] In some embodiments, the second driving environment image information includes a first vehicle environment image and a second vehicle environment image taken at the current moment, the first vehicle environment image is obtained by taking a first vehicle front-view camera, and the second vehicle environment image is obtained by taking a first vehicle infrared camera, and the shooting direction of the first vehicle front-view camera is the same as the shooting direction of the first vehicle infrared camera.

[0042] In this way, the second driving environment image information includes the first vehicle environment image and the second vehicle environment image captured at the current moment. The first vehicle environment image is captured by the first vehicle's front-view camera, and the second vehicle environment image is captured by the first vehicle's infrared camera. The shooting direction of the first vehicle's front-view camera is the same as that of the first vehicle's infrared camera. In this way, the second driving environment image, combined with the visible light band image and the corresponding thermal radiation image, can overcome the limitations of a single modality in inclement weather and provide high-quality input for subsequent target detection.

[0043] In some embodiments, when the current weather is target weather, controlling the vehicle according to the second driving environment image information includes:

[0044] When the current weather is the target weather, fusing the first vehicle environment image and the second vehicle environment image to determine a first fused image;

[0045] The vehicle is controlled according to the first fused image.

[0046] In this way, when the current weather is the target weather, the first and second vehicle environment images are fused to determine a first fused image; the vehicle is controlled based on this first fused image. This fusion of visual details and thermal signatures improves target recognition integrity, providing complete and accurate data for subsequent vehicle control.

[0047] In some embodiments, when the current weather is the target weather, fusing the first vehicle environment image and the second vehicle environment image to determine a first fused image includes:

[0048] When the current weather is the target weather, the first vehicle environment image and the second vehicle environment image are fused according to the first pixel weight corresponding to the first vehicle environment image and the second pixel weight matrix corresponding to the second vehicle environment image to determine the first fused image.

[0049] In this way, when the current weather is the target weather, the first vehicle environment image and the second vehicle environment image are fused according to a first pixel weight matrix corresponding to the first vehicle environment image and a second pixel weight matrix corresponding to the second vehicle environment image to determine a first fused image. This method, by performing pixel-by-pixel weighted calculations on the first and second vehicle environment images, enables pixel-level fusion of the two types of images, allowing the visible light image and infrared image to complement each other within the same resolution space. This reduces the computational complexity of high-resolution images and improves target detection in inclement weather.

[0050] In some embodiments, the step of obtaining the first pixel weight matrix and the second pixel weight matrix includes:

[0051] performing a fusion process on a plurality of first vehicle environment image samples and a plurality of second vehicle environment image samples according to a predetermined second pixel weight and a third pixel weight to determine a plurality of second fused images, wherein the first vehicle environment image samples are captured by a second vehicle front-view camera, the second vehicle environment image samples are captured by a second vehicle infrared camera, and the shooting direction of the second vehicle front-view camera is the same as the shooting direction of the second vehicle infrared camera;

[0052] determining a first object detection result based on a first object detection model and a portion of the second fused image, wherein the first object detection model is trained based on at least another portion of the second fused image;

[0053] According to the first target detection result and the pre-calibrated detection target in the second fused image, the second pixel weight and the third pixel weight are updated to determine the first pixel weight matrix and the second pixel weight matrix.

[0054] In this way, multiple first vehicle environment image samples and multiple second vehicle environment image samples are fused based on predetermined second pixel weights and third pixel weights to determine multiple second fused images, where the first vehicle environment image samples are captured by a second vehicle front-view camera and the second vehicle environment image samples are captured by a second vehicle infrared camera, with the second vehicle front-view camera and the second vehicle infrared camera shooting in the same direction. A first target detection result is determined based on a first target detection model and a portion of the second fused image, where the first target detection model is trained based on at least another portion of the second fused image. Based on the first target detection result and the pre-calibrated detected targets within the second fused image, the second pixel weights and third pixel weights are updated to determine a first pixel weight matrix and a second pixel weight matrix. In this way, using the first target detection model to infer the fused image can verify target detection accuracy and ensure that the first pixel weight matrix and the second pixel weight matrix improve detection performance in inclement weather. The quantum particle swarm optimization algorithm can automatically learn the optimal weight allocation for different weather scenarios, making the fused image more suitable for machine recognition and improving the robustness of target detection in inclement weather.

[0055] In some embodiments, when the current weather is the target weather, fusing the first vehicle environment image and the second vehicle environment image to determine a first fused image includes:

[0056] When the current weather is the target weather, preprocessing the first vehicle environment image and the second vehicle environment image respectively to determine a first vehicle environment processed image and a second vehicle environment processed image;

[0057] The fusion process is performed on the first vehicle environment processed image and the second vehicle environment processed image to determine the first fused image.

[0058] In this way, when the current weather is the target weather, the first and second vehicle environment images are preprocessed separately to determine the first and second processed vehicle environment images. The first and second processed vehicle environment images are then fused to determine the first fused image. This ensures that the pixels in the first and second vehicle environment images are aligned at the same point in the scene, preventing target misalignment during the fusion process. Furthermore, unified resolution processing facilitates the subsequent pixel-by-pixel weighting matrix calculation, avoiding fusion errors caused by inconsistent scales, reducing feature extraction complexity, and improving target detection efficiency.

[0059] In some embodiments, when the current weather is the target weather, preprocessing the first vehicle environment image and the second vehicle environment image respectively to determine the first vehicle environment processed image and the second vehicle environment processed image includes:

[0060] When the current weather is the target weather, the first vehicle environment image and the second vehicle environment image are respectively cropped and / or resolution adjusted to determine the first vehicle environment processed image and the second vehicle environment processed image, wherein the field of view corresponding to the first vehicle environment processed image is the same as the field of view corresponding to the second vehicle environment processed image, and the resolution of the first vehicle environment processed image is the same as the resolution of the second vehicle environment processed image.

[0061] In this way, when the current weather is the target weather, the first vehicle environment image and the second vehicle environment image are cropped and / or resolution-adjusted, respectively, to determine the first vehicle environment processed image and the second vehicle environment processed image. The field of view corresponding to the first vehicle environment processed image is the same as the field of view corresponding to the second vehicle environment processed image, and the resolution of the first vehicle environment processed image is the same as the resolution of the second vehicle environment processed image. In this way, through cropping and resolution adjustment, image features and specifications can be unified, ensuring the consistency of the fusion area, eliminating errors caused by inconsistent fields of view and resolution, and facilitating subsequent weight matrix calculations. It can also suppress noise and adapt to the weight matrix, thereby improving subsequent fusion accuracy and detection results.

[0062] In some embodiments, controlling the vehicle according to the first fused image includes:

[0063] determining, based on the first fused image and the third fused image, a second target detection result corresponding to the first fused image, wherein the third fused image is determined by performing the fusion processing on a third vehicle environment image and a fourth vehicle environment image captured at a previous moment, the third vehicle environment image being captured by the first vehicle front-view camera, and the fourth vehicle environment image being captured by the first vehicle infrared camera;

[0064] The vehicle is controlled according to the second target detection result.

[0065] In this way, based on the first and third fused images, a second target detection result corresponding to the first fused image is determined, and the vehicle is controlled based on the second target detection result. In this way, by performing target detection using fused images of the current moment and past moments, the interference of instantaneous noise at the current moment can be reduced, and an accurate detection result at the current moment can be determined, thereby enabling precise vehicle control.

[0066] In some embodiments, the second target detection result includes a detection result for at least one of a vehicle target, a pedestrian target, and a driver target.

[0067] In this way, the second target detection result includes a detection result for at least one of a vehicle target, a pedestrian target, and a driver target. In this way, the detection result of at least one of a vehicle target, a pedestrian target, and a driver target can provide data support for driving safety decisions.

[0068] In some embodiments, determining a second target detection result corresponding to the first fused image based on the first fused image and the third fused image includes:

[0069] The second target detection result is determined according to the first fused image, the third fused image, and a pre-trained second target detection model.

[0070] In this way, a second target detection result is determined based on the first and third fused images and the pre-trained second target detection model. This pre-trained second target detection model, through an adaptive weighting mechanism, collaborates with the backbone network and feature fusion to perform target detection on fused images from both the current moment and past moments, reducing the interference of instantaneous noise at the current moment and determining a precise detection result for the current moment, thereby enabling precise vehicle control.

[0071] In some embodiments, the third fused image includes multiple images, and the second target detection model includes a preset processing network, a backbone network, and a detection network;

[0072] The preset processing network is configured to perform feature extraction and splicing processing on the first fused image and the plurality of third fused images to determine a fifth feature map;

[0073] The backbone network is configured to determine a target feature map based on the fifth feature map;

[0074] The detection network is configured to determine the second target detection result based on the target feature map.

[0075] In this way, the third fused image includes multiple images, and the second target detection model includes a preset processing network, a backbone network, and a detection network. The preset processing network is configured to perform feature extraction and splicing on the first fused image and multiple third fused images to determine a fifth feature map. The backbone network is configured to determine a target feature map based on the fifth feature map. The detection network is configured to determine the second target detection result based on the target feature map. In this way, the second target detection model works through the collaborative operation of multiple modules. The preset processing network fuses multi-moment feature data from the current and previous moments, the backbone network performs multi-scale feature extraction on this multi-moment feature data, and the detection network ultimately outputs a final detection result based on the features, achieving accurate target recognition in inclement weather.

[0076] In some embodiments, the preset processing network includes a second feature extraction module, a third feature extraction module, a third feature splicing module, a third feature extraction module, a fourth feature splicing module, a fourth feature extraction module, and a fifth feature extraction module;

[0077] The second feature extraction module is configured to perform feature extraction processing on the first fused image to determine a sixth feature map corresponding to the first fused image;

[0078] The third feature extraction module is configured to perform feature extraction processing on the third fused image to determine a seventh feature map corresponding to the third fused image;

[0079] The third feature stitching module is configured to stitch the seventh feature map corresponding to each of the third fused images to determine a third stitching feature map;

[0080] The fourth feature extraction module is configured to perform feature extraction processing on the third spliced ​​feature map to determine an eighth feature map;

[0081] The fourth feature splicing module is configured to perform splicing processing on the sixth feature map and the eighth feature map to determine a fourth splicing feature map;

[0082] The fifth feature extraction module is configured to perform feature extraction processing on the fourth splicing feature map to determine the fifth feature map.

[0083] In this way, the preset processing network includes a second feature extraction module, a third feature extraction module, a third feature stitching module, a third feature extraction module, a fourth feature extraction module, a fourth feature extraction module and a fifth feature extraction module; the second feature extraction module is configured to perform feature extraction processing on the first fused image to determine the sixth feature map corresponding to the first fused image; the third feature extraction module is configured to perform feature extraction processing on the third fused image to determine the seventh feature map corresponding to the third fused image; the third feature stitching module is configured to perform stitching processing on the seventh feature map corresponding to each third fused image to determine the third stitching feature map; the fourth feature extraction module is configured to perform feature extraction processing on the third stitching feature map to determine the eighth feature map; the fourth feature stitching module is configured to perform stitching processing on the sixth feature map and the eighth feature map to determine the fourth stitching feature map; the fifth feature extraction module is configured to perform feature extraction processing on the fourth stitching feature map to determine the fifth feature map. In this way, the pre-set processing network, through multi-branch feature extraction, can capture more complex semantic information, enhancing the detection capabilities of objects of varying sizes. Furthermore, by splicing feature images from different time dimensions, it can integrate multi-scale information, such as shallow high-resolution features and deep semantic features, thereby increasing feature richness. The collaboration between the feature extraction module and the feature splicing module maximizes the capture of the essential characteristics of the target, improving the efficiency and accuracy of subsequent object detection.

[0084] In some embodiments, the second feature extraction module includes a first feature extraction unit, a first feature fusion unit, and a second feature extraction unit;

[0085] The first feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the first fused image to determine a first processed image;

[0086] The first feature fusion unit is configured to perform dimension expansion, convolution, and splicing on the first processed image to determine a second processed image;

[0087] The third feature extraction unit is configured to perform convolution, batch normalization and nonlinear activation processing on the second processed image to determine the sixth feature map.

[0088] In this way, the second feature extraction module includes a first feature extraction unit, a first feature fusion unit, and a second feature extraction unit; the first feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation on the first fused image to determine a first processed image; the first feature fusion unit is configured to perform dimensionality expansion, convolution, and splicing on the first processed image to determine a second processed image; and the third feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation on the second processed image to determine a sixth feature map. In this way, the second feature extraction module can perform multi-scale feature extraction and preliminary processing on the first fused image. The feature extraction unit reduces the impact of illumination fluctuations caused by inclement weather, and the residual connection of the feature fusion unit alleviates the vanishing gradient problem of deep networks, ensuring the stability of feature extraction and improving the detection capability of targets of different sizes.

[0089] In some embodiments, the third feature extraction module includes a third feature extraction unit, a second feature fusion unit, and a fourth feature extraction unit;

[0090] The third feature extraction unit is configured to perform convolution, batch normalization and nonlinear activation processing on the third fused image to determine a third processed image;

[0091] The second feature fusion unit is configured to perform dimension expansion, convolution and splicing processing on the third processed image to determine a fourth processed image;

[0092] The fourth feature extraction unit is configured to perform convolution, batch normalization and nonlinear activation processing on the fourth processed image to determine the seventh feature map.

[0093] Thus, the third feature extraction module includes a third feature extraction unit, a second feature fusion unit, and a fourth feature extraction unit; the third feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the third fused image to determine a third processed image; the second feature fusion unit is configured to perform dimensional expansion, convolution, and splicing processing on the third processed image to determine a fourth processed image; and the fourth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the fourth processed image to determine a seventh feature map. In this way, the third feature extraction module can perform multi-scale feature extraction and preliminary processing on the third fused image. The feature extraction unit reduces the impact of illumination fluctuations caused by inclement weather. The residual connection of the feature fusion unit alleviates the vanishing gradient problem of the deep network, ensuring the stability of feature extraction and providing multiple features from previous moments for subsequent splicing processing.

[0094] In some embodiments, the fourth feature extraction module includes a fifth feature extraction unit, a third feature fusion unit, a sixth feature extraction unit, and a fourth feature fusion unit;

[0095] The fifth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the third spliced ​​feature map to determine a fifth processed image;

[0096] The third feature fusion unit is configured to perform dimension expansion, convolution and splicing on the fifth processed image to determine a sixth processed image;

[0097] The sixth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the sixth processed image to determine a seventh processed image;

[0098] The fourth feature fusion unit is configured to perform dimension expansion, convolution and splicing on the seventh processed image to determine the eighth feature map.

[0099] Thus, the fourth feature extraction module includes a fifth feature extraction unit, a third feature fusion unit, a sixth feature extraction unit, and a fourth feature fusion unit; the fifth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the third spliced ​​feature map to determine a fifth processed image; the third feature fusion unit is configured to perform dimensional expansion, convolution, and splicing processing on the fifth processed image to determine a sixth processed image; the sixth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the sixth processed image to determine a seventh processed image; and the fourth feature fusion unit is configured to perform dimensional expansion, convolution, and splicing processing on the seventh processed image to determine an eighth feature map. In this way, the fourth feature extraction module can perform multi-scale feature extraction and preliminary processing on the third spliced ​​image. The feature extraction unit reduces the impact of illumination fluctuations caused by bad weather. The residual connection of the feature fusion unit alleviates the gradient vanishing problem of the deep network, ensures the stability of feature extraction, and provides multiple features of the previous time for subsequent splicing processing.

[0100] In some embodiments, the fifth feature extraction module includes a seventh feature extraction unit, a fifth feature fusion unit, and an eighth feature extraction unit;

[0101] The seventh feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the fourth spliced ​​feature map to determine an eighth processed image;

[0102] The fifth feature fusion unit is configured to perform dimension expansion, convolution, and splicing on the eighth processed image to determine a ninth processed image;

[0103] The eighth feature extraction unit is configured to perform convolution, batch normalization and nonlinear activation processing on the ninth processed image to determine the fifth feature map.

[0104] Thus, the fifth feature extraction module includes a seventh feature extraction unit, a fifth feature fusion unit, and an eighth feature extraction unit; the seventh feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the fourth spliced ​​feature map to determine the eighth processed image; the fifth feature fusion unit is configured to perform dimensional expansion, convolution, and splicing processing on the eighth processed image to determine the ninth processed image; and the eighth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the ninth processed image to determine the fifth feature map. In this way, the fifth feature extraction module can perform multi-scale feature extraction and preliminary processing on the fourth spliced ​​feature map. The feature extraction unit reduces the impact of illumination fluctuations caused by bad weather. The residual connection of the feature fusion unit alleviates the gradient vanishing problem of the deep network, ensuring the stability of feature extraction and providing multiple features of the previous moment for subsequent splicing processing.

[0105] In some embodiments, the backbone network includes a sixth feature fusion unit, a ninth feature extraction unit, a seventh feature fusion unit, a tenth feature extraction unit, an eighth feature fusion unit, and a ninth feature fusion unit;

[0106] The sixth feature fusion unit is configured to perform dimension expansion, convolution, and splicing processing on the fifth feature map according to the attention data corresponding to the fifth feature map to determine a tenth processed image;

[0107] The ninth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the tenth processed image to determine a twelfth processed image;

[0108] The seventh feature fusion unit is configured to perform dimension expansion, convolution and splicing on the eleventh processed image to determine a twelfth processed image;

[0109] The tenth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the twelfth processed image to determine a thirteenth processed image;

[0110] The eighth feature fusion unit is configured to perform dimension expansion, convolution, and splicing on the thirteenth processed image to determine a fourteenth processed image;

[0111] The ninth feature fusion unit is configured to perform convolution, pooling and splicing on the fourteenth processed image to determine the target feature map.

[0112] Thus, the backbone network includes a sixth feature fusion unit, a ninth feature extraction unit, a seventh feature fusion unit, a tenth feature extraction unit, an eighth feature fusion unit and a ninth feature fusion unit; the sixth feature fusion unit is configured to perform dimension expansion, convolution and splicing processing on the fifth feature map according to the attention data corresponding to the fifth feature map to determine the tenth processed image; the ninth feature extraction unit is configured to perform convolution, batch normalization and non-linear activation processing on the tenth processed image to determine the twelfth processed image; the seventh feature fusion unit is configured to perform dimension expansion, convolution and splicing processing on the eleventh processed image to determine the twelfth processed image; the tenth feature extraction unit is configured to perform convolution, batch normalization and non-linear activation processing on the twelfth processed image to determine the thirteenth processed image; the eighth feature fusion unit is configured to perform dimension expansion, convolution and splicing processing on the thirteenth processed image to determine the fourteenth processed image; the ninth feature fusion unit is configured to perform convolution, pooling and splicing processing on the fourteenth processed image to determine the target feature map. In this way, the backbone network dynamically adapts to environmental changes through a combination of attention mechanisms. By alternately cascading feature extraction and fusion units, it extracts multi-scale features from the weighted visible and infrared fusion images, further enhancing target representation. By outputting multi-scale feature maps to the detection head, it supports the classification and localization of objects of varying sizes, improving detection capabilities for objects of varying sizes and enabling accurate detection of vehicles, pedestrians, and other targets.

[0113] In some embodiments, the detection network is configured to determine the second target detection result based on the target feature map, the attention data corresponding to the tenth processed image, the attention data corresponding to the thirteenth processed image, and the attention data corresponding to the target feature map.

[0114] Thus, the detection network is configured to determine a second target detection result based on the target feature map, the attention data corresponding to the tenth processed image, the attention data corresponding to the thirteenth processed image, and the attention data corresponding to the target feature map. In this way, integrating the channel, spatial, and frequency channel attention data allows for dynamic adjustment of feature weights, achieving both focus on key information and noise suppression, resulting in accurate detection results.

[0115] An embodiment of the present application provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the steps of the above method are implemented.

[0116] An embodiment of the present application provides a vehicle, including the above-mentioned electronic device, to implement the steps of the above-mentioned vehicle control method.

[0117] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by one or more processors, the steps of the above method are implemented.

[0118] An embodiment of the present application provides a computer program product, including a computer program / instruction, which implements the steps of the above method when executed by a processor.

[0119] The electronic device, vehicle, computer-readable storage medium, and computer program product provided by the embodiments of the present application determine the current weather recognition result based on the first driving environment image information; when the current weather is the target weather, the vehicle is controlled based on the second driving environment image information, wherein the visibility corresponding to the target weather is less than or equal to a preset threshold. In this way, when the system recognizes that the weather is bad through the images of the vehicle's surrounding environment, it automatically switches to the bad weather mode. By fusing characteristic images in bad weather, target detection is performed on the fused images, and the vehicle is controlled to actively change lanes, actively follow vehicles, etc. based on the target detection results, thereby ensuring the stable operation of the vehicle in the target weather, and further ensuring the safety of the vehicle and its passengers to a certain extent.

[0120] Additional aspects and advantages of the embodiments of the present application will be given in part in the description below, and in part will become obvious from the description below, or will be learned through practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0121] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0122] Figure 1 This is one of the flowcharts of the vehicle control method according to certain embodiments of the present application;

[0123] Figure 2 This is a second flow chart of a vehicle control method according to certain embodiments of the present application;

[0124] Figure 3 is a schematic diagram of an image stitching process of a vehicle control method according to certain embodiments of the present application;

[0125] Figure 4 This is a third flow chart of a vehicle control method according to certain embodiments of the present application;

[0126] Figure 5 is a schematic diagram of the weather recognition model structure of the vehicle control method in certain embodiments of the present application;

[0127] Figure 6This is one of the weather identification flow diagrams of the vehicle control method according to certain embodiments of the present application;

[0128] Figure 7 This is the second schematic diagram of the weather identification process of the vehicle control method in certain embodiments of the present application;

[0129] Figure 8 This is a fourth flow chart of a vehicle control method according to certain embodiments of the present application;

[0130] Figure 9 This is a fifth flow chart of a vehicle control method according to certain embodiments of the present application;

[0131] Figure 10 This is the sixth flow chart of the vehicle control method according to certain embodiments of the present application;

[0132] Figure 11 This is the seventh flow chart of the vehicle control method according to certain embodiments of the present application;

[0133] Figure 12 This is the eighth flow chart of the vehicle control method according to certain embodiments of the present application;

[0134] Figure 13 This is a ninth flowchart of a vehicle control method according to certain embodiments of the present application;

[0135] Figure 14 is a schematic diagram of a preset processing network structure of a vehicle control method according to certain embodiments of the present application;

[0136] Figure 15 is a schematic diagram of the backbone network structure of a vehicle control method according to certain embodiments of the present application;

[0137] Figure 16 It is a schematic diagram of a vehicle control method according to certain embodiments of the present application. DETAILED DESCRIPTION

[0138] The embodiments of the present application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of the present application, and should not be understood as limiting the embodiments of the present application.

[0139] In related technologies, vehicles can use lidar to perform three-dimensional modeling of the vehicle's surrounding environment, thereby sensing the vehicle's external environment and providing data basis for autonomous driving, such as active lane changing and active following.

[0140] LiDAR is easily affected by weather interference and may fail. When a vehicle encounters severe weather such as heavy rain or snow, the reflected signal may deviate, causing the system to generate erroneous environmental perception data, making it impossible for the vehicle to correctly perceive the outside world. This can lead to abnormalities or failures in functions such as active lane change and active following. This not only seriously undermines the user's driving comfort, but may also cause serious traffic accidents such as rear-end collisions and rollovers in high-speed driving scenarios, posing a major threat to the safety of users' lives and vehicle property.

[0141] Based on the above questions, please refer to Figure 1 , an embodiment of the present application provides a vehicle control method, the method comprising:

[0142] 01: Determine the recognition result of the current weather according to the first driving environment image information;

[0143] 02: When the current weather is target weather, the vehicle is controlled according to the second driving environment image information, wherein the visibility corresponding to the target weather is less than or equal to a preset threshold.

[0144] The embodiment of the present application provides a vehicle control device. The vehicle control method of the embodiment of the present application can be implemented by the vehicle control device of the embodiment of the present application. Specifically, the vehicle control device includes a weather recognition module and a control module. The weather recognition module is used to determine the recognition result of the current weather based on the first driving environment image information. The control module is used to control the vehicle based on the second driving environment image information when the current weather is the target weather, wherein the visibility corresponding to the target weather is less than or equal to a preset threshold.

[0145] The embodiment of the present application also provides an electronic device, which includes a memory and a processor. The vehicle control method of the embodiment of the present application can be implemented by the electronic device of the embodiment of the present application. Specifically, a computer program is stored in the memory, and the processor is used to determine the recognition result of the current weather based on the first driving environment image information. The processor is also used to control the vehicle based on the second driving environment image information when the current weather is the target weather, wherein the visibility corresponding to the target weather is less than or equal to a preset threshold.

[0146] Specifically, the first driving environment image information refers to image data of the vehicle's surrounding environment during driving, and is used to determine whether the vehicle's current driving environment is inclement weather. The current weather recognition result refers to the current weather conditions of the vehicle's surrounding environment, including both inclement weather and non-inclement weather. By comprehensively evaluating the surrounding environment image data, it is possible to determine whether the vehicle's current driving environment is inclement.

[0147] Target weather refers to inclement weather with visibility less than or equal to a preset threshold. The second driving environment image information refers to image data and infrared image data of the environment ahead of the vehicle. When the vehicle is in inclement weather with low visibility, such as fog with visibility less than 100 meters, the image data and infrared image data of the environment ahead of the vehicle's route can be combined to perform target detection in the environment ahead, providing a basis for vehicle control.

[0148] In the implementation mode of the present application, during driving, the first driving environment image information of the vehicle is identified. If the identification result is bad weather, in order to avoid the vehicle driving being affected by the bad weather, the vehicle can be controlled according to the second driving environment image information, such as controlling the vehicle to actively change lanes, actively follow the vehicle, etc., to ensure the stable operation of the vehicle in the target weather.

[0149] In summary, in the embodiments of the present application, when the current weather recognition result is determined based on the first driving environment image information and the current weather is determined to be target weather, the vehicle can be controlled based on the second driving environment image information, such as to control the vehicle to actively change lanes or actively follow other vehicles, thereby ensuring the vehicle's stable operation in the target weather and, to a certain extent, protecting the safety of the vehicle and its passengers. Furthermore, when equipment such as lidar fails due to weather interference, the vehicle can determine the current weather recognition result based on the first driving environment image information and control it based on the second driving environment image information, further ensuring the vehicle's stable operation in the target weather.

[0150] In some embodiments, the first driving environment image information includes a plurality of surround view images captured by a plurality of vehicle surround view cameras having different shooting directions.

[0151] Specifically, the camera captures images from the front, rear, left, and right of the vehicle. Surround-view cameras are deployed in the front, rear, left, and right sides of the vehicle to capture multi-directional surround-view images in real time. These multi-directional surround-view images encompass a wide range of environmental information surrounding the vehicle, providing a wider range for identifying inclement weather conditions. This avoids errors caused by solely relying on forward-looking vehicle recognition and improves the accuracy of inclement weather assessments.

[0152] In this way, the first driving environment image information includes multiple surround-view images captured by multiple vehicle surround-view cameras with different shooting directions. This multi-directional surround-view image provides a wider recognition range for inclement weather, avoiding misidentification in a single direction, thereby improving the accuracy of weather judgment.

[0153] See also Figure 2In some embodiments, step 01 (determining the recognition result of the current weather according to the first driving environment image information) includes:

[0154] 011: stitching multiple surround view images to determine a stitched image;

[0155] 012: Determine the recognition result based on the stitched image and multiple surround view images.

[0156] In some embodiments, the weather recognition module is further configured to stitch the multiple surround view images to determine a stitched image. The weather recognition module is further configured to determine a recognition result based on the stitched image and the multiple surround view images.

[0157] In some embodiments, the processor is further configured to stitch the multiple surround view images to determine a stitched image, and further configured to determine a recognition result based on the stitched image and the multiple surround view images.

[0158] Specifically, a stitched image refers to an image that integrates multi-directional environmental information and is obtained by stitching multiple surround view images. Stitching is the process of stitching multiple surround view images in the channel direction. Figure 3 The stitching process is explained using the example of surround view 1-4, which refers to four 3-channel environmental images in different directions. After stitching in the channel direction, the image information in the front, rear, left, and right directions of the vehicle can be integrated into the same image to obtain a 3*4=12-channel stitched image, providing more comprehensive environmental feature input for subsequent severe weather recognition.

[0159] In the process of identifying severe weather, it is necessary to perform feature extraction processing on the stitched image and multiple surround images respectively, and then perform feature fusion processing on the feature extraction results to obtain a fused feature image. Finally, identification is performed based on the fused feature image to determine whether the current weather is severe weather.

[0160] In this way, multiple surround view images are stitched together to create a stitched image. Recognition results are determined based on this stitched image and the multiple surround view images. This stitching of surround view images along the path integrates multi-dimensional environmental information, providing a more comprehensive input for subsequent weather recognition. Weather recognition based on this stitched image and the multiple surround view images can determine whether the current weather is severe.

[0161] See also Figure 4 In some embodiments, step 012 (determining a recognition result based on the stitched image and the multiple surround view images) includes:

[0162] 0121: Determine the recognition result based on the stitched image, multiple surround images, and the pre-trained weather recognition model.

[0163] In some embodiments, the weather recognition module is further configured to determine a recognition result based on the stitched image and the multiple surround images, and a pre-trained weather recognition model 100 .

[0164] In some embodiments, the processor is further configured to determine a recognition result based on the stitched image and the multiple surround images, and a pre-trained weather recognition model.

[0165] Specifically, the weather recognition model 100 performs feature extraction, feature fusion, and weather recognition processing on the stitched image and surround view images, outputting a weather determination result. By feeding the stitched image and multiple surround view images into the pre-trained weather recognition model 100, the model analyzes the vehicle's surroundings and, by identifying target objects within the environment, determines whether the current weather is severe.

[0166] In this way, the recognition result is determined based on the stitched image and multiple surround view images, as well as the pre-trained weather recognition model. By analyzing the stitched image and multiple surround view images using the pre-trained weather recognition model, feature extraction, feature fusion, and weather recognition can be quickly performed, resulting in accurate weather recognition results.

[0167] See also Figure 5 In some embodiments, the weather recognition model 100 includes a first feature extraction module 110, a first feature splicing module 120, and a weather recognition module 130 connected in sequence.

[0168] The first feature extraction module 110 is configured to perform feature extraction processing on the stitched image and the surround view image respectively, and determine a first feature map of the stitched image and a second feature map of the surround view image;

[0169] The first feature splicing module 120 is configured to perform splicing processing on the first feature map and the plurality of second feature maps to determine a first spliced ​​feature map;

[0170] The weather recognition module 130 is configured to determine a recognition result according to the first splicing feature map.

[0171] Specifically, the weather recognition model 100 refers to the weather recognition model 100 pre-trained in the above embodiment. Figure 6 The recognition process of the weather recognition module 130 is explained as follows:

[0172] The first feature extraction module 110 is used to perform feature extraction processing on the stitched image and multiple surround view images, where the feature extraction processing includes convolution processing and pooling processing. The convolution processing can convert the information of the stitched image and multiple surround view images into abstract feature vectors suitable for classification. The pooling processing can compress the size of the feature image, thereby reducing the amount of calculation while retaining key information, facilitating subsequent feature fusion.

[0173] The first feature map is derived from the stitched image through feature extraction and represents global weather characteristics, such as overall lighting changes and the degree of blur in multiple directions. The second feature map is derived from the surround view image through feature extraction and represents local details in each direction, such as the distribution of raindrops on a particular side and differences in fog concentration, complementing the global features.

[0174] By performing feature extraction processing on the stitched image through the first feature extraction module 110, local interference in a single direction can be avoided. By performing feature extraction processing on multiple surround view images through the first feature extraction module 110, the adaptability of the subsequent weather recognition module 130 algorithm to complex weather scenarios can be improved. The first feature stitching module 120 is used to stitch the first feature map and multiple second feature maps. The stitching processing includes channel stitching and convolution processing. The first stitched feature map is obtained by stitching the first feature map and multiple second feature maps, and represents multi-directional detailed information and overall weather conditions.

[0175] The channel stitching process is performed in the channel direction, integrating the global environmental features of the first feature map and the local features in each direction of the multiple second feature maps, and deeply fusing the feature information through convolution processing, so that the features from different sources, namely the global environmental features of the first feature map and the local features in each direction of the multiple second feature maps, are deeply fused in the same feature space to form a more discriminative comprehensive feature. It can reduce the feature dimension and computational complexity, while enhancing the expression of key features such as texture and brightness changes related to bad weather, and suppressing irrelevant noise. The weather recognition module 130 is used to classify the first spliced ​​feature map into bad weather. By identifying weather-related features in the image, calculating the probability of bad weather and non-bad weather, and comparing them, it is determined whether the weather is bad weather and classifies the weather category.

[0176] Thus, the weather recognition model 100 includes a first feature extraction module 110, a first feature stitching module 120, and a weather recognition module 130, which are connected in sequence. The first feature extraction module 110 is configured to perform feature extraction processing on the stitched image and the surround view image, respectively, to determine a first feature map for the stitched image and a second feature map for the surround view image; the first feature stitching module 120 is configured to stitch the first feature map and multiple second feature maps to determine a first stitched feature map; and the weather recognition module 130 is configured to determine a recognition result based on the first stitched feature map. In this way, the weather recognition module 130 can perform feature extraction, feature stitching, and weather recognition processing on the stitched image and the surround view image, thereby accurately determining whether the vehicle is in inclement weather.

[0177] In some embodiments, the first feature extraction module 110 includes a first convolution unit and a first pooling unit;

[0178] The first convolution unit is configured to perform feature extraction processing on the stitched image and the surround view image respectively, and determine a third feature map of the stitched image and a fourth feature map of the surround view image.

[0179] The first pooling unit is configured to perform pooling processing on the third feature map of the stitched image and the fourth feature map of the surround view image, respectively, to determine the first feature map of the stitched image and the second feature map of the surround view image.

[0180] Specifically, the first convolution unit is used to perform feature extraction processing on the stitched image and the surround view image respectively, converting the stitched image and multiple surround view image information into abstract feature vectors suitable for classification, and obtaining the third feature map of the stitched image and the fourth feature map of the surround view image.

[0181] The first pooling unit is used to perform pooling processing on the third feature map of the stitched image and the fourth feature of the surround view image, which can compress the size of the feature image, retain key information while reducing the amount of calculation, and facilitate subsequent feature fusion.

[0182] The first pooling unit is cascaded to the first convolution unit, and the first feature extraction module 110 may include multiple first pooling units and first convolution units. In one example, please refer to Table 1, which shows the specific structure of the first feature extraction module 110.

[0183]

[0184] The following table 1 is used to explain the working process of the first feature extraction module 110:

[0185] The layer1 processing process is as follows:

[0186] First, preliminary feature extraction is performed on the 12-channel stitched image using a 7*7 convolution kernel with 32 channels and a stride of 2, reducing the feature map size from 224*224 to 112*112 and enhancing the feature dimension. A 5*5 convolution kernel with 32 channels and a stride of 2 is then used to further extract features, again reducing the feature map size to 56*56. A 3*3 atrous convolution with a dilation factor of 1, 64 channels, and a stride of 1 is then used to expand the receptive field, capturing a wider range of environmental features, such as the global distribution of rain and fog, without increasing parameters. Finally, batch normalization (BN), RELU activation function, and global average pooling are used to standardize the features and compress their size.

[0187] The layer2 processing process is as follows:

[0188] A 3x3 convolution kernel with 32 channels and a stride of 2 is used to reduce the dimensionality of the feature map, for example, from 56x56 to 28x28. Batch normalization and Reluctant Unit (RELU) activation functions are then used to extract more abstract semantic features, such as the overall blur level in severe weather. Global average pooling is then used to resize the feature map to prepare for subsequent fusion.

[0189] In this way, the first feature extraction module 110 includes a first convolution unit and a first pooling unit; the first convolution unit is configured to perform feature extraction processing on the stitched image and the surround view image respectively to determine the third feature map of the stitched image and the fourth feature map of the surround view image; the first pooling unit is configured to perform pooling processing on the third feature map of the stitched image and the fourth feature map of the surround view image respectively to determine the first feature map of the stitched image and the second feature map of the surround view image. In this way, by extracting the features of the stitched image, it is possible to integrate all-round environmental information, extract global weather features, and avoid local interference in a single direction. By extracting the features of the surround view image, it is possible to capture local detail features in each direction. The first feature extraction module 110 can provide global and local data basis for subsequent weather classification, improve environmental perception capabilities in severe weather, and thus more accurately judge severe weather.

[0190] In some embodiments, the first feature concatenation module 120 includes a concatenation unit and a second convolution unit.

[0191] The splicing unit is configured to splice the first feature map and the plurality of second feature maps to determine a second spliced ​​feature map;

[0192] The second convolution unit is configured to perform convolution processing on the second splicing feature map to determine the first splicing feature map.

[0193] Specifically, the stitching unit is configured to stitch the first feature map and multiple second feature maps in a channel-wise manner to obtain a second stitched feature map. The second average feature map is an image containing multi-directional information, such as multi-directional details such as raindrops on one side and the overall global weather condition of dim light.

[0194] In an example, the first feature extraction module 110 outputs a first feature map with 64 channels and 5 second feature maps with 64 channels. After the splicing unit performs splicing processing in the channel direction, a second spliced ​​feature map with 64*5=320 channels can be obtained.

[0195] The second convolution unit is used to perform convolution processing on the second spliced ​​feature map to reduce the feature dimension, provide enhanced key features, such as texture and brightness changes related to severe weather, deeply fuse global and local features, and obtain a discriminative first spliced ​​feature map.

[0196] In one example, after the second convolution unit performs convolution processing on the second spliced ​​feature map with a channel number of 64*5=320, a first spliced ​​feature map with a channel number of 64 can be obtained.

[0197] Thus, the first feature splicing module 120 includes a splicing unit and a second convolution unit. The splicing unit is configured to splice the first feature map and multiple second feature maps to determine a second spliced ​​feature map. The second convolution unit is configured to convolve the second spliced ​​feature map to determine the first spliced ​​feature map. Thus, after splicing by the splicing unit, the second spliced ​​feature map contains the overall weather conditions. The second convolution unit deeply fuses global and local features, reducing feature dimensionality and computational complexity, while enhancing the expression of key features and suppressing irrelevant noise, providing easier-to-judge data input for subsequent severe weather identification.

[0198] In some embodiments, the weather recognition module 130 includes a third convolutional unit, a fully connected layer unit, and a classification unit connected in sequence.

[0199] The third convolution unit is configured to perform convolution processing on the first spliced ​​feature map to determine a third spliced ​​feature map;

[0200] The fully connected layer unit is configured to determine feature data based on the third concatenated feature map;

[0201] The classification unit is configured to determine a recognition result based on the feature data.

[0202] Specifically, the third convolution unit is used to perform convolution processing on the first spliced ​​feature map, further extract high-level semantic features, and map the underlying texture, brightness and other features into abstract representations related to weather categories. For example, simultaneous blur in multiple directions may correspond to bad weather, thereby obtaining the third spliced ​​feature map.

[0203] The fully connected layer unit consists of two neurons, each connected to all feature nodes in the third concatenated feature map. This unit is used to capture the combined relationship between different feature dimensions. For example, the co-occurrence of dim light and blurred scenery is more likely to correspond to inclement weather. Furthermore, based on the weight matrix, the fully connected layer unit can map the multi-dimensional fusion features into a two-dimensional classification space, achieving a nonlinear mapping between features and weather categories, and outputting an unnormalized classification score. For example, when the weights of the fog concentration feature and the multi-directional blur feature in the feature vector are high, the fully connected layer will output a high score for inclement weather.

[0204] The classification unit is used to classify the classification scores and convert them into probability values. The probability values ​​of severe weather and non-severe weather meet the following formula:

[0205] P bad +P good =1

[0206] The classification unit can use the Softmax function to amplify the difference between high and low scores, making the transition probability distribution steeper, improving class differentiation, and facilitating accurate recognition. The fully connected layer units learn subtle feature differences and can use the Softmax function to output probabilities that reflect these subtle differences, thereby improving the robustness of the model. In one example, when the fully connected layer unit outputs a severe weather score that is 2 points higher than a non-severe weather score, the Softmax output probabilities are approximately 0.88 and 0.12, respectively, a more significant difference than the original scores.

[0207] By calculating the severe weather and non-severe weather scores of the classification units, the final severe weather probability and non-severe weather probability can be determined.

[0208] The following is a complete example to explain the process of weather classification performed by the weather recognition module 130:

[0209] Front-end features include "global dimming," "multi-directional scene blur," and "contrast reduction." In the fully connected layer, the weight of the "blur feature" is 0.6, and the weight of the "light feature" is 0.4. This calculation yields a score of 2.5 for severe weather and 0.8 for non-severe weather. In the classification unit, the Softmax function converts the severe weather score into a probability. The calculation process is as follows:

[0210]

[0211] Non-severe weather calculations are similar to P good =0.23. This probability represents the confidence of the weather recognition model 100, P bad =0.77 indicates that the model has 77% confidence in judging that the current weather is severe weather. good =0.23 indicates that the model has a 23% confidence level in determining that the weather is not severe.

[0212] In this way, the weather recognition module 130 includes a third convolutional unit, a fully connected layer unit, and a classification unit connected in sequence. The third convolutional unit is configured to perform convolution processing on the first spliced ​​feature map to determine a third spliced ​​feature map; the fully connected layer unit is configured to determine feature data based on the third spliced ​​feature map; and the classification unit is configured to determine the recognition result based on the feature data. In this way, the fully connected layer unit maps the abstract feature space into a two-dimensional classification space, achieving a nonlinear mapping between features and weather categories, and obtaining a feature score. The classification unit can convert the score into a probability value, and the probability output provides a basis for determining inclement weather.

[0213] In some embodiments, the recognition result includes a first probability and a second probability, the first probability is used to indicate the probability that the current weather is the target weather, and the second probability is used to indicate the probability that the current weather is not the target weather.

[0214] Specifically, the first probability refers to the probability that the model determines that the current weather is severe, and the second probability refers to the probability that the weather is not severe. The sum of the first probability and the second probability is equal to 1.

[0215] The weather recognition model 100 may be subject to instantaneous interference, such as strong light flashing, sensor noise, etc., and may output an erroneous probability value. Therefore, the recognition result needs to be further processed. By weighting the first probability and the second probability, a more accurate weather recognition result can be obtained.

[0216] The recognition result includes a first probability and a second probability. The first probability indicates the probability that the current weather is the target weather, while the second probability indicates the probability that the current weather is not the target weather. This allows the recognition result to represent the probability of the current weather being inclement or non-inclement. A weighted calculation based on these probabilities can yield accurate weather recognition results.

[0217] In certain embodiments, further comprising:

[0218] Determine a first probability weighted result and a second probability weighted result at the current moment based on a first preset weight corresponding to the first probability, a second preset weight corresponding to the second probability, and the first probability and the second probability corresponding to each piece of first driving environment image information acquired within a preset time period;

[0219] When the first probability weighted results and the second probability weighted results at multiple consecutive moments all meet the preset conditions, the current weather is identified as the target weather.

[0220] In certain embodiments, the weather identification module 130 is further configured to determine a first probability weighted result and a second probability weighted result at the current moment based on a first preset weight corresponding to the first probability, a second preset weight corresponding to the second probability, and the first probability and second probability corresponding to each piece of first driving environment image information acquired within a preset time period. The weather identification module 130 is further configured to identify the current weather as target weather if the first probability weighted result and the second probability weighted result at multiple consecutive moments both meet a preset condition.

[0221] In certain embodiments, the processor is further configured to determine a first probability weighted result and a second probability weighted result at the current moment based on a first preset weight corresponding to the first probability, a second preset weight corresponding to the second probability, and the first probability and second probability corresponding to each piece of first driving environment image information acquired within a preset time period. The processor is further configured to identify the current weather as target weather if the first probability weighted result and the second probability weighted result at multiple consecutive moments both meet a preset condition.

[0222] Specifically, the first preset weight refers to the initial value of the bad weather probability weight. The second preset weight refers to the initial value of the non-bad weather probability weight. The initial values ​​of the first preset weight and the second preset weight can both be set to 0.5.

[0223] The preset duration refers to a preset length of time. The current time may be affected by instantaneous interference such as strong light flashing, resulting in the acquisition of incorrect first and second probabilities. Therefore, the results of multiple moments within one end of time can be combined to eliminate the impact of single-moment misjudgment.

[0224] The weighted first probability refers to the probability after weighting the first probabilities at multiple moments. The weighted second probability refers to the probability after weighting the second probabilities at multiple moments. The system assigns a higher weight to the most recent moment, such as a weight of 0.5 for moment T, to ensure that the system is more sensitive to current weather changes. While retaining historical reference moments, such as a weight of 0.2 for moment T-2, this avoids the randomness caused by complete reliance on a single moment. The weighted first and second probabilities provide a more accurate judgment probability for weather identification. The precondition refers to the situation where the weighted first probability is greater than the weighted second probability, and the sum of the weighted first and second probabilities is 1.

[0225] See also Figure 7The following is a complete example to explain how to determine the first probability weighted result and the second probability weighted result at the current moment. T and the probability of non-severe weather 1-P T In the case of multiple times such as T-1, T-2, etc., the probability of severe weather P is obtained. T-1 、P T-2 and the probability of non-severe weather 1-P T-1 , 1-P T-2 , and perform weight distribution:

[0226] S bad =0.5P T +0.3P T-1 +0.2P T-2

[0227] S good =0.5(1-P T )+0.3(1-P T-1 )+0.2(1-P T-2 )

[0228] Thus, in the subsequent weather determination process, if the first probability weighted result is greater than the second probability weighted result, i.e., S bad >S good , then the current weather is judged to be bad weather, otherwise it is not bad weather.

[0229] In this way, the first and second probability weighting results for the current moment are determined based on the first preset weight corresponding to the first probability, the second preset weight corresponding to the second probability, and the first and second probabilities corresponding to each piece of first driving environment image information acquired within a preset time period. If the first and second probability weighting results for multiple consecutive moments meet the preset conditions, the current weather is identified as the target weather. In this way, the probability weighting of multiple moments can more accurately determine the current weather, improving the robustness of the entire system and user experience in complex weather conditions.

[0230] In some embodiments, the second driving environment image information includes a first vehicle environment image and a second vehicle environment image taken at the current moment. The first vehicle environment image is obtained by taking a picture with the first vehicle front-view camera, and the second vehicle environment image is obtained by taking a picture with the first vehicle infrared camera. The shooting direction of the first vehicle front-view camera is the same as the shooting direction of the first vehicle infrared camera.

[0231] Specifically, the second driving environment image refers to an enhanced image containing visual details and thermal features. It is an image suitable for vehicle target detection in inclement weather conditions and includes a first vehicle environment image and a second vehicle environment image. The first vehicle environment image refers to a visible light band image of the environment in front of the vehicle, such as the color and texture of objects, and is acquired in real time by a forward-looking visible light camera. The second vehicle environment image refers to a thermal radiation image of the environment in front of the vehicle, such as the temperature distribution of objects, and is acquired in real time by a forward-looking infrared camera. The forward-looking infrared camera and the forward-looking visible light camera capture images synchronously and in the same direction to facilitate subsequent image processing.

[0232] By combining visible light band images with corresponding thermal radiation images, visual defects in severe weather conditions can be compensated. For example, visible light images are blurred in heavy fog, while infrared images can penetrate fog to display hot targets. The second driving environment image can improve target detection performance in severe weather conditions and enhance the environmental robustness of the intelligent driving system.

[0233] In this way, the second driving environment image information includes the first vehicle environment image and the second vehicle environment image captured at the current moment. The first vehicle environment image is captured by the first vehicle's front-view camera, and the second vehicle environment image is captured by the first vehicle's infrared camera. The shooting direction of the first vehicle's front-view camera is the same as that of the first vehicle's infrared camera. In this way, the second driving environment image, combined with the visible light band image and the corresponding thermal radiation image, can overcome the limitations of a single modality in inclement weather and provide high-quality input for subsequent target detection.

[0234] See also Figure 8 In some embodiments, step 02 (controlling the vehicle according to the second driving environment image information when the current weather is the target weather) includes:

[0235] 021: When the current weather is the target weather, the first vehicle environment image and the second vehicle environment image are fused to determine a first fused image;

[0236] 022: Control the vehicle according to the first fused image.

[0237] In some embodiments, the vehicle control module is further configured to fuse the first vehicle environment image and the second vehicle environment image to determine a first fused image when the current weather is the target weather. The vehicle control module is further configured to control the vehicle based on the first fused image.

[0238] In some embodiments, the processor is further configured to, when the current weather is the target weather, fuse the first vehicle environment image and the second vehicle environment image to determine a first fused image, and to control the vehicle based on the first fused image.

[0239] Specifically, the first fused image refers to an enhanced image containing visual details and thermal features, which is obtained by weighted fusion processing of the time-synchronized first vehicle environment image and the second vehicle environment image, and is used to control the vehicle. Fusion processing of the first vehicle environment image and the second vehicle environment image can compensate for the defects of visible light. In scenes such as foggy and rainy days, visible light images cause target blurring due to light scattering, while thermal radiation images can highlight the thermal contours of targets such as pedestrians and vehicles through temperature differences. It can also compensate for the limitations of infrared light. Infrared images lack color information and have low resolution, while visible light can provide texture details such as traffic sign color and lane lines. The fusion of visible and infrared light improves target recognition integrity.

[0240] In this way, when the current weather is the target weather, the first and second vehicle environment images are fused to determine a first fused image; the vehicle is controlled based on this first fused image. This fusion of visual details and thermal signatures improves target recognition integrity, providing complete and accurate data for subsequent vehicle control.

[0241] See also Figure 9 In some embodiments, step 021 (when the current weather is the target weather, fusing the first vehicle environment image and the second vehicle environment image to determine a first fused image) includes:

[0242] 0211: When the current weather is the target weather, the first vehicle environment image and the second vehicle environment image are fused according to the first pixel weight matrix corresponding to the first vehicle environment image and the second pixel weight matrix corresponding to the second vehicle environment image to determine the first fused image.

[0243] In some embodiments, the vehicle control module is also used to, when the current weather is the target weather, fuse the first vehicle environment image and the second vehicle environment image based on a first pixel weight matrix corresponding to the first vehicle environment image and a second pixel weight matrix corresponding to the second vehicle environment image to determine a first fused image.

[0244] In some embodiments, the processor is also used to, when the current weather is the target weather, fuse the first vehicle environment image and the second vehicle environment image based on a first pixel weight matrix corresponding to the first vehicle environment image and a second pixel weight matrix corresponding to the second vehicle environment image to determine a first fused image.

[0245] Specifically, the first pixel weight matrix is ​​used to adjust the size of the first vehicle environment image, and the second pixel weight matrix is ​​used to adjust the size of the second vehicle environment image. The first pixel weight matrix is ​​complementary to the second pixel weight matrix.

[0246] By performing matrix multiplication on the first vehicle environment image and the first pixel weight matrix, and by performing matrix multiplication on the second vehicle environment image and the second pixel weight matrix, the first vehicle environment image and the second vehicle environment image after matrix multiplication can have the same resolution, and a first fused image can be obtained by adding them.

[0247] In this way, when the current weather is the target weather, the first vehicle environment image and the second vehicle environment image are fused according to a first pixel weight matrix corresponding to the first vehicle environment image and a second pixel weight matrix corresponding to the second vehicle environment image to determine a first fused image. By performing pixel-by-pixel weighted calculations on the first and second vehicle environment images, the two types of images can be fused at the pixel level, allowing the visible light image and infrared image to complement each other in the same resolution space. This reduces the computational complexity of high-resolution images and improves target detection in inclement weather.

[0248] In some embodiments, the steps of obtaining the first pixel weight matrix and the second pixel weight matrix include:

[0249] performing a fusion process on the plurality of first vehicle environment image samples and the plurality of second vehicle environment image samples according to a predetermined second pixel weight and a third pixel weight to determine a plurality of second fused images, wherein the first vehicle environment image samples are captured by a second vehicle front-view camera, and the second vehicle environment image samples are captured by a second vehicle infrared camera, and a shooting direction of the second vehicle front-view camera is the same as a shooting direction of the second vehicle infrared camera;

[0250] determining a first object detection result according to a first object detection model and a portion of the second fused image, wherein the first object detection model is trained based on at least another portion of the second fused image;

[0251] According to the first target detection result and the pre-calibrated detection target in the second fused image, the second pixel weight and the third pixel weight are updated to determine the first pixel weight matrix and the second pixel weight matrix.

[0252] Specifically, the second pixel weight refers to the initial value of the initial first pixel weight matrix at each pixel position. The third pixel weight refers to the initial value of the initial second pixel weight matrix at each pixel position. The initial values ​​are all set to 0.5, and each pixel in the matrix satisfies the following formula:

[0253] W1[x][y][z]+W2[x][y][z]=1

[0254] The first vehicle environment image samples refer to multiple visible light images of the environment in front of the vehicle. The second vehicle environment image samples refer to multiple infrared images of the environment in front of the vehicle. The second fused image refers to an enhanced sample image containing visual details and thermal features, obtained by fusing the first vehicle environment image samples with the corresponding second vehicle environment image samples.

[0255] The first target detection model is used to train on the fused images of a subset of image samples to obtain the first target detection results. The pre-calibrated detection targets within the second fused image are manually compared and analyzed within each set of infrared and visible light images. The resulting fused target positions and categories are determined manually. These serve as the ground truth labels for the training and validation sets, verifying the accuracy of the detection results.

[0256] In the process of obtaining the first pixel weight matrix and the second pixel weight matrix, the optimal first pixel weight matrix W1 and the second pixel weight matrix W2 can be solved based on the quantum particle swarm optimization algorithm and the deep learning model. First, the detection loss (Loss) of the YOLOv10 model can be used as the optimization target. By performing target detection training on some image samples, the difference between the detection result and the true value is measured, that is, minimizing

[0257] Loss(YOLOv10(A*W1+B*W2))

[0258] Where A refers to the visible light image pixel matrix, and B refers to the infrared image pixel matrix.

[0259] Next, the matrix parameters are represented by particles, that is, each particle corresponds to a set of W1 and W2 parameters, and the particles are used to constrain the matrix. The constraints are as follows:

[0260] W1[x][y][z]+W2[x][y][z]=1

[0261] W1[x][y][z]≥0

[0262] W2[x][y][z]≥0

[0263] The validation set images are then processed in batches, for example, each batch contains 64 images. Each batch is processed in one iteration, and the weights are iterated 100 times before entering the next batch. Each batch calculates the loss value for the current particle position, updates the individual and global optimal positions, and updates the particle positions using quantum behavioral operators to ensure that the weights are within the constraints. The quantum particle swarm optimization algorithm can automatically learn the optimal weight distribution for different weather scenarios. For example, visible light images in foggy weather can blur the target due to light scattering, so the infrared weight can be enhanced to enhance the thermal contours of pedestrians and vehicles, making their target features more prominent. The completion of all batches is considered one round, and a total of 300 rounds of iterations are performed to ultimately solve the optimal first pixel weight matrix W1 and second pixel weight matrix W2.

[0264] In this way, multiple first vehicle environment image samples and multiple second vehicle environment image samples are fused based on predetermined second pixel weights and third pixel weights to determine multiple second fused images, where the first vehicle environment image samples are captured by a second vehicle front-view camera and the second vehicle environment image samples are captured by a second vehicle infrared camera, with the second vehicle front-view camera and the second vehicle infrared camera shooting in the same direction. A first target detection result is determined based on a first target detection model and a portion of the second fused image, where the first target detection model is trained based on at least another portion of the second fused image. Based on the first target detection result and the pre-calibrated detected targets within the second fused image, the second pixel weights and third pixel weights are updated to determine a first pixel weight matrix and a second pixel weight matrix. In this way, using the first target detection model to infer the fused image can verify target detection accuracy and ensure that the first pixel weight matrix and the second pixel weight matrix improve detection performance in inclement weather. The quantum particle swarm optimization algorithm can automatically learn the optimal weight allocation for different weather scenarios, making the fused image more suitable for machine recognition and improving the robustness of target detection in inclement weather.

[0265] See also Figure 10 In some embodiments, step 021 (when the current weather is the target weather, fusing the first vehicle environment image and the second vehicle environment image to determine a first fused image) includes:

[0266] 0212: When the current weather is the target weather, pre-process the first vehicle environment image and the second vehicle environment image respectively to determine the first vehicle environment processed image and the second vehicle environment processed image;

[0267] 0213: Fusing the first vehicle environment processed image and the second vehicle environment processed image to determine a first fused image.

[0268] In some embodiments, the fusion module is used to pre-process the first vehicle environment image and the second vehicle environment image respectively when the current weather is the target weather, and determine the first vehicle environment processed image and the second vehicle environment processed image. The fusion module is also used to fuse the first vehicle environment processed image and the second vehicle environment processed image to determine the first fused image.

[0269] In some embodiments, the processor is further configured to pre-process the first vehicle environment image and the second vehicle environment image respectively when the current weather is the target weather, and determine the first vehicle environment processed image and the second vehicle environment processed image. The processor is further configured to fuse the first vehicle environment processed image and the second vehicle environment processed image to determine the first fused image.

[0270] Specifically, preprocessing includes cropping and resolution unification. Based on the calibration results, cropping is used to crop the non-overlapping fields of view of the first and second vehicle environment images, retaining the overlapping areas. Resolution unification is used to adjust the resolution of the visible light and infrared images, unifying them and the matrix to facilitate subsequent pixel-level fusion.

[0271] The first vehicle environment processed image and the second vehicle environment processed image obtained after preprocessing have uniform size and the same field of view, which removes the edge area of ​​camera edge distortion, reduces noise interference, and avoids fusion deviation caused by field of view difference, so that the weight matrix can also be weighted pixel by pixel.

[0272] In this way, when the current weather is the target weather, the first and second vehicle environment images are preprocessed separately to determine the first and second processed vehicle environment images. The first and second processed vehicle environment images are then fused to determine the first fused image. This ensures that the pixels in the first and second vehicle environment images are aligned at the same point in the scene, preventing target misalignment during the fusion process. Furthermore, unified resolution processing facilitates the subsequent pixel-by-pixel weighting matrix calculation, avoiding fusion errors caused by inconsistent scales, reducing feature extraction complexity, and improving target detection efficiency.

[0273] See also Figure 11 In some embodiments, step 0212 (when the current weather is the target weather, pre-processing the first vehicle environment image and the second vehicle environment image to determine the first vehicle environment processed image and the second vehicle environment processed image) includes:

[0274] 02121: When the current weather is the target weather, the first vehicle environment image and the second vehicle environment image are respectively cropped and / or resolution adjusted to determine the first vehicle environment processed image and the second vehicle environment processed image, wherein the field of view corresponding to the first vehicle environment processed image is the same as the field of view corresponding to the second vehicle environment processed image, and the resolution of the first vehicle environment processed image is the same as the resolution of the second vehicle environment processed image.

[0275] In some embodiments, the fusion module is also used to perform cropping and / or resolution adjustment on the first vehicle environment image and the second vehicle environment image, respectively, when the current weather is the target weather, to determine the first vehicle environment processed image and the second vehicle environment processed image, wherein the field of view corresponding to the first vehicle environment processed image is the same as the field of view corresponding to the second vehicle environment processed image, and the resolution of the first vehicle environment processed image is the same as the resolution of the second vehicle environment processed image.

[0276] In certain embodiments, the processor is further used to perform cropping and / or resolution adjustment on the first vehicle environment image and the second vehicle environment image, respectively, when the current weather is the target weather, to determine the first vehicle environment processed image and the second vehicle environment processed image, wherein the field of view corresponding to the first vehicle environment processed image is the same as the field of view corresponding to the second vehicle environment processed image, and the resolution of the first vehicle environment processed image is the same as the resolution of the second vehicle environment processed image.

[0277] Specifically, the cropping process is used to crop the non-overlapping fields of view of the first and second vehicle environment images, retaining the overlapping areas to ensure consistency in the fused regions. For example, the edge of the visible light image not covered by the infrared camera is removed to avoid fusion deviations caused by field of view differences. The cropping process can also remove distorted areas around the camera edges, thereby reducing noise interference in irrelevant areas and providing accurate image support for subsequent target detection.

[0278] The resolution adjustment process is used to adjust the resolution of the first vehicle environment image and the second vehicle environment image to make their resolutions uniform, eliminate feature dislocation caused by resolution differences, reduce the computational complexity during feature extraction, and improve the accuracy of target detection. In an example, a first vehicle environment image with a resolution of 3860*2160 is captured by a forward-looking visible light camera, and a second vehicle environment image with a resolution of 640*480 is captured by an infrared camera. In the preprocessing stage, the similarities in the field of view of the first vehicle environment image and the second vehicle environment image are found, and the non-intersecting areas of the field of view of the images taken by the two cameras are cropped. In addition, the resolution of the first vehicle environment image of 3860*2160 is unified to the target resolution of 640*480, and the pixel value of the position is calculated by mapping the coordinate points.

[0279] During preprocessing, if the resolution of the first and second vehicle environment images is the same after cropping, no resolution adjustment is required. Otherwise, mechanical resolution adjustment is performed after cropping. If the first and second vehicle environment images have the same field of view and are free of noise, resolution adjustment can be performed directly without cropping.

[0280] In this way, when the current weather is the target weather, the first vehicle environment image and the second vehicle environment image are cropped and / or resolution-adjusted, respectively, to determine the first vehicle environment processed image and the second vehicle environment processed image. The field of view corresponding to the first vehicle environment processed image is the same as the field of view corresponding to the second vehicle environment processed image, and the resolution of the first vehicle environment processed image is the same as the resolution of the second vehicle environment processed image. In this way, through cropping and resolution adjustment, image features and specifications can be unified, ensuring the consistency of the fusion area, eliminating errors caused by inconsistent fields of view and resolution, and facilitating subsequent weight matrix calculations. It can also suppress noise and adapt to the weight matrix, thereby improving subsequent fusion accuracy and detection results.

[0281] See also Figure 11 In some embodiments, step 022 (controlling the vehicle according to the first fused image) includes:

[0282] 0221: Determine a second target detection result corresponding to the first fused image based on the first fused image and the third fused image, wherein the third fused image is determined by fusing a third vehicle environment image and a fourth vehicle environment image captured at a previous moment, the third vehicle environment image being captured by the first vehicle forward-looking camera, and the fourth vehicle environment image being captured by the first vehicle infrared camera;

[0283] 0222: Control the vehicle according to the second target detection result.

[0284] In certain embodiments, the vehicle control module is further configured to determine a second target detection result corresponding to the first fused image based on the first and third fused images, wherein the third fused image is determined by fusing a third and fourth vehicle environment images captured at a previous moment, the third vehicle environment image being captured by the first vehicle forward-facing camera, and the fourth vehicle environment image being captured by the first vehicle infrared camera. The vehicle control module is further configured to control the vehicle based on the second target detection result.

[0285] In certain embodiments, the processor is further configured to determine a second target detection result corresponding to the first fused image based on the first and third fused images, wherein the third fused image is determined by fusing a third and fourth vehicle environment images captured at a previous moment, the third vehicle environment image being captured by the first vehicle forward-facing camera, and the fourth vehicle environment image being captured by the first vehicle infrared camera. The processor is further configured to control the vehicle based on the second target detection result.

[0286] Specifically, the third fused image refers to an enhanced image containing visual details and thermal signatures from a previous moment, obtained by fusing the third vehicle environment image with the fourth vehicle environment image. The third vehicle environment image refers to the visible light image of the environment in front of the vehicle at the previous moment. The fourth vehicle environment image refers to the infrared image of the environment in front of the vehicle at the previous moment.

[0287] Because targets are difficult to identify in bad weather conditions, it is necessary to fuse images at multiple times, that is, to perform target detection processing on the first fused image at the current moment and the third fused image at the previous moment, so that the second target detection result corresponding to the first fused image can be obtained more accurately, thereby providing accurate data basis for controlling the vehicle.

[0288] In this way, based on the first and third fused images, a second target detection result corresponding to the first fused image is determined, and the vehicle is controlled based on the second target detection result. In this way, by performing target detection using fused images of the current moment and past moments, the interference of instantaneous noise at the current moment can be reduced, and an accurate detection result at the current moment can be determined, thereby enabling precise vehicle control.

[0289] In some embodiments, the second target detection result includes a detection result for at least one of a vehicle target, a pedestrian target, and a driver target.

[0290] Specifically, identifying vehicle, pedestrian, and driver targets can provide data support for driving safety decisions. For example, if a pedestrian is detected within 50 meters ahead, the system may trigger adaptive cruise control and automatic emergency braking, causing the vehicle to slow down and issue an alert. The relative position of the vehicle target can also be combined to assist in lane keeping or lane change decisions, avoiding collisions with vehicles in adjacent lanes.

[0291] In this way, the second target detection result includes a detection result for at least one of a vehicle target, a pedestrian target, and a driver target. In this way, the detection result of at least one of a vehicle target, a pedestrian target, and a driver target can provide data support for driving safety decisions.

[0292] See also Figure 12 In some embodiments, step 0221 (determining a second target detection result corresponding to the first fused image based on the first fused image and the third fused image) includes:

[0293] 02211: Determine a second target detection result based on the first fused image, the third fused image, and a pre-trained second target detection model.

[0294] In some embodiments, the target detection module is used to determine a second target detection result based on the first fused image, the third fused image, and a pre-trained second target detection model.

[0295] In some embodiments, the processor is configured to determine a second target detection result based on the first fused image, the third fused image, and a pre-trained second target detection model.

[0296] Specifically, the pre-trained second target detection model is used to perform target detection on the fused image and the previous fused image to obtain a second target detection result.

[0297] The first fused image and the third fused image are input into the second target detection model. The second target detection model can extract the multi-scale features of the fused image and then fuse them by building a backbone network. The weather classification dynamically adjusts the detection head parameters, such as lowering the NMS threshold, to reduce missed target detection in bad weather.

[0298] In this way, a second target detection result is determined based on the first and third fused images and the pre-trained second target detection model. This pre-trained second target detection model, through an adaptive weighting mechanism, collaborates with the backbone network and feature fusion to perform target detection on fused images from both the current moment and past moments, reducing the interference of instantaneous noise at the current moment and determining a precise detection result for the current moment, thereby enabling precise vehicle control.

[0299] In some embodiments, the third fused image includes multiple images, and the second object detection model 200 includes a preset processing network 210, a backbone network 220, and a detection network 230;

[0300] The preset processing network 210 is configured to perform feature extraction and splicing processing on the first fused image and the plurality of third fused images to determine a fifth feature map;

[0301] The backbone network 220 is configured to determine a target feature map based on the fifth feature map;

[0302] The detection network 230 is configured to determine a second object detection result based on the object feature map.

[0303] Specifically, the second target detection model 200 refers to the above-mentioned pre-trained second target detection model 200, and the preset processing network 210 is used to perform multi-branch feature extraction processing and splicing processing on the first fused image and multiple third fused images, which can capture more complex semantic information and enhance the detection ability of targets of different sizes. The backbone network 220 is used to extract features of different scales from the fused image and dynamically screen channel and spatial features. The detection network 230 is used to fuse continuous multi-time features to reduce single-frame noise, filter duplicate frames through non-maximum suppression (NMS), and output the final detection results to provide accurate environmental information for intelligent driving decisions.

[0304] In this way, the third fused image includes multiple images, and the second target detection model 200 includes a preset processing network 210, a backbone network 220, and a detection network 230. The preset processing network 210 is configured to perform feature extraction and splicing processing on the first fused image and multiple third fused images to determine a fifth feature map. The backbone network 220 is configured to determine a target feature map based on the fifth feature map. The detection network 230 is configured to determine the second target detection result based on the target feature map. In this way, the second target detection model 200 works through the collaborative operation of multiple modules. The preset processing network 210 fuses multi-time feature data from the current and previous moments, the backbone network 220 performs multi-scale feature extraction on the multi-time feature data, and the detection network 230 finally outputs the final detection result based on the features, achieving accurate target recognition in inclement weather.

[0305] In some embodiments, the preset processing network 210 includes a second feature extraction module 211 , a third feature extraction module 212 , a third feature splicing module 213 , a fourth feature extraction module 214 , a fourth feature splicing module 215 , and a fifth feature extraction module 216 ;

[0306] The second feature extraction module 211 is configured to perform feature extraction processing on the first fused image to determine a sixth feature map corresponding to the first fused image;

[0307] The third feature extraction module 212 is configured to perform feature extraction processing on the third fused image to determine a seventh feature map corresponding to the third fused image;

[0308] The third feature stitching module 213 is configured to stitch the seventh feature map corresponding to each third fused image to determine a third stitching feature map;

[0309] The fourth feature extraction module 214 is configured to perform feature extraction processing on the third spliced ​​feature map to determine an eighth feature map;

[0310] The fourth feature splicing module 215 is configured to splice the sixth feature map and the eighth feature map to determine a fourth spliced ​​feature map;

[0311] The fifth feature extraction module 216 is configured to perform feature extraction processing on the fourth spliced ​​feature map to determine a fifth feature map.

[0312] Specifically, the second feature extraction module 211 is used to perform feature extraction processing on the enhanced image containing visual details and thermal features at the current moment. The sixth feature map refers to the feature map at the current moment.

[0313] The third feature extraction module 212 is used to perform feature extraction processing on the enhanced images at each previous moment, and can obtain multiple seventh feature maps. The seventh feature map refers to the feature map of a single previous moment, and multiple moments have multiple seventh feature maps.

[0314] The third feature splicing module 213 is used to splice multiple feature images of a single moment in time to obtain a third spliced ​​feature image. The third spliced ​​feature image refers to a feature image containing multiple moments in time.

[0315] The fourth feature extraction module 214 is used to perform feature extraction processing on the feature image containing multiple moments, and can obtain feature maps of previous moments.

[0316] The fourth feature splicing module 215 is used to splice the feature map of the previous moment and the feature map of the current moment to obtain a fourth spliced ​​feature map. The fourth spliced ​​feature map refers to a feature image containing the feature images of the current moment and the previous moment.

[0317] The fifth feature extraction module 216 is used to perform feature extraction processing on the feature image containing the current moment and the previous moment to obtain a team feature map. The fifth feature map refers to the feature image fused at multiple moments.

[0318] The following Figure 14 The processing flow of the preset processing network 210 is explained as follows:

[0319] First, the third feature extraction module 212 performs feature extraction processing on the fused images of previous moments respectively, and can obtain feature images of multiple previous moments; at the same time, the second feature extraction module 211 continues to extract features from the fused image of the current moment, and extracts the feature image of the current moment; secondly, the third feature stitching module 213 stitches the feature images of multiple past moments, and fuses all the features of the previous moments to obtain a stitched image with multiple time dimensions; then, the fourth feature extraction module 214 performs feature extraction processing on the stitched image with multiple time dimensions to obtain the full feature image of the previous moment; then, the fourth feature stitching module 215 stitches the full feature image of the previous moment and the feature image of the current moment to obtain the stitched image of the full moment; finally, the fifth feature extraction module 216 processes the stitched image of the full moment to obtain the full feature image of the full moment, which is finally input into the main network for target detection.

[0320] In this way, the preset processing network 210 includes a second feature extraction module 211, a third feature extraction module 212, a third feature stitching module 213, a third feature extraction module 212, a fourth feature extraction module 214, a fourth feature extraction module 214 and a fifth feature extraction module 215; the second feature extraction module 211 is configured to perform feature extraction processing on the first fused image to determine the sixth feature map corresponding to the first fused image; the third feature extraction module 212 is configured to perform feature extraction processing on the third fused image to determine the seventh feature map corresponding to the third fused image; the third feature stitching module 213 is configured to perform stitching processing on the seventh feature map corresponding to each third fused image to determine the third stitching feature map; the fourth feature extraction module 214 is configured to perform feature extraction processing on the third stitching feature map to determine the eighth feature map; the fourth feature stitching module is configured to perform stitching processing on the sixth feature map and the eighth feature map to determine the fourth stitching feature map; the fifth feature extraction module 215 is configured to perform feature extraction processing on the fourth stitching feature map to determine the fifth feature map. In this way, the preset processing network 210 can capture more complex semantic information through multi-branch feature extraction processing, enhancing the detection ability of objects of different sizes. Furthermore, by splicing feature images from different time dimensions, it can integrate multi-scale information such as shallow high-resolution features and deep semantic features, thereby improving feature richness. Through the collaboration between the feature extraction module and the feature splicing module, the essential features of the target can be maximized, improving the efficiency and accuracy of subsequent object detection.

[0321] In some embodiments, the second feature extraction module 211 includes a first feature extraction unit, a first feature fusion unit, and a second feature extraction unit;

[0322] The first feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the first fused image to determine a first processed image;

[0323] The first feature fusion unit is configured to perform dimension expansion, convolution, and splicing on the first processed image to determine a second processed image;

[0324] The third feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the second processed image to determine a sixth feature map.

[0325] Specifically, the first feature extraction unit is used to extract local features of the first fused image and generate a primary feature map.

[0326] The first feature fusion unit is used to learn more complex feature combinations through residual connections and multi-branch structures. The weapon processes information of different scales through different branches and fuses multi-scale features for output.

[0327] The third feature extraction unit is used to integrate the multi-scale features output by the first feature fusion unit layer, smooth the feature distribution, and adjust the number of channels for the input of subsequent modules.

[0328] In one example, the second feature extraction module 211 performs feature extraction on the first fused image through the sequentially connected CBS module, C2f module, and CBS module, that is, the sequentially connected first feature extraction unit, the first feature fusion unit, and the second feature extraction unit, as follows:

[0329] First, the CBS module performs preliminary convolution on the 640*640*3 input image, using a 3*3 convolution kernel and 32 channels. It outputs a 640*640*32 feature map, extracting basic visual features such as edges and textures. Next, the C2f module downsamples the feature map to 320*320*64 using a multi-branch residual structure and feature splitting, enhancing feature representation while reducing the number of parameters. Finally, the CBS module further normalizes and activates the feature map output by C2f, outputting a 320*320*32 feature map to prepare for subsequent feature concatenation.

[0330] Thus, the second feature extraction module 211 includes a first feature extraction unit, a first feature fusion unit, and a second feature extraction unit; the first feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the first fused image to determine a first processed image; the first feature fusion unit is configured to perform dimensional expansion, convolution, and splicing processing on the first processed image to determine a second processed image; and the third feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the second processed image to determine a sixth feature map. In this way, the second feature extraction module 211 can perform multi-scale feature extraction and preliminary processing on the first fused image. The feature extraction unit reduces the impact of illumination fluctuations caused by inclement weather, and the residual connection of the feature fusion unit alleviates the vanishing gradient problem of the deep network, ensuring the stability of feature extraction and improving the detection capability of targets of different sizes.

[0331] In some embodiments, the third feature extraction module 212 includes a third feature extraction unit, a second feature fusion unit, and a fourth feature extraction unit;

[0332] The third feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the third fused image to determine a third processed image;

[0333] The second feature fusion unit is configured to perform dimension expansion, convolution, and splicing on the third processed image to determine a fourth processed image;

[0334] The fourth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the fourth processed image to determine a seventh feature map.

[0335] Specifically, the third feature extraction unit is used to extract local features of the third fused image and generate a primary feature map.

[0336] The second feature fusion unit is used to learn more complex feature combinations through residual connections and multi-branch structures. The weapon processes information of different scales through different branches and fuses multi-scale features for output. The fourth feature extraction unit is used to integrate the multi-scale features output by the second feature fusion unit layer, smooth the feature distribution, and adjust the number of channels for the input of subsequent modules. The processing process of the third feature extraction module 212 is similar to that of the second feature extraction module 211. The difference is that the third feature extraction module 212 is used to process multiple fused images at the previous moment, which will not be repeated here.

[0337] Thus, the third feature extraction module 212 includes a third feature extraction unit, a second feature fusion unit, and a fourth feature extraction unit; the third feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the third fused image to determine a third processed image; the second feature fusion unit is configured to perform dimensional expansion, convolution, and splicing processing on the third processed image to determine a fourth processed image; and the fourth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the fourth processed image to determine a seventh feature map. In this way, the third feature extraction module 212 can perform multi-scale feature extraction and preliminary processing on the third fused image. The feature extraction unit reduces the impact of illumination fluctuations caused by inclement weather, and the residual connection of the feature fusion unit alleviates the gradient vanishing problem of the deep network, ensuring the stability of feature extraction and providing multiple features at previous moments for subsequent splicing processing.

[0338] In some embodiments, the fourth feature extraction module 214 includes a fifth feature extraction unit, a third feature fusion unit, a sixth feature extraction unit, and a fourth feature fusion unit;

[0339] The fifth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the third spliced ​​feature map to determine a fifth processed image;

[0340] The third feature fusion unit is configured to perform dimension expansion, convolution, and splicing on the fifth processed image to determine a sixth processed image;

[0341] The sixth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation on the sixth processed image to determine a seventh processed image;

[0342] The fourth feature fusion unit is configured to perform dimension expansion, convolution and splicing on the seventh processed image to determine an eighth feature map.

[0343] Specifically, the fifth feature extraction unit is used to extract the third splicing feature map, that is, the local features of all previous moments, to generate a primary feature map. The third feature fusion unit is used to learn more complex feature combinations through residual connections and multi-branch structures, and to process different scale information through different branches, and to fuse multi-scale features for output; the sixth feature extraction unit is used to perform normalization and activation again to prepare for subsequent splicing questions. The fourth feature fusion unit is used to learn more complex feature combinations through residual connections and multi-branch structures, and to process different scale information through different branches, and to fuse multi-scale features for output. The fourth feature extraction module 214 is similar to the processing process of the above-mentioned second feature extraction module 211, except that the fourth feature extraction module 214 is used to process the spliced ​​image of the previous moment, which will not be repeated here.

[0344] Thus, the fourth feature extraction module 214 includes a fifth feature extraction unit, a third feature fusion unit, a sixth feature extraction unit, and a fourth feature fusion unit; the fifth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the third spliced ​​feature map to determine a fifth processed image; the third feature fusion unit is configured to perform dimensional expansion, convolution, and splicing processing on the fifth processed image to determine a sixth processed image; the sixth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the sixth processed image to determine a seventh processed image; and the fourth feature fusion unit is configured to perform dimensional expansion, convolution, and splicing processing on the seventh processed image to determine an eighth feature map. In this way, the fourth feature extraction module 214 can perform multi-scale feature extraction and preliminary processing on the third spliced ​​image. The feature fusion unit reduces the impact of illumination fluctuations caused by bad weather. The residual connection of the feature splicing unit alleviates the gradient vanishing problem of the deep network, ensuring the stability of feature extraction and providing multiple features at previous moments for subsequent splicing processing.

[0345] In some embodiments, the fifth feature extraction module 215 includes a seventh feature extraction unit, a fifth feature fusion unit, and an eighth feature extraction unit;

[0346] The seventh feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the fourth spliced ​​feature map to determine an eighth processed image;

[0347] The fifth feature fusion unit is configured to perform dimension expansion, convolution, and splicing on the eighth processed image to determine a ninth processed image;

[0348] The eighth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the ninth processed image to determine a fifth feature map.

[0349] Specifically, the seventh feature extraction unit is used to extract local features of the fourth spliced ​​feature map and generate a primary feature map. The fifth feature fusion unit is used to learn more complex feature combinations through residual connections and multi-branch structures. The weapon processes information of different scales through different branches and fuses multi-scale features for output. The eighth feature extraction unit is used to integrate the multi-scale features output by the fifth feature fusion unit layer, smooth the feature distribution, and adjust the number of channels for the input of subsequent modules. The processing process of the fifth feature extraction module 215 is similar to that of the above-mentioned second feature extraction module 211, and will not be repeated here.

[0350] Thus, the fifth feature extraction module 215 includes a seventh feature extraction unit, a fifth feature fusion unit, and an eighth feature extraction unit; the seventh feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the fourth spliced ​​feature map to determine an eighth processed image; the fifth feature fusion unit is configured to perform dimensional expansion, convolution, and splicing processing on the eighth processed image to determine a ninth processed image; and the eighth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the ninth processed image to determine the fifth feature map. In this way, the fifth feature extraction module 215 can perform multi-scale feature extraction and preliminary processing on the fourth spliced ​​feature map. The feature fusion unit reduces the impact of illumination fluctuations caused by bad weather. The residual connection of the feature splicing unit alleviates the gradient vanishing problem of the deep network, ensuring the stability of feature extraction and providing multiple features of the previous time for subsequent splicing processing.

[0351] In some embodiments, the backbone network 220 includes a sixth feature fusion unit 221 , a ninth feature extraction unit 222 , a seventh feature fusion unit 223 , a tenth feature extraction unit 224 , an eighth feature fusion unit 225 , and a ninth feature fusion unit 226 ;

[0352] The sixth feature fusion unit 221 is configured to perform dimension expansion, convolution, and splicing processing on the fifth feature map according to the attention data corresponding to the fifth feature map to determine a tenth processed image;

[0353] The ninth feature extraction unit 222 is configured to perform convolution, batch normalization, and nonlinear activation processing on the tenth processed image to determine a twelfth processed image;

[0354] The seventh feature fusion unit 223 is configured to perform dimension expansion, convolution, and splicing on the eleventh processed image to determine a twelfth processed image;

[0355] The tenth feature extraction unit 224 is configured to perform convolution, batch normalization, and nonlinear activation processing on the twelfth processed image to determine a thirteenth processed image;

[0356] The eighth feature fusion unit 225 is configured to perform dimension expansion, convolution, and splicing on the thirteenth processed image to determine a fourteenth processed image;

[0357] The ninth feature fusion unit 226 is configured to perform convolution, pooling and splicing on the fourteenth processed image to determine a target feature map.

[0358] Specifically, attention data refers to the weight matrix generated by a variety of attention mechanisms, which is used to dynamically adjust the importance of different channels or spatial positions in the feature map. Its core function is to suppress noise and enhance target features in bad weather. The sixth feature fusion unit 221, the seventh feature fusion unit 223, and the eighth feature fusion unit 225 are all used to achieve lightweight feature enhancement and dimensionality reduction through multi-branch residual structure and feature diversion strategy. For example, the 320*320*32 feature map is processed into 320*320*64, which improves the feature expression ability while reducing the number of parameters. The ninth feature extraction unit 222 and the tenth feature extraction unit 224 are both used to perform preliminary convolution processing on the input image and extract basic visual features. The ninth feature fusion unit 226 is used to capture contextual information of different scales through parallel large-core pooling (5*5, 9*9, 13*13), fix the output dimension to adapt to subsequent detection heads, and improve adaptability to targets of different sizes.

[0359] The following Figure 15 The backbone network 220 is explained as follows:

[0360] First, the C2f module, or the sixth feature fusion unit 221, fuses shallow features (low-level edges) with deep features (high-level semantics), such as "wheel" and "face," to generate a more abstract feature map. The CBS module, or the tenth feature extraction unit 224, then performs preliminary feature extraction on the image, using, for example, a 3x3 convolution to extract edges and textures, such as vehicle outlines and pedestrian body edges. The image is then processed again through the C2f module, alternating between feature extraction and fusion. Finally, the SPPF module, or the ninth feature fusion unit 222, pools and fuses feature maps of different scales, such as 5x5, 9x9, and 13x13 windows, to generate feature maps containing multi-scale context, such as 20x20, 40x40, and 80x80 resolutions. The 80x80 feature map retains details such as the texture of pedestrian clothing and the shape of traffic signs, and is used for detecting small objects. The 20x20 feature map contains semantic information such as the overall vehicle outline and pedestrian posture, and is used for detecting large objects.

[0361] Thus, the backbone network 220 includes a sixth feature fusion unit 221, a ninth feature extraction unit 222, a seventh feature fusion unit 223, a tenth feature extraction unit 224, an eighth feature fusion unit 225, and a ninth feature fusion unit 226; the sixth feature fusion unit 221 is configured to perform dimension expansion, convolution, and splicing processing on the fifth feature map according to the attention data corresponding to the fifth feature map to determine the tenth processed image; the ninth feature extraction unit 222 is configured to perform convolution, batch normalization, and nonlinear activation processing on the tenth processed image to determine the twelfth processed image; The seventh feature fusion unit 223 is configured to perform dimensional expansion, convolution, and splicing on the eleventh processed image to determine the twelfth processed image; the tenth feature extraction unit 224 is configured to perform convolution, batch normalization, and nonlinear activation on the twelfth processed image to determine the thirteenth processed image; the eighth feature fusion unit 225 is configured to perform dimensional expansion, convolution, and splicing on the thirteenth processed image to determine the fourteenth processed image; and the ninth feature fusion unit 226 is configured to perform convolution, pooling, and splicing on the fourteenth processed image to determine the target feature map. In this way, the backbone network 220 dynamically adapts to environmental changes through a combination of attention mechanisms. By alternately cascading feature extraction and feature fusion units, multi-scale features can be extracted from the weighted visible light and infrared fusion image, further enhancing target representation. By outputting multi-scale feature maps to the detection head, the classification and localization of targets of different sizes are supported, thereby improving the detection capability of targets of different sizes and achieving accurate detection of targets such as vehicles and pedestrians.

[0362] In some embodiments, the detection network 230 is configured to determine a second target detection result based on the target feature map, the attention data corresponding to the tenth processed image, the attention data corresponding to the thirteenth processed image, and the attention data corresponding to the target feature map.

[0363] Specifically, the attention data corresponding to the tenth processed image refers to channel attention, which is used to analyze the importance of each channel in the feature map, suppress weather noise-related channels, and enhance the target feature channels. The attention data corresponding to the thirteenth processed image refers to spatial attention, which is used to focus on the spatial area where the target is located and ignore background interference. The attention data corresponding to the target feature map refers to frequency channel attention, which is used to combine frequency domain analysis to distinguish the frequency distribution of weather noise and target features.

[0364] The attention data is input into an object detection head, such as a YOLOv8 model, to determine a final object detection result, i.e., a second object detection result. The detection network 230 combines the attention data with the object feature map to perform object detection, thereby reducing category confusion and improving object detection accuracy, thereby providing accurate data for vehicle control.

[0365] Thus, the detection network 230 is configured to determine a second target detection result based on the target feature map, the attention data corresponding to the tenth processed image, the attention data corresponding to the thirteenth processed image, and the attention data corresponding to the target feature map. In this way, by integrating the channel, spatial, and frequency channel attention data, feature weights can be dynamically adjusted to achieve focus on key information and noise suppression, thereby outputting accurate detection results.

[0366] The following Figure 16 The overall process of the vehicle control method according to the embodiment of the present application is explained as follows:

[0367] The surround view image refers to the first driving environment image information. The forward-view camera image and the infrared camera image refer to the second driving image information. The classification algorithm refers to the weather recognition model 100 in the embodiment of the present application. First, after pre-processing the surround view image, the classification algorithm can be used to identify the severe weather conditions in the current vehicle driving environment, that is, the probabilities of severe weather and non-severe weather. Then, the start-stop module performs a weighted calculation on the severe weather probability to determine whether to start or exit the severe weather mode. The fusion algorithm then performs a weighted fusion of the environmental images, including the first driving environment image and the second driving image information, to obtain a fused image with multiple features to facilitate subsequent target detection processing.

[0368] Finally, target detection processing is performed by fusing multiple feature images based on the detection algorithm, i.e., the second target detection model 200, thereby outputting a detection result and improving the accuracy of target detection. This application also provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program implements the steps of the vehicle control method described above.

[0369] It is understood that a computer program includes computer program code. The computer program code may be in source code form, object code form, executable file, or some intermediate form. Computer-readable storage media may include any entity or device capable of carrying computer program code, recording media, USB flash drives, removable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution media.

[0370] In the description of this specification, the descriptions with reference to the terms "particularly", "further", "particularly", "understandably", etc. are intended to mean that the specific features, structures, materials or characteristics described in conjunction with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms are not intended to refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.

[0371] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code that includes one or more executable requests for implementing a specific logical function or step of a process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed in a different order than shown or discussed, including performing functions in a substantially simultaneous manner or in a reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.

[0372] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A vehicle control method, characterized in that: include: determining a recognition result of the current weather according to the first driving environment image information; When the current weather is target weather, the vehicle is controlled according to the second driving environment image information, wherein the visibility corresponding to the target weather is less than or equal to a preset threshold.

2. The method according to claim 1, characterized in that The first driving environment image information includes a plurality of surround view images captured by a plurality of vehicle surround view cameras with different shooting directions.

3. The method according to claim 2, characterized in that The determining of the recognition result of the current weather according to the first driving environment image information includes: performing stitching processing on the plurality of surround view images to determine a stitched image; The recognition result is determined according to the spliced ​​image and the plurality of surround-view images.

4. The method according to claim 3, characterized in that Determining the recognition result according to the stitched image and the plurality of surround-view images includes: The recognition result is determined based on the spliced ​​image, the plurality of surround view images, and a pre-trained weather recognition model.

5. The method according to claim 4, characterized in that The weather recognition model includes a first feature extraction module, a first feature splicing module and a weather recognition module connected in sequence; The first feature extraction module is configured to perform feature extraction processing on the stitched image and the surround view image respectively to determine a first feature map of the stitched image and a second feature map of the surround view image; The first feature splicing module is configured to splice the first feature map and a plurality of the second feature maps to determine a first spliced ​​feature map; The weather recognition module is configured to determine the recognition result according to the first splicing feature map.

6. The method according to claim 5, characterized in that The first feature extraction module includes a first convolution unit and a first pooling unit; The first convolution unit is configured to perform feature extraction processing on the stitched image and the surround view image respectively to determine a third feature map of the stitched image and a fourth feature map of the surround view image; The first pooling unit is configured to perform pooling processing on the third feature map of the stitched image and the fourth feature map of the surround view image, respectively, to determine the first feature map of the stitched image and the second feature map of the surround view image.

7. The method according to claim 5, characterized in that The first feature splicing module includes a splicing unit and a second convolution unit; The splicing unit is configured to splice the first feature map and a plurality of the second feature maps to determine a second spliced ​​feature map; The second convolution unit is configured to perform convolution processing on the second splicing feature map to determine the first splicing feature map.

8. The method according to claim 5, characterized in that The weather recognition module includes a third convolution unit, a fully connected layer unit, and a classification unit connected in sequence; The third convolution unit is configured to perform convolution processing on the first splicing feature map to determine a third splicing feature map; The fully connected layer unit is configured to determine feature data according to the third spliced ​​feature map; The classification unit is configured to determine the recognition result according to the feature data.

9. The method according to any one of claims 1 to 8, characterized in that The recognition result includes a first probability and a second probability, the first probability is used to indicate the probability that the current weather is the target weather, and the second probability is used to indicate the probability that the current weather is not the target weather.

10. The method according to claim 9, characterized in that The method further comprises: Determining a first probability weighted result and a second probability weighted result at a current moment based on a first preset weight corresponding to the first probability, a second preset weight corresponding to the second probability, and the first probability and the second probability respectively corresponding to each piece of first driving environment image information acquired within a preset time period; When the first probability weighted results and the second probability weighted results at multiple consecutive moments both meet a preset condition, the current weather is identified as the target weather.

11. The method according to claim 1, wherein The second driving environment image information includes a first vehicle environment image and a second vehicle environment image taken at the current moment. The first vehicle environment image is taken by the first vehicle front-view camera, and the second vehicle environment image is taken by the first vehicle infrared camera. The shooting direction of the first vehicle front-view camera is the same as the shooting direction of the first vehicle infrared camera.

12. The method according to claim 11, characterized in that When the current weather is the target weather, controlling the vehicle according to the second driving environment image information includes: When the current weather is the target weather, fusing the first vehicle environment image and the second vehicle environment image to determine a first fused image; The vehicle is controlled according to the first fused image.

13. The method according to claim 12, characterized in that When the current weather is the target weather, fusing the first vehicle environment image and the second vehicle environment image to determine a first fused image includes: When the current weather is the target weather, the first vehicle environment image and the second vehicle environment image are fused according to the first pixel weight corresponding to the first vehicle environment image and the second pixel weight matrix corresponding to the second vehicle environment image to determine the first fused image.

14. The method according to claim 13, characterized in that The step of obtaining the first pixel weight matrix and the second pixel weight matrix includes: performing a fusion process on a plurality of first vehicle environment image samples and a plurality of second vehicle environment image samples according to a predetermined second pixel weight and a third pixel weight to determine a plurality of second fused images, wherein the first vehicle environment image samples are captured by a second vehicle front-view camera, the second vehicle environment image samples are captured by a second vehicle infrared camera, and the shooting direction of the second vehicle front-view camera is the same as the shooting direction of the second vehicle infrared camera; determining a first object detection result based on a first object detection model and a portion of the second fused image, wherein the first object detection model is trained based on at least another portion of the second fused image; According to the first target detection result and the pre-calibrated detection target in the second fused image, the second pixel weight and the third pixel weight are updated to determine the first pixel weight matrix and the second pixel weight matrix.

15. The method according to claim 12, characterized in that When the current weather is the target weather, fusing the first vehicle environment image and the second vehicle environment image to determine a first fused image includes: When the current weather is the target weather, pre-processing the first vehicle environment image and the second vehicle environment image respectively to determine a first vehicle environment processed image and a second vehicle environment processed image; The fusion process is performed on the first vehicle environment processed image and the second vehicle environment processed image to determine the first fused image.

16. The method according to claim 15, characterized in that The preprocessing of the first vehicle environment image and the second vehicle environment image to determine the first vehicle environment processed image and the second vehicle environment processed image when the current weather is the target weather includes: When the current weather is the target weather, the first vehicle environment image and the second vehicle environment image are respectively cropped and / or resolution adjusted to determine the first vehicle environment processed image and the second vehicle environment processed image, wherein the field of view corresponding to the first vehicle environment processed image is the same as the field of view corresponding to the second vehicle environment processed image, and the resolution of the first vehicle environment processed image is the same as the resolution of the second vehicle environment processed image.

17. The method according to claim 12, wherein: The controlling the vehicle according to the first fused image includes: determining, based on the first fused image and the third fused image, a second target detection result corresponding to the first fused image, wherein the third fused image is determined by performing the fusion processing on a third vehicle environment image and a fourth vehicle environment image captured at a previous moment, the third vehicle environment image being captured by the first vehicle front-view camera, and the fourth vehicle environment image being captured by the first vehicle infrared camera; The vehicle is controlled according to the second target detection result.

18. The method according to claim 17, characterized in that The second target detection result includes a detection result for at least one of a vehicle target, a pedestrian target, and a driver target.

19. The method according to claim 17, wherein The determining, based on the first fused image and the third fused image, a second target detection result corresponding to the first fused image includes: The second target detection result is determined according to the first fused image, the third fused image, and a pre-trained second target detection model.

20. The method according to claim 19, characterized in that The third fused image includes multiple images, and the second target detection model includes a preset processing network, a backbone network and a detection network; The preset processing network is configured to perform feature extraction and splicing processing on the first fused image and the plurality of third fused images to determine a fifth feature map; The backbone network is configured to determine a target feature map based on the fifth feature map; The detection network is configured to determine the second target detection result based on the target feature map.

21. The method according to claim 20, characterized in that The preset processing network includes a second feature extraction module, a third feature extraction module, a third feature splicing module, a third feature extraction module, a fourth feature splicing module, a fourth feature extraction module and a fifth feature extraction module; The second feature extraction module is configured to perform feature extraction processing on the first fused image to determine a sixth feature map corresponding to the first fused image; The third feature extraction module is configured to perform feature extraction processing on the third fused image to determine a seventh feature map corresponding to the third fused image; The third feature stitching module is configured to stitch the seventh feature map corresponding to each of the third fused images to determine a third stitching feature map; The fourth feature extraction module is configured to perform feature extraction processing on the third spliced ​​feature map to determine an eighth feature map; The fourth feature splicing module is configured to perform splicing processing on the sixth feature map and the eighth feature map to determine a fourth splicing feature map; The fifth feature extraction module is configured to perform feature extraction processing on the fourth splicing feature map to determine the fifth feature map.

22. The method according to claim 21, characterized in that The second feature extraction module includes a first feature extraction unit, a first feature fusion unit and a second feature extraction unit; The first feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the first fused image to determine a first processed image; The first feature fusion unit is configured to perform dimension expansion, convolution, and splicing on the first processed image to determine a second processed image; The third feature extraction unit is configured to perform convolution, batch normalization and nonlinear activation processing on the second processed image to determine the sixth feature map.

23. The method according to claim 21, characterized in that The third feature extraction module includes a third feature extraction unit, a second feature fusion unit and a fourth feature extraction unit; The third feature extraction unit is configured to perform convolution, batch normalization and nonlinear activation processing on the third fused image to determine a third processed image; The second feature fusion unit is configured to perform dimension expansion, convolution and splicing on the third processed image to determine a fourth processed image; The fourth feature extraction unit is configured to perform convolution, batch normalization and nonlinear activation processing on the fourth processed image to determine the seventh feature map.

24. The method according to claim 21, characterized in that The fourth feature extraction module includes a fifth feature extraction unit, a third feature fusion unit, a sixth feature extraction unit and a fourth feature fusion unit; The fifth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the third spliced ​​feature map to determine a fifth processed image; The third feature fusion unit is configured to perform dimension expansion, convolution and splicing on the fifth processed image to determine a sixth processed image; The sixth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the sixth processed image to determine a seventh processed image; The fourth feature fusion unit is configured to perform dimension expansion, convolution and splicing on the seventh processed image to determine the eighth feature map.

25. The method according to claim 21, characterized in that The fifth feature extraction module includes a seventh feature extraction unit, a fifth feature fusion unit and an eighth feature extraction unit; The seventh feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the fourth spliced ​​feature map to determine an eighth processed image; The fifth feature fusion unit is configured to perform dimension expansion, convolution, and splicing on the eighth processed image to determine a ninth processed image; The eighth feature extraction unit is configured to perform convolution, batch normalization and nonlinear activation processing on the ninth processed image to determine the fifth feature map.

26. The method according to claim 20, wherein The backbone network includes a sixth feature fusion unit, a ninth feature extraction unit, a seventh feature fusion unit, a tenth feature extraction unit, an eighth feature fusion unit and a ninth feature fusion unit; The sixth feature fusion unit is configured to perform dimension expansion, convolution, and splicing processing on the fifth feature map according to the attention data corresponding to the fifth feature map to determine a tenth processed image; The ninth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the tenth processed image to determine a twelfth processed image; The seventh feature fusion unit is configured to perform dimension expansion, convolution and splicing on the eleventh processed image to determine a twelfth processed image; The tenth feature extraction unit is configured to perform convolution, batch normalization, and nonlinear activation processing on the twelfth processed image to determine a thirteenth processed image; The eighth feature fusion unit is configured to perform dimension expansion, convolution, and splicing on the thirteenth processed image to determine a fourteenth processed image; The ninth feature fusion unit is configured to perform convolution, pooling and splicing on the fourteenth processed image to determine the target feature map.

27. The method according to claim 26, characterized in that The detection network is configured to determine the second target detection result based on the target feature map, the attention data corresponding to the tenth processed image, the attention data corresponding to the thirteenth processed image, and the attention data corresponding to the target feature map.

28. An electronic device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the method according to any one of claims 1 to 27 is implemented.

29. A vehicle, characterized in that: An electronic device comprising the electronic device described in claim 28.

30. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by one or more processors, the method according to any one of claims 1 to 27 is implemented.

31. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 27 is implemented.