CTI-YOLO-based vehicle detection method under complex weather conditions
By improving the YOLOv8 model, integrating MSRAM, GLSA and TADHead to build a CTI-YOLO detection model, the problem of traditional vehicle detection methods being poor in complex weather was solved, and efficient and robust vehicle detection results were achieved.
Patent Information
- Application Number
- CN202510252781.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-17
AI Technical Summary
Traditional vehicle detection methods perform poorly in complex weather conditions, resulting in reduced image quality, blurred targets, reduced contrast and increased noise, which cannot meet the real-time requirements.
Using a vehicle detection method based on CTI-YOLO, the CTI-YOLO detection model is constructed by improving the YOLOv8 model and integrating MSRAM, GLSA and TADHead to enhance feature extraction and detection capabilities and adapt to complex weather conditions.
It improves the robustness and accuracy of vehicle detection, reduces the computational complexity, meets the real-time requirements, and can effectively identify vehicles under complex weather conditions.
Smart Images

Figure CN120164175A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target detection in computer vision, and in particular to a vehicle detection method under complex weather conditions based on CTI-YOLO. Background Art
[0002] Weather conditions in the real world are diverse and unpredictable. In bad weather such as rain, fog, snow, and dust, image quality often degrades significantly, resulting in blurred targets, reduced contrast, and increased noise. These problems make traditional vehicle detection methods perform poorly or even fail completely in complex weather conditions. For example, autonomous vehicles may not be able to accurately identify surrounding vehicles and pedestrians in rainy or foggy weather, causing serious safety hazards. Therefore, it is of great practical significance to study how to achieve robust vehicle detection under complex weather conditions.
[0003] Although existing solutions have alleviated the detection problem in complex weather conditions to a certain extent, they still have many limitations. For example, some methods use image enhancement techniques (such as defogging and deraining) to improve image quality, but these methods often have high computational complexity and are difficult to meet real-time requirements. Therefore, developing a vehicle detection method that can adapt to complex weather conditions and has efficient computing performance has become a hot topic in current research.
[0004] Vehicle detection in complex weather conditions is a very challenging research direction, and its importance is self-evident. By combining the latest deep learning technology, improving the performance of the YOLO model in complex weather conditions can not only solve key problems in practical applications, but also provide new ideas and methods for the development of technology in related fields. This research direction is both of practical significance and full of scientific research innovation potential. Summary of the invention
[0005] The present invention provides a vehicle detection method under complex weather conditions based on CTI-YOLO, aiming to solve the problem that the traditional vehicle detection method in the prior art performs poorly under complex weather conditions. The present invention adopts the following technical solutions:
[0006] A vehicle detection method under complex weather conditions based on CTI-YOLO, comprising the following steps:
[0007] Step 1: Collect original images of road vehicles under complex weather conditions and perform preprocessing to obtain a preprocessed complex weather vehicle image dataset;
[0008] Step 2: Improve the YOLOv8 model and integrate MSRAM, GLSA and TADHead to build the CTI-YOLO detection model;
[0009] Step 3: Train the model to obtain an efficient training model, and use the trained CTI-YOLO detection model to detect the preprocessed road vehicle image dataset.
[0010] Step 4: Implement deployment and visualize the actual application effect of the method.
[0011] Furthermore, in step 1, collecting original road vehicle images and preprocessing them includes the following steps:
[0012] Step 1.1: Collect traffic road vehicle images under different lighting conditions, weather conditions and time periods from actual traffic scenes; combine public road traffic datasets, use simulation tools and complex weather generation algorithms to simulate and generate vehicle images similar to actual traffic scenes to supplement the sample size;
[0013] Step 1.2: Label the specific area of the vehicle and generate a file containing the vehicle type and location; classify the image by vehicle type.
[0014] Furthermore, in step 2, creating a vehicle detection model CTI-YOLO under complex weather conditions includes the following steps:
[0015] Step 2.1: The multi-scale residual feature aggregation module (MSRAM) replaces the C2f module in the network structure. The multi-scale residual feature aggregation module uses a variety of convolution kernel sizes for feature extraction and introduces some residual connections, which can retain input features and enhance the utilization of original information;
[0016] Step 2.2: The global to local spatial aggregation (GLSA) module is introduced, which has the ability to aggregate and represent global and local spatial features, which is beneficial for locating large and small objects respectively;
[0017] Step 2.3: A feature-sharing dynamic detection head (TADHead) is proposed, which achieves a lightweight design through shared convolution. At the same time, deformable convolution (DCNv2) is introduced to solve the problem that shared parameters may lead to insufficient model expression ability.
[0018] Furthermore, the step 2.1 includes the following steps:
[0019] Multi-scale features are captured through convolution kernels of different sizes; the features are split into different parts in the channel dimension and grouped convolution is performed; after the multi-dimensional feature extraction is completed, the module aggregates the features on different paths through feature splicing and fusion operations; by introducing residual connections, the gradient flow is further enhanced and the global features of the input are retained. The specific formula is as follows:
[0020] f=Conv k*k (x)
[0021] f 1,1 ,f 1,2 =Chunk(f1,dim=1,groups=2)
[0022] f 2,1 =Conv 5*5 (f 1,1 ,groups=1)
[0023] f out =Cat([f3,f 2,2 ,f 1,2 ], dim=1)
[0024] y=Conv 1*1 (f out )+x
[0025] Among them, Conv k*k represents a convolution operation with a kernel size of k*k; Cat([·]) represents a concatenation operation in the channel dimension; f3, f 2,2 , f 1,2 Respectively make feature representations of different scales; x represents the input feature information.
[0026] Furthermore, the step 2.2 includes the following steps:
[0027] The global information of the entire image is extracted through some global perception mechanisms; next, different local areas of the image are perceived and analyzed by using convolution, dilated convolution or other local perception mechanisms to capture finer-grained detail information; after the global and local information perception, GLSA aggregates the global information with the local information to better capture the boundaries, shapes and positions of the detected objects. The specific formula is as follows:
[0028]
[0029] in is the input of global information extraction, Yes, global information extraction output; is the input of local information perception, is the output of GSA.
[0030] Furthermore, in step 2.3, the detection head extracts multi-scale features from the feature extractor to obtain joint features, and generates the offset and mask required by the deformable convolution (DCNv2), and further outputs the multi-scale detection frame. In order to solve the problem of inconsistent scales of the detected targets caused by the use of a shared convolution structure, a scaling layer is used to scale the features. The deformable convolution formula is as follows:
[0031]
[0032] Among them, let w k and p k denote the weight of the kth position and the predefined offset Δp respectively. k and Δm k are the learnable offset and modulation scalar at the kth position, respectively.
[0033] Furthermore, in step 3, the data set is trained to obtain a trained efficient detection model of the improved method; the original image of the traffic scene is input into the trained model to obtain the test result of the model, and a predicted image with a target prediction box and confidence is output.
[0034] Furthermore, in step 4, the method is deployed on a vehicle monitoring system, the vehicle detection results are visualized, and the vehicle detection method is applied to actual traffic scenarios.
[0035] Compared with the prior art, the present invention has the following advantages and beneficial effects: the multi-scale residual feature aggregation module uses a variety of convolution kernel sizes for feature extraction and introduces partial residual connections, which can retain input features and enhance the utilization of original information; the global to local space aggregation module has the ability to aggregate and represent global and local space features, which is conducive to locating large targets and small targets respectively, and reducing the computational complexity of traffic scene image processing; a feature sharing dynamic detection head is proposed, which realizes lightweight design through shared convolution, and introduces deformable convolution to solve the problem that shared parameters may lead to insufficient model expression ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art are briefly introduced below.
[0037] Figure 1 It is a schematic diagram of the process of step 1 to step 4 in the present invention.
[0038] Figure 2 It is the overall network framework of the vehicle detection method of the present invention.
[0039] Figure 3 This is the MSRAM network structure in step 2 of the present invention.
[0040] Figure 4 This is the GLSA network structure in step 2 of the present invention.
[0041] Figure 5 It is the TADHead network structure in step 2 of the present invention.
[0042] Figure 6It is an autonomous driving vehicle deployment experimental platform used in the present invention. DETAILED DESCRIPTION
[0043] The present invention is described in detail below in conjunction with specific embodiments.
[0044] like Figure 1 As shown, a vehicle detection method under complex weather conditions based on CTI-YOLO includes the following steps:
[0045] Step 1: Collect original images of road vehicles under complex weather conditions and perform preprocessing to obtain a preprocessed complex weather vehicle image dataset;
[0046] Step 2: Improve the YOLOv8 model and integrate MSRAM, GLSA and TADHead to build the CTI-YOLO detection model;
[0047] Step 3: Train the model to obtain an efficient training model, and use the trained CTI-YOLO detection model to detect the preprocessed road vehicle image dataset.
[0048] Step 4: Implement deployment and visualize the actual application effect of the method.
[0049] Furthermore, in step 1, collecting original road vehicle images and preprocessing them includes the following steps:
[0050] Step 1.1: Collect traffic road vehicle images under different lighting conditions, weather conditions and time periods from actual traffic scenes; combine public road traffic datasets, use simulation tools and complex weather generation algorithms to simulate and generate vehicle images similar to actual traffic scenes to supplement the sample size;
[0051] Step 1.2: Label the specific area of the vehicle and generate a file containing the vehicle type and location; classify the image by vehicle type.
[0052] Further, in step 2, if Figure 2 As shown in the figure, creating a vehicle detection model CTI-YOLO under complex weather conditions includes the following steps:
[0053] Step 2.1: The multi-scale residual feature aggregation module (MSRAM) replaces the C2f module in the network structure. The multi-scale residual feature aggregation module uses a variety of convolution kernel sizes for feature extraction and introduces some residual connections, which can retain input features and enhance the utilization of original information;
[0054] Step 2.2: The global to local spatial aggregation (GLSA) module is introduced, which has the ability to aggregate and represent global and local spatial features, which is beneficial for locating large and small objects respectively;
[0055] Step 2.3: A feature-sharing dynamic detection head (TADHead) is proposed, which achieves a lightweight design through shared convolution. At the same time, deformable convolution (DCNv2) is introduced to solve the problem that shared parameters may lead to insufficient model expression ability.
[0056] Further, in step 2.1, Figure 3 As shown, the following steps are included:
[0057] Multi-scale features are captured through convolution kernels of different sizes; the features are split into different parts in the channel dimension and grouped convolution is performed; after the multi-dimensional feature extraction is completed, the module aggregates the features on different paths through feature splicing and fusion operations; by introducing residual connections, the gradient flow is further enhanced and the global features of the input are retained. The specific formula is as follows:
[0058] f=Conv k*k (x)
[0059] f 1,1 ,f 1,2 =Chunk(f1,dim=1,groups=2)
[0060] f 2,1 =Conv 5*5 (f 1,1 ,groups=1)
[0061] f out =Cat([f3,f 2,2 ,f 1,2 ], dim=1)
[0062] y=Conv 1*1 (f out )+x
[0063] Among them, Conv k*k represents a convolution operation with a kernel size of k*k; Cat([·]) represents a concatenation operation in the channel dimension; f3, f 2,2 , f 1,2 Respectively make feature representations of different scales; x represents the input feature information.
[0064] Further, in step 2.2, if Figure 4 As shown, the following steps are included:
[0065] The global information of the entire image is extracted through some global perception mechanisms; next, different local areas of the image are perceived and analyzed by using convolution, dilated convolution or other local perception mechanisms to capture finer-grained detail information; after the global and local information perception, GLSA aggregates the global information with the local information to better capture the boundaries, shapes and positions of the detected objects. The specific formula is as follows:
[0066]
[0067] in is the input of global information extraction, Yes, global information extraction output; is the input of local information perception, is the output of GSA.
[0068] Further, in step 2.3, if Figure 5 As shown in the figure, the detection head extracts multi-scale features from the feature extractor, obtains joint features, generates the offset and mask required by the deformable convolution (DCNv2), and further outputs the multi-scale detection frame. In order to solve the problem of inconsistent scales of the detected targets caused by the use of a shared convolution structure, a scaling layer is used to scale the features. The deformable convolution formula is as follows:
[0069]
[0070] Among them, let w k and p k denote the weight of the kth position and the predefined offset Δp respectively. k and Δm k are the learnable offset and modulation scalar at the kth position, respectively.
[0071] Furthermore, in step 3, the data set is trained to obtain a trained efficient detection model of the improved method; the original image of the traffic scene is input into the trained model to obtain the test result of the model, and a predicted image with a target prediction box and confidence is output.
[0072] Further, in step 4, if Figure 6 As shown, the method is deployed on a vehicle monitoring system, the vehicle detection results are visualized, and the vehicle detection method is applied to actual traffic scenarios.
[0073] The examples of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above examples. Various changes can be made within the knowledge scope of ordinary technicians in the field without departing from the purpose of the present invention, which should also be regarded as the protection scope of the present invention.
Claims
1. A vehicle detection method under complex weather conditions based on CTI-YOLO, characterized in that: The following steps are involved: Step 1: Collect original images of road vehicles under complex weather conditions and perform preprocessing to obtain a preprocessed complex weather vehicle image dataset; Step 2: Improve the YOLOv8 model and integrate MSRAM, GLSA and TADHead to build the CTI-YOLO detection model; Step 3: Train the model to obtain an efficient training model, and use the trained CTI-YOLO detection model to detect the preprocessed road vehicle image dataset. Step 4: Implement deployment and visualize the actual application effect of the method.
2. According to claim 1, a vehicle detection method under complex weather conditions based on CTI-YOLO is characterized in that: In step 1, collecting original road vehicle images and preprocessing them includes the following steps: Step 1.1: Collect traffic road vehicle images under different lighting conditions, weather conditions and time periods from actual traffic scenes; combine public road traffic datasets, use simulation tools and complex weather generation algorithms to simulate and generate vehicle images similar to actual traffic scenes to supplement the sample size; Step 1.2: Label the specific area of the vehicle and generate a file containing the vehicle type and location; classify the image by vehicle type.
3. The vehicle detection method under complex weather conditions based on CTI-YOLO according to claim 1, characterized in that: In step 2, creating a vehicle detection model CTI-YOLO under complex weather conditions includes the following steps: Step 2.1: The multi-scale residual feature aggregation module (MSRAM) replaces the C2f module in the network structure. The multi-scale residual feature aggregation module uses a variety of convolution kernel sizes for feature extraction and introduces some residual connections, which can retain input features and enhance the utilization of original information; Step 2.2: The global to local spatial aggregation (GLSA) module is introduced, which has the ability to aggregate and represent global and local spatial features, which is beneficial for locating large and small objects respectively; Step 2.3: A feature-sharing dynamic detection head (TADHead) is proposed, which achieves a lightweight design through shared convolution. At the same time, deformable convolution (DCNv2) is introduced to solve the problem that shared parameters may lead to insufficient model expression ability.
4. The vehicle detection method under complex weather conditions based on CTI-YOLO according to claim 3 is characterized in that: The step 2.1 includes the following steps: Capture multi-scale features through convolution kernels of different sizes; The features are split into different parts in the channel dimension and grouped convolution is performed. After the multi-dimensional feature extraction is completed, the module aggregates the features on different paths through feature splicing and fusion operations. By introducing residual connections, the gradient flow is further enhanced and the global features of the input are retained. The specific formula is as follows: f=Conv k*k (x) f 1,1 ,f 1,2 =Chunk(f1,dim=1,groups=2) f 2,1 =Conv 5*5 (f 1,1 ,groups=1) f out =Cat([f3,f 2,2 ,f 1,2 ],dim=1) y=Conv 1*1 (f out )+x Among them, Conv k*k represents the convolution operation with a kernel size of k*k; Cat([·]) represents the concatenation operation in the channel dimension; f3, f 2,2 , f 1,2 Respectively make feature representations of different scales; x represents the feature information of the input.
5. The vehicle detection method under complex weather conditions based on CTI-YOLO according to claim 3 is characterized in that: The step 2.2 includes the following steps: The global information of the entire image is extracted through some global perception mechanisms; next, different local areas of the image are perceived and analyzed by using convolution, dilated convolution or other local perception mechanisms to capture finer-grained detail information; after the global and local information perception, GLSA aggregates the global information with the local information to better capture the boundaries, shapes and positions of the detected objects. The specific formula is as follows: To G (F i 1 )=Softmax(Transpose(C 1*1 (F i 1 ))) L sa (F i 2 )=To L (F i 2 )⊙F i 2 +F i 2 F i 1 ,F i 2 =Split(F i ) Among them, F i 1 is the input of global information extraction, G sa (F i 1 ) is the global information extraction output; F i 2 is the input of local information perception, L sa (F i 2 ) is the output of GSA.
6. The vehicle detection method under complex weather conditions based on CTI-YOLO according to claim 3 is characterized in that: In step 2.3, the detection head extracts multi-scale features from the feature extractor, obtains joint features, generates offsets and masks required by deformable convolution (DCNv2), and further outputs multi-scale detection frames. In order to solve the problem of inconsistent scales of detected targets caused by the use of a shared convolution structure, a scaling layer is used to scale the features. The deformable convolution formula is as follows: Among them, let w k and p k denote the weight of the kth position and the predefined offset Δp respectively. k and Δm k are the learnable offset and modulation scalar at the kth position, respectively.
7. The vehicle detection method under complex weather conditions based on CTI-YOLO according to claim 1, characterized in that: In step 3, the data set is trained to obtain a trained efficient detection model of the improved method; the original image of the traffic scene is input into the trained model to obtain the test result of the model, and a predicted image with a target prediction box and confidence is output.
8. The vehicle detection method under complex weather conditions based on CTI-YOLO according to claim 1, characterized in that: In step 4, the method is deployed on a vehicle monitoring system, the vehicle detection results are visualized, and the vehicle detection method is applied to actual traffic.
Citation Information
Cited By
Rice brown spot detection and counting method, system and equipment
CN121482047A
A method, system and equipment for detecting and counting brown spot disease in rice.
CN121482047B