Remote wireless vehicle detection method and system
Through the improved YOLO11 detection model and system architecture, the accuracy and robustness of vehicle detection in severe weather are solved, and accurate vehicle identification in sandstorms and foggy weather is achieved, reducing system costs and maintenance difficulties.
Patent Information
- Application Number
- CN202510298977.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-04
AI Technical Summary
The existing vehicle detection system is poorly robust in severe weather and small target detection conditions, difficult to achieve accurate identification, and inconvenient installation and maintenance.
The remote wireless vehicle detection method based on the YOLO11 detection model is adopted, through the collaborative work of cloud servers and terminal servers, combined with the SCINet-Enhance module, DWR module, FCA-DFM module and SPPF-SE module, we will perform image preprocessing, inference detection and post-processing to improve the vehicle recognition accuracy in sandstorms and dusty weather.
Without adding auxiliary light sources, accurate detection of vehicles in bad weather is achieved, hardware and maintenance costs are reduced, and vehicle identification accuracy and system robustness are improved.
Smart Images

Figure CN120259987A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicle identification and detection, and in particular to a remote wireless vehicle detection method and system. Background Art
[0002] With the continuous development of the new generation of information technology, smart cities and smart transportation are gradually being improved and implemented. As an important part of the construction of smart cities, smart transportation to a certain extent reflects the perfection of smart cities. In the production and scientific research fields such as traffic control and smart transportation, the identification and statistics of vehicle flow and vehicle quantity provide important data support for relevant production management and problem analysis. Therefore, the accuracy of the identification and statistics of vehicle flow and vehicle quantity is very crucial.
[0003] Currently, vehicle monitoring systems for identifying and counting vehicle flow and vehicle quantity are mostly installed outdoors, and it is inevitable to encounter harsh working conditions that are not conducive to detection, such as sand and dust, fog, and low light. Moreover, due to the distance, the detected targets (various moving vehicles) are often small. In the above situations, the existing vehicle monitoring systems / devices on the market have poor robustness and often need to be equipped with auxiliary light sources for illumination, which not only increases the system purchase cost but also brings inconvenience to the installation and maintenance of the system / device. However, in the face of bad weather, even with the addition of auxiliary light sources, the current existing technologies cannot achieve accurate identification.
[0004] To solve such problems, industry insiders have also conducted a large number of explorations. For example, the invention patent with the patent application number CN118447451A and the invention name "An Airport Runway Foreign Object Detection Method for Small Targets in Complex Environments" simulates foggy images through a generative adversarial network and a variational autoencoder, but the detection effect in low light conditions and sand and dust weather still needs to be improved; the patent application number is CN116959099A, and the invention name is "An Abnormal Behavior Recognition Method Based on a Time-Accumulated Neural Network". The solution disclosed in this patent has an inhibitory effect on noise, but in the face of insufficient image brightness and narrow dynamic range in low light environments, as well as low contrast, image blurring, and non-Gaussian noise in haze and sand and dust environments, this solution may have a poor solution effect or cannot be solved. In short, the solution disclosed in this patent cannot effectively solve the problems of complex static background noise or the restoration of details in images.
[0005] In summary, currently, if we want to improve the accuracy and robustness of vehicle identification and detection in complex environments to enhance the safety and reliability of smart transportation and provide reliable data support for production and scientific research, it is urgent to develop a vehicle detection method that can adapt to outdoor harsh weather, as well as a vehicle detection system or device that is convenient for installation and maintenance. Summary of the Invention
[0006] To solve the problems existing in the existing vehicle recognition and detection means under bad weather and small target detection conditions, the present invention proposes a remote wireless vehicle detection method and system based on the YOLO11 detection model.
[0007] To achieve the above object, the technical solution adopted by the present invention is: to provide a remote wireless vehicle detection method, which includes the following steps:
[0008] S1, collect the original image, collect the depression angle image of the lane of the detection section, and use the depression angle image as the original image;
[0009] S2, original image transmission, use wireless transmission to upload and store the collected original image to the cloud server;
[0010] S3, original image transfer, the terminal server communicates with the cloud server and obtains the original image from the cloud server;
[0011] S4, image preprocessing, the terminal server preprocesses the obtained original image, including adjusting the size, converting the format and normalizing the original image to obtain the input image for inference and detection;
[0012] S5, build a vehicle detection model, build a vehicle detection model based on YOLO11 and run on the terminal server. The vehicle detection model includes a SCINet-Enhance module arranged at the input end of the YOLO11 standard detection model; the SCINet-Enhance module receives the input image and optimizes the input image, and then transmits the optimized image to the YOLO11 standard detection model;
[0013] S6, inference and detection, the input image generated in step S4 is transmitted to the vehicle detection model built in step S5 for image inference and detection to obtain the inference and detection result;
[0014] S7, data post-processing, the terminal server statistically analyzes the inference and detection results obtained in step S6 and presents the statistical analysis results on the corresponding display device.
[0015] In addition, based on the above method and implementation requirements, a remote wireless vehicle detection system is also provided. The system includes a data acquisition module, a data transmission module, a cloud server, and a terminal server. The data acquisition module includes an embedded chip and a camera, and is used to collect the depression angle images of vehicles on the monitored section, that is, the original images. The data transmission module includes a data communication chip, and is used to wirelessly transmit the original images collected by the data acquisition module to the cloud server. The cloud server is used to receive and store the original images transmitted by the data transmission module, and is also used to provide the original images for the terminal server. The terminal server is used to preprocess the original images, perform inference detection on the images obtained after preprocessing, and preliminarily screen and statistically analyze the detection results.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: This method improves the inaccurate recognition caused by sand and dust, fog, and requires a lower light source brightness for recognition. At the same time, when the detected target is small, accurate recognition can still be achieved. The remote wireless vehicle detection system provided by the present invention is convenient for installation and maintenance, and has low hardware costs and low later maintenance costs.
[0017] To achieve better technical effects, further improvements are made on the basis of the above solution, and the specific improvement solutions are as follows.
[0018] Further, step S4 is divided into the following steps:
[0019] S41, Size adjustment, adjust the size of the obtained original image to 640×640;
[0020] S42, Format adjustment, adjust the format of the image processed in step S41, adjust the size format to channels, height, and width, adjust the storage format to RGB format, and adjust the data format to tensor format;
[0021] S43, Normalization processing, perform normalization processing on the image after the format adjustment in step S42, and normalize the image range to 0-1.
[0022] Further, in step S5, the network structure of the YOLO11 standard detection model in the vehicle detection model is adjusted. The specific adjustment method is as follows: Use two DWR modules to replace the first C3K2 module and the second C3K2 module in the backbone part of the YOLO11 standard detection model;
[0023] Behind the SPPF module in the YOLO11 standard detection model, use the FCA-DFM module to replace the C2PSA module in the standard YOLO11 detection model; FCA-DFM is used to enhance the expression ability of attention in foggy and dusty weather;
[0024] Add a set of Upsample + Concat + C3K2 module combinations behind the FCA-DFM module; that is, there are three sets of sequentially connected Upsample + Concat + C3K2 module combinations in the vehicle detection model.
[0025] Replace the DWR module (the first DWR module) of the first C3K2 module in the YOLO11 standard detection model with the Concat module of the Upsample + Concat + C3K2 module group at the transition between the neck and the head; that is, the output feature map of the DWR module is passed to the Concat module and used as part of the input feature maps of the Concat module.
[0026] Furthermore, add an SE module on the basis of the SPPF module of the YOLO11 standard detection model (that is, add an SE module at the output end of the SPPF), and combine it into an SPPF-SE module; the SPPF-SE module is used to enhance the feature expression ability between channels and improve the small target detection performance.
[0027] Furthermore, set the C3K value of the third C3K2 module, the fourth C3K2 module, the eighth C3K2 module, and the C3K2 module of the Upsample + Concat + C3K2 module group connected to the head (Head) in the corresponding YOLO11 standard detection model to be all True; set the C3K value of the fifth C3K2 module, the sixth C3K2 module, and the seventh C3K2 module in the corresponding YOLO11 standard detection model to be all False.
[0028] Furthermore, add a set of Conv + Concat + C3K2 module groups to the neck (Neck) of the YOLO11 standard detection model, and set the C3K value of the C3K2 module in the Conv + Concat + C3K2 module to be True.
[0029] Furthermore, in step S6, the post-processing module performs a non-maximum suppression strategy on the preliminary detection results to screen out effective detection results.
[0030] Furthermore, the terminal server is provided with an image preprocessing module, an image inference module, and an image postprocessing module; the image preprocessing module is used to preprocess the original image obtained from the cloud server, and use the image obtained after preprocessing as the input image and transfer it to the image inference module; the image inference module is a vehicle detection model established based on YOLO11, which performs inference detection on the input image through the vehicle detection model and transfers the inference detection result to the postprocessing module; the postprocessing module is used to perform statistical analysis, marking, and presentation on the inference detection result obtained by the image inference module. Description of the Drawings
[0031] Figure 1 is a flowchart of the vehicle detection method according to Embodiment 1 of the present invention;
[0032] Figure 2 is a network structure block diagram of the YOLO11 standard detection model;
[0033] Figure 3 is a network structure block diagram of the vehicle detection model in Embodiment 1;
[0034] Figure 4 is a network structure block diagram of the SCI-Enhance module in Embodiment 1;
[0035] Figure 5 is a network structure block diagram of the vehicle detection model according to Embodiment 2 of the present invention;
[0036] Figure 6 is the principle or structure diagram of the DWR module in Embodiment 2;
[0037] Figure 7 is the principle or structure diagram of the FCA-DFM module in Embodiment 2;
[0038] Figure 8 is the principle or structure diagram of the C3K2 module in the YOLO11 standard detection model;
[0039] Figure 9 is a network structure block diagram of the vehicle detection model in Embodiment 3;
[0040] Figure 10 is the principle or structure diagram of the SPPF-SE module in Embodiment 3;
[0041] In the figure:
[0042] Backbone: The first C3K2 module 3; the second C3K2 module 5; the third C3K2 module 7; the fourth C3K2 module 9; the fifth C3K2 module 14; the sixth C3K2 module 17; the seventh C3K2 module 20; the eighth C3K2 module 23; the SPPF module 10; the C2PSA module 11.
[0043] Neck: For the first group of Upsample+Concat+C3K2 module groups, the corresponding serial numbers of each module are 12, 13, and 14; for the second group of Upsample+Concat+C3K2 module groups, the corresponding serial numbers of each module are 15, 16, and 17 respectively; for the first group of Conv+Concat+C3K2 module groups, the corresponding serial numbers of each module are 18, 19, and 20 respectively; for the second group of Conv+Concat+C3K2 module groups, the corresponding serial numbers of each module are 21, 22, and 23 respectively.
[0044] SCI-Enhance module A; the first DWR module B; the second DWR module C; the SPPF-SE module D; the FCA-DFM module E; the added Upsample+Concat+C3K2 module group F; the added Conv+Concat+C3K2 module group G. (Some modules are labeled in the figure, refer to the appendix Figure 2 As shown, since the specific implementation does not involve it, it will not be listed one by one here) Specific implementation
[0045] The present invention is mainly based on the improvement and optimization of the YOLO11 detection model, and realizes the accurate detection of vehicle flow (vehicles) under harsh weather and small target detection conditions without the need to add auxiliary light sources. Among them, the YOLO11 standard detection model is as shown in the appendix Figure 2 As shown, from Figure 2 it can be seen that the YOLO11 standard detection model is mainly composed of three parts (or three layers): the backbone, the neck, and the head; starting from the beginning of the backbone are the first Conv module, the second Conv module, the first C3K2 module 3, the third Conv module, the second C3K2 module 5, the fourth Conv module, the third C3K2 module 7, the fifth Conv module, the fourth C3K2 module 9, the SPPF module 10, and the C2PSA module 11 connected to the neck; the neck includes the first Upsample+Concat+C3K2 (i.e., the fifth C3K2 module 14) module group, the second Upsample+Concat+C3K2 (i.e., the sixth C3K2 module 17) module group, the first Conv+Concat+C3K2 (i.e., the seventh C3K2 module 20) module group, and the second Conv+Concat+C3K2 (i.e., the eighth C3K2 module 23) module group connected in sequence.
[0046] As shown in the appendix Figure 2As shown, the second C3K2 module 5 is connected to the Concat module in the second Upsample+Concat+C3K2 module group, that is, the output feature map of the second C3K2 module 5 is used as a partial input feature map of the Concat module in the second Upsample+Concat+C3K2 module group; similarly, the output feature map of the third C3K2 module 7 is used as a partial input feature map of the Concat module in the first Upsample+Concat+C3K2 module group; the output feature map of the C2PSA module 11 is used as a partial input feature map of the Concat module in the second Conv+Concat+C3K2 module group; the output feature map of the C3K2 module (i.e., the fifth C3K2 module 14) in the first Upsample+Concat+C3K2 module group is used as a partial input feature map of the Concat module in the first Conv+Concat+C3K2 module group.
[0047] The C3K2 module is an important feature extraction component in YOLO11. It is an improved design based on the traditional C3 module. It provides more powerful feature extraction capabilities by combining variable convolution kernels (such as 3x3, 5x5, etc.) and channel separation strategies. It is especially suitable for more complex scenes and deep feature extraction tasks. The C3K2 module usually divides the input features into two parts. One part is directly passed through ordinary convolution operations, and the other part is extracted through multiple C3Ks (when the C3K parameter is set to True) or Bottleneck structures for deep feature extraction. Finally, the two parts of the features are concatenated and fused through 1x1 convolution. This structure can maintain lightweight while effectively extracting deep features. For details about C3K2, please refer to the attached. Figure 8 As the prior art, it will not be described here in detail.
[0048] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0049] Implementation column 1.
[0050] Combined with Figures 1-4 It can be seen that the remote wireless vehicle detection method provided in Example 1 of the present invention includes the following steps: S1, collecting original images, collecting depression-angle images of lanes of the detection section (in the specific implementation process, the scene to be dealt with may also be a formal non-lane, road, such as a parking lot, etc.), and using the depression-angle image as the original image; S2, original image transmission, uploading the collected original image to a cloud server by wireless transmission and storing it;
[0051] S3, original image transfer, the terminal server communicates with the cloud server and obtains the original image from the cloud server;
[0052] S4. Image preprocessing: The terminal server preprocesses the acquired original image, including resizing, format conversion, and normalization of the original image to obtain an input image for inference and detection.
[0053] S5. Build a vehicle detection model: Based on the YOLO11 standard detection model, a vehicle detection model is built and run on the terminal server. As shown in the appendix Figure 3 The vehicle detection model includes a SCINet-Enhance module and the YOLO11 standard detection model, where the SCINet-Enhance module is set at the input end of the YOLO11 standard detection model. The SCINet-Enhance module receives the input image, optimizes it, and then passes the optimized image (i.e., the output feature map of the SCINet-Enhance module) to the YOLO11 standard detection model.
[0054] S6. Inference and detection: The input image generated in step S4 is passed to the vehicle detection model built in step S5 for image inference and detection to obtain inference and detection results.
[0055] S7. Data post-processing: The terminal server statistically analyzes the inference and detection results in step S6 and presents the statistical analysis results on the corresponding display device.
[0056] In the specific implementation process, step S4 is refined as follows:
[0057] S41. Resizing: Resize the acquired original image to 640×640.
[0058] S42. Format adjustment: Adjust the format of the image processed in step S41, change the size format to channels, height, and width, the storage format to RGB format, and the data format to tensor format.
[0059] S43. Normalization processing: Normalize the image after format adjustment in step S42 to normalize the image range to 0-1.
[0060] In step S5, an SCI-Enhance module A is added to the established vehicle detection model. This module consists of an input convolutional layer, multiple convolutional blocks and residual connections, a channel attention mechanism, an output convolutional layer, and a final output and residual connection. The principles of each part of the SCI-Enhance module A will be described in detail below.
[0061] (1) Input convolutional layer
[0062] The input image is subjected to feature extraction through a convolutional layer, and its formula (1) is as follows:
[0063] F0 = LeakyReLU(Conv2D(Input)) (1)
[0064] Where:
[0065] F0 is the initial feature map;
[0066] is the input image;
[0067] Conv2D is a two-dimensional convolution operation;
[0068] LeakyReLU is an activation function with a negative slope of 0.1.
[0069] In the SCINet-Enhance module, the main purpose of the input convolutional layer is to use Conv2D to convert the input image from 3 channels to C channels for simultaneous use and extract preliminary features (preliminary features include three categories: edges and contours, textures and patterns, and color distributions). The LeakyReLU activation function is used in the input convolutional layer. Compared with the standard ReLU, LeakyReLU has a small slope in the negative value region, which can retain more gradient information during training, thereby improving the performance of the model.
[0070] (2) Multi-layer convolutional blocks and residual connections
[0071] The initial feature map F0 passes through multiple convolutional blocks. Each convolutional block consists of a convolutional layer, batch normalization (Batch Norm), and an activation function, and a residual connection is introduced. For the l-th layer (l = 1, 2, 3), the formula (2) for the output feature map of its l-th layer is as follows:
[0072] F l = LeakyReLU(BatchNorm(Conv2D(F l-1 )))+F l-1 (2)
[0073] Where, F l-1 is the output feature map of the previous layer (i.e., the l-1 layer).
[0074] By adding a residual connection between convolutional blocks, in this way, the model can be more easily optimized and avoid the problem of gradient disappearance; in addition, the residual connection can also accelerate convergence.
[0075] (3) Channel Attention mechanism
[0076] After the multi-layer convolutional blocks, a channel attention mechanism is introduced to adaptively adjust the importance of different channels. The calculation steps of the channel attention mechanism are as follows:
[0077] A. Global Average Pooling and Global Max Pooling
[0078] Global average pooling and global max pooling are performed on the feature map, as shown in Equation (3):
[0079] F avg = AP(F L )
[0080] F max = MP(F L ) (3)
[0081] Where:
[0082] is the final output feature map of the convolutional block;
[0083] AP is the global average pooling operation;
[0084] MP is the global max pooling operation.
[0085] B. Shared Multilayer Perceptron (MLP)
[0086] The features after pooling in the previous step are transformed through a shared multilayer perceptron (MLP), as shown in Equation (4):
[0087] MLP(F avg / max ) = Conv2D(LeakyReLU(Conv2D(F avg / max ))) (4)
[0088] Where:
[0089] F avg / max is to perform global average pooling and global max pooling on the feature map separately in the channel dimension.
[0090] Specifically, the MLP consists of two 1x1 convolutional layers. The first convolutional layer reduces the number of channels from C to The second convolutional layer restores the number of channels to C; where reduction is the channel reduction factor.
[0091] C. Generating Attention Weights
[0092] The outputs of average pooling and max pooling are added after being processed by the MLP, and the Sigmoid activation function is applied to generate the final channel attention weights, as shown in Equation (5):
[0093] A = σ(MLP(F avg ) + MLP(F max )) (5)
[0094] Wherein:
[0095] σ represents the Sigmoid activation function;
[0096] is the channel attention weight.
[0097] D. Applying the attention weight
[0098] Apply the channel attention weight to the feature map, as shown in Equation (6):
[0099]
[0100] Wherein, represents the per-channel element-wise multiplication operation.
[0101] Add a ChannelAttention class to the SCINet-Enhance module to implement the channel attention mechanism; this mechanism generates two feature maps through global average pooling and global max pooling, fuses them through convolutional operations, and finally uses the Sigmoid activation function to generate weights. This mechanism helps the model automatically focus on more important channel features and improves the model's ability to adjust the importance weights of different channels.
[0102] (4) Output convolutional layer
[0103] The feature map F' processed by the channel attention mechanism generates the final output feature map through the output convolutional layer, as shown in Equation (7):
[0104] O = σ(Conv2D(F')) (7)
[0105] Wherein:
[0106] Conv2D is a two-dimensional convolutional operation;
[0107] σ is the Sigmoid activation function, which is used to limit the output within the range of [0, 1].
[0108] (5) Final output and residual connection
[0109] Finally, perform a residual connection between the output feature map O and the input image Input, and limit the result, as shown in Equation (8):
[0110] Output = Clamp(O + I, 0.0001, 1) (8)
[0111] Wherein:
[0112] The Clamp function limits the output value between [0.0001, 1] to ensure that the pixel values of the output feature map are within a reasonable range;
[0113] The Output is the output feature map of the SCINet-Enhance module.
[0114] Example 2.
[0115] Combined with the attached Figures 5-8 and compared with the attached Figure 1 Example 2 will be described in detail.
[0116] In step S5, the vehicle detection model established in Example 2, in addition to adding the SCINet-Enhance module at the starting end of the backbone of the YOLO11 standard detection model (i.e., the input end of the YOLO11 standard detection model), also made the following adjustments to the YOLO11 standard detection model (network structure):
[0117] (1) Use two DWR modules (the first DWR module B and the second DWR module C) to replace the first C3K2 module 3 and the second C3K2 module 5 in the backbone part of the YOLO11 standard detection model respectively;
[0118] (2) Behind the SPPF module 10 in the YOLO11 standard detection model, use the FCA-DFM module E to replace the C2PSA module 11 in the standard YOLO11 detection model; the FCA-DFM module E is used to enhance the expression ability of attention in foggy and dusty weather;
[0119] (3) Add a group of Upsample+Concat+C3K2 module combinations behind the FCA-DFM module E; that is, there are three groups of sequentially connected Upsample+Concat+C3K2 module combinations in the vehicle detection model.
[0120] Referring to the attached Figure 5 in Example 2: The output feature map of the first DWR module B serves as the input feature map of the Concat module in the third Upsample+Concat+Conv module group; the output feature map of the second DWR module C serves as the input feature map of the Concat module in the second group of Upsample+Concat+Conv module groups; the feature map of the third C3K2 module 7 serves as the input feature of the Concat module in the first group of Upsample+Concat+Conv module groups (consistent with the original YOLO11 standard detection model here).
[0121] The DWR module uses a two-step residual feature extraction method (regional residualization–semantic residualization) to effectively improve the efficiency of multi-scale information capture in real-time semantic segmentation. The specific principle of the DWR module is as shown in the attached Figure 6 shown.
[0122] Among them, the FCA-DFM module E is a neural network module that combines channel attention and spatial attention mechanisms for enhancing feature representation, and its network structure is as shown in the appendix Figure 7 shown. The following will elaborate on each module in the appendix Figure 7 in detail
[0123] Adaptive_Avg_Pool: refers to the adaptive average pooling function. Adaptive Average Pooling is a pooling operation mainly used in convolutional neural networks (CNNs) in deep learning. Its role is to generate a fixed-size output by pooling the input tensor, regardless of the input size. Adaptive average pooling does not require specifying a fixed window size, but instead dynamically calculates the window size and stride by specifying the target size of the output to adapt to the input dimensions
[0124] Therefore, this module is particularly suitable for processing input images (data) of variable sizes. Assume the input size is H in ×W in , and the output target size is H out ×W out , and the window size and stride will be dynamically determined according to formula (9):
[0125]
[0126] where:
[0127] k h and k w are the height and width of the pooling window
[0128] Adaptive_Max_pool module: refers to the adaptive max pooling function. Adaptive Max Pooling is a pooling operation similar to adaptive average pooling, but it adopts the max pooling strategy. Its main role is to scale the input tensor to the specified output size while retaining the most significant eigenvalue in the input. The calculation formula for the window size and stride is the same as that of the above-mentioned adaptive average pooling
[0129] Conv1D module: is a one-dimensional convolution. Conv1D can effectively extract local features from the input one-dimensional data. By sliding the convolutional kernel (filter) over the sequence for convolution operations, Conv1D can capture local patterns, trends, and correlations in the data
[0130] Conv2D module: It is a two-dimensional convolution, referring to a layer that applies convolution operations in the spatial dimensions of two-dimensional data (such as images). Different from one-dimensional convolution (Conv1D) for processing sequential data, Conv2D is mainly used to process two-dimensional data structures with height and width, such as the RGB channels of color images.
[0131] Concat module: It is an operation that concatenates tensors along a specific dimension, usually used to fuse feature maps from different sources for subsequent processing.
[0132] Sigmod: It is a commonly used activation function in neural networks that maps variables to the range [0, 1]. The expression of this activation function is as shown in formula (10):
[0133]
[0134] The FCA-DFM module E mainly consists of two branches: the channel attention branch and the spatial attention branch; this module (FCA-DFM) fuses the outputs of the channel attention branch and the spatial attention branch to obtain the final attention-enhanced features; the principle of the FCA-DFM module E will be introduced in detail below, as follows:
[0135] (1) Channel attention branch
[0136] The channel attention branch enhances feature representation by focusing on the importance of different channels. The specific steps are as follows (a, b, c, d):
[0137] a. Global average pooling
[0138] Perform global average pooling on the input feature map to obtain the global description of each channel, as shown in formula (11):
[0139]
[0140] Where:
[0141] is the input feature map;
[0142] N is the batch size;
[0143] C is the number of channels;
[0144] H and W are the height and width of the feature map respectively;
[0145] The feature map after global average pooling.
[0146] b. Feature transformation
[0147] Transform the features after global average pooling through two different convolution operations:
[0148] First, perform a dimensionality transformation on F avg (X), changing from to See Equation (12):
[0149]
[0150] Then, apply a one-dimensional convolution (Conv1d), see Equation (13):
[0151]
[0152] Directly apply a 1×1 fully-connected convolution (Conv2d) to F avg (X) to obtain F' avg (X), see Equation (14):
[0153]
[0154] Where:
[0155] F’ avg (X) is the feature map obtained by directly applying a 1×1 fully-connected convolution to F avg (X);
[0156] Then, change x2 to See Equation (15):
[0157]
[0158] c. Attention map calculation
[0159] Calculate two attention maps using x1 and x2, see Equation (16):
[0160]
[0161] Where:
[0162] σ represents the Sigmoid activation function;
[0163] d is the channel dimension.
[0164] d. Attention map fusion
[0165] Fuse the two attention maps, see Equation (17):
[0166]
[0167] Where:
[0168] w is a learnable parameter, constrained between (0,1) by the Sigmoid activation function;
[0169] The channel is the output feature map after the weighted sum of the two attention mechanisms;
[0170] represents element-wise multiplication.
[0171] Then, the fused attention map is further convolved and activated, as shown in Equation (18):
[0172] channel_attn = σ(Conv1d(channel)) (18)
[0173] where:
[0174] channel_attn is the output after the weighted fused attention passes through a one-dimensional convolution and an activation function;
[0175] (2) Spatial attention branch
[0176] The spatial attention branch enhances the feature representation by focusing on the importance of the spatial positions of the feature map. The specific steps are as follows:
[0177] a. Spatial pooling
[0178] Perform (adaptive) global average pooling on the input feature map, as shown in Equation (19):
[0179]
[0180] At the same time, perform global max pooling on the input feature map, as shown in Equation (20):
[0181]
[0182] where:
[0183] spatial_avg is the output feature map after global average pooling.
[0184] b. Feature fusion
[0185] Concatenate the results of global average pooling and global max pooling along the channel dimension to obtain:
[0186]
[0187] Fuse the features with a 7×7 convolutional kernel and apply the Sigmoid activation function to obtain the spatial attention weights, as shown in Equation (21):
[0188]
[0189] (3) Attention mechanism fusion and feature enhancement
[0190] Combine channel attention and spatial attention to re-weight the input feature map, as shown in Equation (22):
[0191]
[0192] Where:
[0193] denotes element-wise multiplication;
[0194] channel_attn is broadcast in the spatial dimension to match the size of X.
[0195] The main advantages of using the FCA-DFM module E in the model are as follows:
[0196] First, by multiplying the input features by both channel attention and spatial attention simultaneously, finer-grained feature adjustment is achieved. This dual attention mechanism can more effectively highlight important features and suppress unimportant features, thereby improving the performance of the model.
[0197] Second, extract spatial features through global average pooling and global max pooling, and use a convolutional layer to combine the results of these two pooling methods to generate a spatial attention map. This enables the model to not only focus on the importance of each channel but also on the importance of different spatial positions in the feature map, thereby enhancing the model's sensitivity and expressive ability to spatial information.
[0198] Third, use a large kernel convolution (with kernel size = 7) in the spatial attention branch. The larger convolution kernel can capture a wider range of spatial context information and enhance the expressive ability of spatial attention.
[0199] In Embodiment 2, the C3K values in the third C3K2 module 7, the fourth C3K2 module 9, the eighth C3K2 module 23 in the corresponding YOLO11 standard detection model, and the C3K2 module in the Upsample+Concat+C3K2 module group connected to the head are all set to True; the C3K values in the fifth C3K2 module 14, the sixth C3K2 module 17, and the seventh C3K2 module 20 in the corresponding YOLO11 standard detection model are all set to False; and in this embodiment, a group of Conv+Concat+C3K2 module groups can also be added to the neck of the YOLO11 standard detection model, and the C3K value of its C3K2 module is set to True. To achieve the selective introduction of residuals. At the same time, through the adoption of three groups of Upsample+Concat+C3K2 module groups and three groups of Conv+Concat+C3K2 module groups, the optimization of the feature pyramid is realized. From the original three levels of P3, P4, and P5 corresponding to the YOLO11 standard detection model, a four-layer feature pyramid of P2-P5 is constructed through three upsampling operations (i.e., three groups of Upsample+Concat+C3K2 module groups); the shallower P2 layer contains richer spatial detail information, which is particularly suitable for the detection of small targets.
[0200] Embodiment 3.
[0201] Next, it will be combined with the attached Figure 8 and the attached Figure 9 and with reference to the attached Figure 1 to introduce Embodiment 3 in detail.
[0202] Embodiment 3 is a further improvement based on Embodiment 2. It mainly replaces the SPPF module 10 with the SPPF-SE module D, that is, an SE module is connected to the tail (output end) of the SPPF module 10; the specific functions of the SE module are as follows:
[0203] (1) Enhance the feature expression ability. The SE module can dynamically adjust the importance of each channel by recalibrating the weights between channels; this enables the network to pay more attention to the important features relevant to the current task, suppress irrelevant or redundant features, and thus improve the overall feature expression ability;
[0204] (2) Enhance the generalization ability of the model. By adaptively adjusting the feature channels, the SE module helps the model better adapt to different data distributions and changes, and improves the generalization ability of the model on unseen data.
[0205] Next, the principle and functions of the SPPF-SE module D will be introduced in detail, and the specific situation is as follows:
[0206] (1) Convolution operation
[0207] Input feature map First, the number of channels is reduced from C1 to
[0208] where:
[0209] is the output feature map of the initial convolution.
[0210] (2) Multi-scale max pooling
[0211] Next, three consecutive max pooling operations are performed, each on the output of the previous layer, as shown in Equation (23):
[0212] Y i = MP(Y i-1 )(i = 1, 2, 3) (23)
[0213] where:
[0214] Y i is the output feature map after the i-th pooling.
[0215] After each pooling operation, the output feature map maintains the spatial size H×W, but captures different context information through pooling kernels of different scales.
[0216] The three consecutive max pooling operations are a downsampling operation used to reduce the size of the feature map. The main role of the max pooling layer is to reduce the dimension of the feature map by taking the maximum value in a local area, extract significant features, thereby reducing the computational amount and parameter scale, and at the same time improving the robustness of the model to changes such as translation and noise. It effectively retains key information, reduces the model complexity, helps prevent overfitting, and enhances the expression ability of features. It is a commonly used downsampling method in convolutional neural networks.
[0217] (3) Feature concatenation
[0218] The initial convolution output feature map Y0 and the feature maps Y1, Y2, Y3 after three poolings are concatenated in the channel dimension to obtain Y concat ,
[0219] (4) Upsampling convolution
[0220] The concatenated features are passed through another 1×1 convolutional layer Conv2 to increase the number of channels from 4C′ to C2, as shown in Equation (24):
[0221] Y conv2 = Conv2(Y concat ),
[0222] (5) Squeeze
[0223] Perform global average pooling on the input feature map to generate a statistical description for each channel See Equation (25);
[0224]
[0225] where: c belongs to 1, 2, … C2.
[0226] The purpose of squeezing is to capture the global context information of each channel through global information aggregation.
[0227] (6) Excitation
[0228] Generate channel weights by passing the compressed description z through two fully connected layers and an activation function Let the size of the intermediate hidden layer be Channel weight generation, see Equation (26):
[0229] s = σ(W2 · δ(W1 · z)) (26)
[0230] where:
[0231] The weight matrix of the first fully connected layer;
[0232] The weight matrix of the second fully connected layer;
[0233] δ(·) is the ReLU activation function;
[0234] σ(·) is the Sigmoid activation function.
[0235] The purpose of excitation is to generate a weight coefficient for each channel by adaptively learning the relationship between channels, which is used to recalibrate the channel response of the original feature map.
[0236] The specific steps are as follows:
[0237] A. The first fully connected layer and ReLU activation, see Equation (27):
[0238] a = δ(W1 · z) (27)
[0239] where:
[0240]
[0241] B. The second fully connected layer and Sigmoid activation, see Equation (28):
[0242] s = δ(W2·a) (28)
[0243] where:
[0244]
[0245] (7) Scaling
[0246] Apply the generated channel weight s to the original features Figure X , to achieve channel-level dynamic recalibration, see formula (29):
[0247] X' c,i,j = s c ·X c,i,j (29)
[0248] where:
[0249] c = 1, 2, …, C2;
[0250] i = 1, 2, …, H;
[0251] j = 1, 2, …, W.
[0252] That is, as shown in formula (30):
[0253]
[0254] where, represents channel-wise multiplication.
[0255] SPPF (Spatial Pyramid Pooling-Fast) in the YOLO11 standard detection model is used to enhance the feature extraction ability; without significantly increasing the computational cost, it expands the receptive field of the network and captures features at different scales.
[0256] The main goal of the SE module is to enhance the expression ability of the convolutional neural network for key features. By adaptively adjusting the weights of each channel, the model pays more attention to important channel features; by introducing the SE module into the SPPF module 10, the features of different channels are weighted.
[0257] The advantages of the SE module are as follows.
[0258] (1) Lightweight: The number of parameters of the SE module is relatively small, usually only increasing a small amount of computational overhead. Therefore, it can be easily integrated into various network architectures without significantly increasing the model complexity.
[0259] (2) Modular design: The SE module, as an independent module, can be flexibly inserted into the existing convolutional blocks to enhance their feature expression ability.
[0260] (3) In the image classification task, the SE module improves the model's ability to distinguish different classes by emphasizing key features, thereby enhancing the classification accuracy.
[0261] Embodiment 4.
[0262] Based on the remote wireless vehicle detection method provided by the present invention above, the developed remote wireless vehicle detection system mainly consists of a data acquisition module, a data transmission module, a cloud server, a terminal server, and other components.
[0263] Among them, the data acquisition module includes an embedded chip and a camera, which are used to collect the depression angle images of vehicles on the monitored section, that is, the original images; in this embodiment, the embedded chip is selected as Raspberry Pi 4 Model B, which is equipped with a 1.5GHz quad-core processor, 8GB RAM, supports gigabit Ethernet and Wi-Fi 802.11ac, and has good scalability; the camera is selected as Raspberry Pi Camera Module V2, which supports 1080p video and can be well compatible with the processor chip of Raspberry Pi.
[0264] The data transmission module includes a data communication chip, which is used to wirelessly transmit the original images collected by the data acquisition module to the cloud server; in this embodiment, the data communication chip uses the Edimax EW-7811 wireless network chip, which is adapted to Raspberry Pi and can transmit data through Wifi; to achieve the wireless transmission of the original images.
[0265] The cloud server is used to receive and store the original images transmitted by the data transmission module, and is also used to provide the original images for the terminal server; in this embodiment, the cloud server uses the Alibaba Cloud server ECS; the Alibaba Cloud server ECS is a secure, reliable, elastic and scalable cloud computing service, which can significantly reduce the implementation cost and improve the operation efficiency.
[0266] The terminal server is used to preprocess the original images, perform inference detection on the images obtained after preprocessing, and screen and statistically analyze the inference detection results.
[0267] The terminal server is provided with an image preprocessing module, an image inference module, and an image postprocessing module; the image preprocessing module is used to preprocess the original images obtained from the cloud server, and use the images obtained after preprocessing as input images and transfer them to the image inference module.
[0268] The image inference module builds a vehicle detection model based on YOLO11. It performs inference detection on the input image data through the vehicle detection model and passes the inference detection results to the post-processing module; the post-processing module is used to perform statistical analysis, marking, and presentation on the inference detection results obtained by the image inference module.
[0269] In this embodiment, the terminal server is a local working computer (or host computer), and the configuration of the terminal server is determined according to specific performance requirements.
[0270] When the post-processing module performs post-processing on the inference detection results, it includes the following steps:
[0271] (1) Perform Non-Maximum Suppression (NMS) to filter the detection results. Non-Maximum Suppression is to suppress elements that are not maxima.
[0272] Non-Maximum Suppression plays a very crucial role in object detection. During the training process, the object detection algorithm adjusts the deep learning network parameters according to the given true valid values to fit the object features of the dataset. After training is completed, the parameters of the neural network are fixed, thus playing a role in directly predicting objects in new images.
[0273] In actual object prediction, general object detection algorithms will generate a very large number of object bounding boxes, among which many overlapping boxes locate the same object. As the last step of object detection, NMS obtains the true object bounding boxes.
[0274] The steps of NMS are as follows:
[0275] Step 1: Sorting. Sort all candidate bounding boxes according to the score of each candidate bounding box (usually the confidence of the object category), and the bounding box with a higher score is ranked in the front;
[0276] Step 2: Select the bounding box with the highest score. Select the candidate bounding box with the highest score and keep it as the final detection result;
[0277] Step 3: Suppress overlapping bounding boxes. Calculate the overlap degree between the currently selected candidate bounding box and other candidate bounding boxes. Usually, the Intersection over Union (IoU) is used to measure it. The calculation of IoU is shown in formula (31):
[0278]
[0279] If the IoU value of a certain candidate bounding box with the currently selected bounding box exceeds the preset threshold (set to 0.45 in the present invention), then this bounding box is considered redundant and needs to be suppressed;
[0280] Step 4: Repeat the process, repeat Step 2 and Step 3 until all the boxes are processed;
[0281] Step 5: Output the result, the last remaining box is the final result after non-maximum suppression;
[0282] (2) Adjust the predicted box back to the original image size
[0283] In this embodiment, the mAP50 index of the vehicle detection model is 0.707; mAP50 is an important index for evaluating the accuracy of the object detection model, which represents the average precision calculated under the condition that the IoU threshold is 0.5. It combines Precision and Recall, and gives a comprehensive value of the overall performance of the model by calculating the average precision under all categories.
[0284] The above detailed description is a specific description of the feasible embodiments of the present invention, and this embodiment is not intended to limit the patent scope of the present invention. Any equivalent implementation or change without departing from the present invention shall be included in the patent scope of this case.
Claims
1. A remote wireless vehicle detection method, characterized in that, Including the following steps: S1. Collect the original image, and collect the depression angle image of the lane of the detection section. The depression angle image is used as the original image; S2. Transmission of the original image. The collected original image is uploaded and stored in the cloud server by using a wireless transmission method; S3. Transfer of the original image. The terminal server communicates with the cloud server and obtains the original image from the cloud server; S4. Image preprocessing. The terminal server preprocesses the obtained original image, including adjusting the size, converting the format, and normalizing the original image to obtain the input image for inference and detection; S5. Build a vehicle detection model. Based on YOLO11, build a vehicle detection model loaded and running on the terminal server. The vehicle detection model includes a SCINet-Enhance module set at the input end of the YOLO11 standard detection model; the SCINet-Enhance module receives the input image, optimizes the input image, and then transfers the optimized image to the YOLO11 standard detection model; S6. Inference and detection. Transfer the input image obtained in step S4 to the vehicle detection model built in step S5 for image inference and detection to obtain the inference and detection results; S7. Data post-processing. The terminal server statistically analyzes the inference and detection results in step S6 and presents the statistical analysis results on the corresponding display device.
2. The remote wireless vehicle detection method according to claim 1, characterized in that Step S4 includes the following sub-steps: S41. Size adjustment. Adjust the size of the obtained original image to 640×640; S42. Format adjustment. Adjust the format of the image processed in step S41. Adjust the size format of the image to channels, height, and width, adjust the storage format to RGB format, and adjust the data format to tensor format; S43. Normalization processing. Perform normalization processing on the image after the format adjustment in step S42 to normalize the image range to 0-1.
3. The remote wireless vehicle detection method according to claim 1, wherein: In step S5, adjust the network structure of the YOLO11 standard detection model in the vehicle detection model. The specific adjustment method is as follows: Use two DWR modules to replace the first C3K2 module and the second C3K2 module in the backbone of the YOLO11 standard detection model; Behind the SPPF module in the YOLO11 standard detection model, use the FCA-DFM module to replace the C2PSA module in the YOLO11 standard detection model; Add a group of Upsample+Concat+C3K2 module combinations behind the FCA-DFM module; The DWR module replacing the first C3K2 module is connected to the Concat module in the Upsample+Concat+C3K2 module group connected to the head; 4. The remote wireless vehicle detection method according to claim 3, wherein: Add an SE module on the basis of the SPPF module in the YOLO11 standard detection model; the SE module is added at the output end of the SPPF module to form an SPPF-SE module.
5. The remote wireless vehicle detection method according to claim 3 or 4, characterized in that: Set the C3K values in the third C3K2 module, the fourth C3K2 module, the eighth C3K2 module in the corresponding YOLO11 standard detection model, and the C3K2 modules in the Upsample+Concat+C3K2 module group connected to the head to be all True; set the C3K values in the fifth C3K2 module, the sixth C3K2 module, and the seventh C3K2 module in the corresponding YOLO11 standard detection model to be all False.
6. The remote wireless vehicle detection method according to claim 5, characterized in that: Add a group of Conv+Concat+C3K2 module groups to the neck of the YOLO11 standard detection model, and set the C3K value of the C3K2 module in the Conv+Concat+C3K2 module to be True.
7. The remote wireless vehicle detection method according to claim 1, characterized in that: In step S6, the post-processing module performs a non-maximum suppression strategy on the inference detection results to screen out effective detection results.
8. A remote wireless vehicle detection system, characterized in that: It includes a data acquisition module, a data transmission module, a cloud server, and a terminal server; The data acquisition module includes an embedded chip and a camera, and is used to collect the depression angle images of vehicles on the monitored section, that is, the original images; The data transmission module includes a data communication chip, and is used to upload the original images collected by the data acquisition module to the cloud server in a wireless transmission manner; The cloud server is used to receive and store the original images uploaded by the data transmission module, and is also used to provide the original image data for the terminal server; The terminal server is used to preprocess the original images, perform inference detection on the images obtained after preprocessing, and screen, statistically analyze the inference detection results.
9. The wireless vehicle detection system according to claim 8, wherein: The terminal server is provided with an image preprocessing module, an image inference module, and an image postprocessing module; The image preprocessing module is used to preprocess the original images obtained from the cloud server, and use the images obtained after the preprocessing as input images and transfer them to the image inference module; The image inference module is a vehicle detection model established based on YOLO11. It performs inference detection on the input images through the vehicle detection model and transfers the inference detection results to the postprocessing module; The postprocessing module is used to statistically analyze, mark, and present the inference detection results obtained by the image inference module.
Citation Information
Patent Citations
Abnormal behavior recognition method based on space-time diagram convolutional neural network
CN116959099A
Airport runway foreign matter detection method for small target in complex environment
CN118447451A