Blind sidewalk illegal parking identification method based on target detection
By improving the YOLOv11 and GRFB_UNet networks, combining the AMFA and ELA attention mechanisms, the ByteTrack tracking algorithm is used to solve the real-time and complexity of blind track illegal parking detection, and efficient illegal parking identification and license plate acquisition are achieved.
Patent Information
- Application Number
- CN202510419811.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art relies on the original image quality in blind path illegal parking detection, with low real-time performance, lack of means of discrimination for short-term or long-term illegal parking, and the complex algorithm architecture leads to large parameters and high edge deployment costs.
The YOLOv11 network model is improved, combined with GRFB_UNet, a multi-scale convolution adaptive aggregation feature module (AMFA) and an improved ELA attention mechanism are designed, and the ByteTrack tracking algorithm is used, and the NWD loss function and SCDown downsampling is used to generate background images through inter-frame differential method to achieve lightweight and real-time detection.
It improves the accuracy and recall rate of blind path illegal parking detection, reduces model parameters and calculation overhead, realizes real-time monitoring and identification of illegal parking vehicles, obtains license plate information, and reduces the deployment difficulty and cost of edge equipment.
Smart Images

Figure CN120356165A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of traffic detection, and particularly to a method for identifying illegal parking on blind paths based on object detection. Background Technique
[0002] As a dedicated passage for visually impaired people to assist in walking, the construction and management of blind paths are of crucial importance. Blind paths can effectively improve the travel safety of blind people and enable them to complete their daily lives more independently. Since blind paths are usually located on the outermost side of the sidewalk, they are often easily affected by illegal vehicle parking, resulting in the occupation or damage of blind paths. This not only affects the use effect of blind paths but also brings great troubles and potential safety hazards to the travel of visually impaired people.
[0003] In the prior art, Yang Zuliang et al. extracted moving objects through a mixture Gaussian model, tracked the objects using the Meanshift algorithm, and finally used a convolutional neural network to determine whether a stationary object is a vehicle. Compared with traditional methods, the accuracy of this method has been improved by 63%; Tang Jie calculated the gray histogram of the set no-parking area. When the gray level changes and tends to be stable, the SSD algorithm is used to detect whether the object is a vehicle to achieve illegal parking detection in the specified area. These methods rely heavily on the quality of the original image, have low real-time performance, lack means to distinguish between vehicles passing by briefly or parking illegally for a long time, and the overly complex algorithm architecture leads to an increase in the number of parameters, thus increasing the difficulty and cost of edge deployment. Summary of the Invention
[0004] Object of the Invention: To solve the problems mentioned in the background technique, the present invention discloses a method for identifying illegal parking on blind paths based on object detection. Based on the improved YOLOv11 and GRFB_UNet, a network model is constructed. While reducing the number of model parameters and computational overhead, it can monitor the blind path area in real time, automatically detect and record illegally parked vehicles, and at the same time distinguish whether the vehicle passes by briefly or parks illegally for a long time.
[0005] Technical Solution:
[0006] The present invention discloses a method for identifying illegal parking on blind paths based on object detection, and the method includes the following steps:
[0007] Step 1: Collect image data of parked vehicles and blind paths on the street, classify them using a python program, make a road vehicle and blind path dataset, and shoot road videos;
[0008] Step 2: Improve YOLOv11 and combine GRFB_UNet to construct a blind path illegal parking recognition model;
[0009] Step 2.1: Introduce the ELA attention mechanism into the parallel channel attention branch, and fuse channel + spatial attention through three learnable weights, and improve it to AELA and add it after each C3k2 in the neck network of YOLOv11;
[0010] Step 2.2: Design a multi-scale convolutional adaptive aggregation feature module AMFA and add it to the 11th layer of the backbone network;
[0011] Step 3: Input the training set image data in the road vehicle and blind path datasets into the improved YOLOv11 model in Step 2 for training and obtain the optimal model parameters. Input the road video to generate a road vehicle-free background map based on the background extraction algorithm that fuses the inter-frame difference method and the Gaussian mixture model. Generate a mask segmentation map using GRFB_UNet. The improved YOLOv11 model detects vehicles in real time and generates prediction boxes. Map the YOLO prediction boxes to the current frame segmentation mask, calculate the blind path occlusion ratio to determine illegal parking;
[0012] Step 4: Use ByteTrack for tracking trajectories, use a timer to handle the judgment of temporarily passing through the blind path and illegally parking on the blind path. When the timer reaches the limit, save the evidence of illegal parking and upload it to the system, and use CNOCR to identify the license plate of the illegally parked target.
[0013] Furthermore, the ELA introduced in Step 2.1 introduces the spatial pyramid pooling technology, which can capture the feature information of the target at different scales:
[0014] y h = σ(G n (F h (z h )))
[0015] y w = σ(G n (F w (z w )))
[0016] where σ represents the non-linear activation function, F h and F w represent 1D convolutions, the convolution kernel size is set to 7, and the position attention representations in the horizontal and vertical directions are represented by y h and y w respectively, z h and z w are to perform global average pooling on the input feature map in the height and width directions, compress the feature vector, perform a 1D convolution operation on the compressed feature vector, and perform normalization; apply the Sigmoid function to generate the attention weight;
[0017] Multiply the generated attention weight by the input feature map:
[0018] Y = x c × y h × y w
[0019] Wherein, x c is the input feature map.
[0020] Furthermore, the improved AELA described in step 2.1 specifically includes: adding a channel attention branch in parallel, adopting dynamic calculation of the convolution kernel size, using dilated convolution to replace and expand the receptive field, using adaptive grouping to replace fixed grouping, introducing three learnable weight parameters, using the weight parameters to fuse two kinds of attention. In the initialization stage, the dynamic convolution kernel calculates the kernel_size according to the number of channels and ensures it is odd. Using the channel attention mechanism, after global average pooling to compress the spatial dimension, grouped convolution maintains channel independence, and Sigmoid activation generates the channel attention map; in the forward propagation stage, the input x is processed through two branches of channel and space. The channel attention is generated by average pooling and 1D convolution, and the spatial attention is processed separately in the horizontal and vertical directions. Group normalization is implemented using dilated convolution and the adaptive channel grouping strategy, and finally merged. The three weight parameters are adjusted by Sigmoid, and the spatial attention and channel attention are fused and connected to the feature map with a residual connection.
[0021] Furthermore, the specific structure of the AMFA module is as follows:
[0022] Let the module input tensor be [B, C, H, W], where B represents the batch size, C represents the number of channels, and H and W represent the height and width of the feature map; the module first saves the original tensor and obtains local information through a depth convolution of 5×5 according to the branches, resulting in a tensor [B, C, H, W]. Design multi-branch depthwise separable strip convolutions with specific sizes of 5×1, 1×5, 7×1, 1×7, 11×1, 1×11, 21×1, and 1×21. Through multiple branches, different feature representations can be obtained. Given the input [B, C, H, W], after passing through the multi-branch strip convolutions, an output tensor [B, C, H, W] is obtained. Global average pooling is performed on the four outputs to obtain the global spatial features for each channel, compressed to 1*1*C, i.e., [B, C, 1, 1]. A 1x1 convolutional kernel is used to perform a convolution operation on the reshaped feature vector to generate the weights for each channel, resulting in an output [B, C, 1, 1]; the four weights are concatenated to form [B, C, 4, 1], and then the self-gating Sigmoid function is used to obtain the weight representation of itself, still maintaining [B, C, 4, 1]. Then, the Softmax function is used to reallocate the weights, obtaining [B, C, 4, 1]. Finally, the importance of the corresponding input channels is adjusted by the weights, multiplied by the weights to obtain a tensor [B, C, H, W], and after adding and fusing multi-scale features, [B, C, H, W] is obtained. Finally, a 1×1 convolution is used to model the relationships between different channels, maintaining [B, C, H, W], and the attention is multiplied by the original input, with the output tensor being [B, C, H, W].
[0023] Furthermore, the blind sidewalk illegal parking recognition model uses SCDown to replace the conventional YOLOv11Conv downsampling for a lightweight network; it uses the NWD loss function to replace the original CIOU Loss.
[0024] Furthermore, the specific tracking and recognition process in step 4 is as follows:
[0025] The multi-object tracking ByteTrack algorithm is used to track the detected vehicle bounding boxes, and a unique tracking ID is assigned to each vehicle. When a vehicle is first determined to be illegally parked, its start time and the initial detection bounding box are recorded, and the timer starts. To avoid repeatedly refreshing the timer during the tracking process and ensure that the timer only starts for the first determined illegal parking behavior, in each subsequent frame, the intersection over union (IoU) between the current detection bounding box and the initial detection bounding box is calculated. If the IoU value decreases significantly, it indicates that the vehicle has moved a large distance, so it is determined that the vehicle is not parked for a long time, and the timer is reset; on the contrary, if the IoU value remains high, it means that the vehicle's position has basically not changed and it is in a parked state. In each frame, it is checked whether the parking time of the vehicle exceeds the set illegal parking time threshold. Once the threshold is reached, that frame is saved as evidence of illegal parking. Subsequently, the license plate of the illegally parked vehicle is detected, and CNOCR recognition is performed to obtain the license plate information.
[0026] Beneficial effects:
[0027] 1. The present invention designs a multi-scale convolutional adaptive aggregation feature module AMFA, which controls the importance of different channels through a Sigmoid gating unit and fuses multi-scale feature information. The EMA-SlideLoss is used as the classification loss function to enhance the classification ability of the model. The NWD loss function is used as the bounding box regression loss function to improve the small target detection ability. The SCDown downsampling module is introduced to reduce the number of model parameters and computational overhead. A multi-scale convolutional adaptive aggregation feature (AMFA) module is designed, and the Sigmoid function of the gating unit is used to obtain weights to control the importance of different channels, generate weights, and fuse the information of different scale features; the ELA attention mechanism is improved, and while maintaining the lightweight characteristics, the model performance is significantly improved. These improvements enhance the model's ability to extract key features and learn difficult-to-classify samples, achieving higher accuracy and recall rates and fewer parameters.
[0028] 2. The present invention uses the ByteTrack algorithm to track the detected vehicles and assigns a unique tracking ID to each vehicle. A parking violation determination mechanism is designed to judge whether a vehicle has parked for a long time by calculating the intersection over union (IoU) of the detection boxes. The function of timing the parking violation time is realized, and evidence is saved when the parking violation time threshold is reached. License plate detection and OCR recognition are performed on the parked vehicles with violations to obtain license plate information.
[0029] 3. The present invention uses GRFB_UNet to generate a background mask segmentation map: a mixture Gaussian model combined with the frame difference method is used to generate a background map, GRFB_UNet segments and identifies the blind sidewalk pixels, combines with the detection boxes of YOLOv11, and calculates the occlusion ratio of the blind sidewalk to realize the real-time detection of blind sidewalk occlusion. Description of the drawings
[0030] Figure 1 It is the flowchart of the method of the present invention;
[0031] Figure 2 It is the architecture diagram of the improved YOLOv11 network model of the present invention;
[0032] Figure 3 It is the principle structure diagram of the improved AELA attention module of the present invention;
[0033] Figure 4 It is the custom AMFA structure diagram of the present invention;
[0034] Figure 5 It is the effect diagram of the test and segmentation mask map of the embodiment of the present invention;
[0035] Figure 6 This is the overall algorithm architecture diagram of the present invention; Specific implementation manners
[0036] In order to make the objectives, features, and advantages of the present invention clearer, the present invention will be clearly and completely described below in conjunction with the accompanying drawings and specific implementation manners. In addition, the specific embodiments described in the present invention are only used to explain the present invention and are not used to limit the present invention.
[0037] As Figure 1 shown, the present invention discloses a method for identifying illegal parking on a blind path based on object detection, and the steps are as follows:
[0038] Step 1: Collect fine-grained image data, classify it using a Python program, and make a data set;
[0039] In this embodiment, image data of parked vehicles and blind paths on the street is collected by collecting open-source data of PaddlePaddle. Through the use of the X-AnyLabeling annotation tool for annotation, it is divided into ambulance, bicycle, bicycle group, bus, car, motorcycle, motorcycle group, motorist, police car, tricycle, truck, xiaofang. The Python program divides the pictures into a training set: validation set: test set according to 8:1:1, generates annotation files and mask files corresponding to the pictures, which can be used for model training, and shoots road videos.
[0040] Step 2: Build an improved YOLOv11 and GRFB-UNet network model, and the improved structure is as Figure 2 shown,
[0041] Step 2.1: Replace the EMASlideLoss classification loss function;
[0042] The EMASlideLoss is used to replace the traditional CIoU loss function. By introducing the exponential moving average (EMA) mechanism to smooth the change of the loss value, the stability and generalization ability of training are improved. Targets of different scales may contribute differently to the loss at different training stages. The SlideLoss smoothed by EMA can balance these contributions and improve the performance of the model at different scales. By weighted averaging historical values, higher weights are given to recent data to smooth fluctuations (such as loss and gradient) during training and reduce the influence of noise.
[0043] Step 2.2: Improve and add the attention module AELA mechanism to the neck network;
[0044] The improved AELA structure is as Figure 3As shown, after each C3k2 in the neck network, ELA uses strip pooling in the spatial dimension to obtain feature vectors in the horizontal and vertical directions, maintaining a narrow kernel shape to capture long-range dependencies and prevent irrelevant regions from affecting label prediction, thereby generating rich target location features in their respective directions. ELA independently processes the feature vectors in each of the above directions to obtain attention predictions, and then combines them using a product operation to ensure accurate position information of the region of interest.
[0045] Improved to AELA: In addition to spatial attention, a channel attention branch is added in parallel, and the dynamic calculation of the convolution kernel size is adopted. Dilated convolution is used to replace and expand the receptive field, and adaptive grouping is used to replace fixed grouping. Three learnable weight parameters are also introduced to fuse the two types of attention with the weight parameters:
[0046] In the initialization stage, the dynamic convolution kernel calculates the kernel_size according to the number of channels and ensures it is odd. Using the channel attention mechanism, after global average pooling compresses the spatial dimension, grouped convolution maintains channel independence, and Sigmoid activation generates the channel attention map; in the forward propagation stage, the input x is processed through both the channel and spatial branches. Channel attention is generated through average pooling and 1D convolution, and spatial attention is processed separately in the horizontal and vertical directions. Group normalization is implemented using dilated convolution and the adaptive channel grouping strategy, and finally merged. Then the three weight parameters are adjusted by Sigmoid to fuse the two types of attention, plus a residual connection.
[0047] Att ch =σ(Conv1D group (GAP(X)))
[0048] where GAP represents global average pooling, Conv1D group represents 1D grouped convolution, and σ represents the activation function;
[0049] Att h =σ(GN(Conv1D dilated (Avg W (X)))
[0050] Att w =σ(GN(Conv1D dilated (Avg H (X)))
[0051] where GAP represents global average pooling, Conv1D group represents 1D grouped convolution, and σ represents the activation function; Avg W and Avg HPerforms global average pooling on the input feature map in the height and width directions, Conv1D dilated represents 1D grouped dilated convolution, GN represents group normalization, and σ represents the activation function;
[0052]
[0053] Among them, multiply and fuse the obtained attention weights above.
[0054]
[0055] Among them, α, γ, and β represent using three weight parameters to fuse the spatial attention and channel attention mechanisms and perform a residual connection with the original feature map X.
[0056] Step 2.3: Design a multi-scale convolutional adaptive aggregation feature module AMFA and apply it to the 11th layer of the backbone network;
[0057] The module first obtains local information through depth convolution. To improve the detection ability for targets of different sizes, the multi-branch different-scale convolution structure increases the feature richness, expands the local perception ability of the network, and enhances the semantic information expression of small targets. As shown in Figure 4, assuming the module input tensor is [B, C, H, W], where B represents the batch size, C represents the number of channels, and H and W represent the height and width of the feature map; the module first saves the original tensor and obtains local information through depth convolution with a kernel size of 5×5 according to the branches, getting the tensor [B, C, H, W]. A multi-branch depthwise separable convolution is designed, specifically 5×1, 1×5, 7×1, 1×7, 11×1, 1×11, 21×1, and 1×21. The reason for choosing depthwise separable convolution is that they are lightweight and dilated convolution is added to expand the receptive field, helping the model consider the surrounding context information when identifying small targets. Through multiple branches, different feature representations can be obtained. The input is [B, C, H, W], and after multi-branch separable convolution, the output is a [B, C, H, W] tensor. Global average pooling is performed on the four outputs to obtain the global spatial features of each channel, compressed to 1*1*C, that is, [B, C, 1, 1]. Similar to the operation of channel attention, the pooled features are reshaped into a one-dimensional vector to generate a single value for each channel. A 1x1 convolutional kernel is used to perform convolution on the reshaped feature vector to generate the weights for each channel, getting the output [B, C, 1, 1]; the four weights are concatenated into [B, C, 4, 1], and then the self-gated Sigmoid function is used to obtain the self-weight representation, still maintaining [B, C, 4, 1]. Then, the Softmax function is used to redistribute the weights, getting [B, C, 4, 1]. Finally, the importance of the corresponding input channels is adjusted by the weights, multiplied by the weights to get the tensor [B, C, H, W]. After adding and fusing multi-scale features, [B, C, H, W] is obtained. Finally, 1×1 convolution is used to model the relationship between different channels, maintaining [B, C, H, W], and the attention is multiplied by the original input, and the output tensor is [B, C, H, W].
[0058] Step 2.4: Replace the Conv downsampling with SCDown to lightweight the network;
[0059] The SCDown module is superior to Conv in terms of both the number of parameters and computational complexity. The SCDown module consists of two main convolutional layers, designed to effectively reduce the spatial and channel dimensions of the feature map. The first convolutional layer uses a 1x1 convolutional kernel to compress the number of channels of the input feature map from c1 to c2. This process not only reduces the complexity of subsequent calculations but also allows the model to focus on more critical feature information. The second convolutional layer further processes the feature map with compressed channels. This layer uses a k, d, k convolutional kernel and a stride s and implements channel-wise step convolution to improve computational efficiency. By adjusting the spatial dimensions, the feature map can be effectively downsampled, enhancing the model's ability to capture information at different scales.
[0060] Step 2.5: Replace the original CIOU Loss with the NWD loss function;
[0061] NWD is a loss function specifically designed for tiny object detection, and it demonstrates significant advantages when dealing with small object detection tasks. IOU is calculated based on the intersection over union ratio and is insensitive to the size of the object. This means that when dealing with small objects, even a slight positional deviation between the predicted box and the ground truth box may lead to a significant drop in the IoU value, thus affecting the optimization of the model. The NWD loss function overcomes the limitations of the traditional IoU (Intersection over Union) loss function when dealing with small objects, especially the sensitivity to positional deviation and the scale invariance problem, by introducing the concept of the Wasserstein distance.
[0062] Step 2.6: Use GRFB-UNet for construction;
[0063] GRFB-UNet not only inherits the powerful encoder-decoder structure of the UNet network and effectively preserves low-level detailed information through skip connections, thus ensuring the accuracy and integrity of the segmentation results. At the same time, it further improves the computational efficiency and generalization ability of the model through optimized algorithms and parameter settings. GRFB-UNet significantly reduces the computational complexity while maintaining high accuracy, is suitable for running on resource-constrained devices, has a fast response, and can immediately obtain the segmentation result of the blind path, and determines whether it is illegally parked on the blind path based on the detection result.
[0064] Step 3: Input the training set image data in the dataset into the network model in Step 2 for training, save the model parameters with the highest accuracy on the validation set during the training process, and name its file best.pth; The experimental environment configuration is to select the Ubuntu20.04 operating system, use an RTX 2080Ti (11GB) graphics card, and the deep learning framework is PyTorch. The specific configuration is shown in Table 1 below:
[0065] Table 1
[0066]
[0067] Input the training set image data in the roadside parked vehicle dataset into YOLOv11, train using GPU, set the parameters of the improved network model, with the number of training rounds being 200 rounds, the momentum being 0.937, the initial learning rate being 0.01, the weight decay coefficient being 0.0005, and adopt the SGD optimizer. During the model training process, the model parameters with the best accuracy will be saved and named best.pt.
[0068] Input the training set image data in the blind sidewalk dataset into GRFB_UNet, train using GPU, set the parameters of the improved network model, with the number of training rounds being 1200 rounds, the momentum being 0.9, the initial learning rate being 0.02, the weight decay coefficient being 0.0001, and adopt the SGD optimizer. During the model training process, the model parameters with the best accuracy will be saved and named model_best.pth.
[0069] Step 4: Background and segmentation mask map extraction. For each fixed camera, take pictures of the vehicle-free street background in advance, or process the road video using the mixture Gaussian model combined with the frame difference method to generate a vehicle-free street background map, and perform semantic segmentation once through GRFB-UNet to obtain the background mask segmentation map. In the real-time part, load the optimal weight file best.pt into the improved YOLOv11 model to perform real-time vehicle detection on the current video frame. After obtaining the detection boxes of the vehicles, for special vehicles such as ambulances, police cars, and fire trucks, since special vehicles temporarily occupy the blind sidewalk for emergency tasks, no determination of illegal parking on the blind sidewalk will be made; input the current video frame into the GRFB_UNet model to obtain the segmentation mask map of the current frame. Map the vehicle prediction box coordinates to the current frame segmentation mask map and the background mask segmentation map, and calculate the number of blind sidewalk pixel points in the current frame segmentation mask map and the number of blind sidewalk pixel points in the background mask segmentation map to obtain the occlusion ratio of the blind sidewalk. The test results are as Figure 5 shown.
[0070] The ByteTrack algorithm for multi-object tracking is used to track the detected vehicle bounding boxes, and a unique tracking ID is assigned to each vehicle. When a vehicle is first determined to be illegally parked, its start time and the initial detection bounding box are recorded, and a timer is started simultaneously. In each subsequent frame, the intersection over union (IoU) between the current detection bounding box and the initial detection bounding box is calculated. If the IoU value is low, it indicates that the vehicle has moved significantly and is not parked for a long time, and the timer is reset. If the IoU value remains high, it means that the vehicle is in a parked state, and it is checked whether the parking time exceeds the set illegal parking time threshold. Once the threshold is reached, the frame is saved as evidence of illegal parking. The license plate of the vehicle in the illegal parking evidence frame is detected. The CNOCR algorithm is used to recognize the detected license plate to obtain the license plate information. The overall algorithm architecture of the present invention is as Figure 6 shown.
[0071] To detect the performance of the improved model as an evaluation index. Precision is the most intuitive evaluation index in classification problems, indicating the proportion of samples that are truly positive among the samples correctly classified by the model; Recall represents the ratio of the number of samples correctly identified as positive by the model to the total number of positive samples. The higher the Recall, the more positive samples are predicted correctly by the model, and the better the model's performance. mAP measures the goodness or badness across all types. mAP takes the average of the APs of all categories. mAP 50 : mAP when the IoU threshold is 0.5. mAP 50:95 This form is the mAP at multiple IoU thresholds. It will take 10 IoU thresholds within the q interval [0.5, 0.95] with a step size of 0.05, calculate the mAP at these 10 IoU thresholds respectively, and then take the average. mAP 50:95 The larger it is, the more accurate the predicted bounding box is because it captures more cases with large IoU thresholds.
[0072]
[0073]
[0074] In the above formula, TP represents the number of samples correctly predicted as positive samples in the sample, FP represents the number of negative samples wrongly predicted as positive samples in the sample, and FN is the number of samples predicted as negative examples but actually positive examples. N represents the total number of categories. AP i is the AP0.5 of the i-th category.
[0075] The comparison table of algorithm metrics is shown in Table 2, which proves the feasibility of the improved algorithm of the present invention.
[0076] Table 2
[0077]
[0078] As can be seen from Table 2, the improved algorithm is designed and applied to vehicle detection. By collecting image data in specific scenarios (such as street monitoring), these images are input into an optimized detection model. After loading the optimal weight file best.pth into the model for inference to obtain the targets in the images to be detected, the improved algorithm has an increase in mAP compared to v10n and v11n, and there is also an effect on partial lightweighting, with a partial decrease in the number of parameters and GFLOPS.
[0079] A blind sidewalk illegal parking recognition algorithm method based on object detection proposed by the present invention combines the advantages of YOLOv11 in the field of real-time detection and the powerful semantic segmentation ability of the GRFB-UNet model, providing an efficient and accurate solution for the intelligent monitoring of blind sidewalk illegal parking problems; when there is no background map provided, the ratio of the number of pixels of the blind sidewalk to the sidewalk is used for judgment, and the accuracy is slightly lower than that with a background segmentation mask map; as Figure 5 shown, the test images are used for vehicle detection using the YOLOv11 test set for training, and the illegal parking is judged corresponding to the real-time segmentation mask, and the ratio of the number of pixels of the blind sidewalk to the sidewalk within the box is used for judgment.
[0080] The above embodiments are an implementation manner of the present invention, but the implementation manner of the present invention is not limited thereto. Any modifications, substitutions, and improvements made by those skilled in the art without departing from the principle and spirit of the present invention are included in the protection scope of the present invention.
Claims
1. A method for identifying illegal parking on a blind path based on object detection, characterized in that, The method includes the following steps: Step 1: Collect image data of parked vehicles and blind paths on the street, classify them using a Python program, make a road vehicle and blind path dataset, and shoot a road video; Step 2: Improve YOLOv11 and construct a blind path illegal parking recognition model by combining GRFB_UNet; Step 2.1: Introduce the ELA attention mechanism into the parallel channel attention branch, and fuse channel + spatial attention through three learnable weights, and improve it to AELA and add it after each C3k2 in the neck network of YOLOv11; Step 2.2: Design a multi-scale convolutional adaptive aggregation feature module AMFA and add it to the 11th layer of the backbone network; Step 3: Input the training set image data in the road vehicle and blind path dataset into the improved YOLOv11 model in Step 2 for training and obtain the optimal model parameters. Input the road video to generate a road vehicle-free background map based on the background extraction algorithm that combines frame difference method and Gaussian mixture model. GRFB_UNet generates a mask segmentation map. The improved YOLOv11 model detects vehicles in real time and generates prediction boxes. Map the YOLO prediction boxes to the current frame segmentation mask, calculate the blind path occlusion ratio to determine illegal parking; Step 4: Use ByteTrack to track the trajectory, use a timer to process the judgment of temporarily passing through the blind path and illegally parking on the blind path. When the timer reaches the limit, save the illegal parking evidence and upload it to the system, and use CNOCR to recognize the license plate of the illegal parking target.
2. The method for identifying illegal parking on a blind path based on object detection according to claim 1, wherein, The ELA introduced in Step 2.1 introduces the spatial pyramid pooling technology, which can capture the feature information of the target at different scales: y h = σ(G n (F h (z h ))) y w = σ(G n (F w (z w ))) where, σ represents the non-linear activation function, F h and F w represent 1D convolution, the convolution kernel size is set to 7, and the position attention representations in the horizontal and vertical directions are represented by y h and y w respectively, z h and z w perform global average pooling on the input feature map in the height and width directions, compress the feature vectors, perform 1D convolution operations on the compressed feature vectors, and perform normalization; apply the Sigmoid function to generate the attention weights; Multiply the generated attention weight by the input feature map: Y = x c × y h × y w Among them, x c is the input feature map.
3. The blind sidewalk illegal parking recognition method based on object detection according to claim 1 or 2, characterized in that, The specific improvement of AELA in Step 2.1 specifically includes: adding a parallel channel attention branch, dynamically calculating the convolution kernel size, using dilated convolution to replace and expand the receptive field, using adaptive grouping to replace fixed grouping, introducing three learnable weight parameters, using the weight parameters to fuse the two kinds of attention. In the initialization stage, the dynamic convolution kernel calculates the kernel_size according to the number of channels and ensures it is odd. Use the channel attention mechanism. After using global average pooling to compress the spatial dimension, grouped convolution maintains channel independence, and Sigmoid activation generates the channel attention map; in the forward propagation stage, the input x is processed through two branches of channel and space. The channel attention is generated through average pooling and 1D convolution, and the spatial attention is processed in the horizontal and vertical directions respectively. Use dilated convolution and adaptive channel grouping strategy to achieve group normalization, and finally merge. The three weight parameters are adjusted by Sigmoid, and the spatial attention and channel attention are fused and connected to the feature map with residual.
4. The method for identifying illegal parking on a blind path based on object detection according to claim 1, characterized in that, The specific structure of the AMFA module is as follows: Let the module input tensor be [B, C, H, W], where B represents the batch size, C represents the number of channels, and H and W represent the height and width of the feature map. The module first saves the original tensor and obtains local information through a depth convolution of 5×5 according to the branches, getting the tensor [B, C, H, W]. Design multi-branch depthwise separable strip convolutions with specific sizes of 5×1, 1×5, 7×1, 1×7, 11×1, 1×11, 21×1, and 1×21. Through multiple branches, different feature representations can be obtained. Input [B, C, H, W], after multi-branch strip convolutions, output a [B, C, H, W] tensor. Perform global average pooling on the four to obtain the global spatial features of each channel, compressed to 1*1*C, that is, [B, C, 1, 1]. Use a 1x1 convolutional kernel to perform a convolution operation on the reshaped feature vector to generate the weights of each channel, getting the output [B, C, 1, 1]. Concatenate the four weights into [B, C, 4, 1], then obtain the weight representation of itself through its own gated Sigmoid function, still maintaining [B, C, 4, 1]. Then reallocate the weights through the Softmax function to get [B, C, 4, 1]. Finally, adjust the importance of the corresponding input channels through the weights, multiply by the weights to get the tensor [B, C, H, W], perform addition to fuse multi-scale features, and get [B, C, H, W]. Finally, use a 1×1 convolution to model the relationship between different channels, maintaining [B, C, H, W], multiply the attention by the original input, and the output tensor is [B, C, H, W].
5. The method for identifying illegal parking on a blind path based on object detection according to claim 1, wherein The blind sidewalk illegal parking recognition model uses SCDown to replace the conventional YOLOv11Conv downsampling for a lightweight network; uses the NWD loss function to replace the original CIOU Loss.
6. The method for identifying illegal parking on a blind path based on object detection according to claim 1, wherein, The specific tracking and recognition process in step 4 is as follows: Use the multi-object tracking ByteTrack algorithm to track the detected vehicle bounding boxes and assign a unique tracking ID to each vehicle. When a vehicle is first determined to be illegally parked, record its start time and the initial detection bounding box, and start the timer. To avoid repeatedly refreshing the timer during the tracking process and ensure that the timer only starts for the first determined illegal parking behavior, in each subsequent frame, calculate the intersection over union (IoU) between the current detection bounding box and the initial detection bounding box. If the IoU value decreases significantly, it indicates that the vehicle has moved a large distance, so it is judged that the vehicle is not parked for a long time and the timer is reset; on the contrary, if the IoU value continues to be high, it means that the vehicle's position has basically not changed and it is in a parked state. Check whether the parking time of the vehicle exceeds the set illegal parking time threshold in each frame. Once the threshold is reached, save the frame as evidence of illegal parking. Subsequently, detect the license plate of the illegally parked vehicle and perform CNOCR recognition to obtain the license plate information.
Citation Information
Cited By
Multi-scale road vehicle detection method based on re-parameterized visual converter
CN120564164A
Single-plant-scale tree positioning and identifying method
CN120747757A
Idle parking space detection method based on maximum matching and parking space segmentation
CN120808242A
Portable intelligent device real-time braille point detection method and system for visually impaired people
CN120823610A
Real-time Braille dot detection method and system for portable smart devices for visually impaired individuals
CN120823610B