License plate positioning detection method used in complex road scene and storage medium

The rotating target license plate detection model, which combines adaptive kernel rotation dynamic convolution and task interaction features, solves the problem of low accuracy in license plate detection in complex road environments, and achieves accurate positioning and high-precision detection of license plate direction.

CN120913192APending Publication Date: 2025-11-07UNIFORM ENTROPY TECH (WUXI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511022718.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing rotating target detection methods struggle to accurately distinguish the true direction of license plates in complex road environments, resulting in low license plate detection accuracy.

Method used

A pre-trained rotating target license plate detection model is adopted, which uses adaptive kernel rotation dynamic convolution to extract feature information of license plate images, and performs feature fusion through path aggregation network and feature pyramid network. Combined with task interaction joint features, target detection is performed to improve prediction accuracy.

Benefits of technology

It effectively distinguishes the true direction of license plates, improving the accuracy of license plate detection and prediction precision in complex road environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913192A_ABST
    Figure CN120913192A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of license plate positioning, and particularly discloses a license plate positioning detection method used in a complex road scene and a storage medium, and the method comprises the steps: obtaining the license plate image data information of a driving motor vehicle; inputting the license plate image data information into a pre-trained rotating target license plate detection model for license plate positioning detection; the pre-trained rotating target license plate detection model at least comprises a feature extraction module which is used for carrying out feature extraction on license plate image data information according to adaptive kernel rotating dynamic convolution to obtain to-be-detected license plate features with space structure and direction information; the feature fusion module is used for performing feature fusion on the to-be-detected license plate features based on the path aggregation network and the feature pyramid network to obtain a to-be-detected license plate feature fusion result with multi-scale features; and the detection head prediction module is used for performing task interaction joint feature-based target detection on the to-be-detected license plate feature fusion result to obtain a license plate positioning detection result.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of license plate positioning, and particularly relates to a license plate positioning and detection method for a complex road scene, a computer storage medium and an electronic device. BACKGROUND

[0002] At present, the license plate of a motor vehicle driving on a road is positioned and detected mostly by using a rotating target detection method in the field of computer vision. The rotating target detection method is essentially different from a traditional target detection method. For example, the traditional target detection method assumes that the position of a license plate to be detected in an image is completely aligned with an image coordinate axis, while the rotating target detection optimizes this inaccurate assumption. In the natural world, many targets do not follow this simple rule and are often placed in the picture at unexpected angles, showing a twisted, tilted or even rotated shape. Therefore, the rotating target detection is essentially different from the traditional target detection, and is thus widely applied.

[0003] However, the current rotating target detection method still has many problems when detecting a license plate in a complex road environment, although it has made progress compared with the traditional target detection method. Specifically, firstly, the direction of a license plate of a motor vehicle may change with a viewing angle, light or shooting angle, resulting in different orientations in different images. This requires a detection algorithm to not only capture the appearance information of a target, but also to be able to distinguish and identify its real direction. However, for the rotating target detection task, the existing standard backbone model may not meet the demand in performance. Secondly, in the current license plate detection, the rotating target detection usually adopts a multi-task learning strategy to solve the problem. This method combines two seemingly independent tasks of classification and positioning, and optimizes them simultaneously through a deep learning model. The classification task is responsible for extracting high-level features for classification from an image, while the positioning task focuses more on finding the exact position of each object in the image. However, since the rotating target detection task involves the processing of spatial information, the features learned by the two tasks may differ in spatial distribution. When the two separate network structures are combined for prediction, the prediction result may be misaligned, thereby affecting the accuracy of the license plate detection result.

[0004] Therefore, how to distinguish the real direction of the current rotating target detection task in a complex scene and improve the prediction accuracy so as to improve the detection accuracy of the license plate detection in a complex road environment has become a technical problem to be solved by those skilled in the art. SUMMARY

[0005] The application provides a license plate positioning detection method for a complex road scene, a computer storage medium and an electronic device, and solves the problem of low license plate detection accuracy caused by the inability to distinguish the real direction of the license plate in a complex scene in the related art.

[0006] As a first aspect of the application, a license plate positioning detection method for a complex road scene is provided, comprising:

[0007] Obtaining license plate image data information of a running motor vehicle;

[0008] Inputting the license plate image data information into a pre-trained rotating target license plate detection model for license plate positioning detection, and obtaining a license plate positioning detection result;

[0009] Displaying the license plate positioning detection result to a user;

[0010] The pre-trained rotating target license plate detection model at least includes a feature extraction module, a feature fusion module and a detection head prediction module; the feature extraction module is used for feature extraction of the license plate image data information according to adaptive kernel rotation dynamic convolution, to obtain license plate features to be detected with spatial structure and direction information; the feature fusion module is used for feature fusion of the license plate features to be detected based on a path aggregation network and a feature pyramid network, to obtain license plate feature fusion results with multi-scale features; and the detection head prediction module is used for target detection of the license plate feature fusion results based on task interaction joint features, to obtain a license plate positioning detection result.

[0011] Further, the feature extraction module is used for feature extraction of the license plate image data information according to adaptive kernel rotation dynamic convolution, to obtain license plate features to be detected with spatial structure and direction information, comprising:

[0012] Dynamic prediction is performed on the license plate image data information, to obtain a rotation angle and a combination weight of each convolution kernel capable of representing the license plate features to be detected;

[0013] Adaptive rotation is performed on the convolution kernel according to the rotation angle of each convolution kernel, to obtain a convolution kernel capable of extracting spatial structure and direction information in the license plate image data information;

[0014] Combination calculation is performed on a plurality of convolution kernels with spatial structure and direction information according to the combination weight, to obtain the license plate features to be detected with spatial structure and direction information.

[0015] Further, adaptive rotation is performed on the convolution kernel according to the rotation angle of each convolution kernel, to obtain a convolution kernel capable of extracting spatial structure and direction information in the license plate image data information, comprising:

[0016] The weight parameters of the original convolution kernel are taken as sampling points, and an original kernel coordinate space is constructed by a bilinear interpolation up-sampling method;

[0017] The original kernel coordinate space is rotated around a center point to obtain a new kernel coordinate space;

[0018] In the new kernel coordinate space, the values in the original kernel coordinate space are down-sampled to obtain a convolution kernel capable of extracting spatial structure and directional information in the license plate image data.

[0019] Further, the detection head prediction module is configured to perform target detection on the to-be-detected license plate feature fusion result based on task interaction joint features, including:

[0020] The to-be-detected license plate feature fusion result is input into a shared convolution to generate task interaction joint features;

[0021] According to the task interaction joint features, dynamic decomposition of tasks is performed;

[0022] According to the task interaction joint features, the weight parameters of each task after dynamic decomposition are determined;

[0023] According to the weight parameters of each task after dynamic decomposition, target prediction is performed.

[0024] Further, the dynamic decomposition of tasks according to the task interaction joint features includes:

[0025] The task interaction joint features and their global average pooling features are respectively used for dynamic decomposition of classification tasks and regression tasks to obtain classification features and regression features, wherein the global average pooling features of the task interaction joint features are used to dynamically generate kernel attention weights.

[0026] Further, the target prediction according to the weight parameters of each task after dynamic decomposition includes:

[0027] According to the feature selection weight, feature attention is performed on the classification features to obtain a classification prediction result;

[0028] According to the fusion feature attention, feature attention is performed on the regression features to obtain a regression prediction result;

[0029] According to the classification prediction result and the regression prediction result, task alignment is realized to obtain a license plate positioning detection result.

[0030] Further, the pre-trained rotating target license plate detection model further includes a data preprocessing module configured to perform normalization and data enhancement processing on the license plate image data information.

[0031] Further, the pre-trained rotating target license plate detection model further comprises a post-processing module, which is configured to perform a non-maximum suppression post-processing operation on the license plate positioning detection result, so as to screen a detection frame meeting a preset standard.

[0032] As another aspect of the present application, a computer storage medium is provided, wherein a computer program is stored, and the computer program is executed by a processor to implement the license plate positioning detection method for complex road scenes as described above.

[0033] As another aspect of the present application, an electronic device is provided, comprising a memory and a processor, wherein the processor is in communication connection with the memory, the memory is configured to store a computer program, and the processor is configured to load and execute the computer program to implement the license plate positioning detection method for complex road scenes as described above.

[0034] The license plate positioning detection method for complex road scenes provided by the present application can distinguish the real direction of the current rotating target detection task in the complex scene and improve the result prediction accuracy, thereby improving the detection accuracy of license plate detection in complex road environments. BRIEF DESCRIPTION OF DRAWINGS

[0035] The accompanying drawings are included to provide a further understanding of the present application, and constitute a part of the specification, and are used together with the following detailed description to explain the present application, but do not constitute a limitation of the present application.

[0036] Figure 1 The flowchart of the license plate positioning detection method for complex road scenes provided by the present application.

[0037] Figure 2 The flowchart of the feature extraction provided by the present application.

[0038] Figure 3 The flowchart of the target detection provided by the present application.

[0039] Figure 4 The structural block diagram of the electronic device provided by the present application. DETAILED DESCRIPTION

[0040] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0041] In order for those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.

[0042] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0043] In the present embodiment, a license plate positioning detection method for complex road scenes is provided, Figure 1 is a flow chart of the license plate positioning detection method for complex road scenes provided according to the embodiments of the present application, as Figure 1 shown, comprising:

[0044] S100, acquiring license plate image data information of a running motor vehicle;

[0045] In the embodiments of the present application, the license plate image data information of the motor vehicle running on the traffic road can be specifically collected, which can be a photo or a video.

[0046] S200, inputting the license plate image data information into a pre-trained rotating target license plate detection model for license plate positioning detection, and obtaining a license plate positioning detection result;

[0047] Specifically, the collected license plate image data information is input into a pre-trained rotating target license plate detection model for license plate detection. Specifically, the pre-trained rotating target license plate detection model at least includes a feature extraction module, a feature fusion module and a detection head prediction module; the feature extraction module is used to extract features of the license plate image data information according to adaptive kernel rotation dynamic convolution, and obtain license plate features to be detected with spatial structure and direction information; the feature fusion module is used to fuse the license plate features to be detected based on a path aggregation network and a feature pyramid network, and obtain a license plate feature fusion result with multi-scale features; and the detection head prediction module is used to perform target detection on the license plate feature fusion result based on task interaction joint features, and obtain a license plate positioning detection result.

[0048] It should be understood that, since the feature extraction module in the pre-trained rotating target license plate detection model extracts features of the license plate image data information through adaptive kernel rotation dynamic convolution, the feature information of the rotating target in the license plate image can be well extracted to determine the real direction thereof, and the target detection based on task interaction joint features through the detection head prediction module can improve the prediction accuracy.

[0049] S300, displaying the license plate positioning detection result to the user;

[0050] In summary, the license plate positioning detection method for complex road scenes provided by the present application can obtain license plate image data information of a driving motor vehicle, and input the license plate image data information into a pre-trained rotating target license plate detection model for license plate detection. The feature extraction module in the pre-trained rotating target license plate detection model extracts features of the license plate image data information through adaptive kernel rotation dynamic convolution, which can well extract the feature information of the rotating target in the license plate image to determine the real direction thereof, and the target detection based on task interaction joint features through the detection head prediction module can improve the prediction accuracy. Therefore, the license plate positioning detection method for complex road scenes provided by the present application can distinguish the real direction of the current rotating target detection task in the complex scene and improve the result prediction accuracy, thereby improving the detection accuracy of license plate detection in complex road environments.

[0051] In the embodiment of the present application, before the feature extraction of the license plate image data, the pre-trained rotating target license plate detection model needs to preprocess the license plate image data information to enhance the image features. Specifically, the pre-trained rotating target license plate detection model further includes a data preprocessing module, which is used to normalize and perform data enhancement processing on the license plate image data information.

[0052] Specifically, the data preprocessing in the embodiment of the present application mainly refers to data normalization and Mosaic data enhancement. In the embodiment of the present application, data normalization is to scale the RGB values (generally integers from 0 to 255) of the license plate image data to the range of 0 to 1 in equal proportion. Using the normalized data to train the model can ensure the stability of the training process and prevent the model from appearing large amplitude jitter phenomenon in the training process.

[0053] The Mosaic data enhancement method combines multiple images into one image, thereby increasing the diversity of the data set and improving the generalization ability of the model. At the same time, the Mosaic data enhancement method can also help the model to identify targets in a smaller range and improve the performance of small target detection. Specifically, the Mosaic data enhancement includes the following steps (taking 4 images as an example):

[0054] (1) randomly select 4 images;

[0055] (2) randomly scale and randomly crop the 4 images, and then splice them together to form a new image;

[0056] (3) label the corresponding rotated rectangular bounding box information on the new image after splicing;

[0057] (4) pass the new image and the corresponding label information into the rotated target detection model for training or verification.

[0058] Based on the license plate image data information after the above data preprocessing, feature extraction is performed.

[0059] Specifically, the feature extraction module is configured to perform feature extraction on the license plate image data information according to adaptive kernel rotation dynamic convolution, to obtain features of the license plate to be detected with spatial structure and direction information, as shown in FIG. 1. Figure 2

[0060] S211, performing dynamic prediction on the license plate image data information to obtain a rotation angle and a combination weight of each convolution kernel capable of representing features of the license plate to be detected;

[0061] In the embodiment of the present application, the feature extraction module serves as the backbone feature extraction network of the pre-trained rotated target license plate detection model, and realizes feature extraction based on adaptive kernel rotation dynamic convolution (hereinafter referred to as AKRDC).

[0062] ​The AKRDC convolution breaks the mode of unified processing of all input images by using a fixed convolution kernel through a standard convolution in a traditional deep learning architecture, and each AKRDC convolution internally contains multiple groups of convolution kernels, each of which can be flexibly and dynamically adapted to rotation according to different features of the input feature map, so as to better capture the spatial structure and directional information of the target object to be detected in the image.

[0063] In the embodiment of the present application, the specific steps of AKRDC include:

[0064] (1) Dynamic prediction: AKRDC first dynamically predicts the rotation angle and combination weight of each convolution kernel from the input features in a data-dependent manner through dynamic prediction, the rotation angle is used for adaptive rotation of the subsequent convolution kernel, and the combination weight is the weighting coefficient of the subsequent combination calculation mechanism. Based on this, AKRDC can ensure that the direction of each convolution kernel in the convolution process matches the specific features of the image through dynamic prediction, thereby greatly improving the feature extraction capability of the model for the target instance to be detected.

[0065] In the embodiment of the present application, when implementing dynamic prediction, the image features are taken as input to predict the rotation angle of each convolution kernel in AKRDC and the corresponding combination weight for integrating the convolution kernel. After receiving the input image features, first, the image features are sent into a customized lightweight depth separable convolution, and then pass through LayerNorm and ReLU activation functions; then, a global average pooling operation is performed on the features to make all the features converge into a feature vector; finally, a linear layer with a Softsign activation function is used to predict the rotation angle, and a linear layer with a Sigmoid activation function is used to predict the combination weight.

[0066] Specifically, the dynamic prediction process can be implemented through a dynamic prediction module. In the embodiment of the present application, given the input feature x, the dynamic prediction module (denoted as f) will predict a group of rotation angles θ:{θ1,θ2,…,θ n} and a group of combination weights α:{α1,α2,…,α n} as shown in the following formula.

[0067] θ,α=f(x).

[0068] In order to ensure that the dynamic prediction module can produce small values in the early stage of the training process, the dynamic prediction module adopts a special initialization method, specifically, the entire dynamic prediction module is initialized with a truncated normal distribution with a mean of 0 and a standard deviation of 0.2.

[0069] S212, adaptively rotating the convolution kernel according to the rotation angle of each convolution kernel to obtain a convolution kernel capable of extracting spatial structure and directional information in the license plate image data.

[0070] In the embodiment of the present application, the AKRDC does not regard the weight of the convolution kernel as an independent parameter, but as some sampling points in the kernel coordinate space, because the weight parameter of the convolution kernel needs to be mapped to the kernel space. In order to realize the rotation of the convolution kernel, the mapping of the weight parameter of the convolution kernel across the kernel space needs to be realized, and for this purpose, the weight parameter of the original convolution kernel is mapped across the kernel space by interpolation, preferably, bilinear interpolation can be used.

[0071] In the embodiment of the present application, the convolution kernel is adaptively rotated according to the rotation angle of each convolution kernel to obtain a convolution kernel capable of extracting spatial structure and directional information in the license plate image data information, comprising:

[0072] 1) The weight parameter of the original convolution kernel is taken as a sampling point, and the original kernel coordinate space is constructed by the up-sampling method of bilinear interpolation;

[0073] 2) The original kernel coordinate space is rotated around the center point to obtain a new kernel coordinate space;

[0074] 3) The values in the original kernel coordinate space are down-sampled in the new kernel coordinate space to obtain a convolution kernel capable of extracting spatial structure and directional information in the license plate image data information.

[0075] Through the above steps, the weight parameter of the rotated convolution kernel can be obtained. Based on this, the rotation process of the convolution kernel is essentially a process of sampling new weight parameters from the rotated coordinates in the kernel space.

[0076] In the embodiment of the present application, it is assumed that an AKRDC convolution has n groups of convolution kernels, denoted as W:{W1,W2,…,W n}, and each group of convolution kernels has a shape of (c out ,c in ,k,k), then, given the input feature x, after obtaining the rotation angle θ:{θ1,θ2,…,θ n} through θ,α=f(x), the n groups of convolution kernels are individually rotated according to the predicted rotation angle, as shown in the following formula:

[0077] W′ i =rotate(W i ,θ i ),i=1,2,…,n,

[0078] In the formula, θ i represents the rotation angle of the convolution kernel W i , W′ i represents the weight parameter of the rotated convolution kernel, and rotate(·) represents the above rotation process.

[0079] S213, combine the plurality of convolution kernels with spatial structure and direction information according to the combination weight to obtain a to-be-detected license plate feature with spatial structure and direction information.

[0080] It should be noted that for the convolution with multiple groups of convolution kernels, a simple and direct processing method is to respectively perform convolution operation on each group of convolution kernels of the convolution with the input feature, and then integrate the output features in an element-by-element addition manner, as shown in the following formula:

[0081] y = a1(W'1*x) + a2(W'2*x) +... + a n (W' n *x),

[0082] Wherein, a: {a1, a2,..., a n} represents the combination weight predicted by the dynamic prediction module, * represents the convolution operation, and y represents the combined output feature.

[0083] It should be understood that if the convolution calculation is performed in the above formula, the calculation amount of the model will be greatly increased, which affects the efficiency of the model. Therefore, the combination calculation mechanism shown in the following formula is used for optimization in the embodiment of the application:

[0084] y = (a1W'1 + a2W'2 +... + a n W' n )*x.

[0085] This means that performing convolution operation on each group of convolution kernels with the input feature and adding these output features (y = a1(W'1*x) + a2(W'2*x) +... + a n (W' n *x)) is equivalent to performing one convolution operation (y = (a1W'1 + a2W'2 +... + a n W' n )*x) using these convolution kernels through the weighted sum of the combination weight. This combination calculation mechanism not only improves the feature representation capability of the network for extracting multiple directional to-be-detected targets, but also remains efficient, because compared with the repeated convolution operation of the formula y = a1(W'1*x) + a2(W'2*x) +... + a n (W' n *x), there is only one convolution operation in the formula y = (a1W'1 + a2W'2 +... + a n W' n )*x.

[0086] In the embodiment of the application, feature fusion is performed on the extracted features. Specifically, feature fusion is performed through a PAN-FPN structure.

[0087] It should be noted that PAN-FPN refers to a network architecture that combines PAN (Path Aggregation Network) and FPN (Feature Pyramid Network). This combination aims to fully utilize the advantages of FPN and PAN to improve the performance of rotating target detection. Specifically, in the PAN-FPN architecture, FPN passes high-level semantic information through a top-down path, while PAN passes low-level positioning information through a bottom-up path. This combination enables the network to utilize both high-level semantic information and low-level positioning information, thereby improving the accuracy of rotating target detection. PAN-FPN has the following advantages:

[0088] (1) Multi-scale feature fusion: PAN-FPN combines the paths of FPN and PAN to achieve multi-scale feature fusion. This fusion not only enhances the semantic information of features but also preserves low-level positioning information, which helps to more accurately detect targets.

[0089] (2) Improved detection accuracy: By combining the paths of FPN and PAN, PAN-FPN can improve detection accuracy while preserving detailed information. This combination makes the network more flexible and accurate when processing targets of different scales.

[0090] In the embodiment of the present application, the detection head prediction module is used to perform target detection based on task interaction joint features on the feature fusion result of the to-be-detected license plate, as shown in Figure 3 , which includes:

[0091] S221, the feature fusion result of the to-be-detected license plate is sent to a shared convolution to generate task interaction joint features;

[0092] In the embodiment of the present application, the task dynamic alignment detection head (hereinafter referred to as TDADH) sends the features on each detection head to a shared convolution to generate task interaction joint features.

[0093] S222, according to the task interaction joint features, the task is dynamically decomposed;

[0094] Specifically, according to the task interaction joint features, the task is dynamically decomposed, which includes:

[0095] Respectively using task interaction joint features and their global average pooling features to perform dynamic decomposition of classification tasks and regression tasks to obtain classification features and regression features, wherein the global average pooling features of the task interaction joint features are used to dynamically generate kernel attention weights.

[0096] Specifically, the dynamic decomposition of the classification task and the regression task is performed using the task interaction joint feature and the global average pooling feature thereof respectively, the average pooling feature of the task interaction joint feature is used to dynamically generate kernel attention weight, and the task interaction joint feature is output after the kernel attention mechanism to output the classification feature and the regression feature after task decomposition respectively.

[0097] S223, determining the weight parameters of each task after dynamic decomposition according to the task interaction joint feature;

[0098] Specifically, the feature selection weight of the classification branch and the offset and mask parameters required by the regression branch DyDCNv2 are generated based on the task interaction joint feature respectively

[0099] S224, target prediction is performed according to the weight parameters of each task after dynamic decomposition.

[0100] In the embodiment of the present application, the target prediction according to the weight parameters of each task after dynamic decomposition comprises:

[0101] 1) performing feature attention on the classification feature according to the feature selection weight to obtain a classification prediction result;

[0102] 2) performing feature attention on the regression feature according to the fusion feature attention to obtain a regression prediction result;

[0103] 3) realizing task alignment according to the classification prediction result and the regression prediction result to obtain a license plate positioning detection result.

[0104] It should be understood that for the classification branch, feature attention is performed using the feature selection weight and finally a classification prediction is generated; and for the regression branch, DyDCNv2 fusion feature attention is used and finally a regression prediction is obtained.

[0105] In the TDADH detection head, the feature selection weight of the classification branch and the offset and mask of the regression branch are generated by using the task interaction joint feature, and they are regarded as feature attention acting on the features of the respective tasks, so that the task alignment can be well realized.

[0106] The specific effects brought by the task dynamic alignment detection head (TDADH) of the present application in the prediction in the embodiment of the present application are described in detail as follows.

[0107] (1) TDADH detection head uses GroupNorm instead of BatchNorm used in the backbone feature extraction network to improve the classification and positioning performance of the detection head.

[0108] (2) TDADH uses shared convolution to share the weight parameters of convolution between each detection head. By using shared convolution, the amount of parameters and calculations required by the model can be significantly reduced. This approach not only reduces the burden of the model, but also helps to optimize the computational efficiency of the model, especially in a device environment with limited hardware resources.

[0109] (3) Considering the inconsistent target scale that may occur when processing multiple detection heads, the TDADH in the embodiment of the application can adjust and scale the features through the Scale layer. This ensures that each detection head can accurately capture and distinguish targets of different scales regardless of their position in the image, thereby improving the overall detection accuracy. In this way, not only does it solve the problem of inconsistent scales, but it also optimizes the model's adaptability to targets of different scales, making it more suitable for complex and varied scene requirements in actual applications.

[0110] (4) In addition to task alignment in the label assignment strategy, TDADH also performs customized task alignment structure on the detection head. To address the problem of the lack of interaction between the two tasks caused by the use of independent classification and positioning branches in existing rotating target detection heads, TDADH learns task interaction joint features from multiple convolution layers. The positioning branch uses DyDCNv2 and task interaction joint features to generate the offset and mask of DyDCNv2, while the classification branch uses task interaction joint features for dynamic feature selection.

[0111] In the embodiment of the application, after obtaining the license plate positioning detection result, the license plate positioning detection result is post-processed. Specifically, the pre-trained rotating target license plate detection model further includes a post-processing module, which is used to perform a non-maximum suppression post-processing operation on the license plate positioning detection result to screen a detection frame that meets a preset standard.

[0112] Specifically, the detection result of the rotating target detection algorithm needs to be post-processed by non-maximum suppression (NMS). NMS removes duplicate detection results for the same target in the same region according to the position information and classification score of the rotating rectangular bounding box predicted by the rotating target detection model. The removal criterion is to remove the detection frame with a lower classification score and leave the detection frame with a higher classification score.

[0113] In summary, the license plate positioning detection method for complex road scenes provided by the application can distinguish the real direction of the current rotating target detection task in a complex scene and improve the result prediction accuracy, thereby improving the detection accuracy of the license plate detection in a complex road environment.

[0114] The following is the experimental verification process and verification results of the license plate positioning detection method for complex road scenes of the application.

[0115] The specific experimental verification is as follows:

[0116] Dataset: The application uses a large-scale rotating target detection dataset DOTA and a self-built license plate positioning dataset to verify. The size of the image in the DOTA dataset is between 800x800 and 20000x20000 pixels, and contains targets of various scales, directions and shapes. The original DOTA dataset contains a total of 2806 images, and the dataset is divided into training set, verification set and test set three parts, and the corresponding number of images is 1411, 458 and 937 respectively. It should be noted that the original DOTA dataset is reconstructed according to the widely recognized method in the embodiment of the application, and the specific reconstruction rule is: the images in the original DOTA dataset are cropped into 1024x1024 sub-images of the same size, and the cropping step is 824, that is, there is at least a 200-pixel overlap between two adjacent sub-images. The reconstructed DOTA dataset contains a total of 31879 images, and the number of images corresponding to the training set, the verification set and the test set is 15749, 5297 and 10833 respectively. In addition, since the test set of the DOTA dataset does not disclose the data label, the experiment of the application uses the training set to train the model and uses the verification set to test the model. The license plate positioning dataset is derived from the video stream of many vehicle inspection stations in a certain area, and the image frames are intercepted from the video to construct the dataset. The dataset is classified into four categories of blue license plate, green license plate, yellow single-row license plate and yellow double-row license plate according to the license plate color and license plate number arrangement.

[0117] Benchmark model: The pre-trained rotating target license plate detection model in the embodiment of the application takes the Yolo model widely used in engineering practice as the benchmark. Specifically, the benchmark model of the pre-trained rotating target license plate detection model is Yolov8-obb. Based on this rotating target detection model, the backbone feature extraction network and the rotating target detection head are improved and optimized, and the improved model is compared with the benchmark model.

[0118] Evaluation index: For the rotating target detection model, the evaluation indexes considered in the embodiment of the application mainly include precision, recall, average precision (AP) and mean average precision (mAP). Among them, the precision (P) refers to the proportion of actual positive samples in the predicted positive samples; the recall (R) refers to the proportion of all actual positive samples that are correctly predicted; the AP is the average value of the area under the PR curve (Precision-Recall Curve), which reflects the average precision of the model at different recall rates; the mAP is the average value of the AP calculated on multiple classes, which is a very important comprehensive performance index in rotating target detection. The higher the mAP is, the better the performance of the model is.

[0119] Experimental results: Tables 1 and 2 are the experimental results of the benchmark model and the method of the present application on the DOTA dataset. From the data in the table, it can be seen that compared with the benchmark model Yolov8-obb, the method of the embodiment of the application has obvious improvement in the indexes of precision (P), recall (R) and average precision (mAP50 and mAP50-95). This means that compared with the benchmark model, the embodiment of the application not only improves the detection accuracy of the model, but also is more conducive to reducing the missed detection of the target instances to be detected, and shows excellent comprehensive detection performance. Table 3 is the experimental results of the benchmark model and the method of the embodiment of the application on the license plate positioning dataset. As can be seen from the table, except that the accuracy (P) has a certain degree of decline, other indexes have a greater improvement, which is sufficient to show that the embodiment of the application has achieved better results in license plate positioning.

[0120] Table 1 Experimental results on DOTA-1.0 dataset

[0121]

[0122] Table 2 Experimental results on DOTA-1.5 dataset

[0123]

[0124]

[0125] Table 3 Experimental results on license plate positioning dataset

[0126]

[0127] It is found through experiments that, compared with the detection frame containing a large amount of background information when the traditional target detection method is used to detect a rotating target, the method of the embodiment of the application can accurately locate the bounding box of the rotating target (license plate) and rarely includes background information into the detection frame of the target to be detected. This is due to the adaptive kernel rotation dynamic convolution proposed in the embodiment of the application, which can adaptively adjust the direction of the convolution kernel according to the orientation of the target to be detected, so as to extract information more in line with the characteristics of the target to be detected, so that the rotating detection frame can fit the actual target. In addition, in the detection result of the embodiment of the application, the class confidence score of the rotating rectangular bounding box is generally good. This is because the task dynamic alignment detection head proposed in the application comprehensively considers the classification information and positioning information, so that the target positioning feature can assist the classification task, so that the classification task can pay attention to the basic features (such as shape and size) of the target to improve the classification effect of the classification task. Therefore, the license plate positioning and detection method for complex road scenes provided by the application can distinguish the real direction of the current rotating target detection task in the complex scene and improve the result prediction accuracy to improve the detection accuracy of the license plate detection in the complex road environment.

[0128] As another embodiment of the application, a computer storage medium is provided for storing a computer program, which is executed by a processor to implement the license plate positioning and detection method for complex road scenes described above.

[0129] In the embodiment of the application, a non-transitory computer readable storage medium is provided, which stores computer executable instructions, and the computer executable instructions can execute the license plate positioning and detection method for complex road scenes in any method embodiment described above. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD) or a solid state drive (SSD), etc. The storage medium can also include a combination of the above types of memories.

[0130] As another embodiment of the application, an electronic device is provided, which includes a memory and a processor, the processor is in communication connection with the memory, the memory is used to store a computer program, and the processor is used to load and execute the computer program to implement the license plate positioning and detection method for complex road scenes described above.

[0131] AsFigure 4 As shown, the electronic device 10 can include at least one processor 11, such as a CPU (Central Processing Unit), at least one communication interface 13, a memory 14, and at least one communication bus 12. The communication bus 12 is used to realize the connection and communication between the components. The communication interface 13 can include a display, a keyboard, and can also include a standard wired interface and a wireless interface. The memory 14 can be a high-speed RAM memory, or a non-volatile memory such as at least one disk memory. The memory 14 can also be at least one storage device located away from the aforementioned processor 11. The memory 14 stores an application program, and the processor 11 calls the program code stored in the memory 14 to execute any of the above method steps.

[0132] The communication bus 12 can be a PCI (peripheral component interconnect) bus or an EISA (extended industry standard architecture) bus, etc. The communication bus 12 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.

[0133] The memory 14 can include a volatile memory such as a RAM (random-access memory), and can also include a non-volatile memory such as a flash memory, a hard disk (HDD) or a solid-state disk (SSD). The memory 14 can also include a combination of the above types of memories.

[0134] The processor 11 can be a CPU (central processing unit), a network processor (NP), or a combination of a CPU and an NP.

[0135] The processor 11 can further include a hardware chip. The hardware chip can be an application-specific integrated circuit (ASIC), a programmable logic device (PLD) or a combination thereof. The PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.

[0136] Optionally, the memory 14 is further configured to store program instructions. The processor 11 can invoke the program instructions to implement the method for license plate positioning detection in a complex road scene as described in the embodiments of the present application. Figure 1 The method for license plate positioning detection in a complex road scene as shown in the embodiments.

[0137] It can be understood that the above embodiments are only exemplary embodiments adopted for illustrating the principles of the present application, and the present application is not limited thereto. Various modifications and improvements can be made by those of ordinary skill in the art without departing from the spirit and essence of the present application, and these modifications and improvements are also considered to be within the protection scope of the present application.

Claims

1. A license plate positioning detection method for complex road scenes, characterized in that, The method comprises: acquiring license plate image data information of a running motor vehicle; inputting the license plate image data information into a pre-trained rotating target license plate detection model for license plate positioning detection to obtain a license plate positioning detection result; displaying the license plate positioning detection result to a user; wherein the pre-trained rotating target license plate detection model at least comprises a feature extraction module, a feature fusion module and a detection head prediction module; the feature extraction module is configured to perform feature extraction on the license plate image data information according to adaptive kernel rotation dynamic convolution to obtain license plate features to be detected with spatial structure and direction information; the feature fusion module is configured to perform feature fusion on the license plate features to be detected based on a path aggregation network and a feature pyramid network to obtain license plate feature fusion results with multi-scale features; and the detection head prediction module is configured to perform target detection on the license plate feature fusion results based on task interaction joint features to obtain the license plate positioning detection result.

2. The method for license plate positioning detection in complex road scenes according to claim 1, characterized in that, The feature extraction module is configured to perform feature extraction on the license plate image data information according to adaptive kernel rotation dynamic convolution to obtain license plate features to be detected with spatial structure and direction information, which comprises: performing dynamic prediction on the license plate image data information to obtain a rotation angle and a combination weight of each convolution kernel capable of representing the license plate features to be detected; performing adaptive rotation on the convolution kernel according to the rotation angle of each convolution kernel to obtain a convolution kernel capable of extracting spatial structure and direction information in the license plate image data information; performing combination calculation on a plurality of convolution kernels with spatial structure and direction information according to the combination weight to obtain the license plate features to be detected with spatial structure and direction information.

3. The method for license plate positioning detection in complex road scenes according to claim 2, characterized in that, The feature extraction module is configured to perform feature extraction on the license plate image data information according to adaptive kernel rotation dynamic convolution to obtain license plate features to be detected with spatial structure and direction information, which comprises: taking weight parameters of an original convolution kernel as sampling points to construct an original kernel coordinate space through an up-sampling method of bilinear interpolation; rotating the original kernel coordinate space around a center point to obtain a new kernel coordinate space; performing down-sampling operation on values in the original kernel coordinate space in the new kernel coordinate space to obtain a convolution kernel capable of extracting spatial structure and direction information in the license plate image data information.

4. The method for license plate positioning detection in complex road scenes according to claim 1, characterized in that, The detection head prediction module is configured to perform target detection on the license plate feature fusion results based on task interaction joint features, which comprises: feeding the license plate feature fusion results into a shared convolution to generate task interaction joint features; performing dynamic decomposition of tasks according to the task interaction joint features; determining weight parameters of each task after dynamic decomposition according to the task interaction joint features; performing target prediction according to the weight parameters of each task after dynamic decomposition.

5. The method for license plate positioning detection in complex road scenes according to claim 4, characterized in that, The detection head prediction module is configured to perform target detection on the license plate feature fusion results based on task interaction joint features, which comprises: performing dynamic decomposition of classification tasks and regression tasks using the task interaction joint features and global average pooling features thereof respectively to obtain classification features and regression features, wherein the global average pooling features of the task interaction joint features are used to dynamically generate kernel attention weights.

6. The method for license plate positioning detection in complex road scenes according to claim 5, characterized in that, According to the weight parameters of each task after dynamic decomposition, target prediction is performed, including: According to the feature selection weight, performing feature attention on the classification features to obtain a classification prediction result; According to the fusion feature attention, performing feature attention on the regression features to obtain a regression prediction result; According to the classification prediction result and the regression prediction result, realizing task alignment to obtain a license plate positioning detection result.

7. The method for license plate positioning detection in complex road scenes according to claim 1, characterized in that, The pre-trained rotating target license plate detection model further comprises a data preprocessing module, which is configured to perform normalization and data enhancement processing on the license plate image data information.

8. The method for license plate positioning detection in complex road scenes according to claim 1, characterized in that, The pre-trained rotating target license plate detection model further comprises a post-processing module, which is configured to perform a non-maximum suppression post-processing operation on the license plate positioning detection result to screen a detection frame meeting a preset standard.

9. A computer storage medium, characterized in that A computer program is stored, and the computer program is executed by a processor to implement the license plate positioning detection method for a complex road scene according to any one of claims 1 to 8.

10. An electronic device, comprising: A memory and a processor are included, the processor is in communication connection with the memory, the memory is configured to store a computer program, and the processor is configured to load and execute the computer program to implement the license plate positioning detection method for a complex road scene according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Improved target detection method based on YOLOv8s

    CN118982734A

  • Ship orientation detection system and method based on rotation convnext and enhanced feature fusion

    CN119693796A