License plate number identification method applied to highway hazardous chemical substance transport vehicle
Image correction and deblurring enhancement were performed using the DFRSC-MB-Taylor Former V2 joint model, and precise localization was achieved using the MBPSO-FOCUS model. The SegVG model was used to generate a license plate region segmentation mask, and the RILPQ-CRNN model was used for license plate character recognition. This solved the problem of insufficient recognition accuracy for hazardous chemical transport vehicles on highways and achieved efficient and accurate license plate number recognition.
Patent Information
- Application Number
- CN202510946820.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-11-21
AI Technical Summary
On highways, license plate recognition for hazardous chemical transport vehicles suffers from problems such as image distortion and blurring. Traditional methods are unable to effectively restore the true state of the vehicle, resulting in insufficient recognition accuracy and a high false judgment rate, which affects management efficiency and accuracy.
The DFRSC-MB-Taylor Former V2 joint model was used for image correction and deblurring enhancement. The MBPSO-FOCUS model was used to accurately locate hazardous chemical transport vehicles. The SegVG model was used to generate license plate region segmentation masks, and the RILPQ-CRNN model was used for license plate character recognition.
It improved the accuracy and efficiency of identifying hazardous chemical transport vehicles, reduced the false identification rate, and achieved effective management and monitoring of hazardous chemical vehicles.
Smart Images

Figure CN120997812A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the field of image processing and recognition, and in particular to a license plate number recognition method for a highway dangerous chemical product transport vehicle. BACKGROUND
[0002] With the full coverage of the highway network in China and the vigorous development of the dangerous chemical product transport industry, the proportion of dangerous chemical product vehicles in highway transportation is increasing year by year. Because these vehicles carry special goods such as flammable, explosive, toxic and harmful goods, their safety management has become the core link of traffic supervision. As an electronic identification of vehicle identity, accurately identifying the license plate of a dangerous chemical product vehicle and completing information input is the basis for realizing transportation supervision, dynamic risk early warning and emergency response.
[0003] Due to the high-speed movement of vehicles in the highway scene, the images captured by the camera are prone to distortion, blurring and other distortion problems, which greatly increases the difficulty of identifying the type and license plate of the dangerous chemical product transport vehicle. The traditional image correction technology cannot effectively restore the true state of the vehicle, resulting in insufficient accuracy of subsequent identification and positioning. Moreover, the corrected image often still has problems such as blurring, noise or low contrast, and traditional image enhancement methods cannot efficiently improve the clarity and quality of the image, thereby affecting the accurate identification of the dangerous chemical product vehicle identification and license plate number.
[0004] In the aspect of dangerous chemical product vehicle license plate recognition, road monitoring images usually contain dangerous chemical product vehicles, ordinary vehicles and backgrounds, and other elements. Traditional positioning models cannot accurately filter out complete dangerous chemical product transport vehicle images, and have problems such as unreasonable feature selection and incomplete filtering of redundant information, resulting in low positioning efficiency and accuracy. In addition, the traditional model has limited recognition accuracy for the license plate area, which is prone to feature misjudgment, resulting in insufficient positioning accuracy and high recognition error rate in complex scenes, which seriously affects the efficiency and accuracy of dangerous chemical product vehicle information input. SUMMARY
[0005] The purpose of the present application is to provide a license plate number recognition method for a highway dangerous chemical product transport vehicle.
[0006] Technical scheme: The license plate number recognition method for a highway dangerous chemical product transport vehicle provided by the present application comprises the following steps:
[0007] (1) Constructing a DFRSC-MB-Taylor Former V2 joint model to preprocess the collected vehicle images;
[0008] (2) Using a FOCUS model improved based on an MBPSO algorithm to accurately identify dangerous chemical product transport vehicles in the preprocessed vehicle images;
[0009] (3) Using the SegVG model, the target area segmentation mask is generated to realize accurate positioning of the license plate area of the dangerous chemical transport vehicle;
[0010] (4) Constructing the RILPQ-CRNN joint model to enhance the extraction ability of the CRNN model to the license plate features, and realizing accurate recognition of the license plate characters of the dangerous chemical transport vehicle.
[0011] Further, the step (1) comprises:
[0012] (1.1) Using the DFRSC model to correct the collected expressway vehicle image;
[0013] (1.2) Using the MB-Taylor Former V2 model to perform deblurring and enhancement processing on the corrected vehicle image.
[0014] Further, the step (1.1) comprises:
[0015] (1.1.1) Inputting multiple frames of RS images; inputting N consecutive multiple frames of rolling shutter RS images X r ∈R N×H×W×3 ; wherein N is the number of RS frames; H is the original RS image height; W is the original RS image width; R is the real number set;
[0016] (1.1.2) Multi-scale feature extraction; using a weight-shared image encoder to extract multi-scale frame-level RS features from N consecutive RS frames, and the formula is:
[0017]
[0018] Wherein, F l is the RS multi-scale feature; L is the feature layer number; l is the feature level;
[0019] (1.1.3) Initial distortion flow estimation; through a global correlation attention mechanism, the distortion flow from the GS image to the RS image is jointly estimated at the lowest resolution layer;
[0020] (1.1.4) Step-by-step refinement; inputting the initial distortion flow predicted at the lowest resolution and the GS features into the decoder, performing layer-by-layer upsampling and fusing the features;
[0021] (1.1.5) Multi-distortion flow decoding; using a distortion flow field prediction strategy at the highest resolution level to distort the RS features to the GS space, predicting multiple distortion flow fields, and the formula is:
[0022] {U 1,r←g , L, U G,r←g} = ConvBlock(Ux ←g , Fg , F warped )
[0023] wherein, U 1,r←g , L, U G,r←g is a predicted plurality of distortion flow fields; U r←g is a refined distortion flow; F g is a GS feature; F warped is a warped RS feature;
[0024] (1.1.6) GS image generation; input the obtained refined GS feature, the plurality of distortion flow fields and the warped RS feature into a decoder to obtain a corrected GS image X g of the transport vehicle shot by the highway camera.
[0025] Further, the step (1.2) comprises:
[0026] (1.2.1) input the image X g ∈R 3×h×w corrected by the DFRSC model;
[0027] (1.2.2) perform shallow feature extraction through a convolutional layer to generate an initial feature F0, and the formula is:
[0028] F0=Conv(X g )=R c×h×w
[0029] wherein, F0is the initial feature; c is the number of channels; h is the height of the initial feature image; w is the width of the initial feature image; and Conv is a convolution operation;
[0030] (1.2.3) generate image features of different scales from the initial feature F0by using a multi-branch depth separable deformation convolution;
[0031] (1.2.4) input the features of different scales into a plurality of Transformer encoder branches respectively, each branch uses a linear self-attention mechanism to reduce the computational complexity, and meanwhile, a dynamic position encoding is integrated to enhance the spatial perception, so as to realize parallel processing and fusion of multi-scale features;
[0032] (1.2.5) use an SKFF module to perform feature fusion on the features output by the branches of the Transformer encoder according to feature weights, so as to obtain fused features;
[0033] (1.2.6) perform down-sampling on the fused branch features by using a pixel de-shuffling operation to obtain deep multi-scale features. Then, the deep multi-scale features are up-sampled by using a pixel shuffling operation through a decoder, so that the resolution of the deep multi-scale features is restored to the original resolution;
[0034] (1.2.7) The features of the current stage encoder are spliced with the up-sampled features using the channel splicing and convolution fusion method;
[0035] (1.2.8) The spliced features are further refined by a residual block to optimize the structure and texture details and generate high-quality repair feature maps;
[0036] (1.2.9) The repair feature maps are reduced from C to 3 channels using a 3x3 convolution layer to generate a residual image R, and the residual image R is superimposed on the input image X g ∈R 3×h×w , thereby obtaining a repaired and enhanced image, and the formula is:
[0037] X=X g +R;
[0038] Where X is the image enhanced by the MB-Taylor Former V2 model; and R is the residual image.
[0039] Further, the step (2) comprises:
[0040] (2.1) Image blocking; the preprocessed image X is divided into N non-overlapping blocks, each image block contains local features of vehicles, sky, road elements, and the segmented image blocks are constructed in the form of a set, and the formula is:
[0041]
[0042] Where X is the input preprocessed image; x is the segmented image block; and N2 is the number of image blocks.
[0043] (2.2) Image feature extraction; the FOCUS model is used to extract features from the preprocessed photos taken by the highway camera, and the MBPSO algorithm is used to select features from the extracted image features to retain the optimal feature subset;
[0044] (2.3) A global similarity filtering layer is used to remove the background in the segmented image and retain the potential vehicle area;
[0045] The optimized image block features B1 are used to calculate the similarity between image blocks using a sliding window mechanism, and the potential vehicle area is retained by filtering large-scale repeated backgrounds according to the similarity threshold;
[0046] (2.4) A text-guided key block selection layer is used to extract key visual features of hazardous chemical transport vehicles;
[0047] (2.5) An adjacent block compression layer is used to remove multiple perspective repeated features of the same hazardous chemical transport vehicle.
[0048] (2.6) To enhance the model's ability to distinguish the features of the dangerous chemical transport vehicle, the compressed visual features B c are fused with the text prompt T through multi-head attention mechanism, and the formula is:
[0049] where Q = TW q , K = B c W k′ ,
[0050] where, is the multi-head attention mechanism; Q is the text prompt as the query; K is the compressed block feature as the key; V is the compressed block feature as the value; is the i-th learnable projection matrix; W q is the learnable projection matrix, W k′ is the learnable projection matrix for projecting the image block features into the key space; is the value projection matrix; d k′ is the dimension of each attention head;
[0051] After concatenating the outputs of each attention head, normalize it, and the formula is:
[0052] O = LayerNorm (Concat (Head1,..., Head h′ ) W o )
[0053] where O is the normalized concatenated multi-head attention mechanism; W o is a learnable matrix; LayerNorm is layer normalization; Concat is the concatenation operation;
[0054] (2.7) Predict whether the image contains a dangerous chemical transport vehicle through the classifier, when the probability is greater than 0.5, determine that there is a dangerous chemical transport vehicle in the image block, and the formula is:
[0055]
[0056] where, is the probability of outputting the vehicle type; Y is the class label of the classification task, Y = 0 indicates that the vehicle in the image is a normal vehicle, Y = 1 indicates that the vehicle in the image is a dangerous chemical transport vehicle; is a learnable matrix; is a learnable vector;
[0057] (2.8) The spatial reconstruction of the retained dangerous chemical transport vehicle image block features after classification is performed, and they are remapped back to the spatial position of the original image, thereby obtaining the complete image X1 of the dangerous chemical transport vehicle.
[0058] Further, the step (2.2) comprises:
[0059] (2.2.1) The segmented image is subjected to feature extraction of the image block using a pre-trained vehicle detection model, and the formula is:
[0060]
[0061] wherein B is the extracted feature set; f is the pre-trained vehicle detection model for extracting the block features; is the i1th image block; is the feature vector of the i1th block;
[0062] (2.2.2) The extracted image features B are selected and optimized by multi-cell cooperation using the MBPSO algorithm,
[0063] The organizational membrane system generally includes three cells, σ1, σ2 for local search of the initial image features B, and σ3 for finding the globally optimal image feature subset, and the formula is:
[0064] Π=(O, σ1, σ2, σ3, syn, i0)
[0065] wherein Π is the organizational membrane system; O is a set of binary coded particles; σ1, σ2 are initial particle group cells, each managing a group of candidate feature subset groups, and each candidate feature subset is a binary vector; σ3 is the globally optimal candidate feature subset; syn is the communication rule, which realizes information interaction between the feature subsets; i0 is the label of the feature subset outputting the final result;
[0066] The initial candidate feature subsets are randomly generated in σ1, σ2, and the position of each candidate feature subset is a binary vector, and the speed is initialized as a random value;
[0067] The speed update formula of each candidate feature subset is:
[0068]
[0069] wherein is the speed of the i2th candidate feature subset at t2+1; ω is the inertia weight, which controls the influence of the historical speed; c1 is the learning factor, which represents the weight of the individual optimum; pbest is the individual optimum; is the state of the i2th candidate feature subset at time t; c2 is a learning factor, representing the weight of the global optimum; gbest is the global optimum; i2 is the number of candidate feature subsets; t2 is the number of iterations; and rand is a random number in [0, 1];
[0070] The position of the candidate feature subset is updated, and the speed is mapped to a probability through a Sigmoid function to determine whether the feature is selected, and the formula is:
[0071]
[0072] wherein, is the j2-dimensional feature selection state of the i2th candidate feature subset; is the j2-dimensional feature selection speed of the i2th candidate feature subset; and j2 is the feature dimension number;
[0073] σ3 receives all individual optimal feature subsets pbest, and selects the one with the highest fitness as the global optimal feature subset gbest, and the formula is:
[0074] r 32 : s 3,k+1 (f 3,k+l w 3,k+1 )→s 3,k+1 {(w 3,k+l ,go),(f 3,k+1 ,go)}
[0075] wherein, r 32 is the rule of cell σ3; s 3,k+l is the new state of σ3 at the k+1th iteration; f 3,k+1 is the updated global optimal fitness value of cell σ3; and w 3,k+1 is the new position and speed set of the feature subset group updated by σ3.
[0076] When the maximum number of iterations t max is reached, the binary vector corresponding to gbest is the optimal feature subset B1, and the formula is: B1 = gbest
[0077] wherein, B1 is the image feature extracted by optimizing the initial feature B using the MBPSO algorithm.
[0078] Further, the step (3) comprises:
[0079] (3.1) inputting the image X1 into the pre-trained DETR model, and using the ResNet and Transformer encoder to extract the visual features of the image, and the formula is:
[0080]
[0081] wherein, is the original visual feature of the i4+1th layer; is the original visual feature of the i4th layer; is the i4th layer in the DETR backbone network; i4 is the total number of layers; v4 represents vision;
[0082] (3.2) The BERT model is used to process the text to obtain the text feature vector The formula is:
[0083]
[0084] wherein, is the original text feature of the i4+1th layer; is the original text feature of the i4th layer; t5 represents text; is the 2i4+1th layer in the DETR backbone network; is the 2i4th layer in the DETR backbone network;
[0085] (3.3) The Triple Alignment module is introduced to iteratively update the query, text and visual features through a three-way attention mechanism;
[0086] (3.4) The text feature vector and the visual feature vector after alignment are combined to generate a multi-modal feature vector;
[0087] (3.5) The output of the encoder is input into the decoder, and the decoder adopts the bbox2seg scheme to convert the target box label into a segmentation mask.
[0088] Further, the step (3.5) comprises:
[0089] (3.5.1) The multi-modal feature output by the encoder is subjected to target box regression using the regression query vector, and the position range of the dial in the image is determined through the Transformer decoder layer and the MLP network;
[0090] (3.5.2) The segmentation query vector with different learnable position encodings is used to segment the target area, and the segmentation mask of the target area is obtained through the Transformer decoder layer and the MLP network; the pixels in the license plate area are marked as foreground (1), and the pixels outside the area are marked as background (0);
[0091] (3.5.3) The final output is the license plate image X2 of the dangerous chemical transport vehicle located.
[0092] Further, the step (4) comprises:
[0093] (4.1) Image input; input the image X2 of the license plate number of the positioned hazardous chemical transport vehicle;
[0094] (4.2) Image local feature extraction; perform feature extraction on the input image X2 by using the RILPQ algorithm;
[0095] (4.3) Bidirectional LSTM sequence modeling; use a bidirectional LSTM network to model the feature sequence x' = x'1,..., x'L extracted by the convolutional layer, and output a label distribution y = y1,..., yT for each frame T
[0096] (4.4) CTC transcription layer decoding; input the predicted label result into the transcription layer and decode it by using CTC, and finally obtain the accurate license plate character recognition result by removing repeated characters and white spaces.
[0097] Further, the step (4.2) comprises:
[0098] (4.2.1) Perform feature extraction on the license plate by using the RILPQ feature extraction algorithm, analyze the local neighborhood frequency domain information by using the short-time Fourier transform, and realize rotation and blur invariance by combining the direction quantization;
[0099] (4.2.2) Map the RILPQ(x) feature to a feature map F' of the same size as the input image, and concatenate it with the original image X2 as a multi-channel input, and the formula is:
[0100] X3 = Concat(X2, F')
[0101] Wherein, X3 is the fused multi-channel image; F' is the feature map generated by RILPQ;
[0102] (4.2.3) Input the fused multi-channel image X3 into the convolutional layer sequence of CRNN, perform multi-layer convolution and pooling operation, and extract the enhanced features after the RILPQ algorithm, and the formula is:
[0103]
[0104] Wherein, is the output feature map of the L3 layer; is the convolutional weight of the L3 layer; is the convolutional bias of the L3 layer; L3 is the number of convolutional layers;
[0105] (4.2.4) The feature map generated by the convolution layer is processed by Map-to-Sequence to convert the feature map into a feature sequence to extract a feature vector sequence x' = x'1,..., x'N. T as the input of the recurrent layer.
[0106] Advantages: Compared with the prior art, the present application has the following obvious advantages: the present application combines the DFRSC-MB-Taylor Former V2 image preprocessing model, the improved MBPSO-FOCUS hazardous chemical product transportation vehicle positioning model, and the RILPQ-CRNN vehicle license plate number recognition model, etc., to solve the problems existing in the character recognition of the traditional license plate recognition method, improve the recognition accuracy and efficiency, and reduce the misrecognition rate. BRIEF DESCRIPTION OF DRAWINGS
[0107] Figure 1 The framework diagram of the license plate number recognition of the hazardous chemical product transportation vehicle on the expressway provided for the embodiment of the present application;
[0108] Figure 2 The framework diagram of the DFRSC-MB-Taylor Former V2 image preprocessing joint model constructed for the present application;
[0109] Figure 3 The flowchart of the improved image positioning of the hazardous chemical product transportation vehicle on the expressway;
[0110] Figure 4 The flowchart of the improved license plate character recognition of the hazardous chemical product transportation vehicle. DETAILED DESCRIPTION
[0111] The technical solutions of the present application will be further described below in combination with the drawings.
[0112] As Figure 1As shown, the embodiment of the application provides a method for highway dangerous chemical transportation vehicle license plate recognition. First, the vehicle image captured by the highway monitoring camera is obtained; a DFRSC-MB-Taylor Former V2 joint model is constructed to preprocess the collected vehicle image, the joint model combines the image correction function of the DFRSC model and the image deblurring enhancement function of the MB-Taylor Former V2 model, which can effectively restore the real state of the vehicle and improve the image clarity and quality; an improved FOCUS model is used to filter invalid information through a three-stage compression strategy, accurately locate the dangerous chemical transportation vehicle, and the specific improvement measures are as follows: in the feature selection stage, the MBPSO algorithm is introduced for feature selection and optimization; the SegVG model is used to extract multi-scale features and fuse them to generate a target region segmentation mask, realizing accurate positioning of the dangerous chemical transportation vehicle license plate region; an improved CRNN model is used to realize the recognition of the dangerous chemical transportation vehicle license plate characters, and the specific improvement measures are as follows: in the feature extraction stage, the RILPQ model is combined to enhance the feature extraction capability of the CRNN model; finally, the license plate recognition result of the dangerous chemical transportation vehicle is entered into the system, which is convenient for effective management and monitoring of the dangerous chemical transportation vehicle.
[0113] The specific implementation process is as follows:
[0114] (1) Obtain the vehicle image captured by the highway monitoring camera;
[0115] Through the highway monitoring camera network, the image X of the vehicle running on the highway is collected r In the collected image, the application considers ordinary vehicles and dangerous chemical transportation vehicles;
[0116] (2) Construct a DFRSC-MB-Taylor Former V2 joint model to preprocess the collected vehicle image;
[0117] As Figure 2 shown, for the problems of image distortion, blur, noise and the like caused by high-speed movement of vehicles in the highway monitoring image, the application constructs a DFRSC-MB-Taylor Former V2 joint model to preprocess the collected vehicle image.
[0118] The joint model combines the image correction function of the DFRSC model and the image deblurring enhancement function of the MB-Taylor Former V2 model, which can effectively restore the real state of the vehicle and improve the image clarity and quality, and the specific steps are as follows:
[0119] (2.1) Use the DFRSC model to correct the collected highway vehicle image;
[0120] In the highway monitoring system, since the camera shooting image will appear distortion or blur and other distortion phenomena, therefore, the DFRSC model is adopted to correct the shooting distorted photo X r , restore the real state of the vehicle, the specific steps are:
[0121] (2.1.1) input multi-frame RS image
[0122] Input N continuous multi-frame rolling shutter RS image X r ∈R N×H×W×3 ;
[0123] Wherein, N is the RS frame number; H is the original RS image height; W is the original RS image width; i is a real set;
[0124] (2.1.2) multi-scale feature extraction
[0125] The image encoder with weight sharing is adopted to extract multi-scale frame-level RS features from N continuous RS frames, and the formula is:
[0126]
[0127] Wherein, F l is the RS multi-scale feature; K is the feature layer number; L is the feature level;
[0128] (2.1.3) initial distortion flow estimation
[0129] Through the global correlation attention mechanism, the distortion flow from the GS image to the RS image is jointly estimated at the lowest resolution layer, and the specific steps are as follows:
[0130] 1) generate GS feature
[0131] The virtual GS feature F L Is predicted by using the convolution block ConvBlock and the lowest resolution RS feature F L , and the time offset t, and the formula is:
[0132]
[0133] Wherein, Is the GS feature; ConvBlock is the convolution block; g is the GS image; t is the exposure time offset between the target GS image and the middle scanning line in the RS frame;
[0134] 2) calculate global correlation matrix
[0135] The GS feature and the RS feature are respectively taken as the query (Query) and the keyword (Key), and the global correlation modeling is calculated to obtain the global correlation matrix M and RS features F L between the attention matrix, which is formulated as:
[0136]
[0137] where M is the attention matrix; softmax is the function; H' is the height of the RS frame of the encoder output; W' is the width of the RS frame of the encoder output; D is the feature dimension; F L is the lowest resolution RS feature, F L ∈R N×H×W×D ;
[0138] 3) Generating initial distortion flow and warped features
[0139] The RS features are warped to the GS space by the attention matrix and aggregated with the 2D coordinate grid of the RS frame to obtain the globally warped RS features and the distortion flow, which is formulated as:
[0140]
[0141] where, is the warped RS feature; is the distortion flow from the GS image to the RS image; G is the coordinate grid of the RS image; is the initial distortion flow from the GS image to the RS image; r is the RS image, i.e., the photographed distorted picture;
[0142] (2.1.4) Step-by-step refinement
[0143] In order to further improve the accuracy of the distortion flow and the GS features, the present application adopts a step-by-step refinement method, which inputs the initial distortion flow and the GS features predicted at the lowest resolution into the decoder, performs layer-by-layer upsampling and fuses the features, and the specific steps are as follows:
[0144] 1) Reverse warping
[0145] At level l, a reverse warping operation is adopted The RS features are warped to the GS space by the current distortion flow, and the formula is:
[0146]
[0147] where, is the warped RS feature; is the reverse warping operation; is the current refined distortion flow;
[0148] 2) Feature fusion and upsampling
[0149] Next, using the Fusion Block, the warped RS features distortion flow and GS features are fused and the next level distortion flow and GS features are obtained by upsampling operation, which is given by:
[0150]
[0151] where, is the next level distortion flow; is the next level GS features;
[0152] (2.1.5) Multi-distortion flow decoding
[0153] To alleviate some of the erroneous estimated displacements in the distortion flow, a distortion flow field prediction strategy is employed at the highest resolution level to warp the RS features to the GS space to predict multiple distortion flow fields, which is given by:
[0154] {U 1,r←g , L, U G,r←g} = ConvBlock(U r←g , F g , F warped )
[0155] where, U 1,r←g , L, U G,r←g are the predicted multiple distortion flow fields; U r←g is the refined distortion flow; F g is the GS features; F warped is the warped RS features;
[0156] (2.1.6) GS image generation:
[0157] The obtained refined GS features, multiple sets of distortion flow and warped RS features are input into the decoder to obtain the corrected GS image X g of the transport vehicle captured by the highway camera.
[0158] (2.2) Adopting the MB-Taylor Former V2 model to process the corrected vehicle image for deblurring enhancement function;
[0159] After correcting the rolling shutter (RS) image, there are degradation problems such as blur, noise, etc., therefore, the present application adopts the MB-Taylor Former V2 model to process the image for deblurring enhancement to improve the clarity and quality of the image, and the specific steps are as follows:
[0160] (2.2.1) Input the image X of the transport vehicle on the highway corrected by the DFRSC modelg ∈R 3×h×w ;
[0161] (2.2.2) shallow feature extraction is performed through a convolutional layer to generate initial features F0, and the formula is:
[0162] F0=Conv(X g )=R c×h×w
[0163] where F0 is the initial feature; c is the number of channels; h is the initial feature image height; w is the initial feature image width; Conv is the convolution operation;
[0164] (2.2.3) a multi-branch deep separable deformable convolution (DSDCN) is used to generate image features of different scales from the initial features F0;
[0165] (2.2.4) different scale features are input into multiple Transformer encoder branches, each branch uses a linear self-attention mechanism to reduce computational complexity, and a dynamic position encoding is integrated to enhance spatial perception, realizing parallel processing and fusion of multi-scale features, and the specific steps are:
[0166] 1) a new feature output is obtained after Taylor expansion self-attention mechanism processing;
[0167] where each branch uses an improved attention mechanism T-MSA++, which combines the first-order term and the remainder approximation of Taylor expansion to replace the traditional Softmax attention mechanism, while maintaining linear complexity and restoring nonlinear attention, and the formula is:
[0168]
[0169] where V i ′ is the image output feature corresponding to the i-th query; f1 is the attention weight of the first-order Taylor expansion term; is the attention weight of the Taylor expansion remainder; Q i is the i-th query vector; K j is the j-th key vector; V j is the j-th value vector; i is the index of the query vector, corresponding to the i-th position in the feature map; j is the index of the key and value vectors, corresponding to the j-th position in the feature map; is the normalized vector of Q i ; s is a learnable modulation factor; φ p is a ReLU-based norm-preserving mapping function used to approximate the high-order remainder of Taylor expansion; is the column vector of φ p ; is the column vector of K jnormalized vector; is a column vector; N1 is the total number of pixels in the feature map;
[0170] 2) Use multi-scale deep convolution (CPE) to encode the relative position of the output feature, so that the model can better capture the spatial relationship, the specific steps are as follows:
[0171] The input feature map is divided into multiple sub-feature maps, and the formula is as follows:
[0172] V I , V II , L = Split (V)
[0173] Where V I , V II is the segmented sub-feature map; V is the input feature map;
[0174] Use deep convolutional layers with different size convolutional kernels to perform convolution operations on the segmented sub-feature maps, and the formula is as follows:
[0175] CPE (V) = Cat (DWC 3×3 (V I ), DWC 5×5 (V II ), …)
[0176] Where CPE is the convolutional position encoding; Cat is the tensor concatenation operation; DWC 3×3 is a deep convolutional layer with a 3x3 convolutional kernel; DWC 5×5 is a deep convolutional layer with a 5x5 convolutional kernel;
[0177] 3) Add the T-MSA++ output and the CPE output to generate the final feature with position information, and the formula is as follows:
[0178] T-MSA++ (Q, K, V) = V + CPE (V)
[0179] Where T-MSA++ (Q, K, V) is the final feature with position information; V' is the feature map representing the attention weighted feature; Q is the query used to calculate the attention weight; K is the key used to calculate the attention weight;
[0180] (2.2.5) Use the SKFF module to fuse the features output by the branches of the Transformer encoder according to the feature weights to obtain the fused features;
[0181] (2.2.6) The fused branch feature is down-sampled by a pixel unshuffle operation to obtain a deep multi-scale feature. Then, the deep multi-scale feature is up-sampled by a pixel shuffle operation through a decoder to restore the resolution of the deep multi-scale feature to the original resolution;
[0182] (2.2.7) The feature of the current stage encoder is spliced with the up-sampled feature by a channel splicing and convolution fusion method;
[0183] (2.2.8) The spliced feature is further refined by a residual block to optimize the structure and texture details and generate a high-quality repair feature map;
[0184] (2.2.9) The repair feature map is reduced from C to 3 channels by a 3x3 convolution layer to generate a residual image R. The residual image R is superimposed on the input image X g ∈R 3×h×w , so as to obtain a repaired and enhanced image, and the formula is:
[0185] X=X g +R;
[0186] Wherein, X is an image enhanced by the MB-Taylor Former V2 model; R is a residual image;
[0187] (3) The FOCUS model improved based on the MBPSO algorithm is used to accurately identify the dangerous chemical transport vehicle in the preprocessed vehicle image;
[0188] In actual application, the pictures taken by the road monitoring system not only include dangerous chemical vehicles, but also include ordinary vehicles and backgrounds. In order to accurately screen the complete image of the dangerous chemical transport vehicle in the image, the FOCUS model improved based on the binary particle swarm optimization (MBPSO) algorithm enhanced by the membrane calculation is used, and a three-stage compression strategy is used to filter invalid information, so that the identification and screening ability of the dangerous chemical transport vehicle is improved. The specific steps are as follows:
[0189] (3.1) Image blocking
[0190] In order to reduce the calculation complexity and improve the generalization ability of the model, the preprocessed image X is divided into N non-overlapping blocks, and each image block may only contain local features of the vehicle, in addition to elements such as sky and road. The segmented image blocks are constructed in the form of a set, and the formula is:
[0191]
[0192] Wherein, X is an input image after preprocessing; x is a segmented image block; N2 is the number of image blocks;
[0193] Image feature extraction
[0194] The FOCUS model is used for feature extraction of the photos taken by the highway camera after preprocessing, and the MBPSO algorithm is used for feature selection of the extracted image features, and the optimal feature subset is reserved. The specific steps are as follows:
[0195] (3.2.1) The image blocks are extracted by using the pre-trained vehicle detection model, and the formula is:
[0196]
[0197] Where B is the extracted feature set; f is the pre-trained vehicle detection model, and the block feature is extracted; is the ith image block; is the feature vector of the ixth block;
[0198] (3.2.2) The MBPSO algorithm is used to select and optimize the extracted image features B through multi-cell cooperation. The specific steps are as follows:
[0199] 1) Construct the organization membrane system
[0200] The organization membrane system usually includes three cells, σ1, σ2 for local search of the initial image feature B, and σ3 for finding the globally optimal image feature subset, and the formula is:
[0201] Π=(O, σ1, σ2, σ3, syn, i0)
[0202] Where Π is the organization membrane system; O is a set of binary coded particles; σ1, σ2 are initial particle group cells, each managing a group of candidate feature subset groups, each candidate feature subset is a binary vector (0 represents eliminating features, 1 represents retaining features); σ3 is the global optimal candidate feature subset; syn is the communication rule, which realizes the information interaction between the feature subsets; i0 is the label of the feature subset outputting the final result;
[0203] 2) Candidate feature subset initialization
[0204] Randomly generate initial candidate feature subsets in σ1, σ2, and the position of each candidate feature subset is a binary vector (0 represents elimination, 1 represents retention), and the speed is initialized to a random value;
[0205] 3) Dynamic optimization of candidate feature subsets
[0206] The speed update formula of each candidate feature subset is:
[0207]
[0208] where, is the velocity of the i2th candidate feature subset at t2+1; ω is the inertia weight, controlling the influence of the historical velocity; c1 is the learning factor, representing the weight of the individual optimum; pbest is the individual optimum; is the state of the i2th candidate feature subset at t; c2 is the learning factor, representing the weight of the global optimum; gbest is the global optimum; i2 is the number of candidate feature subsets; t2 is the iteration number; rand is a random number in [0, 1];
[0209] The position of the candidate feature subset is updated, and the velocity is mapped to a probability through a Sigmoid function to determine whether the feature is selected, and the formula is:
[0210]
[0211] where, is the j2-dimensional feature selection state of the i2th candidate feature subset; is the velocity of the j2-dimensional feature selection of the i2th candidate feature subset; j2 is the feature dimension number;
[0212] wherein the Sigmoid function formula is:
[0213]
[0214] 4) Update the individual optimal image feature subset;
[0215] The σ1, σ2 pass the optimal feature subset pbest and the fitness value to σ3, and the formula is:
[0216]
[0217] wherein r 12 is the rule of cell σ1, used to transmit the individual optimal feature subset of σ1 to cell σ3; s 1,k+1 , s 2,k+1 are the new states of σ1, σ2 at the k+1th iteration, respectively; w 1,k+1 , w 2,k+1 are the new position and velocity sets of the updated candidate feature subset of σ1, σ2; f 1,k+1 , f 2,k+1 are the fitness values of σ1, σ2 at the k+1th iteration, respectively; go is the operator sent to other cells; r 22 is the rule of cell σ2, used to transmit the individual optimal solution of σ2 to cell σ3; k is the current iteration number;
[0218] 5) Select the globally optimal subset of image features;
[0219] σ3 receives all the individual optimal feature subsets pbest, and selects the one with the highest fitness as the globally optimal feature subset gbest. The formula is as follows:
[0220] r 32 :s 3,k+1 (f 3,k+1 w 3,k+1 )→s 3,k+1 {(w 3,k+1 ,go),(f 3,k+1 ,go)}
[0221] Where, r 32 It is the rule of cellular σ3; s 3,k+1 It is the new state of σ3 in the (k+1)th iteration; f 3,k+1 It is the globally optimal fitness value after the cell's σ3 update; w 3,k+1 It is the new set of positions and velocities of the feature sub-clusters after σ3 update;
[0222] 6) Output the optimal feature subset;
[0223] When the maximum number of iterations t is reached max When gbest is selected, the binary vector corresponding to gbest is the optimal feature subset B1, and its formula is:
[0224] B1 = gbest
[0225] Wherein, B1 is the image feature extracted after optimizing the initial feature B using the MBPSO algorithm;
[0226] (3.3) A global similarity filtering layer is used to remove the background, such as the sky and roads, from the segmented image, while retaining the potential vehicle area;
[0227] The optimized image patch features B1 are used to calculate the similarity between image patches using a sliding window mechanism, and large-scale repetitive backgrounds, such as sky and roads, are filtered out based on a similarity threshold to retain potential vehicle regions. The specific steps are as follows:
[0228] The obtained feature vectors are normalized using the following formula:
[0229]
[0230] in, It is a normalized block Features; It represents the block feature; i3 represents the number of the i3th optimized block in the image block set;
[0231] Compute image patches and The similarity between them is calculated using the following formula:
[0232]
[0233] in, It is the cosine similarity between blocks i3 and j3; It is a normalized block Features; j3 represents the number of the j3rd optimized block in the image patch set;
[0234] The dynamic threshold is calculated by dynamically setting the threshold based on the mean and standard deviation of the similarity within the window. The formula is as follows:
[0235] τ g′ =μ(S)+σ(S)
[0236] Where, τ g′ This is a dynamic threshold; blocks exceeding this value are considered redundant. μ(S) is the average block similarity; σ(S) is the standard deviation of block similarity.
[0237] If the average similarity of block i3 within the window is Higher than τ g′ If a block is identified as a redundant image block and removed, the feature of the filtered block is B1.
[0238] (3.4) Use a text-guided key block selection layer to extract key visual features of hazardous chemical transport vehicles;
[0239] Since there are both ordinary vehicles and hazardous chemical transport vehicles in Block B1, this invention uses a semantic description guidance model for hazardous chemical vehicles to focus on key visual features and accurately locate hazard signs and tank structures. The specific steps are as follows:
[0240] Generate linguistic descriptions related to the visual features of hazardous chemical vehicles, and encode T using a pre-trained language model. P As a text embedding, a learnable cue vector T is also introduced. L To enhance the model's adaptability, the formula is:
[0241]
[0242] Where T is a textual cues representing the vehicle's visual features; T L These are learnable cue vectors used for dynamically adapting the model; T P′ t3 is the text-encoded description of hazardous chemicals; t4 is the number of tokens for text prompts; t5 is the number of tokens for learnable prompts; d is the feature dimension.
[0243] The cross-modal attention score between the image block feature and the text feature is calculated, and a softmax function is used for normalization, and the formula is:
[0244]
[0245] Wherein, A is the text-block attention weight matrix; W q is a learnable projection matrix; W k′ is a learnable projection matrix; k' is the number of blocks matched with the dangerous chemical description;
[0246] The correlation score of the image block with the dangerous chemical description is calculated, and the formula is:
[0247]
[0248] Wherein, is the correlation score of block i3 and the dangerous chemical description; is the attention weight of the j3 image block to the i3 image block in the attention matrix;
[0249] According to the attention score, select the k' image block features most relevant to the text description and the dangerous chemical transport vehicle, and the formula is:
[0250]
[0251] Wherein, is the image block feature vector most relevant to the text description;
[0252] (3.5) Use adjacent block compression layer to remove the multi-view redundant features of the same dangerous chemical transport vehicle;
[0253] The Use adjacent block compression module to eliminate local redundancy and maintain vehicle structure integrity, and the specific steps are:
[0254] Calculate the cosine similarity between adjacent image blocks, and the formula is:
[0255]
[0256] Wherein, is the similarity of adjacent blocks and
[0257] According to the similarity, generate a binary mask to determine whether the image block is similar to the adjacent image block, if similar, the mask is 0, otherwise 1, and the formula is:
[0258]
[0259] wherein, is the reserved dissimilar block; is the dynamic threshold;
[0260] the final compressed image block sequence B after multi-stage threshold filtering c only contains the features of the hazardous chemical transport vehicle;
[0261] (3.6) To enhance the discriminant ability of the model for the features of the hazardous chemical transport vehicle, the compressed visual features B c are fused with the text prompt T through a multi-head attention mechanism, and the formula is:
[0262] wherein Q = TW q , K = B c W k′ ,
[0263] wherein, is the multi-head attention mechanism; Q is the text prompt as the query; K is the compressed block feature as the key; V is the compressed block feature as the value; is the i-th learnable projection matrix; W q is the learnable projection matrix for projecting the text prompt into the query space; W k′ is the learnable projection matrix for projecting the image block feature into the key space; is the value projection matrix; d k′ is the dimension of each attention head;
[0264] After concatenating the outputs of each attention head and normalizing, the formula is:
[0265] O = LayerNorm (Concat (Head1,..., Head h′ ) W o )
[0266] wherein, O is the normalized concatenated multi-head attention mechanism; W o is a learnable matrix for mapping the connected output to the final feature dimension; LayerNorm is layer normalization; Concat is the connection operation;
[0267] (3.7) The classifier is used to predict whether the image contains a dangerous goods vehicle. When the probability is greater than 0.5, it is determined that there is a hazardous chemical vehicle in the image block, and the formula is:
[0268]
[0269] wherein, is the probability of output vehicle type; Y is the class label of the classification task, Y=0 represents that the vehicle in the image is a common vehicle, and Y=1 represents that the vehicle in the image is a dangerous chemical transport vehicle; is a learnable matrix used to map the compressed visual feature matrix to a value vector; is a learnable vector used for bias, adding a constant term in the final prediction;
[0270] (3.8) Spatially reconstructing the dangerous chemical transport vehicle image block features reserved after classification, and remapping them back to the spatial position of the original image, so as to obtain the complete dangerous chemical transport vehicle image X1.
[0271] (4) Using the SegVG model, the target region segmentation mask is generated to realize accurate positioning of the dangerous chemical transport vehicle license plate region;
[0272] As Figure 3 shown, in the license plate recognition task, it is necessary to accurately position the license plate number region in the image of the dangerous chemical transport vehicle to avoid interference of invalid regions. Therefore, the SegVG model is used to accurately position the license plate region in the image X1 of the dangerous chemical transport vehicle. By extracting multi-scale features and fusing them, the final target region segmentation mask is generated, which improves the efficiency and accuracy of the whole system, and the specific steps are as follows:
[0273] (4.1) Input the image X1 into the pre-trained DETR model, and use the ResNet and Transformer encoder to extract the visual features of the image, and the formula is:
[0274]
[0275] wherein, is the original visual feature of the i4+1 layer; is the original visual feature of the i4 layer; is the i4 layer in the DETR backbone network; i4 is the total number of layers; v4 represents vision;
[0276] (4.2) The BERT model is used to process the text to obtain the text feature vector The formula is:
[0277]
[0278] wherein, is the original text feature of the i4+1 layer; is the original text feature of the i4 layer; t5 represents text; is the 2i4+1 layer in the DETR backbone network; is the 2i4th layer in the DETR backbone network;
[0279] (4.3) Introduce the Triple Alignment module, iteratively update the query, text and visual features through the Tri-MHA mechanism (Tri-MHA), eliminate the domain differences between them, and align the features, and the formula is:
[0280]
[0281] wherein Z o′ is the original query feature; Z′ o′ is the updated query feature; is the updated text feature; is the updated visual feature; o′ is the representative query;
[0282] (4.4) Merge the aligned text feature vector and the visual feature vector to generate a multi-modal feature vector. Then input it into the Transformer encoder layer to promote further fusion between features, and finally output the encoding results containing text and visual information;
[0283] (4.5) Input the output of the encoder into the decoder, and the decoder adopts the bbox2seg scheme to convert the target box label into a segmentation mask, and the specific steps are:
[0284] (4.5.1) Use the regression query vector to regress the target box of the multi-modal feature output by the encoder, and determine the position range of the dial in the image through the Transformer decoder layer and the MLP network;
[0285] (4.5.2) Use the segmentation query vector with different learnable position encodings to segment the target area, and also get the segmentation mask of the target area through the Transformer decoder layer and the MLP network. Mark the pixels in the license plate area as foreground (1), and mark the pixels outside the area as background (0), so as to realize the pixel-level positioning of the license plate area,
[0286] (4.5.3) Finally output the license plate image X2 of the dangerous chemical transport vehicle located;
[0287] (5) Construct the RILPQ-CRNN joint model to enhance the extraction ability of the CRNN model to the license plate features, and realize accurate recognition of the license plate characters of the dangerous chemical transport vehicle;
[0288] As Figure 4As shown, in order to enhance the robustness and representation ability of the CRNN model to the local features of the license plate image, the application proposes to improve the convolution recurrent neural network (CRNN) model by using a rotation-invariant local phase quantization (RILPQ) feature. The RILPQ feature map and the original image are fused and input into the CRNN, which can provide more rich multi-channel information, improve the accuracy and generalization ability of the CRNN in license plate image recognition, classification and other tasks, and the specific steps are:
[0289] (5.1) Image input
[0290] The license plate number image X2 of the positioned hazardous chemical transport vehicle is input;
[0291] (5.2) Image local feature extraction
[0292] The RILPQ algorithm is used to extract features from the input image X2, and the robustness of the features is enhanced by rotation-invariant processing, and the specific steps are:
[0293] (5.2.1) RILPQ feature extraction is used to analyze the local neighborhood frequency domain information by short-time Fourier transform, and rotation and blur invariance is realized by combining direction quantization to extract features of the license plate, and the specific steps are:
[0294] 1) First, the short-time Fourier transform is used to process the m*m neighborhood N of the image X2 x The image is decomposed into components of different frequencies and directions, so that the local frequency domain F of each pixel x is obtained, and the formula is:
[0295]
[0296] Where F(u, x') is the local frequency domain coefficient; f(x'-y') is the pixel gray value at offset y' in the neighborhood; j5 is the imaginary unit; x' is the pixel point of the license plate picture; y' is the neighborhood pixel coordinate; u is the frequency vector, indicating the frequency domain direction; u T is the transpose of u; N x′ is the local field of pixels;
[0297] 2) Then calculate the typical direction ξ(x') according to the complex moment of the quantization coefficient to ensure rotation invariance, and the formula is:
[0298]
[0299] Where b(x') is the complex value of the typical direction of the pixel point x', which is used for rotation alignment; is the i5th complex value of the quantization coefficient; is the phase of the complex value; M is the total number of quantization coefficients; i5 is the summation index;
[0300] 3) Then rotate the neighborhood to the main direction, calculate the LPQ features, and generate a binary code, the formula of which is:
[0301]
[0302] Among them, F ξ (u, x′) are rotation-invariant LPQ features; f(y′) is a function; ξ is the typical orientation of pixel x′, calculated from b(x′); It is the neighborhood weight function after rotation; R ξ(x′) It is a rotation matrix that rotates the neighborhood to the typical direction ξ;
[0303] Finally, 8-bit binary code is obtained through symbolic quantization and converted into decimal feature RILPQ(x′);
[0304] (5.2.2) The RILPQ(x) feature map is mapped to a feature map F′ of the same size as the input image, and then concatenated with the original image X2 to form a multi-channel input. The formula is as follows:
[0305] X3 = Concat(X2, F′)
[0306] Where X3 is the fused multi-channel image; F′ is the feature map generated by RILPQ;
[0307] (5.2.3) Input the fused multi-channel image X3 into the convolutional layer sequence of the CRNN, perform multi-layer convolution and pooling operations, and extract the features enhanced by the RILPQ algorithm. The formula is as follows:
[0308]
[0309] in, It is the output feature map of the L3 layer; These are the convolution weights of the L3 layer; This is the convolutional bias of the L3 layer; L3 is the number of convolutional layers.
[0310] (5.2.4) Then, the feature maps generated by the convolutional layers are processed using Map-to-Sequence to convert the feature maps into feature sequences and extract the feature vector sequence x′=x′1,...,x′ T , as input to the loop layer;
[0311] (5.3) Bidirectional LSTM Sequence Modeling
[0312] A bidirectional LSTM network is used to extract the feature sequence x′=x′1,...,x′ from the convolutional layers. T Each frame in Each generates a label distribution The specific steps are as follows:
[0313] (5.3.1) When the hidden layer receives a frame in the sequence At that time, a nonlinear function is used to define its internal state. The update is performed using the following formula:
[0314]
[0315] in, It is the internal state of the model at time t6, capturing the sequence information up to the current time step; It is the internal state at time t6-1, i.e., the previous internal state; t6 is the time when each frame in the feature sequence is extracted; g6(·) is a nonlinear function;
[0316] (5.3.2) According to Make predictions Thus, the labels for each frame of the loop layer are obtained as y′=y′1,...,y′ T Its formula is:
[0317]
[0318] in, W1 is the predicted label for each frame; W1 is the weight matrix from the hidden layer to the output layer; b y′ It is the bias term of the output layer; softmax is the classification function;
[0319] (5.4) Decoding the CTC Transcription Layer
[0320] The predicted label results are input into the transcription layer and decoded using CTC. Through steps such as removing duplicate characters and whitespace characters, the accurate license plate character recognition result is finally obtained. The specific steps are as follows:
[0321] (5.4.1) The labels y′=y′1,...,y′ of each frame predicted by the recurrent layer are... T Transform the sequence into a label sequence l′, and calculate the conditional probability of l′ using the following formula:
[0322]
[0323] Where p(l′|y′) is the sum of probabilities of all possible paths π that can map to the label sequence l′; l′ is the label sequence; π represents a path that maps to the final label sequence by removing duplicate and blank labels; y′ is the input sequence; The output label is at time t6. The probability of π; T is the sequence length; β is the mapping function, that is, β maps π to l′;
[0324] (5.4.2) The corresponding sequence of the maximum probability string path is output as the optimal sequence, so as to obtain the license plate of the identified dangerous chemical transport vehicle, and the formula is:
[0325] L * = β (arg max π p (π | y′))
[0326] Wherein, L * is the optimal sequence, that is, the obtained dangerous chemical transport vehicle license plate recognition label;
[0327] (6) The license plate recognition result of the dangerous chemical transport vehicle is input into the system, so as to facilitate the effective management and monitoring of the dangerous chemical transport vehicle.
Claims
1. A method for identifying license plate numbers of vehicles transporting hazardous chemicals on highways, characterized in that, Includes the following steps: (1) Construct a DFRSC-MB-Taylor Former V2 joint model to preprocess the acquired vehicle images; (2) The FOCUS model based on the improved MBPSO algorithm is used to accurately identify hazardous chemical transport vehicles in the preprocessed vehicle images. (3) Using the SegVG model, the target region segmentation mask is generated to achieve accurate positioning of the license plate area of hazardous chemical transport vehicles; (4) Construct a RILPQ-CRNN joint model to enhance the CRNN model's ability to extract license plate features and achieve accurate recognition of license plate characters of hazardous chemical transport vehicles.
2. The method for identifying license plate numbers of hazardous chemical transport vehicles on highways according to claim 1, characterized in that, Step (1) includes: (1.1) The DFRSC model was used to correct the acquired highway vehicle images; (1.2) The MB-Taylor Former V2 model is used to perform deblurring enhancement on the corrected vehicle images.
3. The method for identifying license plate numbers of hazardous chemical transport vehicles on highways according to claim 1, characterized in that, Step (1.1) includes: (1.1.1) Input multiple frames of RS images; input N consecutive multi-frame rolling shutter RS images X captured by a highway camera. r ∈R N×H×W×3 Where N is the number of RS frames; H is the original RS image height; W is the original RS image width; and R is the set of real numbers. (1.1.2) Multi-scale feature extraction; A weight-shared image encoder is used to extract multi-scale frame-level RS features from N consecutive RS frames, using the following formula: Among them, F l This is an RS multi-scale feature; K is the number of feature layers; l is the feature level; (1.1.3) Initial distortion flow estimation: The distortion flow from the GS image to the RS image is jointly estimated at the lowest resolution layer through a global correlation attention mechanism; (1.1.4) Stepwise refinement: Input the initial distortion stream and GS features predicted at the lowest resolution into the decoder, perform layer-by-layer upsampling and feature fusion; (1.1.5) Multi-distortion stream decoding; At the highest resolution level, a distortion flow field prediction strategy is adopted to distort the RS features to the GS space and predict multiple distortion flow fields. The formula is as follows: {And 1,r←g ,L,U G,r←g }=ConvBlock(U r←g ,F g ,F warped ) Among them, U 1,r←g L, U G,r←g It is a prediction of multiple distorted flow fields; U r←g It is a refined distortion stream; F g It is a GS feature; F warped It is a distorted RS feature; (1.1.6) GS Image Generation: The refined GS features, multiple sets of distortion streams, and warped RS features are input into the decoder to obtain the corrected GS image X of the transport vehicle captured by the highway camera. g .
4. The method for identifying license plate numbers of hazardous chemical transport vehicles on highways according to claim 1, characterized in that, Step (1.2) includes: (1.2.1) Input the image X of highway transport vehicles after correction using the DFRSC model. g ∈R 3×h×w ; (1.2.2) Shallow feature extraction is performed through convolutional layers to generate initial feature F0, the formula of which is: F0=Conv(X g )=R c×h×w Where F0 is the initial feature; c is the number of channels; h is the height of the initial feature image; w is the width of the initial feature image; and Conv is the convolution operation. (1.2.3) Multi-branch depthwise separable deformable convolution is used to generate image features of different scales from the initial feature F0; (1.2.4) Features of different scales are input into multiple Transformer encoder branches. Each branch adopts a linear self-attention mechanism to reduce computational complexity, while dynamic position coding is incorporated to enhance spatial awareness, thereby realizing parallel processing and fusion of multi-scale features. (1.2.5) The SKFF module is used to fuse the features output by the branches of the Transformer encoder according to the feature weights to obtain the fused features; (1.2.6) The fused branch features are downsampled using a pixel-wise unscrambling operation to obtain deep multi-scale features. Then, the decoder upsamples them using a pixel-wise unscrambling operation to restore the resolution of the deep multi-scale features to the original resolution. (1.2.7) The features of the encoder at the current stage are combined with the upsampled features by channel splicing and convolutional fusion. (1.2.8) By using residual blocks, the spliced features are further refined into feature maps, and structural and texture details are optimized to generate high-quality repair feature maps; (1.2.9) The number of channels in the repaired feature map is reduced from C to 3 using a 3×3 convolutional layer to generate a residual image R. The residual image R is then superimposed onto the input image X. g ∈R 3×h×w Thus, the restored and enhanced image is obtained, using the following formula: X=X g +R; Where X is the image enhanced using the MB-Taylor Former V2 model; R is the residual image.
5. The method for identifying license plate numbers of hazardous chemical transport vehicles on highways according to claim 1, characterized in that, Step (2) includes: (2.1) Image segmentation; The preprocessed image x is segmented into N non-overlapping blocks. Each image block contains local features of the vehicle, sky, and road elements. The segmented image blocks are constructed in the form of a set, as shown in the formula: Where X is the preprocessed input image; x is the segmented image patch; N2 is the number of image patches; (2.2) Image feature extraction: The FOCUS model is used to extract features from the photos taken by the highway camera after preprocessing, and the MBPSO algorithm is used to select the extracted image features and retain the optimal feature subset. (2.3) A global similarity filtering layer is used to remove the background from the segmented image and retain the potential vehicle region; The optimized image patch features B1 are used to calculate the similarity between image patches using a sliding window mechanism, and large-scale repetitive backgrounds are filtered out based on the similarity threshold to retain potential vehicle regions. (2.4) Use a text-guided key block selection layer to extract key visual features of hazardous chemical transport vehicles; (2.5) Adjacent block compression layers are used to remove multi-view duplicate features of the same hazardous chemical transport vehicle; (2.6) To enhance the model's ability to discriminate the characteristics of hazardous chemical transport vehicles, the compressed visual features B... c The text prompt T is fused with the text prompt T through a multi-head attention mechanism, as shown in the formula: Where Q = TW q K = B c W k′ , in, It is a multi-head attention mechanism; Q is the text prompt as the query; K is the compressed block feature as the key; V is the compressed block feature as the value; W is the i-th learnable projection matrix; q It is a learnable projection matrix, W k′ It is a learnable projection matrix used to project image patch features into the key space; It is a value projection matrix; d k′ It is the dimension of each attention head; After concatenating and normalizing the outputs of each attention point, the formula is: O=LayerNorm(Concat(Head1,...,Head h′ )W o ) Where O represents the normalized, concatenated multi-head attention mechanism; W o It is a learnable matrix; LayerNorm is layer normalization; Concat is a join operation; (2.7) Predict whether an image contains a vehicle carrying hazardous materials using a classifier. If the probability is greater than 0.5, the image patch is determined to contain a vehicle carrying hazardous materials. The formula is as follows: in, Y is the probability of the output vehicle type; Y is the category label of the classification task, Y=0 indicates that the vehicle in the image is a regular vehicle, Y=1 indicates that the vehicle in the image is a hazardous materials transport vehicle; It is a learnable matrix; It is a learnable vector; (2.8) Spatial reconstruction is performed on the image block features of the hazardous chemical transport vehicle retained after classification, and the features are remapped back to the spatial position of the original image to obtain the complete image X1 of the hazardous chemical transport vehicle.
6. The method for identifying license plate numbers of hazardous chemical transport vehicles on highways according to claim 1, characterized in that, Step (2.2) includes: (2.2.1) The image blocks are segmented and their features are extracted using a pre-trained vehicle detection model. The formula is as follows: Among them, the knife is the set of extracted features; f is a pre-trained vehicle detection model that extracts block features; is the i1-th image block; is the feature vector of the i1-th block; (2.2.2) The MBPSO algorithm is used to select and optimize the extracted image features B through multi-cell collaboration. The tissue membrane system typically consists of three cells: σ1 and σ2 are used for local search of the initial image features B, and σ3 is used to find the globally optimal subset of image features. The formula is as follows: Π=(O, σ1, σ2, σ3, syn, i0) Where ∏ is the tissue membrane system; O is the set of particles encoded in binary; σ1 and σ2 are the initial particle swarm cells, each managing a set of candidate feature subsets, and each candidate feature subset is a binary vector; σ3 is the globally optimal candidate feature subset; syn is the communication rule that enables information exchange between feature subsets; and i0 is the feature subset label that outputs the final result. Initial candidate feature subsets are randomly generated from σ1 and σ2, and the position of each candidate feature subset is... It is a binary vector, and its velocity is... Initialize to random values; The speed update formula for each candidate feature subset is: in, ω is the velocity of the i2th candidate feature subset at t2+1; ω is the inertia weight, which controls the influence of historical velocities; c1 is the learning factor, representing the individual optimal weight; pbest is the individual optimal value. c2 is the state of the i2th candidate feature subset at time t; c2 is the learning factor, representing the weight of the global optimum; gbest is the global optimum; i2 is the number of candidate feature subsets; t2 is the number of iterations; rand is a random number in [0, 1]. The positions of the candidate feature subset are updated, and the velocity is mapped to probability using the Sigmoid function to determine whether a feature is selected. The formula is as follows: in, It is the j2-th dimension feature selection state of the i2-th candidate feature subset; It is the speed of selecting the j2nd dimension feature of the i2th candidate feature subset; j2 is the feature dimension number. σ3 receives all the individual optimal feature subsets pbest, and selects the one with the highest fitness as the globally optimal feature subset gbest. The formula is as follows: r 32 :s 3,k+1 (f 3,k+1 w 3,k+1 )→s 3,k+l {(w 3,k+1 ,go),(f 3,k+l ,go)} Where, r 32 It is the rule of cellular σ3; s 3,k+1 It is the new state of σ3 in the (k+1)th iteration; f 3,k+1 It is the globally optimal fitness value after the cell's σ3 update; w 3,k+1 It is the new set of positions and velocities of the feature sub-clusters after σ3 update; When the maximum number of iterations t is reached max When gbest is selected, the binary vector corresponding to gbest is the optimal feature subset B1, and its formula is: B1 = gbest B1 is the image feature extracted after optimizing the initial feature B using the MBPSO algorithm.
7. The method for identifying license plate numbers of hazardous chemical transport vehicles on highways according to claim 1, characterized in that, Step (3) includes: (3.1) Input image X1 into the pre-trained DETR model, and extract the visual features of the image using ResNet and Transformer encoders. The formula is as follows: in, These are the original visual features of layer i4+1; These are the original visual features of layer i4; This is the i4th layer in the DETR backbone network; i4 is the total number of layers; v4 represents vision. (3.2) The BERT model is used to process the text to obtain the text feature vector. The formula is: in, These are the original text features of layer i4+1; t5 represents the original text features of layer i4; t5 represents the text. It is layer 2i4+1 in the DETR backbone network; It is layer 2i4 in the DETR backbone network; (3.3) Introducing the Triple Alignment module, which iteratively updates query, text, and visual features through a three-way attention mechanism; (3.4) Align the text feature vectors and visual feature vectors Merge and connect to generate multimodal feature vectors; (3.5) Input the encoder output into the decoder. The decoder uses the bbox2seg scheme to convert the target box annotations into a segmentation mask.
8. The method for identifying license plate numbers of hazardous chemical transport vehicles on highways according to claim 1, characterized in that, Step (3.5) includes: (3.5.1) Use the regression query vector to perform bounding box regression on the multimodal features output by the encoder, and determine the position range of the dial in the image through the Transformer decoder layer and the MLP network; (3.5.2) The target region is segmented using segmentation query vectors with different learnable positional codes. Similarly, the segmentation mask of the target region is obtained through the Transformer decoder layer and the MLP network. Pixels within the license plate area are marked as foreground (1), and pixels outside the area are marked as background (0). (3.5.3) The final output is the license plate image X2 of the located hazardous chemical transport vehicle.
9. The method for identifying license plate numbers of hazardous chemical transport vehicles on highways according to claim 1, characterized in that, Step (4) includes: (4.1) Image input; Input the license plate image X2 of the located hazardous chemical transport vehicle; (4.2) Image local feature extraction; The RILPQ algorithm is used to extract features from the input image X2; (4.3) Bidirectional LSTM sequence modeling; a bidirectional LSTM network is used to model the feature sequences x′=x′1,...,x′ extracted by the convolutional layers. T Each frame in Each generates a label distribution (4.4) CTC transcription layer decoding: The results of the predicted labels are input into the transcription layer and decoded by CTC. By removing duplicate characters and whitespace characters, the accurate license plate character recognition results are finally obtained.
10. The method for identifying license plate numbers of hazardous chemical transport vehicles on highways according to claim 1, characterized in that, Step (4.2) includes: (4.2.1) RILPQ feature extraction is adopted. Local neighborhood frequency domain information is analyzed by short-time Fourier transform, and rotation and fuzziness invariance are achieved by combining direction quantization to extract features from license plates. (4.2.2) The RILPQ(x) feature map is mapped to a feature map F′ of the same size as the input image, and then concatenated with the original image X2 to form a multi-channel input. The formula is as follows: X3 = Concat(X2, F′) Where X3 is the fused multi-channel image; F′ is the feature map generated by RILPQ; (4.2.3) The fused multi-channel image X3 is input into the convolutional layer sequence of the CRNN, and multiple convolution and pooling operations are performed to extract the features enhanced by the RILPQ algorithm. The formula is as follows: in, It is the output feature map of the L3 layer; These are the convolutional weights of the L3 layer; This is the convolutional bias of the L3 layer; L3 is the number of convolutional layers. (4.2.4) Perform Map-to-Sequence processing on the feature maps generated by the convolutional layers to convert the feature maps into feature sequences and extract the feature vector sequence x′=x′1,...,x′ T , as input to the loop layer.