Vehicle multi-attribute enhanced analysis and intrusion behavior recognition method and device
Through Faster R-CNN and Mask R-CNN combined with ByteTrack tracking algorithm and Transformer attention mechanism, the accuracy and robustness of vehicle multi-attribute recognition and break-in behavior detection in complex scenarios are solved, and efficient vehicle multi-attribute analysis and break-in behavior recognition are achieved.
Patent Information
- Application Number
- CN202510594452.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to achieve fine analysis of vehicle multi-attributes and accurate identification of intrusion behavior in complex scenarios, especially in the case of light changes, target occlusion and multi-vehicle overlapping, with low recognition accuracy and poor robustness.
Faster R-CNN and Mask R-CNN deep learning algorithms are used for vehicle detection and fine-grained segmentation, combined with ByteTrack tracking algorithm and Transformer attention mechanism, the training model is enhanced through Mixup and Mosaic data, and the GIoU loss function is optimized to improve recognition accuracy and stability.
It realizes high-precision vehicle multi-attribute recognition and break-in behavior detection in complex scenarios, meets real-time requirements, and improves the recognition accuracy and robustness of the model in lighting changes, target occlusion and multi-vehicle overlapping scenarios.
Smart Images

Figure CN120259979A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and specifically provides a method and device for enhanced analysis of multiple vehicle attributes and recognition of intrusion behavior. Background Art
[0002] With the rapid development of intelligent monitoring systems, the demand for vehicle attribute recognition and behavior analysis in the fields of security and traffic management is increasing day by day. Traditional vehicle monitoring methods mainly rely on license plate recognition technology, but in complex scenarios (such as changes in lighting, target occlusion, multi-vehicle overlap, etc.), there are problems of low recognition accuracy and poor robustness.
[0003] Although there are already some vehicle detection and recognition methods in the prior art, these methods usually only focus on a single attribute (such as license plate number) or simple tracking functions, and it is difficult to achieve fine analysis of multiple vehicle attributes (such as color, type, brand, etc.) and accurate recognition of intrusion behavior. Summary of the Invention
[0004] The present invention aims at the above-mentioned deficiencies of the prior art and provides a method for enhanced analysis of multiple vehicle attributes and recognition of intrusion behavior with strong practicability.
[0005] A further technical task of the present invention is to provide a device for enhanced analysis of multiple vehicle attributes and recognition of intrusion behavior with reasonable design, safety and applicability.
[0006] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0007] A method for enhanced analysis of multiple vehicle attributes and recognition of intrusion behavior has the following steps:
[0008] S1. Vehicle detection and preliminary analysis of multiple attributes;
[0009] S2. Vehicle fine-grained segmentation and detailed attribute analysis;
[0010] S3. Vehicle intrusion recognition and tracking;
[0011] S4. Vehicle feature enhancement;
[0012] S5. Real-time data enhancement and model training.
[0013] Further, in step S1, it includes:
[0014] S1-1. Extract the vehicle position and preliminary attributes through the object detection Faster R-CNN and multi-label classification network. Use the Faster R-CNN model to detect vehicles in the input image and output the detection results of vehicle positions and confidence levels. Faster R-CNN includes a Region Proposal Network (RPN) and a detection network. The RPN generates candidate regions, and the detection network classifies and regresses the bounding boxes of the candidate regions.
[0015] The RPN generates anchor points on the feature map through a sliding window and predicts the object probability and bounding box offset for each anchor point. The formula is:
[0016]
[0017] where, is the feature map of the input image I, W cls is the classification weight, and σ is the Sigmoid function.
[0018] Bounding box regression:
[0019] t x =(x - x a ) / w a , t y =(y - y a ) / h a
[0020] t w = log(w - w a ), t h = log(h - h a )
[0021] where, (x, y, w, h) are the coordinates of the ground truth box, and (x a , y a , w a , h a ) are the coordinates of the anchor point.
[0022] Detection network output:
[0023] The detection network classifies and regresses the bounding box for each candidate region and outputs the detection result:
[0024] P = Faster R-CNN(I);
[0025] where, P contains the detected vehicle position and its confidence level.
[0026] S1-2. Based on the preliminary classification of attributes from the detection results, use the detected vehicle position information to crop the vehicle region and perform preliminary attribute classification on the vehicle through a multi-label classification network.
[0027] Further, in step S1-2, the output layer of the classification network uses the SoftMax function to calculate the confidence of each category. The formula is as follows:
[0028]
[0029] where x vehicle represents the feature of the cropped vehicle area, y i represents the i-th attribute label, and z i and C respectively represent the original score of the i-th category and the total number of categories.
[0030] Further, in step S2, it includes:
[0031] S2-1. Use Mask R-CNN to perform fine-grained segmentation on the detected vehicle area to obtain the precise boundary and internal area of the vehicle, deepening and supplementing step S1. It is expressed as:
[0032] M = Mask R-CNN(x vehicle )
[0033] where M is the pixel-level segmentation result and detailed attribute information of the vehicle;
[0034] S2-2. Extract detailed attributes based on the segmentation result: Use the segmentation result of Mask R-CNN to further extract the detailed attributes of the vehicle. Feature extraction is implemented through the convolutional neural network CNN. The formula is:
[0035]
[0036] where I segmented represents the segmented vehicle area image, K is the convolutional kernel, and 0 is the output feature map.
[0037] Further, in step S3, it includes:
[0038] S3-1. Track the vehicles in consecutive frames using the ByteTrack algorithm. The formula is:
[0039]
[0040] where and respectively represent the predicted state at time k and the estimated state at time k-1, and F is the state transition matrix;
[0041] S3-2. Intrusion judgment and alarm. According to the tracking result, judge whether the vehicle has intruded into a preset restricted area. If the vehicle position coordinates (x, y) fall within the restricted area, an alarm is triggered. The formula is:
[0042]
[0043] Further, in step S4, it includes:
[0044] S6-1. Adopt the attention mechanism of Transformer to capture the global dependencies between different regions in the image, enhance the representation ability of target features, and the formula is:
[0045]
[0046] where Q, K, and V respectively represent Query, Key, and Value, and d k is the dimension of the key;
[0047] S6-2. Adopt the encoder-decoder structure of U-Net to extract multi-level features. The encoder-decoder structure fuses shallow and deep features through skip connections to further improve the accuracy of feature extraction;
[0048] S6-3. Adopt the GIoU loss function to optimize the positioning accuracy of the detection box, and the formula is:
[0049]
[0050] where A and B represent the predicted box and the ground truth box, and C represents the smallest rectangle box containing A and B.
[0051] Further, in step S5, it includes:
[0052] S5-1. Apply the Mixup and Mosaic algorithms to perform real-time enhancement on the training data to improve the generalization ability of the model. The Mixup formula is:
[0053]
[0054] where x i , x j are two randomly selected images, y i , y j are the corresponding labels, and λ is the mixing ratio;
[0055] S5-2. Use the enhanced data to train the model, and evaluate the model performance through the validation set, and adjust the model parameters to optimize the performance.
[0056] A vehicle multi-attribute enhanced analysis and intrusion behavior recognition device includes: at least one memory and at least one processor;
[0057] The at least one memory is used to store machine-readable programs;
[0058] The at least one processor is configured to call the machine-readable program to execute a method for enhanced analysis of multiple vehicle attributes and intrusion behavior recognition.
[0059] Compared with the prior art, a method and device for enhanced analysis of multiple vehicle attributes and intrusion behavior recognition according to the present invention have the following prominent beneficial effects:
[0060] (1) By integrating deep learning algorithms such as object detection models and Mask R-CNN, it is possible to accurately identify multiple attribute information of vehicles and maintain a high recognition accuracy in complex scenarios (such as light changes, target occlusion, etc.).
[0061] (2) By adopting the ByteTrack object tracking algorithm, it is possible to efficiently track vehicles in consecutive frames and trigger alarms in real time in combination with intrusion judgment logic, meeting the requirements for real-time performance in practical applications.
[0062] (3) Through data augmentation and loss function optimization, the generalization ability and stability of the model in complex scenarios are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0064] Att Figure 1 is a schematic flowchart of a method for enhanced analysis of multiple vehicle attributes and intrusion behavior recognition. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will further elaborate on the present invention in conjunction with specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0066] The following gives a preferred embodiment:
[0067] As Figure 1 shown, a method for enhanced analysis of multiple vehicle attributes and intrusion behavior recognition in this embodiment has the following steps:
[0068] S1. Vehicle detection and preliminary analysis of multiple attributes;
[0069] Including:
[0070] Through object detection (Faster R-CNN) and multi-label classification network, the extraction of vehicle position and preliminary attributes is realized. The Faster R-CNN model is used to detect vehicles in the input image, and the detection results including vehicle position (bounding box coordinates) and confidence are output. Faster R-CNN includes a Region Proposal Network (RPN) and a detection network. The RPN generates candidate regions, and the detection network classifies and regresses the bounding boxes of the candidate regions.
[0071] First, the RPN generates anchor points on the feature map through a sliding window and predicts the object probability and bounding box offset of each anchor point. The formula is:
[0072]
[0073] where, is the feature map of the input image I, W cls is the classification weight, and σ is the Sigmoid function.
[0074] Then, bounding box regression is performed:
[0075] t x =(x - x a ) / w a , t y =(y - y a ) / h a
[0076] t w =log(w - w a ), t h =log(h - h a )
[0077] where, (x, y, w, h) are the coordinates of the ground truth box, and (x a , y a , w a , h a ) are the coordinates of the anchor point.
[0078] Finally, the detection network output is performed:
[0079] The detection network classifies and regresses the bounding boxes of each candidate region and outputs the detection results:
[0080] P = Faster R-CNN(I)
[0081] where, P contains the detected vehicle position and its confidence.
[0082] S1-2. Preliminary classification of attributes based on detection results. Using the detected vehicle position information, the vehicle area is cropped, and a multi-label classification network (implemented based on Keras) is used to perform preliminary attribute classification on the vehicle. The output layer of the classification network uses the SoftMax function to calculate the confidence of each category. The formula is as follows:
[0083]
[0084] Among them, x vehicle represents the feature of the cropped vehicle area, y i represents the i-th attribute label, and z i and C respectively represent the original score of the i-th category and the total number of categories.
[0085] S2. Fine-grained segmentation of the vehicle and detailed attribute analysis;
[0086] Including:
[0087] S2-1. Using Mask R-CNN to perform fine-grained segmentation on the detected vehicle area, obtaining the precise boundary and internal area of the vehicle, and deepening and supplementing the first step. Mask R-CNN fine-grained segmentation: It is an instance segmentation model that can accurately segment the vehicle area and extract detailed attributes. Its pixel-level segmentation ability performs excellently in complex scenarios (such as target occlusion) and has high reliability.
[0088] It can be expressed as:
[0089] M = Mask R-CNN(x vehicle )
[0090] Among them, M is the pixel-level segmentation result and detailed attribute information of the vehicle.
[0091] S2-2. Detailed attribute extraction based on the segmentation result: Using the segmentation result of Mask R-CNN, further extract the detailed attributes of the vehicle, such as vehicle type, color, etc. Feature extraction is achieved through a convolutional neural network (CNN). The formula is:
[0092]
[0093] Among them, I segmented represents the segmented vehicle area image, K is the convolutional kernel, and 0 is the output feature map.
[0094] S3. Vehicle intrusion recognition and tracking;
[0095] Including:
[0096] S3-1. Track the vehicles in consecutive frames using the ByteTrack algorithm, which is based on DeepSORT and Kalman filtering and can achieve stable object tracking in complex scenarios (such as multiple vehicle overlaps). The formula is:
[0097]
[0098] Among them, and represent the predicted state at time k and the estimated state at time k-1 respectively, and F is the state transition matrix.
[0099] S3-2. According to the tracking results, determine whether the vehicle has entered a preset restricted area. If the vehicle position coordinates (x, y) fall within the restricted area, an alarm will be triggered. The formula is:
[0100]
[0101] S4. Vehicle feature enhancement;
[0102] Including:
[0103] S4-1. Adopt the attention mechanism of Transformer, which can capture the global dependencies between different regions in the image and enhance the representation ability of target features. In complex scenarios (such as target occlusion or multiple vehicle overlaps), the attention mechanism can focus on the key regions of the target, reduce background interference, and thus improve the robustness of feature extraction. The formula is:
[0104]
[0105] Among them, Q, K, and V represent Query, Key, and Value respectively, and d k is the dimension of the key.
[0106] S4-2. The encoder-decoder structure of U-Net extracts multi-level features and restores spatial information through the decoder, which can effectively handle target occlusion and multiple vehicle overlap problems. In scenes with light changes, the multi-scale feature extraction ability of U-Net can enhance the recognition ability of the target. The encoder-decoder structure fuses shallow and deep features through skip connections to further improve the accuracy of feature extraction.
[0107] S4-3. Adopt the GIoU loss function to optimize the positioning accuracy of the detection box. Especially in the target overlap scenario, it can more accurately locate the target boundary. The formula is:
[0108]
[0109] Among them, A and B represent the predicted bounding box and the ground truth bounding box, and C represents the smallest rectangle enclosing A and B.
[0110] S5, Real-time Data Augmentation and Model Training;
[0111] Including:
[0112] S5-1, Mixup and Mosaic Data Augmentation: Apply the Mixup and Mosaic algorithms to perform real-time augmentation on the training data to improve the generalization ability of the model. The Mixup formula is:
[0113]
[0114] Among them, x i , x j are two randomly selected images, y i , y j are the corresponding labels, and λ is the mixing ratio.
[0115] S5-2, Model Training and Validation: Use the augmented data to train the model, and evaluate the model performance through the validation set, and adjust the model parameters to optimize the performance.
[0116] Based on the above method, a vehicle multi-attribute enhancement analysis and intrusion behavior recognition device in this embodiment includes: at least one memory and at least one processor;
[0117] At least one memory for storing machine-readable programs;
[0118] At least one processor for calling the machine-readable program to execute a vehicle multi-attribute enhancement analysis and intrusion behavior recognition method.
[0119] The above specific implementation manners are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific implementation manners. Any technical solution that conforms to the technical solutions described in the above specific implementation manners of the present invention and any appropriate changes or replacements made by those of ordinary skill in the art shall fall within the patent protection scope of the present invention.
[0120] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirits of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for multi-attribute enhanced analysis and intrusion behavior recognition of a vehicle, characterized in that It has the following steps: S1. Vehicle detection and preliminary multi-attribute analysis; S2. Vehicle fine-grained segmentation and detailed attribute analysis; S3. Vehicle intrusion recognition and tracking; S4. Vehicle feature enhancement; S5. Real-time data enhancement and model training.
2. The vehicle multi-attribute enhanced analysis and intrusion behavior recognition method according to claim 1, wherein In step S1, it includes: S1-1. Through the object detection Faster R-CNN and multi-label classification network, realize the extraction of vehicle position and preliminary attributes. Use the Faster R-CNN model to detect vehicles in the input image and output the detection results of vehicle position and confidence. Faster R-CNN includes a Region Proposal Network (RPN) and a detection network. The RPN generates candidate regions, and the detection network classifies and performs bounding box regression on the candidate regions; The RPN generates anchor points on the feature map through a sliding window and predicts the object probability and bounding box offset of each anchor point. The formula is: Among them, is the feature map of the input image I, and W cls is the classification weight, and σ is the Sigmoid function; Bounding box regression: t x = (x - x a ) / w a , t y = (y - y a ) / h a t w = log(w - w a ), t h = log(h - h a ) Among them, (x, y, w, h) are the coordinates of the ground truth box, and (x a , y a , w a , h a ) are the coordinates of the anchor point; The detection network output: The detection network classifies and performs bounding box regression on each candidate region and outputs the detection result: P = Faster R-CNN(I); where P contains the detected vehicle position and its confidence; S1-2. Preliminary attribute classification based on the detection result. Use the detected vehicle position information to crop the vehicle region and perform preliminary attribute classification on the vehicle through a multi-label classification network.
3. A method for multi-attribute enhanced analysis and intrusion behavior recognition of a vehicle according to claim 2, characterized in that, In step S1-2, the output layer of the classification network uses the SoftMax function to calculate the confidence of each category. The formula is as follows: Among them, x vehicle represents the vehicle area feature cut out, y i represents the i-th attribute label, z i and C respectively represent the original score and the total number of categories of the i-th category.
4. A method for vehicle multi-attribute enhanced analysis and intrusion behavior recognition according to claim 3, characterized in that, In step S2, it includes: S2-1. Use Mask R-CNN to perform fine-grained segmentation on the detected vehicle region to obtain the precise boundary and internal region of the vehicle, deepening and supplementing step S1. It is expressed as: M = Mask R-CNN(x vehicle ) where M is the pixel-level segmentation result and detailed attribute information of the vehicle; S2-2. Detailed attribute extraction based on the segmentation result: Use the segmentation result of Mask R-CNN to further extract the detailed attributes of the vehicle. Feature extraction is realized through a Convolutional Neural Network (CNN). The formula is: Among them, I segmented represents the segmented vehicle area image, K is the convolution kernel, and 0 is the output feature map.
5. A method for multi-attribute enhanced analysis and intrusion behavior recognition of a vehicle according to claim 4, characterized in that, In step S3, it includes: S3-1. Track the vehicles in consecutive frames using the ByteTrack algorithm. The formula is: wherein, and respectively represent the predicted state at time k and the estimated state at time k-1, and F is the state transition matrix; S3-2. Intrusion judgment and alarm. According to the tracking result, judge whether the vehicle intrudes into a preset restricted area. If the vehicle position coordinates (x, y) fall within the restricted area, an alarm is triggered. The formula is:
6. A method for multi-attribute enhanced analysis and intrusion behavior recognition of a vehicle according to claim 5, characterized in that, In step S4, it includes: S6-1. Adopt the attention mechanism of Transformer to capture the global dependencies between different regions in the image and enhance the representation ability of target features. The formula is: Among them, Q, K, and V represent Query, Key, and Value respectively, and d k is the dimension of the key; S6-2. Adopt the encoder-decoder structure of U-Net to extract multi-level features. The encoder-decoder structure fuses shallow and deep features through skip connections to further improve the accuracy of feature extraction; S6-3. Adopt the GIoU loss function to optimize the positioning accuracy of the detection box. The formula is: where A and B represent the predicted box and the ground truth box, and C represents the smallest rectangle box containing A and B.
7. A method for vehicle multi-attribute enhanced analysis and intrusion behavior recognition according to claim 6, characterized in that In step S5, it includes: S5-1. Apply the Mixup and Mosaic algorithms to perform real-time enhancement on the training data to improve the generalization ability of the model. The Mixup formula is as follows: where x i , x j are two randomly selected images, y i , y j are the corresponding labels, and λ is the mixing ratio; S5-2. Use the enhanced data to train the model, and evaluate the model performance through the validation set, and adjust the model parameters to optimize the performance.
8. A vehicle multi-attribute enhanced analysis and intrusion behavior recognition device, characterized in that, Including: At least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program and execute the method according to any one of claims 1 to 7.