A vehicle violation detection method, device and processing equipment based on multimodal
By integrating vehicle detection frames, wheel contact points, and highway information into a multimodal detection model, the problem of misjudgment in drone violation detection is solved, and high-precision vehicle violation detection is achieved.
Patent Information
- Application Number
- CN202311239324.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-22
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2043-09-22
AI Technical Summary
Existing drone-based highway violation detection methods have the problem of insufficient accuracy when judging whether a vehicle has committed a traffic violation, especially due to interference from the drone's perspective and road conditions, which leads to misjudgment.
A multimodal detection model is adopted, including a first detection model for detecting vehicles and obtaining vehicle detection frames, a second detection model for detecting the contact points of wheels on the ground, and a third detection model for detecting highway information. These models are integrated to determine whether the vehicle has violated traffic regulations.
It achieves high-precision detection of vehicle violations, reduces misjudgments, and improves the accuracy and sophistication of detection.
Smart Images

Figure CN117274840B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of vehicle violation detection, and specifically to a multimodal vehicle violation detection method, device and processing equipment. Background Art
[0002] In recent years, with the development of urbanization and industrialization, traffic volume and the number of cars have continued to increase, and traffic violations on highways have become increasingly prominent. Traditional fixed cameras and sensors have limitations in violation detection, such as limited coverage and poor real-time performance. To improve the efficiency and accuracy of violation detection, the transportation industry needs to seek more advanced technologies to regulate and manage traffic order and ensure safe and smooth traffic.
[0003] Highway inspections and road violation detection are limited by current methods, such as: 1. Limited field of view, which only allows monitoring of specific areas; 2. Difficulty adapting to changes in traffic environments, such as changes in traffic volume and number of lanes on different road sections; 3. Data transmission and processing usually require a certain amount of time, which may affect the real-time nature of violation detection. In this context, the use of drones and supporting highway intelligent inspection strategies has become a relatively effective solution.
[0004] While studying the existing drone-based intelligent inspection strategy for highways, the present applicant found that when judging whether there is a traffic violation based on images collected by drones, the judgment target is usually achieved by whether the image area of the vehicle detected in the image is in the image area of the emergency lane or whether the image areas of the two overlap. However, this type of judgment strategy still has the problem of insufficient accuracy in details. The drone's perspective, highway road conditions or vehicle conditions are all unstable, which will bring a certain degree of interference to image detection, thereby leading to incorrect judgment. Summary of the Invention
[0005] The present application provides a multimodal vehicle violation detection method, apparatus, and processing equipment, which are used to determine the key contact point information of the vehicle wheels on the ground based on the target highway image collected by a drone and the detected vehicle detection frame, and then determine whether there is a vehicle violation in combination with the detected target highway information, and perform multimodal and detailed vehicle violation detection processing, thereby achieving high-precision vehicle violation detection effect.
[0006] In a first aspect, the present application provides a multimodal vehicle violation detection method, the method comprising:
[0007] Obtaining a target highway image of a vehicle violation to be detected, wherein the target highway image is specifically acquired by a drone;
[0008] Inputting the target highway image into a pre-configured multimodal detection model for detection processing, wherein the multimodal detection model includes a first detection model, a second detection model, and a third detection model. The first detection model is used to detect vehicles in the input image and obtain vehicle detection frames. The second detection model is used to detect contact points of wheels on the ground in a vehicle screenshot obtained from the input image using the vehicle detection frames and obtain contact key point information. The third detection model is used to detect highways in the input image and obtain highway information.
[0009] According to the target contact key point information and target highway information output by the multimodal detection model, it is judged whether the corresponding vehicle has committed a vehicle violation.
[0010] In a second aspect, the present application provides a multimodal vehicle violation detection device, comprising:
[0011] an acquisition unit, configured to acquire an image of a target highway where a vehicle violation is to be detected, wherein the image of the target highway is specifically acquired by a drone;
[0012] a detection unit configured to input a target highway image into a preconfigured multimodal detection model for detection processing, wherein the multimodal detection model includes a first detection model, a second detection model, and a third detection model; the first detection model is configured to detect vehicles in the input image and obtain a vehicle detection frame; the second detection model is configured to detect contact points of wheels on the ground in a vehicle screenshot obtained from the input image via the vehicle detection frame and obtain contact key point information; and the third detection model is configured to detect highways in the input image and obtain highway information;
[0013] The judgment unit is used to judge whether the corresponding vehicle has violated traffic regulations based on the target contact key point information and target highway information output by the multimodal detection model.
[0014] In a third aspect, the present application provides a processing device comprising a processor and a memory, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, the method provided in the first aspect of the present application or any possible implementation of the first aspect of the present application is executed.
[0015] In a fourth aspect, the present application provides a computer-readable storage medium, which stores multiple instructions, and the instructions are suitable for a processor to load to execute the method provided in the first aspect of the present application or any possible implementation of the first aspect of the present application.
[0016] From the above content, it can be concluded that this application has the following beneficial effects:
[0017] For the target of vehicle violation detection based on drones, the present application configures a multimodal detection model, which includes a first detection model, a second detection model and a third detection model. The first detection model is used to detect vehicles in the input image and obtain a vehicle detection frame. The second detection model is used to detect the contact points of the wheels on the ground in the vehicle screenshot obtained from the input image through the vehicle detection frame and obtain contact key point information. The third detection model is used to detect highways in the input image and obtain highway information. Based on this model fusion architecture, starting from the target highway image collected by the drone, the target contact key point information of the vehicle wheels on the ground can be determined based on the detected vehicle detection frame. Then, combined with the detected target highway information, it is determined whether there is a vehicle violation, and multimodal and delicate vehicle violation detection processing is performed, thereby achieving high-precision vehicle violation detection effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0019] Figure 1 A flowchart of a multi-modal vehicle violation detection method according to the present invention;
[0020] Figure 2 A schematic diagram of a scenario for drone image acquisition in this application;
[0021] Figure 3 A schematic diagram of a scene before and after the sample image of this application is annotated;
[0022] Figure 4 A schematic diagram of a scenario for preprocessing vehicle screenshots for this application;
[0023] Figure 5 A schematic diagram of a scenario for key contact point detection and processing in this application;
[0024] Figure 6 This is a schematic diagram of a scenario for highway semantic segmentation in this application;
[0025] Figure 7 A schematic diagram of a scenario for detecting and processing vehicle violations in this application;
[0026] Figure 8 A schematic diagram of a scenario in the detection area of this application;
[0027] Figure 9 This is a structural diagram of a multi-modal vehicle violation detection device of the present application;
[0028] Figure 10 This is a structural diagram of the processing equipment for this application. DETAILED DESCRIPTION
[0029] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0030] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or that are inherent to these processes, methods, products or devices. The naming or numbering of steps in this application does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The process steps that have been named or numbered can be changed in the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved.
[0031] The division of modules in this application is a logical division. In actual application, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection between modules can be electrical or other similar forms, which are not limited in this application. Moreover, the modules or submodules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed into multiple circuit modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application.
[0032] Before introducing the multimodal vehicle violation detection method provided by this application, the background content involved in this application is first introduced.
[0033] The multimodal vehicle violation detection method, device and computer-readable storage medium provided in this application can be applied to processing equipment to start from the target highway image collected by the drone, and based on the detected vehicle detection frame, determine the key contact point information of the vehicle wheels on the ground, and then combine the detected target highway information to determine whether there is a vehicle violation, and perform multimodal and detailed vehicle violation detection processing, thereby achieving high-precision vehicle violation detection effect.
[0034] The multimodal vehicle violation detection method mentioned in this application can be implemented by a multimodal vehicle violation detection device, or a different type of processing device such as an unmanned aerial vehicle control device, a server, a physical host, or a user equipment (UE) that integrates the multimodal vehicle violation detection device. The multimodal vehicle violation detection device can be implemented in hardware or software, and the UE can be a terminal device such as a smartphone, tablet computer, laptop computer, desktop computer, or personal digital assistant (PDA). The processing device can be set up in a device cluster.
[0035] Specifically, the processing device that executes the multimodal vehicle violation detection method provided in this application, considering that the solution of this application involves collecting highway images through a drone, the processing device itself can be combined with the control part of the drone. For example, the processing device can directly be a drone controller, console, or other different forms of drone control devices.
[0036] Of course, the present application solution can also use an existing drone control system to complete the collection of highway images. That is, the present application does not require adaptive settings of the existing drone control system in terms of relevant software and hardware. It is only necessary to obtain the highway images required by the present application through the existing drone control system. In this case, the processing device that executes the multimodal vehicle violation detection method provided by the present application can be any device with data processing capabilities, such as a server device, a physical host or a UE.
[0037] It can be seen that the processing equipment that executes the multimodal vehicle violation detection method provided by this application is relatively flexible in terms of equipment type and equipment deployment form in actual situations, and this application does not make any specific restrictions on it.
[0038] Next, we will introduce the multimodal vehicle violation detection method provided by this application.
[0039] First, see Figure 1 , Figure 1 A flow chart of the multimodal vehicle violation detection method of the present application is shown. The multimodal vehicle violation detection method provided by the present application may specifically include the following steps S101 to S103:
[0040] Step S101, obtaining a target highway image where a vehicle violation is to be detected, wherein the target highway image is specifically acquired by a drone;
[0041] It can be understood that the vehicle violation detection and processing performed by this application is considered from the image level, and the vehicle violation detection target is achieved through image detection and processing. In the case that this application is specifically aimed at the vehicle violation detection target on the highway, the first thing to be obtained in the specific detection processing of this application is the highway image. For the convenience of explanation, the currently obtained highway image of the vehicle violation to be detected is recorded as the target highway image.
[0042] Among them, the target highway images processed by this application are collected by drones when performing their arranged aerial photography tasks. The drones can fly freely under control and have a wider field of view and coverage, which helps to achieve more comprehensive violation detection. The drones themselves have flexible maneuverability and can adjust the flight path and monitoring area according to actual conditions to better adapt to different environments. In addition, drones can also collect images in a timely manner and achieve faster violation detection through real-time transmission and analysis technology.
[0043] In this regard, it is obvious that reference can be made to the existing technology for collecting highway images based on drones. This involves the basic operations of drones, and considering that it is not the focus of the solution of this application, this application will not provide a detailed explanation here.
[0044] In addition, it can be understood that the acquisition and processing of the target highway image here can be either real-time acquisition and processing based on drones or extraction and processing of ready-made images in specific applications. These two types of image acquisition methods can be flexibly adjusted according to actual needs.
[0045] Step S102: Input the target highway image into a pre-configured multimodal detection model for detection processing. The multimodal detection model includes a first detection model, a second detection model, and a third detection model. The first detection model is used to detect vehicles in the input image and obtain vehicle detection frames. The second detection model is used to detect contact points of wheels on the ground in a vehicle screenshot obtained from the input image using the vehicle detection frames and obtain contact key point information. The third detection model is used to detect highways in the input image and obtain highway information.
[0046] In order to achieve the goal of detecting vehicle violations based on image detection, this application constructs a model fusion architecture, which consists of a first, second and third detection models to form a multimodal detection model.
[0047] The first detection model and the second detection model are arranged in a supporting manner. The first detection model detects the vehicle in the input image through the target detection algorithm and can output the corresponding vehicle detection frame. At this time, the vehicle detection frame can be used as a basis to obtain the corresponding vehicle screenshot from the original input image of the first detection model (that is, the vehicle image is cropped from the original image using the vehicle detection frame). The second detection model then detects the wheels and the ground in the vehicle screenshot and detects the contact points between the two to obtain the contact key point information required by the multimodal detection model.
[0048] The third detection model detects highways in the input image of the multimodal detection model. Specifically, it can detect elements involved in highways, such as the road surface, solid lines, dotted lines and other elements of the highway, and obtain another output information required by the multimodal detection model, which is the highway information.
[0049] It is understandable that before being put into specific applications, the detection model also needs to involve model training processing, which generally includes:
[0050] After configuring the sample images marked with the corresponding detection results, the sample images are input into the model in sequence, and the model performs the corresponding detection processing to realize forward / forward propagation. Then, the corresponding loss function is calculated for the detection results output by the model, and the model parameters / related weights are optimized with the loss function calculation results to realize backpropagation. In this way, when the training requirements such as training time, training times, and detection accuracy are met, the model training can be completed.
[0051] Step S103 : judging whether the corresponding vehicle has committed a traffic violation based on the target contact key point information and the target highway information output by the multimodal detection model.
[0052] It can be understood that after obtaining the target contact key point information and the target highway information obtained by the multimodal detection model for processing the target highway image, it is determined whether there is a vehicle violation.
[0053] Among them, it can be understood that the target contact key point information provides the specific position information of the vehicle's wheels, which can be combined with the information of the highway itself (target highway information) identified from the image to carry out more accurate position judgment and determine whether the corresponding vehicle violation behavior is matched.
[0054] For the existing technology, the vehicle violation behavior is judged based on whether the vehicle and the emergency lane are included or overlapped in the image. Typically, for example, it can detect whether there is overlap between the vehicle's detection frame and the emergency lane's detection frame, or whether there is overlap between the wheel image and the emergency lane image, and so on.
[0055] However, the unstable drone perspective, highway road conditions or vehicle conditions will all interfere to a certain extent, leading to misjudgment problems. For example, if due to the tilted perspective, a vehicle approaching the emergency lane does not actually exceed the emergency lane, but the image of the vehicle and the emergency lane overlap, it will be mistakenly identified as a vehicle violation.
[0056] Similarly, even if it only focuses on whether the image of the wheel contains / overlaps with the image of the emergency lane, this problem still exists.
[0057] In contrast, the present application does not simply focus on the image of the wheel, but continues to detect and process the image of the wheel (the vehicle screenshot obtained through the vehicle detection frame), and analyzes the key contact points between the wheel and the highway ground, so as to obtain more accurate location information to provide more accurate data reference for vehicle violations, thereby performing multimodal and delicate vehicle violation detection and processing, and thus achieving high-precision vehicle violation detection effects.
[0058] from Figure 1It can be seen from the illustrated embodiment that, for the target of vehicle violation detection based on drones, the present application configures a multimodal detection model, which includes a first detection model, a second detection model and a third detection model. The first detection model is used to detect vehicles in the input image and obtain a vehicle detection frame. The second detection model is used to detect the contact points of the wheels on the ground in the vehicle screenshot obtained from the input image through the vehicle detection frame and obtain contact key point information. The third detection model is used to detect highways in the input image and obtain highway information. Based on this model fusion architecture, starting from the target highway image collected by the drone, the target contact key point information of the vehicle wheels on the ground can be determined based on the detected vehicle detection frame. Then, combined with the detected target highway information, it is determined whether there is a vehicle violation, and multimodal and delicate vehicle violation detection processing is performed, thereby achieving high-precision vehicle violation detection effect.
[0059] Continue to the above Figure 1 Each step of the illustrated embodiment and its possible implementation in practical applications are described in detail.
[0060] It is easy to understand when collecting images. The images taken should face the highway road surface that needs to be inspected as directly as possible, or in other words, the drone camera's viewing angle should face the highway road surface that needs to be inspected as directly as possible.
[0061] As a practical implementation method, Figure 2 The diagram shows a scenario diagram of a drone capturing images in the present application. For the implementation effect of the solution of the present application, when the drone captures images (including the target highway images here and the sample images involved later), the height above the ground is between 60 meters and 80 meters (i.e., the distance range of [60,80]), and the camera pitch angle is between -90° and -40° (i.e., the angle range of [-90,-40]).
[0062] It can be understood that the configuration conditions for the drone camera to collect images set here can bring better processing effects for the detection and training of the model involved in this application. Therefore, in practical applications, it is preferred to control the collection of relevant images under the above-mentioned image acquisition conditions.
[0063] In addition, as mentioned above, the present application may also involve the training processing of a multimodal detection model. Specifically, it may also involve the training processing of three detection models in the fusion model architecture of the multimodal detection model.
[0064] Furthermore, as a practical implementation, the multimodal vehicle violation detection of the present application may also include:
[0065] Obtain sample images collected by the drone;
[0066] Label the sample image with the highway environment image I and the vehicle detection frame target i , vehicle detection frame target i Corresponding vehicle screenshot I i 、Vehicle Screenshot I i Corresponding contact key point information key ij And highway semantic segmentation mask M, where target i = [p1, p2], i represents the vehicle index, p1 is the upper left point of the detection box, p2 is the lower right point of the detection box, I i A screenshot of the vehicle with vehicle index i, with key contact information key ij where j∈[1,4] represents the contact points between the left front wheel, right front wheel, right rear wheel and left rear wheel of the vehicle and the ground. In the highway semantic segmentation mask M, the background class value is 0, the highway pavement class value is 1, the solid line class value is 2, and the dashed line class value is 3.
[0067] Take the highway environment image I and the vehicle detection frame target i Based on the training of the first detection model, and the vehicle screenshot I i And contact key point information key ij The second detection model is trained based on the highway environment image I and the highway semantic segmentation mask M.
[0068] It can be seen that the setting here introduces how to train the first, second and third detection models in this application, and it can also specifically involve the application of target detection and semantic analysis.
[0069] The annotation processing involved can be achieved through relevant automatic annotation tools. Of course, in specific applications, manual annotation can also be used to achieve it, and it can be set according to actual needs.
[0070] For easier understanding, you can also refer to Figure 3 A schematic diagram of a scene before and after the annotation of the sample image of this application is shown to provide a more vivid understanding of the annotation processing involved here.
[0071] exist Figure 3In the image, the entire road surface identified by the highway can be marked as 1, and the area outside the entire road surface is marked as 0. The dotted line on the road surface (corresponding to the lane line of the ordinary lane on the highway) is marked as 1, and the solid line on the road surface (corresponding to the lane line of the ordinary lane on the left side of the highway and the lane line of the emergency lane on the right side) is marked as 2. Different semantic information can also be marked with different colors. These semantic annotation information can be combined with the key contact points between the wheels and the ground to determine whether they match the corresponding vehicle violation behavior.
[0072] In addition, there may be other roads next to the highway, which may also cause additional solid lines, such as Figure 3 If the ramp shown is marked, it should be ignored in the subsequent detection algorithm to avoid interference.
[0073] Furthermore, corresponding to the training process of the first detection model, as a practical implementation method, the highway environment image I and the vehicle detection frame target i The process of training the first detection model as a basis may specifically include the following:
[0074] 1. For the highway environment image I and the vehicle detection frame target i The training data composed of [I,target i ], perform preprocessing including resizing and normalization;
[0075] It can be understood that preprocessing is mainly to obtain training data with higher quality and easier training. In addition to resizing and normalization, other preprocessing steps may also be involved.
[0076] 2. Through the backbone network, the training data [I, target i ] (the highway environment image I after preprocessing) is used for feature extraction, and then cross-scale feature fusion is performed through the Neck network to obtain the fused feature map;
[0077] It can be understood that cross-scale and cross-layer feature fusion can increase the representation capability of the network so that it can process targets of different sizes at the same time, thereby accelerating model training and improving model recognition accuracy.
[0078] 3. Input the feature map into the modeling module Detection Head to predict the detection box and realize forward propagation;
[0079] It can be understood that this application adopts a target detection algorithm to identify road vehicles, constructs a target detection network with an end-to-end one-stage idea, divides the image into multiple grids through the gridding idea, and each grid is responsible for predicting one or more objects. It has good robustness to overlapping targets and is therefore suitable for highway environments where excessive traffic may cause road congestion.
[0080] In this case, the backbone network, Neck network and modeling module Detection Head are all specific configuration contents of the target detection algorithm (model).
[0081] Among them, the modeling module Detection Head includes a series of convolutional layers and pooling layers, as well as convolution kernels for predicting the position and category of the bounding box. In each Detection Head, the model predicts a series of bounding boxes based on the preset anchor points (Anchor). Each bounding box consists of the bounding box coordinates, confidence (indicating whether there is an object in the detection box) and the probability of different categories.
[0082] 4. For the prediction results, first calculate the coordinate loss Localization Loss, confidence loss Confidence Loss and class loss Class Loss (combined with the vehicle detection frame target after preprocessing i Calculated), and then weighted summed by the following formula to get the final loss Loss:
[0083] Loss=0.05*Localization Loss+1*Confidence Loss+0.5*Class Loss;
[0084] To expand on this, the coordinate loss Localization Loss is used to measure the difference between the position and size of the detection frame and the real frame. Its mathematical definition is:
[0085] Loc Loss = IoU-(cc hat ) 2 / c+(rr hat ) 2 / r,
[0086] Among them, IoU is the intersection over union of the predicted detection box and the real detection box; c and r are the position information of the center point of the predicted detection box relative to the width and height of the image, c hat and r hat It is the position information of the center point of the real detection box relative to the width and height of the image.
[0087] Confidence Loss is used to measure the difference between the confidence of the predicted detection box and the actual situation. The model calculates the confidence loss by using binary cross entropy loss (BCELoss), but only applies to the bounding box containing the target. Its mathematical definition is:
[0088]
[0089] Among them, y represents the probability of whether the predicted detection box actually contains the target, and p represents the probability that the predicted detection box is predicted to contain the target.
[0090] Class loss is used to measure the difference between the target category probability predicted by the prediction detection box and the true category. The model also uses binary cross entropy loss to calculate the category loss.
[0091] 5. Optimize the model parameters (related weights) through the final loss Loss to achieve backpropagation, and complete the model training after multiple rounds of iterative training meet the convergence requirements.
[0092] The trained model can be put into practical use as the first detection model. The model can output information such as the target category detected by the input image, the coordinates of the detection box, and the confidence level. This information can be used to draw the detection box and identify the detected target.
[0093] Furthermore, corresponding to the training process of the second detection model, as a practical implementation method, the vehicle screenshot I i And contact key point information key ij The process of training the second detection model as a basis may specifically include the following:
[0094] 1. Take a screenshot of the vehicle I i And contact key point information key ij The training data composed of i ,key ij ], perform preprocessing including resizing and normalization;
[0095] It can be understood that preprocessing is mainly to obtain training data with higher quality and easier training. In addition to resizing and normalization, other preprocessing steps may also be involved.
[0096] And for the target from the vehicle detection box i To get a screenshot of the vehicle I i , then take a screenshot of the vehicle I i During the preprocessing process such as resizing, you can also refer to Figure 4A schematic diagram of a scenario for preprocessing vehicle screenshots of the present application is shown for a more vivid understanding.
[0097] 2. Through the HRNet feature extraction module, the training data [I i ,key ij ] in the training image (vehicle screenshot after preprocessing I i ) to extract features, and then continue to perform key point regression prediction through the key point prediction module of the top layer head of the model to obtain the required 4 key point positions (key ij ), realize forward propagation, the corresponding mathematical expression is as follows:
[0098] FeatureMap=F1(x),
[0099] PointMap=F2(FeatureMap),
[0100] PointMask=Sigmoid(PointMap),
[0101] Among them, F1 is the feature extraction module, and F2 is the key point prediction module. F2 further extracts features from the feature map FeatureMap output by F1 to obtain the key point score matrix PointMap, and uses the Sigmoid activation function to convert the key point score matrix PointMap into the key point prediction mask PointMask. The key point prediction mask PointMask is a probabilistic expression of the key point;
[0102] The mathematical representation here can also be combined with Figure 5 A schematic diagram of a scenario of key contact point detection processing of the present application is shown for a more vivid understanding.
[0103] It can be understood that in the detection process of contact key points, the present application specifically introduces a high-resolution graph analysis network (High-Resolution Network) architecture to design a key point detection model (the second detection model), whose main modules include a feature extraction module and a key point prediction module. By calculating the contact key point coordinates (key points) of the four outermost wheels of the vehicle and the road surface, the key points of the vehicle are detected. ij ) to obtain the actual projection of the vehicle on the road.
[0104] 3. Calculate the loss Loss based on the prediction results (combined with the contact key point information key after preprocessing) ij calculated);
[0105] It can be understood that model training involves the application of loss function, so the model training process here also needs to calculate the corresponding loss according to the preset loss function algorithm.
[0106] The loss function here, in specific operations, can also refer to the loss function configuration content involved in the first detection model mentioned above, or adopt other types of loss functions.
[0107] 4. Optimize model parameters (related weights) through loss, realize back propagation, and complete model training after multiple rounds of iterative training meet the convergence requirements.
[0108] The trained model can be put into practical use as the second detection model, and the model can output the coordinates of the contact key points detected by the input image.
[0109] Furthermore, corresponding to the training process of the third detection model, as a practical implementation method, the process of training the third detection model based on the highway environment image I and the highway semantic segmentation mask M can specifically include the following:
[0110] 1. Perform preprocessing on the training data [I, M] consisting of the highway environment image I and the highway semantic segmentation mask M, including resizing and normalization;
[0111] It can be understood that preprocessing is mainly to obtain training data with higher quality and easier training. In addition to resizing and normalization, other preprocessing steps may also be involved.
[0112] 2. The training image in the training data [I, M] (the pre-processed highway environment image I) is fed into the encoder, and the receptive field is enlarged by stacking the dilated convolutional neural network module.
[0113] It can be understood that the semantic segmentation model (the third detection model) used here mainly consists of two modules: encoder and decoder. The encoder is mainly responsible for feature extraction, while the decoder is responsible for feature fusion and pixel classification.
[0114] The setting of superimposing the dilated convolutional DCNN module here allows the feature layer to obtain a larger field of view without losing information.
[0115] 3. The decoder upsamples and fuses the two feature layers of different depths obtained by the encoder. It adjusts the number of channels of the deep feature layer using 1×1 convolution, then upsamples and stacks it with the shallow feature layer. It then forms the final effective feature layer through depthwise separable convolution, and finally outputs the pixel classification mask through 1×1 convolution.
[0116] For the working logic of encoder and decoder, you can also refer to Figure 6 A schematic diagram of a scene of highway semantic segmentation of the present application is shown for a more vivid understanding.
[0117] 4. Calculate the loss Loss based on the obtained prediction results (calculated in combination with the preprocessed highway semantic segmentation mask M);
[0118] It can be understood that model training involves the application of loss function, so the model training process here also needs to calculate the corresponding loss according to the preset loss function algorithm.
[0119] The loss function here, in specific operations, can also refer to the loss function configuration content involved in the first detection model mentioned above, or adopt other types of loss functions.
[0120] 5. Optimize model parameters (related weights) through loss, realize back propagation, and complete model training after multiple rounds of iterative training meet the convergence requirements.
[0121] The trained model can be put into practical use as the third detection model, and the model can output highway information obtained by semantic segmentation of the input image.
[0122] Regarding the detection and processing of vehicle violation behaviors involved in step S103, this application also provides a specific implementation solution.
[0123] Specifically, as a practical implementation method, refer to Figure 7 The schematic diagram of a scenario for detecting and processing vehicle violations in the present application is shown. Step S103 determines whether the corresponding vehicle has committed a vehicle violation based on the target contact key point information and target highway information output by the multimodal detection model. Specifically, the following steps may be performed:
[0124] 1. Based on the key point coordinates of target vehicle i in the target contact key point information Construct the detection area EdgeArea and detection area LineArea respectively, where the key point coordinates of the target vehicle i Where j∈[1,4] represents the key points of contact between the left front wheel, right front wheel, right rear wheel and left rear wheel of the target vehicle and the ground, and the two road edge detection points on the front and rear sides of the target vehicle are Detection area EdgeArea is composed of points Surrounded by the detection area LineArea is composed of points surround;
[0125] It can be understood that the key point coordinates of the target vehicle i The target contact key point information output by the multimodal detection model provides the contact point coordinates of the four wheels of the target vehicle i on the ground, thereby enabling the two road edge detection points involved in this application to be detected. i =[edge1 i ,edge2 i ]) is constructed in addition to the two detection areas (EdgeArea and LineArea), which provide specific data basis for the subsequent judgment of specific vehicle violations.
[0126] For details, please refer to Figure 8 A schematic diagram of a scene of the detection area of the present application is shown for a more vivid understanding.
[0127] 2. Based on the detection areas EdgeArea and LineArea, and in combination with the highway semantic segmentation mask M representing the target highway information, which has a background value of 0, a highway pavement value of 1, a solid line value of 2, and a dashed line value of 3, determine whether the target vehicle has committed a vehicle violation according to the following vehicle violation judgment rules:
[0128] If the highway semantic segmentation mask M representing the target highway information contains a pixel with the number 2 within the detection area LineArea, it is determined that the vehicle has crossed the line.
[0129] If the ratio of pixels with the number 0 to pixels with the number 1 in the highway semantic segmentation mask M representing the target highway information within the detection area EdgeArea is greater than 0.3, and the highway semantic segmentation mask M representing the target highway information does not contain pixels with the number 2 within the detection area EdgeArea and the detection area LineArea, then it is determined that the vehicle is occupying the emergency lane.
[0130] It can be understood that the determination process involved here can also be understood in conjunction with the following Table 1.
[0131] Table 1 - Rules for determining vehicle violations
[0132]
[0133] In the specific vehicle violation detection and processing, it is easy to understand that it is carried out on a vehicle basis. If several vehicles are detected at the beginning, vehicle violation detection and processing are usually performed several times respectively, and all vehicles in the images are checked one by one or in parallel to see whether there are any vehicle violations.
[0134] In the specific detection process, it can be seen from the above content that based on the road surface, solid lines and dotted lines obtained by semantic analysis of the highway, this application can involve the determination / detection of violations of solid lines and vehicles occupying emergency lanes, and provides different determination rules to achieve the determination goals.
[0135] In addition, it can be understood that the key point coordinates mentioned above The setting of the highway semantic segmentation mask M also corresponds to the corresponding contents of the labeling and training processing involved in the previous training model.
[0136] When conducting vehicle violation detection, it can be understood that it may also involve response processing in different aspects such as behavior recording, telephone warnings, and personnel going to the scene. This can be adaptively adjusted according to the specific response settings in the actual situation. Correspondingly, the present application is based on a multimodal vehicle violation detection method, and may also involve corresponding response processing based on the vehicle violation detection results. However, considering that it is not the focus of the solution of this application, it will not be explained in detail here.
[0137] The above is an introduction to the multimodal vehicle violation detection method provided by this application. In order to facilitate better implementation of the multimodal vehicle violation detection method provided by this application, this application also provides a multimodal vehicle violation detection device from the perspective of functional modules.
[0138] See Figure 9 9 is a structural diagram of a multimodal vehicle violation detection device of the present application. In the present application, the multimodal vehicle violation detection device 900 may specifically include the following structure:
[0139] An acquisition unit 901 is configured to acquire an image of a target highway where a vehicle violation is to be detected, wherein the image of the target highway is specifically acquired by a drone;
[0140] Detection unit 902 is configured to input the target highway image into a pre-configured multimodal detection model for detection processing. The multimodal detection model includes a first detection model, a second detection model, and a third detection model. The first detection model is configured to detect vehicles in the input image and obtain vehicle detection frames. The second detection model is configured to detect contact points of wheels on the ground in a vehicle screenshot obtained from the input image using the vehicle detection frames and obtain contact key point information. The third detection model is configured to detect highways in the input image and obtain highway information.
[0141] The judgment unit is used to judge whether the corresponding vehicle has violated traffic regulations based on the target contact key point information and target highway information output by the multimodal detection model.
[0142] In an exemplary implementation, the apparatus further includes a training unit 904:
[0143] Obtain sample images collected by the drone;
[0144] Label the sample image with the highway environment image I and the vehicle detection frame target i , vehicle detection frame target i Corresponding vehicle screenshot I i 、Vehicle Screenshot I i Corresponding contact key point information key ij And highway semantic segmentation mask M, where target i = [p1, p2], i represents the vehicle index, p1 is the upper left point of the detection box, p2 is the lower right point of the detection box, I i A screenshot of the vehicle with vehicle index i, with key contact information key ij where j∈[1,4] represents the contact points between the left front wheel, right front wheel, right rear wheel and left rear wheel of the vehicle and the ground. In the highway semantic segmentation mask M, the background class value is 0, the highway pavement class value is 1, the solid line class value is 2, and the dashed line class value is 3.
[0145] Take the highway environment image I and the vehicle detection frame target i Based on the training of the first detection model, and the vehicle screenshot I i And contact key point information key ij The second detection model is trained based on the highway environment image I and the highway semantic segmentation mask M.
[0146] In another exemplary implementation, the highway environment image I and the vehicle detection frame target i The process of training the first detection model as a basis includes the following:
[0147] For the highway environment image I and the vehicle detection frame target i The training data composed of [I,target i ], perform preprocessing including resizing and normalization;
[0148] Through the backbone network, the training data [I, target i ] to extract features from the training images, and then perform cross-scale feature fusion through the Neck network to obtain the fused feature map;
[0149] The feature map is input into the modeling module Detection Head to predict the detection box and realize forward propagation;
[0150] For the obtained prediction results, first calculate the coordinate loss Localization Loss, confidence loss Confidence Loss and class loss Class Loss respectively, and then perform weighted summation through the following formula to obtain the final loss Loss:
[0151] Loss=0.05*Localization Loss+1*Confidence Loss+0.5*Class Loss;
[0152] The model parameters are optimized through the final loss Loss to achieve back propagation, and the model training is completed after multiple rounds of iterative training meet the convergence requirements.
[0153] In another exemplary implementation, the vehicle screenshot 1 i And contact key point information key ij The process of training the second detection model as a basis includes the following:
[0154] Screenshot of vehicle I i And contact key point information key ij The training data composed of i ,key ij ], perform preprocessing including resizing and normalization;
[0155] The training data[I i ,key ij ] is used to extract features from the training images and continue to perform key point regression prediction through the key point prediction module of the top layer head of the model to obtain the required 4 key point positions and realize forward propagation. The corresponding mathematical representation is as follows:
[0156] FeatureMap=F1(x),
[0157] PointMap=F2(FeatureMap),
[0158] PointMask=Sigmoid(PointMap),
[0159] Among them, F1 is the feature extraction module, and F2 is the key point prediction module. F2 further extracts features from the feature map FeatureMap output by F1 to obtain the key point score matrix PointMap, and uses the Sigmoid activation function to convert the key point score matrix PointMap into the key point prediction mask PointMask. The key point prediction mask PointMask is a probabilistic expression of the key point;
[0160] Calculate the loss Loss based on the obtained prediction results;
[0161] The model parameters are optimized through loss to achieve back propagation, and the model training is completed after multiple rounds of iterative training meet the convergence requirements.
[0162] In another exemplary implementation, the process of training the third detection model based on the highway environment image I and the highway semantic segmentation mask M includes the following steps:
[0163] The training data [I,M] consisting of the highway environment image I and the highway semantic segmentation mask M is preprocessed including resizing and normalization;
[0164] The training images in the training data [I,M] are fed into the encoder, and the receptive field is enlarged by stacking the dilated convolutional CNN module.
[0165] The decoder upsamples and fuses the two feature layers of different depths obtained by the encoder, adjusts the number of channels of the deep feature layer using 1×1 convolution, then upsamples and stacks it with the shallow feature layer, and forms the final effective feature layer through depth-wise separable convolution. Finally, it outputs the pixel classification mask through 1x1 convolution;
[0166] Calculate the loss Loss based on the obtained prediction results;
[0167] The model parameters are optimized through loss to achieve back propagation, and the model training is completed after multiple rounds of iterative training meet the convergence requirements.
[0168] In another exemplary implementation, the determining unit 903 is specifically configured to:
[0169] Based on the key point coordinates of target vehicle i in the target contact key point information Construct the detection area EdgeArea and detection area LineArea respectively, where the key point coordinates of the target vehicle i Where j∈[1,4] represents the key points of contact between the left front wheel, right front wheel, right rear wheel and left rear wheel of the target vehicle and the ground, and the two road edge detection points on the front and rear sides of the target vehicle are Detection area EdgeArea is composed of points Surrounded by the detection area LineArea is composed of points surround;
[0170] Based on the detection area EdgeArea and the detection area LineArea, combined with the highway semantic segmentation mask M representing the target highway information, which has a background value of 0, a highway pavement value of 1, a solid line value of 2, and a dashed line value of 3, the following vehicle violation judgment rules are used to determine whether the target vehicle has committed a vehicle violation:
[0171] If the highway semantic segmentation mask M representing the target highway information contains a pixel with the number 2 within the detection area LineArea, it is determined that the vehicle has crossed the line.
[0172] If the ratio of pixels with the number 0 to pixels with the number 1 in the highway semantic segmentation mask M representing the target highway information within the detection area EdgeArea is greater than 0.3, and the highway semantic segmentation mask M representing the target highway information does not contain pixels with the number 2 within the detection area EdgeArea and the detection area LineArea, then it is determined that the vehicle is occupying the emergency lane.
[0173] In another exemplary implementation, when the drone collects images, the altitude from the ground is between 60 meters and 80 meters, and the pitch angle of the camera is between -90° and -40°.
[0174] This application also provides a vehicle violation detection and processing device based on multi-modality from the perspective of hardware structure, see Figure 10 , Figure 10 The present invention shows a schematic diagram of a processing device. Specifically, the present invention may include a processor 1001, a memory 1002, and an input / output device 1003. The processor 1001 is configured to execute a computer program stored in the memory 1002. Figure 1 The steps of the vehicle violation detection method based on multimodality in the corresponding embodiment; or, the processor 1001 is used to implement the following when executing the computer program stored in the memory 1002 Figure 9 The memory 1002 is used to store the functions of each unit in the corresponding embodiment. Figure 1 The computer program required for the multimodal vehicle violation detection method in the corresponding embodiment.
[0175] For example, the computer program may be divided into one or more modules / units, one or more of which are stored in the memory 1002 and executed by the processor 1001 to complete the present application. One or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in a computer device.
[0176] The processing device may include, but is not limited to, a processor 1001, a memory 1002, and an input / output device 1003. Those skilled in the art will appreciate that the illustrations are merely examples of processing devices and do not limit the processing device. The processing device may include more or fewer components than shown, or a combination of certain components, or different components. For example, the processing device may also include a network access device, a bus, etc., and the processor 1001, the memory 1002, the input / output device 1003, etc. may be connected via a bus.
[0177] The processor 1001 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. The processor is the control center of the processing device and connects various parts of the entire device using various interfaces and lines.
[0178] Memory 1002 can be used to store computer programs and / or modules. Processor 1001 implements various functions of the computer device by running or executing computer programs and / or modules stored in memory 1002 and accessing data stored in memory 1002. Memory 1002 may primarily include a program storage area and a data storage area. The program storage area may store an operating system, at least one application required for a function, and the data storage area may store data generated based on the use of the processing device. Furthermore, memory may include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0179] When the processor 1001 is used to execute the computer program stored in the memory 1002, it can specifically implement the following functions:
[0180] Obtaining a target highway image of a vehicle violation to be detected, wherein the target highway image is specifically acquired by a drone;
[0181] Inputting the target highway image into a pre-configured multimodal detection model for detection processing, wherein the multimodal detection model includes a first detection model, a second detection model, and a third detection model. The first detection model is used to detect vehicles in the input image and obtain vehicle detection frames. The second detection model is used to detect contact points of wheels on the ground in a vehicle screenshot obtained from the input image using the vehicle detection frames and obtain contact key point information. The third detection model is used to detect highways in the input image and obtain highway information.
[0182] According to the target contact key point information and target highway information output by the multimodal detection model, it is judged whether the corresponding vehicle has committed a vehicle violation.
[0183] Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working process of the multi-modal vehicle violation detection device, processing equipment and corresponding units described above can refer to the following. Figure 1 The description of the multimodal vehicle violation detection method in the corresponding embodiment will not be repeated here.
[0184] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0185] To this end, the present application provides a computer-readable storage medium, which stores a plurality of instructions, which can be loaded by a processor to execute the present application as follows: Figure 1 The steps of the multimodal vehicle violation detection method in the corresponding embodiment can be referred to as follows for specific operations. Figure 1 The description of the multimodal vehicle violation detection method in the corresponding embodiment will not be repeated here.
[0186] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0187] Due to the instructions stored in the computer readable storage medium, the present application can be executed as follows: Figure 1 The steps of the vehicle violation detection method based on multimodality in the corresponding embodiment can thus be implemented as follows: Figure 1 The beneficial effects that can be achieved by the multimodal vehicle violation detection method in the corresponding embodiment are detailed in the previous description and will not be repeated here.
[0188] The above is a detailed introduction to the multimodal vehicle violation detection method, device, processing equipment and computer-readable storage medium provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A vehicle violation detection method based on multimodality, characterized in that: The method comprises: Acquire a target highway image where the vehicle violation to be detected is located, wherein the target highway image is specifically acquired by a drone; Inputting the target highway image into a preconfigured multimodal detection model for detection processing, wherein the multimodal detection model includes a first detection model, a second detection model, and a third detection model, wherein the first detection model is used to detect vehicles in the input image and obtain a vehicle detection frame, the second detection model is used to detect contact points of wheels on the ground in a vehicle screenshot obtained from the input image via the vehicle detection frame and obtain contact key point information, and the third detection model is used to detect highways in the input image and obtain highway information; Determining whether a corresponding vehicle has committed a traffic violation based on the target contact key point information and the target highway information output by the multimodal detection model; The determining, based on the target contact key point information and the target highway information output by the multimodal detection model, whether the corresponding vehicle has committed a traffic violation includes: Based on the key point coordinates of the target vehicle i in the target contact key point information Construct the detection area EdgeArea and the detection area LineArea respectively, where the key point coordinates of the target vehicle i Where j∈[1,4] represents the key points of contact between the left front wheel, right front wheel, right rear wheel and left rear wheel of the target vehicle and the ground, and the two road edge detection points on the front and rear sides of the target vehicle are The detection area EdgeArea is composed of points The detection area LineArea is surrounded by points surround; Based on the detection area EdgeArea and the detection area LineArea, combined with the content of the highway semantic segmentation mask M representing the target highway information, which has a background value of 0, a highway pavement value of 1, a solid line value of 2, and a dashed line value of 3, the following vehicle violation determination rules are used to determine whether the target vehicle has committed the vehicle violation: If the highway semantic segmentation mask M representing the target highway information contains a pixel with the number 2 within the detection area LineArea, it is determined that the vehicle has crossed the line. If the ratio of pixels with the number 0 to pixels with the number 1 in the highway semantic segmentation mask M representing the target highway information within the detection area EdgeArea is greater than 0.3, and the highway semantic segmentation mask M representing the target highway information does not have a pixel with the number 2 within the detection area EdgeArea and the detection area LineArea, it is determined that the vehicle is occupying the emergency lane.
2. The method according to claim 1, characterized in that The method further comprises: Acquire a sample image captured by the drone; Label the sample image with the highway environment image I and the vehicle detection frame target in the image i , the vehicle detection frame target i Corresponding vehicle screenshot I i 、Screenshot of the vehicle I i Corresponding contact key point information key ij The highway semantic segmentation mask M corresponding to the highway environment image I, where target i = [p1, p2], i represents the vehicle index, p1 is the upper left point of the detection box, p2 is the lower right point of the detection box, I i A screenshot of the vehicle with vehicle index i; Take the highway environment image I and the vehicle detection frame target i Based on the training of the first detection model, and the vehicle screenshot I i And the contact key point information key ij The second detection model is trained based on the highway environment image I, and the third detection model is trained based on the highway environment image I and the highway semantic segmentation mask M corresponding to the highway environment image I.
3. The method according to claim 2, characterized in that Take the highway environment image I and the vehicle detection frame target i The process of training the first detection model as a basis includes the following: The highway environment image I and the vehicle detection frame target i The training data composed of [I,target i ], perform preprocessing including resizing and normalization; The backbone network is used to train the training data [I, target i ] to extract features from the training images, and then perform cross-scale feature fusion through the Neck network to obtain the fused feature map; The feature map is input into the modeling module Detection Head to predict the detection box and realize forward propagation; For the obtained prediction results, first calculate the coordinate loss Localization Loss, confidence loss Confidence Loss and class loss Class Loss respectively, and then perform weighted summation through the following formula to obtain the final loss Loss: Loss=0.05*Localization Loss+1*Confidence Loss+0.5*Class Loss; The model parameters are optimized through the final loss Loss to achieve back propagation, and the model training is completed after multiple rounds of iterative training meet the convergence requirements.
4. The method according to claim 2, characterized in that Take the vehicle screenshot I i And the contact key point information key ij The process of training the second detection model as a basis includes the following: Screenshot I of the vehicle i And the contact key point information key ij The training data composed of i ,key ij ], perform preprocessing including resizing and normalization; The training data [I i ,key ij ] is used to extract features from the training images and continue to perform key point regression prediction through the key point prediction module of the top layer head of the model to obtain the required 4 key point positions and realize forward propagation. The corresponding mathematical representation is as follows: FeatureMap=F1(x), PointMap=F2(FeatureMap), PointMask=Sigmoid(PointMap), Among them, F1 is the feature extraction module, F2 is the key point prediction module, F2 further extracts features from the feature map FeatureMap output by F1 to obtain the key point score matrix PointMap, and uses the Sigmoid activation function to convert the key point score matrix PointMap into a key point prediction mask PointMask, which is a probabilistic expression of the key point; Calculate the loss Loss based on the obtained prediction results; The model parameters are optimized through the loss Loss to achieve back propagation, and the model training is completed after multiple rounds of iterative training meet the convergence requirements.
5. The method according to claim 2, characterized in that The process of training the third detection model based on the highway environment image I and the highway semantic segmentation mask M corresponding to the highway environment image I includes the following steps: Performing preprocessing including resizing and normalization on the training data [I, M] consisting of the highway environment image I and the highway semantic segmentation mask M corresponding to the highway environment image I; The training images in the training data [I, M] are fed into the encoder, and the receptive field is enlarged by stacking the dilated convolutional neural network (DCNN) module. The decoder upsamples and fuses the two feature layers of different depths obtained by the encoder, adjusts the number of channels of the deep feature layer using 1×1 convolution, then upsamples and stacks it with the shallow feature layer, and forms the final effective feature layer through depth-wise separable convolution, and finally outputs the pixel classification mask map through 1x1 convolution; Calculate the loss Loss based on the obtained prediction results; The model parameters are optimized through the loss Loss to achieve back propagation, and the model training is completed after multiple rounds of iterative training meet the convergence requirements.
6. The method according to claim 1, characterized in that When the drone collects images, the altitude from the ground is between 60 meters and 80 meters, and the pitch angle of the camera is between -90° and -40°.
7. A vehicle violation detection device based on multimodality, characterized in that: The device comprises: an acquisition unit, configured to acquire an image of a target highway where a vehicle violation is to be detected, wherein the image of the target highway is specifically acquired by a drone; a detection unit configured to input the target highway image into a preconfigured multimodal detection model for detection processing, wherein the multimodal detection model includes a first detection model, a second detection model, and a third detection model, wherein the first detection model is configured to detect vehicles in the input image and obtain a vehicle detection frame, the second detection model is configured to detect contact points of wheels on the ground in a vehicle screenshot obtained from the input image via the vehicle detection frame and obtain contact key point information, and the third detection model is configured to detect highways in the input image and obtain highway information; a judgment unit, configured to judge whether a corresponding vehicle has committed a traffic violation based on the target contact key point information and the target highway information output by the multimodal detection model; The determining, based on the target contact key point information and the target highway information output by the multimodal detection model, whether the corresponding vehicle has committed a traffic violation includes: Based on the key point coordinates of the target vehicle i in the target contact key point information Construct the detection area EdgeArea and the detection area LineArea respectively, where the key point coordinates of the target vehicle i Where j∈[1,4] represents the key points of contact between the left front wheel, right front wheel, right rear wheel and left rear wheel of the target vehicle and the ground, and the two road edge detection points on the front and rear sides of the target vehicle are The detection area EdgeArea is composed of points The detection area LineArea is surrounded by points surround; Based on the detection area EdgeArea and the detection area LineArea, combined with the content of the highway semantic segmentation mask M representing the target highway information, which has a background value of 0, a highway pavement value of 1, a solid line value of 2, and a dashed line value of 3, the following vehicle violation determination rules are used to determine whether the target vehicle has committed the vehicle violation: If the highway semantic segmentation mask M representing the target highway information contains a pixel with the number 2 within the detection area LineArea, it is determined that the vehicle has crossed the line. If the ratio of pixels with the number 0 to pixels with the number 1 in the highway semantic segmentation mask M representing the target highway information within the detection area EdgeArea is greater than 0.3, and the highway semantic segmentation mask M representing the target highway information does not have a pixel with the number 2 within the detection area EdgeArea and the detection area LineArea, it is determined that the vehicle is occupying the emergency lane.
8. A processing device, characterized in that The method comprises a processor and a memory, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, the method according to any one of claims 1 to 6 is executed.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Vehicle detection method, terminal and computer readable storage medium
CN114219856A
Method and device for detecting violation behaviors and terminal equipment
CN115661764A