An intersection safety evaluation method based on a drone and a large model
By combining video collected by drones with improved target detection and tracking models and large-scale models, a method for evaluating intersection safety was constructed. This method solves the problem that existing technologies cannot fully reflect the potential risks of intersections, and achieves a more accurate and adaptable traffic safety evaluation.
Patent Information
- Application Number
- CN202511254060.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-09-04
AI Technical Summary
Existing technologies for traffic safety assessment at urban intersections lack the technical means to integrate multimodal information and possess strong reasoning and learning capabilities. This makes it difficult to comprehensively and in real time reflect the potential risks and overall safety level of intersections, especially in complex traffic scenarios where there is a lack of ability to deeply understand complex interactive semantics and dynamic evolution patterns.
By using drones to collect traffic operation videos and combining them with an improved target detection and tracking model and a large model, a method for evaluating intersection safety is constructed. Through potential conflict indicators and a comprehensive safety level evaluation model, the complex nonlinear dynamic interaction behaviors among pedestrians, non-motorized vehicles, and motorized vehicles are quantified. The large model is used to intelligently adjust weight priorities to achieve a comprehensive quantitative assessment of multi-dimensional risks.
It significantly improves the accuracy and adaptability of intersection traffic safety assessment, and can more comprehensively reflect potential risks and overall safety levels. It makes up for the problems of poor real-time performance and large sample bias in traditional methods, and improves the accuracy of target recognition and trajectory tracking in complex environments.
Smart Images

Figure CN120766167B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent traffic detection, in particular to a crossing safety evaluation method based on a UAV and a large model. BACKGROUND
[0002] With the continuous acceleration of urbanization, the urban traffic network is becoming increasingly complex. Road intersections, as the hub nodes in the urban traffic network, have the characteristics of large traffic flow and complex driving routes, and are the key areas of traffic accidents, among which the mixed traffic conflicts involving pedestrians, non-motor vehicles and motor vehicles are particularly prominent.
[0003] Traditional urban traffic safety evaluation methods mainly rely on fixed monitoring equipment or historical traffic accident data for statistical analysis. There are problems such as limited observation angle, data update lag and insufficient spatial coverage, which makes it difficult to comprehensively and real-time reflect the dynamic traffic running state of the intersection. In addition, fixed cameras are often limited by installation location and angle, and have insufficient capture ability for traffic behavior in the edge area and vertical direction of the intersection, affecting the accuracy of traffic safety evaluation.
[0004] In recent years, UAVs have gradually been introduced into urban traffic management and road condition monitoring scenarios due to their advantages of being mobile, having a wide field of view, and being easy to deploy. In the field of traffic safety, UAV video data provides a new way to obtain macro traffic flow information and fine-grained traffic behavior characteristics. Compared with ground monitoring systems, the height advantage of UAVs enables them to have a bird's eye view, covering the entire intersection area and capturing the spatial distribution and motion trajectories of multiple types of traffic participants, especially suitable for traffic behavior perception and conflict analysis in complex road conditions. However, the rich and heterogeneous data such as images and trajectories obtained by UAVs contain more complex behavior interaction semantics and potential risk features than traditional observation data, which puts higher requirements on the feature extraction ability and intelligent level of subsequent data analysis models.
[0005] Currently, some research has attempted to use image data collected by UAVs to carry out traffic safety evaluation research, but most of them focus on single-dimensional indicators (such as flow, speed, etc.), and fail to fully utilize the global spatio-temporal information under the high-altitude perspective of UAVs. More importantly, in the urban intersection scenario, there are non-linear complex interaction behaviors of multiple types of traffic participants (pedestrians, motor vehicles, non-motor vehicles), and the risk sources have significant dynamics and uncertainty. Existing methods lack the ability to deeply understand these complex interaction semantics, dynamic evolution rules, and comprehensive quantitative evaluation of multi-dimensional risks, especially lacking technical means that can fuse multi-modal information, have strong reasoning and learning ability for end-to-end intelligent analysis and decision-making.
[0006] Therefore, in the prior art, when the safety level of the intersection is evaluated, the conflict index is relatively simple, and in the scene of the safety evaluation of the urban intersection, the potential risk and the overall safety level of the intersection need to be comprehensively and comprehensively reflected.
[0007] Therefore, how to more comprehensively and comprehensively evaluate the traffic safety level of the intersection is a problem that those skilled in the art need to solve. SUMMARY
[0008] The purpose of the present application is to provide an intersection safety evaluation method based on a drone and a large model, aiming to solve or improve at least one of the above technical problems.
[0009] To achieve the above purpose, the present application provides the following scheme:
[0010] An intersection safety evaluation method based on a drone and a large model, comprising:
[0011] Using a drone to shoot a traffic operation video of a to-be-detected intersection;
[0012] Inputting the traffic operation video into a pre-constructed target detection and tracking model to obtain intersection information and traffic flow data, and calculating motion information of traffic participants;
[0013] According to the traffic flow data and the motion information, the traffic conflict data is counted, and the potential conflict index of the intersection is constructed; the potential conflict index includes the interactive conflict potential index ICP, the motion disorder index MDI, the driving abnormal behavior index ABI, the congestion pressure index CPI and the violation interference index VII;
[0014] According to the potential conflict index, a comprehensive safety level evaluation model framework is constructed, a large model is used to combine the intersection information and the traffic flow data to improve the comprehensive safety level evaluation model framework, and a comprehensive safety level evaluation model is obtained;
[0015] The potential conflict index is input into the comprehensive safety level evaluation model for calculation, and the comprehensive safety evaluation level of the intersection is output.
[0016] Further, using a drone to shoot a traffic operation video of a to-be-detected intersection specifically includes the following steps:
[0017] Using the drone to take off at the to-be-detected intersection and fly to a set height;
[0018] Fixing the drone at the intersection center, recording the longitude and latitude coordinates, adjusting the camera angle, setting the video shooting parameters, and shooting the video of the intersection.
[0019] Further, the pre-constructed target detection and tracking model specifically includes:
[0020] An RT-DETR-based target detection and tracking model is constructed, and the backbone part and the AIFI part are improved.
[0021] Further, the improvement of the backbone part specifically includes:
[0022] The Block module in the RT-DETR backbone network is replaced with an M-Block module, and the introduced M-Block module follows the Metaformer structure, and the specific expression is:
[0023] ;
[0024] In the formula, Y is the output feature map; is the mapping function of the M-Block module; is a multi-scale spatial interaction module the gain of the intermediate variable; is the output of the feature enhancement unit ; is the input feature map, B is the batch size of the feature map X, C is the channel number of the feature map X, H and W are the height and width of the feature map X respectively; is a normalization operation; is a scaling factor used to adjust the corresponding intensity of the residual; is a spatial interaction module; is a feature enhancement unit.
[0025] Further, the improvement of the AIFI part specifically includes:
[0026] The FNN module of the RT-DETR backbone network is replaced with an SP-FFN module, and the specific expression of the introduced SP-FFN module is:
[0027] The input feature map is expanded in the channel through a 1×1 convolution kernel to obtain an intermediate feature and generate two sub-channels:
[0028] ,
[0029] In the formula, is the feature map after channel expansion; is a 1×1 convolution operation on the input feature map ; is the number of expanded channels; is the original channel number; r is the channel expansion ratio factor
[0030] ;
[0031] In the formula, feature map split sub-feature map split function split dimension
[0032] sub-feature map feature respectively, and nonlinear fusion is performed, and the specific expression is:
[0033]
[0034] In the formula, depth separable convolution operation nonlinear fusion feature map; ⊙ represents element-wise multiplication; GELU is a Gaussian error linear unit activation function and sub-feature map feature sub-feature map after depth separable convolution
[0035] to obtain the output feature map Y, and the expression is:
[0036]
[0037] In the formula, output feature map; Q is a frequency domain quantization matrix two-dimensional real inverse fast Fourier transform real two-dimensional Fourier transform; Fold is a concatenation operation; ⊙ is element-wise multiplication 1x1 convolution dimension reduction to restore the original channel number PXP block expansion is performed.
[0038] Further, the traffic flow data includes a traffic participant id, a traffic participant category, traffic participant bounding box coordinates, traffic participant bounding box height and width, and a frame in which the traffic participant bounding box is located; the traffic participant category includes a pedestrian, a non-motor vehicle, and a motor vehicle; the motion information includes a speed, an acceleration, and a motion direction.
[0039] Further, according to the traffic flow data and the motion information, traffic conflict data is counted, and a potential conflict index of an intersection is constructed, specifically including the following steps:
[0040] The TTC value of the event extracted in the traffic operation video is counted <a preset time threshold, a potential traffic conflict object pair set is established, and then divided into four categories according to the traffic conflict participant type: motor vehicle-motor vehicle, pedestrian-non-motor vehicle, non-motor vehicle-motor vehicle, and pedestrian-motor vehicle.
[0041] The specific expression of the interactive conflict potential index ICP is:
[0042] ;
[0043] where i, j are two traffic participants; is the time interval of traffic participants arriving at the potential conflict point; ε is a constant; C is the set of potential traffic conflict objects predicted;
[0044] The specific expression of the motion disorder index MDI is:
[0045] ;
[0046] where, is the probability distribution of the traffic participants in each direction;
[0047] The specific expression of the driving abnormal behavior index ABI is:
[0048] ;
[0049] where N is the total number of vehicles in the statistical period, I is an indicator function, is a preset acceleration threshold; is the acceleration of the ith vehicle;
[0050] The specific expression of the congestion pressure index CPI is:
[0051] ;
[0052] ;
[0053] where, is the number of traffic participants at time t; A is the physical area occupied by the traffic participants in a fixed time window, is the union area of all traffic participant target detection boxes in the ith frame, and T is the observation duration;
[0054] The specific expression of the violation interference index VII is:
[0055] ;
[0056] where, is the number of times of target intrusion into the no-entry area per unit time, and T is the observation duration.
[0057] Further, the construction of the comprehensive safety level evaluation model framework according to the potential conflict index includes the following steps:
[0058] The interaction conflict potential index ICP, the motion disorder index MDI, the driving abnormal behavior index ABI, the congestion pressure index CPI, and the violation interference index VII are combined into a conflict index vector, and the specific expression is:
[0059] ;
[0060] In the formula, X is a conflict index vector; ICP is an interactive conflict potential index; MDI is a motion disorder index; ABI is an abnormal behavior index; CPI is a congestion pressure index; VII is a violation interference index;
[0061] The default weight is initialized at the same time, and the specific expression is as follows:
[0062] ;
[0063] In the formula, is an initial weight vector; is a weight value of the interactive conflict potential index ICP; is a weight value of the motion disorder index MDI; is a weight value of the abnormal behavior index ABI; is a weight value of the congestion pressure index CPI; is a weight value of the violation interference index VII;
[0064] The comprehensive safety level evaluation model framework is generated.
[0065] Further, the comprehensive safety level evaluation model framework is improved by using a large model in combination with intersection information and traffic flow data, and the comprehensive safety level evaluation model specifically includes the following steps:
[0066] The traffic flow data and intersection information are input to the large model; the intersection information includes intersection geometric properties, intersection function properties, intersection flow data, and intersection conflict data;
[0067] The weight settings in the comprehensive safety level evaluation model framework are given by using the large model reasoning, and the comprehensive safety level evaluation model is generated.
[0068] According to the specific embodiments provided by the present application, the following technical effects are disclosed:
[0069] 1. The present application breaks away from the dependence of traditional methods on single indicators (such as accident rate) or simple conflict indicators (such as TTC and PET), and quantifies the complex nonlinear dynamic interaction behavior between pedestrians, non-motor vehicles and motor vehicles through a multi-dimensional intersection safety evaluation index system including potential conflict, motion disorder, abnormal behavior, congestion pressure and violation interference, so as to accurately capture potential conflict patterns, and thus more comprehensively and comprehensively reflect the potential risks and overall safety level of the intersection.
[0070] 2. The improved target detection tracking model significantly improves the identification and trajectory tracking accuracy of small-sized targets such as pedestrians and non-motor vehicles in complex environments such as dusk and heavy fog, providing a reliable data foundation for subsequent traffic conflict data statistics and index calculation, effectively making up for the poor real-time performance and large sample bias of traditional methods relying on historical accident data statistics.
[0071] 3. The introduction of a large model (LLM) intelligently adjusts the weight priority of each potential conflict indicator in the final evaluation according to the actual traffic state and user demand of the intersection, captures the dynamic correlation and risk evolution law between each potential conflict indicator in complex traffic scenarios, realizes multi-dimensional risk comprehensive quantitative evaluation, and significantly improves the accuracy, adaptability and interpretability of the evaluation results compared with the traditional artificial preset weight method. BRIEF DESCRIPTION OF DRAWINGS
[0072] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0073] Figure 1 The flowchart of the method of the present application.
[0074] Figure 2 The improved M-Block structure diagram in the present application.
[0075] Figure 3 The improved SP-FFN structure diagram in the present application.
[0076] Figure 4 The effect comparison diagram before and after the model improvement of the method of the present application.
[0077] Figure 5 The unmanned aerial vehicle aerial photograph of the investigation site in the present application.
[0078] Figure 6 The intersection traffic flow data diagram in the present application. DETAILED DESCRIPTION
[0079] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0080] The application aims to provide an intersection safety evaluation method based on a UAV and a large model, aiming to solve or improve at least one of the above technical problems.
[0081] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below in combination with the drawings and specific embodiments.
[0082] Noun explanation:
[0083] RT-DETR: a real-time target detection model based on the Transformer architecture.
[0084] Multi-Scale Spatial Interaction Module (MSIM): a technology designed to enhance the ability of deep learning models to process multi-scale information.
[0085] SimpleGate: a lightweight nonlinear gating mechanism that can adjust the spatial feature saliency and facilitate the model to extract meaningful spatial information.
[0086] As shown in Figure 1 The application provides an intersection safety evaluation method based on a UAV and a large model, which comprises:
[0087] Step 1: Construct a target detection and tracking model based on the UAV perspective.
[0088] The RT-DETR backbone network adopts a four-layer (BasicBolck) structure. In order to enhance the feature extraction ability of different sizes of targets under the UAV perspective, the backbone network and the AIFI module are improved in the application.
[0089] In the backbone network part, an enhanced structure block M-Block is used, which integrates the frequency domain enhancement mechanism and the multi-branch spatial perception mechanism. The improved M-Block module follows the Metafromer structure and has two components, the Multi-Scale Spatial Interaction Module (MSIM) and the Feature Enhancement Unit (FEU). The overall process is represented as a function transformation process:
[0090] ,
[0091] In the formula, Y is the output feature map; is the mapping function of the M-Block module; For a multi-scale spatial interaction module Gain after intermediate variable; For a feature enhancement unit Output; For an input feature map, B is the batch size of the feature map X, C is the number of channels of the feature map X, H and W are the height and width of the feature map X respectively; For a normalization operation; For a scaling factor, used to adjust the corresponding intensity of the residual; For a spatial interaction module; For a feature enhancement unit.
[0092] The multi-scale spatial interaction module (MSIM) extracts significant region features from the spatial dimension, and the specific steps are as follows:
[0093] The feature map X is subjected to two-dimensional normalization processing, and the specific expression is as follows:
[0094] ,
[0095] In the formula, is the feature map after two-dimensional normalization; μ and respectively represent the mean and variance of the channel dimension; and ε is a numerical stability constant.
[0096] The two-dimensional normalization processing of the spatial domain feature map has the effect of alleviating the gradient explosion and gradient disappearance phenomenon in the deep network, and improves the training stability.
[0097] The feature map is input into the multi-scale spatial interaction module MSIM, and significant region features are extracted from the spatial dimension, including the following steps:
[0098] The feature map is subjected to depth separable convolution, and the specific expression is as follows:
[0099] And the feature map is subjected to channel expansion:
[0100] ,
[0101] In the formula, is the feature map after depth separable convolution; is a depth separable convolution operation, using a 3×3 convolution kernel;
[0102] The above depth separable convolution is used to reduce the number of parameters.
[0103] The feature map is subjected to channel expansion:
[0104] ,
[0105] wherein, is the feature map after channel expansion; is a 1x1 convolution operation; r is the ratio factor of channel expansion.
[0106] The feature map Z is subjected to multi-branch convolution, and the specific expression is:
[0107] ,
[0108] ,
[0109] wherein, is the feature map output by the i-th branch convolution; is a depthwise separable convolution operation using a specific expansion rate r; Z is the feature map after fusion of the multi-branch convolution; and M is the number of branches of the multi-branch convolution. Different spatial features of different scales are perceived using the multi-branch convolution, and different expansion rates r are used in each branch to achieve modeling of different receptive field ranges.
[0110]
[0111] The feature map Z is subjected to a gated non-linear activation (SimpleGate), and the specific steps are:
[0112] ,
[0113] ,
[0114] wherein, is the sub-feature map after segmentation; represents uniform segmentation of the feature map Z into two parts; and represents element-wise multiplication. is the feature map after the gated non-linear activation.
[0115] A simplified channel attention module (SCA) is constructed using global average pooling and a 1x1 convolution to process the feature map Z_gate, and the specific formula is:
[0116] ,
[0117] ,
[0118] wherein, is the feature map Z_gate subjected to global average pooling; is the feature map after global average pooling of the feature map Z_gate; is a global average pooling operation on the feature map Z_gate; is the feature map after the global average pooling operation on the feature map Z_gate. 1x1 convolution transform is performed; The feature map after passing through the channel attention module.
[0119] The feature map 1x1 compression convolution is performed to obtain the same number of channels as the feature map X:
[0120] ,
[0121] In the formula, The feature map The feature map after convolution compression; 1x1 compression convolution operation is performed on
[0122] The spatial enhancement final output feature map is generated using residual connection combined with a learnable scaling factor β:
[0123] ,
[0124] In the formula, The final output feature map; β is a learnable scaling factor.
[0125] In view of the deficiency of the spatial domain in efficiently expressing image edge structure and global features, a feature enhancement unit (FEU) is used as an additional attention mechanism to improve the robustness of the model to image quality changes.
[0126] The feature enhancement unit (FEU) is used to improve the robustness of the target detection and tracking model to image quality changes, and the specific steps are as follows:
[0127] The feature map is normalized and preprocessed as the original input feature map:
[0128] ,
[0129] In the formula, The feature map after normalization.
[0130] The feature map is converted from the spatial domain to the frequency domain using two-dimensional discrete Fourier transform to enhance the perception ability of the image in low light scenes such as night, and the specific expression is:
[0131] ,
[0132] ,
[0133] In the formula, The feature map is the spectral representation in the frequency domain; and FFT2 represent two-dimensional discrete Fourier transform; and are the magnitude spectrum and the phase spectrum, respectively, defined as:
[0134] ,
[0135] where, is the frequency domain value at frequency ; and is the pixel value of the spatial domain image at position ; is the imaginary unit.
[0136] Attention is enhanced in the frequency magnitude and the spectrum is reconstructed, keeping the phase unchanged and only linearly transforming the magnitude of the spectrum, with the specific expression:
[0137] ,
[0138] where, is the enhanced spectrum representation; is the multi-layer perception; is the magnitude spectrum; is the phase spectrum; is to keep the original phase unchanged and only enhance the magnitude.
[0139] The spectrum representation is transformed back to the spatial domain by a two-dimensional inverse Fourier transform (IFFT):
[0140] ,
[0141] where, is the frequency domain feature transformed feature map; and IFFT2 is the two-dimensional inverse Fourier transform.
[0142] i.e.
[0143] ,
[0144] where, is the pixel value of the spatial domain ; is the value in the frequency domain ; H and W are the height and width of the image, respectively.
[0145] The frequency domain information is fused with the original residual feature, with the specific expression:
[0146] ,
[0147] where, is the final output feature map; and is the original input feature map; is the frequency domain transformed feature map; and γ is a learnable scaling factor.
[0148] The above steps can maintain the computing efficiency while improving the feature expression capability of the target detection and tracking model.
[0149] In the AIFI part, the robustness of the target detection and tracking model in complex environments is improved, and the feedforward neural network module in the AIFI module of the RT-DETR hybrid encoder is improved to SP-FFN, including:
[0150] The input feature map is processed by a 1x1 convolution kernel to expand the channels, obtain intermediate features, and generate two sub-channels:
[0151] ,
[0152] wherein, is the feature map after channel expansion; is a 1x1 convolution operation on the input feature map ; is the number of expanded channels; is the original number of channels; and r is the ratio factor of channel expansion
[0153] ;
[0154] wherein, is the feature map after splitting; is a splitting function; is a splitting dimension;
[0155] The sub-feature map features are respectively processed by depthwise separable convolution, and are nonlinearly fused, and the specific expression is:
[0156] ,
[0157] wherein, is a depthwise separable convolution operation; is the feature map after nonlinear fusion; represents element-wise multiplication; and GELU is a Gaussian error linear unit activation function; and are the sub-feature map features after depthwise separable convolution;
[0158] to obtain an output feature map Y, and the expression is:
[0159] ;
[0160] wherein, is an output feature map; Q is a frequency domain quantization matrix; is a two-dimensional real inverse fast Fourier transform; is a real two-dimensional Fourier transform; Fold is a concatenation operation; and is an element-wise multiplication; is a 1x1 convolution dimension reduction to restore the original channel number; is a block expansion with a PXP size.
[0161] The specific steps are as follows:
[0162] First, the input feature map is expanded in channels by a 1x1 convolution kernel to obtain an intermediate feature and generate two sub-channels:
[0163] ,
[0164] wherein, is a feature map after channel expansion; is a 1x1 convolution operation on the input feature map ; is the number of expanded channels; is the original channel number; and r is a channel expansion ratio factor.
[0165] The feature map F is split into two parts:
[0166] ,
[0167] wherein, is a feature map after splitting; is a splitting function; is a splitting dimension.
[0168] The sub-feature maps are respectively processed by depth separable convolution, and are fused by a nonlinear function, and the specific expression is as follows:
[0169] ,
[0170] ,
[0171] wherein, is a feature map after depth separable convolution; is a feature map after nonlinear fusion; and is an element-wise multiplication; and GELU is a Gaussian error linear unit activation function.
[0172] Deep separable convolution is used to enhance the local receptive field and spatial perception ability; nonlinear fusion retains the nonlinear expression ability in traditional FFN, while enhancing the explicit coupling between features.
[0173] The feature map is reduced in dimension by 1x1 convolution and restored to the original dimension:
[0174] ,
[0175] In the formula, is the feature map after the original dimension is restored.
[0176] In order to extract the frequency information in the local area, the feature map is divided into a plurality of non-overlapping patches, and the specific steps are as follows:
[0177] The feature map is supplemented with a boundary, and the specific expression is as follows:
[0178] ,
[0179] In the formula, is the feature map after the boundary is supplemented; is the boundary supplement operation.
[0180] The feature map is split into patches, and the specific expression is as follows:
[0181] ,
[0182] In the formula, is the split patch set; The feature map is divided into a plurality of non-overlapping patch blocks; PXP is the size of each patch; H / P and W / P are the number of patches along the height and width directions.
[0183] A two-dimensional real fast Fourier transform (Real2DFFT) is performed on each patch in the patch set , and the specific expression is as follows:
[0184] ,
[0185] In the formula, is the transformed spectral representation; and are two-dimensional real fast Fourier transforms.
[0186] Introducing a learnable frequency domain quantization matrix On the spectral representation The weighted screening is performed, and the specific expression is:
[0187] ,
[0188] In the formula, is the spectral representation after weighted screening.
[0189] The above steps enhance the sensitivity of the target detection and tracking model to specific frequency structures, and enable the target detection and tracking model to actively learn which frequency components are more helpful in distinguishing the target features of traffic participants such as contour edges and high-frequency textures.
[0190] On the spectral representation Perform two-dimensional real inverse fast Fourier transform (Inverse FFT) to reconstruct the spatial domain, and the specific expression is:
[0191] ,
[0192] In the formula, is the Patch set of the reconstructed spatial domain; and are two-dimensional real inverse fast Fourier transforms.
[0193] Perform Patch reconstruction:
[0194] ,
[0195] In the formula, is the reconstructed feature map; is the Patch reconstruction operation.
[0196] The feature map After type restoration according to the original data type, the final output feature map is obtained.
[0197] Compared with the original backbone network of RT-DETR, the improved M-Block module realizes adjustable spatial perception enhancement by introducing multi-branch spatial convolution and gating mechanism, and at the same time combines Fourier analysis to regulate the feature frequency components, which significantly improves the image processing ability and further enhances the expression ability and stability of the model.
[0198] On this basis, combined with the improved SP-FFN module, the process of "spatial feature-nonlinear enhancement-frequency domain screening-reconstruction fusion" is completed, the spatio-temporal modeling capability is improved, the identification effect of small size targets such as pedestrians and vehicles in the video taken by the unmanned aerial vehicle is improved, and the identification effect in extreme complex scenes such as night and heavy fog is further improved, and the calculation complexity is low.
[0199] The intersection images collected by the unmanned aerial vehicle are processed, and the traffic participants including pedestrians, motor vehicles and non-motor vehicles are labeled to generate a data set. The data set is divided into a training set, a validation set and a test set according to a ratio of 8:1:1. The data set is used to train and test the target detection and tracking model, and the weight of the target detection and tracking model is saved.
[0200] The model complexity index and detection accuracy index are shown in Tables 1 and 2, wherein RT-DETR is the benchmark model without modification, RT-DETR-M is the benchmark model applying the improved M-Block module, and RT-DETR-M-SPFFN is the benchmark model applying the improved M-Block module and SP-FFN module.
[0201] It can be seen that the introduction of the M-Block module significantly improves the target detection accuracy, especially in the mAP50 and mAP50-95 indicators. This shows that it effectively improves the feature extraction and expression ability of the model through innovative module design; when M-Block is used in combination with SPFFN, this combination strategy shows unique advantages, and the combination not only maintains the reduction of the number of parameters and GFLOP, but also further improves the accuracy (Precision) and recall rate (ReCall) of the model.
[0202] As Figure 4 shown are the comparison detection result graphs before and after the model is improved, Figure 4 image (a) is the detection result of the benchmark model, Figure 4 image (b) is the detection result of the improved model using M-Block and SPFFN modules.
[0203] Table 1: Model complexity index
[0204] ,
[0205] Table 2: Model detection accuracy index
[0206] .
[0207] Step 2: Use the unmanned aerial vehicle to take the traffic operation video of the intersection to be detected.
[0208] As shown in the attachedFigure 5 As shown, the unmanned aerial vehicle is launched at the intersection (a), intersection (b) and intersection (c) during the morning and evening peak hours of weekdays, the flight height is maintained between 80-120 meters, the unmanned aerial vehicle is fixed at the intersection center, the latitude and longitude coordinates are recorded, the camera angle is adjusted to a top-down view, and recording is started after the body is stable, the video shooting length is fifteen minutes, and the video parameter is set to mp4 format, resolution is 1920*1080, and frame rate is 60 frames.
[0209] Step 3: input the traffic operation video collected at the urban intersection into the target detection and tracking model, obtain intersection information and traffic flow data, and calculate the motion information of the traffic participants. Among them, the intersection information includes intersection geometric properties, intersection functional properties, intersection flow data, and intersection conflict data; the traffic flow data includes traffic participant id, traffic participant category, traffic participant detection box coordinates, traffic participant detection box height and width, and traffic participant detection box frame; the traffic participant category includes pedestrians, non-motor vehicles and motor vehicles; the motion information includes speed, acceleration and motion direction.
[0210] Using the target detection and tracking model, each traffic participant in the unmanned aerial vehicle intersection aerial video is detected and tracked, and traffic flow data is obtained, and finally output as a csv file, and the specific output format is:
[0211] ,
[0212] Among them, represents the trajectory coordinate CSV file data of the i-th intersection. In the matrix, represents the ID number of the n-th traffic participant in the intersection video, represents the category of the n-th traffic participant, and respectively represent the horizontal coordinate and vertical coordinate of the upper left corner of the n-th traffic participant detection box, and respectively represent the width and height of the n-th traffic participant detection box, represents the n-th frame of video i. The final traffic flow data is extracted as shown in Figure 6 .
[0213] Further, using the traffic flow data, the speed, acceleration, motion direction and other information of each traffic participant are extracted, each traffic participant is segmented, the running track is divided into n small track segments, each small track segment is approximately a straight line, and the ratio of the Euclidean distance difference of each small track segment to the time frame is calculated. The speed of the small track segment is recorded as the location speed of the small track segment, and the speed is defined as:
[0214] ,
[0215] wherein, is the average speed (km / h) of the crossing through the intersection, and are the coordinates of the start and end points of the i-th trajectory, respectively, and are the time frames of the start and end points of the i-th trajectory, respectively. is the video detection frame rate.
[0216] The acceleration is defined as:
[0217] ,
[0218] The direction of motion is defined as:
[0219] .
[0220] Step 4: According to the traffic flow data and the motion information, statistical traffic conflict data is constructed to build potential conflict indicators for the intersection.
[0221] Extract potential conflict objects in the three intersections to be detected, the specific method is:
[0222] Statistical events with TTC value <2.5s extracted in the video, establish a set of potential traffic conflict objects, and then divide them into four categories: motor vehicle-motor vehicle, pedestrian-non-motor vehicle, non-motor vehicle-motor vehicle, and pedestrian-motor vehicle.
[0223] Define the intersection interaction conflict potential index (Interaction Conflict Potential, ICP) as:
[0224] ,
[0225] In the formula, are two traffic participants; is the time interval of the traffic participants reaching the potential conflict point; is a constant; is the set of predicted potential traffic conflict objects.
[0226] The larger the interaction conflict potential index, the higher the risk of conflict at the intersection.
[0227] Define the intersection movement disorder index (Movement Disorder Index, MDI) as:
[0228] ,
[0229] wherein, The probability distribution of the traffic participants in each direction is calculated as follows: First, the direction samples are extracted from the trajectory coordinate data of intersection A:
[0230] ,
[0231] where, is the displacement vector of the center point of the target detection frame in the adjacent frame.
[0232] Specifically, the is divided into sectors (8 by default), and the valid direction samples are traversed to assign each to its sector , and the direction count is obtained.
[0233] ,
[0234] When the motion directions of different vehicles are more dispersed and uniform, the entropy value is higher, indicating that the traffic flow is chaotic and difficult to predict; on the contrary, if most vehicles travel along a few paths, the entropy value is lower.
[0235] The intersection driving abnormal behavior index (Abnormal Behavior Index, ABI) is defined as:
[0236] ,
[0237] where, is the total number of vehicles in the statistical period, is the indicator function, is the preset acceleration threshold is the acceleration of the i-th vehicle.
[0238] The driving abnormal behavior index reflects the frequency of sudden deceleration behavior of traffic participants, and the larger the value, the more potential conflicts or disturbances.
[0239] The congestion pressure index (Congestion Pressure Index, CPI) of the intersection is defined as:
[0240] ,
[0241] ,
[0242] where, is the number of traffic participants at time t; is the physical area occupied by the traffic participants in the fixed time window, is the union area of all traffic participant target detection frames in the i-th frame, and T is the observation duration.
[0243] The greater the congestion pressure index is, the more congested the intersection is.
[0244] Define the intersection violation interference index (VII) as follows:
[0245] ,
[0246] Wherein, is the number of times of pedestrian or non-motor vehicle intrusion into the no-entry area per unit time, is the observation time. The no-entry area includes motor vehicle lanes, etc.
[0247] The higher the violation interference index value is, the more frequent the violation interference is.
[0248] As shown in Table 3, according to the above formula, the potential conflict index of each intersection is finally obtained:
[0249] Table 3: Initial potential conflict index of intersection
[0250] .
[0251] Step 5: According to the potential conflict index, a comprehensive safety level evaluation model framework is constructed, and a large model is used to combine intersection information and traffic flow data to improve the comprehensive safety level evaluation model framework, and finally obtain the comprehensive safety level evaluation model.
[0252] The comprehensive safety level evaluation model framework is established as follows:
[0253] Define the input vector: the interactive conflict potential index ICP, the motion disorder index MDI, the driving abnormal behavior index ABI, the congestion pressure index CPI and the violation interference index VII of each intersection form a feature vector, and the specific formula is:
[0254] ,
[0255] In the formula, is the conflict index vector; is the interactive conflict potential index ICP; is the motion disorder index MDI; is the driving abnormal behavior index ABI; is the congestion pressure index CPI; is the violation interference index VII;
[0256] At the same time, the public default weight is initialized, and the specific expression is:
[0257] ,
[0258] wherein, is the initial weight vector; is the weight value of the interactive conflict potential index ICP; is the weight value of the motion disorder index MDI; is the weight value of the abnormal behavior index ABI; is the weight value of the congestion pressure index CPI; is the weight value of the violation interference index VII;
[0259] The comprehensive grade evaluation model framework is improved using a large model (LLM) in combination with traffic flow data and intersection information.
[0260] Preferably, the DeepSeekR1 model is used to improve the comprehensive grade evaluation model framework, and the specific steps are as follows:
[0261] The user inputs the corresponding prompt (Prompt) to the large model for initialization, including but not limited to the following formats:
[0262] "You are now an expert in the field of transportation engineering, proficient in the safety evaluation system of urban road intersections, familiar with the threshold setting standards of traffic conflict parameters such as TTC and PET, and also have relevant background knowledge and model parameter tuning experience in deep learning. Your goal is to: analyze and adaptively optimize the weights based on user input multi-source data, accurately perceive the running state of the intersection, and realize precise quantitative evaluation of intersection safety risks.
[0263] Your answers need to specially note the following core principles:
[0264] Principle one: Always prioritize the safety of traffic participants as the fundamental requirement, and consider other factors such as traffic efficiency and potential risks on this basis.
[0265] Principle two: Dynamically adapt to the actual characteristics of each intersection to avoid ambiguity in decision-making, and adopt a conservative strategy if there is no sufficient and reliable data or arguments to support it.
[0266] Principle three: Maintain the logical interpretability of the answers, and provide relevant decision-making basis texts to meet the transparency requirements of safety evaluation.
[0267] After the output is complete, please consult me if there are any suggestions for improvement, and if there are suggestions, please output based on the above principles again. Do not repeat the content, and if you are ready, tell me.
[0268] At the same time, the user inputs the structured feature description information of the intersection to be detected to the large model, including but not limited to:
[0269] Intersection geometric properties: intersection type, number of lanes in each direction, width of sidewalk.
[0270] Intersection functional properties: intersection vehicle type, proximity to school / commercial area, peak hour flow ratio.
[0271] Intersection flow data: traffic participant trajectory coordinate data, historical traffic flow statistics.
[0272] Intersection conflict data: intersection historical accident records, intersection traffic violation handling.
[0273] The large model analyzes the risk mechanism of the intersection to be detected based on the traffic engineering knowledge base and multi-dimensional features, matches similar scenario case libraries, and the model outputs the thinking process, as follows:
[0274] As shown in Figure 5 Intersection (a) is located around the campus and should emphasize the safety behavior evaluation of pedestrians / non-motor vehicles; intersection (b) is in a mixed area of shopping malls and residential areas and should balance attention to potential conflicts and congestion behavior; intersection (c) is a well-developed arterial road and should focus on traffic order and illegal operation, etc.
[0275] After the thinking process is completed, the model finally generates weight adjustment strategies for the three intersections:
[0276] ,
[0277] ,
[0278] .
[0279] Step 6: Input the potential conflict index into the comprehensive safety level evaluation model for calculation, and output the comprehensive safety evaluation level of the intersection.
[0280] Standardize the initial potential conflict index of all intersections, and the specific expression is:
[0281] ,
[0282] In the formula, is the standardized potential conflict index feature vector; is the potential conflict index to be standardized; is the minimum value of all potential conflict indexes; is the maximum value of all potential conflict indexes;
[0283] To prevent the differences in the standardized results caused by sample input, the present embodiment adopts a city-level reference set, takes the 5% and 95% quantiles as the upper and lower bounds of the standardization of each indicator to cover most normal scenarios and improve the robustness of the evaluation effect, and the baseline calibration values are as shown in Table 4.
[0284] Table 4: Baseline calibration values of intersection parameters
[0285]
[0286] Taking the ICP indicator value of intersection (a) as an example, the calculation process is as follows:
[0287] The standardized risk value is:
[0288]
[0289] The standardized safety-type ICP indicator value (retaining three decimal places) is obtained:
[0290]
[0291] The calculation processes of the indicators of other intersections are the same, and the standardized intersection indicator values are finally obtained as shown in Table 5.
[0292] Table 5: Standardized potential conflict indicator values of intersections
[0293]
[0294] According to the weighted processing of the standardized potential conflict indicator characteristic vectors, the intersection safety evaluation score is obtained, and the specific expression is:
[0295]
[0296] In the formula, S is the intersection safety evaluation score; wi is the weight value of the ith potential conflict indicator; xi is the standardized value of the ith potential conflict indicator;
[0297] Among them, the safety level division rule is: 0.0≤S<0.4 is E dangerous, 0.40≤S<0.55 is D relatively dangerous, 0.55≤S<0.7 is C medium, 0.70≤S<0.85 is B relatively safe, and 0.85≤S<1 is A very safe.
[0298] The safety evaluation results of the intersections are shown in Table 6. The final result is that the risk level of intersection (a) is D (relatively dangerous), and after the improvement of the large model, more attention is paid to abnormal violations such as pedestrian crossing and mixed traffic flow, and the evaluation result is consistent with the current problems of the intersection, fully showing the real risk; the safety evaluation level of intersection (b) is C (medium), which shows that the model has made adaptive correction to the default authority on the basis of balancing attention to potential conflicts, making its output more accurate and comprehensive score; the safety evaluation level of intersection (c) is A, because the arterial channelization of this intersection is perfect, and the model fully quantifies the ability of the intersection to solve potential conflicts and face complex traffic flow scenarios on this basis, making its advantages more clear. The overall intersection safety evaluation is consistent with the actual situation, which shows that the intersection safety evaluation method based on unmanned aerial vehicles and large models has reliability.
[0299] Table 6: Intersection safety evaluation results
[0300] ,
[0301] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0302] The principles and implementation modes of the present application are described by applying specific examples in this paper. The above description of the embodiments is only used to help understand the core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed. In conclusion, the content of the specification should not be understood as a limitation of the present application.
Claims
1. A method for intersection safety evaluation based on a UAV and a large model, characterized in that, The application relates to a method for evaluating the comprehensive safety level of an intersection. The method comprises the following steps: a UAV is used to shoot a traffic operation video of an intersection to be detected; the traffic operation video is input into a pre-constructed target detection and tracking model to obtain intersection information and traffic flow data, and the motion information of a traffic participant is calculated; traffic conflict data is counted according to the traffic flow data and the motion information, and potential conflict indicators of the intersection are constructed; the potential conflict indicators include an interactive conflict potential index ICP, a motion disorder index MDI, an abnormal driving behavior index ABI, a congestion pressure index CPI and a violation interference index VII; a comprehensive safety level evaluation model framework is constructed according to the potential conflict indicators, a large model is used to improve the comprehensive safety level evaluation model framework by combining the intersection information and the traffic flow data, and a comprehensive safety level evaluation model is obtained; the potential conflict indicators are input into the comprehensive safety level evaluation model for calculation, and the comprehensive safety evaluation level of the intersection is output; wherein the traffic flow data includes a traffic participant id, a traffic participant category, traffic participant bounding box coordinates, a traffic participant bounding box height and width and a traffic participant bounding box frame; the traffic participant category includes a pedestrian, a non-motor vehicle and a motor vehicle; the motion information includes a speed, an acceleration and a motion direction; wherein the traffic conflict data is counted according to the traffic flow data and the motion information, and the potential conflict indicators of the intersection are constructed, which specifically include the following steps: events with a TTC value less than a preset time threshold in the traffic operation video are counted, a potential traffic conflict object pair set is established, and then the traffic conflict participants are divided into four categories, i.e., motor vehicle-motor vehicle, pedestrian-non-motor vehicle, non-motor vehicle-motor vehicle and pedestrian-motor vehicle; ; where i, j are two traffic participants; is the time interval for the traffic participant to reach the potential conflict point; ε is a constant; C is the set of potential traffic conflict objects predicted. the specific expression of the interactive conflict potential index ICP is as follows: ; wherein, is the probability distribution of the appearance of each directional traffic participant; the specific expression of the motion disorder index MDI is as follows: ; In the formula, N is the total number of vehicles in the statistical period, I is an indicator function, is a preset acceleration threshold value; is the acceleration of the ith vehicle. the specific expression of the abnormal driving behavior index ABI is as follows: ; ; wherein, is the number of traffic participants at time t; A is the physical area occupied by the traffic participants within a fixed time window, denotes the union area of all traffic participant bounding boxes in the i-th frame, and T is the observation duration. the specific expression of the congestion pressure index CPI is as follows: ; In the formula, is the number of times of target intrusion into the forbidden entry area per unit time, and T is the observation duration.
2. The intersection safety evaluation method based on UAV and large model according to claim 1, characterized in that, the specific expression of the violation interference index VII is as follows: the specific steps of using a UAV to shoot a traffic operation video of an intersection to be detected include the following steps: the UAV is taken off at the intersection to be detected and flown to a set height; 3.The intersection safety evaluation method based on UAV and large model according to claim 1, wherein, the UAV is fixed at the center of the intersection, the longitude and latitude coordinates are recorded, the camera angle is adjusted, the video shooting parameters are set, and the intersection is shot. The pre-constructed target detection and tracking model specifically includes:
4. The intersection safety evaluation method based on UAV and large model according to claim 3, characterized in that, a target detection and tracking model based on RT-DETR is constructed, and the backbone part and the AIFI part are improved. The improvement of the backbone part specifically includes: ; wherein Y is an output feature map; is a mapping function of the M-Block module; is a multi-scale spatial interaction module is a gain-adjusted intermediate variable; is a feature enhancement unit is an output of the feature enhancement unit; is an input feature map, B is a batch size of the feature map X, C is a channel number of the feature map X, H and W are height and width of the feature map X, respectively; is a normalization operation; is a scaling factor for adjusting a corresponding intensity of the residual; is a spatial interaction module; is a feature enhancement unit.
5. The intersection safety evaluation method based on UAV and large model according to claim 4, characterized in that, the Block module in the RT-DETR backbone network is replaced by an M-Block module, the introduced M-Block module follows a Metafromer structure, and the specific expression is as follows: The improvement of the AIFI part specifically includes: The input feature map is processed by a 1x1 convolution kernel Channel expansion is performed to obtain intermediate features, and two sub-channels are generated: , wherein, is the feature map after channel expansion; is a 1x1 convolution operation on the input feature map is the feature map after channel expansion; is the number of expanded channels; is the number of original channels; r is the ratio factor of channel expansion ; wherein is a feature map split sub-feature map; is a split function; is a split dimension; Sub-feature map features Respectively, the depth separable convolution processing is performed, and the nonlinear fusion is performed, and the specific expression is: , In the formula, is a depth separable convolution operation; is a nonlinear fused feature map; represents element-wise multiplication; GELU is a Gaussian error linear unit activation function; and are sub-feature maps respectively sub-feature maps after depth separable convolution the FNN module of the RT-DETR backbone network is replaced by an SP-FFN module, and the specific expression of the introduced SP-FFN module is as follows: an output feature map Y is obtained, and the expression is as follows: ; In the formula, is an output feature map; Q is a frequency domain quantization matrix; is a two-dimensional real inverse fast Fourier transform; is a real two-dimensional Fourier transform; Fold is a concatenation operation; and is an element-wise multiplication; is a 1x1 convolution dimension reduction to restore the original channel number; is a block expansion in a PXP size. 6.The intersection safety evaluation method based on UAV and large model according to claim 1, wherein, The step of constructing the comprehensive safety level evaluation model framework according to the potential conflict index comprises the following steps: The interaction conflict potential index ICP, the motion disorder index MDI, the abnormal behavior index ABI, the congestion pressure index CPI, and the violation interference index VII are combined into a conflict index vector, and the specific expression is as follows: ; wherein X is a conflict index vector; is an interaction conflict potential index ICP; is a motion disorder index MDI; is an abnormal behavior index ABI; is a congestion pressure index CPI; is a violation interference index VII; The default weight is initialized, and the specific expression is as follows: ; wherein, is the initial weight vector; is the weight value of the interaction conflict potential index ICP; is the weight value of the motion disorder index MDI; is the weight value of the abnormal behavior index ABI; is the weight value of the congestion pressure index CPI; is the weight value of the violation interference index VII; The comprehensive safety level evaluation model framework is generated. 7.The intersection safety evaluation method based on UAV and large model according to claim 1, wherein, The step of improving the comprehensive safety level evaluation model framework by using a large model in combination with the intersection information and the traffic flow data comprises the following steps: The traffic flow data and the intersection information are input into a large model; the intersection information comprises intersection geometric attributes, intersection functional attributes, intersection flow data, and intersection conflict data; The weight setting in the comprehensive safety level evaluation model framework is given by using a large model for reasoning, and a comprehensive safety level evaluation model is generated.
Citation Information
Patent Citations
A method for evaluating the safety state of a road intersection
CN109146299A
Intersection safety state perception and diagnosis treatment system and method based on unmanned aerial vehicle video
CN115907462A
Intersection safety evaluation method based on unmanned aerial vehicle detection
CN117877272A
Unmanned aerial vehicle aerial photography small target detection method based on improved RT-DETR network
CN118521929A
Power system adaptive evaluation system construction method based on LLM and analytic hierarchy process
CN119337978A