Road target detection method and system fusing multi-source perception and attention enhancement
By generating spatial, static, and dynamic risk weight maps and combining them with image enhancement techniques, the problem of detecting small road obstacles in unstructured construction scenarios was solved, achieving highly robust and real-time road target detection and ensuring traffic safety.
Patent Information
- Application Number
- CN202610557883.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-24
- Publication Date
- 2026-07-14
AI Technical Summary
Existing road target detection technologies suffer from several problems in unstructured construction scenarios, including the easy omission of small road obstacles, the impact of construction on optical image quality, and the lack of perception bias in key areas when fusion of multi-source sensing data, resulting in insufficient detection sensitivity and real-time performance.
By acquiring historical video streams, real-time image frames, and 3D laser point cloud data, spatial weight maps, static risk weight maps, and dynamic risk weight maps are generated. A unified attention guidance map is constructed to perform multimodal small target detection. Combined with image enhancement technology, the detection accuracy and robustness are improved.
It significantly improves the detection accuracy and real-time performance of small-target road obstacles in unstructured road environments, providing timely early warning information and ensuring traffic safety.
Smart Images

Figure CN122392024A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and more specifically, to a method and system for road target detection that integrates multi-source perception and attention enhancement. More specifically, this invention is particularly applicable to scenarios in intelligent transportation systems where robust detection of specific targets in unstructured or semi-structured road environments is required. Background Technology
[0002] Highways and urban expressways are crucial components of modern transportation networks, handling high-density, high-speed traffic. During road construction, expansion, or maintenance, unstructured construction zones inevitably arise. These scenarios typically lack clear, fixed road markings or highway guardrails, instead relying on numerous temporary, movable barriers such as traffic cones, warning signs, water-filled barriers, or makeshift fences to demarcate actual passable areas, construction zones, and median strips. Due to the temporary and volatile nature of the traffic environment, coupled with high vehicle speeds, these temporary barriers (especially small ones) pose a serious threat to traffic safety if they shift, tip over, or are accidentally pulled into the driving lane. Therefore, accurate and robust detection of key targets in these unstructured road environments is fundamental to ensuring traffic safety and efficiency in construction zones.
[0003] Existing road target detection technologies primarily rely on optical image sensors. However, in specific scenarios such as highway maintenance, vision-based perception methods face significant challenges. Firstly, road barriers used for warning and isolation in construction zones are often small targets like traffic cones. Under drone aerial photography or long-distance vehicle-mounted views, these targets occupy very few pixels in the image, making them easily missed in complex backgrounds. Secondly, frequent work activities in construction areas often lead to environmental interference such as dust and sandstorms, significantly reducing the clarity and contrast of optical images. Further degradation of image quality under poor lighting conditions or inclement weather directly results in detection algorithms failing to meet the safety requirements of high-speed traffic in terms of performance, reliability, and real-time performance.
[0004] To compensate for the shortcomings of optical sensors, the industry has introduced multi-source sensors such as LiDAR. LiDAR actively emits lasers to acquire precise three-dimensional spatial coordinates and geometric structures of targets, exhibiting strong robustness to changes in illumination and dust. However, LiDAR point cloud data is typically sparse and lacks rich color and texture information, resulting in inherent limitations when used alone for classifying small targets. Therefore, multi-source sensing data fusion has become a current research hotspot. Existing multimodal fusion strategies, while combining the advantages of different modalities to some extent, are still insufficient when dealing with unstructured construction scenarios. These methods typically perform indiscriminate feature extraction and fusion across the entire scene, lacking a focus on key areas. In environments with numerous temporary roadblocks and complex backgrounds, computational resources are heavily allocated to non-critical areas, while insufficient attention is paid to dynamic small targets that are most prone to accidents and accidentally intrude into passable areas, resulting in a need to improve detection sensitivity and real-time performance. Summary of the Invention
[0005] This invention provides a road target detection method that integrates multi-source perception and attention enhancement, comprising the following steps: Historical high-definition video streams are acquired, and based on vehicle trajectory sequences, passable areas in the bird's-eye view space are determined and a spatial weight map is generated. A static risk weight map is generated based on the kinematic characteristics of the trajectory sequences. Acquire real-time image frames and 3D laser point cloud data, perform image enhancement on the real-time image frames, and obtain the enhanced image; Track real-time short-term trajectories in real-time high-definition image frames, calculate the morphological deviation between the real-time short-term trajectory and the historical prototype trajectory set, and generate a dynamic risk weight map; A unified attention guidance graph is constructed by integrating spatial weight graphs, static risk weight graphs, and dynamic risk weight graphs; the unified attention guidance graph is then projected onto the image space to perform image spatial attention enhancement on the enhanced image. Multimodal small target detection is performed based on attention-enhanced images and real-time 3D laser point cloud data.
[0006] Generating a spatial weighted graph includes: Based on the binary mask of the passable area Aerial view space Semantic region segmentation is performed to obtain the construction area mask. and isolation area mask ; and according to preset weight parameters , , Calculate the spatial weighted graph : in, The coordinates are for the bird's-eye view.
[0007] Generating a static risk weight diagram includes: Calculate trajectory sequence Any point in the bird's-eye view space instantaneous rate at Instantaneous acceleration magnitude and instantaneous curvature ; Speed risk based on normalization Acceleration risk and curvature risk Calculate instantaneous risk score : in, It is a preset risk combination weight.
[0008] Generating a static risk weight diagram also includes: Instantaneous risk score of all trajectory points Generate a risk accumulation chart ; Generate trajectory counting map ; The static risk weight map is obtained by calculating the average risk score. : in, For Gaussian kernel function, To prevent extremely small positive numbers with a denominator of zero, This represents the total number of trajectories.
[0009] The morphological deviation calculation includes: for each real-time short-time trajectory Calculate its relationship with the set of historical prototype trajectories. Each historical prototype trajectory Discrete Fréchet distance between The minimum discrete Fréchet distance is taken as the real-time short-time trajectory. morphological deviation : in The number of clusters for the historical prototype trajectory.
[0010] Generating a dynamic risk weighting diagram includes: incorporating morphological deviation. Greater than the morphological deviation threshold Real-time short-term trajectory identification is a set of instantaneous abnormal trajectory patterns. ; based on Update the anomaly accumulation graph using time decay and group activation. ; Thresholding is applied to the anomaly accumulation graph to obtain a dynamic risk weight graph. : in, A binary mask for the passable area. This is the threshold for counting anomalies in the population.
[0011] based on Using the construction area determined by the spatial weight map as a priori, and through the priori of the hidden passage... Estimating global atmospheric light ; through guided filtering Optimize transmittance By inverting the model Calculate the enhanced image ;in, For pixel coordinates, For the enhanced image to be solved, This is a preliminary estimate of the transmittance. To guide the image, For filter parameters, This is the lower limit of transmittance.
[0012] Constructing a unified attention guidance graph Includes: fusion generation ; And on Normalization yields a unified attention guidance map. ; Will Projecting onto the image space yields the image space attention map. : in, For the spatial coordinates of the bird's-eye view, As weight, For the image after attention enhancement, For three-channel copying of the image spatial attention map, Hadama product.
[0013] The method also includes: obtaining preliminary detection boxes for multimodal small target detection. Extract a subset of the lidar point cloud within it. And calculate point cloud Maximum value of coordinates: ;based on Compared with the preset upright height threshold and the threshold of lodging height Determine the physical state of the target Based on physical state and the target center point The location within the semantic region determines the risk level. It will also output alarm information.
[0014] This invention also proposes a road target detection system that integrates multi-source perception and attention enhancement, the system comprising: Spatial weight generation module: acquires historical high-definition video streams and, based on vehicle trajectory sequences, determines passable areas in the bird's-eye view space and generates a spatial weight map; Static risk weight generation module: Generates a static risk weight map based on the kinematic features of the trajectory sequence; Image enhancement module: acquires real-time image frames and 3D laser point cloud data, performs image enhancement on the real-time image frames, and obtains enhanced images; Dynamic risk weight generation module: tracks real-time short-term trajectories in real-time high-definition image frames, calculates the morphological deviation between the real-time short-term trajectory and the historical prototype trajectory set; and generates a dynamic risk weight map. Attention Enhancement Module: This module integrates spatial weight maps, static risk weight maps, and dynamic risk weight maps to construct a unified attention guidance map. The unified attention guidance map is then projected onto the image space to perform image space attention enhancement on the enhanced image. Target detection module: Performs multimodal small target detection based on attention-enhanced images and real-time 3D laser point cloud data.
[0015] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aforementioned road target detection method that integrates multi-source perception and attention enhancement.
[0016] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned road target detection method that integrates multi-source perception and attention enhancement.
[0017] The road target detection method and system integrating multi-source perception and attention enhancement provided by this invention have significant beneficial effects. This invention is specifically designed for unstructured road environments formed during the new paving, maintenance, or expansion phases of highways or urban roads. It aims to solve the target detection challenges in such scenarios caused by the lack of fixed road markings, reliance on temporary road barriers, and susceptibility to construction interference, significantly improving the accuracy, robustness, and real-time performance of the detection.
[0018] This invention analyzes the trajectory sequences of group vehicles in historical video streams, using the driving consensus of the group of vehicles as a basis to deduce and determine the actual passable areas in the bird's-eye view space. In application scenarios lacking clear road markings and fixed guardrails, this method based on actual traffic experience demonstrates extremely high reliability in determining the boundaries of passable areas. Building upon this, the invention further performs semantic region segmentation of the scene, identifying passable areas, restricted areas, and construction areas, and constructs a static spatial weight map. This allows the system to establish a stable priori understanding of the scene from a macroscopic perspective, prioritizing the allocation of perception resources to high-risk driving areas and restricted areas containing roadblocks, while reasonably reducing attention to non-traffic threat construction areas, thus achieving optimized allocation of perception resources.
[0019] This invention not only focuses on static spatial risks but also delves deeper into risk information within time and behavior. By analyzing the kinematic characteristics of historical trajectory sequences, this invention constructs a static risk weight map. This map can quantitatively identify inherently high-risk road sections within passable areas where drivers generally adopt cautious driving due to road morphology. More importantly, this invention designs a dynamic risk weight map generation mechanism. This mechanism calculates the morphological deviation between real-time short-term trajectories and historical prototype trajectory sets, and combines this with a group anomaly counting threshold to accurately capture group abnormal detour behaviors caused by newly appearing road obstacles (such as knocked-down traffic cones). This design effectively filters out individual behavioral interference caused by abnormal driving by a single driver, ensuring that the dynamic risk map only responds to real, sudden safety hazards with high confidence, greatly improving the timeliness and accuracy of early warnings.
[0020] This invention addresses the challenge of severe interference with optical image quality caused by construction dust by proposing an image enhancement scheme deeply coupled with the scene. This method utilizes the accurately identified construction area from a previous step as a priori source for estimating atmospheric light parameters. This avoids the pitfall of traditional dust removal algorithms misclassifying other bright objects in the scene (such as white vehicles) as atmospheric light. The estimated dust parameters are thus more accurate. Combined with guided filtering to optimize the transmittance map, this effectively removes the impact of dust while preserving the edge contour information of small road obstacles to the maximum extent, providing high-quality input data for subsequent visual detection. Finally, this invention efficiently fuses and applies the aforementioned multiple prior information. A unified attention guidance map is constructed and back-projected onto the high-resolution original image space, performing pixel-level attention enhancement on the enhanced image. This ensures that the prior knowledge of group validation takes effect before the data enters the neural network backbone, forcing the model to give the highest response to pixels in high-risk areas from the initial stage of feature extraction. Attached Figure Description
[0021] Figure 1 This is a flowchart of the road target detection method that integrates multi-source perception and attention enhancement according to the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0023] This embodiment discloses a road target detection method that integrates multi-source perception and attention enhancement, including the following steps: Historical high-definition video streams are acquired, and based on vehicle trajectory sequences, passable areas in the bird's-eye view space are determined and a spatial weight map is generated. A static risk weight map is generated based on the kinematic characteristics of the trajectory sequences. Acquire real-time image frames and 3D laser point cloud data, perform image enhancement on the real-time image frames, and obtain the enhanced image; Track real-time short-term trajectories in real-time high-definition image frames, calculate the morphological deviation between the real-time short-term trajectory and the historical prototype trajectory set, and generate a dynamic risk weight map; A unified attention guidance graph is constructed by integrating spatial weight graphs, static risk weight graphs, and dynamic risk weight graphs; the unified attention guidance graph is then projected onto the image space to perform image spatial attention enhancement on the enhanced image. Multimodal small target detection is performed based on attention-enhanced images and real-time 3D laser point cloud data.
[0024] The technical solution provided in this embodiment is applicable to highways or urban expressways that are in the process of new paving, repair, or expansion. Such work scenarios typically involve encroaching on existing roads for construction, resulting in a typical unstructured or semi-structured road environment. In these scenarios, existing clear fixed road markings (such as lane lines and edge lines) may be obscured by newly paved asphalt, milled away, or rendered inapplicable due to construction needs; simultaneously, fixed guardrails and other boundaries along the highway may also be missing in the construction area.
[0025] To guide high-speed traffic safely around construction zones, numerous temporary, movable road barriers are typically deployed on-site. These barriers include, but are not limited to: traffic cones (also known as cones), water-filled plastic barriers, temporary metal fences, and construction warning signs (including signs indicating construction ahead, slow down, or mannequins). These temporary road barriers have relatively poor physical stability, especially under strong wind disturbances generated by heavy truck traffic at high speeds, or due to accidental collisions between vehicles in narrow passageways. This can easily lead to unexpected displacement, tilting, or even complete collapse of the temporary road barriers.
[0026] The core safety hazard in such scenarios lies in the fact that once these temporary roadblocks deviate from their intended placement, are knocked down, or become entangled in high-speed traffic, they instantly transform from traffic guidance facilities into road obstacles that endanger driving safety. If high-speed vehicles fail to detect and avoid these abnormal roadblocks in time, it can easily lead to serious traffic accidents such as emergency braking, loss of control, or chain collisions. Therefore, the method in this embodiment aims to provide real-time, robust target detection in such dynamic, high-risk unstructured road environments, particularly for small target roadblocks that have undergone abnormal displacement or changes in state, to provide timely early warning information.
[0027] In this embodiment, an industrial-grade drone equipped with a multi-source sensing module is launched. This drone needs to be capable of continuously hovering at high altitudes over unstructured road areas or cruising along a preset route.
[0028] The multi-source sensing module specifically includes a high-resolution visible light camera and a high-line-count lidar. Preferably, the high-resolution visible light camera is a CMOS sensor with 4K resolution and a global shutter. The lidar is preferably a 32-line or 64-line hybrid solid-state lidar to ensure that sufficient density of 3D laser point cloud data can be acquired on the ground even at flight altitude. The UAV continuously flies and takes pictures over the target construction section. During this process, the control system simultaneously acquires and records high-definition video streams and 3D laser point cloud data. To achieve subsequent multimodal data fusion, the acquired multi-source data must undergo rigorous spatiotemporal alignment preprocessing.
[0029] To achieve precise temporal correspondence between image frames and point cloud data at any given time, this embodiment employs a hardware synchronization scheme based on a unified clock source. Specifically, both the high-resolution visible light camera and the LiDAR are connected to a unified timing system, which, for example, acquires a high-precision pulses per second (PPS) signal by receiving GPS signals. Using this PPS signal as a reference, a unified timestamp is applied to all acquired high-definition image frames and each frame of 3D LiDAR point cloud data. By comparing the timestamps, image frame and point cloud data pairs acquired at the same time can be extracted from the data stream.
[0030] Spatial alignment will adjust the coordinate system of the lidar. The three-dimensional points below Transform to high-resolution visible light camera coordinate system pixel coordinates below .
[0031] Intrinsic parameter calibration was performed on a high-resolution visible light camera. Specifically, the Zhang Zhengyou calibration method was used, and the camera's intrinsic parameter matrix was obtained by photographing a checkerboard calibration board in different poses. and distortion coefficient Intrinsic parameter matrix Define the camera coordinate system The projection relationship between three-dimensional points and the pixel coordinate system.
[0032] in, Focal length Principal point coordinates The tilt factor is used. The distortion coefficient is employed. It can perform distortion correction on the acquired raw images.
[0033] Next, extrinsic parameter calibration is performed between the lidar and the camera. Extrinsic parameter calibration determines the lidar coordinate system. To the camera coordinate system rotation matrix Translation vector This embodiment employs a method based on a joint calibration board. In the calibration field, images and point cloud data of the calibration board are simultaneously acquired from multiple angles. Corner points are extracted from the images, and the PnP algorithm is used to estimate the calibration board's position. The pose is determined by extracting planar or edge features from the point cloud to estimate the calibration board's position. The pose is determined by minimizing the reprojection error or feature point alignment error, and then jointly optimizing to obtain a high-precision extrinsic parameter matrix. .
[0034] After completing the spatiotemporal alignment, for any lidar point that matches on the timestamp... The pixel coordinates are projected onto the distorted image plane according to the following formula. : in Is this point in the camera coordinate system? The depth value below.
[0035] Next, this embodiment specifically illustrates the process of generating passable areas based on group trajectories in the method of the present invention. The purpose of generating passable areas is to use the driving consensus of actual passing vehicles in unstructured road environments lacking clear road markings and fixed guardrails to reverse-engineer and determine the truly safe and passable area boundaries under the current road conditions.
[0036] Specifically, from the time-aligned and spatially aligned high-definition video stream, a time window is selected; preferably, all video frame sequences contained within a 5-minute period of stable traffic flow are selected. ,in This represents the total number of frames.
[0037] S2-1 Trajectory sequence extraction based on multi-target tracking, which requires video sequences All motor vehicles are continuously tracked to extract the trajectory of each vehicle. Considering the potential for temporary obstruction due to construction dust in highway maintenance scenarios, or vehicle trajectories intersecting due to irregular placement of temporary roadblocks, this embodiment employs a multi-target tracking algorithm that combines high and low confidence bounding box association. The process specifically includes: S2-1-1 Object Detection: For video sequences Each frame in ( The system utilizes a pre-trained object detector (YOLOX model) for detection. The object detector is trained specifically for the motor vehicle category. For each frame... The detector outputs a set of bounding boxes. and confidence score Test results .
[0038] S2-1-2 Two-Stage Trajectory Association: Setting a High-Confidence Threshold and low confidence threshold .Will All detection box Extract it. Use a Kalman filter to process the existing trajectory. Perform state prediction, that is, predict its state in the 1st... The position of the frame. Then, calculate... Middle detection frame and The Intersection over Union (IoU) ratio between predicted bounding boxes is used to update successfully matched trajectories. For trajectories that did not match in stage (i)... and The remaining low-confidence detection boxes , i.e., confidence level satisfy Then perform the association matching again.
[0039] In unstructured construction sections, vehicles may cause detectors to output brief low-confidence results due to avoidance maneuvers or changes in lighting. Traditional trackers would lose targets at these points, resulting in broken tracks. The two-stage association strategy employed in this embodiment can recover occluded or blurred vehicles using low-confidence bounding boxes. This is crucial for obtaining complete and continuous vehicle tracks, as incomplete tracks will lead to breaks in the subsequently generated passable area model.
[0040] S2-1-3 Trajectory Output: Repeat the above process for all video frames. For vehicles that are successfully tracked continuously, record their image coordinates in each frame to form a trajectory sequence. ,in This represents the total number of vehicles tracked within that time window. Each trajectory... ( ) represents .in, This indicates that the j-th trajectory in the i-th frame is formed by ( The coordinates are represented by ) These are the two coordinate components of the j-th trajectory in the i-th frame. These represent the start and end frames of the j-th trajectory, respectively.
[0041] S2-2 Trajectory projection and density map generation, extracting the trajectory These are coordinate sequences on a two-dimensional image plane. To eliminate the perspective effect caused by the drone's viewpoint and to obtain the actual traffic density in physical space, these trajectories need to be projected onto a unified bird's-eye view space.
[0042] S2-2-1 Perform perspective transformation calculation: Highways or expressways have relatively flat terrain, therefore they can be transformed through a... homography transformation matrix To achieve image coordinates To BEV coordinates The mapping.
[0043] S2-2-2 Generating a trajectory heatmap: Creating a two-dimensional grid This is used to represent the BEV space. The set... All trajectory points Transform to BEV space to obtain the corresponding .
[0044] To generate a smooth trajectory heatmap ,exist any position on Its trajectory density The calculation is as follows: in, It is a Gaussian kernel function. It is a continuous grayscale image, and its brightness reflects the frequency of vehicle traffic.
[0045] S2-3. Accessible area extraction: trajectory heatmap This represents the consensus on the passage of a group of vehicles. The next step is to extract the high-density core areas from this density map to obtain the passable areas.
[0046] S2-3-1 Set a density threshold This threshold represents the minimum standard of consensus among the group, that is, at least how many vehicles must pass through an area to be considered passable. (For density maps) Binarization is performed to obtain a binary image. : S2-3-2 Morphology-Based Boundary Extraction: Binarized Images Small holes or isolated noise points may exist. This embodiment uses morphological closing operations, first performing dilation and then erosion to fill these holes and smooth the edges of the region, obtaining the final binarized mask of the passable region. .
[0047] Based on the identified passable area, the remaining area within the bird's-eye view (BEV) space is automatically semantically segmented, ultimately generating a spatial weight map that provides attention guidance for subsequent detection tasks.
[0048] S3-1 Generate BEV orthophoto To obtain the visual features of non-traffic areas in the BEV space, it is necessary to generate a BEV orthophoto, i.e., a BEV color map.
[0049] S3-1-1 Select a reference frame, i.e., from the high-definition video stream. Select a representative reference frame from the data. Preferably, Image frames taken when the drone's attitude is stable and the lighting conditions are good.
[0050] S3-1-2 Inverse Perspective Transform Calculation: Traversing the BEV Mesh Each coordinate point on ,pass Back-project it onto the reference frame In the image coordinate system, the corresponding pixel coordinates are obtained. : in , , These are homogeneous coordinate components. After normalization, we obtain... : S3-1-3 color mapping, exist The pixel color value at that location is assigned to the BEV mesh. On Point. Traverse all Then, the BEV orthophoto was generated. .
[0051] S3-2 Semantic Region Segmentation: Based on Binarization Mask of Passable Regions and BEV orthophoto The BEV space is semantically divided into three mutually exclusive regions: passable region, construction region, and isolation region.
[0052] S3-2-1 Passable area is from The mask is defined directly, that is The area.
[0053] S3-2-2 Non-passable area is defined as all areas outside the passable area, and its mask is... for: The non-trafficible area is further divided into a construction area and a restricted area. The construction area is the main working surface for road construction, typically consisting of newly laid asphalt and exposed soil. These characteristics make it visually (in color and texture) significantly different from the trafficable area. In this invention, the two main visual features of the construction area—newly laid asphalt and large areas of exposed loess—are themselves products of large-scale homogenization operations. For newly laid asphalt, asphalt pavement laying is a highly standardized industrial process. Large pavers continuously and uniformly spread the same batch and proportion of asphalt mixture onto the predetermined subgrade. This new pavement is completely consistent in physical composition, color, and texture on a macroscopic scale. For large areas of exposed loess, when carrying out subgrade work or road expansion, whether it's large-scale excavation or backfilling and compaction, the exposed working surface consists of the same type of geological material or engineering filler with consistent properties. After being leveled, this single material naturally presents a consistent hue (such as yellowish-brown). This invention then utilizes the visual consistency of the construction area for extraction.
[0054] S3-2-3 Extract the construction area. This embodiment utilizes... Image of non-traffic areas Segmentation is performed.
[0055] S3-2-3-1 Uniform Region Segmentation Based on Superpixels: To avoid pixel-level noise interference and obtain uniformity features of local regions, the BEV orthophoto is segmented... Superpixel segmentation is performed. This embodiment uses the SLIC algorithm. The average size of the generated superpixels is... The SLIC algorithm in the Lab* color space and Local clustering is performed in space to generate a set A compact and well-defined superpixel This invention only concerns areas located in non-traffic zones. Superpixels within. Calculate each superpixel. ( )and overlap rate To satisfy all superpixels The filtered superpixels form a set of non-travelable areas. , This is the overlap threshold.
[0056] S3-2-3-2 Extracting Superpixel Features: For Each superpixel in Extract feature vectors that characterize its average color and texture uniformity. .
[0057] calculate All pixels within L a The three-dimensional color features are obtained by multiplying the average value of the color space by b*. .calculate All pixels within Variance of (brightness) values . The smaller the value, the more uniform the texture within the superpixel.
[0058] The two are combined to form the final four-dimensional feature vector of the superpixel. .
[0059] S3-2-3-3 Clustering and Recognition Based on GMM, All feature vectors As a dataset, it is clustered using a Gaussian Mixture Model (GMM).
[0060] In this invention, these superpixels are mainly composed of It is a mixture of several categories. (Settings) The Gaussian Mixture Model (GMM) can learn three categories: Category 1 (loess construction area), Category 2 (asphalt construction area), and Category 3 (isolation area / debris area). The GMM is trained using the expectation-maximization algorithm to obtain... Parameters of Gaussian components . It is the first Mixed weights of the components. It is the first The four-dimensional mean vector of each component, i.e., the class center. It is the first Each component Covariance matrix. Set a texture variance threshold. To satisfy all The amount Classified as a construction area component, its index set is denoted as .
[0061] S3-2-3-3 Generate construction area mask ,for Each superpixel in Calculate its belonging to The posterior probability of each component Find the component index to which it most likely belongs. .if If the superpixel most likely belongs to a construction area component, then the superpixel is... All pixels within Marked as construction area.
[0062] S3-2-4 Extracting the isolation area: In this embodiment, the remaining portion of the non-traffic area that is not identified as a construction area is defined as the isolation area, and its mask is... for: By this definition, , and This constitutes the BEV space A complete and mutually exclusive semantic division.
[0063] Based on the above semantic segmentation, a static spatial weight graph is constructed. Spatial weighting maps provide prior attention guidance for small-object obstacle detection models, concentrating computational resources on high-risk areas.
[0064] This embodiment assigns different risk weight values to three semantic regions. A higher weight value indicates a higher risk of the presence of the target to be inspected (an abnormal "temporary roadblock") in that region, or a higher priority for detecting that region. The passable area is the high-speed traffic area. Based on scenario safety hazard analysis, any temporary roadblock entering this area will pose the most serious direct threat. Therefore, this area is assigned the highest weight. The isolation area is where temporary roadblocks should be located. The purpose of detecting this area is to determine the state of the roadblock, i.e., whether it is tilted or even completely collapsed, and whether it is about to deviate from its intended placement position. Its risk level and detection priority are the second highest. The construction area is the construction work surface and should generally not have temporary roadblocks used to guide traffic. Therefore, it is assigned the lowest weight.
[0065] A specific set of weight values is set, resulting in the final spatial weighted graph. In any value at The calculation is as follows: in, , These are the weight parameters for the passable area, the restricted area, and the construction area, respectively.
[0066] Spatial weight graph The aim is to focus the detection efforts on macro-level issues within a specific area. However, within traversable areas... Internally, the inherent risks of different road sections are not equal. For example, during highway maintenance, the inherent risks of S-curves or narrow bottlenecks created by temporary roadblocks are far greater than those of straight sections. The purpose of this step is to analyze the driving consensus of a group of vehicles within the passable area. Internally, it identifies local high-risk areas that generally make drivers feel nervous or require cautious driving, and generates a static risk weight map.
[0067] S4-1: Extract trajectory motion features. To analyze driving behavior, it is necessary to extract the motion features of each trajectory in the bird's-eye view space. Perform kinematic analysis. Before analyzing kinematic characteristics, it is necessary to determine their scale relative to the physical world, defining a physical scale factor. , which represents the actual physical distance corresponding to a unit coordinate length in the bird's-eye view space. Through homography transformation matrix Based on known physical reference calibrations, lane width is preferably used as the physical reference calibration.
[0068] At the same time, define the inter-frame time interval. Its value is the reciprocal of the frame rate of the high-definition video stream.
[0069] Instantaneous velocity and acceleration calculation: for trajectory any point on (in Its instantaneous velocity vector and instantaneous rate The calculation is as follows: Its instantaneous acceleration magnitude ,Pick : The absolute value of acceleration is used because both rapid acceleration and rapid deceleration (braking) reflect drastic changes in driving behavior and are considered risk indicators.
[0070] Instantaneous curvature calculation; curvature reflects the degree of bending of the trajectory. For Point at Analyze its two velocity vectors before and after. and The angle between .
[0071] in It is an inverse cosine function. Instantaneous curvature. This can be expressed as the change in angle per unit distance traveled: Here, when When the value is close to 0, to avoid division by zero, Set it to 0.
[0072] S4-2: Constructing a point-by-point risk quantification model. Based on the aforementioned kinematic characteristics, a point-by-point risk quantification model is constructed. Driving behaviors with low speed, high acceleration (acceleration / deceleration), and high curvature (sharp turns) correspond to higher risks. To eliminate the influence of dimensions, three normalization thresholds are set: maximum speed threshold... Maximum acceleration threshold Maximum curvature threshold .
[0073] Speed risk (Low speed, high risk): Acceleration risk (High speed, high risk): Curvature risk (High curvature, high risk): The trajectory points are obtained by weighted summation of the three normalized risk characteristics. Instantaneous risk score : in, It is a preset risk combination weight.
[0074] In unstructured construction sections, (Curvature weight) is assigned a higher value because of high curvature. This always means that the vehicle is forced to detour or make an S-shaped turn due to a temporary roadblock, which is a direct source of safety hazards. and This reflects the driver's level of hesitation and caution when performing this detour, i.e., deceleration and braking.
[0075] S4-3: Generate a static risk weight graph All trajectories All point-by-point risk scores Two-dimensional grid aggregated into bird's-eye view space Generate a risk heatmap.
[0076] Risk Accumulation Chart : Trajectory Counting Map : in, It is a Gaussian kernel function.
[0077] Static Risk Weights Chart Defined as Each position on the grid Average risk score: in It is a very small positive number used to prevent the denominator from being zero. The purpose of using mean squared values is to eliminate the influence of traffic flow itself, ensuring that the map reflects the inherent driving difficulty / risk of that location, rather than how many cars pass through. A region with low traffic flow but where all cars make sharp turns and slow down (high...) The risk is far higher in an area with high traffic volume but all vehicles traveling at a constant speed in a straight line (low risk). ).
[0078] Static risks only exist in passable areas. Only then does it have meaning. Therefore, and Pixel-by-pixel multiplication yields the final static risk weight map. : Next, this embodiment describes the online real-time detection stage of the method of the present invention.
[0079] S5-1: Real-time data acquisition. During the online detection phase, an industrial-grade drone equipped with a multi-source sensing module continuously hovers over the target unstructured road area. The control system simultaneously acquires current high-definition image frames and 3D laser point cloud data. In this embodiment, Real-time high-definition image frames acquired at any time are denoted as At the same time The real-time laser point cloud collected is denoted as The collected and After spatiotemporal alignment preprocessing.
[0080] S5-2: Image Enhancement Based on Construction Area Highway maintenance is often accompanied by "construction dust." This leads to the loss of high-definition image frames acquired in real time. Clarity issues are frequently encountered, including decreased image contrast, color distortion, and blurred details. This degradation in image quality severely interferes with subsequent visual detection of small roadblocks.
[0081] The purpose of this step is to utilize the prior knowledge obtained, namely the construction area. Location, real-time images of dust pollution Perform targeted image enhancement to generate a clear, enhanced image. .
[0082] in, In pixels The color values of the real-time image observed at the location. is the color value of the enhanced image to be solved. A is the global atmospheric light parameter, i.e., the global ambient light caused by dust scattering. This is a transmittance map, representing the transmittance at the pixel level. Original lighting in the scene What percentage of the dust penetrated to reach the high-resolution visible light camera? The smaller the value, the denser the dust.
[0083] S5-2-1: Estimation of Atmospheric Light A Based on Construction Area Traditional dehazing / dust removal algorithms typically assume that the brightest or whitest areas in an image are atmospheric light. However, in this invention, there may be bright, non-dust-generating objects in the scene (such as white vehicles or reflective temporary metal fences), causing serious errors in estimation by traditional methods.
[0084] This invention utilizes the extracted construction area This issue is addressed using a strong prior. The construction area mask is used within the BEV space. Projected onto the current real-time image On the frame, the image space construction area mask is obtained. The construction area is a major source of construction dust, and this area is typically large, open, and lacks highly reflective objects. Therefore, the pixels in this area are used to estimate global atmospheric scattered light. The most reliable source of information.
[0085] This embodiment is in The dark channel prior theory is applied within the region. Calculation. All pixels within the area Dark passage : in Therefore A local window centered on the center. yes of (r,g,b) channels.
[0086] exist In the middle, select the first The brightest pixel. In the original image. In the dataset, find the original RGB values corresponding to the brightest and darkest channel pixels, and calculate their average value. Define this average value as the global atmospheric light A.
[0087] S5-2-2: Transmittance Estimation and Optimization: After obtaining A, a preliminary estimate of the transmittance is made. : in It has been estimated. of Channel components. It is a retention factor. A preliminary estimate. There is a noticeable blocky effect, and its edges do not match the true edges of the objects. On highways, the edge contour information of small road obstacles to be detected is extremely valuable. If it is smoothed during the dust removal process, it will become undetectable.
[0088] Therefore, this embodiment uses guided filtering. Optimize.
[0089] Using original high-definition image frames As a guide image. Although the resolution is low, the edge structure information is still present. It is the filter window radius. It is the regularization parameter. The guided filtering algorithm utilizes... Use edge structure information in the middle to guide A smooth process to ensure While being smooth, its edges and By keeping the edges of objects consistent, a refined transmittance map can be obtained. .
[0090] S5-2-3: Image restoration, after obtaining... and Then, the final enhanced image is calculated using the inverted atmospheric scattering model. ,Right now : in, It is a set lower limit for transmittance to prevent when When the value is extremely small, the denominator is zero, which avoids local overexposure of the image and excessive amplification of noise.
[0091] Spatial weight graph and static risk weight diagram These are all statistical results based on historical data. They reflect the inherent characteristics of unstructured road environments. However, the core safety hazard in highway maintenance lies in its unpredictability, such as an ice cream cone... They are constantly being scraped by vehicles, completely overturned, and encroaching on passable areas. Such dynamic anomalies cannot be predicted by historical maps.
[0092] When a new obstacle (such as a fallen temporary roadblock) suddenly appears, drivers of following vehicles are forced to take evasive action (i.e., detour). By detecting this collective, localized, geometrically anomalous detour behavior, a dynamic risk weight map can be generated in real time, marking the areas most likely to be dangerous.
[0093] S6-1: Construct a historical prototype trajectory baseline. To achieve anomaly detection in this step, a normal trajectory shape baseline that represents the driving consensus of the historical group of vehicles is needed. This baseline is constructed during the offline scene modeling stage.
[0094] Use the space already projected onto the bird's-eye view. Collection of historical trajectories A trajectory clustering algorithm is used to analyze the historical trajectory set. Divided into A cluster of trajectories. For Each of the clusters ( Extract a central orbit as the historical prototype trajectory of the cluster. Generate and store a collection of historical prototype trajectories. This set It represents the geometric shape of the normal driving path in the consensus of a group of vehicles under the current unstructured road conditions.
[0095] S6-2: Real-time short-time trajectory extraction and projection, during the online real-time detection phase, At any given moment, the system maintains a record containing past events. A high-definition video stream frame sequence over time, denoted as A multi-target tracking algorithm is used to... The system tracks motor vehicles in real time and generates a set of currently tracked, short-term trajectories. ,in This is the index of the currently tracked vehicle. These trajectories... The points on the top are projected onto the bird's-eye view space. The BEV point sequence is obtained. .
[0096] S6-3: Calculation of trajectory morphology deviation based on Fréchet distance, for Each real-time trajectory Calculate it with All historical prototype trajectories Discrete Fréchet distance between .set up It is by A sequence of points .set up It is by A sequence of points .
[0097] It is obtained through dynamic programming calculation. Let... for The former points and The former The Fréchet distance between any two points is recursively defined as follows: in It is Euclidean distance. Right now .
[0098] Each real-time trajectory final form deviation Take the distance from the most similar historical prototype: The value is a measure of shape, not position. If a real-time trajectory... It simply travels normally at the edge of the passable area (i.e., parallel to the prototype trajectory, but with some translation), and its shape is similar to the historical prototype trajectory of its lane. Highly similar, therefore its It will be very low. Conversely, if It is an abnormal detour trajectory generated to avoid obstacles; its shape is unlike any straight path. They are not similar, therefore their It will be very high. This method can effectively avoid misjudgments caused by normal translation within the lane.
[0099] S6-4: Generation of dynamic risk weight graph based on population anomalies, setting a morphological deviation threshold. .exist At that moment, All of the above satisfy trajectory Trajectories identified as transient anomalous morphological patterns are denoted as [the set of these trajectories]. This embodiment maintains a space related to the bird's-eye view. Anomaly accumulation graph of uniform size The image The value represents The number of abnormal trajectories accumulated by the point in recent times. The update process includes two steps: time decay and swarm activation.
[0100] Time decay, the previous moment Cumulative graph Multiplied by a time decay factor : This ensures that sporadic, old individual behavioral abnormalities will automatically decay and be cleared over time.
[0101] Group activation, generation Instantaneous activation graph at time This image is solely by [author's name / author's name]. Trajectory generation in: By analyzing the anomaly accumulation graph Thresholding is performed to obtain the final dynamic risk weight graph. .
[0102] This is the threshold for counting anomalies in the population. It is a Gaussian kernel function.
[0103] It is a high-risk, binarized map. It will only be used in areas within traversable zones and those that have been at least [exploded / blocked] in the near future. Only areas where vehicles have traversed in an abnormal manner will be activated. This means that interference from individual behavior is filtered out, and only newly emerging road obstacles verified by group consensus are responded to.
[0104] S7-1: Multimodal Small Target Detection Guided by Multiple Priors S7-1-1: Constructing a unified attention guidance graph The three independent BEV prior maps are merged into a unified attention guidance map. .
[0105] in: As the basic weight. Used to enhance The inherent risk sections within the area. Used to activate regions with sudden anomalies. and It is used to adjust the weights between static and dynamic risks.
[0106] To facilitate subsequent application to the image, After normalization, the final BEV boot graph is obtained. : The range was normalized to .
[0107] S7-1-2: Projection of BEV guide map to image space Create an image with enhancement With the same resolution Empty attention map Step traversal Each pixel on use calculate Corresponding BEV coordinates .exist Search The corresponding attention value. To avoid Falling Between grid points, bilinear interpolation is used for sampling to obtain accurate attention weights. . The sampled weights Attention map assigned to the image space. Output a value that is the same as... High-resolution, high-precision image spatial attention maps .
[0108] S7-1-3: Image Spatial Attention Enhancement Single channel Copy as three channels .
[0109] in Hadama product.
[0110] for A pixel with a very high value has its original pixel value. It will be magnified. And for... A pixel with a value of zero will retain its pixel value (multiplied by 0). This approach directly encodes the prior knowledge formed by the driving consensus and dynamic risks of a group of vehicles into... In terms of pixel intensity, the subsequent backbone network is forced to give the highest intensity feature extraction response to pixels in high-risk areas from the very beginning.
[0111] S7-1-4: Multimodal Feature Extraction This embodiment employs a two-stream neural network architecture. The image stream uses a convolutional neural network as its backbone to extract... Multiscale feature maps And generate BEV feature maps of the image. Laser point cloud flow uses PointPillars to sparsely... Convert to bird's-eye view space The pseudo-image feature map below is denoted as .
[0112] S7-1-5: Feature Fusion and Detection Head Will and Perform splicing along the channel dimension: Using YOLO-style BEV detection head pairs Perform regression and classification. Output a preliminary set of detection boxes. ,in It is the detected target index. It is the BEV space bounding box of the target. It is its category confidence level.
[0113] S7-2: Hazardous Target Identification and Output S7-2-1: Target State Analysis Based on LiDAR A tilted or even completely overturned traffic cone or water barrier carries a vastly different level of risk compared to an upright one. (Visual) LiDAR is easily affected by viewing angle and lighting when determining 3D pose, while LiDAR... It has a high-precision sensing capability for geometric height.
[0114] For each confidence level detected in S7-1-4 Above the threshold goal :extract All that fall within its 3D bounding box Internal point cloud subset .calculate All points of Maximum value of coordinates (height) : Set the normal upright height threshold for this category of targets. and the threshold of lodging height Define the goal. The physical state.
[0115] S7-2-2: Risk Level Determination Based on Semantic Location Obtain the target BEV center point coordinates Use the generated raw semantic mask for querying: IF THEN (The target has intruded into the passable area, posing a serious threat of traffic accidents.)
[0116] ELSE IF : IF THEN (The target is located in the isolation area but has collapsed and is very likely to be caught in the traffic flow.)
[0117] ELSE ( THEN (The target is located in the isolation area and is in normal condition; continuous monitoring is required.)
[0118] ELSE ( ): THEN (The target is located in a construction area and is not a traffic threat.)
[0119] S7-2-3: Structured alarm output. The method of this invention ultimately outputs a list of structured alarm information to the intelligent transportation system or the background monitoring center.
[0120] This embodiment also proposes a road target detection system that integrates multi-source perception and attention enhancement. The system includes: Spatial weight generation module: acquires historical high-definition video streams and, based on vehicle trajectory sequences, determines passable areas in the bird's-eye view space and generates a spatial weight map; Static risk weight generation module: Generates a static risk weight map based on the kinematic features of the trajectory sequence; Image enhancement module: acquires real-time image frames and 3D laser point cloud data, performs image enhancement on the real-time image frames, and obtains enhanced images; Dynamic risk weight generation module: tracks real-time short-term trajectories in real-time high-definition image frames, calculates the morphological deviation between the real-time short-term trajectory and the historical prototype trajectory set; and generates a dynamic risk weight map. Attention Enhancement Module: This module integrates spatial weight maps, static risk weight maps, and dynamic risk weight maps to construct a unified attention guidance map. The unified attention guidance map is then projected onto the image space to perform image space attention enhancement on the enhanced image. Target detection module: Performs multimodal small target detection based on attention-enhanced images and real-time 3D laser point cloud data.
[0121] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aforementioned road target detection method that integrates multi-source perception and attention enhancement.
[0122] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned road target detection method that integrates multi-source perception and attention enhancement.
[0123] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0124] In this specification, the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the descriptions of the embodiments described later are relatively simple, and relevant parts can be referred to the descriptions of the foregoing embodiments.
[0125] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A road target detection method integrating multi-source perception and attention enhancement, characterized in that, Object detection includes: Historical high-definition video streams are acquired, and based on vehicle trajectory sequences, passable areas in the bird's-eye view space are determined and a spatial weight map is generated. A static risk weight map is generated based on the kinematic characteristics of the trajectory sequences. Acquire real-time image frames and 3D laser point cloud data, perform image enhancement on the real-time image frames, and obtain the enhanced image; Track real-time short-term trajectories in real-time high-definition image frames, calculate the morphological deviation between the real-time short-term trajectory and the historical prototype trajectory set, and generate a dynamic risk weight map; A unified attention guidance graph is constructed by integrating spatial weight graphs, static risk weight graphs, and dynamic risk weight graphs; the unified attention guidance graph is then projected onto the image space to perform image spatial attention enhancement on the enhanced image. Multimodal small target detection is performed based on attention-enhanced images and real-time 3D laser point cloud data.
2. The road target detection method integrating multi-source perception and attention enhancement according to claim 1, characterized in that, Generating a spatial weighted graph includes: Based on the binary mask of the passable area Aerial view space Semantic region segmentation is performed to obtain the construction area mask. and isolation area mask ; and according to preset weight parameters , , Calculate the spatial weighted graph : in, The coordinates are for the bird's-eye view.
3. The road target detection method integrating multi-source perception and attention enhancement according to claim 1, characterized in that, Generating a static risk weight diagram includes: Calculate trajectory sequence Any point in the bird's-eye view space instantaneous rate at Instantaneous acceleration magnitude and instantaneous curvature ; Speed risk based on normalization Acceleration risk and curvature risk Calculate instantaneous risk score : in, It is a preset risk combination weight.
4. The road target detection method integrating multi-source perception and attention enhancement according to claim 3, characterized in that, Generating a static risk weight diagram also includes: Instantaneous risk score of all trajectory points Generate a risk accumulation chart ; Generate trajectory counting map ; The static risk weight map is obtained by calculating the average risk score. : in, For Gaussian kernel function, To prevent extremely small positive numbers with a denominator of zero, This represents the total number of trajectories.
5. The road target detection method integrating multi-source perception and attention enhancement according to claim 1, characterized in that, The morphological deviation calculation includes: for each real-time short-time trajectory Calculate its relationship with the set of historical prototype trajectories. Each historical prototype trajectory Discrete Fréchet distance between The minimum discrete Fréchet distance is taken as the real-time short-time trajectory. morphological deviation : in The number of clusters for the historical prototype trajectory.
6. The road target detection method integrating multi-source perception and attention enhancement according to claim 5, characterized in that, Generating a dynamic risk weighting diagram includes: incorporating morphological deviation. Greater than the morphological deviation threshold Real-time short-term trajectory identification is a set of instantaneous abnormal trajectory patterns. ; based on Update the anomaly accumulation graph using time decay and group activation. ; Thresholding is applied to the anomaly accumulation graph to obtain a dynamic risk weight graph. : in, A binary mask for the passable area. This is the threshold for counting anomalies in the population.
7. The road target detection method integrating multi-source perception and attention enhancement according to claim 1, characterized in that, Image enhancement includes: based on Using the construction area determined by the spatial weight map as a priori, and through the priori of the hidden passage... Estimate global atmospheric light ; through guided filtering Optimize transmittance By inverting the model Calculate the enhanced image ;in, For pixel coordinates, For the enhanced image to be solved, This is a preliminary estimate of the transmittance. To guide the image, For filter parameters, This is the lower limit of transmittance.
8. The road target detection method integrating multi-source perception and attention enhancement according to claim 1, characterized in that, Constructing a unified attention guidance graph include: ; And on Normalization yields a unified attention guidance map. ; Will Projecting onto the image space yields the image space attention map. : in, Spatial coordinates for the bird's-eye view. As weight, For the image after attention enhancement, For three-channel copying of the image spatial attention map, Hadama product.
9. The road target detection method integrating multi-source perception and attention enhancement according to claim 1, characterized in that, The method also includes: preliminary detection boxes obtained from multimodal small target detection. Extract a subset of its internal lidar point cloud. And calculate point cloud Maximum value of coordinates: ;based on Compared with the preset upright height threshold and the threshold of lodging height Determine the physical state of the target Based on physical state and the target center point The location within the semantic region determines the risk level. It will also output alarm information.
10. A road target detection system integrating multi-source perception and attention enhancement, characterized in that... The system includes: Spatial weight generation module: acquires historical high-definition video streams and, based on vehicle trajectory sequences, determines passable areas in the bird's-eye view space and generates a spatial weight map; Static risk weight generation module: Generates a static risk weight map based on the kinematic features of the trajectory sequence; Image enhancement module: acquires real-time image frames and 3D laser point cloud data, performs image enhancement on the real-time image frames, and obtains enhanced images; Dynamic risk weight generation module: tracks real-time short-term trajectories in real-time high-definition image frames, calculates the morphological deviation between the real-time short-term trajectory and the historical prototype trajectory set; and generates a dynamic risk weight map. Attention Enhancement Module: This module integrates spatial weight maps, static risk weight maps, and dynamic risk weight maps to construct a unified attention guidance map. The unified attention guidance map is then projected onto the image space to perform image space attention enhancement on the enhanced image. Target detection module: Performs multimodal small target detection based on attention-enhanced images and real-time 3D laser point cloud data.