Intersection detection early warning method and system
By fusing point cloud and visual data to identify and predict the trajectory of targets at intersections, and combining this with a risk analysis model to calculate potential collision times and generate warning commands, this approach solves the problems of imaging distortion and poor environmental adaptability in traditional intersection safety technologies, achieving high-precision driving safety warnings.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGQIU YUDONG HIGHWAY SURVEY & DESIGN CO LTD
- Filing Date
- 2026-01-29
- Publication Date
- 2026-05-01
AI Technical Summary
Traditional intersection safety technologies suffer from imaging distortion, poor environmental adaptability, and insufficient timely warnings and collaborative protection capabilities of existing electronic sensing systems in complex scenarios. As a result, the reliability of identifying and tracking targets such as pedestrians and non-motorized vehicles is not high, making it difficult to support accurate risk prediction.
The system uses point cloud and visual data fusion to identify the state of moving targets. It extracts precise spatial motion parameters of the targets from point cloud data, combines them with a visual recognition model to obtain refined attribute classifications of the targets, constructs the target's motion trajectory, and uses a risk analysis model to predict potential collision times, generating warning commands for two-way prompts.
It improves the accuracy and real-time performance of traffic safety warnings at intersections, reduces the risk of collisions caused by blind spots or misjudgments, and achieves comprehensive early warning of dynamic risks at intersections.
Smart Images

Figure CN121963532A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of road traffic monitoring and traffic safety, and in particular to a method and system for intersection detection and early warning. Background Technology
[0002] With rapid urbanization and a surge in traffic volume, at-grade intersections, as key conflict points in road networks, face severe challenges to their operational safety. Limited visibility due to buildings, greenery, or terrain obstructions is an inherent safety hazard at many intersections, seriously affecting the observation and decision-making of drivers and pedestrians. Traditionally, convex wide-angle mirrors have been used to expand the field of view, but they suffer from inherent defects such as image distortion leading to misjudgment of distance, easy contamination of the mirror surface, and failure in inclement weather, making it difficult to meet the requirements for all-weather, highly reliable safety warnings. Therefore, developing a new intersection safety technology that can overcome the limitations of traditional technologies and achieve accurate perception and real-time warnings has become an urgent task for improving road traffic safety.
[0003] Existing intersection safety technologies are mainly divided into two categories: passive reflection and active sensing. The former, represented by convex mirrors, relies on optical reflection to indirectly expand the field of view, but it has inherent problems such as image distortion and poor environmental adaptability; the latter gradually adopts electronic sensing methods such as video surveillance and millimeter-wave radar to achieve direct detection of traffic participants.
[0004] However, each single-sensor solution has its shortcomings: vision is susceptible to lighting and weather conditions, and lacks stability; radar has limited capabilities in target classification and fine-grained attitude perception. This results in low reliability in the identification and tracking of targets such as pedestrians and non-motorized vehicles, making it difficult to support accurate risk prediction. Summary of the Invention
[0005] In view of this, this application proposes an intersection detection and early warning method and system, which improves the accuracy and real-time performance of intersection traffic safety early warning, and solves the problems of imaging distortion, poor environmental adaptability, and untimely early warning and insufficient collaborative protection capability of existing electronic sensing systems in complex scenarios in traditional intersection safety solutions.
[0006] This application provides a method and system for detecting and warning at intersections, which adopts the following technical solution: A method for detecting and warning at intersections, comprising: When at least one moving target object is detected at the intersection, the point cloud data and initial image of the moving target object are acquired respectively. The target state information is obtained by recognizing and analyzing the initial image using point cloud data; The point cloud data and target state information are fused to obtain the trajectory of the target object; A pre-defined risk analysis model is used to predict the trajectory of the target object, thereby obtaining the future movement path of the moving target object; The minimum collision time between any two moving target objects is calculated based on the future movement path. If the minimum collision time is less than the preset warning threshold, a warning instruction is generated. The warning command is sent to the roadside display device and / or vehicle terminal to provide two-way warning prompts.
[0007] By adopting the above technical solution, the real-time status of moving targets is identified by fusing point cloud and visual data, the target's motion trajectory is constructed by combining trajectory fusion analysis, the potential collision time is calculated by the risk prediction model, and finally collaborative safety protection is achieved by issuing two-way warning commands. This improves the accuracy and real-time performance of traffic safety warnings at intersections, thereby reducing the risk of collision accidents caused by blind spots or misjudgments. It achieves comprehensive early warning of dynamic risks at intersections in the dimensions of perception, prediction and response.
[0008] Preferably, target state information is obtained by identifying and analyzing the initial image using point cloud data, including: Based on point cloud data, the spatial position and dynamic parameters of the moving target object are extracted to obtain the first motion feature information; Based on the initial image, a preset visual recognition model is used to extract and classify features to obtain target attribute classification information; The target state information is obtained by integrating the first motion feature information and the target attribute classification information.
[0009] By adopting the above technical solution, the precise spatial motion parameters of the target are extracted based on point cloud data, and the refined attribute classification of the target is obtained by combining the visual recognition model. Then, the three-dimensional motion features and two-dimensional visual features are integrated across modal information, so as to realize the complementary state representation of the moving target in three dimensions: spatial position, motion state and attribute category.
[0010] Preferably, the first motion feature information and the target attribute classification information are integrated to obtain target state information, including: Extract the real-time position coordinates of the moving target object from the first motion feature information to obtain the first spatial positioning data; Extract the target position of the moving target object from the initial image to obtain the second spatial positioning data; Spatial matching is performed on the first spatial positioning data and the second spatial positioning data to obtain target matching pairs; Based on the target matching pair, the first motion feature information is associated and matched with the target attribute classification information to obtain the target state information.
[0011] By adopting the above technical solution, the spatial positioning information of the target in point cloud and image data is extracted and matched to establish accurate cross-modal target association. Then, the successfully matched motion features and attribute classification information are integrated accordingly, realizing accurate alignment and deep fusion of multi-source heterogeneous perception data at the target level. This constructs target state information with both accurate spatial motion parameters and rich semantic attributes, improving the accuracy and robustness of target state perception in complex traffic environments.
[0012] Preferably, point cloud data and target state information are fused to obtain the trajectory of the target object, including: Extract the motion sequence of the moving target object in the point cloud data in multiple frames to obtain the original motion trajectory data; The original motion trajectory data and target state information are spatiotemporally correlated and state corrected to obtain a motion state sequence; The motion state sequence is smoothed and interpolated to generate the trajectory of the target object.
[0013] By adopting the above technical solution, the original motion sequence of the target is extracted based on multi-frame point cloud data, and the target state information is combined to perform spatiotemporal correlation and state correction. Then, the continuity and smoothness of the trajectory are optimized by filtering and interpolation, so as to realize the accurate reconstruction of the target motion state in the spatiotemporal dimension and generate a high-precision target motion trajectory.
[0014] Preferably, the motion state sequence is smoothed and interpolated to generate the trajectory of the target object, including: Obtain lane topology information and drivable area at the intersection; By using lane topology information, a preliminary smoothed trajectory is obtained by performing trajectory smoothing filtering on the motion state sequence. The initial smoothed trajectory is constrained within the drivable area for trajectory interpolation and boundary correction to generate the trajectory of the target object.
[0015] By adopting the above technical solution, the trajectory is smoothed and filtered by combining the prior lane topology of the intersection, which effectively suppresses the trajectory jitter caused by perceived noise. Then, the trajectory is interpolated and the boundary is corrected by the physical constraints of the drivable area, which ensures that the generated target trajectory conforms to the geometric and rule constraints of the actual road. This generates a trajectory that is both smooth and conforms to the physical laws of the traffic scene, increasing the rationality and accuracy of trajectory prediction and risk analysis.
[0016] Preferably, the pre-defined risk analysis model includes a trajectory feature encoding layer, a spatiotemporal attention mechanism layer, and a conflict probability decoding layer; A pre-defined risk analysis model is used to predict the trajectory of the target object, resulting in the future movement path of the moving target object, including: The trajectory feature encoding layer extracts features from the trajectory of the target object to generate a trajectory feature vector; The spatial-temporal attention mechanism layer is used to calculate the time-dimensional dependency of trajectory feature vectors to obtain a contextual trajectory representation. By using a conflict probability decoding layer to predict the context trajectory representation, multiple predicted location sequences and their existence probabilities are obtained. The predicted position sequence with the highest probability of existence will be used as the future movement path of the moving target object.
[0017] By adopting the above technical solution, a trajectory feature encoding layer is used to extract multi-dimensional spatiotemporal features of target motion. A spatiotemporal attention mechanism is combined to capture long-distance temporal dependencies and interactive effects. Then, a conflict probability decoding layer is used to generate multimodal prediction paths and their probability distributions. Finally, the path with the highest probability is selected as the final prediction result, thereby achieving a deep understanding of the target's future movement behavior and improving the accuracy of future movement path prediction.
[0018] Preferably, the minimum collision time between any two moving target objects is calculated based on the future movement path. If the minimum collision time is less than a preset warning threshold, a warning instruction is generated, including: By using any two future paths, the spatiotemporal domain occupied by the moving target object within the preset prediction period can be determined. Calculate the time period during which two spatiotemporal occupied domains first overlap in the spatiotemporal coordinate system or when the distance between them is less than a safety threshold. The difference between the start time of the overlapping period and the current time is determined as the minimum collision time. If the minimum collision time is less than the preset warning threshold, a warning command will be generated.
[0019] By adopting the above technical solution, the dynamic occupancy domain of the target in the spatiotemporal dimension is constructed by predicting the path, and the critical moment when the spatiotemporal occupancy domains of multiple targets overlap or approach each other is calculated. This transforms the abstract path prediction into a quantitative collision time indicator, thereby realizing the early quantitative assessment of potential collision risks.
[0020] An intersection detection and early warning system includes: The data acquisition module is used to acquire point cloud data and initial images of the moving target objects when at least one moving target object is detected at the intersection. The state recognition module is used to identify and analyze the initial image using point cloud data to obtain target state information; The trajectory fusion module is used to fuse point cloud data with target state information to obtain the trajectory of the target object; The risk prediction module is used to predict the trajectory of the target object using a preset risk analysis model, and obtain the future movement path of the moving target object. The collision analysis module is used to calculate the minimum collision time between any two moving target objects based on the future movement path. If the minimum collision time is less than the preset warning threshold, a warning instruction is generated. The warning execution module is used to send warning commands to roadside display devices and / or vehicle terminals for two-way warning prompts.
[0021] By adopting the above technical solution, the data acquisition module realizes the synchronous acquisition of multi-source heterogeneous sensing data, the state recognition module completes the accurate analysis of target state across modalities, the trajectory fusion module constructs a highly reliable continuous target trajectory, the risk prediction module proactively infers the future behavior of the target, the collision analysis module quantifies the assessment of potential collision risks, and the early warning execution module realizes the collaborative release of early warning information, thereby improving the accuracy and real-time performance of intersection traffic safety early warning, thus reducing the risk of collision accidents caused by blind spots or prediction errors, and realizing comprehensive early warning of dynamic risks at intersections in the dimensions of perception, prediction, and response.
[0022] An electronic device includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps of an intersection detection and early warning method.
[0023] A computer-readable storage medium having a computer program stored thereon, which, when executed, implements an intersection detection and early warning method.
[0024] In summary, this application includes at least one of the following beneficial technical effects: 1. When at least one moving target object is detected at the intersection, point cloud data and an initial image of the moving target object are acquired. The initial image is then analyzed using the point cloud data to obtain target state information. The point cloud data and target state information are fused to obtain the target object's trajectory. A preset risk analysis model is used to predict the target object's trajectory, resulting in the future movement path of the moving target object. The minimum collision time between any two moving target objects is calculated based on the future movement path. If the minimum collision time is less than a preset warning threshold, a warning command is generated and sent to the roadside display device and / or vehicle terminal for two-way warning. By fusing point cloud and visual data to accurately perceive the target state and predict its trajectory, and calculating the potential collision time in real time to generate warning commands, the accuracy and real-time performance of intersection traffic safety warnings are improved. This reduces the risk of collision accidents caused by blind spots or prediction errors, and solves the problems of imaging distortion, poor environmental adaptability, and untimely warnings and insufficient collaborative protection capabilities of existing electronic perception systems in complex scenarios in traditional intersection safety solutions.
[0025] 2. By extracting precise spatial motion parameters of the target through point cloud data, and combining them with a visual recognition model to obtain refined attribute classification of the target, the three-dimensional motion features and two-dimensional visual features are then integrated across modal information, achieving complementary state representation of the moving target in three dimensions: spatial position, motion state, and attribute category. This overcomes the problem of limited perception capability of a single sensor under complex lighting, weather, or occlusion conditions. 3. By constructing a risk analysis model that includes a trajectory feature encoding layer, a spatiotemporal attention mechanism layer, and a conflict probability decoding layer, deep feature extraction and multimodal path prediction are performed on the trajectory of the target object. Combined with lane topology and drivable area constraints, a reasonable and smooth future movement path is generated, which improves the accuracy of trajectory prediction and scene adaptability. Attached Figure Description
[0026] Figure 1 This is a flowchart of the steps of an intersection detection and early warning method provided in Embodiment 1 of this application; Figure 2 This is a structural block diagram of an intersection detection and early warning system provided in Embodiment 3 of this application; Figure 3 This is a structural block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. It should be noted that in the optional embodiments of this application, the object information and other related data involved require the permission or consent of the object when the embodiments of this application are applied to specific products or technologies, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. That is to say, if the embodiments of this application involve data related to the object, it needs to be obtained with the authorization and consent of the object, the authorization and consent of the relevant departments, and in compliance with the relevant laws, regulations, and standards of the country and region. If personal information is involved in the embodiments, the acquisition of all personal information requires the consent of the individual. If sensitive information is involved, the separate consent of the information subject is required, and the embodiments also need to be implemented with the authorization and consent of the object. Example
[0028] Please see Figure 1 This application provides a method for intersection detection and early warning, comprising: Step 101: When at least one moving target object is detected at the intersection, acquire the point cloud data and initial image of the moving target object respectively.
[0029] A moving target object refers to a traffic participant in a continuous state of motion within the passable area of an intersection. Specifically, it includes motor vehicles, non-motor vehicles, pedestrians, and other moving entities that have legally entered the right-of-way. Preferably, the passable area of an intersection includes motor vehicle lanes, non-motor vehicle lanes, and pedestrian crossings.
[0030] Point cloud data refers to environmental spatial information acquired through lidar scanning and represented as a set of three-dimensional coordinate points. Preferably, point cloud data includes, but is not limited to, geometric and physical attributes such as the target's position, shape, and reflection intensity.
[0031] An initial image refers to a raw digital image or video frame that has not undergone depth processing and is synchronously acquired by a visual sensor deployed at the intersection. Preferably, the initial image includes, but is not limited to, information such as the texture, color, and semantic appearance of the target.
[0032] In this embodiment, the entire intersection area is monitored in real time by simultaneously deploying LiDAR and high-definition cameras on the roadside. When any moving target is detected entering the monitoring area, the LiDAR and camera are triggered simultaneously to collect the point cloud data and initial image of the target at the same time, and assign them a unified timestamp and spatial reference identifier to ensure the alignment of multi-source data in the spatiotemporal dimensions.
[0033] Step 102: Use point cloud data to identify and analyze the initial image to obtain target state information.
[0034] Target state information refers to the comprehensive state description of a moving target object generated by fusing the precise three-dimensional geometric features provided by point cloud data with the rich semantic features provided by the initial image. This includes, but is not limited to, the target's precise three-dimensional position, speed, heading angle, physical size, target category (moving target object), and confidence level.
[0035] In the embodiments of this application, the recognition process based on the initial image is constrained, corrected and enhanced by utilizing the precise three-dimensional spatial structure information represented by point cloud data, so as to overcome the limitations of pure vision methods in depth estimation, occlusion processing and extreme lighting conditions.
[0036] Preferably, step 102 includes the following sub-steps: S21. Based on point cloud data, extract the spatial position and dynamic parameters of the moving target object to obtain the first motion feature information.
[0037] Spatial location refers to the three-dimensional coordinates of a moving target object in the world coordinate system or a preset fixed reference coordinate system. It is characterized by the coordinate values of the target's centroid or three-dimensional geometric center and is used to describe the target's absolute or relative geographical location in an intersection environment.
[0038] Dynamic parameters refer to physical quantities that describe the change of the motion state of a moving target object over time. They include at least instantaneous velocity, instantaneous acceleration, and direction of motion (heading angle), and are used to characterize the instantaneous motion trend and dynamic characteristics of the target.
[0039] Preferably, the first motion feature information is a set of kinematic parameters for three-dimensional physical space measurement, including but not limited to the target's three-dimensional position coordinates, instantaneous velocity vector, acceleration, heading angle, and three-dimensional bounding box size in the world coordinate system.
[0040] In this embodiment, by clustering and segmenting point cloud data from multiple consecutive frames and tracking the target, each independent moving target point cloud cluster is identified. Then, based on the centroid displacement of the same target point cloud cluster between adjacent frames, its instantaneous velocity and acceleration are calculated. Its heading angle is determined by analyzing the principal direction of the point cloud cluster. The minimum bounding box of the point cloud cluster is fitted by combining the three-dimensional distribution of the point cloud cluster to obtain its size information. Finally, the first motion feature information containing accurate spatial position and complete dynamic parameters is generated, realizing the real-time quantitative representation of the motion state of the moving target in three-dimensional space.
[0041] S22. Based on the initial image, a preset visual recognition model is used to extract and classify features to obtain target attribute classification information.
[0042] The pre-set visual recognition model refers to a deep learning network that has been pre-trained and optimized on a large dataset of traffic scene images.
[0043] Preferably, the preset visual recognition model is a conventional image recognition model, which can be built based on a convolutional neural network (CNN) or Transformer architecture. It has the ability to perform end-to-end feature extraction, object detection and fine classification of input images, and can identify and distinguish different categories of traffic participants. No specific limitation is made here.
[0044] Target attribute classification information refers to the semantic description of each detected moving target object in the image after the initial image is analyzed by a visual recognition model. It includes at least the category label of the moving target object, the two-dimensional bounding box position in the image coordinate system, and the classification confidence.
[0045] In this embodiment, the acquired initial image is input into a preset visual recognition model. The model extracts deep semantic features of the image through its deep convolutional layers, and the detection head outputs bounding box proposals and their category probability distributions for all potential target objects in the image. By performing post-processing such as non-maximum suppression on the bounding boxes, the final category of each target, its precise pixel position in the image, and its corresponding confidence score are finally determined, thereby obtaining target attribute classification information rich in semantic information. This enables rapid semantic understanding of various moving targets in intersection scenes and improves the category recognition accuracy for small-scale or shape-variable targets such as pedestrians and non-motorized vehicles.
[0046] S23. Integrate the first motion feature information and the target attribute classification information to obtain the target state information.
[0047] Preferably, step S23 includes the following steps: S231. Extract the real-time position coordinates of the moving target object from the first motion feature information to obtain the first spatial positioning data.
[0048] Real-time position coordinates refer to the position data calculated in real time based on point cloud data through a tracking algorithm. They are used to characterize the three-dimensional spatial position of a moving target object in the world coordinate system or a preset fixed reference system at the current moment. In this embodiment, the real-time position coordinates are represented by the three-dimensional coordinate values of the geometric center.
[0049] In this embodiment of the application, the three-dimensional centroid coordinates of each target at the current moment are directly extracted from the first motion feature information. Preferably, the coordinates have been aligned and filtered by timestamps to form first spatial positioning data that can be directly used for spatial calculation and association.
[0050] S232. Extract the target position of the moving target object in the initial image to obtain the second spatial positioning data.
[0051] The target position refers to the position information of the moving target object in the two-dimensional plane of the image after the initial image is analyzed and processed by a preset visual recognition model, and is represented by the bounding box coordinates in the pixel coordinate system.
[0052] In this embodiment, based on the target attribute classification information, the two-dimensional bounding box information corresponding to each identified target is extracted. Then, using the pre-calibrated camera parameters, the pixel coordinates of the midpoint of the bottom edge of the bounding box are converted into three-dimensional ground coordinates in the world coordinate system through inverse perspective mapping, thereby generating second spatial positioning data derived from visual perception.
[0053] S233. Spatial matching is performed on the first spatial positioning data and the second spatial positioning data to obtain a target matching pair.
[0054] In this embodiment, a distance-based association matching algorithm is adopted. The first spatial positioning dataset and the second spatial positioning dataset are used as input. The Euclidean distance between each pair of coordinate points in the two datasets in a unified world coordinate system is calculated and a distance matrix is constructed. Then, a nearest neighbor greedy matching strategy is applied. Under the premise of satisfying the preset maximum matching distance threshold, the second spatial positioning data that is spatially closest to each first spatial positioning data originating from the point cloud is found, thereby forming a one-to-one target matching pair and completing the spatial association between the lidar target and the visual target.
[0055] S234. Based on the target matching pair, the first motion feature information is associated and matched with the target attribute classification information to obtain the target state information.
[0056] In the embodiments of this application, for each matching pair, the first motion feature information and the target attribute classification information extracted from the initial image are associated and fused to generate a structured target state information record, thereby realizing a multi-dimensional and high-precision state description of the same moving target object.
[0057] Step 103: Fuse the point cloud data with the target state information to obtain the trajectory of the target object.
[0058] In this embodiment, the continuous frame high-precision position information provided by point cloud data is used as the basis, and the semantic and dynamic attributes in the target state information are integrated. Through multi-source data complementarity and state correction, a smooth, continuous and semantically rich target historical motion trajectory is constructed.
[0059] Preferably, step 103 includes the following sub-steps: S31. Extract the motion sequence of the moving target object in the point cloud data in multiple frames to obtain the original motion trajectory data.
[0060] A motion sequence refers to a set of motion states of the same moving target object arranged in chronological order at multiple consecutive acquisition times, wherein the state at each time moment contains at least the three-dimensional spatial coordinates of the target at that time moment.
[0061] In this embodiment of the application, based on the target ID that has been tracked in the point cloud data, the three-dimensional position coordinates, velocity and heading angle information of the target in each of the past N consecutive frames (e.g. the most recent 30 frames, corresponding to a historical period of about 1 second) are extracted, and the set of state points (motion sequence) arranged in time order is used as the original motion trajectory data describing the recent motion of the target.
[0062] S32. Spatiotemporally correlate and correct the original motion trajectory data with the target state information to obtain the motion state sequence.
[0063] In this embodiment, the original motion trajectory data from the point cloud is aligned and fused with the target state information based on the target ID and timestamp. Specifically, the corresponding time position in the original trajectory is corrected using the more accurate real-time position in the target state information, and semantic attributes such as category and size in the target state information are fused to generate a spatiotemporally aligned and state-complete motion state sequence, achieving semantic enhancement and state optimization of historical trajectory information. Each state point in this sequence contains comprehensive information such as corrected position, velocity, heading, and target category.
[0064] S33. Perform smoothing filtering and trajectory interpolation on the motion state sequence to generate the trajectory of the target object.
[0065] Preferably, step S33 includes the following steps: S331. Obtain lane topology information and drivable area of the intersection.
[0066] Lane topology information refers to structured data on each lane within an intersection and their interconnections, including at least the lane's geometric centerline, lane width, lane type, lane connectivity rules, and the location information of traffic control signs.
[0067] The lane types include straight, left turn, and right turn; Traffic control signs include stop lines and yield signs.
[0068] The drivable area refers to the set of physical areas within an intersection that all traffic participants are permitted to pass through according to traffic rules. Specifically, it includes motor vehicle lanes, non-motor vehicle lanes, pedestrian crossings, and areas that can be used for sharing or borrowing lanes in accordance with regulations.
[0069] In this embodiment of the application, by reading a pre-stored high-precision digital map of the intersection or communicating with roadside intelligent facilities, complete structured description information including lane geometry, connectivity, traffic rules, and the boundaries of passable areas can be obtained in real time.
[0070] S332. The motion state sequence is smoothed by using lane topology information to obtain a preliminary smooth trajectory.
[0071] In this embodiment, the position points in the motion state sequence are projected onto the lane centerline reference system corresponding to the lane topology information. A state space model is constructed using lane geometric constraints, and a Kalman filter algorithm based on lane constraints is used to smooth the sequence, suppressing abnormal fluctuations perpendicular to the lane direction. Through the strong correlation between trajectory data and road physical structure, the rationality and accuracy of trajectory smoothing are improved, making the generated preliminary smooth trajectory more consistent with the actual driving behavior of the vehicle.
[0072] For example, when a vehicle travels along a curved lane, the original perception data may contain a few jumps perpendicular to the lane direction due to noise, such as coordinates shifting to adjacent lanes or outside the road. By filtering the position points within the lane geometry, these unreasonable lateral jumps can be effectively eliminated, and the smoothed trajectory can closely follow the direction of the lane centerline.
[0073] S333. Constrain the initial smooth trajectory within the drivable area, perform trajectory interpolation and boundary correction, and generate the trajectory of the target object.
[0074] In this embodiment, the geometric boundary of the drivable area is used as a hard constraint to adaptively adjust the initial smooth trajectory. First, in the missing segments of the trajectory time series, interpolation is performed to fill in the missing points based on the connectivity of the drivable area. Then, the positions of the trajectory points are checked and adjusted to ensure that they are always within the boundary of the drivable area corresponding to the target type. This ensures the physical rationality of the trajectory and avoids the problem of trajectory deviation into non-drivable areas (such as green belts and sidewalks) caused by perception errors or prediction deviations. As a result, a high-quality target object trajectory that is continuous in time and strictly conforms to road traffic rules is generated.
[0075] Step 104: Use a preset risk analysis model to predict the trajectory of the target object and obtain the future movement path of the moving target object.
[0076] In this embodiment of the application, the trajectory of the target object is input into a preset risk analysis model. This model is built on a deep learning architecture. Through deep feature learning of historical trajectories and spatiotemporal context analysis, it predicts the possible movement path of the target and its probability distribution in the future, thereby achieving a forward-looking risk situation assessment and realizing in-depth mining and high-precision prediction of the target's movement intention in complex traffic scenarios.
[0077] Preferably, the preset risk analysis model includes a trajectory feature encoding layer, a spatiotemporal attention mechanism layer, and a conflict probability decoding layer.
[0078] Preferably, step 104 includes the following sub-steps: S41. The trajectory of the target object is extracted through the trajectory feature encoding layer to generate a trajectory feature vector.
[0079] In this embodiment, the trajectory feature encoding layer is composed of a temporal convolutional network, which performs nonlinear transformation and deep feature extraction on the trajectory of the target object, and encodes the variable-length trajectory sequence into a fixed-dimensional trajectory feature vector rich in motion patterns and intent information. This achieves automated and efficient extraction of trajectory temporal features and enhances the model's ability to capture nonlinear motion patterns (such as acceleration and turning).
[0080] S42. The spatial-temporal attention mechanism layer is used to calculate the dependence of trajectory feature vectors in the time dimension to obtain the context trajectory representation.
[0081] In this embodiment, the spatiotemporal attention mechanism layer receives the trajectory feature vector sequence and calculates the correlation weights between features at different time steps through a multi-head self-attention mechanism. This enables the model to adaptively focus on the most critical moments in the historical trajectory for predicting future motion, such as the starting points of sudden braking and turning. By fusing this key information, a contextual trajectory representation containing long-term time dependence and motion causal relationships is generated. This achieves a refined understanding of historical motion patterns and strengthens key information, improving the model's ability to analyze and predict trajectories with complex temporal variation characteristics. Consequently, it improves the accuracy and real-time performance of intersection traffic safety warnings.
[0082] S43. Predict the context trajectory representation through the conflict probability decoding layer to obtain multiple predicted position sequences and their existence probabilities.
[0083] In this embodiment, the conflict probability decoding layer employs a variational autoencoder to map the context trajectory representation to the state space of future time steps. This layer uses a feedforward neural network or a temporal deconvolutional network as its core to output multiple possible future position sequences and their corresponding existence probabilities, i.e., the confidence of each predicted path. This quantifies the uncertainty of the target's future movement, realizes multimodal probabilistic modeling of future motion possibilities, and improves the system's coverage and prediction robustness for various potential behavioral intentions such as lane changing and turning in complex interaction scenarios.
[0084] S44. Use the predicted position sequence with the highest probability as the future movement path of the moving target object.
[0085] In this embodiment, from the multiple predicted position sequences output by the collision probability decoding layer, the sequence with the highest probability of existence is selected as the most likely future movement path of the current moving target object. The spatiotemporal point sequence of this path is used as the input for subsequent collision time calculation. By transforming the probabilistic prediction result into a deterministic decision output, while taking into account the uncertainty of prediction, a clear and executable optimal path estimate is provided for the real-time early warning system, ensuring the efficiency of risk analysis and the clarity of decision-making basis.
[0086] Step 105: Calculate the minimum collision time between any two moving target objects based on the future movement path. If the minimum collision time is less than the preset warning threshold, generate a warning command.
[0087] Preferably, step 105 includes the following sub-steps: S51. Using any two future paths, determine the spatiotemporal occupancy domain of the moving target object within a preset prediction period.
[0088] In the embodiments of this application, for each target, based on the position of each predicted point on its future movement path and its corresponding target physical size (such as bounding box), a continuous spatiotemporal volume is constructed along the time axis within a preset future time interval. The spatiotemporal volume represents the range of the three-dimensional spatial region occupied by the target during its future movement, i.e., the spatiotemporal occupancy domain of the target, which improves the accuracy of the description of the actual occupancy range of the target.
[0089] S52. Calculate the time period when two spatiotemporal occupied domains first overlap in the spatiotemporal coordinate system or when the distance is less than the safety threshold.
[0090] In this embodiment of the application, the spatiotemporal occupancy domains of two targets are placed in a unified spatiotemporal coordinate system for Boolean operation calculation, and the time interval in which the two first intersect (overlap) in the time and space dimensions or the surface distance is first less than a preset safety threshold is detected, so as to achieve accurate spatiotemporal positioning of potential collision risk points.
[0091] S53. The difference between the start time of the overlapping period and the current time is determined as the minimum collision time.
[0092] In this embodiment of the application, when there is an overlapping time period, the starting time point of the time period is extracted, that is, the time when the two targets first enter the risk state in space and time, the difference between the starting time point and the current system time is calculated, and this difference is defined as the minimum collision time between the target pair.
[0093] S54. If the minimum collision time is less than the preset warning threshold, a warning command is generated.
[0094] In this embodiment, the calculated minimum collision time is compared with a preset warning threshold. If the minimum collision time is less than the threshold, a high collision risk is determined, and a warning instruction containing the risk target, risk type, and suggested countermeasures is generated to improve the timeliness and accuracy of the warning response.
[0095] Step 106: Send the warning command to the roadside display device and / or vehicle terminal to provide two-way warning prompts.
[0096] In this embodiment, after a warning command is generated, the system transmits it in parallel to the LED display device on the roadside of the intersection and the vehicle-mounted terminal deployed in the vehicle via wired or wireless communication. The roadside display device displays targeted graphic and text warning information in real time to alert all traffic participants within the intersection area; the vehicle-mounted terminal receives the warning command through a vehicle-to-infrastructure (V2I) communication interface and provides directional warning prompts to the driver via sound, light, or voice through the in-vehicle human-machine interface, achieving synchronous two-way delivery of warning information in both the roadside environment and inside the vehicle. By constructing a three-dimensional collaborative warning network, the warning coverage is expanded and the warning intensity is enhanced, thereby improving the safety protection capability of the intersection.
[0097] The implementation principle of this application embodiment is as follows: By simultaneously acquiring point cloud data and initial images of moving targets using roadside LiDAR and visual sensors deployed at intersections, the precise three-dimensional geometric information provided by the point cloud is used to correct and enhance the visual recognition results, thereby obtaining high-precision target state information. Then, the point cloud time series data and target state information are fused to construct a continuous, smooth target trajectory that conforms to lane constraints. This trajectory is input into a deep learning-based risk analysis model to predict the target's future multimodal movement path and calculate the minimum collision time, achieving a quantitative assessment and forward warning of potential collision risks. Finally, the warning command is synchronously sent to the roadside display device and the vehicle terminal through vehicle-road cooperative communication, forming a full-link monitoring and warning system of "perception-prediction-warning", realizing all-weather three-dimensional active protection against traffic safety risks at intersections. Example
[0098] Please see Figure 2 This application provides an intersection detection and early warning system, comprising: The data acquisition module 201 is used to acquire point cloud data and initial image of the moving target object when at least one moving target object is detected at the intersection.
[0099] The state recognition module 202 is used to identify and analyze the initial image through point cloud data to obtain target state information.
[0100] The trajectory fusion module 203 is used to fuse point cloud data with target state information to obtain the trajectory of the target object.
[0101] The risk prediction module 204 is used to predict the trajectory of the target object using a preset risk analysis model, so as to obtain the future movement path of the moving target object.
[0102] The collision analysis module 205 is used to calculate the minimum collision time between any two moving target objects based on the future movement path. If the minimum collision time is less than the preset warning threshold, a warning command is generated.
[0103] The warning execution module 206 is used to send warning commands to the roadside display device and / or vehicle terminal for two-way warning prompts.
[0104] Since the above is a system corresponding to one intersection detection and early warning method, and its implementation principle is the same as that of another intersection detection and early warning method, for the sake of convenience and brevity, those skilled in the art can clearly understand that the specific working process of the system and modules described above can be referred to the corresponding process in the aforementioned method embodiments, and will not be repeated here. Example
[0105] An electronic device according to an embodiment of the present invention includes: a memory and a processor, wherein the memory stores a computer program; when the computer program is executed by the processor, the processor performs the intersection detection and early warning method as described in any of the above embodiments.
[0106] The memory can be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. The memory has storage space for program code used to perform any of the method steps described above. For example, the storage space for program code may include individual program codes for implementing the various steps in the methods described above. This program code can be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, compact discs (CDs), memory cards, or floppy disks. The program code may be compressed, for example, in a suitable form. When run by a computing processing device, this code causes the computing processing device to perform the various steps in the methods described above. Example
[0107] This invention provides a computer-readable storage medium storing a computer program thereon, which, when executed, implements the intersection detection and early warning method as described in any embodiment of this invention.
[0108] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0109] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units through some interfaces, and may be electrical, mechanical, or other forms.
[0110] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0111] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0112] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0113] It should be noted that the terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The implementations described in the following exemplary embodiments do not represent all implementations consistent with this disclosure.
[0114] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.
[0115] The embodiments described in this specific implementation are preferred embodiments of this application and are not intended to limit the scope of protection of this application. Identical components are represented by the same reference numerals. Therefore, all equivalent changes made to the structure, shape, and principle of this application should be covered within the scope of protection of this application.
Claims
1. A method for detecting and warning at intersections, characterized in that, include: When at least one moving target object is detected at the intersection, the point cloud data and initial image of the moving target object are acquired respectively. The target state information is obtained by identifying and analyzing the initial image using the point cloud data; The point cloud data and the target state information are fused to obtain the trajectory of the target object; A preset risk analysis model is used to predict the trajectory of the target object, thereby obtaining the future movement path of the moving target object; The minimum collision time between any two moving target objects is calculated using the future movement path. If the minimum collision time is less than a preset warning threshold, a warning instruction is generated. The warning command is sent to the roadside display device and / or vehicle terminal for two-way warning prompts.
2. The intersection detection and early warning method according to claim 1, characterized in that, The step of identifying and analyzing the initial image using the point cloud data to obtain target state information includes: Based on the point cloud data, the spatial position and dynamic parameters of the moving target object are extracted to obtain the first motion feature information; Based on the initial image, a preset visual recognition model is used to extract and classify features to obtain target attribute classification information; The first motion feature information and the target attribute classification information are integrated to obtain the target state information.
3. The intersection detection and early warning method according to claim 2, characterized in that, The step of integrating the first motion feature information and the target attribute classification information to obtain the target state information includes: Extract the real-time position coordinates of the moving target object from the first motion feature information to obtain the first spatial positioning data; Extract the target position of the moving target object from the initial image to obtain the second spatial positioning data; Spatial matching is performed between the first spatial positioning data and the second spatial positioning data to obtain a target matching pair; Based on the target matching pair, the first motion feature information is associated and matched with the target attribute classification information to obtain the target state information.
4. The intersection detection and early warning method according to claim 1, characterized in that, The process of fusing the point cloud data with the target state information to obtain the target object trajectory includes: Extract the motion sequence of the moving target object in the point cloud data in multiple frames of data to obtain the original motion trajectory data; The original motion trajectory data and the target state information are spatiotemporally correlated and state corrected to obtain a motion state sequence; The motion state sequence is smoothed and interpolated to generate the trajectory of the target object.
5. The intersection detection and early warning method according to claim 4, characterized in that, The step of smoothing and interpolating the motion state sequence to generate the target object trajectory includes: Obtain the lane topology information and drivable area of the intersection; The motion state sequence is smoothed using the lane topology information to obtain a preliminary smoothed trajectory. The initial smoothed trajectory is constrained within the drivable area for trajectory interpolation and boundary correction to generate the target object trajectory.
6. The intersection detection and early warning method according to claim 1, characterized in that, The preset risk analysis model includes a trajectory feature encoding layer, a spatiotemporal attention mechanism layer, and a conflict probability decoding layer; The trajectory of the target object is predicted using a preset risk analysis model to obtain the future movement path of the moving target object, including: The trajectory feature encoding layer is used to extract features from the trajectory of the target object to generate a trajectory feature vector; The spatiotemporal attention mechanism layer calculates the temporal dependency of the trajectory feature vector to obtain the contextual trajectory representation. The context trajectory representation is predicted by the conflict probability decoding layer to obtain multiple predicted position sequences and their existence probabilities; The predicted position sequence with the highest probability of existence is taken as the future movement path of the moving target object.
7. The intersection detection and early warning method according to claim 1, characterized in that, The step involves calculating the minimum collision time between any two moving target objects using the future movement path. If the minimum collision time is less than a preset warning threshold, a warning instruction is generated, including: By using any two of the future paths, the spatiotemporal occupancy domain of the moving target object within a preset prediction period is determined; Calculate the overlap period when the two spatiotemporal occupied domains first overlap in the spatiotemporal coordinate system or when the distance is less than a safety threshold. The difference between the start time of the overlapping period and the current time is determined as the minimum collision time. If the minimum collision time is less than the preset warning threshold, a warning command is generated.
8. An intersection detection and early warning system, characterized in that, include: The data acquisition module is used to acquire point cloud data and initial images of the moving target objects when at least one moving target object is detected at the intersection. The state recognition module is used to identify and analyze the initial image through the point cloud data to obtain target state information; The trajectory fusion module is used to fuse the point cloud data with the target state information to obtain the trajectory of the target object; The risk prediction module is used to predict the trajectory of the target object using a preset risk analysis model, so as to obtain the future movement path of the moving target object; The collision analysis module is used to calculate the minimum collision time between any two moving target objects based on the future movement path. If the minimum collision time is less than a preset warning threshold, a warning instruction is generated. The warning execution module is used to send the warning command to the roadside display device and / or vehicle terminal for two-way warning prompts.
9. An electronic device, characterized in that, The method includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor causes the processor to perform the steps of the intersection detection and early warning method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, it implements the intersection detection and early warning method as described in any one of claims 1-7.