Real-time POI space-time positioning method based on vehicle-mounted video, medium and equipment
Through the real-time POI spatiotemporal positioning method based on vehicle video, the YOLOv5 network and binocular depth estimation algorithm are used, combined with Kalman filtering and least squares method, the high-precision POI feature recognition and spatiotemporal positioning problems in GPS-free environment are solved, and stable positioning and trajectory calculations are realized in complex environments.
Patent Information
- Application Number
- CN202510496356.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-01
AI Technical Summary
The existing vehicle navigation and autonomous driving systems are difficult to achieve high-precision POI feature recognition and spatio-temporal positioning in environments without GPS signals or unstable signals. Especially in complex environments, the recognition accuracy is insufficient, and the vehicle speed estimation is not accurate, resulting in a decrease in positioning accuracy.
The real-time POI spatio-temporal positioning method based on vehicle video is adopted, and the POI feature recognition is performed through the YOLOv5 network combined with the SE attention mechanism and DIoU-NMS. The video frame is converted into a depth map using the binocular depth estimation calculation method, the POI distance is calculated by combining the POI features and depth maps, and the speed trajectory is optimized by Kalman filtering and least squares method to calculate the time when the vehicle arrives at the POI in real time.
With no GPS support, high-precision POI feature recognition and spatio-temporal positioning are achieved, which can accurately identify road features in complex environments, and through the combination of continuous video frames and vehicle speed and POI features, it provides stable positioning and trajectory calculations, adapt to dynamic environment changes, and improves positioning accuracy and robustness.
Smart Images

Figure CN120411229A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a real-time POI spatio-temporal positioning method, medium, and device based on in-vehicle video. Background Art
[0002] With the rapid development of intelligent transportation and autonomous driving technologies, in-vehicle navigation systems and autonomous driving systems have become important components of modern traffic management and safety assurance. As one of the core technologies in this field, in-vehicle video analysis systems are gradually integrated into multiple aspects such as smart city construction, traffic flow monitoring, and accident warning. The real-time positioning and navigation capabilities of in-vehicle systems play a crucial role in improving traffic management efficiency, ensuring traffic safety, and promoting the development of autonomous driving technologies.
[0003] Currently, the positioning methods commonly used in in-vehicle navigation and autonomous driving systems mainly rely on GPS (Global Positioning System) for real-time positioning. However, GPS signals are prone to loss in complex environments such as urban canyons and tunnels, resulting in a significant decline in positioning accuracy. In addition, the reliability of GPS signals is also limited in some high-security areas (such as military and sensitive regions), unable to meet the requirements of high-precision positioning. To solve this problem, many researchers and engineers have started to explore methods for achieving high-precision positioning without GPS signals. Against this background, feature recognition and spatio-temporal positioning of points of interest (POIs) based on in-vehicle video data have become a research hotspot. POI features, such as toll stations, tunnels, bridges, road intersections, etc., are crucial markers for positioning on the road. By identifying and analyzing these features, the vehicle's position can be inferred, thereby improving the positioning accuracy of in-vehicle systems.
[0004] Currently, most of the POI recognition methods used in in-vehicle video analysis systems and autonomous driving systems rely on object detection algorithms such as YOLO, Faster R-CNN, etc. These methods can effectively identify static objects (such as road signs, traffic lights, etc.) and some dynamic objects (such as other vehicles, pedestrians, etc.) in images or videos, and have been applied to a certain extent in in-vehicle systems. However, the existing object recognition methods still have the following problems:
[0005] Limited recognition accuracy: Existing POI recognition algorithms usually have difficulty ensuring high-precision recognition under complex environmental conditions (such as low light, blur, occlusion, etc.). For example, the recognition of POI features such as toll stations, tunnels, and bridges in in-vehicle videos may be affected by environmental factors (such as reflection, weather changes, etc.), resulting in a reduction in the accuracy of recognition results.
[0006] Unable to adapt to dynamic environments: Most existing POI recognition algorithms are trained for static environments. When encountering dynamic scenarios (such as the appearance of temporary obstacles or traffic incidents in front of a vehicle), the recognition ability and adaptability of the algorithms are poor, and they are unable to track and process the changing environment in real time.
[0007] Spatio-temporal positioning refers to accurately estimating the time when a vehicle arrives at a specific location (i.e., a POI) in a vehicle-mounted system by combining the POI recognition results and other sensor data. Although traditional positioning technologies (such as GPS) can provide certain spatial positioning information for vehicles, in practical applications, the GPS positioning method still has many limitations:
[0008] Dependence on GPS signals: Most existing systems rely on GPS signals for positioning. However, in environments such as urban canyons, tunnels, and underground parking lots, GPS signals may be blocked or interfered with, resulting in a decrease in positioning accuracy or even the inability to obtain effective positioning information. In addition, in highly secure fields such as military and scientific research, the use of GPS signals is often restricted, so it is impossible to rely on traditional GPS technology for positioning.
[0009] Insufficient positioning accuracy: Traditional spatio-temporal positioning methods rely heavily on GPS and other sensor data. In environments without GPS signals or with unstable signals, the positioning accuracy drops significantly, and accurate POI positioning and time prediction cannot be achieved. Even in open environments, GPS signals cannot provide sufficiently accurate positioning results, especially for some applications that require centimeter-level accuracy (such as autonomous driving and precise navigation).
[0010] Inaccurate vehicle speed estimation: In existing vehicle-mounted systems, speed estimation usually relies on GPS or other vehicle-mounted sensors. This can lead to errors in speed estimation in cases where signals are limited or sensors are inaccurate, thereby affecting the accuracy of spatio-temporal positioning.
[0011] Most existing research focuses on the recognition of POI features and does not consider how to combine the POI recognition results with the spatio-temporal positioning of the vehicle. Even if POI features are recognized, how to further estimate the accurate position and arrival time of the vehicle based on these features remains a difficult problem. Most research only stays at the stage of static target recognition and fails to fully explore the potential value of this information in spatio-temporal positioning. Therefore, existing systems cannot provide high-precision spatio-temporal positioning based on POI recognition. Summary of the Invention
[0012] The object of the present invention is to propose a real-time POI spatio-temporal positioning method based on in-vehicle video to solve the problem of POI spatio-temporal positioning in the absence of GPS positioning signals, including the following steps:
[0013] S1. Obtain in-vehicle video stream data of various traffic scenarios and traffic environments, and label POI feature images;
[0014] S2. Perform POI feature recognition on the in-vehicle video stream data based on the labeled POI feature images; meanwhile, use a binocular depth estimation algorithm to convert the video frames of the in-vehicle video stream data into depth maps;
[0015] S3. Combine the POI features and the depth maps to calculate the POI distances;
[0016] S4. Calculate the vehicle speed based on the change in POI distances and time information between consecutive video frames;
[0017] S5. Use the POI distances and the vehicle speed to calculate the timestamp when the vehicle arrives at the POI in real time.
[0018] Furthermore, use the YOLOv5 network to perform POI feature recognition on the in-vehicle video stream data.
[0019] Furthermore, insert the SE (Squeeze-and-Excitation) attention mechanism after each convolutional layer in the YOLOv5 network, and use DIoU-NMS (Distance-Intersection over Union Non-Maximum Suppression) to replace NMS (Non-Maximum Suppression) in the prediction stage of the YOLOv5 network.
[0020] Furthermore, use a binocular depth estimation algorithm to convert the video frames of the in-vehicle video stream data into depth maps, specifically:
[0021] Obtain the internal and external parameters of the camera through binocular calibration. The internal parameters describe the internal characteristics and distortion of the camera, and the external parameters describe the position and attitude of the camera;
[0022] Use the internal and external parameters of binocular calibration and the relative position relationship of the binocular cameras to eliminate distortion and perform row alignment on the left and right video frames;
[0023] Calculate the disparity information between the pixel points of the left and right cameras through the SGBM (Semi-Global Block Matching) stereo matching algorithm to obtain a disparity map;
[0024] Based on the disparity map obtained by stereo matching, use the geometric relationship of the binocular cameras to calculate the depth information of each pixel point using geometric methods;
[0025] The formula for calculating pixel depth using the disparity map:
[0026]
[0027] Among them, depth represents pixel depth, b is the baseline length between the left and right camera centers, d is the parallax, f is the focal length, and c is the distance between the left and right cameras. xr represents the column coordinates of the right camera principal point, c xl Column coordinates representing the principal point of the left camera.
[0028] Furthermore, the POI distance is calculated by combining the POI features and the depth map, which is expressed as:
[0029]
[0030] Among them, D POI represents the POI distance, N represents the total number of pixels of the POI feature, Z i Represents the depth value of the i-th pixel, W i Represents the weight value of the i-th pixel.
[0031] Furthermore, W i Calculated by the following formula:
[0032]
[0033] Among them, α represents the parameter that controls the ratio of gradient weight and spatial weight, represents the gradient magnitude of the i-th pixel, d i represents the distance from the i-th pixel to the POI center, σ represents the standard deviation of the Gaussian function, represents the gradient magnitude of the j-th pixel.
[0034] Furthermore, the vehicle speed is calculated based on the POI distance change and time information between consecutive video frames, specifically:
[0035] The inter-frame speed is calculated according to the following formula:
[0036]
[0037] ΔD t =|D t+1 -D t |
[0038] Among them, V t represents the speed of the vehicle in the tth frame, ΔD t Indicates the change in POI distance between the t+1th frame and the tth frame of the continuous video frame, D t+1 Denotes the POI distance of the t+1th frame, D t represents the POI distance of the t-th frame;
[0039] Kalman filtering is used to smooth the inter-frame velocity, and the formula is:
[0040]
[0041] Among them, represents the speed estimation of the vehicle at the t-th frame, represents the predicted speed of the vehicle at the t-th frame, K t represents the Kalman gain, V t represents the speed of the vehicle at the t-th frame, represents the speed estimation of the vehicle at the (t - 1)-th frame, Δt represents the time for the vehicle to reach the POI, a t-1 represents the acceleration of the vehicle at the (t - 1)-th frame, represents the prediction error covariance, R t represents the covariance of the observation noise;
[0042] Optimize the speed trajectory through the least squares optimization algorithm:
[0043]
[0044] Among them, V1, V2,..., V N represents the instantaneous speed of the vehicle calculated within N frames.
[0045] Furthermore, calculate the time for the vehicle to reach the POI according to the following formula:
[0046]
[0047] Among them, Δt represents the time for the vehicle to reach the POI, V t represents the speed at the t-th frame, a t represents the acceleration at the t-th frame, D t represents the POI distance at the t-th frame.
[0048] The present invention also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned real-time POI spatio-temporal positioning method based on in-vehicle video is implemented.
[0049] The present invention also provides an electronic device, including a processor and a memory, the processor is connected to the memory, among which, the memory is used to store a computer program, the computer program includes computer-readable instructions, and the processor is configured to call the computer-readable instructions to execute the above-mentioned real-time POI spatio-temporal positioning method based on in-vehicle video.
[0050] The beneficial effects brought by the technical solution provided by the present invention are:
[0051] The present invention proposes a real-time POI spatio-temporal positioning method based on in-vehicle video. POIs are obtained through in-vehicle video, and video frames of in-vehicle video stream data are converted into depth maps. Combining POI features and depth maps, the distance of the POI and the vehicle speed are calculated, and the time for the vehicle to reach the POI is calculated in real time. By combining consecutive video frames, vehicle position and speed with the spatial position of POI features, high-precision positioning and trajectory calculation are achieved. Combining in-vehicle video data and spatio-temporal information, it is possible to accurately identify POI features on the road and analyze the spatio-temporal relationship between the vehicle and these features without GPS positioning support. Description of the Drawings
[0052] Figure 1 is a flowchart of the real-time POI spatio-temporal positioning method based on in-vehicle video according to an embodiment of the present invention;
[0053] Figure 2 is a schematic diagram of stereo rectification according to an embodiment of the present invention, Figure 2 in which (a) is a schematic diagram before stereo rectification, Figure 2 in which (b) is a schematic diagram after stereo rectification;
[0054] Figure 3 is a block diagram of an electronic device in an exemplary embodiment according to an embodiment of the present invention;
[0055] Figure 4 is an example diagram of a binocular ranging algorithm according to an embodiment of the present invention, Figure 4 in which (a) is a disparity map, Figure 4 in which (b) is a depth map. Detailed Embodiments
[0056] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described below in conjunction with the accompanying drawings.
[0057] The flowchart of the real-time POI spatio-temporal positioning method based on in-vehicle video according to an embodiment of the present invention is as Figure 1 , and specifically includes the following steps:
[0058] S1. Obtain in-vehicle video stream data of various traffic scenarios and traffic environments, and label POI feature images.
[0059] To ensure the diversity and richness of the constructed dataset, data is obtained from the following three data sources: ① From the road video data (bag data packet) collected by its own visual sensor, images containing corresponding POI features are screened and extracted; ② POI feature images that meet the requirements are searched from publicly available image libraries; ③ From open-source outdoor target detection datasets, such as the nuscenes and Kitti datasets, POI feature images that meet the requirements are selected.
[0060] POI in the present invention refers to landmarks on the road that are helpful for real-time positioning of vehicles, such as toll booths, tunnels, bridges, overpasses, gantries, road intersections, etc.;
[0061] Spatiotemporal positioning refers to the real-time calculation and feedback of the time when the vehicle arrives at the POI.
[0062] S2. POI feature recognition is performed on the vehicle-mounted video stream data based on the annotated POI feature image; at the same time, a binocular depth estimation algorithm is used to convert the video frames of the vehicle-mounted video stream data into a depth map.
[0063] In one embodiment of the present invention, the YOLOv5 network is used to identify POI features in vehicle-mounted video stream data. The YOLOv5 network architecture primarily consists of an input network, a backbone network (for computing convolutional features), a feature fusion network, and an output network (for outputting object category and location information). YOLOv5 boasts fast training speed, high image processing efficiency, and superior detection accuracy, making it suitable for scenarios with rapidly changing environments and demanding real-time performance.
[0064] The YOLOv5 network architecture consists of a backbone network, a feature fusion layer, and a head. The YOLOv5 network's backbone structure boasts exceptional multi-layer feature extraction capabilities. The input raw image is first downsampled at intervals using the Focus architecture, effectively reducing feature loss and increasing computational speed. Next, the CBL architecture is combined with the CSP1_x architecture to achieve efficient and reliable feature extraction and optimization. Finally, the SPP (Spatial Pyramid Pooling) architecture utilizes multi-scale pooling layers to enhance key features, further reducing data volume and improving algorithm efficiency.
[0065] The YOLOv5 network's neck fuses multi-level features extracted by the extraction module through FPN (Feature Pyramid Networks), and enhances the fluidity and aggregation of features with PANet (Path Aggregation Network), improving multi-scale object detection performance. FPN transmits high-level semantic features through a top-down path and fuses low-level features with upsampled features through lateral connections, ensuring that low-level features retain spatial detail while acquiring high-level semantic information. PANet builds on this by introducing a bottom-up information flow path, enhancing the semantic information of low-level features and the spatial information of high-level features, further improving the detection capabilities of multi-scale POI features.
[0066] The head of the YOLOv5 network divides the original image into grids of three scales through multi-scale grid detection, and each grid detects the objects within it, thereby enhancing the network's detection ability for objects of different scales. Its end-to-end structure generates prediction results at the network output layer, improving the detection efficiency and speed. The bounding boxes are predicted through the anchor mechanism, simplifying model learning and improving prediction accuracy. Finally, YOLO outputs prediction results at three scales, including the center coordinates, length and width, confidence level of the bounding boxes, and the class probabilities of each POI feature.
[0067] In another embodiment of the present invention, the YOLOv5 network is improved in the following two aspects: (1) An attention mechanism is added to the YOLOv5 network, and an SE (Squeeze-and-Excitation) module is inserted after each convolutional layer in the YOLOv5 network; (2) DIoU-NMS is used to replace NMS (Non Maximum Suppression) during the prediction stage of the YOLOv5 network.
[0068] To improve the performance of YOLOv5 in complex environments, especially in POI feature recognition, an SE attention mechanism is added to YOLOv5. This improvement enhances the expression of important features and suppresses irrelevant information by adaptively adjusting the channel feature responses, thereby improving the detection accuracy and robustness.
[0069] The SE attention mechanism is a channel attention mechanism that aims to automatically adjust the importance of each channel through two steps: "squeeze" and "excitation". Its basic principle is to compress the feature map through global average pooling to extract the global information of each channel, and then use a fully connected layer to excite each channel to obtain the weight coefficient of each channel. These coefficients are used to weight the features of each channel, thereby enhancing the response of important features.
[0070] An SE module is inserted after each convolutional layer in the YOLOv5 network, and the execution of this module mainly involves the following two steps:
[0071] (1) Channel squeeze:
[0072] For each convolutional feature map of YOLOv5, first apply the global average pooling operation to generate a global descriptor, which represents the global information of each channel in this feature map. This step can capture the overall features of each channel and remove the redundant information in the spatial dimension.
[0073] (2) Channel excitation:
[0074] Through a two-layer fully connected network, the channel descriptors are first mapped to a smaller dimension, then non-linearly transformed through the ReLU activation function, and finally the output is mapped between 0 and 1 through the sigmoid activation function to represent the importance of each channel. Finally, the excited channel weights are multiplied with the original feature map channel by channel to enhance the important features.
[0075] The SE attention mechanism improves the recognition accuracy of POI features in low-light or blurred environments by strengthening the expression of key information. It can adapt to dynamic changes, enhance the recognition ability of environmental changes such as temporary obstacles or traffic events, and improve the system response speed. At the same time, the SE mechanism enhances the attention to detailed features by adaptively adjusting the channel weights, helping to capture subtle changes in the road scene more precisely.
[0076] To further improve the detection accuracy and recall rate of YOLOv5 in processing complex scenes and overlapping targets in vehicle-mounted video analysis, the traditional non-maximum suppression (NMS) algorithm is improved, and DIoU-NMS is used to replace NMS.
[0077] Non-maximum suppression is a commonly used post-processing method in object detection. It selects the optimal box from the candidate boxes according to the confidence score and removes redundant boxes with high overlap. Traditional NMS relies on the intersection over union (IoU) to judge the overlap between boxes. When the IoU exceeds the threshold, the box with a lower confidence is suppressed. However, in scenes with dense and severely overlapping targets, traditional NMS may incorrectly suppress important targets and affect the detection effect.
[0078] Based on traditional NMS, DIoU-NMS introduces the distance information between the center points of the object bounding boxes, further optimizing the processing of the overlapping area. It reduces redundant boxes and improves the detection accuracy and recall rate. In DIoU-NMS, IoU is an index to measure the overlap degree of two candidate boxes, but the Euclidean distance between the center points of the two boxes is further added to improve the distinguishability of the targets. For two boxes B1 and B2, calculate their IoU and center distance, and the formulas are as follows:
[0079]
[0080] Among them, IoU(B1,B2) represents the intersection over union of B1 and B2, d is the Euclidean distance between the center points of the two boxes, and c is the diagonal length of the smallest closed bounding box containing the two boxes.
[0081] In DIoU-NMS, the suppression of candidate bounding boxes is based on a weighted combination of their IoU values and the distance between their center points. Different from traditional NMS which only filters according to the IoU value, DIoU-NMS determines the priority of bounding boxes by comprehensively considering both the IoU and the center distance. During the calculation process, the bounding box with the highest confidence is first selected as the reference box and compared with other boxes. If the DIoU (Distance-IoU) value is lower than the threshold, the box is retained; otherwise, it is suppressed.
[0082] In scenarios with dense and severely occluded targets, traditional NMS may incorrectly suppress important targets, especially when the target overlaps are large. By incorporating the distance between the center points of the targets, DIoU-NMS can more accurately judge the target relationships, reduce unnecessary suppression, and improve the detection accuracy.
[0083] DIoU-NMS can also effectively reduce the mis-suppression of overlapping targets. Especially in complex backgrounds and high-density target situations, by increasing the separation of targets through the center point distance, DIoU-NMS can better maintain the separability of targets, thereby improving the reliability of detection.
[0084] The binocular depth estimation algorithm is used to convert the video frames of vehicle-mounted video stream data into depth maps, specifically as follows:
[0085] (1) Binocular calibration: The internal parameters (such as focal length, principal point position) and external parameters (position and attitude) of the camera are obtained through binocular calibration. The internal parameters describe the internal characteristics and distortions of the camera, and the external parameters describe the position and attitude of the camera. The Zhang-Zhengyou calibration method is adopted, and the calibration is implemented through the Stereo Camera Calibrator APP in Matlab to ensure accurate acquisition of the geometric information of the scene.
[0086] (2) Stereo rectification: Stereo rectification uses the internal and external parameters obtained from binocular calibration and the relative position relationship (rotation matrix and translation vector) of the binocular cameras to eliminate distortions and align rows for the left and right video frames, so that the imaging origins of the two video frames are the same, the optical axes are parallel, the imaging planes are coplanar, and the epipolar lines are row-aligned. Before rectification, there may be distortions and differences in the image coordinate systems in the binocular cameras, resulting in misaligned images and increasing the matching difficulty. Through stereo rectification, the corresponding rows of the left and right images are aligned, simplifying the matching process and improving the accuracy.
[0087] A schematic diagram before stereo rectification, as shown in Figure 2 (a) shows that before rectification, the projection point P of the left camera l and the projection point P of the right camera r are in different rows and the imaging planes are not coplanar; a schematic diagram after stereo rectification, as shown in Figure 2 (b) shows that the imaging planes of the binocular cameras are coplanar and the two projection points are in the same row, reaching an ideal state.
[0088] (3) Stereo matching: The key to ranging in a binocular vision system lies in obtaining the disparity of the target. The disparity information between the pixel points of the left and right cameras is calculated through a stereo matching algorithm, where disparity = u l - u r , and then the depth of the object is inferred. The SGBM stereo matching algorithm uses the BT cost to calculate the matching cost and combines the cost aggregation method of SGM, which not only ensures the accuracy of the matching but also improves the calculation efficiency. It is suitable for real-time stereo vision measurement depth calculation to obtain a disparity map.
[0089] (4) Depth calculation: Based on the disparity map obtained from stereo matching, using the geometric relationship of the binocular cameras, the depth information of each pixel point is calculated. The basic principle of depth calculation is to use triangulation or other geometric methods to convert the disparity into the distance from the object to the camera.
[0090] Formula for calculating pixel depth using the disparity map:
[0091]
[0092] Among them, depth represents the pixel depth, b is the baseline length between the center points of the left and right cameras, d is the disparity, f is the focal length, c xr represents the column coordinate of the principal point of the right camera, and c xl represents the column coordinate of the principal point of the left camera.
[0093] S3. Combine the POI features and the depth map to calculate the POI distance.
[0094] After obtaining the depth map, the position of the POI in the image can be determined according to the recognition result of the target POI. For each pixel of the POI, the corresponding depth value is extracted through the depth map, and the distance of the POI feature is calculated. Considering that the POI is usually a region rather than a single pixel point, it is necessary to statistically analyze the depth values of multiple pixels within the POI region. The method adopted is to calculate the weighted average value of the depth values within this region. The specific formula is:
[0095]
[0096] Among them, D POI represents the POI distance, N represents the total number of pixels of the POI feature, Z i represents the depth value of the i-th pixel, and W i represents the weight value of the i-th pixel.
[0097] After comprehensively considering feasibility and accuracy, a saliency weight scheme based on pixel gradient is selected, and weight distribution is combined with a Gaussian weighting function. This method can effectively process complex visual features and is especially suitable for the precise positioning of POI features in vehicle-mounted videos.
[0098] (1) Pixel gradient saliency weight
[0099] First, calculate the gradient value of each pixel in the image. The gradient reflects the degree of brightness change in the image and is usually used to describe edges and details. Edge regions usually contain more object information and contribute more to depth calculation. To calculate the gradient, the classical Sobel operator can be used to obtain the gradient magnitude of each pixel. The specific formula is:
[0100]
[0101] where G x and G y are the gradient values of the image in the horizontal and vertical directions respectively.
[0102] (2) Gaussian weighting function
[0103] The Gaussian weighting function is used to weight each pixel in the image, making the pixels closer to the center of the POI have higher weights and the pixels at the edge positions have lower weights. The Gaussian function can smoothly distribute the weight w i and reduce the influence of pixels far from the center of the POI. The specific formula is:
[0104]
[0105] where d i is the distance from the i-th pixel to the center of the POI, and σ is the standard deviation of the Gaussian function, which controls the speed of weight decay. A smaller σ value will make the weight distribution more concentrated at the center of the POI.
[0106] Combining saliency and Gaussian weighting
[0107] To comprehensively consider the influence of gradient saliency and spatial position, we combine the two to calculate the final weight of each pixel. This scheme can effectively balance edge saliency and spatial position, ensuring that important regions contribute more to depth calculation. The final weight W i is the weighted sum of the gradient weight and the Gaussian weight:
[0108]
[0109] where is the normalized gradient saliency weight, ensuring that the sum of the weights is 1. α ∈ [0, 1] controls the ratio of the gradient weight to the spatial weight, with a default value of 0.5. represents the gradient magnitude of the i-th pixel, d i represents the distance from the i-th pixel to the center of the POI, σ represents the standard deviation of the Gaussian function, represents the gradient magnitude of the j-th pixel.
[0110] S4. Calculate the vehicle speed based on the change in the distance between POIs and the time information between consecutive video frames.
[0111] When a POI feature is detected and its confidence exceeds a preset threshold, the vehicle speed will be calculated based on the position change and time difference of the POI feature in consecutive video frames: Assume the image frames captured by the in-vehicle video system are I1, I2,... I N , in the t-th and (t + 1)-th frames, the actual distances D t and D t+1 (in meters) of these POI positions are obtained respectively through the binocular depth estimation method. The specific formula is:
[0112]
[0113] ΔD t =|D t+1 -D t |
[0114] where, V t represents the speed of the vehicle at the t-th frame, ΔD t represents the change in the distance between POIs between the (t + 1)-th and t-th frames of consecutive video frames, D t+1 represents the POI distance at the (t + 1)-th frame, and D t represents the POI distance at the t-th frame;
[0115] To improve the accuracy of speed estimation and eliminate the noise in instantaneous speed calculation, especially in the presence of dynamic objects or occlusions, the calculated speed data needs to be smoothed. The Kalman Filter is used to smooth the speed. The Kalman Filter is a recursive estimation method based on the state space model, which can effectively fuse sensor data and suppress noise. Assume the speed V t of the vehicle is the state variable of the system. The process of the Kalman Filter includes two main steps: prediction and update.
[0116] (1) Prediction step
[0117] Predict the speed at the current moment based on the speed estimate at the previous moment
[0118]
[0119] where, a t-1 is the acceleration at the previous moment, which can usually be estimated by the speed change of consecutive frames
[0120] (2) Update step
[0121] Estimate the speed V based on the current frame t and the predicted speed Update the estimated value of the speed:
[0122]
[0123] where K t is the Kalman gain, reflecting the weights of the prediction and the observed data, and the calculation formula is:
[0124]
[0125] where represents the estimated speed of the vehicle at the t-th frame, represents the predicted speed of the vehicle at the t-th frame, K t represents the Kalman gain, V t represents the speed of the vehicle at the t-th frame, represents the estimated speed of the vehicle at the (t - 1)-th frame, Δt represents the time when the vehicle is expected to reach the POI, a t-1 represents the acceleration of the vehicle at the (t - 1)-th frame, represents the prediction error covariance, R t represents the covariance of the observation noise; the updated speed is the vehicle speed after smoothing. Through Kalman filtering, the errors introduced by dynamic target interference or video noise can be effectively eliminated, making the speed estimation more stable and reliable.
[0126] To further improve the accuracy, a least squares optimization algorithm is also introduced to optimize the speed trajectory in consecutive frames and correct the errors caused by dynamic targets and occlusions. Assume that the instantaneous speeds of the vehicle V1, V2,..., V N are calculated within N frames. The speed trajectory is optimized by minimizing the following error function:
[0127]
[0128] where is the speed smoothed by Kalman filtering, V t is the calculated instantaneous speed. By minimizing the error, a more accurate vehicle speed and trajectory can be optimized, and the deviation caused by dynamic targets and video noise can be corrected.
[0129] By utilizing the distance change and time difference of the POI in consecutive frames, combining Kalman filtering and least squares optimization, high-precision vehicle speed estimation can be achieved. This method effectively eliminates the errors caused by dynamic target interference, video noise, and environmental changes, and can provide smooth, stable, and reliable speed data, further supporting accurate spatio-temporal positioning tasks.
[0130] S5. Use the POI distance and vehicle speed to calculate the timestamp when the vehicle arrives at the POI in real time.
[0131] In the spatio-temporal positioning task, accurately predicting the timestamp when the vehicle arrives at the target POI is one of the core objectives. Based on the vehicle's speed information and the POI distance information, this algorithm performs simple time prediction through physical formulas, thereby accurately calculating the time required for the vehicle to travel from the current location to the target feature location, and estimating the time T when the vehicle arrives at the POI accordingly. To improve the prediction accuracy, considering the acceleration or deceleration of the vehicle, the acceleration a is introduced. t The specific formula is:
[0132]
[0133] where, Δt represents the time when the vehicle arrives at the POI, V t represents the speed at the t-th frame, a t represents the acceleration at the t-th frame, D t represents the POI distance at the t-th frame.
[0134] To further improve the prediction accuracy, the estimation of the acceleration a can be optimized by introducing historical speed data. At the same time, a confidence threshold is set. When the confidence of consecutive multi-frame detections is lower than this threshold, the system will stop subsequent POI detections and time estimations. Finally, the system uses the detection result that satisfies the confidence threshold for the last time and the corresponding estimated time as the final prediction value of the time when the vehicle arrives at the POI. t In an exemplary embodiment, there is provided a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned real-time POI spatio-temporal positioning method based on in-vehicle video is implemented.
[0135]
[0136] Figure 3 Please refer to In an exemplary embodiment, there is further provided an electronic device including at least one processor, at least one memory, and at least one communication bus.
[0137] wherein, a computer program is stored on the memory, the computer program includes computer-readable instructions, and the processor calls the computer-readable instructions stored in the memory through the communication bus to execute the above-mentioned real-time POI spatio-temporal positioning method based on in-vehicle video.
[0138] To verify the effectiveness of the method of the present invention, the present invention obtains a dataset of in-vehicle video images of 4400 traffic scenes, covering POI targets such as toll stations, bridges, tunnels, and road intersections, and performs accurate annotation.
[0139] This dataset contains various types of POIs and a large number of complex traffic scenarios. To better simulate various situations in actual applications, the dataset covers a variety of environmental conditions, including:
[0140] Different weather conditions: various extreme weather scenarios such as sunny days, rainy days, and haze;
[0141] Different time periods: including day and night to ensure the ability to handle POI recognition tasks under different lighting conditions;
[0142] Complex traffic environments: including complex scenarios such as high-density urban traffic, rural roads, and tunnels to enhance the robustness of the model;
[0143] Dynamic object interference: Pedestrians, other vehicles and other dynamic targets are specifically added to the dataset to simulate possible interference objects on the actual road and challenge the object detection and recognition capabilities of the model.
[0144] By simulating a variety of complex traffic environments, the dataset enhances the model's adaptability to various interference factors in real-world scenarios. Specifically, it includes: through the simulation of occluders, dynamic targets, and strong light source interference, the model's recognition ability for perspective changes and dynamic objects is improved; through the simulation of low light, bad weather (such as rainy days, snowy days, and haze), and road condition changes (such as potholes and bumps), the model's robustness under different lighting and weather conditions is enhanced.
[0145] Use the existing dataset to train the improved YOLO-v5 object detection algorithm, and view the results such as recognition accuracy and recall rate in training and testing. And make corresponding improvements to continuously improve the detection accuracy. The training parameter settings are as follows: epochs = 400, batch_size = 20, img_size = 640×640. The network is trained for 400 epochs on a total of 3960 images of various categories in the training set, and at the same time tested on 440 images in the test set. The verification results of the test set are shown in Table 1:
[0146] Table 1
[0147]
[0148]
[0149] Based on the POI feature recognition, it is necessary to further provide the corresponding ability of POI spatio-temporal positioning, that is, by calculating the vehicle speed information and the distance information from the current vehicle to the target feature, calculate the exact time when the vehicle travels from the current position to the POI feature position. The example diagram of the binocular ranging algorithm in the embodiment of the present invention is referred to Figure 4 , Figure 4 where (a) in Figure 4In (b) is the depth map.
[0150] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A real-time POI spatio-temporal positioning method based on in-vehicle video, characterized in that, It includes the following steps: S1. Obtain in-vehicle video stream data of various traffic scenarios and traffic environments, and label POI feature images; S2. Perform POI feature recognition on the in-vehicle video stream data based on the labeled POI feature images; meanwhile, use a binocular depth estimation algorithm to convert the video frames of the in-vehicle video stream data into depth maps; S3. Combine the POI features and the depth maps to calculate the POI distances; S4. Calculate the vehicle speed based on the change in POI distances between consecutive video frames and the time information; S5. Use the POI distances and the vehicle speed to calculate the timestamp when the vehicle arrives at the POI in real time.
2. The real-time POI spatio-temporal positioning method based on in-vehicle video according to claim 1, characterized in that, Use the YOLOv5 network to perform POI feature recognition on the in-vehicle video stream data.
3. A real-time POI spatio-temporal positioning method based on in-vehicle video according to claim 2, characterized in that, Insert an SE attention mechanism after each convolutional layer in the YOLOv5 network, and use DIoU-NMS to replace NMS in the prediction stage of the YOLOv5 network.
4. A real-time POI spatio-temporal positioning method based on in-vehicle video according to claim 1, characterized in that Use a binocular depth estimation algorithm to convert the video frames of the in-vehicle video stream data into depth maps, specifically: Obtain the internal and external parameters of the camera through binocular calibration. The internal parameters describe the internal characteristics and distortion of the camera, and the external parameters describe the position and attitude of the camera; Perform distortion elimination and row alignment on the left and right video frames by using the internal and external parameters of binocular calibration and the relative position relationship of the binocular cameras; Calculate the disparity information between the pixel points of the left and right cameras through the SGBM stereo matching algorithm to obtain a disparity map; Based on the disparity map obtained by stereo matching, use the geometric relationship of the binocular cameras and calculate the depth information of each pixel point by a geometric method; The formula for calculating the pixel depth using the disparity map: where depth represents the pixel depth, b is the baseline length between the center points of the left and right cameras, d is the disparity, f is the focal length, c xr represents the column coordinate of the principal point of the right camera, and c xl represents the column coordinate of the principal point of the left camera.
5. The real-time POI spatio-temporal positioning method based on in-vehicle video according to claim 1, wherein Combine the POI features and the depth maps to calculate the POI distances, expressed as: Among them, D POI represents the POI distance, N represents the total number of pixels of the POI feature, Z i represents the depth value of the i-th pixel, W i represents the weight value of the i-th pixel.
6. A real-time POI spatio-temporal positioning method based on in-vehicle video according to claim 5, characterized in that W i Calculated by the following formula: Among them, α represents a parameter for controlling the ratio of the gradient weight and the spatial weight, represents the gradient amplitude of the i-th pixel, d i represents the distance from the i-th pixel to the center of the POI, and σ represents the standard deviation of the Gaussian function, represents the gradient amplitude of the j-th pixel.
7. A real-time POI spatio-temporal positioning method based on in-vehicle video according to claim 1, characterized in that, Calculate the vehicle speed based on the change in POI distances between consecutive video frames and the time information, specifically: Calculate the inter-frame speed according to the following formula: ΔD t = |D t+1 - D t | Among them, V t represents the speed of the vehicle at the t-th frame, and ΔD t represents the change in the POI distance between the (t + 1)-th frame and the t-th frame of consecutive video frames, and D t+1 represents the POI distance of the (t + 1)-th frame, and D t represents the POI distance of the t-th frame; Perform smoothing processing on the inter-frame speed using the Kalman filter, and the formula is: Among them, represents the speed estimate of the vehicle at the t-th frame, represents the predicted speed of the vehicle at the t-th frame, K t represents the Kalman gain, V t represents the speed of the vehicle at the t-th frame, represents the speed estimate of the vehicle at the (t - 1)-th frame, Δt represents the time for the vehicle to reach the POI, a t-1 represents the acceleration of the vehicle at the (t - 1)-th frame, represents the prediction error covariance, R t represents the covariance of the observation noise; Optimize the speed trajectory through the least squares optimization algorithm: Among them, V1, V2, ..., V N represent the instantaneous speed of the vehicle calculated within N frames.
8. A real-time POI spatio-temporal positioning method based on in-vehicle video according to claim 1, characterized in that, Calculate the time when the vehicle arrives at the POI according to the following formula: Among them, Δt represents the time when the vehicle arrives at the POI, V t represents the speed at the t-th frame, a t represents the acceleration at the t-th frame, D t represents the POI distance at the t-th frame.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the method described in any one of claims 1-8.
10. An electronic device, characterized in that, It includes a processor and a memory, the processor is connected to the memory. Among them, the memory is used to store a computer program, the computer program includes computer-readable instructions, and the processor is configured to call the computer-readable instructions to execute the method described in any one of claims 1-8.