Monocular vision-based target ranging and speed measuring method

By combining pinhole ranging, end-to-end ranging, and temporal information into a multi-task network, the problem of unstable ranging in monocular vision under complex scenes is solved, achieving higher accuracy and more stable target ranging and velocity measurement.

CN120926942BActive Publication Date: 2026-07-28WUHAN JIMU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUHAN JIMU INTELLIGENT TECH CO LTD
Filing Date
2025-07-15
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Existing monocular vision-based target ranging methods have low ranging accuracy in complex scenes and are easily affected by target detection box jitter, resulting in unstable ranging.

Method used

By combining pinhole ranging, end-to-end ranging, and temporal information, the target parameter values ​​are output through a multi-task network, and training labels are created using multimodal data. A temporal fusion method is then used for ranging and velocity measurement.

Benefits of technology

It improves the accuracy and robustness of target ranging and velocity measurement, is suitable for complex scenarios, and reduces the incidence of safety accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120926942B_ABST
    Figure CN120926942B_ABST
Patent Text Reader

Abstract

The monocular vision-based target ranging and speed measuring method is suitable for measuring the distance and speed of a target in intelligent driving. A driving vehicle is provided with a collection device for collecting target images. The method comprises the following steps: a model of a multi-task network is constructed, the model can extract high-level semantic features of the target images and output target parameter values, the target parameter values comprise a detection frame of the target, physical width and height of the target and end-to-end ranging values of the target; a training label making strategy is constructed, real parameter values of the target in the images are obtained through a multi-modal mode, and are used for making a large amount of training data for training of the model of the multi-task network; the detection frame of the target and the physical width and height of the target in the target parameter values can obtain geometric ranging values of the target through a small aperture imaging principle; the geometric ranging values and the end-to-end ranging values are fused through a time sequence fusion method to obtain fusion ranging values and speed measuring values of the target, real-time ranging and speed measuring of the target can be realized, and the accuracy and robustness of the ranging and speed measuring can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of intelligent driving, specifically relating to a target ranging and speed measurement method based on monocular vision. Background Technology

[0002] In the field of intelligent driving, accurate ranging and speed measurement of targets is a crucial prerequisite for AEB (Autonomous Emergency Braking) functionality. It enables real-time monitoring of the dynamics of targets ahead to avoid collision risks, a function vital for ensuring traffic safety and improving traffic efficiency. Currently, ranging and speed measurement methods based on monocular vision have certain limitations, with relatively low accuracy in complex scenarios. Traditional ranging methods generally employ the pinhole imaging principle, which heavily relies on the accuracy of the target detection box; any jitter in the detection box will cause instability in the ranging measurement.

[0003] Patent CN111982072A proposes a monocular vision-based target ranging method. It first identifies the target's location region, then selects the midpoint of the lower edge as the target projection ranging reference point, and finally uses a ground point ranging method to obtain the target's distance. shaft and Distance along the axial direction. This method is essentially still a visual ranging method based on geometry. It cannot avoid the problem of not being able to obtain an accurate grounding point due to factors such as occlusion in complex scenes. In addition, this method also relies on the detection box area of ​​target recognition. When the detection box shakes, it will still cause instability in ranging.

[0004] Patent CN202410310655.2 also proposes a target ranging method based on monocular vision. It collects data, trains a model using the LightGBM machine learning method, and then uses this model to predict the target distance end-to-end. This method innovatively uses model-based end-to-end ranging technology. However, it employs a purely data-driven monocular vision depth estimation method without incorporating conventional geometric ranging methods such as the pinhole principle. Therefore, this method is easily limited by the training data, making it difficult to cover a wide range of scenes and thus unsuitable for truly complex real-world scenarios. Summary of the Invention

[0005] In view of this, the target ranging and velocity measurement method based on monocular vision provided by the present invention combines pinhole ranging, end-to-end ranging and time-series information to calculate target parameters, which can improve the accuracy and robustness of ranging and velocity measurement in AEB function.

[0006] To achieve the above-mentioned technical objectives, the specific technical solution adopted by the present invention is as follows:

[0007] A target ranging and velocity measurement method based on monocular vision, applicable to target ranging and velocity measurement in intelligent driving, wherein the driving vehicle is equipped with a target image acquisition device, and the method includes, S101: Construct a model for a multi-task network, which can extract high-level semantic features of the target image and output target parameter values, including the target detection box, the target's physical width and height, and the target's end-to-end ranging value. S102: Construct a training label production strategy to obtain the true parameter values ​​of the target in the image through a multimodal approach, which is used to produce a large amount of training data for training the multi-task network model; S103: The target detection box and the target physical width and height in the target parameter value can be used to obtain the geometric distance value of the target through the pinhole imaging principle. The geometric distance value and the end-to-end distance value are used to perform distance and velocity measurement on the target through a time-series fusion method.

[0008] Beneficial effects of the technology: A multi-task network is designed to simultaneously output target detection bounding boxes, target physical width, target physical height, and end-to-end distance, providing richer target information. An automatic annotation strategy is designed, employing different data generation methods based on millimeter-wave radar and lidar. This automatic annotation strategy can generate large quantities of training data labels for target distance, target physical width, and target physical height in general scenarios for model training. A temporal fusion ranging logic is designed, combining the advantages of pinhole ranging and end-to-end ranging with the temporal information of the target to obtain more robust target ranging values. Based on these ranging values, more stable velocity values ​​are derived, thereby improving the accuracy and stability of target ranging and velocity measurement. Attached Figure Description

[0009] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a flowchart illustrating the principle of target depth tag fabrication based on millimeter-wave radar in this invention. Figure 2 This is a flowchart illustrating the principle of target depth tag fabrication based on lidar in this invention. Figure 3 This is the overall flowchart of the present invention. Detailed Implementation

[0011] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0012] The following specific examples illustrate the implementation of this disclosure. Those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0013] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using other structures and / or functionalities besides one or more of the aspects set forth herein.

[0014] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this disclosure. The drawings only show the components related to this disclosure and are not drawn according to the number, shape and size of the components in actual implementation. In actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0015] like Figures 1 to 3 The target ranging and speed measurement method based on monocular vision shown is suitable for ranging and speed measurement of targets in intelligent driving. The driving vehicle is equipped with a data acquisition device to collect target images, and the method includes: S101: Construct a multi-task network model that can extract high-level semantic features of the target image and output target parameter values, including the target detection box, the target's physical width and height, and the target's end-to-end ranging value, etc. S102: Construct a training label production strategy to obtain the true parameter values ​​of the target in the image through a multimodal approach, which is used to produce a large amount of training data for training the multi-task network model; S103: The target detection box and the target physical width and height in the target parameter value can be used to obtain the geometric distance value of the target through the pinhole imaging principle. The geometric distance value and the end-to-end distance value are used to perform distance and velocity measurement on the target through a time-series fusion method.

[0016] By introducing a multi-task network for end-to-end ranging and a time-series fusion strategy, more accurate and stable target distance and velocity information can be obtained, making it suitable for target ranging and velocity measurement in different scenarios. This can effectively improve the performance of AEB functions such as collision risk warning and reduce the incidence of safety accidents.

[0017] In one specific embodiment, the model of S101 includes a backbone network, a detection branch network, and a deep output head network, wherein, The detection branch network includes a classification output head unit and a bounding box regression output head unit, and the depth output head network includes a target depth output head unit, a target width output head unit, and a target height output head unit; The backbone network receives the target image transmitted by the acquisition device, extracts the high-level semantic features of the target image in matrix form, and adapts and sends them to the detection branch network and the depth output head network. The classification output head unit is used to calculate the category of the target, and the box regression output head unit is used to calculate the detection box of the target; The target depth output head unit is used to calculate the end-to-end target distance, the target width output head unit is used to calculate the physical width of the target, and the target height output head unit is used to calculate the physical height of the target.

[0018] Furthermore, for the detection branch network, in order to train the network parameters, the detection branch network uses the cross-entropy loss function and the CIOU loss function to calculate the loss between the predicted values ​​output by the classification output head unit and the box regression output head unit and the true parameter values. The true parameter values ​​include the true labels of the categories. S301: Define the number of categories as follows The classification output head unit outputs a dimensional probability vector , For candidate boxes belonging to the first The probability of a class; the true label of a class is one. one-hot vector of dimension , For candidate boxes belonging to the first Class, and the remaining elements in the vector ; S302: The predicted value output by the classification output head unit is The cross-entropy loss function is used for class training. The cross-entropy loss function model is expressed as follows: ; S303: Define the predicted value output by the box regression output head unit as... Furthermore, the predicted bounding box is represented by the top-left vertex and the width and height of the box, and is trained using the CIoU loss function, which is expressed as: ,in, For prediction boxes With real frame The intersection over union (IoU) ratio, for and Distance from the center point This is the diagonal length of the smallest bounding box of the two boxes. For aspect ratio consistency parameters and , used to measure and The difference in aspect ratio, For adaptive weight coefficients and Used to balance aspect ratio loss. The impact; S304: For the depth branch, we use Laplacian Aleatoric Uncertainty Loss and L1 loss to measure the difference between the predicted and true values ​​of the "target depth output head" and the "target width and height output head," respectively. The ground truth for the target depth is defined as... Because the target depth has a large numerical range (from a few meters to hundreds of meters), directly using the target depth for loss calculation and gradient iteration can easily lead to training instability. Therefore, it is necessary to use the reciprocal of the predicted target depth for calculation. Specifically, Define the truth value of the target depth as Using the true depth of the target reciprocal Predict the target depth, i.e. The predicted value output by the target depth output head unit is defined as follows: The loss function obtained by training with the Laplacian stochastic uncertainty loss function is as follows: ,in, The log-variance of the model prediction is used to measure the uncertainty of the predicted target depth. The larger the value, the higher the uncertainty of the model prediction, and the larger the error term. The weighting coefficients will be weakened, then This itself will contribute more to the loss, so the model will focus more on learning reasonable target depth uncertainty during training; when The smaller the value, the lower the uncertainty of the model prediction, and the lower the error term. Weighting coefficients It will increase, meaning the error term will increase at this point. This will contribute more to the loss, so the model will focus more on reducing prediction error during training.

[0019] S305: Define the predicted values ​​output by the target width output head unit and the target height output head unit as follows: and The true values ​​of the target's width and height are and The model was trained using the L1-smooth loss function to obtain the model. and model , is represented as: ; ; S306: Superimpose the above loss functions to obtain the total loss function. : ,in, , , , , The weight coefficients for each loss function are generally set using empirical values ​​from experiments.

[0020] In one specific embodiment, the network structure designed in this patent includes a depth branch, which predicts the depth and physical width and height of the target. To generate high-quality training data in large quantities, we designed an automatic labeling method to create training labels based on different modalities of millimeter-wave radar and lidar. For example, if the acquisition device is a millimeter-wave radar, the construction of the training label model in S102 includes... The intrinsic and extrinsic parameter matrices of the vehicle's forward-looking camera, as well as the transformation matrix between the millimeter-wave radar and the forward-looking camera, are obtained through calibration. Static scene data with lane lines and general scene data are collected. The general scene data is used to create training data, including the determination of the default vanishing point and the general scene vanishing point. S401: For the collected static scene data with lane lines, the default vanishing point is calculated. This involves using a large lane line model (such as the LaneNet model) to obtain lane line semantic segmentation points, and then segmenting the lane line points at the pixel level. Fit to a linear equation: ,in, The number of lane lines detected; The default vanishing point is calculated using the least squares method. ; S402: Continuous frame data of a general scene acquired by millimeter-wave radar. Large-scale detection models (such as the Grounding-DINO model) and multi-target tracking models (such as the ByteTrack model) are used to obtain the bounding box information and tracking ID information of each target in the front view. A large-scale lane detection model (such as the LaneNet model) is used to obtain lane line information in each frame, retaining only the two lane lines closest to the left and right sides of the vehicle. and The formula for fitting lane lines is: , ; The calculated intersection of the two lines As the actual vanishing point of the current frame , is represented as: ; If the actual frames of the scene do not contain lane lines, the default vanishing point is used. If the actual frame of the scene contains lane lines, the actual vanishing point is used. Satisfies the default vanishing point At the preset threshold Within the range, when When using the current frame vanishing point Otherwise, use the default vanishing point. It's important to note that the general scenario differs from the static scenario. The static scenario is a scenario with lane lines that is manually selected for better road sections, so all lane lines can be used to calculate robust vanishing points as the default vanishing points. However, in the general scenario, considering factors such as occlusion by surrounding vehicles or edge distortion, only the two lane lines closest to the vehicle are selected to calculate the actual vanishing points.

[0021] Furthermore, the determination of the true parameter values ​​described in S102 includes, S501: The actual parameter values ​​include the target's physical width and height and the distance measurement; the calculation includes... Define the target point in the vehicle camera coordinate system The projection on the image plane is ,Right now, ,in, The intrinsic parameter matrix of the forward-looking camera. , and They are respectively direction and Focal length of direction, For the principal point; Define the physical width of the target as Physical height is And set the initial values ​​for the physical width and height of each category as follows: , Its specific value is derived from the statistical average of the physical width and height of a large amount of data on different types of vehicles; the pixel width of the detection box for the front or rear of the vehicle is defined as... The pixel height of the vehicle body detection frame is For consecutive frames of the front view image, the detection large model can determine the information of the vehicle body detection box and the front and rear detection boxes; Based on the principle of pinhole imaging, the preliminary visual distance of the target can be estimated.

[0022] or , The calculation formula can be based on the pixel width of the vehicle. and pixel car height To calculate. When the front or rear of the vehicle is visible and in consecutive frames. The former method is generally used when the change is stable; the latter method is used in other cases for calculation. The longitudinal distance values ​​of each target point obtained from the millimeter-wave radar are taken and compared with the preliminary visual distance. Fusion matching is performed on consecutive frames, and the matched targets are then matched using the distance values ​​from the millimeter-wave radar. Replace its original initial visual distance .

[0023] Based on the principle of pinhole imaging, the preliminary actual physical width of the target can be estimated. and actual physical height , represented as: or ; Physical height of the target Represented as: or , in, For the height of the camera, This is the y-coordinate value of the bottom edge of the target bounding box. Vanishing point The vertical coordinate value. Both of the above different calculation formulas can be derived from the principle of pinhole imaging. In practice, depending on the characteristics of different scenes (such as target crossing, target occlusion, etc.), the above different formulas can be selected to calculate the physical width of the target; S501: Because the pixel width of distant targets is small, even a small perturbation in pixel width can lead to a large deviation in the physical width of distant targets estimated using the above method. Therefore, correction based on temporal width is necessary. The principle is that the physical width of the same target remains constant as it moves from near to far. Therefore, we can utilize the characteristics of temporal frames to represent the physical width of distant targets using the physical width of nearby targets. The true parameter values ​​include the target depth label value. Therefore, the calculation includes... Define the threshold for long distance as (like For the physical width of a distant target, the target can be used in... It is represented by the physical vehicle width within the range. The logic is as follows: taking the process from the appearance to the disappearance of the target in consecutive frames as the target's lifecycle, the physical vehicle width of the same target in consecutive frames within the lifecycle is calculated and denoted as follows. , … ,in, Given the lifecycle length of the current target, the formula for correcting the width based on the time-series vehicle width is defined as follows: , in, These are weighting coefficients. For this goal in the life cycle The millimeter-wave radar range corresponding to the frame. The formula describes that "the farther the target is, the more participants..." Weighting coefficients in the calculation formula The rule of "the smaller the value" means that by introducing time sequence, "the physical vehicle width of the target at a long distance can be represented by the vehicle width at a close distance", which can greatly improve the problem of inaccurate estimation of the physical vehicle width of distant targets.

[0024] Similarly, the physical vehicle height correction value of the target is obtained. Based on the principle of pinhole imaging, the corrected visual ranging value is obtained. Represented as: or ; Corrected visual range value The target range value is obtained by matching the millimeter-wave radar value with the target range value, and the final target range value is used as the depth label value of the target.

[0025] In one specific embodiment, the acquisition device is a lidar, and the construction of the training label model in S102 includes, The intrinsic and extrinsic parameters of the vehicle's forward-looking camera and the transformation matrix from the LiDAR to the camera were obtained through calibration. Video data of a typical scene was collected, containing both forward-looking images and LiDAR point cloud data. A 4D automatic annotation model was used to obtain 3D bounding boxes and target category information based on the collected point cloud data. A detection model was also used to obtain target bounding boxes for each frame, and a segmentation model was used to obtain target segmentation results for each frame. S601: For 4D automatic annotation of large models, the input is point cloud data of each frame, and the output is 3D detection boxes. Target Category Target tracking ID number For large-scale detection models, the input is the front view image data, and the output is the 2D detection bounding boxes of each target in the current frame image. Target Category information; Considering that point cloud data may exceed the visual boundary on the front view image and retain a lot of information about fully occluded targets, a filtering strategy is added here to appropriately filter the results of 4D automatic annotation. The specific filtering measures are as follows: First, the 3D detection boxes output by the 4D automatic annotation large model are projected onto the front view image to remove 3D detection boxes that are not on the image; second, the minimum bounding rectangles corresponding to the 3D detection boxes are calculated, and 3D boxes with a large degree of overlap and a large distance between them are filtered out.

[0026] For the segmentation model, the input consists of the front view image and the 2D detection bounding boxes of the detection model. This segmentation model uses the 2D detection bounding boxes as cue regions, and its output is the semantic segmentation result of the corresponding instance target (such as car, bus, person, etc.) within the detection bounding box. From this, the mask matrix of the target instance can be constructed, thus obtaining the 2D detection bounding box of each target in each frame. The corresponding mask matrix.

[0027] S602: Actual parameter values ​​include the actual physical width and height of the target. Depth of target Its calculations include, The point cloud within the filtered 3D detection bounding box output by the 4D automatically annotated large model. Projecting onto the image yields a projected point cloud. Take the projected point cloud Minimum bounding rectangle The 2D bounding box of the target instance corresponding to the 3D detection bounding box is used as the target 2D bounding box on the front view image; the minimum bounding rectangle of the mask region of each target instance is taken from the mask matrix of the segmented large model. As a target 2D bounding box representing a target instance; Calculate the current frame and The intersection-union ratio is used to obtain the correspondence between the 2D and 3D detection boxes of the target in the front view image and to perform target matching based on the detection boxes; For a matched target, the width and height of the 3D detection box are used as the physical width of the current target. and physical height The depth value Z of the target is taken as the point closest to the forward-looking camera within the 3D detection bounding box, and is expressed as: , , , in, The width of the 3D bounding box that matches the target. The height of the 3D bounding box that matches the target. , , , Represents the four vertices of the bottom edge of the 3D frame. , , , The ordinate value.

[0028] Based on the above embodiments, the target parameter values ​​include end-to-end distance. S103 includes, The visual distance to a target can be easily calculated based on the pinhole principle. , is represented as , or ; Based on dynamic weight allocation and temporal filtering, combined with target features of the current scene and historical frame data, the end-to-end distance is adaptively adjusted. Visual distance The fusion ratio is adjusted to achieve robust ranging, where, S701: Fusion Strategy and Ranging, where, for the target detection bounding box in the front view, the percentage of overlap between the current target bounding box and other targets is calculated. ,like (Threshold, such as 0.3), is determined to be an occluded scene; Construct occlusion weights , is represented as , , in, For adjustment coefficients, The range is (0,1], and as the formula shows, the more severe the target occlusion, the higher the weight. The lower the value, the better; the model training uses Laplacian uncertainty loss, meaning that this branch predicts both the target depth and a log-variance. It describes the uncertainty of prediction depth and can be viewed as a value similar to a confidence score. When The larger the value, the greater the uncertainty it describes in terms of target depth, which means that the reliability of the predicted target depth value is lower. If... (Threshold, such as 0.6), determine the target as a target with low confidence distance.

[0029] Construct confidence weights , is represented as , , in, For adjustment coefficients, The range is (0,1], and σ is the logarithmic uncertainty of the deep branch prediction in the multi-task network model. Once the value exceeds the threshold, the larger the value, the higher the weight. The lower the value, the better.

[0030] S702: Defines the fusion distance of a single frame. , ,in, , For normalized weights; For the current frame, maintain the most recent Single-frame fusion distance Then, smoothing is performed to predict the distance of the current frame. or ,in, Indicates Kalman filtering, Indicates polynomial fitting; For the current frame, the fusion distance is obtained according to the formula for the single-frame fusion distance. , ; Calculate the fusion distance of the current frame Distance values ​​based on time series prediction deviation , ; S703: Constructing Trend Consistency Weights ,in For adjustment coefficients, The threshold for distance deviation is , when the distance deviation is... When the size is larger, the weight The smaller.

[0031] Define the fusion distance in time series , , That is, we get:

[0032] This fusion ranging strategy improves ranging stability by dynamically fusing geometric priors (i.e., visual distance) and data-driven predictions (i.e., end-to-end distance) and combining them with temporal information, and can effectively cope with target ranging in complex scenarios in intelligent driving.

[0033] S704: Velocity measurement based on time-series information, wherein the continuous ranging values ​​of the target in consecutive frames are represented as follows: ,in, For a fixed window length of consecutive frames in history, Indicates the previous number in the history window The ranging value of the target in a given frame. This indicates the ranging value of the target in the current frame; Based on the differential principle, the fundamental velocity value between adjacent frames is calculated. , , in, The difference in timestamps between two adjacent frames is represented by the speed. Due to factors such as noise, the speed calculated solely by the difference in distance between adjacent frames is subject to a certain degree of disturbance, which may cause the overall target speed to fluctuate significantly. Therefore, the speed calculated by the above formula is only used as the basic speed value. The smoothed speed value will be obtained by using the Kalman filter smoothing method below.

[0034] A Kalman filter model is established to approximate the target motion as a uniform or uniformly accelerated model, and the model state vector is... Its component elements represent distance, velocity, and acceleration, respectively, and the state equation is defined as: ,in, The state matrix, , for The state variable at time t, For process noise, Represented as a normal distribution, , , , The process noise variance representing distance, velocity, and acceleration; The observation equation is defined as follows: ,in, For the observation matrix, For the observed value of velocity, To observe noise; The state equation and observation equation are transformed into state prediction steps and state update steps, where, State prediction steps: , ; in, , These are the state prediction value and the covariance prediction value, respectively. For covariance; Status update steps: ; ; ; , , These are the Kalman gains. Status update value, covariance update value, It is the identity matrix; The iteration of the target velocity is obtained through the state prediction step. And converted into speed prediction values Then, the base velocity value calculated based on the distance difference in the current frame is used. Perform the above By updating the state during the filtering process, a smoother and more stable estimate of the target velocity can be calculated. .

[0035] The multi-task network constructed by this method can simultaneously output the target's detection bounding box, physical width and height, and depth. This allows it to combine the advantages of geometric vision ranging and end-to-end ranging with temporal information to obtain more accurate and stable ranging values. This method exhibits strong generalization ability, overcomes the limitations of single ranging methods, and is applicable to a wider range of scenarios.

[0036] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. A target ranging and speed measurement method based on monocular vision, applicable to target ranging and speed measurement in intelligent driving, wherein the driving vehicle is equipped with a target image acquisition device, characterized in that, Its methods include, S101: Construct a model for a multi-task network, which can extract high-level semantic features of the target image and output target parameter values, including the target detection box, the target's physical width and height, and the target's end-to-end ranging value. S102: Construct a training label creation strategy. Obtain the true parameter values ​​of targets in the image through a multimodal approach to create a large batch of training data for training the multi-task network model. Design a multimodal approach to create training label data to generate ground truth data for model training. This includes designing different combinations of millimeter-wave radar and cameras, as well as combinations of LiDAR and cameras. When the acquisition device is millimeter-wave radar, the multimodal label creation strategy includes obtaining the intrinsic and extrinsic parameter matrices of the vehicle's front-view camera through calibration, as well as the transformation matrix between the millimeter-wave radar and the front-view camera. Collect static scene data with lane lines and general scene data. The general scene data is used to create training data, including determining the default vanishing point and the vanishing point of the general scene. S401: For the collected static scene data with lane lines, the default vanishing point is calculated. This involves using a large lane line model to obtain lane line semantic segmentation points, and then dividing the pixel-level lane line points... Fit to a linear equation: ,in, The number of lane lines detected; The default vanishing point is calculated using the least squares method. ; S402: Continuous frame data of a general scene acquired by millimeter-wave radar. Large-scale detection and multi-target tracking models are used to obtain the bounding box information and tracking ID information of each target in the front view. A lane line detection model is used to obtain lane line information in each frame, retaining only the two lane lines closest to the vehicle's left and right sides. and The formula for fitting lane lines is: and ; The calculated intersection of the two lines As the actual vanishing point of the current frame , is represented as: ; If the actual frames of the scene do not contain lane lines, the default vanishing point is used. If the actual frame of the scene contains lane lines, the actual vanishing point is used. Satisfies the default vanishing point At the preset threshold Within the range, when When using the current frame vanishing point Otherwise, use the default vanishing point. ; S103: The target detection box and physical width and height in the target parameter values ​​can be used to obtain the geometric range value of the target through the pinhole imaging principle. The geometric range value and the end-to-end range value are then used to perform range and velocity measurement on the target through a time-series fusion method. The target parameter values ​​include the end-to-end distance. ,include, The visual distance to a target can be calculated based on the pinhole principle. , is represented as , or ; Based on dynamic weight allocation and temporal filtering, combined with target features of the current scene and historical frame data, the end-to-end distance is adaptively adjusted. Visual distance The fusion ratio is adjusted to achieve robust ranging, where, S701: Fusion Strategy and Ranging, where, for the target detection bounding box in the front view, the percentage of overlap between the current target bounding box and other targets is calculated. ,like This is determined to be an obstruction of the scene; Construct occlusion weights , is represented as , , in, For adjustment coefficients, The range is (0,1]; Construct confidence weights , is represented as , , in, For adjustment coefficients, The range is (0,1], and σ is the log uncertainty of the deep branch prediction in the multi-task network model; S702: Defines the fusion distance of a single frame. , ,in, , For normalized weights; For the current frame, maintain the most recent Single-frame fusion distance Then, smoothing is performed to predict the distance of the current frame. or ,in, Indicates Kalman filtering, Indicates polynomial fitting; For the current frame, the fusion distance is obtained according to the formula for the single-frame fusion distance. , ; Calculate the fusion distance of the current frame Distance values ​​based on time series prediction deviation , ; S703: Constructing Trend Consistency Weights ,in For adjustment coefficients, The threshold for the distance deviation; Define the fusion distance in time series ; S704: Velocity measurement based on time-series information, wherein the continuous ranging values ​​of the target in consecutive frames are represented as follows: ,in, For a fixed window length of consecutive frames in history, Indicates the previous number in the history window The ranging value of the target in a given frame. This indicates the ranging value of the target in the current frame; Based on the differential principle, the fundamental velocity value between adjacent frames is calculated. , , in, This represents the timestamp difference between two adjacent frames; A Kalman filter model is established to approximate the target motion as a uniform or uniformly accelerated model, and the model state vector is... Its component elements represent distance, velocity, and acceleration, respectively, and the state equation is defined as: ,in, The state matrix, , for The state variable at time t, For process noise, Represented as a normal distribution, , , , The process noise variance representing distance, velocity, and acceleration; The observation equation is defined as follows: ,in, For the observation matrix, For the observed value of velocity, To observe noise; The state equation and observation equation are transformed into state prediction steps and state update steps, where, State prediction steps: , , in, , These are the state prediction value and the covariance prediction value, respectively. For covariance; Status update steps: ; ; ; , , These are the Kalman gains. Status update value, covariance update value, It is the identity matrix; The iteration of the target velocity is obtained through the state prediction step. And converted into speed prediction values Then, the base velocity value calculated based on the distance difference in the current frame is used. Perform the above By updating the state during the filtering process, a smoother and more stable estimate of the target velocity can be calculated. .

2. The target ranging and velocity measurement method according to claim 1, characterized in that, The model in S101 includes a backbone network, a detection branch network, and a deep output head network. The detection branch network includes a classification output head unit and a bounding box regression output head unit, and the depth output head network includes a target depth output head unit, a target width output head unit, and a target height output head unit; The backbone network receives the target image transmitted by the acquisition device, extracts the high-level semantic features of the target image in matrix form, and adapts and sends them to the detection branch network and the depth output head network. The classification output head unit is used to calculate the category of the target, and the box regression output head unit is used to calculate the detection box of the target; The target depth output head unit is used to calculate the end-to-end target distance, the target width output head unit is used to calculate the physical width of the target, and the target height output head unit is used to calculate the physical height of the target.

3. The target ranging and velocity measurement method according to claim 2, characterized in that, The detection branch network uses the cross-entropy loss function and the CIOU loss function to calculate the loss between the predicted values ​​output by the classification output head unit and the box regression output head unit and the true parameter values. The true parameter values ​​include the true labels of the categories. S301: Define the number of categories as follows The classification output head unit outputs a dimensional probability vector , For candidate boxes belonging to the first The probability of a class; the true label of a class is one. one-hot vector of dimension , For candidate boxes belonging to the first Class, and the remaining elements in the vector ; S302: The predicted value output by the classification output head unit is The cross-entropy loss function is used for class training. The cross-entropy loss function model is expressed as follows: ; S303: Define the predicted value output by the box regression output head unit as... Furthermore, the predicted bounding box is represented by its top-left vertex and width and height, and is trained using the CIoU loss function, which is expressed as: ,in, For prediction boxes With real frame The intersection and union ratio, for and Distance from the center point This is the diagonal length of the smallest bounding box of the two boxes. For aspect ratio consistency parameters and , For adaptive weight coefficients and ; S304: Define the true value of the target depth as Using the true depth of the target reciprocal Predict the target depth, that is, The predicted value output by the target depth output head unit is defined as follows: The Laplacian stochastic uncertainty loss function is used for training to obtain... ,in, The log-variance of the model predictions; S305: Define the predicted values ​​output by the target width output head unit and the target height output head unit as follows: and The true values ​​of the target's width and height are and The model was trained using the L1-smooth loss function to obtain the model. and model , is represented as: , ; S306: Superimpose the above loss functions to obtain the total loss function. : ,in, , , , , These are the weighting coefficients corresponding to each loss function.

4. The target ranging and velocity measurement method according to claim 3, characterized in that, The determination of the true parameter values ​​described in S102 includes, S501: The actual parameter values ​​include the target's physical width and height and the distance measurement; the calculation includes... Define the target point in the vehicle camera coordinate system The projection on the image plane is ,Right now, ,in, The intrinsic parameter matrix of the forward-looking camera. , and They are respectively direction and Focal length of direction, For the principal point; Define the physical width of the target as Physical height is And set the initial values ​​for the physical width and height of each category as follows: , Define the pixel width of the detection bounding box for the front or rear of the vehicle as... The pixel height of the vehicle body detection frame is For consecutive frames of the front view image, the detection large model can determine the information of the vehicle body detection box and the front and rear detection boxes; Based on the principle of pinhole imaging, the preliminary visual distance of the target can be estimated. : or ; Take the longitudinal distance value of each target point obtained by the millimeter-wave radar and compare it with the preliminary visual distance. Fusion matching is performed on consecutive frames, and the matched targets are then matched using the distance values ​​from the millimeter-wave radar. Replace its original initial visual distance ; Based on the principle of pinhole imaging, the preliminary actual physical width of the target can be estimated. and actual physical height , represented as: or ; Physical height of the target Represented as: or , in, For the height of the camera, This is the y-coordinate value of the bottom edge of the target bounding box. Vanishing point The ordinate value; S501: The true parameter values ​​include the target depth label values, and the calculation includes... Define the threshold for long distance as The lifecycle of a target is defined as the process from its appearance to its disappearance within consecutive frames. The physical vehicle width of the same target within this lifecycle in consecutive frames is calculated and denoted as follows: , … ,in, Given the lifecycle length of the current target, the formula for correcting the width based on the time-series vehicle width is defined as follows: , in, These are weighting coefficients. For this goal in the life cycle The millimeter-wave radar range corresponding to the frame; Similarly, the physical vehicle height correction value of the target is obtained. Based on the principle of pinhole imaging, the corrected visual ranging value is obtained. Represented as: or ; Corrected visual range value The target range value is obtained by matching the millimeter-wave radar value with the target range value, and the final target range value is used as the depth label value of the target.

5. The target ranging and velocity measurement method according to claim 4, characterized in that, When the data acquisition device is a lidar, the multimodal tag creation strategy in S102 includes: The intrinsic and extrinsic parameters of the vehicle's forward-looking camera and the transformation matrix from the LiDAR to the camera were obtained through calibration. Video data of a typical scene was collected, containing both forward-looking images and LiDAR point cloud data. A 4D automatic annotation model was used to obtain 3D bounding boxes and target category information based on the collected point cloud data. A detection model was also used to obtain target bounding boxes for each frame, and a segmentation model was used to obtain target segmentation results for each frame. S601: For 4D automatic annotation of large models, the input is point cloud data of each frame, and the output is 3D detection boxes. Target Category Target tracking ID number For large-scale detection models, the input is the front view image data, and the output is the 2D detection bounding boxes of each target in the current frame image. Target Category information; S602: Actual parameter values ​​include the actual physical width and height of the target. Depth of target Its calculations include, The point cloud within the filtered 3D detection bounding box output by the 4D automatically annotated large model. Projecting onto the image yields a projected point cloud. Take the projected point cloud Minimum bounding rectangle The 2D bounding box of the target instance corresponding to the 3D detection bounding box is used as the target 2D bounding box on the front view image; the minimum bounding rectangle of the mask region of each target instance is taken from the mask matrix of the segmented large model. As a target 2D bounding box representing a target instance; Calculate the current frame and The intersection-union ratio is used to obtain the correspondence between the 2D and 3D detection boxes of the target in the front view image and to perform target matching based on the detection boxes; For a matched target, the width and height of the 3D detection box are used as the physical width of the current target. and physical height The depth value of the target is taken as the point closest to the front-view camera within the 3D detection bounding box. , is represented as: , , , in, The width of the 3D bounding box that matches the target. The height of the 3D bounding box that matches the target. , , , Represents the four vertices of the bottom edge of the 3D frame. , , , The ordinate value.