Roadside monocular camera real-time vehicle speed extraction method based on MF fusion algorithm

By applying the MF fusion algorithm in roadside monocular camera video, combined with depth estimation calculation method and optical flow method, the difficulty in extracting vehicle speed caused by picture occlusion and distortion is solved, and the accurate calculation of vehicle speed and the real-time and reliability of traffic flow monitoring are achieved.

CN119941799AActive Publication Date: 2025-05-06CHANGAN UNIV +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510016692.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

In the field of traffic monitoring, it is difficult to extract vehicle speed due to screen occlusion and image distortion in roadside monocular camera videos, and traditional internal reference correction methods are difficult to adapt to the variable viewing angles and distortion characteristics.

Method used

Using the method based on the MF fusion algorithm, the depth estimation calculation method is performed through the MiDaS model to predict the depth information of each frame of the image, and the motion information of pixels in adjacent frames is captured by the Farneback optical flow method to calculate the actual speed of the vehicle.

Benefits of technology

Without relying on camera internal reference, the accurate calculation of vehicle speed is achieved, which significantly improves the real-time and reliability of roadside traffic flow monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941799A_ABST
    Figure CN119941799A_ABST
Patent Text Reader

Abstract

The invention discloses a roadside monocular camera real-time vehicle speed extraction method based on an MF fusion algorithm, and the method comprises the following steps: S1, installing a monocular camera at a roadside, and collecting a road video image; s2, carrying out framing processing on the road video image to form a frame-by-frame image; s3, performing a monocular depth estimation algorithm through the MiDaS model; s4, obtaining two adjacent frame-by-frame images It and It + 1, and calculating the optical flow between the adjacent images by using a Farneback optical flow method; and S5, combining the optical flow vector with the depth map, and calculating the pixel speed to obtain the traffic flow running speed. According to the roadside monocular camera real-time vehicle speed extraction method based on the MF fusion algorithm, the pixel depth of each frame of image is predicted through a depth estimation algorithm, and the motion information of pixels in adjacent frames is captured in combination with an optical flow method, so that the actual speed of the vehicle is calculated; and the real-time performance and the reliability of roadside traffic flow monitoring are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of road monitoring, and in particular to a method for extracting real-time vehicle speed from a roadside monocular camera based on an MF fusion algorithm. Background Art

[0002] In the field of traffic monitoring, especially for roadside monitoring systems, vehicle speed extraction faces multiple technical challenges that affect the accuracy and reliability of traffic data. First, roadside monitoring videos often encounter severe image occlusion, especially during high traffic hours, when the overlap of other vehicles makes it difficult to identify and track specific vehicles. This occlusion makes it difficult for the monitoring system to effectively assess the actual speed of each vehicle, which in turn affects the comprehensive analysis of traffic flow.

[0003] Secondly, image distortion in surveillance videos is often unavoidable due to the camera’s installation position and viewing angle. This distortion manifests itself as nonlinear characteristics in roadside surveillance, affecting the target’s depth estimation and velocity calculation. Although camera intrinsic correction can often be used to correct image distortion, in practice, due to the particularity of roadside settings, traditional intrinsic correction methods often cannot accurately adapt to the changing viewing angles and distortion characteristics. Especially when using wide-angle lenses, the degree of distortion can vary significantly depending on the viewing angle, resulting in poor correction results. In addition, the lack of fixed reference objects in the surveillance scene limits the effective calibration process and makes it difficult to obtain camera intrinsic parameters.

[0004] Therefore, extracting accurate speed information from monocular surveillance videos has become an urgent task. To overcome these challenges, developing new technologies that can handle nonlinear distortion in dynamic scenes and directly extract vehicle speed data from videos has become a focus of current research. These technologies need to fully consider the unique environment of roadside surveillance to ensure that reliable traffic data analysis can still be provided in complex situations. Summary of the invention

[0005] The purpose of the present invention is to provide a real-time vehicle speed extraction method for a roadside monocular camera based on the MF fusion algorithm, which can realize the accurate calculation of the vehicle speed without relying on the camera's internal parameters, predict the pixel depth of each frame image through the depth estimation algorithm, and capture the motion information of pixels in adjacent frames in combination with the optical flow method, so as to calculate the actual speed of the vehicle.

[0006] The present invention provides a method for extracting real-time vehicle speed from a roadside monocular camera based on an MF fusion algorithm, comprising the following steps:

[0007] S1: Install a monocular camera on the roadside to collect road video images;

[0008] S2: Processing the road video image by frame division to form frame-by-frame images;

[0009] S3: Use the MiDaS model to perform a monocular depth estimation algorithm to predict the depth information of each frame-by-frame image and obtain a depth map of each frame-by-frame image;

[0010] S4: Obtain two adjacent frame-by-frame images I t and I t+1 , use the Farneback optical flow method to calculate the optical flow between adjacent images, analyze the pixel changes between adjacent frame-by-frame images, and generate the optical flow vector of each pixel;

[0011] S5: By combining the optical flow vector with the depth map, the vehicle running speed is obtained by calculating the pixel speed.

[0012] Preferably, in step S3,

[0013] The frame-by-frame images are adjusted to the size H*W required by the MiDaS model; the adjusted frame-by-frame images I prep It is expressed as:

[0014] I prep =Resize(I,H,W);

[0015] Among them, I prep Represents the adjusted frame-by-frame image; Resize means adjusting the image to the specified height H and width W, keeping the number of channels I of the image unchanged; I: represents the number of channels of the input image, representing the three color channels of the RGB image; H: represents the height of the image; W: represents the width of the image;

[0016] For the adjusted frame-by-frame image I prep Normalization is performed, and the normalized frame-by-frame image I norm It is expressed as:

[0017]

[0018] Among them, I norm represents the normalized frame-by-frame image; μ is the mean used in training;

[0019] The normalized frame-by-frame image I norm Input the pre-trained MiDaS model M, and use the convolutional neural network CNN to extract the normalized frame-by-frame image I norm Extract features and predict the depth information of each pixel. The output of the MiDaS model M is a depth map D:

[0020] D=M(I norm );

[0021] Where D is the depth map; M is the pre-trained MiDaS model;

[0022] Scale the depth map D back to the original image size H orgin *W orgin , the final depth map D final :

[0023] Among them, H orgin is the original image height; W orgin is the original image width;

[0024] Use color mapping to adjust the final depth map D final For visualization, map the depth values ​​into color space.

[0025] Preferably, in step S3, the normalized frame-by-frame images I are matched. norm Each pixel in the final depth map D final The depth value information is used to calculate the distance of the detected target relative to the camera:

[0026] Distance=D final (x,y);

[0027] Wherein, Distance is the distance information of the detection target to the monocular camera; x represents the horizontal coordinate of the pixel in the image; y represents the vertical coordinate of the pixel in the image.

[0028] Preferably, in step S4,

[0029] Calculate frame-by-frame image I t The spatial gradient and temporal gradient of the frame-by-frame image I t Use the convolution operation to calculate the horizontal and vertical gradients and get I x and I y ;

[0030]

[0031] Among them, I x Indicates that the convolution operation calculates the gradient in the horizontal direction; I y Indicates that the convolution operation calculates the gradient in the vertical direction;

[0032] Calculate adjacent frame-by-frame images I t and frame-by-frame images I t+1 The time gradient between

[0033]

[0034] in, Represents adjacent frame-by-frame images I t and frame-by-frame images I t+1 The time gradient between

[0035] A fixed-size window is selected as the neighborhood window for each pixel;

[0036] The brightness of the pixels in the selected neighborhood window is fitted by the least square method;

[0037] P(x,y)=a0+a1x+a2y+a3x 2 +a4y 2 +a5xy;

[0038] Among them, P(x,y) represents the brightness value of the pixel at position (x,y) in the image, that is, the grayscale intensity of the pixel; a0 is a constant term, which represents the average brightness of the pixels in the neighborhood; a1 is the linear change coefficient in the x direction, which represents the change in brightness in the horizontal direction; a2 is the linear change coefficient in the y direction, which represents the change in brightness in the vertical direction; a3 is the quadratic term coefficient in the x direction, which represents the curved change in brightness in the horizontal direction; a4 is the quadratic term coefficient in the y direction, which represents the curved change in brightness in the vertical direction; a5 is the mixed term coefficient, which represents the interactive effect of brightness in the horizontal and vertical directions.

[0039] Preferably, in step S2, an optical flow equation for each pixel is constructed according to the optical flow constraint:

[0040]

[0041] Among them, u represents the component of the optical flow in the horizontal direction, which usually represents the horizontal movement speed of the pixel in the image; v represents the component of the optical flow in the vertical direction, which usually represents the vertical movement speed of the pixel in the image;

[0042] Iteratively solve the motion vector for each pixel and update the estimated value using the least squares method until convergence;

[0043] Summarize the motion vector of each pixel into a velocity field;

[0044] V(x,y)=(μ(x,y),v(x,y));

[0045] Among them, V(x,y) represents the optical flow vector at the position (x,y) in the image, that is, the motion vector of the pixel; u(x,y) represents the horizontal optical flow component at the position (x,y) in the image; v(x,y) represents the vertical optical flow component at the position (x,y) in the image.

[0046] Preferably, in step S5, the parallax displacement is obtained by the magnitude of the optical flow vector; the parallax displacement of the pixel point is expressed as the modulus of the optical flow vector:

[0047]

[0048] Preferably, in step S5, the frame rate of the camera is f, and the time interval between adjacent frames is

[0049] The actual displacement Δs(x,y) is calculated by multiplying the pixel disparity displacement by the depth value:

[0050] Δs(x,y)=d(x,y)·D t (x,y);

[0051] Among them, D t (x, y) represents the depth change of the pixel point (x, y), that is, the change in depth or distance of the position over time;

[0052] The speed of each pixel is obtained by dividing the actual displacement by the time interval:

[0053]

[0054] Therefore, the present invention adopts the above-mentioned method for real-time vehicle speed extraction of a roadside monocular camera based on the MF fusion algorithm to achieve accurate calculation of vehicle speed without relying on camera internal parameters, predicts the pixel depth of each frame image through a depth estimation algorithm, and combines the optical flow method to capture the motion information of pixels in adjacent frames to calculate the actual speed of the vehicle, thereby achieving accurate extraction of vehicle speed under nonlinear distortion conditions, and significantly improving the real-time and reliability of roadside traffic flow monitoring.

[0055] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 It is a schematic diagram of the overall structure of a method for extracting real-time vehicle speed from a roadside monocular camera based on the MF fusion algorithm of the present invention;

[0057] Figure 2 It is an image depth map corresponding to the image frame generated by the MiDaS depth estimation algorithm of the roadside monocular camera real-time vehicle speed extraction method based on the MF fusion algorithm of the present invention;

[0058] Figure 3 The present invention provides a flow chart for calculating optical flow vectors using the Farneback optical flow method of a roadside monocular camera real-time vehicle speed extraction method based on the MF fusion algorithm. DETAILED DESCRIPTION

[0059] The technical solution of the present invention is further described below through the accompanying drawings and embodiments.

[0060] Unless otherwise defined, technical or scientific terms used in the present invention shall have the common meanings understood by one having ordinary skills in the field to which the present invention belongs.

[0061] The words "first", "second" and similar terms used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. Words such as "include" or "comprises" and similar terms mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connect" or "connected" and similar terms are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0062] Embodiment 1

[0063] like Figure 1-Figure 3 As shown, the present invention provides a method for extracting real-time vehicle speed from a roadside monocular camera based on an MF fusion algorithm, comprising the following steps:

[0064] S1: A monocular camera is installed on the roadside to cover the lane area to be monitored and collect road video images; the road video images collected by the monocular camera will be used for subsequent image analysis;

[0065] S2: Processing the road video image by frame division to form frame-by-frame images;

[0066] According to the road video image frame rate of the monocular camera (e.g., 30 frames per second), the road video image data is decomposed into separate frame-by-frame images. For a 1-second road video image, 30 frame-by-frame images (e.g., Frame1, Frame2, Frame3...Frame30) will be obtained;

[0067] S3: Use the MiDaS model to perform a monocular depth estimation algorithm to predict the depth information of each frame-by-frame image and obtain a depth map of each frame-by-frame image;

[0068] S31: resize the frame-by-frame images to the size H*W required by the MiDaS model;

[0069] Each frame-by-frame image (e.g., Frame1) is resized to 384×384 pixels, which is the standard input size required by the model. The resized image is denoted as I prep ,The standardized size allows the model to maintain high depth estimation accuracy at different resolutions.

[0070] Adjusted frame-by-frame image I prep It can be expressed as:

[0071] I prep =Resize(I,H,W);

[0072] Among them, I prep Represents the adjusted frame-by-frame image; Resize means adjusting the image to the specified height H and width W, keeping the number of channels I of the image unchanged; I: represents the number of channels of the input image, representing the three color channels of the RGB image; H: represents the height of the image; W: represents the width of the image;

[0073] S32: Analyzing the adjusted frame-by-frame image I prep Normalize the pixel values ​​from [0,255] to [0,1] and subtract the mean (according to the training data); the normalized frame-by-frame image I norm It is expressed as:

[0074]

[0075] Among them, I norm represents the normalized frame-by-frame image; μ is the mean used in training;

[0076] The normalized frame-by-frame image I norm Input the pre-trained MiDaS model M, and use the convolutional neural network CNN to extract the normalized frame-by-frame image I norm The output of the MiDaS model M is a depth map D, which contains the relative depth information of each pixel in the image:

[0077] D=M(I norm );

[0078] Where D is the depth map; M is the pre-trained MiDaS model;

[0079] S34: Scale the depth map D back to the size H of the original image orgin *W orgin , so that the depth data can correspond to the pixels of the original image to the final depth map D final :

[0080] D final =Resize(D,H orgin ,W orgin );

[0081] Among them, H orgin is the original image height; W orgin is the original image width;

[0082] S35: Use color mapping to adjust the final depth map D finalFor visualization, map the depth values ​​into color space.

[0083] S36: Matching the normalized frame-by-frame image I norm Each pixel in the final depth map D final Depth value information, calculate the distance information of the object:

[0084] Distance=D final (x,y);

[0085] Wherein, Distance is the distance information of the detection target to the monocular camera; x represents the horizontal coordinate of the pixel in the image; y represents the vertical coordinate of the pixel in the image.

[0086] S4: Obtain two adjacent frame-by-frame images I t and I t+1 , the Farneback optical flow method is used to calculate the optical flow between adjacent images, analyze the pixel changes between adjacent frame-by-frame images, and generate the optical flow vector of each pixel, thereby capturing the motion information of the object in the image;

[0087] S41: Calculate frame-by-frame image I t The spatial gradient and temporal gradient of the frame-by-frame image I t Use the convolution operation to calculate the horizontal and vertical gradients and get I x and I y :

[0088]

[0089] Among them, I x Indicates that the convolution operation calculates the gradient in the horizontal direction; I y Indicates that the convolution operation calculates the gradient in the vertical direction;

[0090] S42: Calculate adjacent frame-by-frame images I t and frame-by-frame images I t+1 The time gradient between A fixed-size window is selected as the neighborhood window for each pixel;

[0091]

[0092] in, Represents adjacent frame-by-frame images I t and frame-by-frame images I t+1 The time gradient between

[0093] S43: Perform brightness fitting on the pixels in the selected neighborhood window by the least square method to obtain a quadratic polynomial model, corresponding to coefficient a=(a0,a1,a2,a3,a4,a5) T ;

[0094] P(x,y)=a0+a1x+a2y+a3x 2 +a4y 2 +a5xy;

[0095] Among them, P(x,y) represents the brightness value of the pixel at position (x,y) in the image, that is, the grayscale intensity of the pixel; a0 is a constant term, which represents the average brightness of the pixels in the neighborhood; a1 is the linear change coefficient in the x direction, which represents the change in brightness in the horizontal direction; a2 is the linear change coefficient in the y direction, which represents the change in brightness in the vertical direction; a3 is the quadratic term coefficient in the x direction, which represents the curved change in brightness in the horizontal direction; a4 is the quadratic term coefficient in the y direction, which represents the curved change in brightness in the vertical direction; a5 is the mixed term coefficient, which represents the interactive effect of brightness in the horizontal and vertical directions.

[0096] S44: According to the optical flow constraint, construct the optical flow equation for each pixel:

[0097]

[0098] Where [u,v] T Represents the optical flow (or motion) velocity vector of each pixel; u represents the horizontal component of the optical flow, which usually represents the horizontal motion velocity of the pixel in the image; v represents the vertical component of the optical flow, which usually represents the vertical motion velocity of the pixel in the image;

[0099] S45: Iteratively solve the motion vector (u, v) for each pixel point, and update the estimated value using the least squares method until convergence;

[0100] S46: Summarize the motion vector (u, v) of each pixel into a velocity field V (x, y) to represent the motion information of each point in the image.

[0101] V(x,y)=(μ(x,y),v(x,y));

[0102] Among them, V(x,y) represents the optical flow vector at the position (x,y) in the image, that is, the motion vector of the pixel; u(x,y) represents the horizontal optical flow component at the position (x,y) in the image; v(x,y) represents the vertical optical flow component at the position (x,y) in the image.

[0103] S5: By combining the optical flow vector with the depth map, the pixel speed is calculated to obtain the vehicle running speed;

[0104] S51: The parallax displacement is obtained by the magnitude of the optical flow vector.

[0105] The parallax displacement d(x,y) of a pixel (x,y) is expressed as the modulus of the optical flow vector:

[0106]

[0107] S52: Assuming the frame rate of the camera is f (i.e., the number of frames per second), the time interval between adjacent frames is Assuming the frame rate is 30FPS, the time interval between adjacent frames is 0.03s;

[0108] S53: Calculate the actual displacement Δs(x,y) by multiplying the pixel disparity displacement by the depth value:

[0109] Δs(x,y)=d(x,y)·D t (x,y);

[0110] Among them, D t (x, y) represents the depth change of the pixel point (x, y), that is, the change in depth or distance of the position over time;

[0111] S54: The velocity v(x, y) of each pixel is obtained by dividing the actual displacement Δs(x, y) by the time interval Δt:

[0112]

[0113] Therefore, the present invention adopts the above-mentioned method for real-time vehicle speed extraction of a roadside monocular camera based on the MF fusion algorithm to achieve accurate calculation of vehicle speed without relying on camera internal parameters, predicts the pixel depth of each frame image through a depth estimation algorithm, and combines the optical flow method to capture the motion information of pixels in adjacent frames to calculate the actual speed of the vehicle, thereby achieving accurate extraction of vehicle speed under nonlinear distortion conditions, and significantly improving the real-time and reliability of roadside traffic flow monitoring.

[0114] The above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.

Claims

1. A method for extracting real-time vehicle speed from a roadside monocular camera based on MF fusion algorithm, characterized in that: The following steps are involved: S1: Install a monocular camera on the roadside to collect road video images; S2: Processing the road video image by frame division to form frame-by-frame images; S3: Use the MiDaS model to perform a monocular depth estimation algorithm to predict the depth information of each frame-by-frame image and obtain a depth map of each frame-by-frame image; S4: Obtain two adjacent frame-by-frame images I t and I t+1 , use the Farneback optical flow method to calculate the optical flow between adjacent images, analyze the pixel changes between adjacent frame-by-frame images, and generate the optical flow vector of each pixel; S5: By combining the optical flow vector with the depth map, the vehicle running speed is obtained by calculating the pixel speed.

2. According to the method for extracting real-time vehicle speed from a roadside monocular camera based on the MF fusion algorithm of claim 1, it is characterized in that: In step S3, the frame-by-frame image is adjusted to the size H*W required by the MiDaS model; the adjusted frame-by-frame image I prep It is expressed as: I prep =Resize(I,H,W); Among them, I prep Represents the adjusted frame-by-frame image; Resize means adjusting the image to the specified height H and width W, keeping the number of channels I of the image unchanged; I: represents the number of channels of the input image, representing the three color channels of the RGB image; H: represents the height of the image; W: represents the width of the image; For the adjusted frame-by-frame image I prep Normalization is performed, and the normalized frame-by-frame image I norm It is expressed as: Among them, I norm represents the normalized frame-by-frame image; μ is the mean used in training; The normalized frame-by-frame image I norm Input the pre-trained MiDaS model M, and use the convolutional neural network CNN to extract the normalized frame-by-frame image I norm Extract features and predict the depth information of each pixel. The output of the MiDaS model M is a depth map D: D=M(I norm ); Where D is the depth map; M is the pre-trained MiDaS model; Scale the depth map D back to the original image size H orgin *W orgin , get the final depth map D final ; Among them, H orgin is the original image height; W orgin is the original image width; Use color mapping to adjust the final depth map D final For visualization, map the depth values ​​into color space.

3. The method for extracting real-time vehicle speed from a roadside monocular camera based on the MF fusion algorithm according to claim 2 is characterized in that: In step S3, the normalized frame-by-frame image I is matched norm Each pixel in the final depth map D final The depth value information is used to calculate the distance of the detected target relative to the camera: Distance=D final (x,y); Wherein, Distance is the distance information of the detection target to the monocular camera; x represents the horizontal coordinate of the pixel in the image; y represents the vertical coordinate of the pixel in the image.

4. The method for extracting real-time vehicle speed from a roadside monocular camera based on the MF fusion algorithm according to claim 1 is characterized in that: In step S4, the frame-by-frame image I is calculated t The spatial gradient and temporal gradient of the frame-by-frame image I t Use the convolution operation to calculate the horizontal and vertical gradients and get I x and I y : Among them, I x Indicates that the convolution operation calculates the gradient in the horizontal direction; I y Indicates that the convolution operation calculates the gradient in the vertical direction; in, Represents adjacent frame-by-frame images I t and frame-by-frame images I t+1 The time gradient between A fixed-size window is selected as the neighborhood window for each pixel; The brightness of the pixels in the selected neighborhood window is fitted by the least square method; P(x,y)=a0+a1x+a2y+a3x 2 +a4y 2 +a5xy; Among them, P(x,y) represents the brightness value of the pixel at position (x,y) in the image, that is, the grayscale intensity of the pixel; a0 is a constant term, which represents the average brightness of the pixels in the neighborhood; a1 is the linear change coefficient in the x direction, which represents the change in brightness in the horizontal direction; a2 is the linear change coefficient in the y direction, which represents the change in brightness in the vertical direction; a3 is the quadratic term coefficient in the x direction, which represents the curved change in brightness in the horizontal direction; a4 is the quadratic term coefficient in the y direction, which represents the curved change in brightness in the vertical direction; a5 is the mixed term coefficient, which represents the interactive effect of brightness in the horizontal and vertical directions.

5. The method for extracting real-time vehicle speed from a roadside monocular camera based on the MF fusion algorithm according to claim 4 is characterized in that: In step S2, the optical flow equation for each pixel is constructed according to the optical flow constraint: Among them, u represents the component of the optical flow in the horizontal direction, which usually represents the horizontal movement speed of the pixel in the image; v represents the component of the optical flow in the vertical direction, which usually represents the vertical movement speed of the pixel in the image; Iteratively solve the motion vector for each pixel and update the estimated value using the least squares method until convergence; Summarize the motion vector of each pixel into a velocity field; V(x,y)=(μ(x,y),v(x,y)); Among them, V(x,y) represents the optical flow vector at the position (x,y) in the image, that is, the motion vector of the pixel; u(x,y) represents the horizontal optical flow component at the position (x,y) in the image; v(x,y) represents the vertical optical flow component at the position (x,y) in the image.

6. The method for extracting real-time vehicle speed from a roadside monocular camera based on the MF fusion algorithm according to claim 1, characterized in that: In step S5, the parallax displacement is obtained by the magnitude of the optical flow vector; the parallax displacement of the pixel point is expressed as the modulus of the optical flow vector:

7. The method for extracting real-time vehicle speed from a roadside monocular camera based on the MF fusion algorithm according to claim 6, characterized in that: In step S5, the frame rate of the camera is f, and the time interval between adjacent frames is The actual displacement Δs(x,y) is calculated by multiplying the pixel disparity displacement by the depth value: Δs(x,y)=d(x,y)·D t (x,y); Among them, D t (x,y) represents the depth change of pixel (x,y); The speed of each pixel is obtained by dividing the actual displacement by the time interval:

Citation Information

Patent Citations

  • Monocular vision odometer method fusing edge features and deep learning

    CN111311666A

  • Monocular vision inertial positioning method for automatic driving in closed park

    CN113436261A

  • Speed information acquisition method and device, equipment and medium

    CN113450579A

  • Step speed limiting method for multi-lane expressway exit

    CN116052446A

  • Train-mounted monocular vision speed measurement method

    CN117074713A