Method, device, equipment and storage medium for determining pitch angle of in-vehicle camera
Through the combination of optical flow estimation model and coordinate values, the focal length and pitch angle of the vehicle-mounted camera are automatically calculated, which solves the accuracy problem caused by manual measurement and achieves high-precision pitch angle determination.
Patent Information
- Application Number
- CN202210580410.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-26
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-05-26
AI Technical Summary
In the prior art, the determination of the pitch angle of the on-board camera relies on manual measurement, resulting in low accuracy and is greatly affected by the proficiency of the technician and measurement errors.
By obtaining two adjacent video images taken by the on-board camera, the focal length is calculated using the optical flow estimation model, and the pitch angle is automatically determined based on the coordinate values of the optical center and the vanishing point.
It improves the accuracy of the pitch angle of the on-board camera, reduces the error of manual operation, and realizes automated and high-precision pitch angle determination.
Smart Images

Figure CN115018926B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and particularly to a method, device, equipment and storage medium for determining the pitch angle of an in-vehicle camera. Background Art
[0002] Currently, fields such as autonomous driving and assisted driving are becoming increasingly popular, and the calibrated parameters of the in-vehicle cameras installed on vehicles are very important in these fields. For example, the calibrated parameters of the in-vehicle cameras are required for tasks such as lane line detection, vehicle tracking and ranging. The calibrated parameters of the in-vehicle cameras generally include the heading angle, pitch angle and roll angle of the in-vehicle camera. Among them, the pitch angle of the in-vehicle camera is the most important.
[0003] In the related art, after installing the in-vehicle camera, a reference object marked with a reference line is placed in front of the vehicle. Then, a technician finds a target position on the horizontal ground and measures the distance between the target position and the in-vehicle camera, the distance between the reference object and the in-vehicle camera, the height of the reference line, and the height of the in-vehicle camera. Finally, the pitch angle of the in-vehicle camera is calculated based on the measured distance between the target position and the in-vehicle camera, the distance between the reference object and the in-vehicle camera, the height of the reference line, and the height of the in-vehicle camera.
[0004] However, the above process needs to be manually completed by technicians, which has relatively strict requirements for the proficiency of technicians, and there will be errors when different technicians perform manual measurements, resulting in inaccurate determination of the pitch angle of the in-vehicle camera. Summary of the Invention
[0005] The present application provides a method, device, equipment and storage medium for determining the pitch angle of an in-vehicle camera, which can improve the accuracy of the pitch angle of the in-vehicle camera. The technical solution is as follows:
[0006] In a first aspect, a method for determining the pitch angle of an in-vehicle camera is provided. The method includes:
[0007] Obtain a first image and a second image, where the first image is the previous frame of two adjacent video images in the video captured by the in-vehicle camera, and the second image is the latter frame of the two video images;
[0008] Input the first image and the second image into an optical flow estimation model to output an optical flow map through the optical flow estimation model, where the optical flow map is used to indicate the displacement value of pixel points belonging to the same object between the first image and the second image;
[0009] Determine the focal length of the in-vehicle camera according to the optical flow map;
[0010] Obtain the coordinate value of the optical center in the second image and the coordinate value of the vanishing point in the second image;
[0011] Determine the pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image.
[0012] In this application, a first image and a second image are obtained. The first image and the second image are two adjacent video frames in the video captured by the vehicle-mounted camera. Then, the first image and the second image are input into the optical flow estimation model to output an optical flow map through the optical flow estimation model. Since the optical flow map can accurately indicate the displacement value of pixel points belonging to the same object between the first image and the second image, determining the focal length of the vehicle-mounted camera according to the optical flow map can make the determined focal length of the vehicle-mounted camera more accurate. Finally, obtain the coordinate value of the optical center in the second image and the coordinate value of the vanishing point in the second image, and determine the pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image. In this way, the focal length of the vehicle-mounted camera can be accurately determined according to the optical flow map, and the pitch angle of the vehicle-mounted camera can be determined according to the accurately determined focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image, which can improve the accuracy of the pitch angle of the vehicle-mounted camera.
[0013] Optionally, the optical flow estimation model includes a first encoding module, a second encoding module, and a first decoding module;
[0014] The first encoding module includes m first convolutional layers, the second encoding module includes m second convolutional layers. The m first convolutional layers are used to extract the feature map of the first image, and the m second convolutional layers are used to extract the feature map of the second image, where m is an integer greater than or equal to 2;
[0015] The first decoding module includes n optical flow estimation modules. The first optical flow estimation module among the n optical flow estimation modules is used to: output an optical flow map according to the feature map extracted by the m-th first convolutional layer among the m first convolutional layers and the feature map extracted by the m-th second convolutional layer among the m second convolutional layers; the i-th optical flow estimation module among the n optical flow estimation modules is used to: output an optical flow map according to the optical flow map output by the (i - 1)-th optical flow estimation module among the n optical flow estimation modules, the feature map extracted by the (m - i + 1)-th first convolutional layer among the m first convolutional layers, and the feature map extracted by the (m - i + 1)-th second convolutional layer among the m second convolutional layers; the optical flow map output by the n-th optical flow estimation module among the n optical flow estimation modules is the optical flow map output by the optical flow estimation model, where n is an integer greater than or equal to 2, and i is an integer greater than or equal to 2 and less than or equal to n.
[0016] Optionally, outputting an optical flow map based on the optical flow map output by the (i-1)-th optical flow estimation module among the n optical flow estimation modules, the feature map extracted by the (m-i+1)-th first convolutional layer among the m first convolutional layers, and the feature map extracted by the (m-i+1)-th second convolutional layer among the m second convolutional layers includes:
[0017] Upsampling the optical flow map output by the (i-1)-th optical flow estimation module to obtain a target optical flow map;
[0018] Transforming the feature map extracted by the (m-i+1)-th first convolutional layer according to the target optical flow map to obtain a target feature map;
[0019] Determining the matching cost between the target feature map and the feature map extracted by the (m-i+1)-th second convolutional layer;
[0020] Outputting an optical flow map according to the target optical flow map, the matching cost, and the feature map extracted by the (m-i+1)-th second convolutional layer.
[0021] Optionally, the displacement value of the pixel points belonging to the same object between the first image and the second image includes a horizontal displacement value, and determining the focal length of the vehicle-mounted camera according to the optical flow map includes:
[0022] Determining the average value of the horizontal displacement values of all the pixel points indicated by the optical flow map;
[0023] Obtaining the angular velocity of the vehicle equipped with the vehicle-mounted camera;
[0024] Dividing the average value by the angular velocity to obtain a first ratio;
[0025] Determining the negative value of the first ratio as the focal length of the vehicle-mounted camera.
[0026] Optionally, obtaining the coordinate value of the vanishing point in the second image includes:
[0027] Inputting the second image into a vanishing point prediction model to output a vanishing point probability matrix through the vanishing point prediction model, where the probability value in the vanishing point probability matrix is the probability value that each pixel point in the second image is a vanishing point, and the pixel point corresponding to the maximum probability value in the vanishing point probability matrix is the vanishing point in the second image;
[0028] Determining the coordinate value of the pixel point corresponding to the maximum probability value in the vanishing point probability matrix in the second image as the coordinate value of the vanishing point in the second image.
[0029] Optionally, determining the pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image includes:
[0030] Identifying two lane lines of the road included in the second image;
[0031] Determining the intersection coordinate value of the two lane lines in the second image in the second image;
[0032] Determining the pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the intersection coordinate value, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image.
[0033] Optionally, determining the pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the intersection coordinate value, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image includes:
[0034] Determining the initial pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image;
[0035] Determining the corrected pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the intersection coordinate value, and the coordinate value of the vanishing point in the second image;
[0036] Determining the pitch angle of the vehicle-mounted camera according to the initial pitch angle and the corrected pitch angle of the vehicle-mounted camera.
[0037] Optionally, determining the initial pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image includes:
[0038] Determining the initial pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image through the following formula;
[0039]
[0040] where, θ 初 is the initial pitch angle of the vehicle-mounted camera, f is the focal length of the vehicle-mounted camera, v1 is the ordinate value in the coordinate value of the vanishing point in the second image, and v0 is the ordinate value in the coordinate value of the optical center in the second image.
[0041] Optionally, determining the pitch angle of the vehicle-mounted camera according to the initial pitch angle and the corrected pitch angle of the vehicle-mounted camera includes:
[0042] Determine the pitch angle of the vehicle-mounted camera according to the initial pitch angle and the corrected pitch angle of the vehicle-mounted camera through the following formula;
[0043] θ = θ 初 + Δθ
[0044] where θ is the pitch angle of the vehicle-mounted camera, θ 初 is the initial pitch angle of the vehicle-mounted camera, and Δθ is the corrected pitch angle of the vehicle-mounted camera.
[0045] In a second aspect, a device for determining the pitch angle of a vehicle-mounted camera is provided. The device includes:
[0046] A first acquisition module, configured to acquire a first image and a second image. The first image is the previous frame of two adjacent video images in the video captured by the vehicle-mounted camera, and the second image is the subsequent frame of the two video images;
[0047] An optical flow generation module, configured to input the first image and the second image into an optical flow estimation model to output an optical flow map through the optical flow estimation model. The optical flow map is used to indicate the displacement value of pixel points belonging to the same object between the first image and the second image;
[0048] A first determination module, configured to determine the focal length of the vehicle-mounted camera according to the optical flow map;
[0049] A second acquisition module, configured to acquire the coordinate value of the optical center in the second image and the coordinate value of the vanishing point in the second image;
[0050] A second determination module, configured to determine the pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image.
[0051] Optionally, the optical flow estimation model includes a first encoding module, a second encoding module, and a first decoding module;
[0052] The first encoding module includes m first convolutional layers, the second encoding module includes m second convolutional layers. The m first convolutional layers are used to extract the feature map of the first image, and the m second convolutional layers are used to extract the feature map of the second image. m is an integer greater than or equal to 2;
[0053] The first decoding module includes n optical flow estimation modules. The first optical flow estimation module among the n optical flow estimation modules is configured to: output an optical flow map according to the feature map extracted from the m-th first convolutional layer among the m first convolutional layers and the feature map extracted from the m-th second convolutional layer among the m second convolutional layers; the i-th optical flow estimation module among the n optical flow estimation modules is configured to: output an optical flow map according to the optical flow map output by the (i - 1)-th optical flow estimation module among the n optical flow estimation modules, the feature map extracted from the (m - i + 1)-th first convolutional layer among the m first convolutional layers, and the feature map extracted from the (m - i + 1)-th second convolutional layer among the m second convolutional layers; the optical flow map output by the n-th optical flow estimation module among the n optical flow estimation modules is the optical flow map output by the optical flow estimation model, where n is an integer greater than or equal to 2, and i is an integer greater than or equal to 2 and less than or equal to n.
[0054] Optionally, the i-th optical flow estimation module is configured to:
[0055] Upsample the optical flow map output by the (i - 1)-th optical flow estimation module to obtain a target optical flow map;
[0056] Transform the feature map extracted from the (m - i + 1)-th first convolutional layer according to the target optical flow map to obtain a target feature map;
[0057] Determine the matching cost between the target feature map and the feature map extracted from the (m - i + 1)-th second convolutional layer;
[0058] Output an optical flow map according to the target optical flow map, the matching cost, and the feature map extracted from the (m - i + 1)-th second convolutional layer.
[0059] Optionally, the displacement value of the pixel points belonging to the same object between the first image and the second image includes a horizontal displacement value. The first determination module is configured to:
[0060] Determine the average value of the horizontal displacement values of all the pixel points indicated by the optical flow map;
[0061] Obtain the angular velocity of the vehicle equipped with the vehicle-mounted camera;
[0062] Divide the average value by the angular velocity to obtain a first ratio;
[0063] Determine the negative value of the first ratio as the focal length of the vehicle-mounted camera.
[0064] Optionally, the second acquisition module is configured to:
[0065] Input the second image into a vanishing point prediction model to output a vanishing point probability matrix through the vanishing point prediction model. The probability values in the vanishing point probability matrix are the probabilities that each pixel point in the second image is a vanishing point, and the pixel point corresponding to the maximum probability value in the vanishing point probability matrix is the vanishing point in the second image;
[0066] Determine the coordinate value of the pixel point corresponding to the maximum probability value in the vanishing point probability matrix in the second image as the coordinate value of the vanishing point in the second image.
[0067] Optionally, the second determination module includes:
[0068] An identification unit for identifying two lane lines of the road included in the second image;
[0069] A first determination unit for determining the intersection coordinate value of the two lane lines in the second image in the second image;
[0070] A second determination unit for determining the pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the intersection coordinate value, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image.
[0071] Optionally, the second determination unit is used to:
[0072] Determine the initial pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image;
[0073] Determine the corrected pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the intersection coordinate value, and the coordinate value of the vanishing point in the second image;
[0074] Determine the pitch angle of the vehicle-mounted camera according to the initial pitch angle and the corrected pitch angle of the vehicle-mounted camera.
[0075] Optionally, the second determination unit is used to:
[0076] Determine the initial pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image through the following formula;
[0077]
[0078] where θ 初 is the initial pitch angle of the vehicle-mounted camera, f is the focal length of the vehicle-mounted camera, v1 is the ordinate value in the coordinate value of the vanishing point in the second image, and v0 is the ordinate value in the coordinate value of the optical center in the second image.
[0079] Optionally, the second determination unit is configured to:
[0080] Determine the pitch angle of the vehicle-mounted camera according to the initial pitch angle and the corrected pitch angle of the vehicle-mounted camera through the following formula;
[0081] θ = θ 初 + Δθ
[0082] where θ is the pitch angle of the vehicle-mounted camera, θ 初 is the initial pitch angle of the vehicle-mounted camera, and Δθ is the corrected pitch angle of the vehicle-mounted camera.
[0083] In a third aspect, a computer device is provided. The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the above method for determining the pitch angle of a vehicle-mounted camera is implemented.
[0084] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above method for determining the pitch angle of a vehicle-mounted camera is implemented.
[0085] In a fifth aspect, a computer program product containing instructions is provided. When it runs on a computer, it causes the computer to execute the steps of the above method for determining the pitch angle of a vehicle-mounted camera.
[0086] It can be understood that the beneficial effects of the above second aspect, third aspect, fourth aspect, and fifth aspect can refer to the relevant descriptions in the above first aspect, and will not be elaborated here. Description of the Drawings
[0087] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0088] Figure 1 is a flowchart of a method for determining the pitch angle of a vehicle-mounted camera provided by an embodiment of the present application;
[0089] Figure 2 is a schematic structural diagram of an optical flow estimation model provided by an embodiment of the present application;
[0090] Figure 3 is a schematic structural diagram of a vanishing point prediction model provided by an embodiment of the present application;
[0091] Figure 4 It is a schematic structural diagram of an attention sub-module provided by an embodiment of the present application;
[0092] Figure 5 It is a schematic structural diagram of another attention sub-module provided by an embodiment of the present application;
[0093] Figure 6 It is a schematic diagram of the principle of a method for determining the pitch angle of an in-vehicle camera provided by an embodiment of the present application;
[0094] Figure 7 It is a schematic structural diagram of a device for determining the pitch angle of an in-vehicle camera provided by an embodiment of the present application;
[0095] Figure 8 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0096] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0097] It should be understood that the "multiple" mentioned in the present application refers to two or more. In the description of the present application, unless otherwise specified, " / " means "or", for example, A / B may represent A or B; the "and / or" herein is merely a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in order to clearly describe the technical solutions of the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the terms "first" and "second" do not limit the quantity and execution order, and the terms "first" and "second" do not necessarily limit to be different.
[0098] Before explaining the embodiments of the present application in detail, the application scenarios of the embodiments of the present application will be described first.
[0099] The parameters calibrated for an in-vehicle camera generally include the heading angle, pitch angle, and roll angle of the in-vehicle camera. Among them, the pitch angle of the in-vehicle camera is the most important. The pitch angle of the in-vehicle camera refers to the angle between the optical axis of the in-vehicle camera and the horizontal plane. If the pitch angle of the in-vehicle camera is not determined accurately, there will be errors in the processing results of the video images captured by the in-vehicle camera, which will affect the driving system's estimation of the vehicle state on the road, such as the estimation of the vehicle distance, etc., thus affecting the decision-making of the driving system and further affecting the normal driving of the vehicle.
[0100] The method for determining the pitch angle of an in-vehicle camera provided in an embodiment of the present application is applied to a scenario where the pitch angle of an in-vehicle camera is determined. Specifically, first, two adjacent frame video images captured by the in-vehicle camera during the driving of the vehicle are obtained, and then these two adjacent frame video images are input into an optical flow estimation model to obtain an optical flow map, so that the focal length of the in-vehicle camera can be calculated based on this optical flow map. After that, the coordinate value of the vanishing point in the video image is obtained, and finally, the pitch angle of the in-vehicle camera is determined according to the focal length of the in-vehicle camera, the optical center coordinate value of the video image, and the coordinate value of the vanishing point. In this way, the focal length of the in-vehicle camera can be accurately determined according to the optical flow map, and finally, the pitch angle of the in-vehicle camera is determined according to the accurately determined focal length of the in-vehicle camera, the optical center coordinate value, and the coordinate value of the vanishing point, so that the accuracy of the pitch angle of the in-vehicle camera can be improved.
[0101] The method for determining the pitch angle of an in-vehicle camera provided in an embodiment of the present application will be explained in detail below.
[0102] Figure 1 is a flowchart of a method for determining the pitch angle of an in-vehicle camera provided in an embodiment of the present application. This method can be applied to a terminal, such as an in-vehicle terminal. Refer to Figure 1 This method includes the following steps.
[0103] Step 101: The terminal obtains a first image and a second image.
[0104] During the driving of the vehicle, the in-vehicle camera captures a video including a road, and the terminal obtains two adjacent frame video images in the video captured by the in-vehicle camera. The first image is the previous frame video image among the two adjacent frame video images in the video captured by the in-vehicle camera, and the second image is the latter frame video image among these two frame video images. For example, the first image is the video image at time t in the video captured by the in-vehicle camera, and the second image is the video image at time t + 1 in the video captured by the in-vehicle camera.
[0105] Optionally, the terminal can obtain the first image and the second image in real time, or obtain the first image and the second image at intervals of a preset duration. The preset duration can be set in advance. And the preset duration can be set by a technician according to driving requirements.
[0106] That is to say, the first image and the second image can be two adjacent frame video images in the video captured by the in-vehicle camera that the terminal has obtained most recently. In this case, every time the terminal obtains the first image and the second image, the following steps 102 - 105 can be executed to determine the most recent pitch angle of the in-vehicle camera.
[0107] Step 102: The terminal inputs the first image and the second image into an optical flow estimation model to output an optical flow map through the optical flow estimation model.
[0108] The optical flow estimation model is used to estimate the optical flow of the input first image and second image to output an optical flow map. The optical flow map is a visualization image, and the optical flow map is used to accurately indicate the displacement value of pixel points belonging to the same object between the first image and the second image. For example, the optical flow map is used to indicate how much the pixel points belonging to the same object have moved from time t to time t+1.
[0109] Since the ability of neural network-based prediction of optical flow has far exceeded the method of outputting optical flow based on traditional feature matching, and the optical flow map is used to accurately indicate the displacement value of pixel points belonging to the same object between the first image and the second image, therefore, by outputting an optical flow map through this optical flow estimation model, this optical flow map can indicate a more accurate displacement value of pixel points belonging to the same object between the first image and the second image.
[0110] Exemplarily, the optical flow estimation model includes a first encoding module, a second encoding module, and a first decoding module; the first encoding module includes m first convolutional layers, the second encoding module includes m second convolutional layers, the first decoding module includes n optical flow estimation modules, and the first encoding module and the second encoding module process the input image using the same weights (which can also be called parameters). Specifically, the weights in the j-th first convolutional layer are the same as the weights in the j-th second convolutional layer, that is, the j-th first convolutional layer and the j-th second convolutional layer share weights. m is an integer greater than or equal to 2, n is an integer greater than or equal to 2, and m + 1 = n. j is an integer greater than or equal to 1 and less than or equal to m.
[0111] As Figure 2 shown, Figure 2 is a schematic structural diagram of the optical flow estimation model. The input of the optical flow estimation model is the first image 201 and the second image 202. The optical flow estimation model includes m first convolutional layers 203, m second convolutional layers 204, and n optical flow estimation modules 205. The output of the optical flow estimation model is the optical flow map 206. The m first convolutional layers 203 are connected in sequence, and the m second convolutional layers 204 are connected in sequence; the first optical flow estimation module 205 among the n optical flow estimation modules 205 is respectively connected to the m-th first convolutional layer 203 and the m-th second convolutional layer 204. The i-th optical flow estimation module 205 is respectively connected to the (i - 1)-th optical flow estimation module 205, the (m - i + 1)-th first convolutional layer 203, and the (m - i + 1)-th second convolutional layer 204. That is, the i-th optical flow estimation module 205 and the (m - i + 1)-th first convolutional layer 203 are cross-layer connected, and the i-th optical flow estimation module 205 and the (m - i + 1)-th second convolutional layer 204 are also cross-layer connected. i is an integer greater than or equal to 2 and less than or equal to n.
[0112] m first convolutional layers 203 are used to extract the feature maps of the first image 201, and m second convolutional layers 204 are used to extract the feature maps of the second image 202.
[0113] The first optical flow estimation module 205 among the n optical flow estimation modules 205 is used to: output an optical flow map according to the feature map extracted by the m-th first convolutional layer 203 among the m first convolutional layers 203 and the feature map extracted by the m-th second convolutional layer 204 among the m second convolutional layers 204.
[0114] The sizes of the feature map extracted by the m-th first convolutional layer 203 and the feature map extracted by the m-th second convolutional layer 204 are the same as the size of the optical flow map output by the first optical flow estimation module 205.
[0115] Optionally, the operation of the first optical flow estimation module 205 to output an optical flow map according to the feature map extracted by the m-th first convolutional layer 203 and the feature map extracted by the m-th second convolutional layer 204 may be: the first optical flow estimation module 205 determines the matching cost between the feature map extracted by the m-th first convolutional layer 203 and the feature map extracted by the m-th second convolutional layer 204; and outputs an optical flow map according to the matching cost between the feature map extracted by the m-th first convolutional layer 203 and the feature map extracted by the m-th second convolutional layer 204, and the feature map extracted by the m-th second convolutional layer 204.
[0116] The matching cost is used to measure the difference between two feature maps, and the greater the matching cost, the greater the difference between the two feature maps; the smaller the matching cost, the smaller the difference between the two feature maps. That is, the greater the matching cost between the feature map extracted by the m-th first convolutional layer 203 and the feature map extracted by the m-th second convolutional layer 204, the greater the difference between the feature map extracted by the m-th first convolutional layer 203 and the feature map extracted by the m-th second convolutional layer 204; the smaller the matching cost between the feature map extracted by the m-th first convolutional layer 203 and the feature map extracted by the m-th second convolutional layer 204, the smaller the difference between the feature map extracted by the m-th first convolutional layer 203 and the feature map extracted by the m-th second convolutional layer 204.
[0117] In this way, by using the matching cost between the feature map extracted by the m-th first convolutional layer 203 and the feature map extracted by the m-th second convolutional layer 204 to indicate the optical flow map output by the first optical flow estimation module 205, the optical flow map output by the first optical flow estimation module 205 can be made more accurate.
[0118] Among them, the operation of the first optical flow estimation module 205 to determine the matching cost between the feature map extracted by the m-th first convolutional layer 203 and the feature map extracted by the m-th second convolutional layer 204 can be: according to the feature map extracted by the m-th first convolutional layer 203 and the feature map extracted by the m-th second convolutional layer 204, the matching cost between the feature map extracted by the m-th first convolutional layer 203 and the feature map extracted by the m-th second convolutional layer 204 is obtained through the following formula:
[0119]
[0120] Among them, is the matching cost between the feature map extracted by the m-th first convolutional layer 203 and the feature map extracted by the m-th second convolutional layer 204, is the feature map extracted by the m-th first convolutional layer 203, is the feature map extracted by the m-th second convolutional layer 204, and K is the number of channels of the feature map extracted by the m-th first convolutional layer 203 or the feature map extracted by the m-th second convolutional layer 204.
[0121] Among them, the operation of the first optical flow estimation module 205 to output an optical flow map according to the matching cost between the feature map extracted by the m-th first convolutional layer 203 and the feature map extracted by the m-th second convolutional layer 204, and the feature map extracted by the m-th second convolutional layer 204 can be: according to the matching cost between the feature map extracted by the m-th first convolutional layer 203 and the feature map extracted by the m-th second convolutional layer 204, and the feature map extracted by the m-th second convolutional layer 204, the optical flow map is obtained through the following formula:
[0122]
[0123] Among them, flow 1 is the optical flow map output by the first optical flow estimation module 205, and CNN1() represents the first convolutional network.
[0124] The first convolutional network can be set in advance. The first convolutional network is set in the first optical flow estimation module 205, and the input of the first convolutional network is the matching cost between the feature map extracted by the m-th first convolutional layer 203 and the feature map extracted by the m-th second convolutional layer 204, and the feature map extracted by the m-th second convolutional layer 204. The output of the first convolutional network is the optical flow map. That is, by inputting the matching cost between the feature map extracted by the m-th first convolutional layer 203 and the feature map extracted by the m-th second convolutional layer 204, and the feature map extracted by the m-th second convolutional layer 204 into the first convolutional network for processing, the optical flow map can be obtained.
[0125] The $i$-th optical flow estimation module 205 in the $n$ optical flow estimation modules 205 is configured to: output an optical flow map according to the optical flow map output by the $(i - 1)$-th optical flow estimation module 205 in the $n$ optical flow estimation modules 205, the feature map extracted by the $(m - i + 1)$-th first convolutional layer 203 in the $m$ first convolutional layers 203, and the feature map extracted by the $(m - i + 1)$-th second convolutional layer 204 in the $m$ second convolutional layers 204.
[0126] In this way, taking the feature map extracted by the $(m - i + 1)$-th first convolutional layer 203 and the feature map extracted by the $(m - i + 1)$-th second convolutional layer 204 as the inputs of the $i$-th optical flow estimation module 205, and then having the $i$-th optical flow estimation module 205 process the feature map extracted by the $(m - i + 1)$-th first convolutional layer 203 and the feature map extracted by the $(m - i + 1)$-th second convolutional layer 204 can make the output optical flow map more accurate.
[0127] In this case, the optical flow map output by the $n$-th optical flow estimation module 205 in the $n$ optical flow estimation modules 205 is the optical flow map 206 output by the optical flow estimation model.
[0128] Optionally, the operation of the $i$-th optical flow estimation module 205 to output an optical flow map according to the optical flow map output by the $(i - 1)$-th optical flow estimation module 205, the feature map extracted by the $(m - i + 1)$-th first convolutional layer 203, and the feature map extracted by the $(m - i + 1)$-th second convolutional layer 204 can be: the $i$-th optical flow estimation module 205 upsamples the optical flow map output by the $(i - 1)$-th optical flow estimation module 205 to obtain a target optical flow map; according to the target optical flow map, transforms the feature map extracted by the $(m - i + 1)$-th first convolutional layer 203 to obtain a target feature map; determines the matching cost between the target feature map and the feature map extracted by the $(m - i + 1)$-th second convolutional layer 204; and outputs an optical flow map according to the target optical flow map, the matching cost, and the feature map extracted by the $(m - i + 1)$-th second convolutional layer 204.
[0129] In this case, first transforming the feature map extracted by the $(m - i + 1)$-th first convolutional layer 203 and then determining the matching cost between the obtained target feature map and the feature map extracted by the $(m - i + 1)$-th second convolutional layer 204 can make the matching cost more accurate, so that according to the target optical flow map, the matching cost, and the feature map extracted by the $(m - i + 1)$-th second convolutional layer 204, a more accurate optical flow map can be output.
[0130] Optionally, the upsampling multiple of the optical flow map output by the i-th optical flow estimation module 205 for the (i-1)-th optical flow estimation module 205 can be set according to the multiple difference between the size of the input data and the size of the output data of the (m-i+2)-th second convolutional layer. In this way, it can be ensured that the size of the target optical flow map is the same as the size of the feature map extracted by the (m-i+1)-th first convolutional layer 203 and the size of the feature map extracted by the (m-i+1)-th second convolutional layer 204 in subsequent operations. For example: if the multiple difference between the size of the input data and the size of the output data of the (m-i+2)-th second convolutional layer is 2, then the i-th optical flow estimation module 205 can perform 2-fold upsampling on the optical flow map output by the (i-1)-th optical flow estimation module 205.
[0131] Among them, the operation of the i-th optical flow estimation module 205 to transform the feature map extracted by the (m-i+1)-th first convolutional layer 203 according to the target optical flow map to obtain the target feature map can be: according to the target optical flow map, the feature map extracted by the (m-i+1)-th first convolutional layer 203 is transformed through the following formula:
[0132]
[0133] Among them, W i is the target feature map, Wrap() represents the transformation operation, is the feature map extracted by the (m-i+1)-th first convolutional layer 203, flow i-1 is the optical flow map output by the (i-1)-th optical flow estimation module 205, and upsample(flow i-1 ) is the target optical flow map.
[0134] Among them, the operation of the i-th optical flow estimation module 205 to determine the matching cost between the target feature map and the feature map extracted by the (m-i+1)-th second convolutional layer 204 is similar to the operation of determining the matching cost between the feature map extracted by the m-th first convolutional layer 203 and the feature map extracted by the m-th second convolutional layer 204 above, and this application embodiment will not elaborate on this.
[0135] Among them, the operation of the i-th optical flow estimation module 205 to output the optical flow map according to the target optical flow map, this matching cost, and the feature map extracted by the (m-i+1)-th second convolutional layer 204 can be: according to the target optical flow map, this matching cost, and the feature map extracted by the (m-i+1)-th second convolutional layer 204, the optical flow map is obtained through the following formula:
[0136]
[0137] Among them, flow i is the optical flow map output by the i-th optical flow estimation module 205, is the matching cost between the target feature map and the feature map extracted by the (m - i + 1)-th second convolutional layer 204, and CNN2() is the second convolutional network.
[0138] The second convolutional network can be pre-set. The second convolutional network is disposed in the i-th optical flow estimation module 205, and the input of the second convolutional network is the target optical flow map, this matching cost, and the feature map extracted by the (m - i + 1)-th second convolutional layer 204, and the output of the second convolutional network is the optical flow map. That is, by inputting the target optical flow map, this matching cost, and the feature map extracted by the (m - i + 1)-th second convolutional layer 204 into the second convolutional network for processing, the optical flow map can be obtained.
[0139] It should be noted that before using this optical flow estimation model to estimate the optical flow between the first image and the second image to obtain the optical flow map, this optical flow estimation model also needs to be trained.
[0140] Specifically, the terminal can obtain multiple training samples and use these multiple training samples to train the neural network model to obtain this optical flow estimation model.
[0141] These multiple training samples can be pre-set. Each training sample in these multiple training samples includes a sample image and a sample label. The sample image is a video image pair, where the video image pair is two adjacent frames of video images captured by a vehicle-mounted camera, and the sample label is the optical flow map corresponding to the sample image. That is, the input data in each training sample of these multiple training samples is the video image pair, and the sample label is the optical flow map corresponding to this video image pair. The optical flow map corresponding to this video image is used to indicate the displacement value of pixel points belonging to the same object between this video image pair.
[0142] The neural network model can include multiple network layers, and these multiple network layers include an input layer, multiple hidden layers, and an output layer. The input layer is responsible for receiving input data; the output layer is responsible for outputting processed data; the multiple hidden layers are located between the input layer and the output layer and are responsible for processing data, and the multiple hidden layers are invisible to the outside. The structure of this neural network model is the same as that of the optical flow estimation model shown in Figure 2 the optical flow estimation model shown in
[0143] Among them, when the terminal uses multiple training samples to train the neural network model, for each training sample in these multiple training samples, the input data in this training sample can be input into the neural network model to obtain output data; the loss value between this output data and the sample label in this training sample is determined through a loss function; the parameters in the neural network model are adjusted according to this loss value. After the parameters in the neural network model are adjusted based on each training sample in these multiple training samples, the neural network model with the parameters adjusted is this optical flow estimation model.
[0144] Optionally, the training of the optical flow estimation model provided by the embodiments of the present application is divided into two stages: a pre-training stage and an optimization stage. In the pre-training stage, the embodiments of the present application use the mean square error loss function to determine the loss value between the output data and the sample label in this training sample. For example, the loss value between the output data and the sample label in this training sample is determined by the following formula.
[0145]
[0146] where loss pre-train is the loss value between the output data and the sample label in the pre-training stage, F is the number of pixel points in the output data, s is the s-th optical flow estimation module among the n optical flow estimation modules, r represents the r-th pixel point among the pixel points of the output data, is the optical flow of the r-th pixel point in the optical flow map output by the s-th optical flow estimation module, is the optical flow of the r-th pixel point in the sample label, s is an integer greater than or equal to 1 and less than or equal to n, F is a positive integer, and r is an integer greater than or equal to 1 and less than or equal to F.
[0147] In this way, using the mean square error loss function in the pre-training stage can accelerate the convergence speed of the neural network model.
[0148] In the optimization stage, the embodiments of the present application use the mean absolute error loss function to determine the loss value between the output data and the sample label in this training sample. For example, the loss value between the output data and the sample label in this training sample is determined by the following formula.
[0149]
[0150] where loss fine-tune is the loss value between the output data and the sample label in the optimization stage.
[0151] In this way, using the mean absolute error loss function in the optimization stage can further improve the optical flow quality.
[0152] Among them, the operation of the terminal to adjust the parameters in the neural network model according to the loss value can refer to the related technology, and the embodiments of the present application do not elaborate on this in detail.
[0153] For example, the terminal can use the formula to adjust any parameter in the neural network model. Among them, It is the adjusted parameter. z is the parameter before adjustment. α is the learning rate, which can be preset. For example, α can be 0.001, 0.000001, etc. The embodiments of the present application do not make a unique limitation on this. dz is the partial derivative of the loss function with respect to z, which can be obtained according to the loss value.
[0154] Optionally, the training dataset of the optical flow estimation model can be a public dataset or video images captured by an in-vehicle camera collected on the network. The embodiments of the present application do not make a limitation on this.
[0155] Since the optical flow map output by the optical flow estimation model is used to indicate the displacement value of pixel points belonging to the same object between the first image and the second image, the terminal can determine the focal length of the in-vehicle camera according to the optical flow map, that is, continue to execute step 103.
[0156] Step 103: The terminal determines the focal length of the in-vehicle camera according to the optical flow map.
[0157] The displacement value of pixel points belonging to the same object between the first image and the second image includes a horizontal displacement value, that is, the displacement value of pixel points belonging to the same object in the horizontal direction.
[0158] When the vehicle is turning, that is, when the in-vehicle camera rotates horizontally around the vertical direction by a certain angle within a unit time, this angle is reflected in the image as a certain displacement of the row pixels in the image, that is, the displacement value of pixel points belonging to the same object in the horizontal direction (i.e., horizontally). The focal length of the in-vehicle camera is equivalent to a scale factor, and the ratio of this angle to the displacement value generated by the corresponding row pixels in the image is the focal length of the in-vehicle camera.
[0159] In this case, the terminal can determine the focal length of the in-vehicle camera according to the relatively accurate displacement value of pixel points indicated by the accurately obtained optical flow map during the driving of the vehicle, so that the determined focal length of the in-vehicle camera is relatively accurate.
[0160] Specifically, the operation of step 103 can be: the terminal determines the average value of the horizontal displacement values of all pixel points indicated by the optical flow map; obtains the angular velocity of the vehicle equipped with the in-vehicle camera; divides the average value by the angular velocity to obtain a first ratio; and determines the negative of the first ratio as the focal length of the in-vehicle camera.
[0161] This average value is the offset distance between the first image and the second image. During the driving of the vehicle, the angular velocity of the vehicle is measured. The process of determining the focal length of the in-vehicle camera above is the process of determining the focal length of the in-vehicle camera through the following formula:
[0162]
[0163] where f is the focal length of the vehicle-mounted camera, and Δu is the average value, i.e., the distance offset. is the angular velocity.
[0164] Step 104: The terminal obtains the coordinate value of the optical center in the second image and the coordinate value of the vanishing point in the second image.
[0165] The optical center is the center of the vehicle-mounted camera lens and also the center of the second image. Then, the coordinate value of the optical center in the second image is where W and H are the width and height of the second image, respectively.
[0166] The vanishing point is the visual intersection point of two parallel lane lines in the second image, which is usually used for camera parameter calibration.
[0167] Optionally, the operation for the terminal to obtain the coordinate value of the vanishing point in the second image can be: the terminal inputs the second image into the vanishing point prediction model to output a vanishing point probability matrix through the vanishing point prediction model. The probability value in the vanishing point probability matrix is the probability value that each pixel point in the second image is the vanishing point. The pixel point corresponding to the maximum probability value in the vanishing point probability matrix is the vanishing point in the second image; the coordinate value of the pixel point corresponding to the maximum probability value in the vanishing point probability matrix in the second image is determined as the coordinate value of the vanishing point in the second image.
[0168] In this way, during the vehicle driving process, by inputting the second image into the vanishing point prediction model, the coordinate value of the vanishing point in the second image can be automatically obtained.
[0169] Optionally, after the terminal determines the coordinate value of the pixel point corresponding to the maximum probability value in the vanishing point probability matrix as the coordinate value of the vanishing point in the second image, a Gaussian heat map can also be generated according to the coordinate value of the vanishing point in the second image to achieve the visualization effect of the vanishing point.
[0170] For example: The terminal can set the kernel radius of the Gaussian heat map to 9, use the coordinate value of the vanishing point in the second image as the center point of the Gaussian heat map, and then set the Gaussian value of the center point of the Gaussian heat map to 1. The color in the Gaussian heat map corresponding to the Gaussian value of 1 is red, that is, the red area in the Gaussian heat map is the position of the vanishing point.
[0171] Exemplarily, the vanishing point prediction model sequentially includes a third encoding module, an attention module, and a second decoding module, as Figure 3 shown. Figure 3It is a schematic structural diagram of the vanishing point prediction model. The input data of the vanishing point prediction model is input data 301. The vanishing point prediction model includes a third encoding module 302, an attention module 303, and a second decoding module 304. The output data of the vanishing point prediction model is output data 305. The third encoding module 302 is used to perform convolution operations. The attention module 303 is used to extract attention feature maps. The second decoding module 304 is used to perform deconvolution operations. The third encoding module 302, the attention module 303, and the second decoding module are connected in sequence.
[0172] In this case, by setting the attention module 303 in the vanishing point prediction model, the receptive field of the vanishing point prediction model can be increased, the expressiveness of the feature map can be enhanced, and thus the accuracy of vanishing point prediction can be improved.
[0173] Optionally, the third encoding module 302 includes multiple third convolutional layers, and the convolutional kernel sizes of each third convolutional layer can be set to be different or the same. The second decoding module 304 includes multiple deconvolutional layers, and the convolutional kernel sizes of the deconvolutional layers can be set according to the convolutional kernel sizes of the corresponding third convolutional layers in the third encoding module 302. For example: Figure 3 In [example], the third encoding module 302 includes three third convolutional layers, the second decoding module includes 3 deconvolutional layers, the convolutional kernel size of the first third convolutional layer is 7×7, and the convolutional kernel sizes of the second and third third convolutional layers are both 4×4, with a stride set to 2. Since the second decoding module 304 is used to decode the size of the feature map to be the same as the size of the input data of the vanishing point prediction model, the convolutional kernel sizes of the first and second deconvolutional layers are 4×4, with a stride set to 2, and the convolutional kernel size of the third deconvolutional layer is 7×7.
[0174] The attention module 303 includes multiple attention sub-modules, and each attention sub-module in the multiple attention sub-modules can be implemented by the following two possible structures.
[0175] The first possible structure is as Figure 4 shown. Figure 4It is a schematic structural diagram of the attention sub-module. The input data of the attention sub-module is the input data 401. The attention sub-module includes a fourth convolutional layer 402, a horizontal pooling layer 403, a vertical pooling layer 404, a first decoding layer 405, a second decoding layer 406, a first fusion layer 407, an activation layer 408, a second fusion layer 409, and a third fusion layer 410. The output data of the attention sub-module is the attention feature map 411. The fourth convolutional layer 402 is respectively connected to the horizontal pooling layer 403 and the vertical pooling layer 404. The horizontal pooling layer 403 is connected to the first decoding layer 405. The vertical pooling layer 404 is connected to the second decoding layer 406. Both the first decoding layer 405 and the second decoding layer 406 are connected to the first fusion layer 407. The first fusion layer 407 is connected to the activation layer 408. Both the activation layer 408 and the fourth convolutional layer 402 are connected to the second fusion layer 409. Both the second fusion layer 409 and the input data 401 are connected to the third fusion layer 410.
[0176] The fourth convolutional layer 401 is used to perform a convolution operation on the input data 401 of the attention sub-module to obtain a first feature map. The horizontal pooling layer 403 is used to perform a pooling operation on the first feature map using a pooling kernel of size 1×W to obtain a horizontal feature map. The vertical pooling layer 404 is used to perform a pooling operation on the first feature map using a pooling kernel of size H×1 to obtain a vertical feature map. The first decoding layer 405 is used to perform upsampling on the horizontal feature map to obtain a horizontal attention feature map. The second decoding layer 406 is used to perform upsampling on the vertical feature map to obtain a vertical attention feature map. The first fusion layer 407 is used to add the horizontal attention feature map and the vertical attention feature map to obtain a fused feature map. The activation layer 408 is used to perform an activation operation on the fused feature map to obtain an activation feature map including horizontal attention features and vertical attention features. The second fusion layer 409 is used to perform a multiplication operation on the activation feature map and the first feature map to obtain a second feature map. The third fusion layer 410 is used to add the input data 401 of the attention sub-module and the second feature map to obtain the attention feature map 411. Here, W is the width of the first feature map, and H is the height of the first feature map.
[0177] Since the second image contains horizontal semantics and vertical semantics, the horizontal attention mechanism and the vertical attention mechanism are adopted to respectively aggregate context information from the horizontal and vertical directions, model the long-range dependence relationship of features, and extract the horizontal features and vertical features in the first feature map. In this way, the expression ability of features can be enhanced, thereby improving the accuracy of vanishing point prediction.
[0178] The second possible structure is as Figure 5 shown Figure 5It is a schematic structural diagram of the attention sub-module. The input data of the attention sub-module is the input data 501. The attention sub-module includes a fourth convolutional layer 502, a horizontal pooling layer 503, a vertical pooling layer 504, a fifth convolutional layer 505, a sixth convolutional layer 506, a first decoding layer 507, a second decoding layer 508, a first fusion layer 509, a seventh convolutional layer 510, an activation layer 511, a second fusion layer 512, an eighth convolutional layer 513, and a third fusion layer 514. The output data of the attention sub-module is the attention feature map 515. The fourth convolutional layer 502 is respectively connected to the horizontal pooling layer 503 and the vertical pooling layer 504; the horizontal pooling layer 503, the fifth convolutional layer 505, and the first decoding layer 507 are connected in sequence, and the vertical pooling layer 504, the sixth convolutional layer 506, and the second decoding layer 508 are connected in sequence; both the first decoding layer 507 and the second decoding layer 508 are connected to the first fusion layer 509; the first fusion layer 509, the seventh convolutional layer 510, and the activation layer 511 are connected in sequence; both the activation layer 511 and the fourth convolutional layer 502 are connected to the second fusion layer 512; the second fusion layer 512 is connected to the eighth convolutional layer 513, and both the eighth convolutional layer 513 and the input data 501 are connected to the third fusion layer 514.
[0179] The fourth convolutional layer 502 is used to perform a convolution operation on the input data 501 of the attention sub-module to obtain a first feature map; the horizontal pooling layer 503 is used to perform a pooling operation on the first feature map using a pooling kernel of size 1×W to obtain a horizontal feature map; the vertical pooling layer 504 is used to perform a pooling operation on the first feature map using a pooling kernel of size H×1 to obtain a vertical feature map; the fifth convolutional layer 505 is used to perform a convolution operation on the horizontal feature map to obtain a third feature map; the sixth convolutional layer 506 is used to perform a convolution operation on the vertical feature map to obtain a fourth feature map; the first decoding layer 507 is used to perform upsampling on the third feature map to obtain a horizontal attention feature map; the second decoding layer 508 is used to perform upsampling on the fourth feature map to obtain a vertical attention feature map; the first fusion layer 509 is used to add the horizontal attention feature map and the vertical attention feature map to obtain a fusion feature map; the seventh convolutional layer 510 is used to perform a convolution operation on the fusion feature map to obtain a fifth feature map; the activation layer 511 is used to perform an activation operation on the fifth feature map to obtain an activation feature map including horizontal attention features and vertical attention features; the second fusion layer 512 is used to multiply the activation feature map by the first feature map to obtain a second feature map; the eighth convolutional layer 513 is used to perform a convolution operation on the second feature map to obtain a sixth feature map; the third fusion layer 514 is used to add the input data 501 of the attention sub-module and the sixth feature map to obtain the attention feature map 515.
[0180] In this case, a fifth convolutional layer 505, a sixth convolutional layer 506, and an eighth convolutional layer 513 are added to the attention sub-module to perform feature integration on the feature map, and the receptive field of the attention module can be increased, thereby improving the accuracy of vanishing point prediction. Moreover, by adding a seventh convolutional layer 510, the dimension of the feature map can be reduced, thus saving computing resources.
[0181] Optionally, the pooling operations of the horizontal pooling layer 503 and the vertical pooling layer 504 can be average pooling operations or maximum pooling operations, and the embodiments of the present application do not limit this.
[0182] Optionally, the activation function of the activation layer 511 can be a Sigmoid activation function, a Softmax activation function, a Relu (Rectified Linear Unit) activation function, etc., and the embodiments of the present application do not limit this.
[0183] For example: the convolution kernel size of the fourth convolutional layer 502 can be set to 3×3, and the dilation factor size can be set to 2. Then, the first feature map extracted by the fourth convolutional layer 502 is respectively input into the horizontal pooling layer 503 with a pooling kernel size of 1×W and the vertical pooling layer 504 with a pooling kernel size of H×1, and average pooling operations are respectively performed on the first feature map: where p is the index value of the row pixel points in the first feature map, q is the index value of the column pixel points in the first feature map, c is the number of channels of the first feature map, x 横 is the horizontal attention feature map, and x 纵 is the vertical attention feature map. Then, the horizontal attention feature map x 横 is input into the fifth convolutional layer 505 with a convolution kernel size of 3×1 for convolution operation to obtain a third feature map, and the vertical attention feature map x 纵Perform a convolution operation in the sixth convolutional layer 506 with an input convolution kernel size of 1×3 to obtain a fourth feature map. Since the sizes of the third feature map and the fourth feature map are different from the feature map size of the input data 501 of the attention sub-module, a fusion operation cannot be performed subsequently. Therefore, the third feature map is input into the first decoding layer 507 to obtain a horizontal attention feature map, and the fourth feature map is input into the second decoding layer 508 to obtain a vertical attention feature map. In this way, the sizes of the horizontal attention feature map and the vertical attention feature map are the same. Then, the horizontal attention feature map and the vertical attention feature map are added and fused to obtain a fused feature map including global attention features. After that, this fused feature map is input into the seventh convolutional layer 510 with a convolution kernel size of 1×1 to perform a convolution operation to reduce the dimension of the fused feature map. After obtaining the activation feature map including horizontal attention features and vertical attention features, multiply the activation feature map by the first feature map, that is, fuse the obtained attention features and the first feature map, and a feature map including fused attention features, that is, the second feature map, can be obtained. Finally, input the second feature map into an eighth convolutional layer with a convolution kernel size of 3×3 to perform a convolution operation to obtain a sixth feature map, and fuse the sixth feature map with the input data 501 of the attention sub-module, and the attention feature map 515 output by this attention sub-module can be obtained.
[0184] It should be noted that before using the vanishing point prediction model to predict the vanishing point of the second image and output the vanishing point probability matrix, the vanishing point prediction model needs to be trained.
[0185] Specifically, the terminal can obtain multiple training samples and use these multiple training samples to train the neural network model to obtain the vanishing point prediction model.
[0186] These multiple training samples can be preset. Each training sample in these multiple training samples includes a sample image and a sample label. The sample image is an image containing a vanishing point, and the sample label is the coordinate value of the vanishing point contained in the sample image. That is, the input data in each training sample in these multiple training samples is the sample image containing a vanishing point, and the sample label is the coordinate value of the vanishing point contained in the sample image.
[0187] The neural network model can include multiple network layers, and these multiple network layers include an input layer, multiple hidden layers, and an output layer. The input layer is responsible for receiving input data; the output layer is responsible for outputting the processed data; the multiple hidden layers are located between the input layer and the output layer and are responsible for processing data, and the multiple hidden layers are not visible to the outside. The neural network model can be Figure 4 the same as the structure of the vanishing point prediction model shown in Figure 5 or the neural network model can be
[0188] When the terminal uses multiple training samples to train the neural network model, for each training sample in the multiple training samples, the input data in this training sample can be input into the neural network model to obtain output data; the loss value between the output data and the sample label in this training sample is determined through a loss function; and the parameters in the neural network model are adjusted according to the loss value. After adjusting the parameters in the neural network model based on each training sample in the multiple training samples, the neural network model with the parameter adjustment completed is the vanishing point prediction model.
[0189] Optionally, the vanishing point prediction model provided in the embodiments of the present application uses the mean squared error loss function to determine the loss value between the output data and the sample label in this training sample during the training phase. Specifically, the operation of using the mean squared error loss function to determine the loss value between the output data and the sample label in this training sample has been elaborated in detail in step 102, and will not be repeated here.
[0190] Among them, the operation of the terminal adjusting the parameters in the neural network model according to the loss value can refer to related technologies, and the embodiments of the present application will not elaborate on this in detail.
[0191] For example, the terminal can use the formula to adjust any parameter in the neural network model. Among them, is the adjusted parameter. z is the parameter before adjustment. α is the learning rate, and α can be set in advance. For example, α can be 0.001, 0.000001, etc. The embodiments of the present application do not make a unique limitation on this. dz is the partial derivative of the loss function with respect to z, which can be obtained according to the loss value.
[0192] Optionally, the training dataset of the vanishing point prediction model can be a publicly available road dataset, such as: the Cityscapes dataset, or road scene images collected on the Internet.
[0193] Step 105: The terminal determines the pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image.
[0194] In this case, since the focal length of the vehicle-mounted camera is a relatively accurate focal length determined according to the optical flow map, determining the pitch angle of the vehicle-mounted camera according to the relatively accurate focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image can make the determined pitch angle of the vehicle-mounted camera relatively accurate.
[0195] Specifically, the operation of step 105 can be as follows: The terminal identifies two lane lines of the road included in the second image; determines the intersection coordinate values of the two lane lines in the second image in the second image; and determines the pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the intersection coordinate values, the coordinate values of the optical center in the second image, and the coordinate values of the vanishing point in the second image.
[0196] The intersection coordinate values of the two lane lines in the second image in the second image are the coordinate values of the true vanishing point in the second image obtained by identifying the two lane lines, while the coordinate values of the vanishing point in the second image obtained in step 104 above are the coordinate values of the predicted vanishing point in the second image obtained by prediction.
[0197] Optionally, the way for the terminal to identify two lane lines of the road included in the second image can be: First, use an edge detection operator to detect the edge information of the second image, and then extract the two lane lines of the road from the edge information of the second image. Optionally, the edge detection operator can be a canny operator, a sobel operator, etc., and the method for extracting the two lane lines of the road from the edge information of the second image can be a Hough transform, etc., and the embodiments of the present application do not limit this.
[0198] Among them, the operation of the terminal to determine the pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the intersection coordinate values, the coordinate values of the optical center in the second image, and the coordinate values of the vanishing point in the second image can be implemented through the following steps (1)-(3):
[0199] (1) The terminal determines the initial pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate values of the optical center in the second image, and the coordinate values of the vanishing point in the second image.
[0200] Since the coordinate values of the vanishing point in the second image are the coordinate values of the predicted vanishing point in the second image obtained by prediction, the initial pitch angle of the vehicle-mounted camera determined by the coordinate values of the predicted vanishing point in the second image is the predicted initial pitch angle of the vehicle-mounted camera.
[0201] Specifically, as Figure 6 shown, Figure 6 is the schematic diagram for determining the initial pitch angle of the vehicle-mounted camera, Figure 6 which includes the second image 601, the optical center 602 of the second image, the vanishing point 603 of the second image, the vehicle-mounted camera 604, and the optical axis 605. Figure 6 In which, the optical axis 605, the focal length f of the vehicle-mounted camera, and the Y-axis form a right triangle, then the terminal can determine the initial pitch angle of the vehicle-mounted camera 604 through the following formula according to the focal length f of the vehicle-mounted camera 604, the coordinate values of the optical center 602 in the second image, and the coordinate values of the vanishing point 603 in the second image:
[0202]
[0203] where, θ 初 is the initial pitch angle of the vehicle-mounted camera 604, f is the focal length of the vehicle-mounted camera 604, v1 is the ordinate value in the coordinate values of the vanishing point 603 in the second image, and v0 is the ordinate value in the coordinate values of the optical center 602 in the second image.
[0204] (2) The terminal determines the corrected pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the intersection coordinate value, and the coordinate value of the vanishing point in the second image.
[0205] The corrected pitch angle is the angle required to correct the initial pitch angle of the vehicle-mounted camera.
[0206] Since the coordinate value of the vanishing point in the second image is the coordinate value of the predicted vanishing point in the predicted second image, and the intersection coordinate value is the coordinate value of the true vanishing point in the second image, there may be a gap between the coordinate value of the predicted vanishing point and the coordinate value of the true vanishing point in the second image, so there may also be a gap between the determined initial pitch angle of the vehicle-mounted camera and the true pitch angle.
[0207] In this case, according to the focal length of the vehicle-mounted camera, the intersection coordinate value, and the coordinate value of the vanishing point in the predicted second image, the gap between the initial pitch angle and the true pitch angle of the vehicle-mounted camera can be determined, and this gap is the corrected pitch angle of the vehicle-mounted camera. Using the corrected pitch angle in this way can correct the initial pitch angle of the vehicle-mounted camera to obtain the true pitch angle of the vehicle-mounted camera.
[0208] Specifically, the operation of the terminal to determine the corrected pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the intersection coordinate value, and the coordinate value of the vanishing point in the second image can be: the terminal determines the corrected pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the intersection coordinate value, and the coordinate value of the vanishing point in the second image through the following formula:
[0209]
[0210] where, Δθ is the corrected pitch angle of the vehicle-mounted camera, and v2 is the ordinate value in the intersection coordinate value.
[0211] (3) The terminal determines the pitch angle of the vehicle-mounted camera according to the initial pitch angle and the corrected pitch angle of the vehicle-mounted camera.
[0212] After the terminal knows the gap between the initial pitch angle and the true pitch angle of the vehicle-mounted camera, it can correct the initial pitch angle of the vehicle-mounted camera according to this gap, that is, add the corrected pitch angle to the initial pitch angle to correct the initial pitch angle, that is, a more accurate pitch angle of the vehicle-mounted camera can be obtained through the correction process described by the following formula.
[0213] Specifically, according to the initial pitch angle and the calibrated pitch angle of the vehicle-mounted camera, the pitch angle of the vehicle-mounted camera is determined by the following formula:
[0214] θ = θ 初 + Δθ
[0215] where θ is the pitch angle of the vehicle-mounted camera.
[0216] In this case, the terminal first determines the initial pitch angle of the vehicle-mounted camera, then determines the calibrated pitch angle of the vehicle-mounted camera, and then uses the calibrated pitch angle to correct the initial pitch angle, so that the determined pitch angle of the vehicle-mounted camera is relatively accurate, thus improving the accuracy of the pitch angle of the vehicle-mounted camera.
[0217] It should be noted that the method for determining the pitch angle of the vehicle-mounted camera provided by the embodiments of the present application can automatically determine the pitch angle of the vehicle-mounted camera during the driving process of the vehicle. The whole process does not require the intervention of technicians, and can effectively solve the problem of inaccurate determination of the pitch angle caused by factors such as the proficiency of technicians and measurement errors. In addition, since the method for determining the pitch angle of the vehicle-mounted camera provided by the embodiments of the present application does not require too many input parameters, in the subsequent use process, even if the internal parameters of the vehicle-mounted camera are severely distorted due to a series of reasons such as aging and damage, the pitch angle of the vehicle-mounted camera can still be accurately determined.
[0218] In the embodiments of the present application, the terminal acquires a first image and a second image. The first image and the second image are two adjacent video images in the video captured by the vehicle-mounted camera. Then, the first image and the second image are input into an optical flow estimation model to output an optical flow map through the optical flow estimation model. Since the optical flow map can accurately indicate the displacement value of the pixel points belonging to the same object between the first image and the second image, determining the focal length of the vehicle-mounted camera according to the optical flow map can make the determined focal length of the vehicle-mounted camera relatively accurate. Finally, the coordinate value of the optical center in the second image and the coordinate value of the vanishing point in the second image are acquired, and the pitch angle of the vehicle-mounted camera is determined according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image. In this way, the focal length of the vehicle-mounted camera can be accurately determined according to the optical flow map, and the pitch angle of the vehicle-mounted camera can be determined according to the accurately determined focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image, which can improve the accuracy of the pitch angle of the vehicle-mounted camera.
[0219] Figure 7 is a schematic structural diagram of a device for determining the pitch angle of a vehicle-mounted camera provided by the embodiments of the present application. The device for determining the pitch angle of the vehicle-mounted camera can be implemented by software, hardware, or a combination of both to become part or all of a computer device, and the computer device can be the computer device shown in the following Figure 8 as shown. SeeFigure 7 , the device includes: a first acquisition module 701, an optical flow generation module 702, a first determination module 703, a second acquisition module 704, and a second determination module 705.
[0220] The first acquisition module 701 is configured to acquire a first image and a second image. The first image is the previous frame of two adjacent video images in the video captured by the vehicle-mounted camera, and the second image is the subsequent frame of these two video images.
[0221] The optical flow generation module 702 is configured to input the first image and the second image into an optical flow estimation model, so as to output an optical flow map through the optical flow estimation model. The optical flow map is used to indicate the displacement value of pixel points belonging to the same object between the first image and the second image.
[0222] The first determination module 703 is configured to determine the focal length of the vehicle-mounted camera according to the optical flow map.
[0223] The second acquisition module 704 is configured to acquire the coordinate value of the optical center in the second image and the coordinate value of the vanishing point in the second image.
[0224] The second determination module 705 is configured to determine the pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image.
[0225] Optionally, the optical flow estimation model includes a first encoding module, a second encoding module, and a first decoding module.
[0226] The first encoding module includes m first convolutional layers, the second encoding module includes m second convolutional layers. The m first convolutional layers are used to extract the feature map of the first image, and the m second convolutional layers are used to extract the feature map of the second image. m is an integer greater than or equal to 2.
[0227] The first decoding module includes n optical flow estimation modules. The first optical flow estimation module among the n optical flow estimation modules is configured to: output an optical flow map according to the feature map extracted by the m first convolutional layers among the m first convolutional layers and the feature map extracted by the m second convolutional layers among the m second convolutional layers. The i-th optical flow estimation module among the n optical flow estimation modules is configured to: output an optical flow map according to the optical flow map output by the (i - 1)-th optical flow estimation module among the n optical flow estimation modules, the feature map extracted by the (m - i + 1)-th first convolutional layer among the m first convolutional layers, and the feature map extracted by the (m - i + 1)-th second convolutional layer among the m second convolutional layers. The optical flow map output by the n-th optical flow estimation module among the n optical flow estimation modules is the optical flow map output by the optical flow estimation model. n is an integer greater than or equal to 2, and i is an integer greater than or equal to 2 and less than or equal to n.
[0228] Optionally, the (i - 1)-th optical flow estimation module is configured to:
[0229] Upsample the optical flow map output by the (i - 1)-th optical flow estimation module to obtain a target optical flow map;
[0230] According to the target optical flow map, transform the feature map extracted by the (m - i + 1)-th first convolutional layer to obtain a target feature map;
[0231] Determine the matching cost between the target feature map and the feature map extracted by the (m - i + 1)-th second convolutional layer;
[0232] Output an optical flow map according to the target optical flow map, the matching cost, and the feature map extracted by the (m - i + 1)-th second convolutional layer.
[0233] Optionally, the displacement value of the pixel points belonging to the same object between the first image and the second image includes a lateral displacement value, and the first determination module 703 is used for:
[0234] Determine the average value of the lateral displacement values of all the pixel points indicated by the optical flow map;
[0235] Obtain the angular velocity of the vehicle equipped with the vehicle-mounted camera;
[0236] Divide the average value by the angular velocity to obtain a first ratio;
[0237] Determine the negative value of the first ratio as the focal length of the vehicle-mounted camera.
[0238] Optionally, the second acquisition module 704 is used for:
[0239] Input the second image into a vanishing point prediction model to output a vanishing point probability matrix through the vanishing point prediction model. The probability value in the vanishing point probability matrix is the probability value that each pixel point in the second image is a vanishing point, and the pixel point corresponding to the maximum probability value in the vanishing point probability matrix is the vanishing point in the second image;
[0240] Determine the coordinate value of the pixel point corresponding to the maximum probability value in the vanishing point probability matrix in the second image as the coordinate value of the vanishing point in the second image.
[0241] Optionally, the second determination module 705 includes:
[0242] An identification unit for identifying two lane lines of the road included in the second image;
[0243] A first determination unit for determining the intersection coordinate value of the two lane lines in the second image in the second image;
[0244] A second determination unit for determining the pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the intersection coordinate value, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image.
[0245] Optionally, the second determination unit is configured to:
[0246] Determine an initial pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image;
[0247] Determine a calibrated pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the intersection point, and the coordinate value of the vanishing point in the second image;
[0248] Determine the pitch angle of the vehicle-mounted camera according to the initial pitch angle and the calibrated pitch angle of the vehicle-mounted camera.
[0249] Optionally, the second determination unit is configured to:
[0250] Determine the initial pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image through the following formula;
[0251]
[0252] where θ 初 is the initial pitch angle of the vehicle-mounted camera, f is the focal length of the vehicle-mounted camera, v1 is the vertical coordinate value in the coordinate value of the vanishing point in the second image, and v0 is the vertical coordinate value in the coordinate value of the optical center in the second image.
[0253] Optionally, the second determination unit is configured to:
[0254] Determine the pitch angle of the vehicle-mounted camera according to the initial pitch angle and the calibrated pitch angle of the vehicle-mounted camera through the following formula;
[0255] θ = θ 初 + Δθ
[0256] where θ is the pitch angle of the vehicle-mounted camera, θ 初 is the initial pitch angle of the vehicle-mounted camera, and Δθ is the calibrated pitch angle of the vehicle-mounted camera.
[0257] In an embodiment of the present application, a first image and a second image are obtained. The first image and the second image are two adjacent video images in a video captured by an in-vehicle camera. Then, the first image and the second image are input into an optical flow estimation model to output an optical flow map through the optical flow estimation model. Since the optical flow map can accurately indicate the displacement values of pixel points belonging to the same object between the first image and the second image, determining the focal length of the in-vehicle camera according to the optical flow map can make the determined focal length of the in-vehicle camera relatively accurate. Finally, the coordinate value of the optical center in the second image and the coordinate value of the vanishing point in the second image are obtained, and the pitch angle of the in-vehicle camera is determined according to the focal length of the in-vehicle camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image. In this way, the focal length of the in-vehicle camera can be accurately determined according to the optical flow map, and the pitch angle of the in-vehicle camera can be determined according to the accurately determined focal length of the in-vehicle camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image, which can improve the accuracy of the pitch angle of the in-vehicle camera.
[0258] It should be noted that when determining the pitch angle of the in-vehicle camera provided in the above embodiment, only the division of the above functional modules is used as an example for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0259] Each functional unit and module in the above embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present application.
[0260] The in-vehicle camera pitch angle determination device provided in the above embodiment and the in-vehicle camera pitch angle determination method embodiment belong to the same concept. For the specific working processes and technical effects brought by the units and modules in the above embodiment, reference can be made to the method embodiment part, which will not be elaborated here.
[0261] Figure 8 This is a schematic structural diagram of a computer device provided in an embodiment of the present application. As Figure 8 shown, the computer device 8 includes: a processor 80, a memory 81, and a computer program 82 stored in the memory 81 and executable on the processor 80. When the processor 80 executes the computer program 82, the steps in the in-vehicle camera pitch angle determination method in the above embodiment are implemented.
[0262] The computer device 8 can be a general-purpose computer device or a special-purpose computer device. In a specific implementation, the computer device 8 can be a terminal. Optionally, it can be a vehicle-mounted terminal. The embodiments of the present application do not limit the type of the computer device 8. Those skilled in the art can understand that Figure 8 merely examples of the computer device 8 are provided, which do not constitute a limitation on the computer device 8. It may include more or fewer components than those shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0263] The processor 80 can be a central processing unit (CPU). The processor 80 can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0264] In some embodiments, the memory 81 can be an internal storage unit of the computer device 8, such as the hard disk or memory of the computer device 8. In other embodiments, the memory 81 can also be an external storage device of the computer device 8, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 8. Further, the memory 81 can also include both the internal storage unit and the external storage device of the computer device 8. The memory 81 is used to store an operating system, application programs, a boot loader, data, and other programs, etc. The memory 81 can also be used to temporarily store data that has been output or will be output.
[0265] The embodiments of the present application also provide a computer device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the processor executes the computer program, the steps in any of the above method embodiments are implemented.
[0266] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above method embodiments can be implemented.
[0267] An embodiment of the present application provides a computer program product which, when running on a computer, causes the computer to execute the steps in the above-mentioned method embodiments.
[0268] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned method embodiments of the present application, it can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented. Wherein, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code to the photographing device / terminal device, recording medium, computer memory, ROM (Read-Only Memory), RAM (Random Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk and optical data storage device, etc. The computer-readable storage medium mentioned in the present application can be a non-volatile storage medium, in other words, it can be a non-transitory storage medium.
[0269] It should be understood that all or part of the steps of implementing the above embodiments can be realized by software, hardware, firmware or any combination thereof. When implemented by software, it can be realized in the form of a computer program product entirely or partially. The computer program product includes one or more computer instructions. The computer instructions can be stored in the above-mentioned computer-readable storage medium.
[0270] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0271] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0272] In the embodiments provided in the present application, it should be understood that the disclosed device / computer equipment and method can be implemented in other ways. For example, the device / computer equipment embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0273] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0274] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for determining the pitch angle of a vehicle-mounted camera, characterized in that, Applied to a terminal, the method includes: Obtain a first image and a second image, where the first image is the previous frame of two adjacent video images in the video captured by a vehicle-mounted camera, and the second image is the subsequent frame of the two video images; Input the first image and the second image into an optical flow estimation model to output an optical flow map through the optical flow estimation model. The optical flow map is used to indicate the displacement value of pixel points belonging to the same object between the first image and the second image. The displacement value of pixel points belonging to the same object between the first image and the second image includes a horizontal displacement value; Determine the average value of the horizontal displacement values of all pixel points indicated by the optical flow map, obtain the angular velocity of the vehicle equipped with the vehicle-mounted camera, divide the average value by the angular velocity to obtain a first ratio, and determine the negative of the first ratio as the focal length of the vehicle-mounted camera; Obtain the coordinate value of the optical center in the second image and the coordinate value of the vanishing point in the second image; Determine the pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image.
2. The method according to claim 1, wherein The optical flow estimation model includes a first encoding module, a second encoding module, and a first decoding module; The first encoding module includes m first convolutional layers, the second encoding module includes m second convolutional layers. The m first convolutional layers are used to extract the feature map of the first image, and the m second convolutional layers are used to extract the feature map of the second image. m is an integer greater than or equal to 2; The first decoding module includes n optical flow estimation modules. The first optical flow estimation module among the n optical flow estimation modules is used to: output an optical flow map according to the feature map extracted by the m-th first convolutional layer among the m first convolutional layers and the feature map extracted by the m-th second convolutional layer among the m second convolutional layers; The i-th optical flow estimation module among the n optical flow estimation modules is used to: output an optical flow map according to the optical flow map output by the (i - 1)-th optical flow estimation module among the n optical flow estimation modules, the feature map extracted by the (m - i + 1)-th first convolutional layer among the m first convolutional layers, and the feature map extracted by the (m - i + 1)-th second convolutional layer among the m second convolutional layers; The optical flow map output by the n-th optical flow estimation module among the n optical flow estimation modules is the optical flow map output by the optical flow estimation model. n is an integer greater than or equal to 2, and i is an integer greater than or equal to 2 and less than or equal to n.
3. The method according to claim 2, wherein The outputting an optical flow map according to the optical flow map output by the (i - 1)-th optical flow estimation module among the n optical flow estimation modules, the feature map extracted by the (m - i + 1)-th first convolutional layer among the m first convolutional layers, and the feature map extracted by the (m - i + 1)-th second convolutional layer among the m second convolutional layers includes: Upsample the optical flow map output by the (i - 1)-th optical flow estimation module to obtain a target optical flow map; Transform the feature map extracted by the (m - i + 1)-th first convolutional layer according to the target optical flow map to obtain a target feature map; Determine the matching cost between the target feature map and the feature map extracted by the (m - i + 1)-th second convolutional layer; Output an optical flow map based on the target optical flow map, the matching cost, and the feature map extracted by the (m - i + 1)-th second convolutional layer.
4. The method according to any one of claims 1 to 3, characterized in that The obtaining the coordinate value of the vanishing point in the second image includes: Input the second image into a vanishing point prediction model to output a vanishing point probability matrix through the vanishing point prediction model. The probability value in the vanishing point probability matrix is the probability value that each pixel point in the second image is a vanishing point, and the pixel point corresponding to the maximum probability value in the vanishing point probability matrix is the vanishing point in the second image; Determine the coordinate value of the pixel point corresponding to the maximum probability value in the vanishing point probability matrix in the second image as the coordinate value of the vanishing point in the second image.
5. The method according to any one of claims 1 to 3, characterized in that The determining the pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image includes: Identify two lane lines of the road included in the second image; Determine the intersection coordinate value of the two lane lines in the second image in the second image; Determine the pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the intersection coordinate value, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image.
6. The method according to claim 5, wherein The determining the pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the intersection coordinate value, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image includes: Determine the initial pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image; Determine the corrected pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the intersection coordinate value, and the coordinate value of the vanishing point in the second image; Determine the pitch angle of the vehicle-mounted camera according to the initial pitch angle and the corrected pitch angle of the vehicle-mounted camera.
7. The method according to claim 6, wherein The determining the initial pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image includes: Determine the initial pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image through the following formula; where, is the initial pitch angle of the vehicle-mounted camera, is the focal length of the vehicle-mounted camera, is the ordinate value in the coordinate value of the vanishing point in the second image, and is the ordinate value in the coordinate value of the optical center in the second image.
8. The method according to claim 6, wherein The determining the pitch angle of the vehicle-mounted camera according to the initial pitch angle and the corrected pitch angle of the vehicle-mounted camera includes: Determine the pitch angle of the vehicle-mounted camera according to the initial pitch angle and the corrected pitch angle of the vehicle-mounted camera through the following formula; where, is the pitch angle of the vehicle-mounted camera, is the initial pitch angle of the vehicle-mounted camera, and is the corrected pitch angle of the vehicle-mounted camera.
9. A pitch angle determination device for a vehicle-mounted camera, characterized in that, The device includes: A first acquisition module, configured to acquire a first image and a second image, where the first image is the previous frame of two adjacent video images in a video captured by a vehicle-mounted camera, and the second image is the subsequent frame of the two video images; An optical flow generation module, configured to input the first image and the second image into an optical flow estimation model, so as to output an optical flow map through the optical flow estimation model, where the optical flow map is used to indicate the displacement value of pixel points belonging to the same object between the first image and the second image, and the displacement value of the pixel points belonging to the same object between the first image and the second image includes a horizontal displacement value; A first determination module, configured to determine the average value of the horizontal displacement values of all pixel points indicated by the optical flow map, obtain the angular velocity of the vehicle equipped with the vehicle-mounted camera, divide the average value by the angular velocity to obtain a first ratio, and determine the negative of the first ratio as the focal length of the vehicle-mounted camera; A second acquisition module, configured to acquire the coordinate value of the optical center in the second image and the coordinate value of the vanishing point in the second image; A second determination module, configured to determine the pitch angle of the vehicle-mounted camera according to the focal length of the vehicle-mounted camera, the coordinate value of the optical center in the second image, and the coordinate value of the vanishing point in the second image.
10. A computer device, characterized in that, The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the computer program is executed by the processor, it implements the method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Vehicle-mounted monocular camera external parameter self-calibration method
CN106875448A
Method and device for determining position and posture of imaging equipment
CN111220143A