Method, apparatus, medium, device and program product for determining a homography matrix
The camera image frame is processed through neural network models, the plane equation and inter-frame odometer of the road surface are determined, and the homography matrix is calculated in combination with the camera internal reference, which solves the problem of insufficient estimation accuracy and robustness of the road surface homography matrix in the prior art, and improves the accuracy of the estimation results.
Patent Information
- Application Number
- CN202111526426.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-14
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-12-14
AI Technical Summary
In the prior art, when estimating the homography matrix of the pavement, there are problems with low accuracy and robustness, especially when there is weak texture, violent light changes or reflection in the pavement area.
The two image frames captured by the camera are processed through the neural network model, and the plane equations of the inter-frame odometer and the target pavement are determined. Then, based on the inter-frame odometer, the plane equation and pre-stored camera internal reference, the homography matrix of the target pavement is determined.
The consistency between the homography matrix and the target pavement is improved, and the accuracy and robustness of the estimation results are enhanced.
Smart Images

Figure CN114170325B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a method, apparatus, storage medium, electronic device, and computer program product for determining a homography matrix. Background Art
[0002] A homography matrix (Homogrpahy) is used to describe the correspondence relationship of the imaging positions of points on a plane in three-dimensional space in two perspectives. The homography matrix of a road surface can be used to generate a bird's-eye view (Bird Eye View) or an inverse perspective mapping (Inverse Perspective Mapping), which has very important application value for autonomous driving and indoor robots.
[0003] In related technologies, the methods for estimating the homography matrix of a road surface are divided into a feature point-based method and a direct method. The feature point-based method first extracts feature points in the road surface area, then performs feature point matching through matching or tracking, and then restores the road surface homography matrix through the 5-point method or 8-point method combined with random sample consensus; the direct method can optimize the homography matrix by using the correspondence relationship of dense pixels in the image. Summary of the Invention
[0004] Embodiments of the present disclosure provide a method, apparatus, storage medium, electronic device, and computer program product for determining the homography matrix of a road surface.
[0005] According to one aspect of the embodiments of the present disclosure, a method for determining a homography matrix is provided, including: acquiring a first image frame and a second image frame captured by a camera, where the first image frame and the second image frame have a target road surface in the same area; processing the first image frame and the second image frame by using a neural network model to determine an inter-frame odometer; determining a plane equation of the target road surface based on the first image frame; and determining the homography matrix of the target road surface based on the inter-frame odometer, the plane equation, and a pre-stored camera internal parameter.
[0006] According to another aspect of the embodiments of the present disclosure, a method for training a neural network model is provided, including: obtaining multiple groups of sample image pairs in a training set and the sample camera intrinsic parameters of each sample image pair, each group of sample image pairs including a first sample image frame and a second sample image frame, the first sample image frame including a first mask of a sample road surface area, and the second sample image frame including a second mask of the sample road surface area; processing the sample image pairs by using a first initial network branch and a second initial network branch in a pre-constructed initial neural network model to obtain a sample plane equation of the sample road surface and a sample inter-frame odometer; obtaining a sample homography matrix based on the sample camera intrinsic parameters, the sample plane equation, and the sample inter-frame odometer; determining a sample mapped image based on the second sample image frame and the sample homography matrix; determining a global photometric consistency loss based on the sample mapped image and the first sample image frame; determining the sample road surface area in the sample mapped image based on the sample mapped image and the second mask; determining the sample road surface area in the first sample image based on the first mask and the first sample image; determining a road surface photometric consistency loss based on the sample road surface area in the first sample image and the sample road surface area in the sample mapped image; and training the initial neural network model based on the global photometric consistency loss and the road surface photometric consistency loss to obtain a trained neural network model.
[0007] According to another aspect of the embodiments of the present disclosure, a device for determining a homography matrix is provided, including: an image acquisition unit configured to acquire a first image frame and a second image frame captured by a camera, the first image frame and the second image frame having a target road surface in the same area; an inter-frame odometer unit configured to process the first image frame and the second image frame by using a neural network model to determine an inter-frame odometer; a plane equation unit configured to determine a plane equation of the target road surface based on the first image frame; and a matrix determination unit configured to determine a homography matrix of the target road surface based on the inter-frame odometer, the plane equation, and pre-stored camera intrinsic parameters.
[0008] According to another aspect of the embodiments of the present disclosure, there is provided an apparatus for training a neural network model, including: a sample acquisition unit configured to acquire multiple groups of sample image pairs in a training set and the sample camera intrinsic parameters of each sample image pair, each group of sample image pairs including a first sample image frame and a second sample image frame, the first sample image frame including a first mask of a sample road surface area, and the second sample image frame including a second mask of the sample road surface area; a first processing unit configured to process the sample image pairs by using a first initial network branch and a second initial network branch in a pre-constructed initial neural network model to obtain a sample plane equation of the sample road surface and a sample inter-frame odometer; a second processing unit configured to obtain a sample homography matrix based on the sample camera intrinsic parameters, the sample plane equation, and the sample inter-frame odometer; a sample mapping unit configured to determine a sample mapped image based on the second sample image frame and the sample homography matrix; a first loss unit configured to determine a global photometric consistency loss based on the sample mapped image and the first sample image frame; a second loss unit configured to: determine the sample road surface area in the sample mapped image based on the sample mapped image and the second mask; determine the sample road surface area in the first sample image based on the first mask and the first sample image; determine a road surface photometric consistency loss based on the sample road surface area in the first sample image and the sample road surface area in the sample mapped image; a model training unit configured to train the initial neural network model based on the global photometric consistency loss and the road surface photometric consistency loss to obtain a trained neural network model.
[0009] According to another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program for executing the method in any of the above embodiments.
[0010] According to another aspect of the embodiments of the present disclosure, an electronic device includes: a processor; a memory for storing executable instructions of the processor; the processor for executing the method in any of the above embodiments.
[0011] According to another aspect of the embodiments of the present disclosure, there is provided a computer program product including computer program / instructions, wherein when the computer program / instructions are executed by a processor, the method in any of the above embodiments is implemented.
[0012] A method, apparatus, storage medium, and electronic device for determining a homography matrix provided in the above embodiments of the present disclosure can determine the inter-frame odometry and the plane equation of the target road surface between two image frames captured by a camera through a neural network model, and then determine the homography matrix of the target road surface based on the inter-frame odometry, the camera internal parameters, and the plane equation of the road surface. The homography matrix of the road surface obtained in this way improves the consistency between the homography matrix and the geometric information of the target road surface, which helps to improve the accuracy.
[0013] The technical solutions of the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] By describing the embodiments of the present disclosure in more detail with reference to the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure, but do not constitute a limitation to the present disclosure. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0015] Figure 1 It is a schematic diagram of the scenario applicable to the method for determining the homography matrix of the present disclosure;
[0016] Figure 2 It is a flowchart of an embodiment of the method for determining the homography matrix of the present disclosure;
[0017] Figure 3 It is a flowchart of determining the homography matrix in an embodiment of the method for determining the homography matrix of the present disclosure;
[0018] Figure 4 It is a flowchart of determining the plane equation in an embodiment of the method for determining the homography matrix of the present disclosure;
[0019] Figure 5 It is a flowchart of determining the road surface offset in an embodiment of the method for determining the homography matrix of the present disclosure;
[0020] Figure 6 It is a flowchart of an embodiment of the method for training the neural network model of the present disclosure;
[0021] Figure 7 It is a schematic framework diagram of an example of the method for training the neural network model of the present disclosure;
[0022] Figure 8 It is a schematic structural diagram of an embodiment of the apparatus for determining the homography matrix of the present disclosure;
[0023] Figure 9Structural schematic diagram of an embodiment of the apparatus for training a neural network according to the present disclosure;
[0024] Figure 10 It is a structural diagram of an electronic device provided by an exemplary embodiment of the present application. Detailed implementation manners
[0025] Hereinafter, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the exemplary embodiments described herein.
[0026] It should be noted that: Unless otherwise specifically stated, the relative arrangements, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.
[0027] Those skilled in the art can understand that the terms "first", "second", etc. in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, etc., and do not represent any specific technical meaning, nor do they represent an inevitable logical order between them.
[0028] It should also be understood that in the embodiments of the present disclosure, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.
[0029] It should also be understood that for any component, data or structure mentioned in the embodiments of the present disclosure, unless otherwise clearly defined or given a contrary indication in the context, it can generally be understood as one or more.
[0030] In addition, the term "and / or" in the present disclosure is only a description of the association relationship of the associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the associated objects before and after.
[0031] It should also be understood that the present disclosure emphasizes the differences between the various embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated one by one.
[0032] At the same time, it should be understood that for the sake of description, the dimensions of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0033] The following description of at least one exemplary embodiment is actually only illustrative and in no way limits the present disclosure or its application or use.
[0034] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but in appropriate cases, the said technologies, methods, and devices should be regarded as part of the specification.
[0035] It should be noted that like reference numerals and letters refer to like items in the following figures, and thus, once an item is defined in one figure, further discussion thereof in subsequent figures is not required.
[0036] Embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate together with many other general or special computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, or servers include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.
[0037] Terminal devices, computer systems, servers, and other electronic devices can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment. In a distributed cloud computing environment, tasks can be executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.
[0038] Overview of the present disclosure
[0039] In the process of implementing the present disclosure, the inventors found that the method based on feature points has a high dependence on the quality of feature point detection and matching. When there are a large number of weak textures, drastic illumination changes, or reflections in the road surface area, the accuracy and robustness of this method are relatively low. The method based on the direct method has a high dependence on the initial value. When the initial value deviates far, it is difficult to obtain accurate results.
[0040] From the above description, it can be seen that the methods for determining the homography matrix in the related art all have defects in terms of accuracy.
[0041] Exemplary overview
[0042] The present disclosure can utilize a neural network model to determine an inter-frame odometer based on two image frames captured by a camera, and determine a road surface equation of a target road surface based on the first image frame. Then, based on the inter-frame odometer, the road surface equation, and the pre-calibrated camera internal parameters, a homography matrix of the target road surface can be determined, which can improve the consistency between the homography matrix and the geometric information of the target road surface and help improve the accuracy. An example is as Figure 1 shown.
[0043] Figure 1 As shown in ,
[0044] , Exemplary method , ,
[0045] , on the driverless vehicle 100, a camera 110 and an on-vehicle computer (not shown in the figure) are loaded. Computer instructions of the neural network model are pre-stored in the on-vehicle computer. The camera 110 can collect images of the target road surface 120 in real time and store the images in the local storage space. The on-vehicle computer can extract the images captured at the current moment from the storage space as the first image frame 130, and use the historical image closest to the current moment as the second image frame 140. Then, the on-vehicle computer can use the neural network model 150 to process the first image frame 130 and the second image frame 140 to obtain the inter-frame odometer and the plane equation of the target road surface 120, and determine the homography matrix of the target road surface based on the inter-frame odometer, the plane equation, and the camera internal parameters.
[0044] Exemplary method
[0045] Next, referring to Figure 2 , Figure 2 shows a flowchart of an embodiment of the method for determining a homography matrix according to the present disclosure. As Figure 2 shown, the process includes the following steps:
[0046] Step 210: Obtain a first image frame and a second image frame captured by the camera.
[0047] Among them, the first image frame and the second image frame have the target road surface in the same area.
[0048] In this embodiment, the first image frame and the second image frame may be images of the target road surface captured by the camera at different times.
[0049] In a specific example, the camera may be an on-vehicle camera of the driverless vehicle, and the execution subject may be the on-vehicle computer of the driverless vehicle. The camera internal parameters of the on-vehicle camera are pre-stored in the on-vehicle computer. During the driving process of the driverless vehicle, the camera captures the target road surface in real time (for example Figure 1Images of the target road surface 120) area in it are obtained to form an image dataset composed of images of the target road surface area at different times. The acquisition frequency of the camera can be set according to the driving speed of the driverless vehicle and the acquisition range of the camera lens, so that there is the target road surface in the same area between at least two adjacent image frames. Then, two adjacent image frames can be selected from the image dataset as the first image frame and the second image frame. For example, the image frame at the current moment can be used as the first image frame, and the historical image frame closest to the current moment can be used as the second image frame.
[0050] Step 220: Process the first image frame and the second image frame using a neural network model to determine the inter-frame odometer.
[0051] In this embodiment, the inter-frame odometer represents the attitude change when the camera acquires the first image frame and the second image frame. The inter-frame odometer usually includes parameters with six degrees of freedom, representing the attitude change from two dimensions of rotation and translation respectively. As an example, the inter-frame odometer can be a matrix (R, t) composed of quaternions, where R represents the rotation matrix composed of quaternions, and t represents the translation matrix composed of quaternions.
[0052] As an example, the execution entity (such as an in-vehicle computer or other terminal device) can use the first image frame and the second image frame as input data, and process the first image frame and the second image frame using a neural network. First, through semantic segmentation, the semantic labels of each pixel point in the first image frame and the second image frame are predicted, and the two-dimensional coordinates of the center points of each object included in the first image frame and the second image frame are predicted based on the semantic labels of the pixel points; then, the three-dimensional postures of each object in the first image frame and the second image frame are predicted respectively based on the semantic labels of the pixel points and the two-dimensional coordinates of the center points of the objects; finally, by comparing the three-dimensional postures of the same object in the first image frame and the second image frame, the inter-frame odometer of the first image frame and the second image frame is determined.
[0053] Step 230: Based on the first image frame, determine the plane equation of the target road surface.
[0054] In this embodiment, the plane equation represents the geometric information of the pixel points belonging to the target road surface in the first image frame.
[0055] As an example, the execution entity can process the first image frame using a neural network, estimate the normal vector of the target road surface, the three-dimensional coordinates of the pixel points of the target road surface, and the height when the camera acquires the first image frame, and then obtain the plane equation shown in formula (1) below.
[0056]
[0057] In the formula, i represents the moment when the first image frame is acquired. Denote the normal vector of the target road surface at time i, P represents the three-dimensional coordinates of the pixel point, and h i represents the height at which the camera captures the first image frame.
[0058] Step 240: Determine the homography matrix of the target road surface based on the inter-frame odometer, the plane equation, and the pre-stored camera internal parameters.
[0059] In a specific example, the execution entity can be the in-vehicle computer of an autonomous vehicle, where the internal parameters K of the camera are pre-stored. The execution entity can use the image captured by the camera at the current moment as the first image frame, then use the image adjacent to the first image frame in the historical image frames as the second image frame, and then use the neural network model to process the first image frame and the second image frame to predict the inter-frame odometer (R, t) and the plane equation (as shown in Equation (1)), and then determine the homography matrix H of the target road surface through the following Equation (2), and the homography matrix of the target road surface can be determined in real time.
[0060]
[0061] The method for determining the homography matrix provided in this embodiment can determine the inter-frame odometer between two image frames captured by the camera and the plane equation of the target road surface through the neural network model, and then determine the homography matrix of the target road surface based on the inter-frame odometer, the camera internal parameters, and the plane equation of the road surface. The homography matrix of the road surface obtained in this way improves the consistency between the homography matrix and the geometric information of the target road surface, which helps to improve the accuracy.
[0062] Next, refer to Figure 3 , Figure 3 which shows the process of determining the homography matrix in an embodiment of the method for the homography matrix of the present disclosure. In some optional implementation manners of the embodiment shown in Figure 2 , the above step 240 can also adopt the process shown in Figure 3 , and this process includes the following steps:
[0063] Step 310: Determine the current height of the camera and the current normal vector of the target road surface based on the plane equation.
[0064] In this implementation manner, the current height of the camera represents the height at which the camera captures the first image frame. As shown in Equation (1), the execution entity can extract h t as the current height and extract as the current normal vector of the target road surface.
[0065] Step 320: Determine the inter-frame rotation matrix and the inter-frame translation vector based on the inter-frame odometer.
[0066] As an example, the inter-frame odometer can be (R, t), and the execution entity can extract R as the inter-frame rotation matrix and extract t as the inter-frame translation vector.
[0067] Step 330: Determine the first matrix based on the ratio of the product of the current normal vector and the inter-frame translation vector to the current height.
[0068] As an example, the execution entity can determine the first matrix A through the following formula (3):
[0069]
[0070] Step 340: Determine the second matrix based on the difference between the inter-frame rotation matrix and the first matrix.
[0071] As an example, based on formula (3), the execution entity can determine the second matrix B through the following formula (4):
[0072] B = R - A (4)
[0073] Step 350: Determine the homography matrix of the target road surface based on the second matrix, the camera internal parameters, and the inverse matrix of the camera internal parameters.
[0074] As an example, based on formulas (3) and (4), the execution entity can determine the homography matrix H of the target road surface through the following formula (5):
[0075] H = KBK -1 (5)
[0076] From Figure 3 It can be seen that Figure 3 The shown process reflects the steps of determining the current normal vector of the target road surface and the current height of the camera from the plane equation, determining the inter-frame rotation matrix and the inter-frame translation vector from the inter-frame odometer, and determining the homography matrix therefrom. In this process, the consistency between the calculation process of the homography matrix and the geometric information of the target road surface can be ensured, thereby further improving the accuracy of the homography matrix.
[0077] In some alternative implementation manners of this embodiment, Figure 2 the neural network model in the shown process includes a first network branch; the above step 220 can be implemented in the following manner: using the first network branch, encoding the first image frame and the second image frame to obtain a feature vector; decoding the feature vector to obtain the inter-frame odometer.
[0078] In this implementation manner, the first network branch may include an encoder and a decoder. The execution entity may encode the first image frame and the second image frame through the encoder to extract image features, obtaining a high-dimensional feature vector. The feature vector may characterize the image features of the first image frame and the second image frame. For example, it may include semantic information, pose information, etc. of pixel points. Then, the decoder is used to decode the feature vector, and based on the extracted image features, the pose change of the camera is determined to obtain the inter-frame odometer.
[0079] In this implementation manner, the first network branch of the neural network may be used to encode the first image frame and the second image frame, extract the high-dimensional features of the first image frame and the second image frame, and thereby determine the inter-frame odometer, which can improve the adaptability of the neural network to weakly textured or highly reflective images, and further improve the accuracy of the inter-frame odometer.
[0080] In Figure 2 In some optional implementation manners of the illustrated embodiments, the neural network model includes a second network branch for determining the plane equation of the target road surface. Further refer to Figure 4 , Figure 4 shows the process of determining the plane equation in an embodiment of the method for determining the homography matrix of the present disclosure. As shown in Figure 4 shown, the above step 230 may include the following steps:
[0081] Step 410: Process the first image frame by using the second network branch to determine the road surface offset of the target road surface.
[0082] In this implementation manner, the road surface offset characterizes the change degree of the road surface at the current moment (i.e., the moment when the first image frame is acquired) relative to the road surface pose at the initial moment, and may include the height change and rotation change of the road surface. As an example, the height change amount of the target road surface may be characterized by the height residual of the camera, and the rotation change amount of the target road surface may be characterized by the normal vector residual of the target road surface.
[0083] As an example, the execution entity may process the first image frame by using the second branch network in the neural network model, extract the image features of the first image frame, and estimate the height residual and normal vector residual of the target road surface at the current moment based on the image features, thereby obtaining the road surface offset.
[0084] Step 420: Based on the initial normal vector of the pre-stored target road surface, the road surface offset, and the initial height of the pre-stored camera, determine the plane equation of the target road surface.
[0085] In this implementation manner, the initial normal vector characterizes the normal vector of the target road surface at the initial moment. The camera may be pre-calibrated to determine the initial normal vector of the target road surface at the initial moment and the initial height of the camera.
[0086] As an example, the execution entity may first determine the height residual and the normal vector residual of the target road surface from the road surface offset to obtain a height residual matrix and a normal vector residual matrix; then, as shown in the following formula (6), the execution entity may determine the matrix obtained by cross-multiplying the initial normal vector and the normal vector residual matrix as the current normal vector; as shown in the following formula (7), the execution entity may determine the sum of the height residual matrix and the initial height as the current height; thereafter, the execution entity may determine the plane equation of the target road surface through the above formula (1).
[0087]
[0088]
[0089] In the formula, N i represents the current normal vector at the i-th moment, δR i represents the normal vector residual at the i-th moment, represents the initial normal vector, h i represents the current height of the camera at the i-th moment, δh i represents the height residual of the camera at the i-th moment, represents the initial height of the camera.
[0090] Through Figure 4 it can be seen that Figure 4 the process shown reflects estimating the road surface offset of the target road surface through the second network branch in the neural network, and then combining the pre-calibrated initial normal vector and initial height to determine the current normal vector and current height of the target road surface, and further determining the plane equation of the target road surface, which can reduce the data calculation amount in the process of determining the plane equation, improve the calculation speed, and is especially suitable for online real-time correction of the plane equation.
[0091] Next, referring to Figure 5 , Figure 5 shows the process of determining the road surface offset in an embodiment of the method for determining the homography matrix according to the present disclosure. As Figure 5 shown, the above step 410 may further include the following steps:
[0092] Step 510: Extract multiple image features from the first image frame by using convolutional layers with different resolutions in the first network branch, and fuse the multiple image features to obtain fused image features.
[0093] Step 520: Estimate the first offset angle, the second offset angle, and the height offset of the target road surface based on the fused image features.
[0094] Generally, the rotation action in the three-dimensional space can be decomposed into a yaw angle component, a pitch angle component, and a roll angle component.
[0095] In this implementation manner, the first offset angle may represent the yaw angle component of the target road surface normal vector residual, and the second offset angle may represent the pitch angle component of the rotation amount. The height offset may represent the height residual of the target road surface. As an example, the first offset angle and the second offset angle may be matrices composed of quaternions.
[0096] It should be noted that since the roll angle component has no influence on the offset of the normal vector, that is, the roll angle component has nothing to do with the rotational change of the target road surface, therefore, this implementation manner may ignore the roll angle component.
[0097] Step 530: Determine the rotational offset of the target road surface based on the first offset angle and the second offset angle.
[0098] As an example, the execution subject may use the matrix obtained by cross-multiplying the first offset angle and the second offset angle as the rotational offset of the target road surface.
[0099] Step 540: Determine the road surface offset of the target road surface based on the rotational offset and the height offset.
[0100] In Figure 5 the shown process, the first image frame may be processed through the first network branch in the neural network to estimate the first offset angle, the second offset angle, and the height offset of the target road surface, and then the rotational offset is determined based on the first offset angle and the second offset angle, and further the road surface offset is obtained, ignoring the roll angle component that has nothing to do with the rotational change of the target road surface, which can reduce the amount of operation data and improve the operation efficiency.
[0101] Next, refer to Figure 6 , Figure 6 which shows a flowchart of an embodiment of the method for training a neural network model according to the present disclosure. As Figure 6 shown, the process includes the following steps:
[0102] Step 610: Obtain multiple groups of sample image pairs in the training set and the sample camera internal parameters of each sample image pair.
[0103] In this embodiment, each group of sample image pairs includes a first sample image frame and a second sample image frame. The first sample image frame includes a first mask of the sample road surface area, and the second sample image frame includes a second mask of the sample road surface area. The first mask and the second mask may be matrices composed of pixel values. Among them, the pixel values of the pixel points in the sample road surface area are set to 1, and the pixel values of the pixel points in the non-sample area are set to 0.
[0104] As an example, the execution entity for training the neural network can be a terminal device or a server. The execution entity can obtain sample image data and sample camera internal parameters from a public dataset through the network, or select sample image data from the image set captured by the camera, determine two images with the same sample road surface area as a sample image pair, and then perform semantic segmentation on the first sample image frame and the second sample image frame in the sample image pair respectively to determine the first mask and the second mask therein.
[0105] Step 620: Process the sample image pair using the first initial network branch and the second initial network branch in the pre-constructed initial neural network model to obtain the sample plane equation of the sample road surface and the sample inter-frame odometer.
[0106] In this embodiment, the first initial network branch and the second initial network branch represent sub-networks in the initial neural network model that have not been completed training.
[0107] As an example, the first initial network branch can include an encoder and a decoder to be optimized, perform encoding and decoding on the first sample image frame and the second sample image frame to obtain the sample inter-frame odometer, and the second initial network branch can include convolutional layers and pooling layers with multiple resolutions to perform feature extraction and feature mapping on the first sample image frame to obtain the sample plane equation of the sample road surface.
[0108] Step 630: Based on the sample camera internal parameters, the sample plane equation, and the sample inter-frame odometer, obtain the sample homography matrix.
[0109] As an example, the initial neural network model can first extract the sample normal vector and the current height of the sample camera from the sample plane equation. The sample plane equation is shown in formula (8):
[0110]
[0111] In the formula, represents the sample normal vector of the sample road surface at time j, h SAMj represents the current height of the sample camera at the j-th moment (i.e., the camera height when collecting the sample image), P SAM represents the three-dimensional coordinates of the pixel points in the sample image.
[0112] After that, the execution entity can substitute the sample camera internal parameters, the sample inter-frame odometer, and the sample normal vector and the current height of the sample camera into the following formula (9) to obtain the sample homography matrix.
[0113]
[0114] In the formula, H SAM represents the sample homography matrix, K SAMdenotes the intrinsic parameters of the sample camera, (R SAM , t SAM ) denotes the inter-frame odometry of the sample, R SAM denotes the inter-frame rotation matrix of the sample, t SAM denotes the inter-frame translation amount of the sample.
[0115] Step 640: Determine the sample mapped image based on the second sample image frame and the sample homography matrix.
[0116] In this embodiment, the execution subject can perform a perspective transformation on the pixel points in the second sample image frame based on the sample homography matrix to obtain the sample mapped image. The perspective transformation can be characterized by the processing process expressed by the following formula (10) for example.
[0117] P2 = H 2→1 P1 (10)
[0118] In the formula, P2 represents the pixel coordinates of the pixel points in the second sample image frame, H 2→1 represents the sample homography matrix, and P1 represents the pixel coordinates of the pixel points in the sample mapped image.
[0119] Step 650: Determine the global photometric consistency loss based on the sample mapped image and the first sample image frame.
[0120] In this embodiment, the execution subject can determine the global photometric consistency loss based on the pixel values of the corresponding pixel points in the sample mapped image and the first sample image frame.
[0121] As an example, the execution subject can determine the pixel values of the corresponding pixel points in the sample mapped image and the first sample image frame through the following formula (11):
[0122]
[0123] In the formula, represents the pixel value of the nth pixel point in the sample mapped image, where, [P n represents the coordinates of the nth pixel point in the sample mapped image; represents the pixel interpolation of the (n - 1)th pixel point in the second sample image frame, [P n-1 represents the coordinates of the (n - 1)th pixel point in the second sample image frame.
[0124] After that, the execution subject can determine the pixel point pairs from the sample mapped image and the first sample image frame according to the pixel point coordinates. Then, based on the pixel values of the pixel point pairs, the global photometric consistency loss is determined using the norm loss function, and the calculation process is shown in formula (12).
[0125]
[0126] Where E H represents the global photometric consistency loss, represents the pixel at the nth pixel point in the first sample image frame, and m represents the total number of pixel points.
[0127] Step 660: Based on the sample mapping image and the second mask, determine the sample road surface area in the sample mapping image.
[0128] In this embodiment, the execution subject can perform a convolution operation on the sample mapping image and the second mask, and convert the pixel values of the pixel points in the non-sample road surface area in the sample mapping image to 0, so as to obtain the sample road surface area in the sample mapping image.
[0129] Step 670: Based on the first mask and the first sample image, determine the sample road surface area in the first sample image.
[0130] In this embodiment, the execution subject can perform a convolution operation on the first sample image frame and the first mask, and convert the pixel values of the pixel points in the non-sample road surface area in the first sample image frame to 0, so as to obtain the sample road surface area in the first sample image frame.
[0131] Step 680: Based on the sample road surface area in the first sample image and the sample road surface area in the sample mapping image, determine the road surface photometric consistency loss.
[0132] As an example, the execution subject can substitute the pixel values of the pixel points in the sample road surface area in the first sample image and the sample road surface area in the sample mapping image into the above formula (12) to obtain the road surface photometric consistency loss.
[0133] Step 690: Based on the global photometric consistency loss and the road surface photometric consistency loss, train the initial neural network model to obtain the trained neural network model.
[0134] In the method for training the neural network model in this embodiment, the first mask and the second mask in the sample image pair are used as the labeled data of the sample, so as to determine the road surface photometric consistency loss during the training process, and based on the global photometric consistency loss and the road surface photometric consistency, the training process of the initial neural network model is constrained to obtain the trained neural network model. On the one hand, compared with using the homography matrix as the labeled data of the sample in the related technology, the acquisition method of the mask information is simpler. Combining with the weakly supervised training method, it is possible to use data in different scenarios and long-tail data for training, which helps to improve the generalization ability and robustness of the neural network model; on the other hand, the road surface photometric consistency loss determined based on the mask information can more accurately constrain the processing process of the sample road surface area, which helps to improve the accuracy of the neural network model.
[0135] In some alternative implementation manners of this embodiment, the initial neural network model may further include a third initial network branch. Before the above step 690, the method may further include: using the third initial network branch to generate a sample depth map of the first sample image frame; determining a depth smoothness loss based on the sample depth map; and the above step 690 may further include: training the initial neural network model based on the global photometric consistency loss, the road surface photometric consistency loss, and the depth smoothness loss.
[0136] Next, referring to Figure 7 , Figure 7 shows a schematic framework diagram of an example of the method for training a neural network model according to the present disclosure. In Figure 7 the example shown, the initial neural network includes a first initial network branch 730, a second initial network branch 740, a third initial network branch 750, and a homography matrix layer 760. Among them, the first initial network branch 730 can process the first sample image frame 710 and the second sample image frame 720 to obtain a sample inter-frame odometer; the second initial network branch 740 can process the first sample image frame to obtain a sample plane equation; the third initial network branch 750 can process the first sample image frame to obtain a sample depth map 770; and the homography matrix layer 760 can determine a sample homography matrix according to the sample inter-frame odometer, the sample plane equation, and the sample camera internal parameters. After that, the execution subject performs a perspective transformation on the second sample image frame 720 according to the sample homography matrix to obtain a sample mapped image 780. Then, the execution subject can determine the global photometric consistency loss according to the sample mapped image 780 and the first sample image frame 710, determine the depth smoothness loss according to the sample depth map 770, and determine the road surface photometric consistency loss according to the sample mapped image 780, the first sample image frame 710, the first mask 790, and the second mask 791.
[0137] In this implementation manner, a third initial network branch can be added to the initial neural network model to generate a sample depth map of the first sample image, so as to determine the depth smoothness loss in the training stage and add the depth smoothness loss to the constraints in the training stage, which can not only improve the convergence speed of the neural network model, but also improve the performance of the neural network model.
[0138] Exemplary device
[0139] Figure 8 is a schematic structural diagram of an embodiment of the device for determining a homography matrix according to the present disclosure. The device of this embodiment can be used to implement the corresponding method embodiment of the present disclosure. As Figure 8The device shown includes: an image acquisition unit 810 configured to acquire a first image frame and a second image frame captured by a camera, where the first image frame and the second image frame have a target road surface in the same area; an inter-frame odometer unit 820 configured to process the first image frame and the second image frame using a neural network model to determine an inter-frame odometer; a plane equation unit 830 configured to determine a plane equation of the target road surface based on the first image frame; and a matrix determination unit 840 configured to determine a homography matrix of the target road surface based on the inter-frame odometer, the plane equation, and a pre-stored camera internal parameter.
[0140] In this embodiment, the matrix determination unit 840 further includes: a normal vector module configured to determine a current height of the camera and a current normal vector of the target road surface based on the plane equation; an inter-frame odometer module configured to determine an inter-frame rotation matrix and an inter-frame translation vector based on the inter-frame odometer; a first matrix module configured to determine a first matrix based on the ratio of the product of the current normal vector and the inter-frame translation vector to the current height; a second matrix module configured to determine a second matrix based on the difference between the inter-frame rotation matrix and the first matrix; and a matrix determination module configured to determine a homography matrix of the target road surface based on the second matrix, the camera internal parameter, and the inverse matrix of the camera internal parameter.
[0141] In this embodiment, the neural network model includes a first network branch; the inter-frame odometer unit 820 further includes: an encoding module configured to encode the first image frame and the second image frame using the first network branch to obtain a feature vector; and a decoding module configured to decode the feature vector to obtain an inter-frame odometer.
[0142] In this embodiment, the neural network model includes a second network branch; the plane equation unit 830 further includes: an offset module configured to process the first image frame using the second network branch to determine a road surface offset of the target road surface; and an equation determination module configured to determine a plane equation of the target road surface based on a pre-stored initial normal vector of the target road surface, the road surface offset, and a pre-stored initial height of the camera.
[0143] In this embodiment, the offset module further includes: a feature extraction sub-module configured to extract multiple image features from the first image frame using convolutional layers with different resolutions in the first network branch and fuse the multiple image features to obtain a fused image feature; a first prediction sub-module configured to estimate a first offset angle, a second offset angle, and a height offset of the target road surface based on the fused image feature; a rotation sub-module configured to determine a rotation offset of the target road surface based on the first offset angle and the second offset angle; and an offset determination sub-module configured to determine a road surface offset of the target road surface based on the rotation offset and the height offset.
[0144] Next, refer to Figure 9 , Figure 9 which shows a schematic structural diagram of an embodiment of the apparatus for training a neural network model according to the present disclosure, for implementing an embodiment of the method for training a neural network model. As Figure 9 shown, the apparatus includes: a sample acquisition unit 910 configured to acquire multiple groups of sample image pairs in a training set and the sample camera internal parameters of each sample image pair, each group of sample image pairs including a first sample image frame and a second sample image frame, the first sample image frame including a first mask of a sample road surface area, and the second sample image frame including a second mask of the sample road surface area; a first processing unit 920 configured to process the sample image pairs by using a first initial network branch and a second initial network branch in a pre-constructed initial neural network model to obtain a sample plane equation of the sample road surface and a sample inter-frame odometer; a second processing unit 930 configured to obtain a sample homography matrix based on the sample camera internal parameters, the sample plane equation, and the sample inter-frame odometer; a sample mapping unit 940 configured to determine a sample mapped image based on the second sample image frame and the sample homography matrix; a first loss unit 950 configured to determine a global photometric consistency loss based on the sample mapped image and the first sample image frame; a second loss unit 960 configured to: determine the sample road surface area in the sample mapped image based on the sample mapped image and the second mask; determine the sample road surface area in the first sample image based on the first mask and the first sample image; determine a road surface photometric consistency loss based on the sample road surface area in the first sample image and the sample road surface area in the sample mapped image; a model training unit 970 configured to train the initial neural network model based on the global photometric consistency loss and the road surface photometric consistency loss to obtain a trained neural network model.
[0145] In this embodiment, the initial neural network model further includes a third initial network branch, and the apparatus further includes: a third processing unit configured to generate a sample depth map of the first sample image frame by using the third initial network branch; a third loss unit configured to determine a depth smoothness loss based on the sample depth map; and the model training unit 970 is further configured to: train the initial neural network model based on the global photometric consistency loss, the road surface photometric consistency loss, and the depth smoothness loss.
[0146] Exemplary electronic device
[0147] Next, refer to Figure 10 to describe an electronic device according to an embodiment of the present disclosure. Figure 10 shows a block diagram of an electronic device according to an embodiment of the present disclosure. As Figure 10 shown, the electronic device 1000 includes one or more processors 1010 and a memory 1020.
[0148] The processor 1010 can be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 1000 to perform desired functions.
[0149] The memory 1020 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory, for example, can include: random access memory (RAM) and / or cache, etc. The non-volatile memory, for example, can include: read-only memory (ROM), hard disk, and flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage media, and the processor 1010 can run the program instructions to implement the methods for determining a homography matrix, the methods for training a neural network model, and / or other desired functions according to various embodiments of the present disclosure described above. Various contents such as input signals, signal components, and noise components can also be stored in the computer-readable storage media.
[0150] In one example, the electronic device 1000 can further include: an input device 1030 and an output device 1040, etc., and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown). In addition, the input device 1030 can include, for example, a keyboard, a mouse, and so on. The output device 1040 can output various information to the outside. The output device 1040 can include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, and so on.
[0151] Of course, for simplicity, Figure 10 only some of the components related to the present disclosure in the electronic device 1000 are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application scenarios, the electronic device 1000 can further include any other appropriate components.
[0152] Exemplary computer program product and computer-readable storage medium
[0153] In addition to the above methods and devices, the embodiments of the present disclosure can also be a computer program product, which includes computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the methods for determining a homography matrix or the methods for training a neural network model according to various embodiments of the present disclosure described in the "Exemplary Methods" section above in this specification.
[0154] The computer program product may be written in any combination of one or more programming languages for executing the program code of the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0155] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium storing computer program instructions, which when run by a processor cause the processor to execute the steps in the methods for determining a homography matrix or for training a neural network model according to various embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0156] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive listing) of the readable storage medium may include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0157] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, and effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-mentioned specific details are only for illustrative purposes and for the convenience of understanding, rather than limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.
[0158] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system embodiments, since they basically correspond to the method embodiments, they are described relatively simply, and the relevant parts can be referred to the partial description of the method embodiments.
[0159] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including," "comprising," "having," etc. are open-ended terms, meaning "including but not limited to," and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The phrase "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.
[0160] The methods and apparatuses of the present disclosure can be implemented in many ways. For example, the methods and apparatuses of the present disclosure can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of the steps for the methods is for illustrative purposes only, and the steps of the methods of the present disclosure are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure can also be implemented as a program recorded on a recording medium, and these programs include machine-readable instructions for implementing the methods according to the present disclosure. Therefore, the present disclosure also covers the recording medium storing the programs for executing the methods according to the present disclosure.
[0161] It should also be noted that in the apparatuses, equipment, and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.
[0162] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
[0163] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.
Claims
1. A method for determining a homography matrix, comprising: Obtaining a first image frame and a second image frame captured by a camera, wherein the first image frame and the second image frame have a target road surface in the same area; Processing the first image frame and the second image frame by using a neural network model to determine an inter-frame odometer; Based on the first image frame, determining a plane equation of the target road surface; Based on the inter-frame odometer, the plane equation, and pre-stored camera intrinsics, determining the homography matrix of the target road surface; Wherein, determining the inter-frame odometer includes: through semantic segmentation, predicting semantic labels of each pixel point in the first image frame and the second image frame, and predicting two-dimensional coordinates of the center points of each object included in the first image frame and the second image frame based on the semantic labels of each pixel point; predicting three-dimensional poses of each object in the first image frame and the second image frame respectively based on the semantic labels of each pixel point and the two-dimensional coordinates of the center points of each object; and determining the inter-frame odometer of the first image frame and the second image frame based on the three-dimensional poses of the same object in the first image frame and the second image frame.
2. The method according to claim 1, wherein Based on the inter-frame odometer, the plane equation, and pre-stored camera intrinsics, determining the homography matrix of the target road surface includes: Based on the plane equation, determining the current height of the camera and the current normal vector of the target road surface; Based on the inter-frame odometer, determining an inter-frame rotation matrix and an inter-frame translation vector; Determining a first matrix based on the ratio of the product of the current normal vector and the inter-frame translation vector to the current height; Determining a second matrix based on the difference between the inter-frame rotation matrix and the first matrix; Based on the second matrix, the camera intrinsics, and the inverse matrix of the camera intrinsics, determining the homography matrix of the target road surface.
3. The method according to claim 1 or 2, wherein The neural network model includes a first network branch; Processing the first image frame and the second image frame to determine an inter-frame odometer includes: Encoding the first image frame and the second image frame by using the first network branch to obtain feature vectors; Decoding the feature vectors to obtain the inter-frame odometer.
4. The method according to claim 1 or 2, wherein, The neural network model includes a second network branch; Based on the first image frame, determining a plane equation of the target road surface includes: Processing the first image frame by using the second network branch to determine a road surface offset of the target road surface; Based on the initial normal vector of the pre-stored target road surface, the road surface offset, and the initial height of the pre-stored camera, determining the plane equation of the target road surface.
5. The method according to claim 4, wherein Processing the first image frame by using the second network branch to determine a road surface offset of the target road surface includes: Extracting multiple image features from the first image frame by using convolutional layers with different resolutions in the second network branch, and fusing the multiple image features to obtain fused image features; Based on the fused image features, estimating a first offset angle, a second offset angle, and a height offset of the target road surface; Based on the first offset angle and the second offset angle, determining a rotation offset of the target road surface; Determine the road surface offset of the target road surface based on the rotational offset and the height offset.
6. A method for training a neural network model, comprising: Obtain multiple groups of sample image pairs in a training set and the sample camera intrinsics of each sample image pair. Each group of sample image pairs includes a first sample image frame and a second sample image frame. The first sample image frame includes a first mask of a sample road surface area, and the second sample image frame includes a second mask of the sample road surface area; Process the sample image pairs using a first initial network branch and a second initial network branch in a pre-constructed initial neural network model to obtain a sample plane equation and a sample inter-frame odometer of the sample road surface; Based on the sample camera intrinsics, the sample plane equation, and the sample inter-frame odometer, obtain a sample homography matrix; Based on the second sample image frame and the sample homography matrix, determine a sample mapped image; Based on the sample mapped image and the first sample image frame, determine a global photometric consistency loss; Based on the sample mapped image and the second mask, determine the sample road surface area in the sample mapped image; Based on the first mask and the first sample image frame, determine the sample road surface area in the first sample image frame; Based on the sample road surface area in the first sample image frame and the sample road surface area in the sample mapped image, determine the road surface photometric consistency loss; Based on the global photometric consistency loss and the road surface photometric consistency loss, train the initial neural network model to obtain a trained neural network model, and the trained neural network model is used to implement the method for determining the homography matrix according to any one of claims 1-5 above.
7. The method according to claim 6, wherein The initial neural network model further includes a third initial network branch. Before training the initial neural network model based on the global photometric consistency loss and the road surface photometric consistency loss, it further includes: Generate a sample depth map of the first sample image frame using the third initial network branch; Based on the sample depth map, determine a depth smoothness loss; Training the initial neural network model based on the global photometric consistency loss and the road surface photometric consistency loss further includes: Based on the global photometric consistency loss, the road surface photometric consistency loss, and the depth smoothness loss, train the initial neural network model.
8. A device for determining the homography matrix of a road surface, comprising: An image acquisition unit configured to acquire a first image frame and a second image frame captured by a camera, where the first image frame and the second image frame have a target road surface in the same area; An inter-frame odometer unit configured to process the first image frame and the second image frame using a neural network model to determine an inter-frame odometer; A plane equation unit configured to determine a plane equation of the target road surface based on the first image frame; A matrix determination unit configured to determine the homography matrix of the target road surface based on the inter-frame odometer, the plane equation, and the pre-stored camera intrinsics; The inter-frame odometry unit is configured to predict the semantic labels of each pixel point in the first image frame and the second image frame through semantic segmentation, and predict the two-dimensional coordinates of the center points of each object included in the first image frame and the second image frame based on the semantic labels of each pixel point; predict the three-dimensional poses of each object in the first image frame and the second image frame respectively based on the semantic labels of each pixel point and the two-dimensional coordinates of the center points of each object; Determine the inter-frame odometry of the first image frame and the second image frame based on the three-dimensional poses of the same object in the first image frame and the second image frame.
9. An apparatus for training a neural network model, comprising: A sample acquisition unit configured to acquire multiple groups of sample image pairs in a training set and the sample camera internal parameters of each sample image pair, each group of sample image pairs including a first sample image frame and a second sample image frame, the first sample image frame including a first mask of a sample road surface area, and the second sample image frame including a second mask of the sample road surface area; A first processing unit configured to process the sample image pair by using a first initial network branch and a second initial network branch in a pre-constructed initial neural network model to obtain a sample plane equation and a sample inter-frame odometry of the sample road surface; A second processing unit configured to obtain a sample homography matrix based on the sample camera internal parameters, the sample plane equation, and the sample inter-frame odometry; A sample mapping unit configured to determine a sample mapped image based on the second sample image frame and the sample homography matrix; A first loss unit configured to determine a global photometric consistency loss based on the sample mapped image and the first sample image frame; A second loss unit configured to determine the sample road surface area in the sample mapped image based on the sample mapped image and the second mask; Determine the sample road surface area in the first sample image frame based on the first mask and the first sample image frame; Determine the road surface photometric consistency loss based on the sample road surface area in the first sample image frame and the sample road surface area in the sample mapped image; A model training unit configured to train the initial neural network model based on the global photometric consistency loss and the road surface photometric consistency loss to obtain a trained neural network model, and the trained neural network model is used to implement the method for determining a homography matrix according to any one of claims 1-5 above.
10. A computer-readable storage medium storing a computer program for executing the method according to any one of claims 1-7 above.
11. An electronic device, the electronic device comprising: A processor; A memory for storing executable instructions of the processor; The processor is configured to execute the method according to any one of claims 1-7 above.
12. A computer program product comprising a computer program / instructions, wherein, When the computer program / instructions are executed by the processor, the method according to any one of claims 1-7 above is implemented.
Citation Information
Patent Citations
Method, device, equipment and medium for determining pose of camera
CN111325792A
Image depth estimation method and device, readable storage medium and electronic equipment
CN112381868A
Method and device for adjusting homography matrix parameters
CN113592706A