Training Method of Pose Estimation Network Model, Pose Estimation Method and Device
By using predicted depth information and supervised depth information to train the pose estimation network model, the problem of camera pose estimation relies on hardware equipment in the prior art is solved, efficient and accurate pose estimation is achieved, and cost and deployment difficulty is reduced.
Patent Information
- Application Number
- CN202210798980.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-08
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-07-08
AI Technical Summary
In the field of autonomous driving, the prior art relies on hardware equipment such as depth sensors when estimating camera positions in the field of autonomous driving, resulting in high requirements for hardware equipment, high cost and difficult deployment.
A training method for pose estimation network model is proposed. By using predicted depth information and supervised depth information to train the network model, supervised information is simplified, training efficiency is improved, and no need to rely on other hardware equipment other than the camera.
It achieves efficient and accurate camera position estimation, reduces dependence on hardware devices, saves costs and reduces deployment difficulty.
Smart Images

Figure CN115147683B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a training method for a pose estimation network model, a pose estimation method, a device, an electronic device, and a computer-readable storage medium. Background Art
[0002] In the field of autonomous driving, the pose estimation of a camera plays an important role. For example, in the three-dimensional reconstruction technology required for autonomous driving, when performing three-dimensional estimation, the higher the accuracy of the camera pose, the higher the accuracy of the reconstruction. Another example is that when a vehicle with autonomous driving function is driving, since the camera will have large fluctuations due to bumps during the vehicle's movement, it is easy to cause deviations in the pose estimation of the camera relative to the ground, which will have a greater impact on the accuracy of the reconstruction estimation. In view of this, the prior art usually uses hardware devices such as depth sensors to estimate the camera pose, which has high requirements for hardware devices. Summary of the Invention
[0003] To solve the above technical problems, the present disclosure is proposed. Embodiments of the present disclosure provide a training method for a pose prediction network model, a pose prediction method for a camera, a device, an electronic device, and a computer-readable storage medium.
[0004] According to one aspect of the present disclosure, there is provided a training method for a pose estimation network model, including:
[0005] Performing pose prediction on a sample image by using the pose estimation network model to obtain a predicted pose angle of the camera; wherein, the sample image is collected at a preset position of the mobile device by the camera;
[0006] Determining predicted depth information corresponding to the sample image based on the predicted pose angle;
[0007] Determining a model loss based on the supervised depth information corresponding to the sample image and the predicted depth information;
[0008] Training the pose estimation network model based on the model loss.
[0009] According to another aspect of the present disclosure, there is provided a pose estimation method for a camera, including:
[0010] Obtaining a target image, wherein the target image is collected at a preset position of the mobile device by the camera;
[0011] Processing the target image by using the pose estimation network model to obtain at least one pose angle of the camera relative to the mobile device; wherein, the pose angle includes a pitch angle and / or a rotation angle.
[0012] According to another aspect of the present disclosure, there is provided a training device for a pose estimation network model, including:
[0013] A pose estimation module, configured to use the pose estimation network model to perform pose prediction on a sample image to obtain a predicted pose angle of the camera; wherein, the sample image is collected by the camera at a preset position of the mobile device;
[0014] A depth prediction module, configured to determine predicted depth information corresponding to the sample image based on the predicted pose angle determined by the pose estimation module;
[0015] A loss determination module, configured to determine a model loss based on the supervised depth information corresponding to the sample image and the predicted depth information determined by the depth prediction module;
[0016] A model training module, configured to train the pose estimation network model based on the model loss determined by the loss determination module.
[0017] According to another aspect of the present disclosure, there is provided a pose estimation device for a camera, including:
[0018] An image acquisition module, configured to acquire a target image, wherein the target image is collected by the camera at a preset position of the mobile device;
[0019] A pose angle estimation module, configured to use the pose estimation network model to process the target image obtained by the image acquisition module to obtain at least one pose angle of the camera relative to the mobile device.
[0020] According to another aspect of the present disclosure, there is provided a computer-readable storage medium storing a computer program for executing the method according to any one of the above embodiments.
[0021] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0022] A processor;
[0023] A memory for storing executable instructions of the processor;
[0024] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method according to any one of the above embodiments.
[0025] Based on the training method, pose estimation method, device, electronic device, and computer-readable storage medium of a pose estimation network model provided in the above embodiments of the present disclosure, the pose estimation network model is trained using predicted depth information and supervised depth information. Compared with using pose information as the supervised information, the supervised information in this embodiment is simplified, the training efficiency of the network model is improved, and moreover, the trained pose estimation network model does not need to rely on other hardware devices except the camera, achieving the technical effects of cost savings and deployment difficulty reduction.
[0026] The technical solutions of the present disclosure will be further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] By describing the embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure, and do not constitute a limitation to the present disclosure. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0028] Figure 1 It is a schematic flowchart of a method for estimating the pose of a camera provided by an exemplary embodiment of the present disclosure.
[0029] Figure 2 It is a schematic flowchart of a method for training a pose estimation network model provided by an exemplary embodiment of the present disclosure.
[0030] Figure 3 It is the present disclosure Figure 2 A schematic flowchart of step 202 in the embodiment shown.
[0031] Figure 4 It is the present disclosure Figure 3 A schematic flowchart of step 2021 in the embodiment shown.
[0032] Figure 5 It is the present disclosure Figure 2 A schematic flowchart of step 203 in the embodiment shown.
[0033] Figure 6 It is a schematic structural diagram of a training device for a pose estimation network model provided by an exemplary embodiment of the present disclosure.
[0034] Figure 7 It is a schematic structural diagram of a training device for a pose estimation network model provided by another exemplary embodiment of the present disclosure.
[0035] Figure 8 It is a schematic structural diagram of a device for estimating the pose of a camera provided by an exemplary embodiment of the present disclosure.
[0036] Figure 9 It is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. Detailed implementation manners
[0037] Next, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the exemplary embodiments described herein.
[0038] It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present disclosure.
[0039] Those skilled in the art can understand that terms such as "first", "second", etc. in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them.
[0040] It should also be understood that in the embodiments of the present disclosure, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.
[0041] It should also be understood that for any component, data or structure mentioned in the embodiments of the present disclosure, unless otherwise clearly defined or given a contrary indication in the context, it can generally be understood as one or more.
[0042] In addition, the term "and / or" in the present disclosure is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the associated objects before and after.
[0043] It should also be understood that the present disclosure emphasizes the differences between various embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be described one by one.
[0044] At the same time, it should be understood that for the convenience of description, the dimensions of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0045] The following description of at least one exemplary embodiment is actually only illustrative and in no way restricts the present disclosure and its application or use.
[0046] Well-known technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be considered as part of the specification.
[0047] It should be noted that like reference numerals and letters refer to like items in the following figures, and thus, once an item is defined in one figure, further discussion thereof is not required in subsequent figures.
[0048] Embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.
[0049] Electronic devices such as terminal devices, computer systems, servers, etc. can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.
[0050] Summary of the Application
[0051] In the process of implementing the present disclosure, the inventors found that in the prior art, the estimation of the camera pose relies on hardware devices such as depth sensors, and this technical solution has at least the following problems: it has high requirements for hardware devices, and the depth sensor needs to have high precision.
[0052] Exemplary Applications
[0053] Figure 1 It is a flowchart of a method for estimating the pose of a camera provided by an exemplary embodiment of the present disclosure. As Figure 1 shown, the method provided in this embodiment can be applied to electronic devices such as in-vehicle terminals, clients, etc., and includes:
[0054] Step 101, obtain a target image.
[0055] Among them, the target image is acquired by the camera at a preset position of the mobile device.
[0056] Optionally, the mobile device can be a device such as a vehicle. At this time, the pose estimation method provided in this embodiment can be applied in the field of autonomous driving to provide camera pose estimation information in 3D reconstruction for autonomous driving; the camera can include one or more cameras. When multiple cameras are included, the positions of each camera on the mobile device are usually different. For example, it can be a front-view camera set in front of the vehicle, or it can be a surround-view camera set around the vehicle; the target image corresponding to each camera is obtained by image acquisition at the preset position corresponding to each camera among the multiple cameras.
[0057] Step 102: Process the target image by using the pose estimation network model to obtain at least one pose angle of the camera relative to the ground.
[0058] Among them, the pose angle includes a pitch angle and / or a rotation angle.
[0059] Optionally, the pose estimation network model can be trained based on the training method provided in any one of the following embodiments, or the pose estimation network model can be trained based on any applicable training method in the related art, so that the pose estimation network model realizes that the input is the target image and the output is at least one pose angle of the camera corresponding to the target image relative to the ground.
[0060] In this embodiment, the pose estimation network model is used to directly output the pitch angle and / or rotation angle of the camera relative to the ground with the target image as the input. Since the calculation process is simple and the estimation is realized only based on a single-frame image, time-consuming steps such as feature point extraction are omitted, making the pose estimation faster. And the pose angle at the current moment is directly predicted based on a single-frame image, which overcomes the problem of error accumulation when predicting based on multiple frames of images; moreover, this embodiment is only based on the camera and does not rely on other hardware devices such as depth sensors and inertial sensors, which overcomes the problem of dependence on hardware devices in the prior art for estimating the camera pose, saves costs, and reduces the deployment difficulty.
[0061] Exemplary Methods
[0062] Figure 2 It is a schematic flowchart of a training method for a pose estimation network model provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to an electronic device, such as Figure 2 shown, and includes the following steps:
[0063] Step 201: Perform pose prediction on the sample image by using the pose estimation network model to obtain the predicted pose angle of the camera.
[0064] Among them, the sample image is collected by the camera at a preset position of the mobile device.
[0065] The pose estimation network model provided in this embodiment is used to determine, based on an image, the pose angle corresponding to the preset position of the camera on the mobile device when the image is collected, so as to obtain a predicted pose angle. When the number of sample images is multiple, the pose estimation network model is used for prediction respectively to obtain the predicted pose angle corresponding to the preset position where each sample image is located. Among them, the preset positions corresponding to each sample image may be the same or different. For example, multiple sample images are collected by cameras at different preset positions on a vehicle (an example of a mobile device), or multiple sample images are collected by a camera at the same preset position on a vehicle at multiple moments (the pose angle may change at different moments), or multiple sample images are respectively collected by cameras at preset positions on different vehicles; optionally, the predicted pose angle may include a pitch angle and / or a rotation angle; the network structure of the pose estimation network model in this embodiment may be the same as or similar to the pose estimation network model in the related art, and the specific network structure of the pose estimation network model is not limited in this embodiment.
[0066] Step 202: Determine the predicted depth information corresponding to the sample image based on the predicted pose angle.
[0067] In one embodiment, after obtaining the predicted pose angle, the predicted pose angle can be converted into a rotation matrix. Combining the predicted position corresponding to the camera when collecting the sample image, the height information of the camera relative to the ground can be known. Based on this height information and the above rotation matrix, the predicted depth information corresponding to the sample image can be obtained. The predicted depth information corresponding to the sample image is the predicted depth information of the ground corresponding to each pixel in the sample image relative to the camera.
[0068] Step 203: Determine the model loss based on the supervised depth information and the predicted depth information corresponding to the sample image.
[0069] Optionally, the supervised depth information of the sample image can be determined through a preset, or determined by processing the depth map corresponding to the sample image (obtained by a depth sensor) through a neural network. For example, the pixel points representing the ground in the depth map are determined through a semantic segmentation network (the sample image is input into the semantic segmentation network to determine the pixel points representing the ground in the sample image, and in combination with the correspondence between the depth map and the sample image, the pixel points representing the ground are determined from the depth map), and the supervised depth information is determined based on the depth information corresponding to the pixel points representing the ground; the model loss is determined through the difference between the supervised depth information and the predicted depth information, where specifically, any method for calculating the network model loss in the prior art can be used to determine the model loss.
[0070] Step 204: Train the pose estimation network model based on the model loss.
[0071] Optionally, based on the loss training method of any existing neural network model, the pose estimation network model can be trained using the model loss. For example, the gradient descent method can be used. When the preset conditions are met, the trained pose estimation network model is obtained. Optionally, the preset conditions may include, but are not limited to: the model loss is less than a preset value, the number of training times reaches a preset number, the difference between the model losses determined continuously twice is less than a preset difference, etc.
[0072] Based on the training method of a pose estimation network model provided in the above embodiments of the present disclosure, using the predicted depth information and the supervised depth information to train the pose estimation network model, compared with using the pose information as the supervised information, the supervised information in this embodiment is simplified, the training efficiency of the network model is improved, and moreover, the trained pose estimation network model does not need to rely on other hardware devices except the camera, achieving the technical effects of cost savings and reduced deployment difficulty.
[0073] As Figure 3 shown, based on the above Figure 2 shown embodiments, step 202 may include the following steps:
[0074] Step 2021: Determine the projection matrix between the observation coordinate system corresponding to the preset position and the pixel coordinate system corresponding to the sample image based on the predicted pose angle.
[0075] The projection matrix in this embodiment is used to project the observation coordinate system onto the pixel coordinate system. The projection matrix can be determined based on the rotation matrix determined by the predicted pose angle, the set position coordinates corresponding to the camera, the internal parameter matrix of the camera, and the direction rotation matrix representing the conversion from the pixel coordinate system to the observation coordinate system.
[0076] In addition, the xy plane of the above ground is parallel to the observation coordinate system. The observation coordinate system includes an x-axis, a y-axis, and a z-axis. The center line of the mobile device is used as the x-axis, and the moving direction of the mobile device is used as the positive direction of the x-axis; the y-axis is perpendicular to the x-axis and the positive direction is to the left. The y-axis and the x-axis are in the same horizontal plane parallel to the ground; the z-axis is perpendicular to the xy plane and the positive direction is upward. In some alternative examples, when the mobile device is a vehicle, the observation coordinate system is the vehicle coordinate system. The vehicle coordinate system uses the center line of the vehicle as the x-axis, and the driving direction of the vehicle is the positive direction of the x-axis; the y-axis is perpendicular to the x-axis and the positive direction is to the left. The y-axis and the x-axis are in the same horizontal plane and parallel to the ground; the z-axis is perpendicular to the xy plane and the positive direction is upward.
[0077] Step 2022: Determine the predicted depth information corresponding to the ground pixel points in the sample image based on the projection matrix and the two-dimensional coordinate information of the ground pixel points in the sample image.
[0078] Optionally, the pixel points representing the ground in the sample image can be identified through a semantic segmentation network model or other neural network models and determined as ground pixel points. The two-dimensional coordinate information of these ground pixel points in the image coordinate system can be obtained from the sample image. Based on the projection matrix, the two-dimensional coordinate information of the ground pixel points is projected into the observation coordinate system (three-dimensional coordinate system), and thus the predicted depth information corresponding to each ground pixel point can be determined.
[0079] In this embodiment, since the supervised depth information is used as the supervised information for training the pose estimation network model, the predicted depth information corresponding to this supervised depth information is required to calculate the network loss. The result predicted by the pose estimation network model is the predicted pose angle. Therefore, through the method provided in this embodiment, the predicted depth information can be determined based on the predicted pose angle and the conversion between coordinate systems, achieving the information correspondence in the training process of the pose estimation network model, simplifying the supervised information, and improving the training efficiency of the network model.
[0080] As Figure 4 shown, based on the above Figure 3 shown embodiment, step 2021 may include the following steps:
[0081] Step 401: Determine the rotation matrix corresponding to the camera based on the predicted pose angle.
[0082] Optionally, based on the predicted pose angle, the Euler angles are calculated and converted to obtain the rotation matrix. For example, the predicted pose angle includes the predicted pitch angle and the predicted roll angle. At this time, since the yaw angle has less application range and basically does not change in the vehicle-mounted situation, it can be set to 0. The rotation matrix is a matrix used to describe the pose (excluding position) relationship between two different Cartesian coordinate systems in three-dimensional space. In this embodiment, the pose change between the camera coordinate system corresponding to the camera and the world coordinate system is represented by the rotation matrix. For example, the rotation matrix can be expressed as
[0083] where each element in the rotation matrix R is respectively:
[0084] r 11 : cos(yaw)cos(pitch);
[0085] r 12 : cos(yaw)sin(pitch)sin(roll) - cos(roll)sin(yaw);
[0086] r13 : sin(yaw)sin(roll) + cos(yaw)cos(roll)sin(pitch);
[0087] r 21 : cos(pitch)sin(yaw);
[0088] r 22 : cos(yaw)cos(roll) + sin(yaw)sin(pitch)sin(roll);
[0089] r 23 : cos(roll)sin(yaw)sin(pitch) - cos(yaw)sin(roll);
[0090] r 31 : -sin(pitch);
[0091] r 32 : cos(pitch)sin(roll);
[0092] r 33 : cos(pitch)cos(roll).
[0093] Step 402: Determine the coordinate matrix of the camera in the observation coordinate system based on the set position.
[0094] Optionally, since the set position is a known position when the camera captures the sample image, this coordinate matrix is known. For example, it can be expressed as where t x , t y , t z respectively represent the coordinates of the set position on the x-axis, y-axis, and z-axis in the observation coordinate system.
[0095] Step 403: Determine the direction rotation matrix between the observation coordinate system and the pixel coordinate system corresponding to the sample image.
[0096] In this embodiment, the direction rotation matrix represents the attitude change between the observation coordinate system and the pixel coordinate system. Therefore, without considering the position of the origin, only the directions of the x-axis, y-axis, and z-axis are transformed. The x-axis in the observation coordinate system corresponds to the direction in which the camera faces directly forward, the y-axis corresponds to the direction in which the camera faces left, and the z-axis corresponds to the direction in which the camera faces upward; while in the pixel coordinate system, the x-axis corresponds to the direction in which the camera faces right, the y-axis corresponds to the direction in which the camera faces downward, and the z-axis corresponds to the direction in which the camera faces directly forward. Therefore, in an optional example, the direction rotation matrix can be expressed as: The rotation matrix in this direction can be used to transform the axis directions corresponding to the viewing coordinate system to the axis directions corresponding to the pixel coordinate system.
[0097] Step 404: Based on the rotation matrix, the coordinate matrix, the direction rotation matrix, and the preset camera internal parameters, determine the projection matrix between the viewing coordinate system and the pixel coordinate system.
[0098] In this embodiment, when the preset camera internal parameters (known parameters of the camera) are known, the internal parameter matrix corresponding to the camera can be determined. For example, the camera internal parameter matrix is expressed as where f u represents the length of the focal length in the x-axis direction described by pixels, and f v represents the length of the focal length in the y-axis direction described by pixels. c u and c v represent the actual positions of the principal points, with the unit of pixels.
[0099] Given the above rotation matrix, coordinate matrix, direction rotation matrix, and preset camera internal parameters, for a point (x, y) in the image coordinate system and its corresponding three-dimensional coordinates (X 1 , Y 1 , Z 1 ), the following formula (1) is satisfied:
[0100]
[0101] where Z 2 is a proportionality coefficient (which can be preset); KAR represents the successive matrix multiplication of the internal parameter matrix K, the direction rotation matrix A, and the rotation matrix R corresponding to the camera; KART represents the successive matrix multiplication of the internal parameter matrix K, the direction rotation matrix A, the rotation matrix R corresponding to the camera, and the coordinate matrix T; P represents the projection matrix; the projection matrix can be determined through the above formula (1). Combined with this projection matrix, the depth value corresponding to the image coordinate point in the image (such as the ground coordinate point, etc.) can be determined. Optionally, in this embodiment, X 1 represents the predicted depth value corresponding to the image coordinate point with image coordinates x and y. Considering that the height of the ground coordinate point is 0, so Z 1 = 0, which is equivalent to removing the third column of the projection matrix P. At the same time, considering the fourth row in the projection matrix P, etc., then the third column and the fourth row of the projection matrix P can be removed to obtain a 3x3 matrix P 0 . By taking the inverse of P 0 , the following formula (2) can be obtained:
[0102]
[0103] That is, the depth X corresponding to the point (x, y) in the image coordinate system can be obtained.1 。
[0104] Optionally, based on the above embodiments, step 401 may further include:
[0105] Determine the rotation matrix corresponding to the camera based on the predicted pitch angle and the predicted rotation angle.
[0106] Among them, the predicted attitude angle includes the predicted pitch angle and the predicted rotation angle.
[0107] In this embodiment, during the movement of the mobile device, for example, during the driving of a vehicle, the yaw angle changes little in the on-vehicle situation due to its small application range, so it can be set to 0; therefore, the yaw angle in the rotation matrix provided in the above embodiment can be set to 0 to obtain the rotation matrix The projection matrix between the observation coordinate system and the pixel coordinate system can be determined more quickly through the simplified rotation matrix, improving the calculation efficiency of determining the projection matrix.
[0108] Optionally, based on the above Figure 3 shown embodiments, before performing step 2022, it may further include:
[0109] Perform semantic segmentation on the sample image to determine the ground pixel points included in the sample image.
[0110] Optionally, the sample image can be semantically segmented based on a semantic segmentation network model, and the pixel points corresponding to different semantic categories in the sample image (specific semantic categories can be determined according to a preset. After determining multiple semantic categories according to the preset, the semantic segmentation network model is trained based on these categories so that the trained semantic segmentation network model can well segment these semantic categories in the image) are segmented into different parts. For example, all pixel points with the semantic meaning of "ground" are aggregated into a set, and the pixel points in this set are used as the ground pixel points; in this embodiment, the pixel points corresponding to the ground semantics in the sample image are determined through segmentation. Since the height of the determined ground pixel points is 0, the predicted depth information of the ground pixel points in the observation coordinate system can be determined in step 2022 in combination with the projection matrix, so as to determine the predicted depth information corresponding to the supervised depth information based on the predicted attitude angle, and further determine the model loss.
[0111] As Figure 5 shown, based on the above Figure 2 shown embodiments, step 203 may include the following steps:
[0112] Step 2031, determine the reference ground pixel points in the three-dimensional image based on the semantic segmentation result of the sample image and the depth map corresponding to the sample image.
[0113] Optionally, a depth map corresponding to the sample image can be obtained through a depth sensor (e.g., a depth camera, etc.). The process of obtaining the depth map may include: at a preset position corresponding to the camera that obtains the sample image through the depth sensor, while collecting the sample image, collecting a depth map with depth information that is exactly the same as the image content in the sample image, that is, each pixel point in the depth map corresponds one-to-one with each pixel point in the sample image. Optionally, the ground pixel points in the sample image can be directly determined by using the semantic segmentation result obtained in step 202 of the above embodiment. Based on the determined ground pixel points and the pixel correspondence between the sample image and the depth map, the reference ground pixel points with the semantic of "ground" in the three-dimensional image can be determined. Or, the sample image can be semantically segmented based on a semantic segmentation network model, and the pixel points corresponding to different semantic categories in the sample image (specific semantic categories can be determined according to the preset. After determining multiple semantic categories according to the preset, the semantic segmentation network model is trained based on these categories so that the trained semantic segmentation network model can well segment these semantic categories in the sample image) can be segmented into different parts. For example, all pixel points with the semantic of "ground" are aggregated into a set. Based on this set and the pixel correspondence between the sample image and the depth map, the reference ground pixel points with the semantic of "ground" in the three-dimensional image can be determined.
[0114] Step 2032: Determine the supervised depth information based on the depth information corresponding to the reference ground pixel points.
[0115] In this embodiment, since each pixel point in the depth map has depth information, after determining the reference ground pixel points in the depth map based on the semantic segmentation result of the sample image, the supervised depth information corresponding to the reference ground pixel points can be determined based on the depth information of the depth map.
[0116] Step 2033: Determine the model loss based on the supervised depth information and the predicted depth information.
[0117] Since the pixel points in the depth map correspond one-to-one with the pixel points in the sample image, it can be known that the reference ground pixel points also correspond one-to-one with the ground pixel points in the sample image. Therefore, the supervised depth information and the predicted depth information correspond one-to-one. At this time, the model loss can be determined through the difference between the predicted depth information corresponding to the sample image and the supervised depth information. Based on this model loss, the pose estimation network model is trained, realizing the network training of the pose estimation network model supervised by depth information. Since the model loss is determined based on depth information, it overcomes the problem that when directly using the pose angle as the supervised information, it is difficult to obtain the pose angle, and during the movement of the mobile device, the pose angle will change with the movement, resulting in the inability to accurately obtain the supervised information. In this embodiment, the predicted pose angle predicted by the pose estimation network model is processed to obtain the predicted depth information, and the model loss determined by combining the supervised depth information depending only on the depth sensor is independent of other hardware devices such as depth sensors and inertial sensors, which can save costs and reduce the deployment difficulty.
[0118] Any training method of the pose estimation network model and pose estimation method provided by the embodiments of the present disclosure can be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices, servers, etc. Alternatively, any training method of the pose estimation network model and pose estimation method provided by the embodiments of the present disclosure can be executed by a processor. For example, the processor executes any training method of the pose estimation network model and pose estimation method mentioned in the embodiments of the present disclosure by calling the corresponding instructions stored in the memory. This will not be elaborated below.
[0119] Exemplary Devices
[0120] Figure 6 It is a schematic structural diagram of a training device for a pose estimation network model provided by an exemplary embodiment of the present disclosure. As Figure 6 shown, the device provided in this embodiment includes:
[0121] A pose estimation module 61, configured to perform pose prediction on a sample image by using a pose estimation network model to obtain a predicted pose angle of a camera; wherein, the sample image is collected at a preset position of a mobile device by the camera;
[0122] A depth prediction module 62, configured to determine predicted depth information corresponding to the sample image based on the predicted pose angle determined by the pose estimation module 61;
[0123] A loss determination module 63, configured to determine a model loss based on the supervised depth information corresponding to the sample image and the predicted depth information determined by the depth prediction module 62;
[0124] The model training module 64 trains the pose estimation network model based on the model loss determined by the loss determination module 63.
[0125] Based on the training device of a pose estimation network model provided in the above embodiments of the present disclosure, the pose estimation network model is trained using predicted depth information and supervised depth information, which simplifies the supervision information, improves the training efficiency of the network model, and moreover, the trained pose estimation network model does not need to rely on other hardware devices except the camera, achieving the technical effects of cost savings and reduced deployment difficulty.
[0126] Figure 7 It is a schematic structural diagram of a training device of a pose estimation network model provided in another exemplary embodiment of the present disclosure. As Figure 7 shown, the device provided in this embodiment includes:
[0127] The depth prediction module 62 includes:
[0128] The projection matrix determination unit 621 is configured to determine the projection matrix between the observation coordinate system corresponding to the preset position and the pixel coordinate system corresponding to the sample image based on the predicted pose angle;
[0129] The ground depth prediction unit 622 is configured to determine the predicted depth information corresponding to the ground pixel points in the sample image based on the projection matrix and the two-dimensional coordinate information of the ground pixel points in the sample image.
[0130] Optionally, the projection matrix determination unit 621 is specifically configured to determine the rotation matrix corresponding to the camera based on the predicted pose angle; determine the coordinate matrix of the camera in the observation coordinate system based on the set position; determine the direction rotation matrix between the observation coordinate system and the pixel coordinate system corresponding to the sample image; and determine the projection matrix between the observation coordinate system and the pixel coordinate system based on the rotation matrix, the coordinate matrix, the direction rotation matrix, and the preset camera internal parameters.
[0131] When determining the rotation matrix corresponding to the camera based on the predicted pose angle, the projection matrix determination unit 621 is configured to determine the rotation matrix corresponding to the camera based on the predicted pitch angle and the predicted rotation angle; wherein, the predicted pose angle includes the predicted pitch angle and the predicted rotation angle.
[0132] The depth prediction module 62 further includes:
[0133] The semantic segmentation unit 623 is configured to perform semantic segmentation on the sample image to determine the ground pixel points included in the sample image.
[0134] In some alternative embodiments, the loss determination module 63 includes:
[0135] A depth map segmentation unit 631, configured to determine reference ground pixel points in the three-dimensional image based on a semantic segmentation result of the sample image and a depth map corresponding to the sample image;
[0136] A supervision information determination unit 632, configured to determine the supervision depth information based on depth information corresponding to the reference ground pixel points;
[0137] A model loss unit 633, configured to determine the model loss based on the supervision depth information and the predicted depth information.
[0138] Figure 8 It is a schematic structural diagram of a pose estimation device of a camera provided by an exemplary embodiment of the present disclosure. As Figure 8 shown, the device provided in this embodiment includes:
[0139] An image acquisition module 81, configured to acquire a target image.
[0140] Wherein, the target image is acquired by the camera at a preset position of the mobile device.
[0141] An attitude angle estimation module 82, configured to process the target image obtained by the image acquisition module 81 by using a pose estimation network model to obtain at least one attitude angle of the camera relative to the ground.
[0142] In this embodiment, by using the pose estimation network model, the target image is used as an input, and the pitch angle and / or rotation angle of the camera relative to the ground are directly output. Since the calculation process is simple and the estimation is only based on a single-frame image, time-consuming steps such as feature point extraction are omitted, making the pose estimation faster. Moreover, the attitude angle at the current moment is directly predicted from a single-frame image, overcoming the problem of error accumulation when predicting based on multiple-frame images. And, this embodiment is only based on the camera and does not rely on other hardware devices such as depth sensors and inertial sensors, overcoming the dependence on hardware devices when estimating the camera pose in the prior art, saving costs and reducing the deployment difficulty.
[0143] Exemplary Electronic Devices
[0144] Next, reference Figure 9 is made to describe an electronic device according to an embodiment of the present disclosure. The electronic device may be any one or both of the first device 100 and the second device 200, or a stand-alone device independent of them, and the stand-alone device can communicate with the first device and the second device to receive the input signals collected from them.
[0145] Figure 9 The block diagram of an electronic device according to an embodiment of the present disclosure is illustrated.
[0146] As shown Figure 9 in FIG. 1, the electronic device 90 includes one or more processors 91 and a memory 92.
[0147] The processor 91 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 90 to perform desired functions.
[0148] The memory 92 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 91 may run the program instructions to implement the training method and pose estimation method of the pose estimation network model of the various embodiments of the present disclosure described above and / or other desired functions. Various contents such as input signals, signal components, noise components, etc. may also be stored in the computer-readable storage media.
[0149] In one example, the electronic device 90 may further include: an input device 93 and an output device 94, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0150] For example, when the electronic device is the first device 100 or the second device 200, the input device 93 may be the above-mentioned microphone or microphone array for capturing the input signal of the sound source. When the electronic device is a stand-alone device, the input device 93 may be a communication network connector for receiving the collected input signals from the first device 100 and the second device 200.
[0151] In addition, the input device 93 may further include, for example, a keyboard, a mouse, and so on.
[0152] The output device 94 may output various information to the outside, including the determined distance information, direction information, etc. The output device 94 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0153] Of course, for simplicity, Figure 9Only some of the components related to the present disclosure in the electronic device 90 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 90 may further include any other appropriate components.
[0154] Exemplary Computer Program Products and Computer Readable Storage Media
[0155] In addition to the above methods and devices, embodiments of the present disclosure may also be computer program products, which include computer program instructions that, when run on a processor, cause the processor to execute the steps in the training method and pose estimation method of the pose estimation network model according to various embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0156] The computer program product may be written in any combination of one or more programming languages for executing the program code of the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0157] Furthermore, embodiments of the present disclosure may also be computer-readable storage media, on which computer program instructions are stored, and the computer program instructions, when run on a processor, cause the processor to execute the steps in the training method and pose estimation method of the pose estimation network model according to various embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0158] The computer-readable storage media may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0159] The basic principles of the present disclosure have been described in connection with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are merely examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. Additionally, the specific details disclosed above are for illustrative and facilitating understanding purposes only, and not limitations. These details do not limit the present disclosure to necessarily adopting the above specific details for implementation.
[0160] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the phrase "and / or", and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with it.
[0161] It should also be noted that in the devices, equipment, and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.
[0162] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects are very obvious to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
[0163] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.
Claims
1. A training method for a pose estimation network model, comprising: Performing pose prediction on a sample image using the pose estimation network model to obtain the predicted pose angles of the camera; wherein, the sample image is collected by the camera at a preset position of the mobile device; Based on the predicted pose angles, determining the predicted depth information corresponding to the sample image; Based on the supervised depth information corresponding to the sample image and the predicted depth information, determining the model loss; Based on the model loss, training the pose estimation network model; The determining the predicted depth information corresponding to the sample image based on the predicted pose angles includes: Based on the predicted pose angles, determining the projection matrix between the observation coordinate system corresponding to the preset position and the pixel coordinate system corresponding to the sample image; the projection matrix is determined based on the rotation matrix determined by the predicted pose angles, the set position coordinates corresponding to the camera, the internal parameter matrix of the camera, and the direction rotation matrix representing the conversion of the pixel coordinate system to the observation coordinate system; Identifying the pixel points representing the ground in the sample image through a semantic segmentation network model and determining them as ground pixel points; Based on the projection matrix and the two-dimensional coordinate information of the ground pixel points in the sample image, determining the predicted depth information corresponding to the ground pixel points in the sample image.
2. The method according to claim 1, wherein, The determining the projection matrix between the observation coordinate system corresponding to the preset position and the pixel coordinate system corresponding to the sample image based on the predicted pose angles includes: Based on the predicted pose angles, determining the rotation matrix corresponding to the camera; Based on the set position, determining the coordinate matrix of the camera in the observation coordinate system; Determining the direction rotation matrix between the observation coordinate system and the pixel coordinate system corresponding to the sample image; Based on the rotation matrix, the coordinate matrix, the direction rotation matrix, and the preset camera internal parameters, determining the projection matrix between the observation coordinate system and the pixel coordinate system.
3. The method according to claim 2, wherein, The predicted pose angles include a predicted pitch angle and a predicted rotation angle; The determining the rotation matrix corresponding to the camera based on the predicted pose angles includes: Based on the predicted pitch angle and the predicted rotation angle, determining the rotation matrix corresponding to the camera.
4. For the method according to any one of claims 1-3, before determining the predicted depth information corresponding to the ground pixel points in the sample image based on the projection matrix and the two-dimensional coordinate information of the ground pixel points in the sample image, it also comprises: Performing semantic segmentation on the sample image to determine the ground pixel points included in the sample image.
5. The method according to any one of claims 1-3, wherein, The determining the model loss based on the supervised depth information corresponding to the sample image and the predicted depth information includes: Based on the semantic segmentation result of the sample image and the depth map corresponding to the sample image, determining the reference ground pixel points in the three-dimensional image; Based on the depth information corresponding to the reference ground pixel points, determining the supervised depth information; Determine the model loss based on the supervised depth information and the predicted depth information.
6. A method for pose estimation of a camera, comprising: obtaining a target image, where the target image is acquired by the camera at a preset position of a mobile device; processing the target image by using a pose estimation network model to obtain at least one pose angle of the camera relative to the mobile device; wherein the pose angle includes a pitch angle and / or a rotation angle, and the pose estimation network model is trained by using the training method of the pose estimation network model according to any one of the above claims 1-5.
7. A training device for a pose estimation network model, comprising: a pose estimation module, configured to perform pose prediction on a sample image by using a pose estimation network model to obtain a predicted pose angle of the camera; wherein the sample image is acquired by the camera at a preset position of a mobile device; a depth prediction module, configured to determine predicted depth information corresponding to the sample image based on the predicted pose angle determined by the pose estimation module; a loss determination module, configured to determine a model loss based on the supervised depth information corresponding to the sample image and the predicted depth information determined by the depth prediction module; a model training module, configured to train the pose estimation network model based on the model loss determined by the loss determination module; the depth prediction module includes: a projection matrix determination unit, configured to determine a projection matrix between an observation coordinate system corresponding to the preset position and a pixel coordinate system corresponding to the sample image based on the predicted pose angle; the projection matrix is determined based on a rotation matrix determined by the predicted pose angle, a set position coordinate corresponding to the camera, an internal parameter matrix of the camera, and a direction rotation matrix representing the conversion of the pixel coordinate system to the observation coordinate system; a ground depth prediction unit, configured to identify pixel points representing the ground in the sample image through a semantic segmentation network model and determine them as ground pixel points; and determine predicted depth information corresponding to the ground pixel points in the sample image based on the projection matrix and two-dimensional coordinate information of the ground pixel points in the sample image.
8. A pose estimation device for a camera, comprising: an image acquisition module, configured to acquire a target image, where the target image is acquired by the camera at a preset position of a mobile device; a pose angle estimation module, configured to process the target image obtained by the image acquisition module by using a pose estimation network model to obtain at least one pose angle of the camera relative to the mobile device, and the pose estimation network model is trained by using the training method of the pose estimation network model according to any one of the above claims 1-5.
9. A computer-readable storage medium storing a computer program for executing the method according to any one of the above claims 1-6.
10. An electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the method according to any one of the above claims 1-6.
Citation Information
Patent Citations
Pose estimation method and device based on deep neural network
CN110473254A
Model training method and device
CN112241976A