Monocular camera-based dense point cloud generation method and apparatus, and electronic device
The method of generating dense point clouds through a monocular camera solves the problem of insufficient perception capability of sparse point clouds, achieves efficient three-dimensional environment perception and obstacle detection, and reduces costs.
Patent Information
- Application Number
- CN202510502677.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-09-12
AI Technical Summary
In the existing technology, the visual inertial odometry system in the robotics field can only generate sparse point clouds, which is difficult to provide sufficient three-dimensional spatial information, limiting the robot's ability to perceive the surrounding environment. In addition, the cost of generating dense point clouds through multi-sensor fusion is high.
By acquiring target images and image sequences based on a monocular camera, a sparse point cloud is generated, obstacles are identified and projected, and a new point cloud is generated by combining the target point cloud. Finally, a dense point cloud is obtained, and the point cloud is combined using the internal parameters of the monocular camera and the pixel coordinate system.
It provides richer three-dimensional spatial information, improves the robot's perception of the surrounding environment, and can accurately detect and avoid obstacles, while reducing costs and not relying on other sensors such as lidar.
Smart Images

Figure CN120635295A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, device and electronic device for generating dense point clouds based on a monocular camera. Background Art
[0002] Currently, in the robotics field (such as lawn mowers and robot vacuums), visual-inertial odometry (VIO) systems are commonly used to perceive the surrounding environment. However, VIO systems can only generate sparse point clouds, which cannot provide sufficient three-dimensional spatial information. This greatly limits the robot's ability to perceive the surrounding environment and cannot accurately detect and avoid obstacles. However, the cost of fusing multiple sensors (such as vision and lidar) to generate dense point clouds is high. Therefore, how to densify sparse point clouds to improve the robot's perception of the surrounding environment has become a pressing technical problem. Summary of the Invention
[0003] The present application provides a method, device and electronic device for generating dense point clouds based on a monocular camera to solve the technical problem of high cost in the prior art of generating dense point clouds by using multi-sensor fusion.
[0004] In a first aspect, an embodiment of the present application provides a method for generating a dense point cloud based on a monocular camera, the method comprising:
[0005] Acquire a target image captured by a monocular camera and an image sequence containing the target image, and generate a sparse point cloud based on the image sequence;
[0006] Identifying obstacles in the target image to obtain an obstacle area, wherein the target image is formed based on a pixel coordinate system;
[0007] Projecting the sparse point cloud to a pixel coordinate system to obtain a set of projection points corresponding to the sparse point cloud;
[0008] Comparing the projection point set and the obstacle area to obtain a target point cloud, wherein the target point cloud is a set of sparse points corresponding to the projection points falling into the obstacle area;
[0009] A new point cloud is generated based on the target point cloud, and the new point cloud is combined with the sparse point cloud to obtain a dense point cloud.
[0010] Optionally, generating a new point cloud based on the target point cloud includes:
[0011] Obtaining a target projection point corresponding to a target sparse point, wherein the target sparse point is any sparse point in the target point cloud;
[0012] Sampling is performed based on the target projection point to obtain a sampling point set;
[0013] Based on the sampling point set, a sub-point cloud is obtained, wherein the new point cloud includes a plurality of the sub-point clouds, and the sub-point clouds correspond one-to-one to the sparse points in the target point cloud.
[0014] Optionally, the performing sampling based on the target projection point to obtain a sampling point set includes:
[0015] Obtaining the coordinate position of the target projection point in the pixel coordinate system;
[0016] Determine a target sampling area based on the coordinate position, wherein the target sampling area is an area with the target projection point as the center and a preset distance as the radius, and the target sampling area is located in the obstacle area;
[0017] Sampling is performed within the sampling area to obtain a preset number of sampling points, thereby forming the sampling point set.
[0018] Optionally, obtaining a sub-point cloud based on the sampling point set includes:
[0019] Obtaining a depth value corresponding to the target projection point, and determining the depth value corresponding to the target projection point as the depth value of each sampling point in the sampling point set;
[0020] Obtaining an intrinsic parameter matrix corresponding to the monocular camera, wherein the intrinsic parameter matrix is used to characterize geometric characteristics inside the monocular camera;
[0021] Based on the depth value of each sampling point in the sampling point set and the intrinsic parameter matrix, each sampling point in the sampling point set is projected to a camera coordinate system to obtain the sub-point cloud.
[0022] Optionally, projecting the sparse point cloud to a pixel coordinate system to obtain a set of projection points corresponding to the sparse point cloud includes:
[0023] Normalizing the coordinates of each sparse point in the sparse point cloud in the camera coordinate system to obtain a normalized processing result corresponding to each sparse point;
[0024] Determining the projection position of each sparse point in the pixel coordinate system based on the normalization processing result corresponding to each sparse point and the intrinsic parameter matrix corresponding to the monocular camera;
[0025] Based on the projection position of each sparse point in the pixel coordinate system, each sparse point is projected to the pixel coordinate system to obtain the projection point corresponding to each sparse point, thereby obtaining a projection point set corresponding to the sparse point cloud.
[0026] Optionally, before projecting the sparse point cloud to a pixel coordinate system to obtain a set of projection points corresponding to the sparse point cloud, the method further includes:
[0027] Converting the sparse point cloud from a camera coordinate system to a vehicle body coordinate system;
[0028] filtering the sparse point cloud in the vehicle body coordinate system to obtain a filtered sparse point cloud;
[0029] Converting the filtered sparse point cloud from the vehicle body coordinate system back to the camera coordinate system;
[0030] The projecting the sparse point cloud to a pixel coordinate system to obtain a set of projection points corresponding to the sparse point cloud includes:
[0031] The filtered sparse point cloud in the camera coordinate system is projected to a pixel coordinate system to obtain a set of projection points corresponding to the filtered sparse point cloud.
[0032] Optionally, converting the sparse point cloud from a camera coordinate system to a vehicle coordinate system comprises:
[0033] Obtaining an extrinsic parameter matrix corresponding to the monocular camera, wherein the extrinsic parameter matrix is used to represent a rotation matrix and a translation vector from the camera coordinate system to the vehicle coordinate system;
[0034] Based on the extrinsic parameter matrix, the sparse point cloud is transformed from the camera coordinate system to the vehicle body coordinate system.
[0035] In a second aspect, an embodiment of the present application further provides a dense point cloud generation device based on a monocular camera, the device comprising:
[0036] An acquisition module is used to acquire a target image captured by a monocular camera and an image sequence containing the target image, and generate a sparse point cloud based on the image sequence;
[0037] an identification module, configured to identify obstacles in the target image and obtain an obstacle area, wherein the target image is formed based on a pixel coordinate system;
[0038] A projection module, configured to project the sparse point cloud into a pixel coordinate system to obtain a set of projection points corresponding to the sparse point cloud;
[0039] a comparison module, configured to compare the projection point set and the obstacle region to obtain a target point cloud, wherein the target point cloud is a set of sparse points corresponding to the projection points falling into the obstacle region;
[0040] A generation module is used to generate a new point cloud based on the target point cloud, and combine the new point cloud with the sparse point cloud to obtain a dense point cloud.
[0041] In a third aspect, an embodiment of the present application further provides an electronic device, comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;
[0042] Memory for storing computer programs;
[0043] The processor is configured to implement the method for generating a dense point cloud based on a monocular camera as described in any one of the embodiments of the first aspect when executing the program stored in the memory.
[0044] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method for generating a dense point cloud based on a monocular camera as described in any embodiment of the first aspect is implemented.
[0045] The above technical solution provided by the embodiment of the present application has the following advantages compared with the prior art:
[0046] The method provided by an embodiment of the present application obtains a target image captured by a monocular camera and an image sequence containing the target image, and generates a sparse point cloud based on the image sequence; identifies obstacles in the target image to obtain an obstacle area, wherein the target image is formed based on a pixel coordinate system; projects the sparse point cloud to a pixel coordinate system to obtain a set of projection points corresponding to the sparse point cloud; compares the set of projection points with the obstacle area to obtain a target point cloud, wherein the target point cloud is a set of sparse points corresponding to the projection points falling into the obstacle area; generates a new point cloud based on the target point cloud, and combines the new point cloud with the sparse point cloud to obtain a dense point cloud. Through the above method, the set of projection points corresponding to the sparse point cloud can be compared with the obstacle area in the target image, so as to determine the target point cloud where the projection point falls into the obstacle area, and then generate a new point cloud based on the target point cloud to obtain a dense point cloud. This can provide richer three-dimensional spatial information, greatly improve the robot's perception of the surrounding environment, and enable it to accurately detect and avoid obstacles; at the same time, since the present application mainly relies on a monocular camera to implement, there is no need to fuse information from other sensors such as lidar to generate a dense point cloud, so the cost is lower and it has stronger market competitiveness. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0049] Figure 1 A schematic diagram of a flow chart of a method for generating a dense point cloud based on a monocular camera provided in an embodiment of the present application;
[0050] Figure 2 A schematic diagram of a flow chart of another method for generating a dense point cloud based on a monocular camera provided in an embodiment of the present application;
[0051] Figure 3 A schematic diagram of the structure of a dense point cloud generation device based on a monocular camera provided in an embodiment of the present application;
[0052] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0054] See also Figure 1 , Figure 1 The flowchart of a method for generating dense point cloud based on a monocular camera is provided in an embodiment of the present application. Figure 1 As shown, the dense point cloud generation method based on a monocular camera may include the following steps:
[0055] Step S101: Acquire a target image captured by a monocular camera and an image sequence containing the target image, and generate a sparse point cloud based on the image sequence.
[0056] Specifically, the target image can be any frame image captured by a monocular camera. The image sequence refers to a sequence of multiple consecutive frames, and these multiple frames include the target image. In other words, the image sequence can be a sequence consisting of a target image and multiple frames before the target image, or a sequence consisting of a target image and multiple frames after the target image, or a sequence consisting of a target image and multiple frames before and after the target image, and the embodiments of the present application do not make specific limitations.
[0057] The above-mentioned sparse point cloud refers to a relatively sparsely distributed point cloud generated based on an image sequence. Specifically, the sparse point cloud can be generated by fusing inertial information, odometry information, and an image sequence, or by other means, which is not specifically limited in this application. When generating a sparse point cloud by fusing inertial information, odometry information, and an image sequence, the image sequence, inertial information, and odometry information can be first obtained. Then, a feature detection algorithm is used to extract key feature points from the image sequence, and the key feature points are tracked using the optical flow method. Next, the image sequence, inertial information, and odometry information are fused using an extended Kalman filter to obtain the attitude change information of the monocular camera. The inertial information includes information such as acceleration and angular velocity, which can be used to estimate the instantaneous motion and attitude change of the camera. The odometry information includes changes in relative position and attitude and is generally used to supplement the deficiencies in inertial information. The tracked key feature points are then triangulated based on the attitude change information of the monocular camera, and the coordinates of the key feature points in the camera coordinate system are calculated to generate a sparse point cloud. The purpose of triangulation here is to map the feature points in the two-dimensional image to three-dimensional space to generate a sparse point cloud. Step S102: Identify obstacles in the target image to obtain obstacle areas, wherein the target image is formed based on a pixel coordinate system.
[0058] Specifically, a pre-trained target classification model can be used to identify obstacles in the target image, obtain mask information (i.e., classification labels) for different obstacles, and then determine the obstacle area based on the mask information of each obstacle. This obstacle area can refer to the area corresponding to a single obstacle or to the areas corresponding to multiple obstacles, depending on the number of obstacles identified. For example, assuming that obstacles 1, 2, and 3 are identified in the target image, the obstacle area includes the areas corresponding to obstacle 1, obstacle 2, and obstacle 3. Furthermore, assuming that only obstacle 4 is identified in the target image, the obstacle area only includes the area corresponding to obstacle 4.
[0059] Step S103: Project the sparse point cloud to the pixel coordinate system to obtain a set of projection points corresponding to the sparse point cloud.
[0060] Specifically, the above-mentioned camera coordinate system refers to a three-dimensional rectangular coordinate system centered on the camera, which is used to describe the spatial position of an object relative to the optical center of the camera. In this camera coordinate system, the origin is the optical center of the camera (i.e., the center of the lens); the Z axis (i.e., the optical axis) is consistent with the direction of the camera lens and points directly in front of the scene; the X axis and Y axis are usually parallel to the sensor plane of the camera, with X pointing to the right and Y pointing down (or up, depending on the specific definition). The above-mentioned pixel coordinate system refers to a two-dimensional coordinate system that describes the position of pixels in an image, with the upper left corner or center of the digital image as the origin. In this pixel coordinate system, the origin is usually the upper left corner or center of the image. The u-axis (i.e., the horizontal axis) is in the positive direction to the right, corresponding to the column of the image. The v-axis (i.e., the vertical axis) is in the positive direction downward, corresponding to the row of the image. The above-mentioned set of projection points may include multiple projection points, each projection point corresponding one-to-one to each sparse point in the sparse point cloud.
[0061] Step S104 : Compare the projection point set and the obstacle area to obtain a target point cloud, wherein the target point cloud is a set of sparse points corresponding to the projection points falling into the obstacle area.
[0062] After projecting the sparse point cloud from the camera coordinate system to the pixel coordinate system, the projected point set corresponding to the sparse point cloud and the obstacle area can be compared to determine the target point cloud whose projected points fall into the obstacle area from the sparse point cloud.
[0063] Step S105: Generate a new point cloud based on the target point cloud, and combine the new point cloud with the sparse point cloud to obtain a dense point cloud.
[0064] After determining the target point cloud whose projection point falls into the obstacle area from the sparse point cloud, a new point cloud can be generated based on the target point cloud, so that the new point cloud can be combined with the original sparse point cloud to obtain a dense point cloud. Specifically, when generating a new point cloud based on the target point cloud, the new point cloud can be generated based on some of the sparse points in the target point cloud, or based on all of the sparse points in the target point cloud. Moreover, when generating a new point cloud based on the target point cloud, the number of new point clouds generated by different sparse points in the target point cloud can be the same or different, and this embodiment of the application does not specifically limit this.
[0065] Through the above method, the set of projection points corresponding to the sparse point cloud can be compared with the obstacle area in the target image, so as to determine the target point cloud where the projection point falls into the obstacle area, and then generate a new point cloud based on the target point cloud to obtain a dense point cloud. This can provide richer three-dimensional spatial information, greatly improve the robot's perception of the surrounding environment, and enable it to accurately detect and avoid obstacles; at the same time, since the present application mainly relies on a monocular camera to implement, there is no need to fuse information from other sensors such as lidar to generate a dense point cloud, so the cost is lower and it has stronger market competitiveness.
[0066] In an optional embodiment, the above step S105, generating a new point cloud based on the target point cloud, includes:
[0067] Obtain the target projection point corresponding to the target sparse point, where the target sparse point is any sparse point in the target point cloud;
[0068] Sampling is performed based on the target projection point to obtain a sampling point set;
[0069] Based on the sampling point set, a sub-point cloud is obtained, wherein the new point cloud includes multiple sub-point clouds, and the sub-point clouds correspond one-to-one to the sparse points in the target point cloud.
[0070] Specifically, for any sparse point in the target point cloud, its corresponding projection point can be obtained, and then sampling can be performed based on its corresponding projection point to obtain a sampling point set. Then, based on the sampling point set, a sub-point cloud can be obtained. It should be noted that the sampling point set here can include multiple sampling points, and the number of sampling points in the sampling point set can be set according to actual needs, and is not specifically limited in the embodiments of this application. The number of sub-point clouds here can correspond to the number of sparse points in the target point cloud. The sub-point clouds corresponding to each sparse point in the target point cloud together constitute a new point cloud.
[0071] Using the same method as above, each sparse point in the target point cloud can be processed accordingly, and the sub-point cloud corresponding to each sparse point in the target point cloud can be obtained, thereby forming a new point cloud, which is convenient for the subsequent densification of the original sparse point cloud based on the new point cloud. This can provide richer three-dimensional spatial information, greatly enhance the robot's ability to perceive the surrounding environment, and enable it to accurately detect and avoid obstacles.
[0072] In an optional embodiment, the above step of sampling based on the target projection point to obtain a sampling point set includes:
[0073] Get the coordinate position of the target projection point in the pixel coordinate system;
[0074] Based on the coordinate position, determine the target sampling area, where the target sampling area is an area with the target projection point as the center and a preset distance as the radius, and the target sampling area is located in the obstacle area;
[0075] Sampling is performed within the sampling area to obtain a preset number of sampling points, thereby forming a sampling point set.
[0076] Specifically, when sampling based on the target projection point to obtain a set of sampling points, the coordinate position of the target projection point in the pixel coordinate system can be obtained first, and then based on the coordinate position, the target sampling area can be determined, and then sampling is performed within the sampling area to obtain a preset number of sampling points, thereby forming a set of sampling points. It should be noted that the target sampling area is an area with the target projection point as the center and a preset distance as the radius, and the target sampling area is located in the obstacle area. The preset number and preset distance here can be set according to actual needs and are not specifically limited here. As an optional embodiment, the target sampling area can be within the neighboring pixels of the target projection point (such as a 7×7, 5×5 or 3×3 window), and the neighboring pixels are within the obstacle area.
[0077] Using the same method as above, each sparse point in the target point cloud can be processed accordingly, and a preset number of sampling points can be randomly selected from their respective sampling areas to obtain a set of sampling points corresponding to each sparse point, thereby forming a new point cloud. This makes it convenient to densify the original sparse point cloud based on the new point cloud. This can provide richer three-dimensional spatial information, greatly enhance the robot's ability to perceive the surrounding environment, and enable it to accurately detect and avoid obstacles.
[0078] In an optional embodiment, the above steps of obtaining a sub-point cloud based on the sampling point set include:
[0079] Obtain the depth value corresponding to the target projection point, and determine the depth value corresponding to the target projection point as the depth value of each sampling point in the sampling point set;
[0080] Obtain the intrinsic parameter matrix corresponding to the monocular camera, where the intrinsic parameter matrix is used to characterize the internal geometric characteristics of the monocular camera;
[0081] Based on the depth value and intrinsic parameter matrix of each sampling point in the sampling point set, each sampling point in the sampling point set is projected to the camera coordinate system to obtain a sub-point cloud.
[0082] Specifically, when obtaining a sub-point cloud based on a sampling point set, the depth value corresponding to the target projection point can be obtained, and the depth value corresponding to the target projection point can be determined as the depth value of each sampling point in the sampling point set, and the intrinsic parameter matrix corresponding to the monocular camera is obtained. Then, based on the depth value and intrinsic parameter matrix of each sampling point in the sampling point set, each sampling point in the sampling point set is projected to the camera coordinate system to obtain a sub-point cloud. Among them, the depth value corresponding to the target projection point can be obtained by triangulation or other methods, which is not limited in this application. The intrinsic parameter matrix can be used to characterize the internal geometric characteristics of the monocular camera, such as focal length, optical center, etc. The intrinsic parameter matrix can be expressed as:
[0083]
[0084] Among them, f x and f y is the focal length of the camera in the x and y directions, usually calculated from the relationship between the focal length of the camera and the pixel size. x and c y is the optical center of the image, representing the origin of the image coordinate system, which is usually close to the center of the image. For example, assuming that the coordinates of a sampling point in the sampling point set in the pixel coordinate system are (u, v), and the depth value of the sampling point is z, then the three-dimensional coordinates of the sampling point in the camera coordinate system can be inferred by combining the intrinsic parameter matrix:
[0085] X=z·(uc x ) / f x ;
[0086] Y=z·(vc y ) / f y ;
[0087] Among them, f x and f y is the focal length of the camera in the x and y directions, usually calculated from the relationship between the focal length of the camera and the pixel size. x and c y The optical center of the image represents the origin of the image coordinate system, which is usually close to the center of the image. (u, v) represents the coordinates of a sampling point in the pixel coordinate system, and (x, y, z) represents the coordinates of the sampling point in the camera coordinate system.
[0088] Through the above method, the three-dimensional coordinates corresponding to each sampling point in the sampling point set can be obtained, thereby obtaining a sub-point cloud, which is convenient for subsequently obtaining a new point cloud based on multiple sub-point clouds, thereby obtaining a dense point cloud, thereby providing richer three-dimensional spatial information, greatly improving the robot's ability to perceive the surrounding environment, enabling it to accurately detect and avoid obstacles.
[0089] In an optional embodiment, the above step 103 of projecting the sparse point cloud into a pixel coordinate system to obtain a set of projection points corresponding to the sparse point cloud includes:
[0090] Normalize the coordinates of each sparse point in the sparse point cloud in the camera coordinate system to obtain the normalized processing results corresponding to each sparse point;
[0091] Based on the normalization processing results corresponding to each sparse point and the intrinsic parameter matrix corresponding to the monocular camera, the projection position of each sparse point in the pixel coordinate system is determined;
[0092] Based on the projection position of each sparse point in the pixel coordinate system, each sparse point is projected to the pixel coordinate system to obtain the projection point corresponding to each sparse point, thereby obtaining the projection point set corresponding to the sparse point cloud.
[0093] When projecting the sparse point cloud to the pixel coordinate system and obtaining the projection point set corresponding to the sparse point cloud, the coordinates of each sparse point in the sparse point cloud in the camera coordinate system can be normalized to obtain the normalized processing results corresponding to each sparse point. For example, for a sparse point P = [x, y, z] T The coordinates in the camera coordinate system are normalized, and the process can be expressed by the following formula:
[0094]
[0095] Among them, x, y and z represent the coordinates of the sparse point in the x, y and z directions respectively, p norm Represents the normalized processing result corresponding to the sparse point. Then, based on the normalized processing results corresponding to each sparse point and the intrinsic parameter matrix corresponding to the monocular camera, the projection position of each sparse point in the pixel coordinate system is determined. This process can be expressed by the following formula:
[0096]
[0097] Among them, f x and f y is the focal length of the camera in the x and y directions, usually calculated from the relationship between the focal length of the camera and the pixel size. x and c y is the optical center of the image, which indicates the origin of the image coordinate system, usually close to the center of the image. K represents the intrinsic parameter matrix, x, y, and z represent the coordinates of a sparse point in the x, y, and z directions respectively, and p norm Indicates the normalized processing result corresponding to the sparse point, p pixel Represents the coordinates of the sparse point in the pixel coordinate system.
[0098] Through the above method, each sparse point in the sparse point cloud can be projected into the pixel coordinate system to obtain the projection point set corresponding to the sparse point cloud, which is convenient for subsequent position matching of the projection point set with the obstacle area to determine the sparse point cloud belonging to the obstacle (i.e., the target point cloud).
[0099] In an optional embodiment, before the above step S103 of projecting the sparse point cloud into the pixel coordinate system to obtain a set of projection points corresponding to the sparse point cloud, the method further includes:
[0100] Convert the sparse point cloud from the camera coordinate system to the vehicle body coordinate system;
[0101] Filtering the sparse point cloud in the vehicle body coordinate system to obtain a filtered sparse point cloud;
[0102] Convert the filtered sparse point cloud from the vehicle body coordinate system back to the camera coordinate system;
[0103] The above step S103, projecting the sparse point cloud to the pixel coordinate system to obtain a set of projection points corresponding to the sparse point cloud, includes:
[0104] The filtered sparse point cloud in the camera coordinate system is projected to the pixel coordinate system to obtain a set of projection points corresponding to the filtered sparse point cloud.
[0105] Specifically, before projecting the sparse point cloud to the pixel coordinate system, the sparse point cloud can be first converted from the camera coordinate system to the vehicle coordinate system, and then the sparse point cloud in the vehicle coordinate system is filtered to obtain a filtered sparse point cloud, and then the filtered sparse point cloud is converted from the vehicle coordinate system back to the camera coordinate system, and then the filtered sparse point cloud in the camera coordinate system is projected to the pixel coordinate system. Among them, the camera coordinate system takes the optical center of the camera as the origin, usually the Z axis is along the optical axis, the X axis is to the right, and the Y axis is downward. The body coordinate system takes the center of the vehicle (such as the rear axle center or the center of gravity) as the origin, usually the X axis is forward, the Y axis is to the left, and the Z axis is upward. When filtering the sparse point cloud in the vehicle coordinate system, it needs to be done in the vehicle coordinate system, so it is necessary to convert from the camera coordinate system to the vehicle coordinate system before filtering, and convert from the vehicle coordinate system back to the camera coordinate system after filtering. When filtering the sparse point cloud in the vehicle coordinate system, you can apply a pass-through filter to the sparse point cloud in the vehicle coordinate system to remove point clouds that do not meet a specific range. For example, you can filter to retain point clouds that meet the following conditions:
[0106] Vehicle forward direction (i.e. longitudinal distance): 0 m ≤ x ≤ 1.5 m;
[0107] Left and right direction of vehicle body (i.e. lateral distance): -0.5m≤y≤0.5m;
[0108] Vertical direction of the vehicle body (i.e. height distance): -0.5m≤z≤0.5m;
[0109] Through the above method, point clouds that are far away or too high or too low can be effectively removed, thereby optimizing the quality of sparse point clouds and reducing the computational complexity of subsequent processing.
[0110] In an optional embodiment, the above step of converting the sparse point cloud from the camera coordinate system to the vehicle body coordinate system includes:
[0111] Get the extrinsic parameter matrix corresponding to the monocular camera, where the extrinsic parameter matrix is used to represent the rotation matrix and translation vector from the camera coordinate system to the vehicle coordinate system;
[0112] Based on the extrinsic matrix, the sparse point cloud is transformed from the camera coordinate system to the vehicle coordinate system.
[0113] Specifically, the above extrinsic matrix can be used to represent the rotation matrix and translation vector from the camera coordinate system to the vehicle coordinate system. The extrinsic matrix can be expressed as:
[0114]
[0115] in, Represents the extrinsic parameter matrix, R represents the rotation matrix, which is used to represent the rotation of the camera coordinate system relative to the vehicle coordinate system, and t represents the translation vector, which is used to represent the translation of the camera coordinate system relative to the vehicle coordinate system.
[0116] When converting the sparse point cloud from the camera coordinate system to the vehicle coordinate system, we can first obtain the extrinsic parameter matrix corresponding to the monocular camera, and then convert the sparse point cloud from the camera coordinate system to the vehicle coordinate system based on the extrinsic parameter matrix. For example, if the coordinates of a point in the camera coordinate system are P c =[x c ,y c , z c , 1] T , transformed to the vehicle coordinate system through the external parameter matrix:
[0117]
[0118] Among them, P c Represents the coordinates of a point in the camera coordinate system, p body Represents the coordinates of the point in the vehicle coordinate system, Represents the extrinsic parameter matrix, R represents the rotation matrix, which is used to represent the rotation of the camera coordinate system relative to the vehicle coordinate system, and t represents the translation vector, which is used to represent the translation of the camera coordinate system relative to the vehicle coordinate system.
[0119] Through the above method, the sparse point cloud can be converted from the camera coordinate system to the vehicle coordinate system, which facilitates the subsequent filtering of the sparse point cloud in the vehicle coordinate system.
[0120] In an optional embodiment, the dense point cloud generation process based on a monocular camera provided in this application can be as follows: Figure 2 As shown, the following steps may be specifically included:
[0121] Step S201: Use the preset VIO system to output a sparse point cloud.
[0122] Step S202: convert the sparse point cloud from the camera coordinate system to the vehicle body coordinate system.
[0123] Step S203: Filter the sparse point cloud in the vehicle body coordinate system to obtain a filtered sparse point cloud.
[0124] Step S204: convert the filtered sparse point cloud from the vehicle body coordinate system back to the camera coordinate system.
[0125] Step S205 : Projecting the filtered sparse point cloud in the camera coordinate system to the pixel coordinate system to obtain a set of projection points corresponding to the filtered sparse point cloud.
[0126] Step S206: Acquire the target image captured by the monocular camera.
[0127] Step S207: Identify obstacles in the target image to obtain obstacle areas.
[0128] Step S208: Determine whether the projection point in the projection point set is located in the obstacle area.
[0129] Step S209: Select a preset number of sampling points from the obstacle area, and determine the depth value corresponding to the projection point as the depth value of the sampling point.
[0130] Step S210: Based on the depth value and the intrinsic parameter matrix of the sampling point, the sampling point is projected to the camera coordinate system to obtain a new point cloud, and a dense point cloud is obtained based on the new point cloud and the sparse point cloud.
[0131] Step S211: directly discard the projection point.
[0132] Among them, the above steps S201 to S205 can be performed before the above steps S206 to S207, or after the above steps S206 to S207, or simultaneously with the above steps S206 to S207, which is not specifically limited in the embodiments of the present application.
[0133] The monocular camera-based dense point cloud generation method provided in this application generates a sparse point cloud based on a monocular camera, an inertial measurement unit (IMU), and an odometry. After the sparse point cloud is generated, a dense point cloud is obtained through point cloud optimization, segmentation, projection, and densification. As a result, the following effects can be achieved:
[0134] (1) Improve point cloud density: The generated dense point cloud can provide richer three-dimensional spatial information, enhance the robot's perception of the surrounding environment, and is particularly suitable for complex environments.
[0135] (2) Reduce costs: Only a monocular camera, IMU, and odometer are used, without the need for additional expensive sensors (such as lidar), which reduces the hardware cost of the system.
[0136] (3) Realize real-time processing: The entire process can be completed in real time during the operation of the robot, ensuring the response speed and reliability of the system.
[0137] (4) Enhanced obstacle detection: It can accurately identify various obstacles in the image, improve the accuracy of obstacle detection, and avoid collisions or misjudgments.
[0138] See also Figure 3 , Figure 3 This is a schematic diagram of the structure of a dense point cloud generation device based on a monocular camera provided in an embodiment of the present application. Figure 3 As shown, the dense point cloud generation device 300 based on a monocular camera includes:
[0139] An acquisition module 301 is configured to acquire a target image and an image sequence containing the target image captured by a monocular camera, and generate a sparse point cloud based on the image sequence;
[0140] The recognition module 302 is used to identify obstacles in the target image and obtain an obstacle area, wherein the target image is formed based on a pixel coordinate system;
[0141] The projection module 303 is used to project the sparse point cloud into a pixel coordinate system to obtain a set of projection points corresponding to the sparse point cloud;
[0142] A comparison module 304 is configured to compare the projection point set and the obstacle region to obtain a target point cloud, wherein the target point cloud is a set of sparse points corresponding to the projection points falling into the obstacle region;
[0143] The generating module 305 is configured to generate a new point cloud based on the target point cloud, and combine the new point cloud with the sparse point cloud to obtain a dense point cloud.
[0144] Furthermore, the generating module 305 includes:
[0145] A first acquisition submodule is used to acquire a target projection point corresponding to a target sparse point, wherein the target sparse point is any sparse point in the target point cloud;
[0146] The sampling submodule is used to perform sampling based on the target projection point to obtain a set of sampling points;
[0147] The first determination submodule is configured to obtain a sub-point cloud based on the sampling point set, wherein the new point cloud includes a plurality of sub-point clouds, and the sub-point clouds correspond one-to-one to the sparse points in the target point cloud.
[0148] Furthermore, the sampling submodule includes:
[0149] A first acquisition unit is used to acquire the coordinate position of the target projection point in the pixel coordinate system;
[0150] a determination unit, configured to determine a target sampling area based on the coordinate position, wherein the target sampling area is an area centered on the target projection point and having a preset distance as a radius, and the target sampling area is located within the obstacle area;
[0151] The sampling unit is used to perform sampling in the sampling area to obtain a preset number of sampling points, thereby forming a sampling point set.
[0152] Furthermore, the first determining submodule includes:
[0153] A second acquiring unit is configured to acquire a depth value corresponding to a target projection point, and determine the depth value corresponding to the target projection point as a depth value of each sampling point in the sampling point set;
[0154] A third acquisition unit is used to obtain an intrinsic parameter matrix corresponding to the monocular camera, wherein the intrinsic parameter matrix is used to characterize the internal geometric characteristics of the monocular camera;
[0155] The projection unit is used to project each sampling point in the sampling point set to the camera coordinate system based on the depth value and the intrinsic parameter matrix of each sampling point in the sampling point set to obtain a sub-point cloud.
[0156] Furthermore, the projection module 303 includes:
[0157] The processing submodule is used to normalize the coordinates of each sparse point in the sparse point cloud in the camera coordinate system to obtain the normalized processing results corresponding to each sparse point;
[0158] The second determination submodule is used to determine the projection position of each sparse point in the pixel coordinate system based on the normalization processing result corresponding to each sparse point and the intrinsic parameter matrix corresponding to the monocular camera;
[0159] The projection submodule is used to project each sparse point to the pixel coordinate system based on the projection position of each sparse point in the pixel coordinate system, obtain the projection point corresponding to each sparse point, and thus obtain the projection point set corresponding to the sparse point cloud.
[0160] Furthermore, the dense point cloud generation device 300 based on a monocular camera further includes:
[0161] A first conversion module, configured to convert the sparse point cloud from a camera coordinate system to a vehicle coordinate system;
[0162] A filtering processing module is used to filter the sparse point cloud in the vehicle body coordinate system to obtain a filtered sparse point cloud;
[0163] A second conversion module is used to convert the filtered sparse point cloud from the vehicle body coordinate system back to the camera coordinate system;
[0164] The projection module is further used to project the filtered sparse point cloud in the camera coordinate system to the pixel coordinate system to obtain a set of projection points corresponding to the filtered sparse point cloud.
[0165] Furthermore, the first conversion module includes:
[0166] The second acquisition submodule is used to obtain the extrinsic parameter matrix corresponding to the monocular camera, wherein the extrinsic parameter matrix is used to represent the rotation matrix and translation vector from the camera coordinate system to the vehicle coordinate system;
[0167] The first conversion submodule is used to convert the sparse point cloud from the camera coordinate system to the vehicle body coordinate system based on the extrinsic parameter matrix.
[0168] It should be noted that the dense point cloud generation device 300 based on a monocular camera can implement the steps of the dense point cloud generation method based on a monocular camera provided in any of the aforementioned method embodiments, and can achieve the same technical effects, which will not be repeated here.
[0169] like Figure 4 As shown, the embodiment of the present application further provides an electronic device, including a processor 411, a communication interface 412, a memory 413 and a communication bus 414, wherein the processor 411, the communication interface 412, and the memory 413 communicate with each other through the communication bus 414.
[0170] Memory 413, for storing computer programs;
[0171] In one embodiment of the present application, the processor 411 is configured to implement the monocular camera-based dense point cloud generation method provided in any one of the aforementioned method embodiments when executing the program stored in the memory 413 .
[0172] In actual applications, the electronic device can be a lawn mower robot, sweeping robot or other equipment equipped with a monocular camera, IMU and odometer, or it can be a control device installed on a lawn mower robot, sweeping robot or other equipment. The embodiments of this application do not make specific limitations.
[0173] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, a dense point cloud generation method based on a monocular camera as provided in any of the aforementioned method embodiments is implemented.
[0174] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0175] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is intended to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for generating dense point clouds based on a monocular camera, characterized in that: The method comprises: Acquire a target image captured by a monocular camera and an image sequence containing the target image, and generate a sparse point cloud based on the image sequence; Identifying obstacles in the target image to obtain an obstacle area, wherein the target image is formed based on a pixel coordinate system; Projecting the sparse point cloud to a pixel coordinate system to obtain a set of projection points corresponding to the sparse point cloud; Comparing the projection point set and the obstacle area to obtain a target point cloud, wherein the target point cloud is a set of sparse points corresponding to the projection points falling into the obstacle area; A new point cloud is generated based on the target point cloud, and the new point cloud is combined with the sparse point cloud to obtain a dense point cloud.
2. The method according to claim 1, characterized in that Generating a new point cloud based on the target point cloud includes: Obtaining a target projection point corresponding to a target sparse point, wherein the target sparse point is any sparse point in the target point cloud; Sampling is performed based on the target projection point to obtain a sampling point set; Based on the sampling point set, a sub-point cloud is obtained, wherein the new point cloud includes a plurality of the sub-point clouds, and the sub-point clouds correspond one-to-one to the sparse points in the target point cloud.
3. The method according to claim 2, characterized in that The sampling based on the target projection point to obtain a sampling point set includes: Obtaining the coordinate position of the target projection point in the pixel coordinate system; Determine a target sampling area based on the coordinate position, wherein the target sampling area is an area with the target projection point as the center and a preset distance as the radius, and the target sampling area is located in the obstacle area; Sampling is performed within the sampling area to obtain a preset number of sampling points, thereby forming the sampling point set.
4. The method according to claim 2, characterized in that The step of obtaining a sub-point cloud based on the sampling point set includes: Obtaining a depth value corresponding to the target projection point, and determining the depth value corresponding to the target projection point as the depth value of each sampling point in the sampling point set; Obtaining an intrinsic parameter matrix corresponding to the monocular camera, wherein the intrinsic parameter matrix is used to characterize geometric characteristics inside the monocular camera; Based on the depth value of each sampling point in the sampling point set and the intrinsic parameter matrix, each sampling point in the sampling point set is projected to a camera coordinate system to obtain the sub-point cloud.
5. The method according to claim 1, wherein The projecting the sparse point cloud to a pixel coordinate system to obtain a set of projection points corresponding to the sparse point cloud includes: Normalizing the coordinates of each sparse point in the sparse point cloud in the camera coordinate system to obtain a normalized processing result corresponding to each sparse point; Determining the projection position of each sparse point in the pixel coordinate system based on the normalization processing result corresponding to each sparse point and the intrinsic parameter matrix corresponding to the monocular camera; Based on the projection position of each sparse point in the pixel coordinate system, each sparse point is projected to the pixel coordinate system to obtain the projection point corresponding to each sparse point, thereby obtaining a projection point set corresponding to the sparse point cloud.
6. The method according to claim 5, characterized in that Before projecting the sparse point cloud to a pixel coordinate system to obtain a set of projection points corresponding to the sparse point cloud, the method further includes: Converting the sparse point cloud from a camera coordinate system to a vehicle body coordinate system; filtering the sparse point cloud in the vehicle body coordinate system to obtain a filtered sparse point cloud; Converting the filtered sparse point cloud from the vehicle body coordinate system back to the camera coordinate system; The projecting the sparse point cloud to a pixel coordinate system to obtain a set of projection points corresponding to the sparse point cloud includes: The filtered sparse point cloud in the camera coordinate system is projected to a pixel coordinate system to obtain a set of projection points corresponding to the filtered sparse point cloud.
7. The method according to claim 6, characterized in that The converting the sparse point cloud from the camera coordinate system to the vehicle coordinate system includes: Obtaining an extrinsic parameter matrix corresponding to the monocular camera, wherein the extrinsic parameter matrix is used to represent a rotation matrix and a translation vector from the camera coordinate system to the vehicle coordinate system; Based on the extrinsic parameter matrix, the sparse point cloud is transformed from the camera coordinate system to the vehicle body coordinate system.
8. A dense point cloud generation device based on a monocular camera, characterized in that: The device comprises: An acquisition module is used to acquire a target image captured by a monocular camera and an image sequence containing the target image, and generate a sparse point cloud based on the image sequence; an identification module, configured to identify obstacles in the target image and obtain an obstacle area, wherein the target image is formed based on a pixel coordinate system; A projection module, configured to project the sparse point cloud into a pixel coordinate system to obtain a set of projection points corresponding to the sparse point cloud; a comparison module, configured to compare the projection point set and the obstacle region to obtain a target point cloud, wherein the target point cloud is a set of sparse points corresponding to the projection points falling into the obstacle region; A generation module is used to generate a new point cloud based on the target point cloud, and combine the new point cloud with the sparse point cloud to obtain a dense point cloud.
9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; The processor is configured to implement the method for generating a dense point cloud based on a monocular camera as described in any one of claims 1 to 7 when executing a program stored in the memory.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating a dense point cloud based on a monocular camera according to any one of claims 1 to 7 is implemented.