Map construction method and device, electronic equipment and storage medium
By combining object detection and 3D feature point detection results with line detection results for joint optimization, the problem of inaccurate semantic mapping was solved, and high-precision semantic map construction and visualization were achieved, enhancing the comprehensiveness and operability of the map.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
- Filing Date
- 2023-06-08
- Publication Date
- 2026-05-15
AI Technical Summary
In existing technologies, semantic mapping performance is affected by the visual positioning system, resulting in incomplete and low-accuracy maps. Reliance on depth images leads to inaccurate construction.
Semantic information of objects is obtained through object detection. Combined with the results of 3D feature point and line detection, a semantic map is constructed through joint optimization. This map is then fused with a semi-dense map to generate a semi-dense semantic map.
It improves the accuracy and completeness of semantic maps, reduces the dependence on depth images, enhances the comprehensiveness and operability of maps, and supports the construction and visualization of high-precision object semantic maps.
Smart Images

Figure CN116597011B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of visual positioning technology, and more specifically, to a map building method, apparatus, electronic device, and computer-readable storage medium. Background Technology
[0002] Smart devices such as cameras generally have scene understanding capabilities. To improve scene understanding capabilities, it is necessary to improve the effect of semantic mapping.
[0003] In related technologies, semantic mapping can be achieved by combining neural networks and Bayesian models to achieve dense semantic mapping, segmenting based on RGBD image streams to achieve object-level semantic mapping, combining with visual systems, or combining segmentation networks and dense mapping.
[0004] In the above methods, the semantic mapping performance is closely linked to the visual positioning system and is affected by the visual positioning system, thus having certain limitations; at the same time, some semantic mapping relies on depth maps, and the generated maps are not complete enough and have low accuracy. Summary of the Invention
[0005] The purpose of this disclosure is to provide a map construction method and apparatus, electronic device and computer-readable storage medium, thereby overcoming, at least to some extent, the problem of map inaccuracy caused by the limitations and defects of related technologies.
[0006] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part from practice of this disclosure.
[0007] According to a first aspect of this disclosure, a map construction method is provided, comprising: acquiring an image to be processed, and performing object detection on the image to be processed to obtain semantic information of objects in the image to be processed; performing pose estimation on the image to be processed to obtain three-dimensional feature points and camera pose; performing line detection on the image to be processed to obtain line detection results; fusing the semantic information, the three-dimensional feature points and the line detection results to obtain an initial pose of a three-dimensional detection box, and jointly optimizing the camera pose and the initial pose to construct a semantic map; constructing a semi-dense map corresponding to the image to be processed, and fusing the semantic map and the semi-dense map to obtain a semi-dense semantic map of the image to be processed.
[0008] According to a second aspect of this disclosure, a map building apparatus is provided, comprising: a target detection module for acquiring an image to be processed and performing target detection on the image to obtain semantic information of objects in the image; a pose estimation module for performing pose estimation on the image to be processed to obtain three-dimensional feature points and camera pose; a line detection module for performing line detection on the image to be processed to obtain line detection results; a semantic map building module for fusing the semantic information, the three-dimensional feature points, and the line detection results to obtain an initial pose of a three-dimensional detection box, and jointly optimizing the camera pose and the initial pose to build a semantic map; and a map generation module for building a semi-dense map corresponding to the image to be processed, and fusing the semantic map and the semi-dense map to obtain a semi-dense semantic map of the image to be processed.
[0009] According to a third aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the method described in any of the preceding methods and possible implementations thereof by executing the executable instructions.
[0010] According to a fourth aspect of this disclosure, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the method described in any of the preceding claims and possible implementations thereof.
[0011] In the technical solution provided in this disclosure, on the one hand, semantic information of objects in the image to be processed is obtained through object detection. Then, the initial pose of the 3D detection box can be obtained based on the semantic information, 3D feature points, and line detection results, achieving a high-precision reconstruction process of the 3D object detection box. Combining the object's semantic information with the line detection results of the image can improve the accuracy of object detection. Combining the object's semantic information with the camera pose of visual SLAM for joint optimization can improve the precise pose of the 3D detection box of the object semantic map, thus improving the accuracy and precision of semantic map construction. On the other hand, a semi-dense map is constructed for the image to be processed to achieve visualization. A semi-dense semantic map is obtained by fusing the semi-dense map and the semantic map, which can accurately provide the pose of the category of interest and the 3D detection box diagram in the semi-dense semantic map, achieving high-precision construction of the object semantic map, improving the map's completeness and comprehensiveness, and also improving the visualization effect. Furthermore, it avoids the dependence on depth images when constructing maps in related technologies, increasing the comprehensiveness and operability of map construction.
[0012] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0013] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0014] Figure 1 A schematic diagram illustrates an application scenario where the map construction method of the present disclosure embodiments can be applied.
[0015] Figure 2 The illustration shows a flowchart of a map construction method according to an embodiment of the present disclosure.
[0016] Figure 3 The schematic diagram illustrates the structure of the target detection model in an embodiment of this disclosure.
[0017] Figure 4 The schematic diagram illustrates a pose estimation process in an embodiment of this disclosure.
[0018] Figure 5 This illustration schematically shows a diagram of estimating camera pose based on camera motion patterns in an embodiment of the present disclosure.
[0019] Figure 6 The schematic diagram illustrates a flow chart of line detection in an embodiment of this disclosure.
[0020] Figure 7 The illustration shows a schematic diagram of the process for constructing a semi-dense map according to an embodiment of the present disclosure.
[0021] Figure 8 A schematic diagram of a semi-dense map according to an embodiment of the present disclosure is shown.
[0022] Figure 9 A schematic diagram of a semi-dense semantic map according to an embodiment of the present disclosure is shown.
[0023] Figure 10 This diagram illustrates the overall process of map construction according to an embodiment of the present disclosure.
[0024] Figure 11 A schematic block diagram of a map building apparatus according to an embodiment of the present disclosure is shown.
[0025] Figure 12 A block diagram of an electronic device according to an embodiment of the present disclosure is shown schematically. Detailed Implementation
[0026] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0027] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0028] To address the technical problems in related technologies, this disclosure provides a map construction method applicable to indoor semantic mapping, outdoor semantic mapping, or semantic mapping in other scenarios. The intelligent device represented by the terminal can be a mobile robot, smartphone, drone, AR / VR, etc. Once a semi-dense semantic map of the image to be processed is obtained, the intelligent device can better understand its surrounding environment and accurately and efficiently perform functions such as navigation and obstacle avoidance based on the semi-dense semantic map. Figure 1 A schematic diagram of the system architecture of the map building method and apparatus to which embodiments of the present disclosure can be applied is shown.
[0029] like Figure 1 As shown, when the terminal is an intelligent robot, the intelligent robot obtains RGB image data streams through an RGB monocular camera in an indoor home environment. First, it performs target detection to obtain semantic information; second, it implements a monocular vision SLAM system and combines the target detection semantic information to complete backend optimization to achieve a lightweight object semantic map; finally, it implements semi-dense mapping of the monocular RGB image and merges it with the object semantic map to achieve navigation and obstacle avoidance functions of the mobile robot.
[0030] It should be noted that the map building method provided in this embodiment can be executed by a terminal. The terminal can be a smart device with map building functionality, such as a smartphone, computer, tablet, smart speaker, smartwatch, in-vehicle device, wearable device, monitoring device, robot, or other smart device. The map building device can also be installed in the terminal. Alternatively, the map building method can be executed by a server and then deployed to the terminal for execution; no specific limitations are specified here.
[0031] Next, refer to Figure 2 The diagram illustrates each step of the map construction method.
[0032] In step S210, the image to be processed is acquired, and target detection is performed on the image to be processed to obtain the semantic information corresponding to the objects in the image to be processed.
[0033] In this embodiment, the image to be processed can be an RGB image. Image acquisition can be performed using a sensor. The image sensor can include an RGB monocular camera, capable of acquiring RGB images at a frequency of 30Hz. The image to be processed can be an image corresponding to an indoor environment, which may contain one or more objects or other objects, without specific limitations here.
[0034] After acquiring the image to be processed, object detection can be performed on it. Object detection refers to the detection and recognition of objects in the image to be processed. Object detection can be divided into two sub-tasks: object localization (object detection bounding box) and object classification (object label). In this embodiment of the disclosure, object detection can be implemented using an object detection model.
[0035] An example of an object detection model is the YOLOv3 model. The backbone network of the YOLOv3 model is DarkNet-53, with an input size of 416×416. It reduces the feature map size by increasing the stride through convolutional kernels, obtaining feature maps at three different scales to capture both deep and shallow features of the image. This allows for high-precision detection and recognition of objects of varying sizes. The overall framework of the object detection model is shown below. Figure 3 As shown.
[0036] refer to Figure 3 As shown, the object detection model comprises three components: DBL, Res_unit, and Resblock_body. The DBL component consists of three parts: convolution, BN layer normalization, and Leaky ReLU activation function. Res_unit is a residual structure composed of multiple DBLs, making the network structure deeper. Resblock_body consists of pooling, DBL, and Res_unit, i.e. Figure 3The `resn` module in the model represents the number of residual structures (`Res_unit`). This network lacks pooling and fully connected layers; instead, during forward propagation, it uses the stride of the convolutional kernels to scale the input image, resulting in three feature layers at different scales. Based on this object detection model, features can be extracted from the image at different scales, yielding feature maps of multiple scales. By fusing deep and shallow features, the final output consists of three feature maps of different sizes for subsequent detection.
[0037] For example, the image to be processed, acquired by an RGB monocular camera, can be used as input and preprocessed to meet the input requirements of the YOLOv3 network. Next, a target detection model can extract features from the image at different scales, resulting in multiple feature layers of varying scales. Further, the last feature layer is convolved to generate a first feature map Y1; the previous feature layer is concatenated with the intermediate vector (output of the last feature layer) corresponding to the upsampled first feature map to obtain a second feature map Y2; then, the adjacent feature layer above the previous feature layer is concatenated with the intermediate vector (output of the previous feature layer) corresponding to the upsampled second feature map to obtain a third feature map. The size of the third feature map is larger than that of the second feature map and also larger than that of the first feature map. This concatenation refers to tensor concatenation, where feature maps are directly concatenated along the channel dimension to expand the tensor's dimension. Upsampling generates a larger image from a small feature map using interpolation or other methods. The upsampling method used is pooling, which expands the feature map size through element copying without learning parameters.
[0038] After obtaining the first, second, and third feature maps, semantic information of objects in the image to be processed can be obtained through fitting calculations based on these maps. The semantic information may include object bounding boxes and categories. The bounding boxes represent the corresponding detection boxes for each object, and the categories represent the object labels.
[0039] The target detection method in this embodiment uses multiple scales to detect targets of different sizes, avoiding overlap with other algorithms, saving model space, and improving accuracy.
[0040] Next, refer to Figure 2 As shown, in step S220, the pose estimation of the image to be processed is performed to obtain three-dimensional feature points and camera pose.
[0041] In this embodiment of the disclosure, camera pose and 3D feature points can be estimated using a visual SLAM (Simultaneous Localization and Mapping) front-end. For example, the coordinates of the 3D feature points can be obtained through triangulation, and the camera pose can be estimated by further combining the camera motion mode, the 3D feature points, and feature matching pairs.
[0042] When performing camera pose estimation using SLAM, refer to Figure 4 As shown, the main steps include:
[0043] In step S410, the image to be processed is converted into a grayscale image, and feature extraction and feature matching are performed on the grayscale image to obtain feature matching pairs between the current frame and the previous frame;
[0044] In step S420, triangulation is performed based on feature matching pairs to obtain three-dimensional feature points, and pose estimation is performed by combining the camera motion mode, feature matching pairs, and three-dimensional feature points to obtain the camera pose.
[0045] In this embodiment, the RGB image is first converted to a grayscale image. Feature extraction is then performed based on the grayscale image, specifically using the ORB (Oriented Fast and Rotated Brief) feature extraction algorithm. ORB can be used to quickly create feature vectors from keypoints in an image, and these feature vectors can be used to identify objects in the image. ORB first identifies special regions in the image, called keypoints. Keypoints are small, prominent areas in the image, such as corners. Then, ORB calculates the corresponding feature vector for each keypoint. Of course, other feature extraction algorithms such as Fast or Brief can also be used for feature extraction; no specific limitation is made here.
[0046] After obtaining the feature vectors, feature matching can be performed on the feature vectors to obtain feature matching pairs between multiple images.
[0047] Furthermore, based on this feature matching pair, triangulation can be used to recover the 3D feature points. Specifically, triangulation can be performed using formula (1). Here, Pp represents the pixel coordinates of each frame, and Pw represents the world coordinate system coordinates (coordinates of the 3D feature points). K represents the camera's intrinsic parameter matrix and is known, R represents the camera's rotation matrix, and t represents the camera's translation matrix. The rotation and translation matrices constitute the camera pose. The camera pose can be used to describe the camera's position and angles, and may include, for example, six degrees of freedom, position coordinates, and three orientation angles.
[0048] The camera's rotation and translation matrices can be obtained by performing SVD (Singular Value Decomposition) directly using the Perspective-n-Point (PNP) method on given feature matching pairs. The PNP problem describes how to estimate the camera pose when n 3D feature points and their projected positions are known. Since the 3D feature points of the current frame are known, upon the arrival of a new frame, image matching can be used to obtain the corresponding 2D feature points. Then, based on these correspondences between 3D spatial points and 2D feature points, the PnP algorithm can be used to solve for the camera pose of the current frame.
[0049] Based on this, the world coordinate system coordinates Pw, i.e., the coordinates of the three-dimensional feature points, can be obtained according to formula (1). Among them, Q can be determined according to the camera's intrinsic parameter matrix K, the camera's rotation matrix R, and the translation matrix t.
[0050]
[0051] Next, camera pose can be estimated using different methods through camera motion modes. Camera motion modes can include an initial state and a constant-speed motion state. Figure 5 The diagram illustrates a flowchart for estimating camera pose based on camera motion patterns. (Refer to...) Figure 5 The following situations may be included as shown:
[0052] In step S510, in response to the camera motion mode being in the initial state, the camera pose is obtained by performing pose estimation through feature matching.
[0053] For example, the initial state can be a static mode or another mode. In this case, the camera pose can be obtained by performing singular value decomposition using the obtained feature matching pairs through the PNP direct method.
[0054] In step S520, in response to the camera motion mode being uniform motion, the camera pose of the current frame can be determined using the camera pose of the previous frame, and the reprojection error is obtained by projecting the three-dimensional feature points based on the camera pose of the current frame. The camera pose of the current frame is then determined based on the comparison between the reprojection error and the error threshold.
[0055] For example, in a uniform motion state, since there are already several frames of images, there may be a relative motion model between adjacent frames. This relative motion model can be used to match the two-dimensional feature points between two frames of images, and then the transformation relationship between the two frames of images can be represented based on the matching relationship.
[0056] In this scenario, the camera pose of the current frame can be determined based on the camera pose of the previous frame. Then, based on the current frame's camera pose, the 3D feature points are projected onto the current frame for reprojection to obtain the reprojection error. Further, the reprojection error can be compared with an error threshold. The comparison result determines whether the camera pose of the current frame should be determined based on the previous frame's camera pose. If the comparison result indicates that the reprojection error meets the error threshold, the camera pose of the previous frame can be used as the camera pose of the current frame.
[0057] When the comparison result shows that the reprojection error does not meet the error threshold, the camera pose of the previous frame cannot be determined as the camera pose of the current frame. At this time, feature matching can be performed between the current frame image and the reference frame image to obtain the correspondence between the three-dimensional points of the reference frame image and the two-dimensional feature points of the current frame. This correspondence is used to represent the transformation relationship between the reference frame image and the current frame image. The reference frame can be any frame located before the current frame and different from the previous frame. For example, it can be any frame located before the current frame except the previous frame. In addition, the camera pose of the previous frame can be determined as the initial camera pose of the current frame. The camera pose of the reference frame is obtained by multiplying the initial pose by the transformation relationship. Next, the least squares optimization problem constructed by PNP can be applied to solve the problem. The optimization quantity is the initial camera pose of the current frame, which minimizes the error between the initial camera pose of the current frame and the camera pose of the reference frame, and the camera pose with the smallest error is determined as the camera pose of the current frame. The optimization equation can be shown in formula (2), where i represents the i-th point and ξ represents the variable corresponding to the pose.
[0058]
[0059] In this embodiment of the disclosure, different methods are used to estimate the camera pose by means of camera motion modes, which can improve the accuracy of camera pose estimation.
[0060] Next, continue to refer to Figure 2 As shown, in step S230, line detection is performed on the image to be processed to obtain line detection results.
[0061] In this embodiment of the disclosure, line detection can be performed on the grayscale image corresponding to the image to be processed. Line detection can be used to detect straight lines in a grayscale image. Specifically, LSD line detection can be performed on the grayscale image. For example, refer to... Figure 6 As shown, performing line detection may include the following steps:
[0062] In step S610, the gradient and direction of each pixel in the image to be processed are determined, and all pixels are sorted according to their gradient values.
[0063] In step S620, region growing is performed at the pixel corresponding to the maximum gradient value, and multiple valid pixels in the region where the pixel is located are determined within the neighborhood of the pixel.
[0064] In step S630, the pixel cloud density of the region where multiple valid pixels are located is determined, and line detection is completed based on the comparison result of pixel cloud density and density threshold.
[0065] In some embodiments, the grayscale image can first be downsampled to reduce the jagged edges. Next, the gradient and direction of each pixel in the grayscale image are calculated, and the pixels are filtered, removing those with gradient values less than a preset value. This preset value can be set according to actual needs. Then, all remaining pixels are sorted according to their gradient values; this sorting can be pseudo-sorting.
[0066] Furthermore, region growing is performed at the pixel corresponding to the maximum gradient value. For example, the pixel with the maximum gradient value in the sorted list obtained from the pseudo-sorting is used as the seed point, and multiple valid pixels in the region containing that pixel are determined within its neighborhood. This neighborhood can be an 8-neighborhood. Within the 8-neighborhood, all pixels whose gradient direction error with the pixel corresponding to the maximum gradient value satisfies an error threshold can be found and considered as valid pixels in that region. The error threshold can be, for example, 22.5 degrees, or it can be determined according to actual needs.
[0067] Further, the pixel cloud density of the region containing the effective pixel corresponding to the maximum gradient value is calculated, and the pixel cloud density is compared with a density threshold to obtain the comparison result. The pixel cloud density reflects the distribution of pixels within the region. The density threshold can be determined according to actual needs, for example, it can be 0.7 or other values. If the comparison result shows that the pixel cloud density is greater than 0.7, it is considered a valid line detection, and the line detection result can be stored in the output matrix, thus completing the line detection. By performing line detection, the lines contained in the image to be processed can be obtained.
[0068] Next, continue to refer to Figure 2 As shown, in step S240, semantic information, three-dimensional feature points and line detection results are fused to obtain the initial pose of the three-dimensional detection box, and the camera pose and the initial pose are jointly optimized to construct a semantic map.
[0069] In this embodiment, the target detection result can be fused with the 3D map point cloud generated by visual SLAM to complete the data association and obtain associated data, which is used to determine whether objects need to be merged. For example, since there is a matching relationship between 3D and 2D feature points, and each 2D feature point contains semantic information, the 3D feature points can be associated with the semantic information based on the matching relationship between the 3D and 2D feature points, thereby determining the semantic information of the 3D feature points. For example, the category and detection box of each 3D feature point can be determined. Further, the 3D feature points can be merged according to the category in the semantic information of the 3D feature points to obtain associated data. That is, 3D feature points of the same category can be merged to obtain associated data. For example, 3D feature points of the category "table" can be merged, and 3D feature points of the category "chair" can be merged.
[0070] Furthermore, the associated data and line detection results can be fused to determine the initial pose of the 3D bounding box. The initial pose of the 3D bounding box can include yaw angle and scale, specifically the orientation, yaw angle, and position of the 3D bounding box, where the position can be in 3D coordinates. For a given object, there is one bounding box and line detection results. If the number of line detections within the bounding box is less than a preset number, it means that the bounding box and line detection results cannot be associated. After associating the bounding box and line detection results, rule data can be filtered out to determine the target line. The orientation and yaw angle of the 3D bounding box are obtained based on the angle of the target line, and the position of the 3D bounding box is obtained based on the position of the target line.
[0071] Building upon this, a semantic map can be constructed by jointly optimizing the camera pose and the initial pose of the 3D bounding box. This semantic map can be a lightweight object semantic map, capable of using any type of map as a carrier to map semantics. For example, a joint optimization function can be used to perform joint bundle adjustment (BA) optimization on the initial pose of the 3D bounding box and the camera pose. BA optimization can simultaneously adjust both the camera pose and feature point positions. BA typically constructs a least-squares problem, adjusting both the camera pose and feature point coordinates by minimizing the reprojection error.
[0072] Specifically, the joint optimization function can be determined by the error in yaw angle, the error in scale, and the error in camera pose. The joint optimization function can be calculated using formula (3):
[0073] f = arg min∑(e(θ) y Formula (3) is: )+e(s))+arg min∑e(p)
[0074] Where, e(θ) y ) refers to the yaw angle error, e(s) refers to the scale error, which is the error between the projection of the object's 3D detection box onto the 2D image and the detection of the line segment parallel to it, and e(p) refers to the camera pose error of the visual SLAM system.
[0075] In this embodiment, semantic information of objects in the image to be processed is acquired, and three-dimensional feature points are associated with the semantic information to obtain associated data. This associated data is then fused with the line detection results of the image to be processed to obtain the initial pose of the three-dimensional detection box. Finally, joint optimization is performed based on the camera pose and the initial pose of the three-dimensional detection box to generate an object semantic map. This reduces the cost of the acquisition equipment, improves the accuracy of object detection, and can improve the precise pose of the three-dimensional detection box of the object semantic map through different dimensions, thereby improving the accuracy and precision of the semantic map.
[0076] Next, in step S250, a semi-dense map corresponding to the image to be processed is constructed, and the semantic map and the semi-dense map are fused to obtain a semi-dense semantic map of the image to be processed.
[0077] In this embodiment of the disclosure, in addition to obtaining a semantic map, a semi-dense map can also be constructed for the image to be processed. For example, a semi-dense map corresponding to the image to be processed can be constructed by combining the image to be processed with the line segments corresponding to the image to be processed.
[0078] In some embodiments, firstly, straight line segments can be extracted from the image to be processed, and the line segments can be matched to obtain the correspondence between line segments in adjacent frames. Based on the correspondence, optimization is performed to obtain three-dimensional straight line segments. The channel information of the image to be processed is then fused into each three-dimensional straight line segment to obtain a semi-dense map. The straight line segment extraction can be used to extract straight line segments contained in the image to be processed. Straight line segment extraction can be accomplished by combining edge detection algorithms and straight line segment fitting methods. The edge detection algorithm can be of any type and is not specifically limited here. Next, the line segments in adjacent frames of the image to be processed can be matched to obtain the correspondence between the line segments in adjacent frames. This correspondence can be used to represent the association between line segments, and then three-dimensional straight line segments can be obtained based on this correspondence. Based on the three-dimensional straight line segments, the channel information of each channel of the image to be processed is combined to obtain a semi-dense map. The channel information of each channel can be RGB information.
[0079] Specifically, refer to Figure 7 As shown, constructing a semi-dense map mainly involves the following steps:
[0080] In step S710, the RGB image captured by the RGB camera is used as input.
[0081] In step S720, straight line segments are extracted from the RGB image. This extraction can be achieved by combining edge detection algorithms with straight line segment fitting methods. After obtaining the straight line segments, they can be verified using the Helmholtz method.
[0082] In step S730, the verified line segments are matched to obtain the correspondence between line segments in adjacent frame images, and BA optimization is performed to obtain the information of the three-dimensional line segments. The information of the three-dimensional line segments includes the coordinates of the line segments, etc.
[0083] In step S740, the channel information of the image to be processed is fused to each three-dimensional line segment. That is, the three-channel information of the RGB image is fused to each three-dimensional line segment. For example, the three-channel information of each three-dimensional line segment can be different. By introducing three-channel information to each three-dimensional line segment, each three-dimensional line segment can be visualized.
[0084] In step S750, a semi-dense map is created. The semi-dense map can be used for visualization, see reference... Figure 8 As shown, a semi-dense map can exist and be displayed in the form of a point cloud.
[0085] After obtaining the semi-dense map, it can be fused with the semantic map to obtain a semi-dense semantic map corresponding to the image to be processed. The semi-dense map exists in the form of a point cloud, where each point cloud contains pixels, but the semantic information of each pixel is unknown, i.e., the pixel category is unknown. The semi-dense map can be used for visualization. Based on this, for each 3D point in the point cloud, the semantic information of each 3D point can be retrieved from the corresponding txt file of the semantic map. This semantic information can then be used to fill the corresponding position in the semi-dense map to generate the semi-dense semantic map.
[0086] refer to Figure 9 The diagram shows a semi-dense semantic map obtained by fusing a semi-dense map and a semantic map. It can be seen that the accuracy of the semi-dense semantic map meets the requirements, and it provides the pose and 3D detection bounding box of the category of interest, which improves the visualization effect and enables the construction of a high-precision object semantic map for indoor scenes.
[0087] Figure 10 This schematically illustrates the overall flowchart of map construction. (Refer to...) Figure 10 As shown, it mainly includes several modules such as object detection, visual SLAM front-end, line detection, and back-end optimization. The main steps include:
[0088] In step S1000, an RGB image is acquired. For example, an RGB monocular camera can be used to acquire the image at a frequency of 30Hz.
[0089] In step S1002, the RGB image is input into the target detection model.
[0090] In step S1004, the detection bounding boxes and categories of objects in the RGB image are obtained to complete the target detection function.
[0091] In step S1006, the RGB image is converted to a grayscale image.
[0092] In step S1008, feature extraction and feature matching are performed on the grayscale image.
[0093] In step S1010, the three-dimensional feature points and camera pose are obtained.
[0094] In step S1012, line detection is performed on the grayscale image to obtain the line detection result.
[0095] In step S1014, associated data is obtained by performing data association based on three-dimensional feature points and semantic information.
[0096] In step S1016, the initial pose of the 3D detection box is obtained based on the associated data and the line detection results.
[0097] In step S1018, the initial pose and the camera pose are jointly optimized to obtain a semantic map.
[0098] Based on the above steps, intelligent mobile robots and other terminals, in an indoor home environment, an outdoor environment, or a smartphone scenario, fuse object detection, visual SLAM, object semantic mapping, and semi-dense mapping. Using an RGB monocular camera to acquire RGB image data streams, object detection is performed first to obtain the semantic information of objects in the image, i.e., object detection boxes and semantic categories. Next, visual SLAM and lightweight object mapping are performed on the images acquired by the RGB camera. Visual SLAM completes feature extraction, feature matching, triangulation to obtain 3D feature points, and estimates camera pose. Lightweight object mapping combines object detection information and visual SLAM information to complete image preprocessing, line detection, and data association of multi-source information. It also achieves joint BA optimization of camera pose, 3D object pose, and scale, ultimately realizing lightweight object semantic mapping. Finally, to visualize the mapping results, monocular semi-dense mapping is performed on the RGB images, thus completing the entire object semantic mapping construction. By fusing semi-dense maps with object semantic mapping, a semi-dense semantic map is obtained, which can be used to realize navigation and obstacle avoidance functions for mobile robots.
[0099] In this embodiment, data can be acquired using a low-cost monocular RGB camera, reducing the cost of the acquisition equipment. A low-power, low-performance target detection model is used to obtain semantic information of objects in the image to be processed, resulting in high-precision 3D object bounding box reconstruction. Combining the aforementioned semantic information with the line detection results of the image to be processed improves the accuracy of object detection. Combining the semantic information of the object with the camera pose from visual SLAM for combined BA optimization improves the accuracy of the initial pose of the 3D bounding box in the object semantic map. Furthermore, an accurate semantic map can be obtained based on the initial pose of the 3D bounding box and the camera pose. Simultaneously, a semi-dense map is constructed from the RGB image to improve visualization. By fusing the semi-dense map and the semantic map, a high-precision object semantic map for indoor scenes can be constructed. This provides the possibility for high-precision object semantic mapping in indoor home scenes and improves the accuracy and reliability of supporting navigation and obstacle avoidance functions of terminal devices.
[0100] Next, in this embodiment of the disclosure, a map building apparatus is also provided, with reference to... Figure 11 As shown, the map building device 1100 mainly includes a target detection module 1101, a pose estimation module 1102, a line detection module 1103, a semantic map building module 1104, and a map generation module 1105, wherein:
[0101] The target detection module 1101 is used to acquire the image to be processed and perform target detection on the image to be processed to obtain the semantic information of the objects in the image to be processed.
[0102] The pose estimation module 1102 is used to perform pose estimation on the image to be processed to obtain three-dimensional feature points and camera pose.
[0103] Line detection module 1103 is used to perform line detection on the image to be processed and obtain line detection results;
[0104] The semantic map construction module 1104 is used to fuse the semantic information, the three-dimensional feature points and the line detection results to obtain the initial pose of the three-dimensional detection box, and to jointly optimize the camera pose and the initial pose to construct a semantic map.
[0105] The map generation module 1105 is used to construct a semi-dense map corresponding to the image to be processed, and to fuse the semantic map and the semi-dense map to obtain a semi-dense semantic map of the image to be processed.
[0106] In one exemplary embodiment of this disclosure, the object detection module includes: extracting features at different scales from the image to be processed using an object detection model to obtain feature layers at multiple scales; and obtaining the semantic information based on the feature layers at multiple scales, wherein the semantic information includes a detection box and a category.
[0107] In an exemplary embodiment of this disclosure, the pose estimation module includes: a feature matching pair acquisition module, configured to convert the image to be processed into a grayscale image, and perform feature extraction and feature matching on the grayscale image to obtain a feature matching pair; and a motion mode pose estimation module, configured to perform triangulation based on the feature matching pair to obtain three-dimensional feature points, and combine the camera motion mode, the feature matching pair, and the three-dimensional feature points to perform pose estimation to obtain the camera pose.
[0108] In an exemplary embodiment of this disclosure, the motion mode pose estimation module includes: an initialization estimation module, configured to, in response to the camera motion mode being in an initialization state, perform pose estimation through the feature matching pair to obtain the camera pose; and a uniform motion estimation module, configured to, in response to the camera motion mode being in a uniform motion state, determine the camera pose of the current frame using the camera pose of the previous frame, and project three-dimensional feature points based on the camera pose of the current frame to obtain a reprojection error, and determine the camera pose of the current frame based on the comparison result of the reprojection error and an error threshold.
[0109] In an exemplary embodiment of this disclosure, the uniform velocity estimation module includes: a current frame pose preliminary determination module, used to multiply the camera pose of the previous frame by the result of the relative motion model between the previous frame and the current frame as the camera pose of the current frame; the relative motion model is used to represent the transformation relationship between the previous frame and the current frame.
[0110] In an exemplary embodiment of this disclosure, the uniform speed estimation module includes: a first determining module, configured to determine the camera pose of the previous frame as the camera pose of the current frame in response to the comparison result indicating that the reprojection error meets the error threshold; and a second determining module, configured to perform feature matching between the current frame and a reference frame to obtain the transformation relationship between the three-dimensional feature points of the reference frame and the two-dimensional feature points of the current frame in response to the comparison result indicating that the reprojection error does not meet the error threshold, and perform pose optimization based on the transformation relationship and the initial camera pose of the current frame to obtain the camera pose of the current frame.
[0111] In one exemplary embodiment of this disclosure, the line detection module includes: a gradient value determination module, configured to determine the gradient and direction corresponding to each pixel in the image to be processed, and sort all pixels according to the gradient value; an effective pixel determination module, configured to perform region growing at the pixel corresponding to the maximum gradient value, and determine multiple effective pixels in the region where the pixel is located within the neighborhood of the pixel; and a line detection execution module, configured to determine the pixel cloud density of the region where the effective pixel corresponding to the maximum gradient value is located, and perform line detection based on the comparison result of the pixel cloud density and a density threshold to obtain the line detection result.
[0112] In one exemplary embodiment of this disclosure, the semantic map construction module includes: a data association module, used to determine the semantic information of the three-dimensional feature points based on the matching relationship between the three-dimensional feature points and the two-dimensional feature points, and to merge the three-dimensional feature points according to the semantic information to obtain associated data; and an initial pose determination module, used to fuse the associated data and the line detection results to determine the initial pose of the three-dimensional detection box.
[0113] In one exemplary embodiment of this disclosure, the initial pose determination module includes: a data filtering module, used to filter out target straight lines when the number of straight line detection results within the detection box corresponding to the object is greater than a number threshold; and an orientation angle determination module, used to obtain the orientation, yaw angle, and position of the three-dimensional detection box based on the angle and position of the target straight line.
[0114] In an exemplary embodiment of this disclosure, the semantic map construction module includes: a joint optimization function determination module, used to perform logical processing on the yaw angle error in the initial pose, the scale error in the initial pose, and the camera pose error to determine a joint optimization function; and a joint optimization module, used to optimize based on the joint optimization function to obtain the semantic map.
[0115] In an exemplary embodiment of this disclosure, the map generation module includes: a line segment relationship extraction module, used to extract line segments from the image to be processed and match the line segments to obtain the correspondence between line segments in adjacent frame images; a three-dimensional line segment acquisition module, used to optimize based on the correspondence to obtain three-dimensional line segments; and a channel information fusion module, used to fuse the channel information of the image to be processed into each three-dimensional line segment to obtain a semi-dense map.
[0116] It should be noted that the specific details of each part of the above-mentioned map building device have been described in detail in the implementation of the map building method. For any undisclosed details, please refer to the implementation of the method section, and therefore will not be repeated here.
[0117] Exemplary embodiments of this disclosure also provide an electronic device. This electronic device may be the terminal described above. Generally, the electronic device may include a processor and a memory, the memory being used to store executable instructions of the processor, the processor being configured to perform the map construction method described above by executing the executable instructions.
[0118] The following is based on Figure 12 Taking the mobile terminal 1200 as an example, the construction of this electronic device will be described by way of example. Those skilled in the art will understand that, apart from components specifically designed for mobile purposes, Figure 12 The structure can also be applied to fixed types of equipment.
[0119] like Figure 12 As shown, the mobile terminal 1200 may specifically include: a processor 1201, a memory 1202, a bus 1203, a mobile communication module 1204, an antenna 1, a wireless communication module 1205, an antenna 2, a display screen 1206, a camera module 1207, an audio module 1208, a power module 1209, and a sensor module 1210.
[0120] Processor 1201 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, an encoder, a decoder, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). The method in this exemplary embodiment can be executed by an AP, GPU, or DSP. When the method involves neural network-related processing, it can be executed by an NPU. For example, the NPU can load neural network parameters and execute neural network-related algorithm instructions.
[0121] An encoder encodes (compresses) images or videos to reduce data size for easier storage or transmission. A decoder decodes (decompresses) the encoded data to restore the original image or video data. The mobile terminal 1200 can support one or more encoders and decoders, such as image formats like JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), and BMP (Bitmap), and video formats like MPEG (Moving Picture Experts Group) 1, MPEG10, H.1063, H.1064, and HEVC (High Efficiency Video Coding).
[0122] The processor 1201 can be connected to the memory 1202 or other components via the bus 1203.
[0123] The memory 1202 can be used to store computer executable program code, which includes instructions. The processor 1201 executes various functional applications and data processing of the mobile terminal 1200 by running the instructions stored in the memory 1202. The memory 1202 can also store application data, such as images, videos, and other files.
[0124] The communication function of mobile terminal 1200 can be implemented through mobile communication module 1204, antenna 1, wireless communication module 1205, antenna 2, modem processor, and baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Mobile communication module 1204 can provide 3G, 4G, 5G and other mobile communication solutions for mobile terminal 1200. Wireless communication module 1205 can provide wireless communication solutions such as wireless LAN, Bluetooth, and near-field communication for mobile terminal 1200.
[0125] The display screen 1206 is used to implement display functions, such as displaying user interfaces, images, and videos. The camera module 1207 is used to implement shooting functions, such as capturing images and videos, and may include a color temperature sensor array. The audio module 1208 is used to implement audio functions, such as playing audio and capturing voice. The power module 1209 is used to implement power management functions, such as charging the battery, supplying power to the device, and monitoring battery status. The sensor module 1210 may include one or more sensors to implement corresponding sensing and detection functions. For example, the sensor module 1210 may include an inertial sensor, which is used to detect the motion posture of the mobile terminal 1200 and output inertial sensing data.
[0126] It should be noted that the present disclosure also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist alone and not assembled into the electronic device.
[0127] Computer-readable storage media can be, for example—but not limited to—electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0128] A computer-readable storage medium can transmit, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof. The computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the electronic device to perform the methods described in the following embodiments.
[0129] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0130] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0131] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0132] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
[0133] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A map construction method, characterized in that, include: The image to be processed is acquired, and the target detection of the image to be processed is performed to obtain the semantic information of the objects in the image to be processed. The image to be processed is an RGB image; Pose estimation is performed on the image to be processed to obtain three-dimensional feature points and camera pose; Line detection is performed on the image to be processed to obtain the line detection result; The semantic information, the three-dimensional feature points, and the line detection results are fused to obtain the initial pose of the three-dimensional detection box. The camera pose and the initial pose are then jointly optimized to construct a semantic map. A semi-dense map corresponding to the image to be processed, existing in the form of a point cloud, is constructed. For each three-dimensional point in the point cloud, the semantic information of each three-dimensional point is found from the semantic map. The semantic information of each three-dimensional point is filled into the corresponding position of the semi-dense map to fuse the semantic map and the semi-dense map to obtain the semi-dense semantic map of the image to be processed. The construction of the semi-dense map corresponding to the image to be processed includes: Extract the line segments from the image to be processed, and match the line segments to obtain the correspondence between line segments in adjacent frame images; Based on the aforementioned correspondence, a three-dimensional line segment is obtained through optimization. The channel information of the image to be processed is fused to the three-dimensional straight line segments to obtain a semi-dense map.
2. The map construction method according to claim 1, characterized in that, The step of performing target detection on the image to be processed to obtain semantic information of objects in the image to be processed includes: The image to be processed is subjected to feature extraction at different scales using a target detection model to obtain feature layers at multiple scales. The semantic information is obtained based on the feature layers at the multiple scales; the semantic information includes the object's bounding box and category.
3. The map construction method according to claim 1, characterized in that, The step of performing pose estimation on the image to be processed to obtain three-dimensional feature points and camera pose includes: The image to be processed is converted into a grayscale image, and features are extracted and matched on the grayscale image to obtain feature matching pairs between the current frame and the previous frame. Based on the feature matching pairs, triangulation is performed to obtain three-dimensional feature points. Then, the camera pose is estimated by combining the camera motion mode, the feature matching pairs, and the three-dimensional feature points.
4. The map construction method according to claim 3, characterized in that, The step of combining the camera motion pattern, the feature matching pairs, and the 3D feature points to perform pose estimation to obtain the camera pose includes: In response to the camera motion mode being in the initial state, the camera pose is obtained by performing pose estimation through the feature matching pair; In response to the camera motion mode being uniform motion, the camera pose of the current frame is determined using the camera pose of the previous frame, and the reprojection error is obtained by projecting the three-dimensional feature points based on the camera pose of the current frame. The camera pose of the current frame is then determined based on the comparison between the reprojection error and the error threshold.
5. The map construction method according to claim 4, characterized in that, The step of determining the camera pose of the current frame using the camera pose of the previous frame includes: The camera pose of the previous frame is obtained by left-multiplying the camera pose of the previous frame by the relative motion model between the previous and current frames; the relative motion model is used to represent the transformation relationship between the previous and current frames.
6. The map construction method according to claim 4, characterized in that, Determining the camera pose of the current frame based on the comparison result of the reprojection error and the error threshold includes: If the comparison result indicates that the reprojection error meets the error threshold, the camera pose of the previous frame is determined as the camera pose of the current frame. In response to the comparison result that the reprojection error does not meet the error threshold, feature matching is performed between the current frame and the reference frame to obtain the transformation relationship between the three-dimensional feature points of the reference frame and the two-dimensional feature points of the current frame. Based on the transformation relationship and the initial camera pose of the current frame, pose optimization is performed to obtain the camera pose of the current frame.
7. The map construction method according to claim 1, characterized in that, The process of performing line detection on the image to be processed to obtain line detection results includes: Determine the gradient and direction of each pixel in the image to be processed, and sort all pixels according to their gradient values; Perform region growing at the pixel corresponding to the maximum gradient value, and determine multiple effective pixels in the region where the pixel is located within the neighborhood of the pixel. The pixel cloud density of the region where multiple valid pixels are located is determined, and line detection is performed based on the comparison result between the pixel cloud density and the density threshold to obtain the line detection result.
8. The map construction method according to claim 1, characterized in that, The process of fusing the semantic information, the three-dimensional feature points, and the line detection results to obtain the initial pose of the three-dimensional detection box includes: The semantic information of the three-dimensional feature points is determined based on the matching relationship between the three-dimensional feature points and the two-dimensional feature points. The three-dimensional feature points are then merged according to the semantic information to obtain associated data. The associated data and line detection results are fused to determine the initial pose of the 3D detection box.
9. The map construction method according to claim 8, characterized in that, The step of fusing the associated data and line detection results to determine the initial pose of the 3D detection box includes: If the number of line detection results within the detection box corresponding to the object is greater than the number threshold, the target line is filtered out. Based on the angle and position of the target line, the orientation, yaw angle, and position of the three-dimensional detection frame are obtained.
10. The map construction method according to claim 1, characterized in that, The step of jointly optimizing the camera pose and the initial pose to construct a semantic map includes: Logical processing is performed on the yaw angle error, the scale error, and the camera pose error in the initial pose to determine the joint optimization function; The semantic map is obtained by optimizing based on the joint optimization function.
11. A map building apparatus, characterized in that, include: The target detection module is used to acquire the image to be processed and perform target detection on the image to obtain the semantic information of the objects in the image to be processed. The image to be processed is an RGB image; The pose estimation module is used to estimate the pose of the image to be processed, and obtain the three-dimensional feature points and the camera pose. A line detection module is used to perform line detection on the image to be processed and obtain line detection results; The semantic map construction module is used to fuse the semantic information, the three-dimensional feature points and the line detection results to obtain the initial pose of the three-dimensional detection box, and to jointly optimize the camera pose and the initial pose to construct a semantic map. The map generation module is used to construct a semi-dense map in the form of point cloud corresponding to the image to be processed. For each three-dimensional point in the point cloud, the semantic information of each three-dimensional point is found from the semantic map, and the semantic information of each three-dimensional point is filled into the corresponding position of the semi-dense map to fuse the semantic map and the semi-dense map to obtain the semi-dense semantic map of the image to be processed. The construction of the semi-dense map corresponding to the image to be processed includes: Extract the line segments from the image to be processed, and match the line segments to obtain the correspondence between line segments in adjacent frame images; Based on the aforementioned correspondence, a three-dimensional line segment is obtained through optimization. The channel information of the image to be processed is fused to the three-dimensional straight line segments to obtain a semi-dense map.
12. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the map construction method according to any one of claims 1-10 by executing the executable instructions.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the map construction method according to any one of claims 1-10.