Map data processing method, device, storage medium and electronic device
By generating and optimizing maps, the problem of low efficiency and accuracy during fusion and update of three-dimensional maps in the prior art is solved, and efficient map data fusion suitable for multiple scenarios is achieved.
Patent Information
- Application Number
- CN202111520245.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-13
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-12-13
AI Technical Summary
In the prior art, when integrating and updating three-dimensional maps, the map data processing efficiency and accuracy are low, making it difficult to be applicable to the integration and updating of multiple scenarios.
By obtaining map data for the first scene and the second scene, generating corresponding maps and placing them in the same coordinate system, determining the matching relationship between the positioning data, optimizing the positioning data, and fusing two sets of map data to generate new map data.
It improves the efficiency and accuracy of fusion and update of three-dimensional maps, and is suitable for map data fusion in multiple scenarios, expanding the scope of application.
Smart Images

Figure CN114241039B_ABST
Abstract
Description
Background Art
[0002] With the rapid development of computer vision, three-dimensional maps are widely used in scenarios such as visual navigation, mobile robots, and autonomous driving. Among them, three-dimensional maps are usually obtained by capturing images of the scene and reconstructing them based on the captured images. In practical applications, visual positioning often occurs after image acquisition and map construction. However, over time, the scene may change, such as construction, posters, seasonal changes, weather changes, etc., which may cause the three-dimensional map information to lag, resulting in reduced positioning accuracy or positioning failure. Therefore, the three-dimensional map needs to be continuously updated to reduce the impact on visual positioning.
[0003] When performing 3D map fusion and update, the existing technology usually needs to re-collect scene images at a fixed position, determine new point cloud information, and replace the old point cloud information with the new point cloud information to achieve the update of the 3D map. However, this method has high requirements for the re-collected images and new point cloud information. For example, it is necessary to limit the positions of the two image collections to be exactly the same, and the area for updating the point cloud information cannot be too large. The efficiency and accuracy of map data processing are low, and it is difficult to apply to more 3D map fusion and update scenarios. Summary of the invention
[0004] The present disclosure provides a map data processing method, a map data processing device, a computer-readable storage medium and an electronic device, thereby improving, at least to a certain extent, the problem of low map data processing efficiency and accuracy when merging and updating three-dimensional maps in the prior art.
[0005] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by the practice of the present disclosure.
[0006] According to a first aspect of the present disclosure, a map data processing method is provided, comprising: obtaining first map data of a first scene, and second map data of a second scene associated with the first scene; the first map data comprising a set of first three-dimensional points, a plurality of first images, and first pose data corresponding to the first images, and the second map data comprising a set of second three-dimensional points, a plurality of second images, and second pose data corresponding to the second images; generating a first atlas according to the first pose data and a first matching relationship between the first images; generating a second atlas according to the second pose data and a second matching relationship between the second images; placing the first atlas and the second atlas in the same coordinate system, and determining a third matching relationship between the first pose data and the second pose data according to the positional relationship between the first atlas and the second atlas; optimizing at least one of the first pose data and the second pose data based on the third matching relationship, and fusing the first map data and the second map data according to the optimized first pose data and the second pose data to obtain third map data.
[0007] According to a second aspect of the present disclosure, a map data processing device is provided, comprising: a map acquisition module, for acquiring first map data of a first scene, and second map data of a second scene associated with the first scene; the first map data comprises a set of first three-dimensional points, a plurality of first images, and first pose data corresponding to the first images, and the second map data comprises a set of second three-dimensional points, a plurality of second images, and second pose data corresponding to the second images; a map generation module, for generating a first map based on the first pose data and a first matching relationship between the first images; and generating a second map based on the second pose data and a second matching relationship between the second images; a relationship determination module, for placing the first map and the second map in the same coordinate system, and determining a third matching relationship between the first pose data and the second pose data according to the positional relationship between the first map and the second map; and a pose optimization module, for optimizing at least one of the first pose data and the second pose data based on the third matching relationship, and fusing the first map data and the second map data according to the optimized first pose data and the second pose data to obtain third map data.
[0008] According to a third aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the map data processing method of the first aspect and possible implementation methods thereof are implemented.
[0009] According to a fourth aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor, wherein the processor is configured to execute the map data processing method of the first aspect and possible implementation thereof by executing the executable instructions.
[0010] The technical solution disclosed in this disclosure has the following beneficial effects:
[0011] Acquire first map data of a first scene and second map data of a second scene associated with the first scene; the first map data includes a set of first three-dimensional points, a plurality of first images and first pose data corresponding to the first images, and the second map data includes a set of second three-dimensional points, a plurality of second images and second pose data corresponding to the second images; generate a first atlas according to the first pose data and a first matching relationship between the first images; generate a second atlas according to the second pose data and a second matching relationship between the second images; place the first atlas and the second atlas in the same coordinate system, and determine a third matching relationship between the first pose data and the second pose data according to the positional relationship between the first atlas and the second atlas; optimize at least one of the first pose data and the second pose data based on the third matching relationship, and fuse the first map data and the second map data according to the optimized first pose data and the second pose data to obtain the third map data. On the one hand, this exemplary embodiment proposes a new map data processing method, which generates a first atlas through the first pose data and the first matching relationship in the first map data, and generates a second atlas through the second pose data and the second matching relationship in the second map data. From the perspective of topological structure, the third matching relationship between the first map data and the second map data is determined based on the first atlas and the second atlas, and then the pose data is optimized according to the third matching relationship to ensure the validity and reliability of the pose data before the map data is fused, eliminate interference and errors to a certain extent, and further ensure the accuracy of the third map data generated by the fusion of the first map data and the second map data; on the other hand, when the second scene is associated with the first scene, the map data fusion processing of the first map data of the first scene and the second map data of the second scene can be performed through this exemplary embodiment. Compared with the prior art, when performing map fusion, images with exactly the same acquisition positions are required, and the fusion area cannot be too large. This exemplary embodiment has no strict requirements on the acquisition positions of the first image and the second image, and can be applied to the fusion of the first map data and the second map data in various scenes, and has a wider range of applications.
[0012] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.
[0014] Figure 1 A structural diagram of an electronic device according to the exemplary embodiment is shown;
[0015] Figure 2 A flowchart showing a map data processing method in this exemplary embodiment is shown;
[0016] Figure 3 A schematic diagram showing a triangulation principle;
[0017] Figure 4 A schematic diagram showing a first spectrum and a second spectrum in this exemplary embodiment;
[0018] Figure 5 A sub-flow chart showing a map data processing method in this exemplary embodiment;
[0019] Figure 6 Another sub-flow chart of a map data processing method in this exemplary embodiment is shown;
[0020] Figure 7 Another sub-flow chart of a map data processing method in this exemplary embodiment is shown;
[0021] Figure 8 Another sub-flow chart of a map data processing method in this exemplary embodiment is shown;
[0022] Figure 9-12 A plurality of schematic diagrams showing traversal of the second graph in this exemplary embodiment are shown;
[0023] Figure 13-14 A plurality of schematic diagrams showing unmatched second nodes in this exemplary embodiment;
[0024] Fig.15 Another sub-flow chart of a map data processing method in this exemplary embodiment is shown;
[0025] Fig.16 A schematic diagram showing a third matching relationship in this exemplary embodiment;
[0026] Fig.17 Another sub-flow chart of a map data processing method in this exemplary embodiment is shown;
[0027] Fig.18 A schematic diagram of an application for map data processing in this exemplary embodiment is shown;
[0028] Fig.19 A flowchart of a map data processing method in this exemplary embodiment is shown;
[0029] Fig. 20 A structural block diagram of a map data processing device in this exemplary embodiment is shown. DETAILED DESCRIPTION
[0030] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as being limited to the examples set forth herein; on the contrary, these embodiments are provided so that the present disclosure will be more comprehensive and complete, and the concepts of the example embodiments are fully conveyed to those skilled in the art. The described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.
[0031] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0032] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the steps. For example, some steps may be decomposed, while some steps may be combined or partially combined, so the actual execution order may change according to the actual situation.
[0033] The exemplary embodiments of the present disclosure provide a map data processing method. Its application scenarios include but are not limited to: obtaining scene images before the mall environment changes, establishing an original scene map, collecting a segment of scene images that can be successfully located and scene images that cannot be successfully located after the mall environment changes, re-building the map, and obtaining a new scene map, by fusing the original scene map and the new scene map, updating the original scene of the mall to adapt to the scene change; or collecting images of different sub-scenes of the same target scene, building maps and fusing them separately, and generating map data of the overall target scene, wherein different sub-scenes have scene intersections, such as fusing the map data of scenes in different areas of the mall, and there are intersections between scenes in different areas, generating map data of the entire mall, etc.
[0034] The exemplary embodiment of the present disclosure provides an electronic device for implementing a map data processing method. It may be a terminal or a server in the cloud, including but not limited to a computer, a smart phone, a wearable device (such as augmented reality glasses), a robot, a drone, etc. Generally, the electronic device includes a processor and a memory, the memory is used to store executable instructions of the processor, and may also store application data, such as image data, video data, etc.; the processor is configured to execute the map data processing method by executing the executable instructions.
[0035] Below Figure 1 Taking the mobile terminal 100 in FIG. 1 as an example, the structure of the above electronic device is exemplarily described. It should be understood by those skilled in the art that, in addition to the components specifically used for mobile purposes, Figure 1 The construction in can also be applied to fixed type equipment.
[0036] like Figure 1 As shown, the mobile terminal 100 may specifically include: a processor 110, an internal memory 121, an external memory interface 122, a USB (Universal Serial Bus) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 171, a receiver 172, a microphone 173, an earphone interface 174, a sensor module 180, a display screen 190, a camera module 191, an indicator 192, a motor 193, a button 194 and a SIM (Subscriber Identification Module) card interface 195, etc.
[0037] The processor 110 may include one or more processing units. For example, the processor 110 may include an AP (Application Processor), a modem processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor and / or an NPU (Neural-Network Processing Unit), etc.
[0038] The encoder can encode (i.e. compress) the image or video data, for example, encode the captured scene image to form the corresponding code stream data to reduce the bandwidth occupied by data transmission; the decoder can decode (i.e. decompress) the code stream data of the image or video to restore the image or video data, for example, decode the code stream data of the scene image to obtain complete image data, so as to facilitate the execution of the map data processing method of this exemplary embodiment. The mobile terminal 100 can support one or more encoders and decoders. In this way, the mobile terminal 100 can process images or videos in a variety of encoding formats, such as: JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), BMP (Bitmap) and other image formats, MPEG (Moving Picture Experts Group) 1, MPEG2, H.263, H.264, HEVC (High Efficiency Video Coding) and other video formats.
[0039] In one implementation, the processor 110 may include one or more interfaces, and may be connected to other components of the mobile terminal 100 via different interfaces.
[0040] The internal memory 121 may be used to store computer executable program codes, which include instructions. The internal memory 121 may include volatile memory and non-volatile memory. The processor 110 executes various functional applications and data processing of the mobile terminal 100 by running the instructions stored in the internal memory 121.
[0041] The external memory interface 122 can be used to connect an external memory, such as a Micro SD card, to expand the storage capacity of the mobile terminal 100. The external memory communicates with the processor 110 through the external memory interface 122 to implement data storage functions, such as storing images, videos and other files.
[0042] The USB interface 130 is an interface that complies with USB standard specifications and can be used to connect a charger to charge the mobile terminal 100 , or to connect headphones or other electronic devices.
[0043] The charging management module 140 is used to receive charging input from a charger. While charging the battery 142, the charging management module 140 can also power the device through the power management module 141; the power management module 141 can also monitor the status of the battery.
[0044] The wireless communication function of the mobile terminal 100 can be implemented by antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modulation and demodulation processor and baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the mobile terminal 100. The wireless communication module 160 can provide wireless communication solutions including WLAN (Wireless Local Area Networks) (such as Wi-Fi (Wireless Fidelity) network), BT (Bluetooth), GNSS (Global Navigation Satellite System), FM (Frequency Modulation), NFC (Near Field Communication), IR (Infrared) and the like applied to the mobile terminal 100.
[0045] The mobile terminal 100 can implement a display function and display a user interface through the GPU, the display screen 190 and the AP, etc. For example, when the user turns on the shooting function, the mobile terminal 100 can display a shooting interface and a preview image on the display screen 190 .
[0046] The mobile terminal 100 can realize the shooting function through the ISP, camera module 191, encoder, decoder, GPU, display screen 190 and AP, etc. For example, the user can start the related services of mapping or visual positioning to trigger the shooting function. At this time, the camera module 191 can be used to capture scene images and perform positioning.
[0047] The mobile terminal 100 can implement audio functions through the audio module 170, the speaker 171, the receiver 172, the microphone 173, the earphone interface 174 and the AP.
[0048] The sensor module 180 may include a depth sensor 1801 , a pressure sensor 1802 , a gyroscope sensor 1803 , an air pressure sensor 1804 , etc., to implement corresponding sensing detection functions.
[0049] The indicator 192 may be an indicator light, which may be used to indicate the charging status, power change, message, missed call, notification, etc. The motor 193 may generate a vibration prompt, or may be used for touch vibration feedback, etc. The buttons 194 include a power button, a volume button, etc.
[0050] The mobile terminal 100 may support one or more SIM card interfaces 195 for connecting a SIM card to implement functions such as calls and mobile communications.
[0051] Figure 2 An exemplary process of a map data processing method is shown, including the following steps S210 to S240:
[0052] Step S210, obtaining first map data of a first scene and second map data of a second scene associated with the first scene; the first map data includes a set of first three-dimensional points, multiple first images and first pose data corresponding to the first images, and the second map data includes a set of second three-dimensional points, multiple second images and second pose data corresponding to the second images.
[0053] Among them, the first scene and the second scene can be any scene, such as a shopping mall, a street, etc. And the second scene has an association relationship with the first scene, and the second scene has an association relationship with the first scene, which means that the second scene has some of the same scenes as the first scene, that is, the intersection of the first scene and the second scene is not empty, for example, the first scene is a shopping mall before a store is renovated, and the second scene is a scene after the store is renovated, wherein the stores before and after the renovation have some of the same scenes, or the shopping mall scenes other than the renovated stores are the same, etc.; or the first scene is a shopping mall without scene arrangement before holding an event, and the second scene is a shopping mall with scene arrangement during or after holding an event, wherein the scene arrangement areas are partially the same; or a shopping mall includes store A, store B, and store C, the first scene is a scene including store A and store B, and the second scene is a scene including store B and store C, etc.
[0054] The first map data includes a set of first three-dimensional points, a first image, and first pose data corresponding to the first image, wherein the set of first three-dimensional points refers to three-dimensional point cloud data about the first scene, such as a three-dimensional point cloud map, and the first pose data refers to the pose data of the camera when each first image is collected, such as a rotation matrix or a translation matrix, etc., and different first images correspond to different first pose data; the second map data includes a set of second three-dimensional points, a second image, and second pose data corresponding to the second image, wherein the set of second three-dimensional points refers to three-dimensional point cloud data about the second scene, and the second pose data refers to the pose data of the camera when each second image is collected, and different second images correspond to different second pose data, and the second image may include an image that is successfully located in the first map data and an image that cannot be located in the first map data. In addition, the first map data may also include information about feature points in the first image, such as location information of the feature points or description information such as local descriptors or global descriptors, etc., and the second map data is similar. In this exemplary embodiment, multiple first images may be collected for the first scene, multiple second images may be collected for the second scene, and the first scene may be 3D reconstructed based on the first image to obtain the first map data, and the second scene may be 3D reconstructed based on the second image to obtain the second map data. For example, the terminal may turn on the video shooting function, collect multiple consecutive frames of the first image before the mall environment changes, and construct a 3D point cloud map of the scene before the mall environment changes, and collect the second image after the mall environment changes, that is, collect the supplementary image, and construct a 3D point cloud map of the scene after the mall environment changes. In addition, the first map data about the first scene and the first map data about the second scene may also be directly obtained from other terminals or map data sources.
[0055] In this exemplary embodiment, the construction of the first map data can adopt the SFM (Structure-From-Motion) algorithm, etc., through the collection and extraction of feature points in the image, image pose estimation, point cloud triangulation and global optimization and other processing processes, the two-dimensional first image is converted into three-dimensional information in the world coordinate system, so as to perform three-dimensional reconstruction processing on the first scene and obtain a three-dimensional map of the first scene, usually three-dimensional point cloud data. The construction of the second map data is similar to the construction of the first map data, and the construction process of the first map data is taken as an example for specific description.
[0056] The first map data obtained by three-dimensionally reconstructing the first scene may include the following process:
[0057] First, it is necessary to perform feature point extraction and feature point matching on the first image to determine an initial first image matching pair.
[0058] Among them, feature points refer to representative and highly recognizable points or areas in the image, such as corners and boundaries in the image. In the first image, gradients at different positions can be detected, and feature points can be extracted at positions with larger gradients. Generally, feature points need to be described after they are extracted, such as by using an array to describe the pixel distribution characteristics around the feature points, which is called the description information (or descriptor, descriptor) of the feature points. The description information of the feature points can be regarded as local description information of the first image. This exemplary embodiment can use algorithms such as FAST (Features From Accelerated Segment Test), BRIEF (Binary Robust Independent Elementary Features), ORB (Oriented FAST and Rotated BRIEF), SIFT (Scale-Invariant Feature Transform), SURF (Speeded UpRobust Features), SuperPoint (feature point detection and descriptor extraction based on self-supervised learning), and R2D2 (Reliable and Repeatable Detector and Descriptor) to extract feature points and describe the feature points.
[0059] In this exemplary embodiment, the total number of matching point pairs of feature points of each first image can be determined based on the matching relationship between each first image and other first images. For example, the first image A has a matching relationship with m other first images, and the first image A and the m first images constitute m pairs of first image matching pairs. Each pair of first image matching pairs has a corresponding matching point pair of feature points, and the number of matching point pairs is N. i , then we can determine the total number of matching point pairs between the first image A and the m first images as Among them, the matching point pair refers to the feature points in the two first images that are relatively similar, which is considered to be the projection of the same object point in the three-dimensional space of the first scene on the two first images. By calculating the similarity of the description information of the two feature points, it can be determined whether the two feature points are a matching point pair. In general, the description information of the feature points can be represented as a vector, and the similarity between the description information vectors of the two feature points is calculated, such as by measuring by Euclidean distance, cosine similarity, etc. If the similarity is higher than a preset similarity threshold, it is determined that the corresponding two feature points have a matching relationship to form a matching point pair. The similarity threshold is a standard for measuring whether two feature points are similar enough, and can be set according to experience or actual conditions. Then, all the first images can be sorted in descending order according to the total number of matching point pairs of each first image to determine the image sequence as the first first image in the first image matching pair, and then the image sequence of the first first image is traversed, and the image sequence of the second first image of the current first image is determined by sorting in descending order of the number of feature point matches between the current first image and other first images. Finally, the image sequence of the second first image is traversed, and if the current first first image and the current second first image satisfy the geometric constraint relationship, they are determined to be an initial first image matching pair.
[0060] After obtaining the initial first image matching pair, the image pose estimation can be performed on the initial first image matching pair, the relative pose between the two first images in the initial first image matching pair can be determined, and the matching point pairs of the first images can be triangulated to obtain three-dimensional point cloud data. Finally, the three-dimensional point cloud can be optimized and filtered through conditions such as epipolar constraints and viewing angle constraints to find the next frame of the first image to be reconstructed.
[0061] It should be noted that the first image of the next frame to be reconstructed needs to meet the following requirements: the first image of the frame has not been reconstructed before; the first image of the frame has enough visible points in the existing three-dimensional point cloud data; the number of attempts to reconstruct the first image of the frame does not exceed a preset threshold, such as 3 times. The first image that meets the above requirements will be pushed into the sequence of the first image of the next frame to be reconstructed. Finally, the two-dimensional-three-dimensional matching pair between the first image to be reconstructed and the existing three-dimensional point cloud data can be searched, and the PnP algorithm (Perspective-n-Point, a method for solving 3D-2D point pair motion) is used to match the n feature points in the current three-dimensional point cloud data with the n feature points in the first image to be reconstructed, and then the preliminary pose of the first image to be reconstructed is solved.
[0062] Furthermore, triangulation processing may be performed to determine the coordinates of the three-dimensional feature points.
[0063] After determining the preliminary pose of the first image, the newly added 2D-2D matching point pairs can be triangulated to generate new 3D points. Figure 3 Take the example to explain the triangulation principle. Figure 3 Three first images are shown, assuming that the homogeneous coordinates of the three-dimensional point P in the world coordinate system are X=[x, y, z, 1] T , correspondingly, the projection points in the two first images are p1 and p2, and their coordinates in their respective camera coordinate systems are The camera projection matrices corresponding to the two first images are P1 and P2, where P1 = [P 11 ,P 12 ,P 13 ] T , P2=[P 21 ,P 22 ,P 23 ] T , P 11 , P 12 , P 13 Corresponding to the 1st to 3rd rows of the projection matrix P1, P 21 , P 22 , P 23 Corresponding to the 1st to 3rd rows of the projection matrix P2, under ideal conditions, there are for By cross-multiplying it by itself on both sides, we get:
[0064]
[0065] Right now:
[0066]
[0067] You can get:
[0068]
[0069] Among them, formula (3) can be obtained by linear transformation of formula (1) and (2). Therefore, under the camera perspective corresponding to one first image, two constraints can be obtained. Combined with the camera perspective corresponding to another first image, it can be obtained that: AX = 0, where:
[0070]
[0071] A is the linear constraint matrix of the three-dimensional point P. For the above equation, when the number of viewing points is small and there are no outliers, the matrix can be directly decomposed to obtain the coordinates of the three-dimensional point P, for example, SVD (Singular Value Decomposition) decomposition can be used. When there are outliers, other methods such as RANSAC (Random sample consensus) method can also be used for estimation.
[0072] Finally, the 3D point cloud BA (Bundle Adjustment) is optimized.
[0073] After each frame of the first image is reconstructed, local BA optimization can be performed. When the number of reconstructed images increases by a fixed number, or the number of reconstructed 3D points increases by a fixed number, a global BA optimization can be performed on all previous reconstructions to obtain complete 3D point cloud data, i.e., the first map data.
[0074] Step S220, generating a first atlas according to the first pose data and the first matching relationship between the first images; generating a second atlas according to the second pose data and the second matching relationship between the second images.
[0075] Wherein, the first map data may include multiple first images, each of which may correspond to the first pose data. In this exemplary embodiment, after obtaining the first map data, the first atlas may be generated according to the first matching relationship between the first images in the multiple first images, and the first pose data corresponding to the first images. Wherein, the first matching relationship refers to the matching relationship between the first images, such as the matching relationship determined based on the global description information of the first images or the number of matching points, etc. The first atlas refers to a data structure used to reflect the association relationship between the first images corresponding to the first images through different first pose data. The second map data may include multiple second images, each of which may correspond to the second pose data. Similarly, the second atlas may be generated according to the second matching relationship between the second images, and the second pose data corresponding to the second images. Wherein, the second atlas refers to a data structure used to reflect the association relationship between the second images corresponding to the second images through different second pose data.
[0076] In an exemplary embodiment, the above step S220 may include:
[0077] A first graph is generated with the first pose data as the first node and the first matching relationship between the first images as the edge between the first nodes; a second graph is generated with the second pose data as the second node and the second matching relationship between the second images as the edge between the second nodes.
[0078] A graph refers to a mesh data structure, which consists of a non-empty vertex set V and a set E describing the relationship between vertices, wherein the data in the vertex set V is used to constitute different entities of the graph, i.e., the nodes of the graph, and the data in the relationship set E is used to describe the relationship between different nodes, i.e., an arc or edge from node to node in the graph. In this exemplary embodiment, after acquiring the first map data and the second map data, the terminal can use the first pose data as the first node and the first matching relationship between the first images as the edge between the first nodes to generate the first graph, and use the second pose data as the second node and the second matching relationship between the second images as the edge between the second nodes to generate the second graph. Among them, the first node is a node constituting the first graph, which is different from the second node, and its "first" and "second" do not refer to number or order.
[0079] Next, combine Figure 4 , the generation process of the first map and the second map is illustrated by example, Figure 4 As shown, white circles are first nodes, each of which represents the first pose data corresponding to a first image, that is, a camera pose; black circles are second nodes, each of which represents the second pose data corresponding to a second image. The first images with matching relationships are connected by dotted lines, for example, the first nodes represented by the first pose data corresponding to two first images whose similarity exceeds a preset similarity threshold can be connected by dotted lines; the second images with matching relationships are connected by solid lines. Based on this, Figure 4 The first spectrum M and the second spectrum N are shown.
[0080] In order to facilitate the subsequent optimization of the posture data, the coordinate transformation parameters of the first map data and the second map data can be determined first, and based on the coordinate transformation parameters, the first posture data in the first map data and the second posture data in the second map can be converted to the same coordinate system. Then, the first map is generated according to the converted first posture data and the first matching relationship, and the second map is generated according to the second posture data and the second matching relationship, so that the first map and the second map are placed in the same coordinate system.
[0081] In addition, after the first map and the second map are generated, the first map and the second map can be placed in the same coordinate system. Specifically, in an exemplary embodiment, Figure 5 As shown, the map data processing method may further include:
[0082] Step S510, obtaining a first 3D point subset from the first map data, and obtaining a second 3D point subset from the second map data, wherein the first 3D point subset and the second 3D point subset have a matching relationship;
[0083] Step S520, determining coordinate transformation parameters between the first map data and the second map data according to the three-dimensional distribution characteristics of the first three-dimensional point subset and the second three-dimensional point subset; the coordinate transformation parameters are used to place the first map and the second map in the same coordinate system.
[0084] The first map data includes a first 3D point subset, which is represented by p'={p1'...p i '…p n '}, the second map data includes a second 3D point subset, represented by q'={q1'...q i '…q n '}, the first 3D point subset has a matching relationship with the second 3D point subset, for example, p i ' and q i ' is a pair of matching points.
[0085] Since the 3D points in the first map data and the second map data are in different coordinate systems, the position representations of the 3D points are different, for example, the coordinates of the two 3D points in the matching point pair are different. However, the overall distribution characteristics of the 3D points in the 3D point subset are the same or similar, such as the density distribution, shape distribution, or distance distribution between 3D points, etc. Therefore, the present invention can determine the coordinate transformation parameters of the first map data and the second map data according to the 3D distribution characteristics of the first 3D point subset and the second 3D point subset, so as to place the first map and the second map in the same coordinate system through the coordinate transformation parameters.
[0086] In an exemplary embodiment, if Figure 6 As shown, the above step S510 may include:
[0087] Step S610, obtaining a 2D-2D matching point pair between the first image and the second image;
[0088] Step S620, determining a 3D-3D matching point pair between the first 3D point and the second 3D point according to the 2D-2D matching point pair;
[0089] Step S630: forming a first 3D point subset with the first 3D points in the 3D-3D matching point pairs, and forming a second 3D point subset with the second 3D points in the 3D-3D matching point pairs.
[0090] In this exemplary embodiment, since the first scene and the second scene have an intersection, there is an image in the first image that has a matching relationship with the second image, such as a relatively similar image. In this exemplary embodiment, the first image that matches the second image can be determined by image retrieval. Further, a 2D-2D matching point pair between the first image and the second image can be obtained.
[0091] In an exemplary embodiment, step S610 includes:
[0092] For at least one second image to be matched in the second image, at least one first image similar to it is searched in the first image, and feature points of the at least one first image and the second image to be matched are matched to obtain a two-dimensional-two-dimensional matching point pair.
[0093] That is, one or more second images can be selected from the second image as the second image to be matched, and searched among all the first images to determine the first image similar to the second image to be matched. The present disclosure does not limit the specific method for detecting image similarity. Exemplarily, the similarity between the global description information of the second image to be matched and the global description information of the first image can be detected. Global description information refers to information formed by extracting features from the entire image. For example, a CNN (such as NetVLAD, etc.) including a VLAD (Vector of Locally Aggregated Descriptors) layer can be used to extract global description information from the first image and the second image to be matched. For example, it can be a 4096-dimensional global description vector. Then, the similarity between the global description vector of the second image to be matched and the global description vector of the first image is calculated. For example, it can be measured by Euclidean distance, cosine similarity, etc. If the similarity is higher than a preset similarity threshold, it is determined that the second image to be matched matches the first image or meets the similarity condition; alternatively, the L2 norm can be used to calculate the similarity of the global description information of the second image to be matched and the first image. The smaller the L2 norm of the two global description information, the more similar the corresponding second image to be matched and the first image are, and so on.
[0094] Furthermore, feature point matching can be performed on at least one first image and the second image to be matched. The specific matching method can be to represent the description information of the feature points to be matched as vectors, calculate the similarity between the description information vectors of the two feature points, etc., thereby determining the two-dimensional-two-dimensional matching point pairs between the first image and the second image.
[0095] In this exemplary embodiment, the first map data includes a first three-dimensional point and a first image, and the second map data includes a second three-dimensional point and a second image, so each two-dimensional point in the two-dimensional image can have a corresponding three-dimensional point in the three-dimensional point set. Based on this, this exemplary embodiment can first convert a two-dimensional-two-dimensional matching point pair into a three-dimensional-two-dimensional matching point pair, or determine a three-dimensional-two-dimensional matching point pair by calculating the posture data of the second image, and then determine a three-dimensional-three-dimensional matching point pair between the first three-dimensional point and the second three-dimensional point based on the three-dimensional-two-dimensional matching point pair. Finally, based on the three-dimensional-three-dimensional matching point pair, a first three-dimensional point subset can be determined in the set of the first three-dimensional point, and a second three-dimensional subset can be determined in the set of the second three-dimensional point.
[0096] Specifically, in an exemplary embodiment, Figure 7 As shown, the above step S620 may include:
[0097] Step S710, determining the pose data of the second image in a first world coordinate system according to the two-dimensional-two-dimensional matching point pairs; the first world coordinate system is the world coordinate system of the first map data;
[0098] Step S720, using the pose data of the second image in the first world coordinate system, back-projecting the plurality of first three-dimensional points to obtain three-dimensional-two-dimensional matching point pairs between the plurality of first three-dimensional points and their corresponding two-dimensional points in the second image;
[0099] Step S730: Map the two-dimensional point in the three-dimensional-two-dimensional matching point pair to a second three-dimensional point to obtain a three-dimensional-three-dimensional matching point pair.
[0100] After determining the 2D-2D matching point pair, this exemplary embodiment can calculate the pose data of the second image in the first world coordinate system based on the 2D-2D matching points. For example, the PnP algorithm can be used to calculate the pose data of the second image, wherein the first world coordinate system is the coordinate system where the set of first 3D points in the first map data is located. Then, using the pose data of the second image in the first world coordinate system, multiple first 3D points can be back-projected to obtain 3D-2D matching point pairs between multiple first 3D points and their corresponding 2D points in the second image. Finally, based on the correspondence between each 2D point in the 3D-2D matching point pair and the 3D point in the set of second 3D points, the 2D points in the 3D-2D matching point pair can be mapped to the second 3D point to obtain a 3D-3D matching point pair.
[0101] In order to obtain more accurate posture data of the second image, a clustering method can also be used to cluster the multiple posture data obtained by calculation. In an exemplary embodiment, Figure 8As shown, in the above step S710, determining the pose data of the second image in the first world coordinate system according to the two-dimensional-two-dimensional matching point pairs may include the following steps:
[0102] Mapping the two-dimensional points of the first image in the two-dimensional-two-dimensional matching point pairs to first three-dimensional points, and determining a plurality of candidate poses of the second image in the first world coordinate system based on a matching relationship between the two-dimensional points of the second image and the first three-dimensional points;
[0103] Clustering is performed on the plurality of candidate poses, and pose data of the second image in the first world coordinate system is determined according to the best cluster.
[0104] In practical applications, N pose data of the second image in the first world coordinate system can usually be calculated, and these N pose data are candidate poses. Further, multiple candidate poses are clustered, and the pose data of the second image in the first world coordinate system is determined according to the optimal class, wherein the optimal class refers to the class with the most pose data in the clustering result. For example, the N pose data calculated are clustered, and the class with the largest number is taken as the optimal class, and then all candidate poses in the optimal class are weighted averaged, and the calculation result is used as the pose data of the second image in the first world coordinate system.
[0105] This exemplary embodiment can obtain the most likely camera pose by clustering the pose data. This exemplary embodiment does not specifically limit the clustering method, for example, DBSCAN (Density-Based Spatial Clustering of Applications with Noise, density-based clustering algorithm) clustering method can be used.
[0106] In an exemplary embodiment, step S620 may further include:
[0107] A first three-dimensional point corresponding to a two-dimensional point in the first image associated with the optimal class is obtained to obtain a plurality of first three-dimensional points for back-projection.
[0108] In the above-mentioned process of clustering multiple candidate postures, an optimal class with the largest number of classes can be determined. In order to obtain more accurate three-dimensional-two-dimensional matching point pairs, this exemplary embodiment can obtain first three-dimensional points corresponding to two-dimensional points in the first image associated with the optimal class, so as to obtain three-dimensional-two-dimensional matching point pairs between the multiple first three-dimensional points and their corresponding two-dimensional points in the second image when back-projecting the multiple first three-dimensional points.
[0109] In an exemplary embodiment, the coordinate transformation parameters may include rotation parameters and translation parameters; Figure 8As shown, the above step S520, determining the coordinate transformation parameters between the first map data and the second map data according to the three-dimensional distribution characteristics of the first three-dimensional point subset and the second three-dimensional point subset, may include the following steps:
[0110] Step S810, determining a first center point of the first three-dimensional point subset and a second center point of the second three-dimensional point subset;
[0111] Step S820, decentralized the first three-dimensional point subset based on the first center point, and decentralized the second three-dimensional point subset based on the second center point;
[0112] Step S830, determining a rotation parameter according to the covariance between the decentralized first three-dimensional point subset and the second three-dimensional point subset;
[0113] Step S840, determining a scale relationship according to the decentralized first three-dimensional point subset and the second three-dimensional point subset, and determining a translation parameter according to the scale relationship, the rotation parameter, the first center point and the second center point.
[0114] After obtaining the first three-dimensional point subset p'={p1'...p i '…p n '} and the second three-dimensional point subset q'={q1'…q i '…q n '}After that, the first center point of the first three-dimensional point subset and the second center point of the second three-dimensional point subset can be determined, wherein the center point can refer to the centroid point of the three-dimensional point subset, with P c Denotes the first center point of the first three-dimensional point subset, with q c represents the second center point of the second 3D point subset. Then, the first 3D point subset is decentralized based on the first center point. For example, the coordinates of each 3D point in the first 3D point subset are subtracted from the coordinates of the first center point, so that all 3D points in the first 3D point subset are translated toward the origin, and a new decentralized first 3D point subset p = {p1…p i …p n}; Similarly, the second three-dimensional point subset can be decentralized based on the second center point to obtain a new decentralized second three-dimensional point subset q = {q1…q i …q n}.
[0115] Then, the rotation parameter can be determined according to the covariance between the decentralized first three-dimensional point subset and the second three-dimensional point subset, which can specifically include first determining the relative scale of the decentralized first three-dimensional point subset and the second three-dimensional point subset as:
[0116]
[0117] Then construct the following covariance matrix:
[0118]
[0119] By decomposing the above covariance matrix, for example, by SVD, we can get: H = UΣV T , when R=VU T When RH=VΣV T , let A=VΣ 1 / 2 , then RH=AA T , therefore, it can be proved that R=VU T is the rotation matrix, which is the rotation parameter.
[0120] Finally, based on the determination of the scale relationship, rotation parameters, and the first center point and the second center point, the translation parameters, such as the translation vector t, can be determined by the following formula:
[0121] t=sRq c -p c (7)
[0122] Among them, s represents the scale relationship, such as the relative scale, R represents the rotation matrix, and q c and P c Represent the first center point and the second center point respectively.
[0123] Furthermore, the first map data and the second map data may be placed in the same coordinate system by using the determined rotation parameters and translation parameters, thereby achieving rigid registration of the first map data and the second map data.
[0124] Step S230, placing the first map and the second map in the same coordinate system, and determining a third matching relationship between the first pose data and the second pose data according to the positional relationship between the first map and the second map.
[0125] In this exemplary embodiment, the first matching relationship refers to the matching relationship between multiple first images. When the first pose data is used as the first node in the first atlas, it can be reflected in the edge connection relationship between the first nodes representing the first pose data; the second matching relationship refers to the matching relationship between multiple second images. When the second pose data is used as the second node in the second atlas, it can be reflected in the edge connection relationship between the second nodes representing the second pose data. The third matching relationship is different from the first matching relationship and the second matching relationship. It refers to the matching relationship between the first image and the second image, that is, the matching relationship between the first node in the first atlas and the second node in the second atlas, which can be determined by the positional relationship between the first atlas and the second atlas. Specifically, the first atlas and the second atlas can be placed in the same coordinate system, and the third matching relationship can be determined by calculating the matching degree between the first node and the second node.
[0126] In an exemplary embodiment, in the above step S230, determining the third matching relationship between the first pose data and the second pose data according to the positional relationship between the first atlas and the second atlas may include:
[0127] The second graph is traversed to obtain a third matching relationship according to the matching relationship between each second node and the first nodes within a preset range around it.
[0128] This exemplary embodiment can traverse each second node in the second graph to determine a third matching relationship based on the matching relationship between each second node and the first nodes within a preset range around it. Specifically, a second node can be initialized in the second graph, for example, a second node can be randomly determined, or a second node can be determined according to a preset order rule, etc. Starting from the second node, the first node that best matches the second node is found in the first graph, and other second nodes are traversed in sequence.
[0129] In an exemplary embodiment, traversing the second graph may include:
[0130] Obtaining an initial matching node pair between the first graph and the second graph; the initial matching node pair includes a first node and a second node;
[0131] The second graph is traversed with the second node in the initial matching node pair as the root node.
[0132] The initial matching node pair refers to the first pair of the first node and the second node having a matching relationship after the matching begins. Fig. 9 As shown in the figure, the first node M1 in the first graph M has a matching relationship with the second node N1 in the second graph N, and the first node M1 and the second node N1 are the initial matching node pair. Furthermore, the second node in the initial matching node pair can be used as the root node to traverse the second graph, that is, to search for other first nodes with matching relationships in the first graph. The root node refers to the starting node of traversing the second graph, that is, in the second graph, when searching upward from the second node as the root node, there is no parent node with a matching relationship with the first node. Figure 10-12 As shown, the traversal starts with the second node N1 as the root node, and the first node M2 and the second node N2, the first node M3 and the second node N3, the first node M4 and the second node N4, and the first node M5 and the second node N5 having a matching relationship are obtained.
[0133] In an exemplary embodiment, the above-mentioned obtaining initial matching node pairs between the first graph and the second graph may include:
[0134] Determine a second reference image in the second image, and search the first reference image in the first image for the first reference image with the highest matching degree with the second reference image;
[0135] A first node corresponding to the first reference image and a second node corresponding to the second reference image form an initial matching node pair.
[0136] Among them, the second reference image may refer to the second image corresponding to the first second node that performs matching calculation when starting to traverse the second graph. This exemplary embodiment may perform a global search in the first image to search for the first reference image that has the highest matching degree with the second reference image. Specifically, the first reference image may be determined by calculating the similarity between the second reference image and the first image. When the similarity exceeds a preset similarity threshold, it is considered that the first image matches the second reference image, and the first image is determined to be the first reference image. Furthermore, the first node corresponding to the first reference image and the second node corresponding to the second reference image may form an initial matching node pair, such as Fig. 9 The first node M1 and the second node N1 shown form an initial matching node pair.
[0137] In this exemplary embodiment, Figure 9-12 As shown in the dotted ellipse area. If matching nodes can be found in this area, it means that the initial node matching is accurate at the topological structure level. If matching node pairs cannot be found, it means that the initial matching confidence is low, and it is likely that noise or similar scenes have caused an incorrect match. In this case, the initial first node or second node can be reselected.
[0138] When the second node in the initial matching node pair is used as the root node and the second graph is traversed, if the child node of the root node does not find the matching first node within the search range, such as Fig.13 The second node N6 in the left rectangular dotted area shows that no matching first node is found in the area. At this time, the second node N6 is traversed and can be marked as v, but no match is generated. When there is a first node in the search area, but no matching relationship is generated, such as Fig.14 As shown in the second node N7 and the first node M7 in the right rectangular dotted area, at this time, there is also no matching information, marked as v. After all the second nodes in the second graph are traversed, new edges will be generated between the images with matching relationships, that is, new edges will be generated between the first node and the second node with matching relationships, for example, a new edge will be generated between the first node M1 and the second node N1.
[0139] It should be noted that in this exemplary embodiment, in addition to traversing the second graph and obtaining the third matching relationship based on the matching relationship between each second node and the first nodes within a preset range around it, the first graph can also be traversed to determine the third matching relationship based on the matching relationship between each first node and the second nodes within a preset range around it. The method flow is similar to the above-mentioned traversal process and will not be repeated here.
[0140] Step S240: Based on the third matching relationship, at least one of the first pose data and the second pose data is optimized, and the first map data and the second map data are fused according to the optimized first pose data and the second pose data to obtain third map data.
[0141] Considering that the first map data can be obtained by three-dimensional reconstruction based on the first image, and the second map data can also be obtained by three-dimensional reconstruction based on the second image, there is a process of converting two-dimensional information into three-dimensional information. If the first atlas and the second atlas are directly placed in the same coordinate system, a certain error may be generated when converting three-dimensional information. Therefore, this exemplary embodiment can optimize at least one of the first pose data and the second pose data based on the third matching relationship, and fuse the first map data and the second map data based on the optimized first pose data and the second pose data to obtain the third map data. The third map data refers to the map data after the fusion of the first map data and the second map data, which can be three-dimensional point cloud data.
[0142] In an exemplary embodiment, if Fig.15 As shown, in the above step S240, based on the third matching relationship, optimizing at least one of the first pose data and the second pose data may include the following steps:
[0143] Step S1510, determining a first pose transformation parameter corresponding to the third matching relationship according to the first pose data corresponding to the first node in the third matching relationship and the second pose data corresponding to the second node;
[0144] Step S1520, determining a second posture transformation parameter corresponding to the third matching relationship according to the edge between the first node and the second node in the third matching relationship;
[0145] Step S1530: Optimize at least one of the first pose data and the second pose data based on the difference between the first pose transformation parameter and the second pose transformation parameter.
[0146] In this exemplary embodiment, the third matching relationship can be obtained by traversing the second graph, for example, Figure 9-12After the traversal process of the second graph, a new edge will be generated between the first node and the second node that generate the matching relationship, and the edge can represent the third matching relationship.
[0147] like Fig.16 As shown, the nodes included in the region 1610 and the region 1620 are all first nodes or second nodes having a matching relationship after traversal, wherein the node included in the region 1610 is the first node, and the node included in the region 1620 is the second node. The first nodes having a matching relationship have a connection relationship, which is a first matching relationship, the second nodes having a matching relationship have a connection relationship, which is a second matching relationship, and a new connection relationship is generated between the first nodes and the second nodes having a matching relationship, which is a third matching relationship.
[0148] Then, the first pose transformation parameter corresponding to the third matching relationship can be determined according to the first pose data corresponding to the first node in the third matching relationship and the second pose data corresponding to the second node. Specifically, the first pose transformation parameter corresponding to the third matching relationship can be determined by X i Indicates the first position data corresponding to the first node, through X j Indicates the second pose data corresponding to the second node, through X i -1 X j Characterize the first pose transformation parameter, which can be the pose transformation parameter obtained by summing all relationships. According to the edge between the first node and the second node in the third matching relationship, determine the second pose transformation parameter corresponding to the third matching relationship, and pass T ij Finally, based on the difference between the first pose transformation parameter and the second pose transformation parameter, at least one of the first pose data and the second pose data can be optimized, for example, by making the difference between the two approach 0, so as to optimize the first pose data or the second pose data.
[0149] In an exemplary embodiment, if Fig.17 As shown, the map data processing method may further include:
[0150] Step S1710, determining the confidence of the third matching relationship according to the number of matching inliers between the first image corresponding to the first node and the second image corresponding to the second node in the third matching relationship;
[0151] Then step S1530 may include:
[0152] Step S1720: Optimize at least one of the first pose data and the second pose data based on the difference between the first pose transformation parameter and the second pose transformation parameter and the confidence of the third matching relationship.
[0153] This exemplary embodiment can also optimize at least one of the first pose data and the second pose data in combination with the confidence of the third matching relationship. Specifically, the optimization equation can be constructed by the following formula:
[0154]
[0155] Among them, X i Indicates the first pose data corresponding to the first node, X j represents the second pose data corresponding to the second node, T ij represents the second pose transformation parameter calculated by the third matching relationship, that is, Fig.16 The newly generated edge between the first node and the second node, Ω ij Represents the second pose transformation parameter T ij The confidence level, in this exemplary embodiment, can be determined based on the number of matching inner points between the first image corresponding to the first node and the second image corresponding to the second node in the third matching relationship. The higher the number of inner points, the greater the confidence level, and the lower the number of inner points, the smaller the confidence level. The above formula (8) can be used to achieve joint optimization of the first pose data or the second pose data or the first pose data and the second pose data, eliminating the error caused by directly placing the first atlas and the second atlas in the same coordinate system.
[0156] In summary, in this exemplary embodiment, first map data of a first scene and second map data of a second scene associated with the first scene are obtained; the first map data includes a set of first three-dimensional points, multiple first images and first pose data corresponding to the first images, and the second map data includes a set of second three-dimensional points, multiple second images and second pose data corresponding to the second images; a first atlas is generated based on the first pose data and a first matching relationship between the first images; a second atlas is generated based on the second pose data and a second matching relationship between the second images; the first atlas and the second atlas are placed in the same coordinate system, and the third matching relationship between the first pose data and the second pose data is determined based on the positional relationship between the first atlas and the second atlas; based on the third matching relationship, at least one of the first pose data and the second pose data is optimized, and the first map data and the second map data are fused based on the optimized first pose data and the second pose data to obtain the third map data. On the one hand, this exemplary embodiment proposes a new map data processing method, which generates a first atlas through the first pose data and the first matching relationship in the first map data, and generates a second atlas through the second pose data and the second matching relationship in the second map data. From the perspective of topological structure, the third matching relationship between the first map data and the second map data is determined based on the first atlas and the second atlas, and then the pose data is optimized according to the third matching relationship to ensure the validity and reliability of the pose data before the map data is fused, eliminate interference and errors to a certain extent, and further ensure the accuracy of the third map data generated by the fusion of the first map data and the second map data; on the other hand, when the second scene is associated with the first scene, the map data fusion processing of the first map data of the first scene and the second map data of the second scene can be performed through this exemplary embodiment. Compared with the prior art, when performing map fusion, images with exactly the same acquisition positions are required, and the fusion area cannot be too large. This exemplary embodiment has no strict requirements on the acquisition positions of the first image and the second image, and can be applied to the fusion of the first map data and the second map data in various scenes, and has a wider range of applications.
[0157] In an exemplary embodiment, the map data processing method may further include:
[0158] Acquire first pose data and second pose data having a similar relationship, and form a first image corresponding to the first pose data and a second image corresponding to the second pose data into an image pair to be compared;
[0159] In response to the similarity between the two images in the image pair to be compared being less than a preset threshold, the image in the image pair to be compared that is shot later and the first 3D point or the second 3D point corresponding to the image are deleted from the third map data.
[0160] Through the above steps S210 to S240, the first map data and the second map data can be fused to obtain the third map data. However, the first node and some three-dimensional points in the first map data are still retained in the third map data. If they are not removed, they will become larger and larger as the number of fusions increases, and contain too much interference information, which will affect the efficiency and accuracy of the positioning algorithm. Therefore, this exemplary embodiment can generate more accurate third map data by deleting the interfering first three-dimensional points or second three-dimensional points.
[0161] Specifically, the confidence of the first posture data or the second posture data can be determined according to the acquisition time of the first image or the second image. The newer the acquisition time of the posture data in the map data, the higher the confidence is set, and the older the acquisition time of the posture data, the lower the confidence is set. The first posture data is the first node, and the second posture data is the second node.
[0162] By acquiring first pose data and second pose data having a similar relationship, a first image corresponding to the first pose data and a second image corresponding to the second pose data are formed into an image pair to be compared, and the similarity of the two images in the image pair to be compared is compared. When the similarity is less than a preset threshold, the image with a later shooting time in the image pair to be compared and the first three-dimensional point or the second three-dimensional point corresponding to the image can be deleted from the third map data.
[0163] Comparing the similarity between the second posture data in the second map data and the first posture data in the first map data is to compare the similarity between the two images in the image pair. Specifically, the similarity comparison can be performed directly by extracting the global description information of the first image corresponding to the first posture data, or the global description information of the second image corresponding to the second posture data. If the similarity between the two images in the image pair to be compared meets a certain threshold, it means that the scene changes in the two images to be compared are small and do not need to be deleted. If the similarity between the two images in the image pair to be compared is weak, the three-dimensional points of the image with a closer shooting time can be retained according to the confidence of the posture data, and then the posture data with the lowest confidence and its corresponding three-dimensional points can be deleted. After the deletion is completed, a BA optimization can be performed to obtain an updated point cloud map.
[0164] like Fig.18 As shown, Fig.18 (a) is the map data of the outdoor and a small part of the indoor of a store, which can be regarded as the first map data. According to the needs, the indoor map data is re-collected, which can be regarded as the second map, such as Fig.18 As shown in (b), the two map data are successfully merged and updated into new map data using the map data processing method in this exemplary embodiment, as shown in FIG. Fig.18(c) shows that the updated map data preserves the indoor and outdoor information, and removes the indoor poster change area, and updates to the latest map.
[0165] Fig.19 A flow chart of a map data processing method in this exemplary embodiment is shown, which may specifically include four parts, namely: a map data reconstruction module 1910, a point cloud registration module 1920, a joint pose optimization module 1930, and a map update module 1940.
[0166] Among them, the map data reconstruction module 1910 is used to perform three-dimensional reconstruction based on the collected images to generate three-dimensional point cloud data, and can specifically include a feature point extraction and matching unit 1911, an image pose estimation unit 1912, a triangulation unit 1913 and a BA optimization unit 1914.
[0167] The point cloud registration module 1920 is used to determine the coordinate transformation parameters so as to place the first map data and the second map data in the same coordinate system, or to place the first atlas and the second atlas in the same coordinate system. Specifically, it may include a visual positioning unit 1921, a 3D-3D point pair matching unit 1922, and a rigid transformation unit 1923. The rigid transformation unit 1923 is a unit that performs coordinate system conversion through coordinate transformation parameters.
[0168] The posture joint optimization module 1930 is used to optimize at least one of the first posture data and the second posture data, and may specifically include a topology search unit 1931, used to traverse the second graph to obtain a third matching relationship based on the matching relationship between each second node and the first node within a preset range around it, a posture transformation parameter determination unit 1932, and a joint optimization unit 1933.
[0169] The map update module 1940 is used to update the fused third map data to remove interfering three-dimensional points, and may specifically include: a time confidence determination unit 1941, an image similarity calculation unit 1942 and a three-dimensional point cloud removal unit 1943.
[0170] The exemplary embodiment of the present disclosure also provides a map data processing device. Fig. 20As shown, the map data processing device 2000 may include: a map acquisition module 2010, used to acquire first map data of a first scene, and second map data of a second scene associated with the first scene; the first map data includes a set of first three-dimensional points, a plurality of first images and first pose data corresponding to the first images, and the second map data includes a set of second three-dimensional points, a plurality of second images and second pose data corresponding to the second images; a map generation module 2020, used to generate a first map based on the first pose data and a first matching relationship between the first images; and to generate a second map based on the second pose data and a second matching relationship between the second images; a relationship determination module 2030, used to place the first map and the second map in the same coordinate system, and determine a third matching relationship between the first pose data and the second pose data based on the positional relationship between the first map and the second map; a pose optimization module 2040, used to optimize at least one of the first pose data and the second pose data based on the third matching relationship, and fuse the first map data and the second map data according to the optimized first pose data and the second pose data to obtain the third map data.
[0171] In an exemplary embodiment, the atlas generation module includes: a atlas generation unit, used to generate a first atlas with the first pose data as the first node and the first matching relationship between the first images as the edge between the first nodes; and to generate a second atlas with the second pose data as the second node and the second matching relationship between the second images as the edge between the second nodes.
[0172] In an exemplary embodiment, the relationship determination module includes: a traversal unit, configured to traverse the second graph to obtain a third matching relationship based on a matching relationship between each second node and a first node within a preset range around it.
[0173] In an exemplary embodiment, the traversal unit includes: a node pair acquisition subunit, used to obtain an initial matching node pair between a first graph and a second graph; the initial matching node pair includes a first node and a second node; and a graph traversal subunit, used to traverse the second graph with the second node in the initial matching node pair as a root node.
[0174] In an exemplary embodiment, the node pair acquisition subunit includes: a reference image search subunit, used to determine the second reference image in the second image, and to search the first reference image with the highest matching degree with the second reference image in the first image; and a node pair formation subunit, used to form an initial matching node pair with the first node corresponding to the first reference image and the second node corresponding to the second reference image.
[0175] In an exemplary embodiment, the posture optimization module includes: a first parameter determination unit, used to determine the first posture transformation parameter corresponding to the third matching relationship based on the first posture data corresponding to the first node in the third matching relationship and the second posture data corresponding to the second node; a second parameter determination unit, used to determine the second posture transformation parameter corresponding to the third matching relationship based on the edge between the first node and the second node in the third matching relationship; a posture optimization unit, used to optimize at least one of the first posture data and the second posture data based on the difference between the first posture transformation parameter and the second posture transformation parameter.
[0176] In an exemplary embodiment, the map data processing device may also include: a confidence determination module, used to determine the confidence of the third matching relationship based on the number of matching internal points between the first image corresponding to the first node and the second image corresponding to the second node in the third matching relationship; a posture optimization module, used to optimize at least one of the first posture data and the second posture data based on the difference between the first posture transformation parameters and the second posture transformation parameters, and the confidence of the third matching relationship.
[0177] In an exemplary embodiment, the map data processing device may also include: a subset acquisition module, used to acquire a first three-dimensional point subset from the first map data, and to acquire a second three-dimensional point subset from the second map data, wherein the first three-dimensional point subset and the second three-dimensional point subset have a matching relationship; a map coordinate transformation module, used to determine the coordinate transformation parameters between the first map data and the second map data according to the three-dimensional distribution characteristics of the first three-dimensional point subset and the second three-dimensional point subset; the coordinate transformation parameters are used to place the first map and the second map in the same coordinate system.
[0178] In an exemplary embodiment, the subset acquisition module includes: a 2D-2D matching point pair acquisition unit, used to acquire 2D-2D matching point pairs between a first image and a second image; a 3D-3D matching point pair determination unit, used to determine a 3D-3D matching point pair between a first 3D point and a second 3D point based on the 2D-2D matching point pairs; and a subset travel unit, used to form a first 3D point subset with the first 3D point in the 3D-3D matching point pair, and to form a second 3D point subset with the second 3D point in the 3D-3D matching point pair.
[0179] In an exemplary embodiment, a 3D-3D matching point pair determination unit includes: a posture data determination subunit, used to determine the posture data of the second image in the first world coordinate system based on the 2D-2D matching point pair; the first world coordinate system is the world coordinate system of the first map data; a back-projection subunit, used to use the posture data of the second image in the first world coordinate system to back-project multiple first 3D points to obtain 3D-2D matching point pairs between the multiple first 3D points and their corresponding 2D points in the second image; a mapping subunit, used to map the 2D points in the 3D-2D matching point pairs to second 3D points to obtain 3D-3D matching point pairs.
[0180] In an exemplary embodiment, a pose data determination subunit includes: a candidate pose determination subunit, which is used to map the two-dimensional points of the first image in the two-dimensional-two-dimensional matching point pairs to first three-dimensional points, and determine multiple candidate poses of the second image in the first world coordinate system based on the matching relationship between the two-dimensional points of the second image and the first three-dimensional points; a candidate pose clustering subunit, which is used to cluster multiple candidate poses and determine the pose data of the second image in the first world coordinate system based on the optimal cluster.
[0181] In an exemplary embodiment, a 3D-3D matching point pair determination unit includes: a first 3D point acquisition subunit, used to obtain a first 3D point corresponding to a 2D point in a first image associated with an optimal class, so as to obtain a plurality of first 3D points for back projection.
[0182] In an exemplary embodiment, a two-dimensional-two-dimensional matching point pair acquisition unit includes: an image search subunit, which is used to search for at least one first image similar to at least one second image to be matched in the second image, and perform feature point matching between the at least one first image and the second image to be matched to obtain a two-dimensional-two-dimensional matching point pair.
[0183] In an exemplary embodiment, the coordinate transformation parameters include rotation parameters and translation parameters; the atlas coordinate transformation module includes: a center point determination unit, used to determine the first center point of the first three-dimensional point subset and the second center point of the second three-dimensional point subset; a decentering unit, used to decenter the first three-dimensional point subset based on the first center point, and to decenter the second three-dimensional point subset based on the second center point; a rotation parameter determination unit, used to determine the rotation parameters according to the covariance between the decentralized first three-dimensional point subset and the second three-dimensional point subset; a translation parameter determination unit, used to determine the scale relationship according to the decentralized first three-dimensional point subset and the second three-dimensional point subset, and determine the translation parameters according to the scale relationship, the rotation parameters, the first center point and the second center point.
[0184] In an exemplary embodiment, the map data processing device may also include: a module for determining an image pair to be compared, used to obtain first pose data and second pose data having a similar relationship, and form an image pair to be compared with a first image corresponding to the first pose data and a second image corresponding to the second pose data; a three-dimensional point deletion module, used to delete an image with a later shooting time in the image pair to be compared and the first three-dimensional point or the second three-dimensional point corresponding to the image from the third map data in response to the similarity between the two images in the image pair to be compared being less than a preset threshold.
[0185] The specific details of each part of the above device have been described in detail in the implementation method part, so they will not be repeated here.
[0186] The exemplary embodiments of the present disclosure further provide a computer-readable storage medium, which can be implemented in the form of a program product, including program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above “Exemplary Method” section of this specification, for example, to execute Figure 2 , Figure 5 , Figure 6 , Figure 7 , Figure 8 , Fig.15 or Fig.17 The program product may be in the form of a portable compact disk read-only memory (CD-ROM) and include program code, and may be run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto, and in this document, a readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, apparatus, or device.
[0187] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory, a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0188] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, in which readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Readable signal media may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0189] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.
[0190] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0191] It will be appreciated by those skilled in the art that various aspects of the present disclosure may be implemented as systems, methods or program products. Therefore, various aspects of the present disclosure may be specifically implemented in the following forms, namely: complete hardware implementation, complete software implementation (including firmware, microcode, etc.), or implementations combining hardware and software aspects, which may be collectively referred to herein as "circuit", "module" or "system". Those skilled in the art will readily think of other implementations of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptive changes of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The specification and implementation are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims.
[0192] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A map data processing method, characterized in that: include: Acquire first map data of a first scene, and second map data of a second scene associated with the first scene; The first map data includes a set of first three-dimensional points, a plurality of first images, and first pose data corresponding to the first images, and the second map data includes a set of second three-dimensional points, a plurality of second images, and second pose data corresponding to the second images; Generate a first atlas according to the first posture data and a first matching relationship between the first images; Generate a second atlas according to the second posture data and a second matching relationship between the second images; Placing the first map and the second map in the same coordinate system, and determining a third matching relationship between the first pose data and the second pose data according to a positional relationship between the first map and the second map; Based on the third matching relationship, at least one of the first posture data and the second posture data is optimized, and the first map data and the second map data are fused according to the optimized first posture data and the second posture data to obtain third map data.
2. The method according to claim 1, characterized in that generating a first atlas according to the first posture data and a first matching relationship between the first images; Generating a second atlas according to the second posture data and the second matching relationship between the second images includes: Generate a first graph with the first posture data as a first node and the first matching relationship between the first images as an edge between the first nodes; A second graph is generated with the second posture data as a second node and the second matching relationship between the second images as an edge between the second nodes.
3. The method according to claim 2, characterized in that The determining a third matching relationship between the first posture data and the second posture data according to the positional relationship between the first atlas and the second atlas includes: The second graph is traversed to obtain the third matching relationship according to the matching relationship between each second node and the first nodes within a preset range around it.
4. The method according to claim 3, characterized in that: The traversing the second graph includes: Obtaining an initial matching node pair between the first graph and the second graph; the initial matching node pair includes a first node and a second node; The second graph is traversed with the second node in the initial matching node pair as the root node.
5. The method according to claim 4, characterized in that The obtaining of initial matching node pairs between the first graph and the second graph includes: Determine a second reference image in the second image, and search the first reference image having the highest matching degree with the second reference image in the first image; A first node corresponding to the first reference image and a second node corresponding to the second reference image form the initial matching node pair.
6. The method according to claim 2, characterized in that The optimizing at least one of the first posture data and the second posture data based on the third matching relationship includes: Determine a first pose transformation parameter corresponding to the third matching relationship according to the first pose data corresponding to the first node in the third matching relationship and the second pose data corresponding to the second node; Determine, according to the edge between the first node and the second node in the third matching relationship, a second posture transformation parameter corresponding to the third matching relationship; At least one of the first pose data and the second pose data is optimized based on a difference between the first pose transformation parameter and the second pose transformation parameter.
7. The method according to claim 6, characterized in that The method further comprises: Determining the confidence of the third matching relationship according to the number of matching inner points between the first image corresponding to the first node and the second image corresponding to the second node in the third matching relationship; The optimizing at least one of the first pose data and the second pose data based on the difference between the first pose transformation parameter and the second pose transformation parameter comprises: At least one of the first pose data and the second pose data is optimized based on the difference between the first pose transformation parameter and the second pose transformation parameter and the confidence of the third matching relationship.
8. The method according to claim 1, characterized in that The method further comprises: Acquire a first 3D point subset from the first map data, and acquire a second 3D point subset from the second map data, wherein the first 3D point subset and the second 3D point subset have a matching relationship; Coordinate transformation parameters between the first map data and the second map data are determined based on the three-dimensional distribution characteristics of the first three-dimensional point subset and the second three-dimensional point subset; the coordinate transformation parameters are used to place the first map and the second map in the same coordinate system.
9. The method according to claim 8, characterized in that The acquiring a first 3D point subset from the first map data and acquiring a second 3D point subset from the second map data comprises: Acquire a two-dimensional-two-dimensional matching point pair between the first image and the second image; Determine a 3D-3D matching point pair between a first 3D point and a second 3D point according to the 2D-2D matching point pair; The first 3D point subset is formed by the first 3D point in the 3D-3D matching point pair, and the second 3D point subset is formed by the second 3D point in the 3D-3D matching point pair.
10. The method according to claim 9, characterized in that The step of determining a 3D-3D matching point pair between a first 3D point and a second 3D point according to the 2D-2D matching point pair comprises: Determine, according to the two-dimensional-two-dimensional matching point pair, the pose data of the second image in a first world coordinate system; the first world coordinate system is the world coordinate system of the first map data; Back-projecting the plurality of first three-dimensional points using the position data of the second image in the first world coordinate system to obtain three-dimensional-two-dimensional matching point pairs between the plurality of first three-dimensional points and their corresponding two-dimensional points in the second image; The two-dimensional point in the three-dimensional-two-dimensional matching point pair is mapped to a second three-dimensional point to obtain the three-dimensional-three-dimensional matching point pair.
11. The method according to claim 10, characterized in that Determining the pose data of the second image in the first world coordinate system according to the two-dimensional-two-dimensional matching point pair includes: Mapping the two-dimensional points of the first image in the two-dimensional-two-dimensional matching point pairs into first three-dimensional points, and determining a plurality of candidate poses of the second image in the first world coordinate system based on a matching relationship between the two-dimensional points of the second image and the first three-dimensional points; The plurality of candidate poses are clustered, and pose data of the second image in the first world coordinate system is determined according to an optimal cluster.
12. The method according to claim 11, characterized in that The step of determining a 3D-3D matching point pair between a first 3D point and a second 3D point according to the 2D-2D matching point pair further includes: A first three-dimensional point corresponding to a two-dimensional point in a first image associated with the optimal class is obtained to obtain the plurality of first three-dimensional points for back-projection.
13. The method according to claim 9, characterized in that The obtaining of a two-dimensional-two-dimensional matching point pair between the first image and the second image includes: For at least one second image to be matched in the second image, at least one first image similar to it is searched in the first image, and feature points of the at least one first image and the second image to be matched are matched to obtain the two-dimensional-two-dimensional matching point pair.
14. The method according to claim 8, characterized in that The coordinate transformation parameters include rotation parameters and translation parameters; the determining the coordinate transformation parameters between the first map data and the second map data according to the three-dimensional distribution characteristics of the first three-dimensional point subset and the second three-dimensional point subset includes: Determining a first center point of the first three-dimensional point subset and a second center point of the second three-dimensional point subset; Decentralizing the first three-dimensional point subset based on the first center point, and decentralizing the second three-dimensional point subset based on the second center point; Determine the rotation parameter according to the covariance between the decentralized first three-dimensional point subset and the second three-dimensional point subset; A scale relationship is determined according to the decentralized first three-dimensional point subset and the second three-dimensional point subset, and the translation parameter is determined according to the scale relationship, the rotation parameter, the first center point, and the second center point.
15. The method according to claim 1, characterized in that The method further comprises: Acquire first posture data and second posture data having a similar relationship, and form a first image corresponding to the first posture data and a second image corresponding to the second posture data into an image pair to be compared; In response to the similarity between the two images in the image pair to be compared being less than a preset threshold, the image in the image pair to be compared that is shot later and the first 3D point or the second 3D point corresponding to the image are deleted from the third map data.
16. A map data processing device, characterized in that: include: A map acquisition module, used to acquire first map data of a first scene and second map data of a second scene associated with the first scene; The first map data includes a set of first three-dimensional points, a plurality of first images, and first pose data corresponding to the first images, and the second map data includes a set of second three-dimensional points, a plurality of second images, and second pose data corresponding to the second images; A map generation module, used for generating a first map according to the first posture data and a first matching relationship between the first images; Generate a second atlas according to the second posture data and a second matching relationship between the second images; a relationship determination module, used for placing the first map and the second map in the same coordinate system, and determining a third matching relationship between the first posture data and the second posture data according to the positional relationship between the first map and the second map; A posture optimization module is used to optimize at least one of the first posture data and the second posture data based on the third matching relationship, and fuse the first map data and the second map data according to the optimized first posture data and the second posture data to obtain third map data.
17. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 15 is implemented.
18. An electronic device, characterized in that: include: processor; A memory, configured to store executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 15 by executing the executable instructions.
Citation Information
Patent Citations
Map reconstruction method and device, computer readable medium and electronic equipment
CN112927362A
Method for generating map topology and computing device for executing the method
KR102271745B1