Identification code map reconstruction method and device, electronic equipment and medium
By expanding real-life equipment and identification code technology, the existing large-space mapping technology is solved, and the problem of difficult reconstruction of existing large-space mapping technology is achieved in large scene areas or outdoor environments, achieving efficient and accurate map reconstruction.
Patent Information
- Application Number
- CN202510195420.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-24
AI Technical Summary
Existing large-space mapping technology based on vision is difficult to effectively reconstruct maps in the case of too large scene area or outdoor environment, and SLAM and SFM algorithms have insufficient speed and accuracy.
By extending real-life devices, we use identification codes (such as QR codes, ArUco codes, etc.) to determine the coordinate conversion matrix, filter the image groups, and optimize the coordinate conversion matrix and global three-dimensional coordinates to achieve efficient map reconstruction.
It improves the speed and accuracy of map reconstruction, can effectively complete map reconstruction in large spaces or outdoor scenes, and reduces dependence on the site environment.
Smart Images

Figure CN120198608A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine vision technology, and in particular to a method, device, electronic device and medium for reconstructing an identification code map. Background Art
[0002] Vision-based large-space mapping technology is a technology that uses a vision sensor (such as a camera) to obtain environmental images and processes these images through algorithms to construct a three-dimensional map of a large space. For vision-based large-space mapping technology, existing solutions often use SLAM (Simultaneous Localization and Mapping) technology and SFM (Structure from Motion) technology to map large-space scenes. When using SLAM technology for map reconstruction, it is necessary to combine a head-mounted IMU for multi-sensor fusion mapping to meet the accuracy requirements of large-space map reconstruction. However, such algorithms are relatively complex and require dense texture patterns to be arranged in large-space scenes to ensure the normal completion of large-space site reconstruction by the SLAM algorithm. In addition, for scenes with an area exceeding 1000+ square meters or outdoor scenes, the SLAM algorithm will be unable to complete site reconstruction due to the excessive size of the scene or the change of outdoor scene lighting, and is easily affected by the site environment. The solution for map reconstruction using SFM technology can realize monocular image acquisition and mapping without restrictions on the mapping area. However, this solution cannot recover scale information using monocular image data, and the mapping speed of the SFM solution is relatively slow. It often takes more than dozens of hours to complete the mapping of a 1000-square-meter large-space site, which is 10-20 times slower than the SLAM solution. In addition, the SFM solution also requires dense texture patterns to be arranged in large-space scenes to ensure mapping quality. Summary of the Invention
[0003] Embodiments of this application provide a method, device, electronic device and medium for reconstructing an identification code map to improve the speed and accuracy of map reconstruction while meeting the requirements of the map reconstruction range.
[0004] According to one aspect of this application, a method for reconstructing an identification code map is provided. The method includes:
[0005] Collecting images of a target scene through an extended reality device to obtain an image set, and determining a coordinate transformation matrix corresponding to each identification code according to the image coordinates and local three-dimensional coordinates of each identification code in the image set; wherein, identification codes are distributed and set in the target scene, and the image set contains images of all identification codes in the target scene;
[0006] Grouping the images in the image set that contain the same identification code into an image group, and for each image group, screening and retaining a preset number of frames of images;
[0007] Traverse the image group. For the target image in the currently traversed image group, unify the local three-dimensional coordinates of the feature points of each identification code into the same coordinate system according to the coordinate transformation matrix corresponding to each identification code, obtain the global three-dimensional coordinates of the feature points of each identification code, and optimize the coordinate transformation matrix of the images in the currently traversed image group and the global three-dimensional coordinates obtained by traversing the current image group based on an optimization algorithm; wherein, the target image is a frame of image selected from the image group.
[0008] If the traversal of the image group is completed, optimize the coordinate transformation matrix of the images in all image groups and all the global three-dimensional coordinates obtained by traversing the image group based on an optimization algorithm, and obtain the map reconstruction result based on the optimized global three-dimensional coordinates.
[0009] According to one aspect of the present application, there is provided an identification code map reconstruction device, and the device includes:
[0010] A coordinate transformation matrix determination module, configured to collect images of a target scene through an extended reality device to obtain an image set, and determine a coordinate transformation matrix corresponding to each identification code according to the image coordinates and local three-dimensional coordinates of each identification code in the image set; wherein, identification codes are distributed in the target scene, and the image set includes images of all identification codes in the target scene.
[0011] An image group determination module, configured to classify the images in the image set that contain the same identification code into one image group, and for each image group, screen and retain a preset number of frames of images.
[0012] A local optimization module, configured to traverse the image group. For the target image in the currently traversed image group, unify the local three-dimensional coordinates of the feature points of each identification code into the same coordinate system according to the coordinate transformation matrix corresponding to each identification code, obtain the global three-dimensional coordinates of the feature points of each identification code, and optimize the coordinate transformation matrix of the images in the currently traversed image group and the global three-dimensional coordinates obtained by traversing the current image group based on an optimization algorithm; wherein, the target image is a frame of image selected from the image group.
[0013] A global optimization module, configured to, if the traversal of the image group is completed, optimize the coordinate transformation matrix of the images in all image groups and all the global three-dimensional coordinates obtained by traversing the image group based on an optimization algorithm, and obtain the map reconstruction result based on the optimized global three-dimensional coordinates.
[0014] According to another aspect of the present application, there is provided an electronic device, and the electronic device includes:
[0015] At least one processor; and
[0016] A memory communicatively connected to at least one processor; wherein,
[0017] The memory stores a computer program executable by at least one processor, and the computer program is executed by at least one processor to enable the at least one processor to execute the identification code map reconstruction method according to any embodiment of the present application.
[0018] According to another aspect of the present application, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the identification code map reconstruction method according to any embodiment of the present application when executed.
[0019] The technical solution of the embodiments of the present application obtains an image set by collecting images of a target scene through an extended reality device, determines a coordinate transformation matrix corresponding to each identification code according to the image coordinates and local three-dimensional coordinates of each identification code in the image set; classifies the images containing the same identification code in the image set into an image group, and for each image group, filters and retains a preset number of frames of images, which can establish the three-dimensional coordinates of the identification code through a limited number of images, improve the map reconstruction speed, traverse the image group, and for the target image in the currently traversed image group, unify the local three-dimensional coordinates of the feature points of each identification code into the same coordinate system according to the coordinate transformation matrix corresponding to each identification code to obtain the global three-dimensional coordinates of the feature points of each identification code, and optimize the coordinate transformation matrix of the images in the currently traversed image group and the global three-dimensional coordinates obtained by traversing the current image group based on an optimization algorithm, so as to optimize the global three-dimensional coordinates through the correlation between different images, improve the accuracy of the global three-dimensional coordinates of the identification code unified into the same coordinate system in each image, and through subsequent global optimization, realize the secondary optimization of all three-dimensional coordinates, and further improve the accuracy of the reconstructed identification code map.
[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 It is a flowchart of an identification code map reconstruction method provided by an embodiment of the present application;
[0023] Figure 2Schematic diagram of the identification code provided by the embodiment of the present application;
[0024] Figure 3 Flowchart of a method for reconstructing an identification code map provided by another embodiment of the present application;
[0025] Figure 4 Flowchart of a method for reconstructing an identification code map provided by yet another embodiment of the present application;
[0026] Figure 5 Flowchart of the specific implementation manner provided by the embodiment of the present application;
[0027] Figure 6 Schematic diagram of the structure of an identification code map reconstruction device provided by the embodiment of the present application;
[0028] Figure 7 Schematic diagram of the structure of an electronic device provided by the embodiment of the present application. Specific implementation manner
[0029] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0030] It should be noted that the terms "first", "second", "third", "fourth", "actual", "preset", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0031] Figure 1 Flowchart of a method for reconstructing an identification code map provided by the embodiment of the present application. The embodiment of the present application is applicable to the situation of reconstructing an identification code map. This method can be executed by an identification code map reconstruction device, which can be implemented in the form of hardware and / or software, and the identification code map reconstruction device can be configured in an electronic device. As Figure 1As shown, the method includes:
[0032] S110. Collect an image set of a target scene through an extended reality device, and determine a coordinate transformation matrix corresponding to each identification code according to the image coordinates and local three-dimensional coordinates of each identification code in the image set; wherein, identification codes are distributed in the target scene, and the image set includes images of all the identification codes in the target scene.
[0033] Among them, the extended reality device, i.e., the XR device, is a hardware device integrating virtual reality (VR), augmented reality (AR), and mixed reality (MR) technologies. In the embodiments of the present application, there may be at least one extended reality device, and the solutions in the embodiments of the present application are executed for each extended reality device. The extended reality device can be driven by a user wearing it on the head, holding it in the hand, or other carrying methods to move, or carried by a robot to move, and perform a scanning image collection on the target scene. Identification codes are set at different positions in the target scene. The identification code is a synthetic or fiducial marker used in computer vision applications, and the form of the identification code is not limited. For example, it can be a QR code, an ArUco code, or an AprilTag code, etc. Identification codes, such as ArUco and AprilTag, are binary square fiducial markers. The main advantage of these markers is that a single marker can provide enough correspondences (such as four corners) to obtain the pose of the extended reality device.
[0034] In the embodiments of the present application, the target scene is imaged at different positions and postures through the extended reality device, so that the collected images contain identification codes, and the set of identification codes included in multiple frames of images collected by the mobile extended reality device contains all the identification codes in the target scene, so as to more comprehensively and accurately perform map reconstruction on the identification codes in the target scene. The set of all images collected by the extended reality device is used as the image set.
[0035] Exemplarily, for each image in the image set, the image contains an identification code. When setting the identification code, the identification code identifier and size information of the identification code are known, that is, the local three-dimensional coordinates of the feature points in the identification code are known. For example, Figure 2As shown, the identification of the identification code, i.e., ID is 0. Assuming the side length of the identification code is L, and taking the four vertices of the identification code as feature points, the local three-dimensional coordinates of the four feature points are A(0, 0, 0), B(L, 0, 0), C(0, L, 0), and D(L, L, 0) respectively. The identification code in the image can be recognized to determine the image coordinates of the identification code in the image. Based on the image coordinates and local three-dimensional coordinates of the identification code, the coordinate transformation matrix corresponding to the identification code can be determined, and this transformation matrix actually reflects the position and attitude information when the image collector in the extended reality device performs image acquisition. Specifically, the cv::solvePnP function of OpenCV can be called to determine the coordinate transformation matrix corresponding to the identification code, and the reprojection error corresponding to this transformation matrix can be output. The reprojection error reflects the difference between the pixel coordinates obtained by projecting a three-dimensional space point onto the image plane through the camera model and the actual observed pixel coordinates. The smaller the reprojection error, the more accurate the coordinate transformation matrix is reflected.
[0036] S120. Group the images in the image set that contain the same identification code into an image group, and for each image group, filter and retain a preset number of frames of images.
[0037] Exemplarily, if there are the same identification codes in different images, that is, it may be obtained by performing image acquisition on the same identification code at different positions and postures. In this case, only one global three-dimensional coordinate determination process is required, and there is no need to execute the global three-dimensional coordinate determination process for each frame of image. The images in the image set that contain the same identification code can be grouped into an image group. To avoid excessive number of images in the image group affecting the processing efficiency, a preset number of frames of images can be filtered and retained, thereby effectively improving the map reconstruction speed and reducing unnecessary computing power consumption while ensuring the accuracy.
[0038] Specifically, containing the same identification code means that the number of contained identification codes is the same and the identification of the identification code is also the same. For example, if image 1 contains identification code a, identification code b, and identification code c, and image 2 also contains identification code a, identification code b, and identification code c, then it is determined that image 1 and image 2 contain the same pair of identification codes and are grouped into the same image group. If image 3 contains identification code a, identification code b, and identification code d, even though the number of identification codes is the same, but the identification of the identification code is different, then image 3 cannot be grouped with image 1 and image 2 into an image group.
[0039] To filter and retain a preset number of frames of images, it can be filtered according to the actual situation. For example, images with an image quality evaluation score higher than a preset threshold, the clarity of the identification code area higher than the clarity threshold, the area of the identification code area higher than the area threshold, etc. can be filtered. The preset number can be selected according to the actual situation to reduce computing power consumption while ensuring the map reconstruction accuracy. For example, it can be selected as 5.
[0040] S130. Traverse the image group. For the target image in the currently traversed image group, unify the local three-dimensional coordinates of the feature points of each identification code into the same coordinate system according to the coordinate transformation matrix corresponding to each identification code, obtain the global three-dimensional coordinates of the feature points of each identification code, and optimize the coordinate transformation matrix of the images in the currently traversed image group and the global three-dimensional coordinates obtained by traversing the current image group based on an optimization algorithm; wherein, the target image is a frame of image selected from the image group.
[0041] Exemplarily, traverse each image group. For each image group, unify the local three-dimensional coordinates of all the identification codes included in the image group into the same coordinate system to obtain the global three-dimensional coordinates of the feature points of each identification code. Specifically, the local three-dimensional coordinates of each identification code are known, and the corresponding coordinate transformation matrix is also known. The feature points of each identification code can be unified into the same coordinate system according to the local three-dimensional coordinates of the feature points of each identification code and the corresponding coordinate transformation matrix. Since the same identification codes are included in each image of each image group, it is only necessary to select one image, and unify the local three-dimensional coordinates of the feature points of each identification code into the same coordinate system according to the coordinate transformation matrix corresponding to the identification code in this image to obtain the global three-dimensional coordinates of the feature points of each identification code.
[0042] After determining the global three-dimensional coordinates of the feature points of each identification code in the target image, the coordinate transformation matrix of all the images in the image group and the global three-dimensional coordinates can be optimized based on an optimization algorithm, so that the global three-dimensional coordinates and the coordinate transformation matrix are more accurate. The beneficial effect of the above solution is that it can optimize the global three-dimensional coordinates based on the limited images in the image group, accurately determine the coordinates of the feature points of the identification code, and can only process the limited images without processing all the images, effectively improving the processing speed and avoiding unnecessary computing power consumption.
[0043] S140. If the traversal of the image group is completed, optimize the coordinate transformation matrix of the images in all the image groups and all the global three-dimensional coordinates obtained by traversing the image group based on an optimization algorithm, and obtain the map reconstruction result based on the optimized global three-dimensional coordinates.
[0044] Exemplarily, after traversing all the images in the image group, that is, after the global three-dimensional coordinates of all the identification codes in the target scene have been determined, the global optimization can be further performed on all the global three-dimensional coordinates to improve the global accuracy of the coordinates. Specifically, the coordinate transformation matrix of the images in all the image groups and all the global three-dimensional coordinates can be optimized based on the optimization algorithm to obtain the more accurate global three-dimensional coordinates of all the feature points of the identification codes. Through the above scheme of first performing local optimization and then global optimization, after obtaining the global three-dimensional coordinates of the feature points of the identification code in an image group, the coordinate optimization in one stage can be realized, the processing speed and coordinate accuracy can be improved, and finally the global optimization is performed to further improve the coordinate accuracy.
[0045] In the technical solution of the embodiment of the present application, at least two linear array cameras are made to face the checkerboard calibration board for image acquisition to obtain the target images respectively collected by each linear array camera; for the target images collected by two adjacent linear array cameras, according to the pixel points with the maximum pixel gradient in the target images and the adjacent pixel points of the pixel points with the maximum pixel gradient, the adjacent checkerboard demarcation points are determined as the feature points, so as to make full use of the pixel value distribution characteristics of the checkerboard and the gradient characteristics of the pixel points to accurately determine the feature points located at the checkerboard demarcation points. According to the image coordinates of the feature points in the target images and the plane coordinates of the feature points in the plane coordinate system established based on the checkerboard calibration board, the perspective transformation matrix for the transformation between the image coordinates and the plane coordinates is determined, so as to more accurately determine the transformation relationship between the image coordinates, which can realize linear and non-linear transformations, effectively eliminate the accuracy loss of non-linearity, and improve the accuracy of the linear array camera and the fusion degree of the spliced images.
[0046] Figure 3 The flowchart of a method for reconstructing an identification code map provided in another embodiment of the present application is based on the above embodiment for optimization. For the solutions not described in detail in the embodiment of the present application, please refer to the above embodiment. As Figure 3 shown, the method of the embodiment of the present application specifically includes the following steps:
[0047] S210. Use the extended reality device to perform image acquisition on the target scene to obtain an image set. Among them, identification codes are distributed in the target scene, and the image set contains images of all the identification codes in the target scene.
[0048] S220. Perform identification code detection on each image in the image set to determine the image coordinates of the feature points in each identification code in the image.
[0049] Exemplarily, it is possible to detect the identification codes in each image in the image set, obtain each identification code, and identify the image coordinates of the feature points in the image for the pre-specified feature points. The cv::aruco::detectMarkers function of the Opencv library or the apriltag_detector_detect function of apriltag can be called to detect and identify the identification codes existing in the image. Assuming that the pre-specified feature points are the four vertices of a square identification code, the image coordinates of the four vertices of each identification code are identified. The output result may include the identification code identifier of the identification code and the image coordinates of the feature points of the identification code.
[0050] S230. Determine the coordinate transformation matrix corresponding to each identification code according to the image coordinates and local three-dimensional coordinates of the feature points in the identification code.
[0051] Call the cv::solvePnP function of OpenCV for rough positioning, that is, determine the coordinate transformation matrix corresponding to each identification code. Specifically, for each identification code, input the image coordinates and local three-dimensional coordinates of the feature points of the identification code into the called function, and output the corresponding coordinate transformation matrix and reprojection error.
[0052] S240. Identify the identification code identifiers of each identification code in the images of the image set, and for each image, store the identification code identifier included in the image in a set class and set an image category code for the set class.
[0053] Exemplarily, for each image in the image set, identify the identification code identifiers of each identification code in the image, store the identification code identifiers of each identification code in an image in a set class, and set an image category code for the set class. For example, obtain all the recognized QR code IDs_i in image i, store IDs_i in the c++ set class Set, and use the Set object as the image category code for this image.
[0054] S250. Group the images in the image set, use the image category code as the key index and the image identifier as the value, and index the image identifiers with the same image category code into one image group.
[0055] Exemplarily, it is possible to group the images in the image set, use the image category code as the key index and the image identifier as the value, and group the image identifiers with the same image category code into one image group, that is, group the images with the same identification code into one image group.
[0056] S260. For each image in each image group, respectively determine the evaluation score of the image according to the size information of the identification code included in the image and / or the reprojection error of the coordinate transformation matrix corresponding to the identification code.
[0057] Exemplarily, a group of images may contain multiple images. An excessive number of images may affect the data processing speed. Therefore, the images can be screened to appropriately reduce the number of images and improve the map reconstruction speed without affecting the coordinate accuracy. For each image in the group of images, an evaluation can be performed to determine an evaluation score. The evaluation score of the image can be determined based on the size information of the identification code contained in the image and / or the reprojection error of the coordinate transformation matrix corresponding to the identification code. Since the identification code in the image and the coordinate transformation matrix corresponding to the identification code are required in the process of map reconstruction, the accuracy of identifying the identification code and the accuracy of the coordinate transformation matrix corresponding to the identification code affect the accuracy of map reconstruction. The size information of the identification code can reflect the accuracy of identifying the identification code. Generally, the accuracy of identifying the identification code is positively correlated with the size information of the identification code. The reprojection error reflects the accuracy of the coordinate transformation matrix, and the accuracy of the coordinate transformation matrix is negatively correlated with the magnitude of the reprojection error. The evaluation score of the image can be determined based on the size information of the identification code and / or the reprojection error of the coordinate transformation matrix corresponding to the identification code.
[0058] In an embodiment of the present application, determining the evaluation score of the image according to the size information of the identification code contained in the image and / or the reprojection error of the coordinate transformation matrix corresponding to the identification code includes:
[0059] Determining the average image area of the identification code in the image and / or the average reprojection error of the coordinate transformation matrix corresponding to the identification code;
[0060] Determining the evaluation score of the image according to the average image area and / or the average reprojection error; wherein, the evaluation score of the image is positively correlated with the average image area and negatively correlated with the average reprojection error.
[0061] Exemplarily, the image may include multiple identification codes. The side lengths of each identification code can be recognized in the image, the image area of each identification code can be calculated based on the side lengths, and the average value of the image areas of each identification code can be calculated. A coordinate transformation matrix can be calculated based on the image coordinates and local three-dimensional coordinates of each identification code. Each coordinate transformation matrix corresponds to a reprojection error, and the average value of the reprojection errors of the coordinate transformation matrices corresponding to each identification code can be calculated. According to the average value of the image area and / or the average value of the reprojection error, the evaluation score of the image is determined. Specifically, the weight corresponding to the image area and the weight corresponding to the average value of the reprojection error can be set, and the average value of the image area and the average value of the reprojection error are weighted and summed to obtain the evaluation score of the image. Before performing the weighted summation process on the average value of the image area and the average value of the reprojection error, a normalization process can also be performed, so that the average value of the image area and the average value of the reprojection error are in the same dimension. Score = w1*mArea / max(maxMArea,mArea)+w2*(1-mPError / max(minMPError,mPError)); where mArea is the average value of the image area, mPError is the average value of the reprojection error, max() is the operation of taking the maximum value; w1 is the weight of the average value of the image area, w2 is the weight of the average value of the reprojection error; maxMArea is the preset maximum area, for example, it can be taken as the range of 60*60 pixels. minMPError is the preset minimum reprojection error, for example, it can be taken as 0.4.
[0062] S270. Select and retain a preset number of images from the image group according to the evaluation score.
[0063] Exemplarily, since the higher the evaluation score, the more conducive it is to improving the accuracy of map reconstruction, the preset number of images with the highest evaluation scores can be selected for retention. The preset number can be selected as 5, for example.
[0064] S280. Traverse the image group. For the target image in the currently traversed image group, unify the local three-dimensional coordinates of the feature points of each identification code into the same coordinate system according to the coordinate transformation matrix corresponding to each identification code, obtain the global three-dimensional coordinates of the feature points of each identification code, and optimize the coordinate transformation matrix of the images in the currently traversed image group and the global three-dimensional coordinates obtained by traversing the current image group based on the optimization algorithm; where the target image is a frame of image selected from the image group.
[0065] In the embodiment of the present application, traversing the image group includes:
[0066] Regarding the image group with the most identification codes included as the image group obtained in the first traversal;
[0067] In the process of selecting the image group for the next traversal from the remaining image groups, the image group with the largest number of overlapping identification codes with the identification codes whose global three-dimensional coordinates have been determined currently is used as the image group for the next traversal.
[0068] The image group with the largest number of included identification codes is used as the image group obtained in the first traversal, which can preferentially determine the global three-dimensional coordinates of as many identification codes as possible. In the process of selecting the image group for the next traversal from the remaining image groups, the image group with the largest number of overlapping identification codes with the identification codes whose global three-dimensional coordinates have been determined currently is used as the image group for the next traversal, which can select the image group with the greatest correlation with the previously traversed image groups, that is, the image group with the highest identification code overlap degree, so as to achieve strong correlation of the identification codes and perform more accurate unification and fusion of the three-dimensional coordinates.
[0069] S290. If the traversal of the image groups is completed, optimize the coordinate transformation matrix of the images in all the image groups and all the global three-dimensional coordinates obtained by traversing the image groups based on the optimization algorithm, and obtain the map reconstruction result based on the optimized global three-dimensional coordinates.
[0070] The embodiment of the present application provides an identification code map reconstruction method, which detects the identification codes of each image in the image set and determines the image coordinates of the feature points in each identification code in the image; determines the coordinate transformation matrix corresponding to each identification code according to the image coordinates of the feature points in the identification code and the local three-dimensional coordinates, and can determine the conversion relationship between the coordinates in the image and the actual three-dimensional coordinates through the identification code. By identifying the identification code identifiers of each identification code in the images of the image set, for each image, store the identification code identifier included in the image in a set class, set an image category code for the set class, group the images in the image set, use the image category code as the key index, and use the image identifier as the value, and index the image identifiers with the same image category code into one image group to realize the classification and merging of the images, perform unified operations on the images containing the same identification code, effectively improve the processing speed, and avoid unnecessary waste of computing power caused by repeated processing of the images. By respectively determining the evaluation score of each image according to the size information of the identification code included in the image and / or the reprojection error of the coordinate transformation matrix corresponding to the identification code for each image in each image group; according to the evaluation score, select a preset number of images to be retained from the image group, which can appropriately screen the images in the image group, reduce the number of images participating in coordinate optimization while retaining enough images for correlation to improve the coordinate accuracy, so as to improve the processing speed.
[0071] Figure 4The flowchart of a method for reconstructing an identification code map provided by another embodiment of this application. The embodiments of this application are optimized based on the above embodiments. For the solutions not described in detail in the embodiments of this application, refer to the above embodiments. As Figure 4 shown, the method of the embodiment of this application specifically includes the following steps:
[0072] S310. Use an extended reality device to collect images of a target scene to obtain an image set, and determine the coordinate transformation matrix corresponding to each identification code according to the image coordinates and local three-dimensional coordinates of each identification code in the image set; wherein, identification codes are distributed in the target scene, and the image set contains images of all identification codes in the target scene.
[0073] S320. Group the images containing the same identification code in the image set into an image group, and for each image group, filter and retain a preset number of frames of images.
[0074] S330. Traverse the image group. For the target image in the currently traversed image group, select a target identification code from the identification codes contained in the target image, and use the local three-dimensional coordinates of the feature points of the target identification code as the global three-dimensional coordinates of the feature points of the target identification code.
[0075] Exemplarily, in the process of determining the global three-dimensional coordinates of the identification code, as long as the corresponding global three-dimensional coordinates of all the identification codes in the image are determined, so only the identification codes in one image are needed to determine the global three-dimensional coordinates. It is possible to select a target image from the image group during the traversal of the image group, and for all the identification codes in the target image, determine the global three-dimensional coordinates of the identification codes. Specifically, select a target identification code from the identification codes contained in the target image, and use the local three-dimensional coordinates of the feature points of the target identification code as the global three-dimensional coordinates of the feature points of this identification code, that is, establish a reference origin, and the global three-dimensional coordinates of the feature points of other identification codes are established based on this. The target image can be the image with the highest evaluation score in the image group.
[0076] In the embodiment of this application, selecting a target identification code from the identification codes contained in the target image includes:
[0077] Obtain the reprojection error corresponding to the output during the process of determining the coordinate transformation matrix corresponding to each identification code;
[0078] Use the identification code corresponding to the coordinate transformation matrix with the smallest reprojection error as the target identification code.
[0079] Exemplarily, during the process of determining the coordinate transformation matrix corresponding to each identification code, a reprojection error will be output correspondingly. The identification code corresponding to the coordinate transformation matrix with the smallest reprojection error is used as the target identification code, that is, the identification code with the most accurate correspondence between the image coordinates and the local three-dimensional coordinates is used as the target identification code.
[0080] S340. Based on the origin in the global three-dimensional coordinates of the feature points corresponding to the target identification code as a reference, according to the coordinate transformation matrix of the target identification code, the coordinate transformation matrices of other identification codes except the target identification code, and the local three-dimensional coordinates of the feature points of other identification codes, determine the global three-dimensional coordinates of the feature points of other identification codes.
[0081] Exemplarily, the origin in the global three-dimensional coordinates of the feature points obtained from the target identification code can be used as a reference, that is, as the global origin. According to the coordinate transformation matrix of the target identification code, the coordinate transformation matrices of other identification codes except the target identification code, and the local three-dimensional coordinates of the feature points of other identification codes, determine the global three-dimensional coordinates of the feature points of other identification codes, so as to unify the three-dimensional coordinates of each identification code into the same coordinate system to form a global coordinate.
[0082] In the embodiment of the present application, according to the coordinate transformation matrix corresponding to the target identification code, the coordinate transformation matrices of other identification codes except the target identification code, and the local three-dimensional coordinates of the feature points of other identification codes, determining the global three-dimensional coordinates of the feature points of other identification codes includes:
[0083] Taking the product of the inverse matrix of the coordinate transformation matrix corresponding to the target identification code, the coordinate transformation matrix corresponding to the other identification code, and the local three-dimensional coordinates of the feature points of the other identification code as the global three-dimensional coordinates of the feature points of the other identification code.
[0084] Exemplarily, the global three-dimensional coordinates of the feature points of other identification codes can be determined by the following formula:
[0085] iP_align_j = inv(bestT_cam_marker_i) * T_cam_marker_j * P_marker;
[0086] Among them, inv() is the matrix inverse operation; bestT_cam_marker_i is the coordinate transformation matrix of the target image, and the coordinate transformation matrix corresponding to each identification code of the target image can be taken, and the coordinate transformation matrix with the smallest reprojection error; T_cam_marker_j is the coordinate transformation matrix corresponding to the j-th identification code of the target image; P_marker is the local three-dimensional coordinates of the j-th identification code of the target image; iP_align_j is the global three-dimensional coordinates of the j-th identification code in the target image.
[0087] S350. Optimize the coordinate transformation matrix of the images in the currently traversed image group and the global three-dimensional coordinates obtained by traversing the current image group based on an optimization algorithm.
[0088] S360. If the traversal of the image group is completed, optimize the coordinate transformation matrix of the images in all image groups and all the global three-dimensional coordinates obtained by traversing the image groups based on an optimization algorithm, and obtain a map reconstruction result based on the optimized global three-dimensional coordinates.
[0089] As a non-limiting implementation, optimizing the coordinate transformation matrix of the images in the currently traversed image group and the global three-dimensional coordinates obtained by traversing the current image group based on an optimization algorithm includes:
[0090] For each image in the currently traversed image group, use the coordinate transformation matrix with the smallest reprojection error in the coordinate transformation matrices corresponding to each identification code in the image as the coordinate transformation matrix of the image.
[0091] Input the coordinate transformation matrices of the images in the image group and the global three-dimensional coordinates obtained by traversing the current image group into a bundle adjustment algorithm to optimize the global three-dimensional coordinates.
[0092] Exemplarily, for each traversed image group, obtain the global three-dimensional coordinates of all identification codes in the image group, and input the global three-dimensional coordinates of all identification codes in the image group and the coordinate transformation matrices corresponding to all images into the bundle adjustment algorithm to achieve the optimization of the global three-dimensional coordinates. This optimization process is only for the global three-dimensional coordinates of all identification codes in one image group, which is a local optimization process with a small amount of data to be optimized, effectively improving the optimization efficiency.
[0093] As a non-limiting implementation, optimizing the coordinate transformation matrix of the images in all image groups and all the global three-dimensional coordinates obtained by traversing the image groups based on an optimization algorithm includes:
[0094] For each image in each image group, use the coordinate transformation matrix with the smallest reprojection error in the coordinate transformation matrices corresponding to each identification code in the image as the coordinate transformation matrix of the image.
[0095] Input the coordinate transformation matrices of the images in all image groups and all the global three-dimensional coordinates into a bundle adjustment algorithm to optimize the global three-dimensional coordinates.
[0096] Exemplarily, after all image groups have been traversed, all the three-dimensional coordinates obtained by traversing all the image groups and the coordinate transformation matrices of all the images are input into the bundle adjustment algorithm to optimize the global three-dimensional coordinates. This optimization process is an optimization for all the global three-dimensional coordinates and is a global optimization process, which can further improve the accuracy of the three-dimensional coordinates.
[0097] The embodiment of the present application provides a method for reconstructing an identification code map, which selects a target identification code from the identification codes included in the target image, and uses the local three-dimensional coordinates of the feature points of the target identification code as the global three-dimensional coordinates of the feature points of the target identification code; based on the origin in the global three-dimensional coordinates of the feature points of the target identification code as a reference, according to the coordinate transformation matrix of the target identification code, the coordinate transformation matrices of other identification codes except the target identification code, and the local three-dimensional coordinates of the feature points of other identification codes, the global three-dimensional coordinates of the feature points of other identification codes are determined, and the global three-dimensional coordinates of the identification code feature points are accurately determined through a limited number of images, and the coordinate optimization is performed through the association between the images, improving the map reconstruction accuracy.
[0098] Figure 5 It is a flowchart of a specific implementation manner provided by the embodiment of the present application, as Figure 5 shown, a specific implementation process includes:
[0099] 1. The XR device collects mapping data
[0100] Using the XR device camera, collect the large-space two-dimensional code environment image data to obtain the mapping image set I, as well as the internal parameter K of the XR device camera and the camera distortion parameter D.
[0101] 2. Image preprocessing
[0102] Detect and identify the two-dimensional codes in the image, and at the same time use the two-dimensional codes for visual positioning to obtain the camera pose (position and attitude) when taking the image.
[0103] (1) Two-dimensional code detection and identification
[0104] Call the cv::aruco::detectMarkers function of the Opencv library or the apriltag_detector_detect function of apriltag to detect and identify the two-dimensional codes existing in the image. Image i belongs to the mapping image set I, and image i is used as the input for two-dimensional code detection and identification, and the output results include the image coordinates Boxes_i of the four vertices of Ni two-dimensional codes and the identified two-dimensional code identifiers IDs_i.
[0105] Where i is the image number
[0106] (2) Coarse localization of QR codes
[0107] Using the image coordinates Boxes_i of the four vertices of the QR codes detected and recognized in (1), and calling the cv::solvePnP function in OpenCV for coarse localization. For the Ni QR codes in image i, the cv::solvePnP function needs to be called one by one for coarse localization to obtain the coordinate transformation matrix T_cam_markers_i under the condition of Ni QR codes and the reprojection error ProjErrors_i corresponding to the Ni coordinate transformation matrices.
[0108] 3. Automatic grouping of images
[0109] Images in the same image group have the same number and ID of QR codes, and N images are retained in each image group; when the number of images in the image group is greater than N, the redundant images will be filtered out, and when the number of images in the image group is less than or equal to N, all images will be retained. N can be set to 5.
[0110] 31) Image category encoding
[0111] Step1: Obtain all the recognized QR code IDs_i in image i.
[0112] Step2: Store IDs_i into the C++ set class Set, and use the Set object as the image category encoding.
[0113] 32) Image grouping and filtering
[0114] Step1: Create a dictionary for automatic image grouping. Implement it using the C++ dictionary class Map, where the Key is the Set object and the Value is an array of image numbers.
[0115] Step2: Store the category encoding Set objects and image numbers of all images into the dictionary. Images with the same category encoding will be automatically grouped into one group, and the images in the image group have the same number and ID of QR codes.
[0116] 33) Filtering of grouped images
[0117] To control the number of images in each image group, it is necessary to filter the image groups with more than N images in the image group. The specific filtering strategy is as follows:
[0118] Step1: Calculate the average area mArea of all QR codes in the image;
[0119] Step2: Calculate the average reprojection error mPError of the localization of all QR codes in the image;
[0120] Step 3: Calculate the image score score, with the formula as follows: Score = w1 * mArea / max(maxMArea, mArea) + w2 * (1 - mPError / max(minMPError, mPError));
[0121] Among them, mArea is the average value of the image area, mPError is the average value of the reprojection error, and max() is the operation of taking the maximum value; w1 is the weight of the average value of the image area, and w2 is the weight of the average value of the reprojection error; maxMArea is the preset maximum area, for example, it can take the range of 60 * 60 pixels. minMPError is the preset minimum reprojection error, for example, it can take 0.4.
[0122] Step 4: Sort the images according to the image score score, and only keep the top N images, and the remaining images will be directly discarded.
[0123] 4. Incremental QR code reconstruction
[0124] 41) QR code reconstruction initialization
[0125] A) Select the initialization image group: Select the image group with the largest number of QR codes as the initialization image group ImgGroup_init.
[0126] B) Obtain the best coordinate transformation matrix: For image i, select the coordinate transformation matrix with the smallest reprojection error as the best coordinate transformation matrix bestT_cam_marker_i of image i.
[0127] C) Multi-QR code coordinate alignment: Unify all the QR codes in the image group ImgGroup_init to one coordinate system. When there are multiple QR codes in the image group ImgGroup_init, it is necessary to unify the 3D coordinate systems of multiple QR codes to one coordinate system in order to complete the initialization of the 3D coordinates of the QR codes. Figure 3 D coordinate initialization.
[0128] The specific steps for unifying the 3D coordinates of multiple QR codes in image i are as follows:
[0129] Step 1 Determine the origin of the unified coordinate system: Use the image with the highest score in the image group as the target image, and use the vertex A of the QR code with the smallest reprojection error in the target image as the origin (0, 0, 0).
[0130] Step 2 Coordinate unification: Convert the 3D coordinates of the remaining QR codes to the unified coordinate system through the following formula
[0131] iP_align_j = inv(bestT_cam_marker_i) * T_cam_marker_j * P_marker
[0132] Among them, inv() is the matrix inversion operation; bestT_cam_marker_i is the coordinate transformation matrix of the target image, which can be the coordinate transformation matrix corresponding to each QR code in the target image, and the coordinate transformation matrix with the smallest reprojection error; T_cam_marker_j is the coordinate transformation matrix corresponding to the j-th QR code in the target image; P_marker is the local three-dimensional coordinate of the j-th QR code in the target image; iP_align_j is the global three-dimensional coordinate of the j-th QR code in the target image.
[0133] D) Local BA optimization: Use the BA algorithm to locally optimize the global three-dimensional coordinates of the QR code and the coordinate transformation matrix to make the global three-dimensional coordinates of the QR code and the coordinate transformation matrix more accurate. The currently determined global three-dimensional coordinates form the incremental reconstruction map IncreaseMap.
[0134] 42) Obtain the next reconstruction image group
[0135] Select the image group ImgGroup_k with the most identical QR code associations with the incremental reconstruction map IncreaseMap as the next reconstruction image group.
[0136] 42) Incremental reconstruction of the image group
[0137] A) Obtain all QR codes GroupMarkers in ImgGroup_k
[0138] B) Alignment of multi-QR code coordinates. Process the QR codes GroupMarkers in ImgGroup_k according to the steps in 41) A and B to obtain the global three-dimensional coordinates.
[0139] 5. Global BA optimization
[0140] Use the BA algorithm to further globally optimize the global three-dimensional coordinates of the QR code and the coordinate transformation matrix to make the global three-dimensional coordinates and the coordinate transformation matrix more accurate.
[0141] 6. Result output
[0142] Output the global three-dimensional coordinates of the QR code map and the QR code ID corresponding to the global three-dimensional coordinates.
[0143] Figure 6The following is a schematic structural diagram of an identification code map reconstruction device provided by an embodiment of the present application. The device can execute the identification code map reconstruction method provided by any embodiment of the present application, and has corresponding functional modules and beneficial effects for executing the method. As Figure 6 shown, the device includes:
[0144] A coordinate transformation matrix determination module 410, configured to collect images of a target scene through an augmented reality device to obtain an image set, and determine a coordinate transformation matrix corresponding to each identification code according to the image coordinates and local three-dimensional coordinates of each identification code in the image set; wherein, identification codes are distributed and arranged in the target scene, and the image set includes images of all identification codes in the target scene;
[0145] An image group determination module 420, configured to classify images in the image set that contain the same identification code into one image group, and for each image group, screen and retain a preset number of frames of images;
[0146] A local optimization module 430, configured to traverse the image group, and for a target image in the currently traversed image group, unify the local three-dimensional coordinates of the feature points of each identification code into the same coordinate system according to the coordinate transformation matrix corresponding to each identification code, to obtain the global three-dimensional coordinates of the feature points of each identification code, and optimize the coordinate transformation matrix of the images in the currently traversed image group and the global three-dimensional coordinates obtained by traversing the current image group based on an optimization algorithm; wherein, the target image is a frame of image selected from the image group;
[0147] A global optimization module 440, configured to, if the traversal of the image group ends, optimize the coordinate transformation matrix of the images in all image groups and all global three-dimensional coordinates obtained by traversing the image group based on an optimization algorithm, and obtain a map reconstruction result based on the optimized global three-dimensional coordinates.
[0148] In an embodiment of the present application, the coordinate transformation matrix determination module determines a coordinate transformation matrix corresponding to each identification code according to the image coordinates and local three-dimensional coordinates of each identification code in the image set, including:
[0149] Detect identification codes for each image in the image set to determine the image coordinates of feature points in each identification code in the image;
[0150] Determine a coordinate transformation matrix corresponding to each identification code according to the image coordinates and local three-dimensional coordinates of the feature points in the identification code.
[0151] In an embodiment of the present application, the image group determination module 420 classifies images in the image set that contain the same identification code into one image group, including:
[0152] Identify the identification codes of each identification code in the images of the image set, and for each image, store the identification codes included in the image in a set class, and set an image category code for the set class;
[0153] Group the images in the image set, use the image category code as the key index and the image identifier as the value, and index the image identifiers with the same image category code into one image group.
[0154] In the embodiment of the present application, the image group determination module 420 screens and retains a preset number of frames of images for each image group, including:
[0155] For each image in each image group, respectively determine the evaluation score of the image according to the size information of the identification code included in the image and / or the reprojection error of the coordinate transformation matrix corresponding to the identification code;
[0156] According to the evaluation score, select a preset number of images from the image group for retention.
[0157] In the embodiment of the present application, the image group determination module 420 determines the evaluation score of the image according to the size information of the identification code included in the image and / or the reprojection error of the coordinate transformation matrix corresponding to the identification code, including:
[0158] Determine the average image area of the identification code in the image and / or the average reprojection error of the coordinate transformation matrix corresponding to the identification code;
[0159] According to the average image area and / or the average reprojection error, determine the evaluation score of the image; wherein, the evaluation score of the image is positively correlated with the average image area and negatively correlated with the average reprojection error.
[0160] In the embodiment of the present application, the local optimization module 430 unifies the local three-dimensional coordinates of the feature points of each identification code into the same coordinate system according to the coordinate transformation matrix corresponding to each identification code, and obtains the global three-dimensional coordinates of the feature points of each identification code, including:
[0161] Select a target identification code from the identification codes included in the target image, and use the local three-dimensional coordinates of the feature points of the target identification code as the global three-dimensional coordinates of the feature points of the target identification code;
[0162] Based on the origin in the global three-dimensional coordinates of the feature points of the target identification code as a reference, according to the coordinate transformation matrix of the target identification code, the coordinate transformation matrices of other identification codes except the target identification code, and the local three-dimensional coordinates of the feature points of other identification codes, determine the global three-dimensional coordinates of the feature points of other identification codes.
[0163] In the embodiment of the present application, the local optimization module 430 selects a target identification code from the identification codes included in the target image, including:
[0164] Obtain the reprojection error corresponding to the output during the process of determining the coordinate transformation matrix corresponding to each identification code;
[0165] Use the identification code corresponding to the coordinate transformation matrix with the minimum reprojection error as the target identification code.
[0166] In the embodiment of the present application, the local optimization module 430 determines the global three-dimensional coordinates of the feature points of other identification codes according to the coordinate transformation matrix corresponding to the target identification code, the coordinate transformation matrices corresponding to other identification codes except the target identification code, and the local three-dimensional coordinates of the feature points of other identification codes, including:
[0167] Take the product of the inverse matrix of the coordinate transformation matrix corresponding to the target identification code, the coordinate transformation matrices corresponding to other identification codes, and the local three-dimensional coordinates of the feature points of other identification codes as the global three-dimensional coordinates of the feature points of other identification codes.
[0168] In the embodiment of the present application, the local optimization module 430 optimizes the coordinate transformation matrix of the images in the currently traversed image group and the global three-dimensional coordinates obtained by traversing the current image group based on an optimization algorithm, including:
[0169] For each image in the currently traversed image group, use the coordinate transformation matrix with the minimum reprojection error among the coordinate transformation matrices corresponding to each identification code in the image as the coordinate transformation matrix of the image;
[0170] Input the coordinate transformation matrices of the images in the image group and the global three-dimensional coordinates obtained by traversing the current image group into a bundle adjustment algorithm to optimize the global three-dimensional coordinates.
[0171] In the embodiment of the present application, the global optimization module 440 optimizes the coordinate transformation matrix of the images in all image groups and all the global three-dimensional coordinates obtained by traversing the image groups based on an optimization algorithm, including:
[0172] For each image in each image group, use the coordinate transformation matrix with the minimum reprojection error among the coordinate transformation matrices corresponding to each identification code in the image as the coordinate transformation matrix of the image;
[0173] Input the coordinate transformation matrices of the images in all image groups and all the global three-dimensional coordinates into a bundle adjustment algorithm to optimize the global three-dimensional coordinates.
[0174] In the embodiment of the present application, the local optimization module 430 traverses the image group, including:
[0175] Take the image group with the most included identification codes as the image group obtained in the first traversal;
[0176] In the process of selecting the image group for the next traversal from the remaining image groups, take the image group with the largest number of overlapping identification codes with the identification codes whose global three-dimensional coordinates have been determined currently as the image group for the next traversal.
[0177] An identification code map reconstruction device provided by an embodiment of the present application can execute an identification code map reconstruction method provided by any embodiment of the present application, and has corresponding functional modules and beneficial effects for executing the method.
[0178] Figure 7 The structural schematic diagram of an electronic device 10 that can be used to implement the embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present application described herein and / or required by the present application.
[0179] As Figure 7 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0180] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless identification code map reconstruction transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0181] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the identification code map reconstruction method.
[0182] In some embodiments, the identification code map reconstruction method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the identification code map reconstruction method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the identification code map reconstruction method by any other suitable means (e.g., by means of firmware).
[0183] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0184] The computer program for implementing the method of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable identification code map reconstruction devices, such that when the computer program is executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0185] In the context of this application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0186] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0187] The systems and techniques described herein can be implemented in a computing system that includes backend components (such as, for example, a data server), or a computing system that includes middleware components (such as, for example, an application server), or a computing system that includes frontend components (such as, for example, a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (such as, for example, a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0188] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0189] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in this application can be executed in parallel, sequentially or in a different order, as long as the information expected by the technical solution of this application can be achieved, and this is not limited herein.
[0190] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of this application shall be included within the protection scope of this application.
Claims
1. A method for reconstructing an identification code map, characterized in that: The method comprises: An image set is obtained by performing image acquisition on a target scene through an extended reality device, and a coordinate conversion matrix corresponding to each identification code is determined according to the image coordinates and local three-dimensional coordinates of each identification code in the image set; wherein identification codes are distributed in the target scene, and the image set contains images of all identification codes in the target scene; Classifying the images in the image set that contain the same identification code into an image group, and for each image group, screening and retaining a preset number of frames of images; The image group is traversed, and for a target image in the currently traversed image group, the local three-dimensional coordinates of the feature points of each identification code are unified into the same coordinate system according to the coordinate transformation matrix corresponding to each identification code, so as to obtain the global three-dimensional coordinates of the feature points of each identification code, and the coordinate transformation matrix of the image in the currently traversed image group and the global three-dimensional coordinates obtained by traversing the current image group are optimized based on an optimization algorithm; wherein the target image is a frame of image selected from the image group; If the traversal of the image group is completed, the coordinate transformation matrices of the images in all the image groups and all the global three-dimensional coordinates obtained by traversing the image group are optimized based on the optimization algorithm, and the map reconstruction result is obtained based on the optimized global three-dimensional coordinates.
2. The method according to claim 1, characterized in that include: According to the image coordinates and local three-dimensional coordinates of each identification code in the image set, the coordinate transformation matrix corresponding to each identification code is determined, including: Performing identification code detection on each image in the image set to determine the image coordinates of feature points in each identification code in the image; The coordinate conversion matrix corresponding to each identification code is determined according to the image coordinates and the local three-dimensional coordinates of the feature points in the identification code.
3. The method according to claim 1, characterized in that The images in the image set containing the same identification code are grouped into an image group, comprising: Identify the identification code of each identification code in the image of the image set, store the identification code contained in each image into a set class, and set an image category code for the set class; The images in the image set are grouped, the image category code is used as a key index, the image identifier is used as a value, and the image identifiers with the same image category code obtained by indexing are grouped into one image group.
4. The method according to claim 1, characterized in that: For each image group, a preset number of frames are selected and retained, including: For each image in each image group, determining an evaluation score of the image according to size information of the identification code contained in the image and / or a reprojection error of a coordinate transformation matrix corresponding to the identification code; A preset number of images are selected from the image group and retained according to the evaluation scores.
5. The method according to claim 4, characterized in that Determining an evaluation score of the image according to size information of the identification code contained in the image and / or a reprojection error of a coordinate transformation matrix corresponding to the identification code includes: Determine an image area mean of the identification code in the image and / or a reprojection error mean of a coordinate transformation matrix corresponding to the identification code; An evaluation score of the image is determined according to the image area mean and / or the reprojection error mean; wherein the image evaluation score is positively correlated with the image area mean and negatively correlated with the reprojection error mean.
6. The method according to claim 1, characterized in that According to the coordinate conversion matrix corresponding to each identification code, the local three-dimensional coordinates of the feature points of each identification code are unified into the same coordinate system to obtain the global three-dimensional coordinates of the feature points of each identification code, including: Selecting a target identification code from the identification codes included in the target image, and using the local three-dimensional coordinates of the feature points of the target identification code as the global three-dimensional coordinates of the feature points of the target identification code; Based on the origin in the global three-dimensional coordinates of the feature points of the target identification code as a reference, the global three-dimensional coordinates of the feature points of other identification codes are determined according to the coordinate transformation matrix of the target identification code, the coordinate transformation matrix of other identification codes except the target identification code, and the local three-dimensional coordinates of the feature points of other identification codes.
7. The method according to claim 6, characterized in that Selecting a target identification code from the identification codes contained in the target image comprises: Obtaining a reprojection error corresponding to an output in the process of determining a coordinate transformation matrix corresponding to each identification code; The identification code corresponding to the coordinate transformation matrix with the smallest reprojection error is used as the target identification code.
8. The method according to claim 6, characterized in that Determining the global three-dimensional coordinates of the feature points of other identification codes according to the coordinate conversion matrix corresponding to the target identification code, the coordinate conversion matrices corresponding to other identification codes except the target identification code, and the local three-dimensional coordinates of the feature points of other identification codes includes: The product of the inverse matrix of the coordinate transformation matrix corresponding to the target identification code, the coordinate transformation matrix corresponding to the other identification codes and the local three-dimensional coordinates of the feature points of the other identification codes is used as the global three-dimensional coordinates of the feature points of the other identification codes.
9. The method according to claim 1, characterized in that: Based on the optimization algorithm, the coordinate transformation matrix of the image in the current traversed image group and the global three-dimensional coordinates obtained by traversing the current image group are optimized, including: For each image in the currently traversed image group, the coordinate transformation matrix with the smallest reprojection error among the coordinate transformation matrices corresponding to the identification codes in the image is used as the coordinate transformation matrix of the image; The coordinate transformation matrix of each image in the image group and the global three-dimensional coordinates obtained by traversing the current image group are input into a bundle adjustment algorithm to optimize the global three-dimensional coordinates.
10. The method according to claim 1, characterized in that The coordinate transformation matrix of all images in the image group and all global three-dimensional coordinates obtained by traversing the image group are optimized based on the optimization algorithm, including: For each image in each image group, the coordinate transformation matrix with the smallest reprojection error among the coordinate transformation matrices corresponding to each identification code in the image is used as the coordinate transformation matrix of the image; The coordinate transformation matrix of each image in the entire image group and all global three-dimensional coordinates are input into the bundle adjustment algorithm to optimize the global three-dimensional coordinates.
11. The method according to claim 1, characterized in that Traversing the image group, including: The image group containing the most identification codes is taken as the image group obtained by the first traversal; In the process of selecting the image group for the next traversal from the remaining image groups, the image group whose identification codes overlap the most with the identification codes of the currently determined global three-dimensional coordinates is selected as the image group for the next traversal.
12. An identification code map reconstruction device, characterized in that: The device comprises: A coordinate conversion matrix determination module is used to acquire an image set by performing image acquisition on a target scene through an extended reality device, and determine a coordinate conversion matrix corresponding to each identification code according to the image coordinates and local three-dimensional coordinates of each identification code in the image set; wherein the identification codes are distributed in the target scene, and the image set contains images of all the identification codes in the target scene; An image group determination module, used to classify the images in the image set containing the same identification code into one image group, and for each image group, screen and retain a preset number of frames of images; A local optimization module, used to traverse the image group, and for a target image in the currently traversed image group, unify the local three-dimensional coordinates of the feature points of each identification code into the same coordinate system according to the coordinate transformation matrix corresponding to each identification code, obtain the global three-dimensional coordinates of the feature points of each identification code, and optimize the coordinate transformation matrix of the image in the currently traversed image group and the global three-dimensional coordinates obtained by traversing the current image group based on the optimization algorithm; wherein the target image is a frame of image selected from the image group; The global optimization module is used to optimize the coordinate transformation matrix of all images in the image group and all global three-dimensional coordinates obtained by traversing the image group based on the optimization algorithm when the image group traversal is completed, and obtain the map reconstruction result based on the optimized global three-dimensional coordinates.
13. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the identification code map reconstruction method described in any one of claims 1-11.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the identification code map reconstruction method described in any one of claims 1-11 when executed.