Method and apparatus for determining the transformation relationship between the coordinate system of a 3D scene model and the physical coordinate system.
By selecting feature points from multiple 2D images and using camera position information to automatically calculate the transformation relationship between the 3D scene model and the physical coordinate system, the tedious and time-consuming problem in the existing technology is solved, and automatic registration between the 3D scene model and the physical world is realized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-29
- Publication Date
- 2026-04-03
AI Technical Summary
In existing technologies, determining the transformation relationship between the coordinate system of a 3D scene model and the physical coordinate system mainly relies on manual annotation and surveying, which is a tedious and time-consuming process.
By selecting feature points from multiple 2D images and utilizing the camera's position information in the model coordinate system and the physical coordinate system, the transformation relationship between the model coordinate system and the physical coordinate system is automatically calculated, including rotation, translation, and scaling parameters.
It achieves automatic registration between 3D scene models and the physical world, simplifies the coordinate system transformation process, and improves efficiency and accuracy.
Smart Images

Figure CN114693782B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision, and more particularly to a method, electronic device, and storage medium for converting between a 3D scene model coordinate system and a physical coordinate system. Background Technology
[0002] With the development of technology, 3D scene models have been widely used in the field of computer vision. In some applications using computer vision for localization or navigation, real-time localization and navigation of the moving target in the 3D scene are often achieved by comparing and matching 2D images of the moving target (e.g., a robot equipped with a camera) with a pre-established 3D scene model. The 3D scene model itself has its own coordinate system (hereinafter referred to as the model coordinate system), and the physical world or real-world scene corresponding to the 3D scene model also has a coordinate system (hereinafter referred to as the physical coordinate system). The transformation relationship between these two coordinate systems (e.g., translation, scaling, and rotation) must be obtained to achieve the localization or navigation of the moving target. Currently, the transformation relationship between the model coordinate system and the physical coordinate system is basically determined by manual annotation and surveying, which is tedious and time-consuming. Summary of the Invention
[0003] The present invention provides a method, electronic device and medium for determining the transformation relationship between the coordinate system of a 3D scene model and the physical coordinate system. It can automatically calculate the transformation relationship between the model coordinate system and the physical coordinate system, that is, realize the automatic registration between the 3D scene model and the physical world.
[0004] The above objective is achieved through the following technical solution:
[0005] According to a first aspect of the present invention, a method for determining the transformation relationship between a model coordinate system and a physical coordinate system of a three-dimensional scene is provided, comprising: for a three-dimensional scene model constructed based on multiple two-dimensional images, selecting at least three two-dimensional images from the multiple two-dimensional images and obtaining the position information of the camera corresponding to each selected two-dimensional image in the physical coordinate system; for each selected two-dimensional image, determining the position information of the camera corresponding to it in the model coordinate system; and determining the transformation relationship between the model coordinate system and the physical coordinate system based on the determined position information of the camera corresponding to each two-dimensional image in the model coordinate system and the physical coordinate system.
[0006] In some embodiments of the present invention, determining the position information of the camera corresponding to each selected two-dimensional image in the model coordinate system may include: selecting at least four feature points from the two-dimensional image and determining the model coordinates of each feature point in the model coordinate system and the pixel coordinates of the feature points in the two-dimensional image; and calculating the position information of the camera corresponding to the two-dimensional image in the model coordinate system based on the model coordinates and pixel coordinates of each selected feature point.
[0007] In some embodiments of the present invention, the camera positions corresponding to the selected two-dimensional images are not collinear. In some embodiments, the camera positions corresponding to the selected two-dimensional images are not coplanar. In some embodiments, the camera positions corresponding to the selected two-dimensional images are neither collinear nor coplanar.
[0008] In some embodiments of the present invention, determining the transformation relationship between the model coordinate system and the physical coordinate system may include determining the rotation, translation, and scaling parameters between the model coordinate system and the physical coordinate system based on the position information of the camera corresponding to each determined two-dimensional image in the model coordinate system and the physical coordinate system.
[0009] In some embodiments of the present invention, obtaining the position information of the camera corresponding to each of the selected two-dimensional images in the physical coordinate system may include determining the position information of the camera corresponding to each of the selected two-dimensional images in the physical coordinate system from the position information of the camera in the physical coordinate system of the three-dimensional scene when each two-dimensional image was captured, which was pre-marked and stored during the process of constructing a three-dimensional scene model based on multiple two-dimensional images.
[0010] According to a second aspect of the present invention, a system for determining the transformation relationship between a model coordinate system and a physical coordinate system of a three-dimensional scene is also provided, comprising: a physical coordinate determination module, configured to select at least three two-dimensional images from the plurality of two-dimensional images for a three-dimensional scene model constructed based on a plurality of two-dimensional images and obtain the position information of the camera corresponding to each of the selected two-dimensional images in the physical coordinate system; a model coordinate determination module, configured to determine the position information of the camera corresponding to each selected two-dimensional image in the model coordinate system; and a transformation relationship determination module, configured to determine the transformation relationship between the model coordinate system and the physical coordinate system based on the determined position information of the camera corresponding to each two-dimensional image in the model coordinate system and the physical coordinate system.
[0011] According to a third aspect of the present invention, an electronic device is provided, including a processor and a memory, wherein the memory stores a computer program that, when executed by the processor, can be used to implement the method described in the first aspect of the present invention.
[0012] According to a fourth aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored thereon, which, when executed by a processor, is capable of implementing the method described in the first aspect of the present invention.
[0013] Compared with the prior art, the solution of the embodiments of this application utilizes the correspondence between the position information of the camera in the physical coordinate system and the position information in the model coordinate system of the multiple two-dimensional images used to construct the three-dimensional scene model to determine the transformation relationship between the model coordinate system and the physical coordinate system, so as to facilitate the automatic registration of the three-dimensional scene model and the physical world quickly. Attached Figure Description
[0014] The embodiments of the present invention will be further described below with reference to the accompanying drawings, wherein:
[0015] Figure 1 A flowchart illustrating a method for determining the transformation relationship between the model coordinate system and the physical coordinate system of a three-dimensional scene according to an embodiment of the present invention is shown.
[0016] Figure 2 This is a schematic diagram showing the camera poses corresponding to the 2D images marked in the example 3D scene;
[0017] Figure 3 A functional block diagram of an apparatus for determining the transformation relationship between the model coordinate system and the physical coordinate system of a three-dimensional scene according to an embodiment of the present invention is shown. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0019] In the field of computer vision, 3D scene models are typically constructed based on multiple 2D images of the scene. One method for reconstructing a 3D scene model using 2D images involves using a dedicated camera capable of precisely calibrating its own position and pose (hereinafter collectively referred to as pose) to capture a large number of 2D images of the scene and recording the pose information of the dedicated camera when each image is captured. Then, the images can be spatially sorted based on this pose information, and a 3D scene reconstruction can be performed. The inventors of this application also disclosed a scene reconstruction method based on 2D images in another patent application No. 202010758289.9. This method involves pre-arranging at least one visual marker (e.g., an optical communication device, a QR code, a graphic symbol, etc.) in the scene to be reconstructed. By capturing scene images including the visual marker and analyzing these images, the pose information of the camera relative to the visual marker is determined. Then, based on the pose information of the visual marker in the physical coordinate system, the actual pose information of the camera when capturing the scene image is determined. Thus, based on the camera pose information associated with each scene image, multiple scene images are spatially sorted, and a 3D scene model is built based on the sorted scene images.
[0020] It can be seen that when using two-dimensional images to construct a three-dimensional scene model, it is often necessary to determine the pose information of the camera in the physical coordinate system of the scene when each two-dimensional image is captured. The inventors have discovered that such information can not only be used to spatially sort the various two-dimensional images during the construction of the three-dimensional scene model, but also to register the three-dimensional scene model with the physical world after the three-dimensional scene model is constructed.
[0021] Figure 1 A flowchart illustrating a method for determining the transformation relationship between a 3D scene model coordinate system and a physical coordinate system according to an embodiment of the present invention is provided. Figure 1 As shown, in step 101, for a 3D scene model constructed using multiple 2D images, at least three 2D images are selected from the multiple 2D images, and the position information of the camera corresponding to each of the selected 2D images in the physical coordinate system is obtained. As mentioned above, the pose information of the camera in the physical coordinate system when capturing each 2D image has already been marked during the process of constructing the 3D scene model using multiple 2D images. In one embodiment, for the selected at least three 2D images, it is required that the positions of the cameras corresponding to each of them are not collinear. In another embodiment, the positions of the cameras corresponding to each of the selected 2D images are neither collinear nor coplanar.
[0022] Next, in step 102, for each selected 2D image, the position information of the corresponding camera in the model coordinate system is determined. For example, at least four feature points can be selected from the 2D image, and the model coordinates of these feature points in the model coordinate system of the 3D scene and the imaging positions (i.e., pixel coordinates) of these feature points in the 2D image are determined. The model coordinates of these feature points can be obtained through the established 3D scene model, and the pixel coordinates of these feature points are, for example, the positions of these feature points in a 2D coordinate system established with the top-left corner of the image as the origin. Then, the pose information of the camera corresponding to the 2D image is calculated based on the model coordinates and pixel coordinates of the selected feature points. For example, the PnP (Perspective-n-Point) algorithm can be used to solve for the camera pose when the image was captured, given that the spatial positions of multiple points and their imaging positions in the image are known. Optionally, the BA (Bundle Adjustment) optimization algorithm can also be used simultaneously to obtain a more accurate camera pose. In fact, for a 3D scene model created based on multiple 2D images, the above algorithm can be used to mark the position and pose of the camera corresponding to each 2D image in the model coordinate system of the 3D scene model. For example, as... Figure 2 The three-dimensional scene model shown has pyramid-shaped markers used to represent the position and orientation of the camera corresponding to each two-dimensional image in the scene's model coordinate system.
[0023] Continue to refer to Figure 1 In step 103, based on the position information of the cameras corresponding to each selected 2D image in the model coordinate system and the physical coordinate system of the 3D scene, the transformation relationship between the model coordinate system and the physical coordinate system is determined. That is, the transformation relationship between the two coordinate systems is determined using the corresponding positional relationship of the cameras corresponding to each 2D image in the two coordinate systems.
[0024] In 3D space, the transformation relationship between two coordinate systems includes translation, scaling, and rotation, each represented by three parameters, for a total of nine parameters. That is, the transformation relationship between two coordinate systems in 3D space can be represented by three translation parameters, three scaling parameters, and three rotation parameters. Typically, to maintain consistency with the real scene, 3D scene models have the same scaling ratio on all three axes, resulting in seven parameters (three translation parameters, one scaling parameter, and three rotation parameters). This requires the corresponding positional relationships of at least three points in the two coordinate systems (they cannot be collinear) to solve for the transformation coefficients (i.e., at least three 2D images are needed, and the corresponding camera positions are not collinear). When the scaling ratios are different, at least four corresponding positional relationships are required, and these points cannot be coplanar (i.e., at least four 2D images are needed, and the corresponding camera positions are not coplanar). How to use the known coordinate correspondence of at least three points in the two coordinate systems to solve for the rotation, translation, and scaling parameters in the affine transformation matrix formula between the two coordinate systems is existing technology and will not be elaborated upon here.
[0025] Through the aforementioned embodiments, for a three-dimensional scene model constructed using multiple two-dimensional images, the transformation relationship between the model coordinate system and the physical coordinate system of the three-dimensional scene model can be determined automatically and quickly, realizing automatic registration between the model coordinate system and the physical coordinate system. This is beneficial to the promotion and popularization of positioning and navigation applications in three-dimensional scenes.
[0026] Figure 3 This is a functional block diagram of an apparatus 300 for determining the transformation relationship between a 3D scene model coordinate system and a physical coordinate system according to an embodiment of the present invention. Although the components are described in a functionally separate manner in this block diagram, such description is for illustrative purposes only. The components shown in the figure can be arbitrarily combined or divided into independent software, firmware, and / or hardware components. Moreover, regardless of how such components are combined or divided, they can be executed on the same computing device or distributed across multiple computing devices, wherein the multiple computing devices may be connected by one or more networks.
[0027] like Figure 3As shown, the device 300 includes a physical coordinate determination module 301, a model coordinate determination module 302, and a transformation relationship determination module 303. The physical coordinate module 301, as described above in conjunction with step S101, selects at least three 2D images from the plurality of 2D images for a 3D scene model constructed based on multiple 2D images and obtains the position information of the camera corresponding to each selected 2D image in the physical coordinate system. The model coordinate determination module 302, as described above in conjunction with step S102, determines the position information of the corresponding camera in the model coordinate system for each selected 2D image. The transformation relationship determination module 303, as described above in conjunction with step S103, determines the transformation relationship between the model coordinate system and the physical coordinate system based on the determined position information of the camera corresponding to each 2D image in both the model coordinate system and the physical coordinate system.
[0028] In one embodiment of the present invention, the invention can be implemented in the form of a computer program. The computer program can be stored in various computer-readable storage media (e.g., hard disk, optical disk, flash memory, etc.), and when the computer program is executed by a processor, it can be used to implement the method of the present invention.
[0029] In another embodiment of the invention, the invention can be implemented as an electronic device. This electronic device includes a processor and a memory, in which a computer program is stored. When executed by the processor, the computer program can be used to implement the method of the invention.
[0030] References to “various embodiments,” “some embodiments,” “one embodiment,” or “embodiment” throughout this document refer to a particular feature, structure, or property described in connection with said embodiment that is included in at least one embodiment. Therefore, the appearance of the phrases “in various embodiments,” “in some embodiments,” “in one embodiment,” or “in an embodiment” throughout this document does not necessarily refer to the same embodiment. Furthermore, a particular feature, structure, or property may be combined in any suitable manner in one or more embodiments. Thus, a particular feature, structure, or property shown or described in connection with one embodiment may be combined, in whole or in part, with features, structures, or properties of one or more other embodiments without limitation, provided that such combination is not illogical or inoperable. Expressions such as “according to A,” “based on A,” “through A,” or “using A” appearing throughout this document are non-exclusive; that is, “according to A” may cover “according to A only” or “according to A and B”, unless specifically stated or clearly understood from the context to mean “according to A only.” For clarity, this application describes some illustrative operational steps in a certain order. However, those skilled in the art will understand that each of these operational steps is not essential, and some steps may be omitted or replaced by other steps. These operational steps also do not necessarily have to be performed sequentially as shown. Instead, some of these operational steps can be performed in different orders or in parallel as needed, provided that the new execution method is not illogical or inoperable.
[0031] This has described several aspects of at least one embodiment of the invention. It will be understood that various changes, modifications, and improvements will be readily apparent to those skilled in the art. Such changes, modifications, and improvements are intended within the spirit and scope of the invention. While the invention has been described through some embodiments, it is not limited to the embodiments described herein, and includes various changes and variations made without departing from the scope of the invention.
Claims
1. A method for determining the transformation relationship between the model coordinate system and the physical coordinate system of a 3D scene, comprising: For a 3D scene model constructed based on multiple 2D images, at least three 2D images are selected from the multiple 2D images and the position information of the camera corresponding to each of the selected 2D images in the physical coordinate system of the 3D scene is obtained, wherein the positions of the cameras corresponding to the selected 2D images are not collinear. For each selected 2D image, determine the position information of its corresponding camera in the model coordinate system of the 3D scene; Based on the position information of the camera corresponding to each two-dimensional image in the model coordinate system and the physical coordinate system of the three-dimensional scene, the transformation relationship between the model coordinate system and the physical coordinate system of the three-dimensional scene is determined.
2. The method according to claim 1, wherein determining the position information of the corresponding camera in the model coordinate system for each selected two-dimensional image includes: Select at least four feature points from the two-dimensional image, and determine the model coordinates of each feature point in the model coordinate system and the pixel coordinates of the feature points in the two-dimensional image; The position information of the camera in the model coordinate system corresponding to the selected feature points is calculated based on the model coordinates and pixel coordinates of each selected feature point.
3. The method according to claim 1, wherein the at least three two-dimensional images comprise at least four two-dimensional images, wherein the camera positions corresponding to the selected two-dimensional images are not coplanar.
4. The method according to any one of claims 1-3, wherein determining the transformation relationship between the model coordinate system and the physical coordinate system comprises: Based on the position information of the camera corresponding to each determined two-dimensional image in the model coordinate system and the physical coordinate system, the rotation, translation and scaling parameters between the model coordinate system and the physical coordinate system are determined.
5. The method according to any one of claims 1-3, wherein obtaining the position information of the camera corresponding to each of the selected two-dimensional images in the physical coordinate system includes: The position information of the camera in the physical coordinate system corresponding to each selected 2D image is determined by the position information of the camera in the physical coordinate system when each 2D image is captured, which is pre-marked and stored during the process of constructing a 3D scene model based on multiple 2D images.
6. A system for determining the transformation relationship between the model coordinate system and the physical coordinate system of a 3D scene, comprising: The physical coordinate determination module is used to select at least three two-dimensional images from the multiple two-dimensional images for a three-dimensional scene model constructed based on multiple two-dimensional images and obtain the position information of the camera corresponding to each selected two-dimensional image in the physical coordinate system of the three-dimensional scene, wherein the positions of the cameras corresponding to the selected two-dimensional images are not collinear. The model coordinate determination module is used to determine the position information of the corresponding camera in the model coordinate system of the 3D scene for each selected 2D image; The transformation relationship determination module is used to determine the transformation relationship between the model coordinate system and the physical coordinate system of the three-dimensional scene based on the position information of the camera corresponding to each two-dimensional image in the model coordinate system and the physical coordinate system of the three-dimensional scene.
7. A computer-readable storage medium storing a computer program that, when executed by a processor, can be used to implement the method of any one of claims 1-5.
8. An electronic device comprising a processor and a memory, the memory storing a computer program that, when executed by the processor, is capable of implementing the method of any one of claims 1-5.
Citation Information
Patent Citations
Spatial positioning method and device, equipment, storage medium and navigation stick
CN111821025A
Surgical navigation system precision detection method
CN112006779A