Method and system for ar-assisted three-dimensional model scale recovery

CN115601496BActive Publication Date: 2026-08-21HANGZHOU YIXIAN XIANJIN TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210994762.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-18
Publication Date
2026-08-21
Estimated Expiration
2042-08-18

AI Technical Summary

Technical Problem

[0006]本申请实施例提供了一种AR辅助三维模型尺度恢复的方法、系统、计算机设备和计算机可读存储介质,以至少解决相关技术中三维模型尺度恢复方法应用便捷性较差的问题

Benefits of technology

[0017]相比于相关技术,本申请实施例提供的AR辅助三维模型尺度恢复的方法,首先获取目标场景数据,其次,对经过预处理的三维模型,将其与图像在虚拟相机中叠加显示,通过虚拟相机进行对齐处理,将三维模型与图像中的场景角度完全重合,并记录模型坐标系与相机坐标系之间的第一转换矩阵,以及,根据第一转换矩阵确定图像在模型坐标系下的位置坐标;进一步的,基于该位置坐标和图像中各对象的真实物理位置,确定世界坐标系对模型坐标系的第二转换矩阵,进而基于第二转换矩阵中的尺度缩放信息,将三维模型恢复至真实尺寸。通过本申请,解决了三维模型真实尺度恢复便捷性较差的问题,无需标定装置也不限制于应用场景,用户通过一台终端设备,利用AR作为辅助,即可通过交互操作实现三维模型的真实尺度恢复。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115601496B_ABST
    Figure CN115601496B_ABST
Patent Text Reader

Abstract

The application relates to a method for AR-assisted scale recovery of a three-dimensional model, wherein the method comprises: acquiring target scene data; after preprocessing the three-dimensional model, superimposing the three-dimensional model and an image in a virtual camera; aligning the three-dimensional model and the image according to an interactive instruction input by a user through an AR editor, recording a first conversion matrix between a model coordinate system and a camera coordinate system, and determining position coordinates of the image in the model coordinate system according to the first conversion matrix; determining a second conversion matrix between a world coordinate system and the model coordinate system based on optimized position coordinates and a real physical position of an object in the image, wherein the second conversion matrix contains scale zoom information; and recovering the three-dimensional model to a real size based on the scale zoom information. Through the application, the problem of poor convenience of scale recovery of a three-dimensional model is solved, a calibration device is not required, and the application scene is not limited; a user can realize real scale recovery of a three-dimensional model through interactive operation with the aid of AR.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of 3D map reconstruction, and in particular to a method and system for using AR-assisted 3D model scale recovery. Background Technology

[0002] 3D reconstruction refers to the creation of a 3D mathematical model of a 3D object (or scene) that is suitable for computer representation and processing. It is a key technology for creating virtual reality that expresses the objective world in a computer.

[0003] In the process of 3D model reconstruction, restoring the true scale of the scene and objects is an important step. Related technologies utilize the following methods to restore the scale of 3D models: 1: By placing a calibration device of a specific size in space, after taking an image containing the calibration device, the scale of the entire scene can be recovered using the known size of the calibration device; 2: Use binocular or multi-camera cameras to capture images simultaneously to obtain depth information, and then use this depth information to restore the scale of the 3D map; 3: Use measuring tools such as rulers to actually measure the real size of the target scene, and then perform scale restoration of the 3D model; 4. Utilize multiple sensors capable of transmitting and receiving GPS signals to obtain actual size information, and then reconstruct a 3D map; 5: Acquire on-site images by using cameras mounted on mobile devices (such as unmanned vehicles). Since the mounting height of the camera on the mobile device is known, this height information can be used to reconstruct the size of the 3D map.

[0004] However, various problems still exist in the above methods: Method 1 requires the creation of a calibration device that matches the target scene. In addition, in practical applications, this calibration device is only suitable for scenes that match its size, which has many limitations. Method 2 requires the use of multiple cameras to take pictures simultaneously, and the relative poses between the multiple cameras need to be calibrated; Method 3 is only applicable to scenarios where dimensions are easy to measure (e.g., tall buildings). Methods 4 and 5, however, have limitations in their application scenarios.

[0005] Currently, no effective solution has been proposed to address the issue of the poor convenience of methods for restoring the true scale of 3D models in related technologies. Summary of the Invention

[0006] This application provides an AR-assisted method, system, computer device, and computer-readable storage medium for 3D model scale recovery, to at least solve the problem of poor application convenience of 3D model scale recovery methods in related technologies.

[0007] In a first aspect, embodiments of this application provide an AR-assisted method for scale recovery of a three-dimensional model, the method comprising: Acquire target scene data, wherein the target scene data includes: a 3D model, an image, and camera parameters and pose information when the image was captured; After preprocessing the 3D model, it is overlaid and displayed on the image in a virtual camera, wherein the virtual camera is constructed based on the camera parameters; Based on the interactive commands input by the user through the AR editor, the 3D model and the image are aligned, and the first transformation matrix between the model coordinate system and the camera coordinate system is recorded. Based on the first transformation matrix, the position coordinates of the image in the model coordinate system are determined. The position coordinates are optimized, and based on the optimized position coordinates and the actual physical position of the object in the image, a second transformation matrix between the world coordinate system and the model coordinate system is determined, wherein the second transformation matrix contains scale scaling information; Based on the scale information, the 3D model is restored to its true size.

[0008] In some embodiments, preprocessing the 3D model includes: Determine the direction of the gravity axis of the three-dimensional model, and rotate the three-dimensional model so that its gravity axis coincides with any coordinate axis of the three-dimensional coordinate system; Construct the bounding box of the 3D model and scale the shape of the 3D model to a size close to its real size.

[0009] In some embodiments, aligning the 3D model with the image includes: The initialization process includes: aligning the top and bottom edges of the bounding box of the 3D model with the top and bottom edges of the field of view of the virtual camera, and rotating the 3D model so that its Z-axis is parallel to the gravity axis of the image coordinate system. The position alignment process includes: using the virtual camera, moving the 3D model in the front-back depth direction and the up-down direction to align its position with the image; The rotation alignment process includes: rotating the 3D model around the gravity axis using the virtual camera so that it completely coincides with the scene angle in the image.

[0010] In some embodiments, at least one set of images is aligned with a 3D model in the virtual camera to obtain at least one set of alignment results.

[0011] In some embodiments, the position coordinates are optimized, including: Based on the first transformation matrix, calculate the 2D projection coordinates of each vertex in the 3D model in the image, and determine the association information between the 2D projection coordinates and each vertex; The location coordinates are optimized based on the associated information, and the reliability of the optimized location coordinates is determined.

[0012] In some embodiments, optimizing the location coordinates based on the association information includes: Acquire multiple sets of related information, as well as multiple sets of images of the target scene taken from different angles; Feature points are extracted from the multiple sets of images, and the images are matched based on the feature points to obtain the 2D association relationship between the multiple sets of images; Based on the 2D association relationship and the position coordinates, triangulation is performed to obtain the 3D position of the feature point; Based on the 3D position and the parameters of the virtual camera, the pose of the image and the 3D points in the image are optimized using the SFM algorithm to obtain the optimized image position coordinates.

[0013] In some embodiments, determining the reliability of the optimized position coordinates includes: Based on the optimized position coordinates, the 2D projection coordinates of each vertex in the 3D model in the current image are recalculated; Determine whether the deviation between the recalculated 2D projection coordinates and the 2D projection coordinates before optimization is greater than a preset deviation threshold. If so, output an unreliable optimization result response. If not, based on the actual physical location of the object in the image and the location coordinates, determine the second transformation matrix between the world coordinate system and the model coordinate system.

[0014] Secondly, embodiments of this application provide an AR-assisted 3D model scale recovery system, the system comprising: an acquisition module, a preprocessing module, an alignment module, and a reconstruction and recovery module, wherein... The acquisition module is used to acquire target scene data, wherein the target scene data includes: a 3D model, an image, and camera parameters and pose information when the image was captured; The preprocessing module is used to preprocess the 3D model and then overlay it with the image in a virtual camera, wherein the virtual camera is constructed based on the camera parameters; The alignment module is used to align the 3D model with the image according to the interactive instructions input by the user through the AR editor, record the first transformation matrix between the model coordinate system and the camera coordinate system, and determine the position coordinates of the image in the model coordinate system according to the first transformation matrix. The reconstruction and restoration module is used to optimize the position coordinates, determine a second transformation matrix between the world coordinate system and the model coordinate system based on the optimized position coordinates and the actual physical position of the object in the image, wherein the second transformation matrix contains scale scaling information, and restore the three-dimensional model to its true size based on the scale scaling information.

[0015] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in the first aspect above.

[0016] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect above.

[0017] Compared to related technologies, the AR-assisted 3D model scale restoration method provided in this application first acquires target scene data. Then, it overlays the pre-processed 3D model with an image in a virtual camera, aligning the 3D model with the scene angle in the image using the virtual camera. A first transformation matrix between the model coordinate system and the camera coordinate system is recorded, and the image's position coordinates in the model coordinate system are determined based on the first transformation matrix. Further, based on these position coordinates and the actual physical positions of objects in the image, a second transformation matrix between the world coordinate system and the model coordinate system is determined. Finally, based on the scale information in the second transformation matrix, the 3D model is restored to its true size. This application solves the problem of poor convenience in restoring the true scale of 3D models. It requires no calibration device and is not limited to any particular application scenario. Users can achieve true scale restoration of 3D models through interactive operation using a terminal device and AR as an aid. Attached Figure Description

[0018] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram of the pose of a terminal during movement according to an embodiment of the present invention. Figure 2This is a schematic diagram illustrating the application environment of an AR-assisted 3D model scale recovery method according to an embodiment of this application; Figure 3 This is a flowchart of an AR-assisted 3D model scale recovery method according to an embodiment of this application; Figure 4 This is a schematic diagram illustrating the alignment of the gravity axis and Z-axis of a three-dimensional model according to an embodiment of this application. Figure 5 This is a schematic diagram illustrating the overlay display of a three-dimensional model and a real image according to an embodiment of this application; Figure 6 This is a schematic diagram of rotating a model on the gravity axis of an image according to an embodiment of this application; Figure 7 This is a structural block diagram of an AR-assisted 3D model scale recovery system according to an embodiment of this application; Figure 8 A schematic diagram of the internal structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.

[0020] Obviously, the accompanying drawings described below are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar scenarios based on these drawings without any inventive effort. Furthermore, it is understood that although the efforts made in this development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, any changes to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.

[0021] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0022] In this document, it should be understood that the terms used may be technical means used to implement part of this invention or other summary technical terms. For example, terms may include: Target scene: refers to the scene to be scaled back, which can be, but is not limited to, various types of 3D scenes such as objects, rooms, buildings, indoor spaces, and blocks.

[0023] A 3D model is a polygonal representation of an object (scene), typically displayed using a computer or other video equipment. In this embodiment, the 3D model can be a virtual model of an object, an offline building, or an offline scene, usually represented as a point cloud, vector, etc.

[0024] Real scale: refers to the actual size of a scene (or object) in the real world. The goal of this application is to obtain a 3D model with real scale from a known 3D model that has no real scale through a series of operations.

[0025] Position: Given a three-dimensional coordinate system (Cartesian coordinate system), the position of the target in that coordinate system, usually represented by (x, y, z).

[0026] Pose: Position and orientation (facing), for example, in two dimensions it is generally (x, y, yaw), and in three dimensions it is generally (x, y, z, yaw, pitch, roll). The last three elements describe the object's orientation. Figure 1 This is a schematic diagram of the pose of a terminal during movement according to an embodiment of the present invention. Figure 1 As shown, yaw is the heading angle, rotating around the Z-axis; pitch is the pitch angle, rotating around the Y-axis; and roll is the roll angle, rotating around the X-axis.

[0027] Augmented Reality (AR): A technology that uses precise calculations of the position and angle of camera images, combined with image analysis techniques, to allow the virtual world on the screen to combine and interact with real-world scenes.

[0028] AR Editor: In this embodiment, AR editor refers to a software developed in the AR-related module to complete the corresponding operations. The AR editor can be developed independently or based on certain engines, and its form includes, but is not limited to, software, web pages, mobile applications, etc.

[0029] Virtual camera: refers to the window in the AR editor for observing the virtual world. For example, Unity requires at least one camera, and multiple cameras can be used to observe the world from different perspectives. After adding scripts, you can perform operations on the camera, such as rotation, translation, etc.

[0030] This application provides an AR-assisted 3D model scale recovery method, which can be applied to applications such as... Figure 2 In the application environment shown, Figure 2 This is a schematic diagram illustrating the application environment of an AR-assisted 3D model scale recovery method according to an embodiment of this application, such as... Figure 2 As shown, the user captures image data of the target scene (object) through the terminal. Further, the terminal overlays the aforementioned real image onto a virtual 3D model. The user then aligns the image with the 3D model using an AR editor. After alignment, the terminal runs a preset processing flow to obtain a transformation matrix between the world coordinate system of the object in the image and the model coordinate system. This transformation matrix reflects scale information, which can then be used to restore the 3D model to its true scale.

[0031] It should be noted that in this application, the aforementioned terminal can be a computer device such as a smartphone, tablet computer, or PC, the three-dimensional model can be a three-dimensional model of an object, indoor space, building, or street, and the device for capturing images includes, but is not limited to, a mobile phone, camera, or drone.

[0032] Figure 3 This is a flowchart of an AR-assisted 3D model scale recovery method according to an embodiment of this application, such as... Figure 3 As shown, the process includes the following steps: S301, acquire target scene data, which includes: 3D model, image, and camera parameters and pose information when the image was captured; The above data can be obtained, but is not limited to, through smartphones, cameras, and other terminal devices. Camera parameters can be provided by any camera-capable terminal device; image and pose information can be obtained for different targets using different methods, for example: Object class: The user holds the phone and takes pictures around the target object. During the shooting process, the phone runs a tracking algorithm (VIO), and each picture will have pose information (6dof pose) at a real physical scale. Large-scale construction projects: 1. The mobile phone takes pictures of the building from different angles and records the sensor information of the mobile phone, including GPS and orientation sensor. By combining this sensor information, the mobile phone's (6DOF pose) can be obtained. 2. When drones or other devices equipped with GPS take photos, the images they capture include POS data. This also includes GPS data and IMU data, specifically the exterior orientation elements from oblique photogrammetry: (latitude, longitude, elevation, heading (Phi), pitch (Omega), and roll (Kappa)). Furthermore, when acquiring target scene data, the real physical location information of the image can be directly obtained, which is also the scale information source in this application; however, it should be noted that the real physical location information obtained here is usually a coarse value.

[0033] Optionally, in some cases, a 3D model of the scene can be obtained through methods such as SFM (struct-from-motion), and then the 3D model can be mapped to these coarse physical location information to provide a relatively better initial value for subsequent processing.

[0034] S302, After preprocessing the 3D model, it is overlaid with the image in a virtual camera and displayed. The virtual camera is constructed based on the camera parameters mentioned above. The purpose of preprocessing is to perform preliminary alignment of the 3D model with the real environment, including alignment of shape outline and aspect ratio, so as to obtain a 3D model that is relatively close to the real scale.

[0035] Furthermore, the aforementioned AR editor is software developed to perform corresponding AR editing operations. Through this AR editor, the aforementioned image can be read and displayed; and a virtual camera can be constructed based on the camera parameters used when the image was captured. A 3D model can be placed within the virtual camera, and the real graphics can be overlaid onto the virtual 3D model. Users can manipulate the 3D model through the AR editor to rotate and align its position.

[0036] It should be noted that the definition and function of "virtual camera" have already been given in the terminology explanation section above. The type of virtual camera may differ depending on the engine environment, but its function remains the same. This embodiment will not elaborate on how to construct a virtual camera based on camera parameters.

[0037] S303, based on the interactive instructions input by the user through the AR editor, aligns the 3D model with the image, records the first transformation matrix between the model coordinate system and the camera coordinate system, and determines the position coordinates of the image in the model coordinate system based on the first transformation matrix; The alignment process described above can be performed by the user on the terminal's visual interface. The user can rotate the 3D model and move it freely in the depth and vertical directions using touch and other means, so that the image and the 3D model completely overlap.

[0038] Furthermore, the aforementioned camera coordinate system is the coordinate system of the virtual camera, which remains unchanged during the alignment process and can be directly obtained from the application; while the model coordinate system can be determined based on the accumulated movement and rotation paths during the alignment process.

[0039] Furthermore, after obtaining the first transformation matrix between the two coordinate systems, since this first transformation matrix reflects the transformation relationship between the coordinate points of the image and the 3D model in the camera coordinate system, the position coordinates of the image in the model coordinate system can be obtained based on this transformation matrix.

[0040] Optionally, the contour projection of the 3D model at the image location can be marked on the image to facilitate subsequent processing.

[0041] S304, optimize the position coordinates, and determine the second transformation matrix between the world coordinate system and the model coordinate system based on the optimized position coordinates and the actual physical position of the object in the image. The second transformation matrix contains scale information. Position coordinate optimization can be achieved based on the SFM algorithm. Furthermore, the position of the image in the model coordinate system is currently known, and the true physical scale of the target scene has been obtained in step S201.

[0042] Based on these two parameters, a three-dimensional similarity transformation (similarity3, sim3) can be performed. Then, by solving RANSAC, the second transformation matrix of the real-world coordinate system relative to the model coordinate system can be obtained. Since the real physical location in the image is known, and the position coordinates of the image on the model after alignment are also known, the second transformation matrix obtained based on the three-dimensional similarity transformation can reflect the scale scaling information between the real world and the model.

[0043] It should be noted that solving the RANSAC algorithm involves using multiple pairs of matching points to perform a similarity transformation, thereby solving for the rotation matrix, translation vector, and scale between the two coordinate systems. The specific implementation steps for solving the transformation matrix using the RANSAC algorithm, as well as the mathematical operations involved, are conventional techniques in this field and will not be elaborated upon in this embodiment.

[0044] S305 restores the 3D model to its true size based on scale information.

[0045] Through the above steps S301 to S305, compared with the existing methods for restoring the true scale of 3D models, this embodiment creatively utilizes AR to assist in the restoration of the true scale. Real scene images are superimposed on the virtual 3D model. After aligning the two, a preset calculation process is run to obtain the transformation matrix between the model coordinate system and the world coordinate system. Then, the true physical size of the scene is restored based on the transformation matrix.

[0046] This application allows users to capture images in real-time on-site using mobile phones or other terminals, and conveniently and independently perform alignment operations with AR assistance, thereby achieving scale restoration of the model. This technical solution requires no calibration equipment and is not limited by specific application scenarios, greatly improving the convenience of restoring the true scale of 3D models. Furthermore, compared to methods based on calibration devices or image feature calculations, this solution is less affected by interference from the distribution of feature points in the image, resulting in a more uniform and stable alignment area and higher accuracy in the calculation results.

[0047] In some embodiments, before aligning the 3D model with the image, in order to reduce the number of alignment steps for the user and improve efficiency, the 3D model can be preprocessed, specifically including the following steps: Step 1. Determine the gravity axis orientation of the 3D model of the target scene. Optional methods include: manual determination or automatic determination by an algorithm. Step 2. Rotate the model so that its gravity axis coincides with any coordinate axis of the three-dimensional coordinate system. Optionally, in this embodiment, the gravity axis is rotated to coincide with the Z-axis (of course, it can also be the X or Y axis). Figure 4 This is a schematic diagram illustrating the alignment of the gravity axis and Z-axis of a three-dimensional model according to an embodiment of this application. Step 3. Construct the bounding box of the 3D model and scale its shape to a size close to its actual dimensions. Constructing a model's bounding box can be done manually by selecting the box or through automatic calculation. Furthermore, different types of 3D models require different methods for shape alignment, specifically: For very large buildings, optionally, the top view of the building model obtained in step 2 can be compared with the outline in the satellite map and used as the initial value for the next step. For object or room-level models, optionally, the model can be scaled to a size close to the object's outline based on intuitive estimation, and used as the initial value for the next input.

[0048] In some embodiments, a virtual camera is constructed using the camera parameters obtained in step S301, and after the real image and the virtual model are overlaid and displayed in the virtual camera, the user aligns the 3D model with the image, specifically including: The initialization process includes: Align the top and bottom edges of the 3D model's bounding box with the top and bottom edges of the virtual camera's field of view, and rotate the 3D model so that its Z-axis is parallel to the gravity axis of the image coordinate system. Figure 5 This is a schematic diagram illustrating the overlay display of a 3D model and a real image according to an embodiment of this application, such as... Figure 5As shown, the overlay effect of the image and the model is abstractly illustrated. Furthermore, the gravity axis of the image and the Z-axis of the 3D model are set to be parallel. The dotted lines represent the coordinate axes of the 3D model, and the solid arrows indicate the direction of the gravity axis of the image. In this embodiment, the 3D model is initially placed on the optical center line of the virtual camera. By moving the 3D model along the optical center line, its upper and lower edges are aligned with the upper and lower edges of the virtual camera's field of view.

[0049] It should be noted that, compared to directly scaling the model, the above method allows us to calculate the distance from the virtual camera to the model directly based on the virtual camera's field of view (FOV parameter) after the model has moved precisely to the top and bottom alignment along the optical center line. This is because the property of near objects appearing larger than distant objects is utilized within the camera's field of view.

[0050] The position alignment process includes: Using a virtual camera, the 3D model is moved in the front-back depth direction and the up-down direction to align its position with the image; The rotation alignment process includes: Using a virtual camera, the 3D model is rotated around the gravity axis until it perfectly aligns with the scene angle in the image. Figure 6 The diagram shown is a schematic of rotating a three-dimensional model on the gravity axis of an image according to an embodiment of this application.

[0051] It should be noted that in this embodiment, the order of position alignment and rotation alignment is strictly defined. With the above design, compared with alignment by freely moving the camera, the implementation logic in this embodiment is easier, and due to the limitation of operational flexibility, it is also easier to ensure alignment quality.

[0052] Finally, after alignment, the first transformation matrix between the model coordinate system and the camera coordinate system is recorded, and the position coordinates of the image in the model coordinate system are determined based on the first transformation matrix.

[0053] It should be noted that, for a given target scene, to improve alignment quality, at least 3-5 images evenly distributed in space should be selected for the above alignment operation to obtain the corresponding alignment result. Theoretically, the more images aligned, the better the subsequent scale restoration effect.

[0054] In some embodiments, in order to obtain more accurate "position coordinates", after determining the position coordinates of the image in the model coordinate system according to the transformation matrix, it needs to be optimized; Optionally, position coordinate optimization can be performed based on the SFM algorithm, specifically: First, based on the first transformation matrix, the 2D projection coordinates of each vertex in the 3D model (including each point in the point cloud, or face, patch, and vector vertices in the model) in the image are calculated, and multiple sets of association information between the 2D projection coordinates and each vertex are determined. That is, one 2D projection coordinate corresponds to one vertex, and this correspondence between 2D projection points and model vertices is called association information.

[0055] Secondly, the location coordinates are optimized based on multiple sets of related information, and the accuracy of the optimized location coordinates is judged.

[0056] Among these, optimizing location coordinates based on multiple sets of associated information includes: Acquire multiple sets of related information and multiple sets of images of the target scene taken from different angles, extract feature points from multiple sets of images, and perform image matching based on feature points to obtain the 2D relationship between multiple sets of images; Based on the 2D association relationship and the above position coordinates, triangulation is performed to obtain the 3D position of the feature point on the 3D model. Based on the 3D position and virtual camera parameters, the SFM algorithm is used to optimize the image pose and 3D points in the image to obtain the optimized image position coordinates.

[0057] Furthermore, judging the accuracy of the optimized position coordinates includes: Based on the optimized position coordinates, recalculate the 2D projection coordinates of each vertex in the 3D model in the current image; Determine whether the deviation between the recalculated 2D projection coordinates and the 2D projection coordinates before optimization is greater than a preset deviation threshold. If so, output an unreliable response for the optimization result. If not, the optimization result is considered reliable, and the following process continues: based on the real physical location and position coordinates of each object in the image, determine the second transformation matrix from the world coordinate system to the model coordinate system, and perform real scale restoration according to the second transformation matrix; and transform the position of the image in the model coordinate system to the new model coordinate system.

[0058] This embodiment also provides an AR-assisted 3D model scale recovery system, which is used to implement the above embodiments and preferred embodiments, and will not be repeated as already described. As used below, the terms "module," "unit," "subunit," etc., can refer to a combination of software and / or hardware that performs a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also conceivable.

[0059] Figure 7This is a structural block diagram of an AR-assisted 3D model scale recovery system according to an embodiment of this application. The system includes: an acquisition module 70, a preprocessing module 71, an alignment module 72, and a reconstruction and recovery module 73, wherein... The acquisition module 70 is used to acquire target scene data, wherein the target scene data includes: a 3D model, an image, and camera parameters and pose information when the image was captured; The preprocessing module 71 is used to preprocess the 3D model and then overlay it with the image in a virtual camera, wherein the virtual camera is constructed based on camera parameters; The alignment module 72 is used to align the 3D model with the image according to the interactive instructions input by the user through the AR editor, record the first transformation matrix between the model coordinate system and the camera coordinate system, and determine the position coordinates of the image in the model coordinate system according to the first transformation matrix. The reconstruction and restoration module 73 is used to optimize the position coordinates, determine the second transformation matrix between the world coordinate system and the model coordinate system based on the optimized position coordinates and the actual physical position of the object in the image, wherein the second transformation matrix contains scale scaling information, and restore the 3D model to its true size based on the scale scaling information.

[0060] In one embodiment, Figure 8 A schematic diagram of the internal structure of the electronic device according to an embodiment of this application is shown below. Figure 8 The diagram illustrates an electronic device, which can be a server, and its internal structure can be shown as follows. Figure 8 The electronic device includes a processor, a network interface, internal memory, and non-volatile memory connected via an internal bus. The non-volatile memory stores an operating system, computer programs, and a database. The processor provides computing and control capabilities, the network interface communicates with external terminals via a network, the internal memory provides an environment for the operating system and computer programs to run, the computer programs are executed by the processor to implement an AR-assisted 3D model scale reconstruction method, and the database stores data.

[0061] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0062] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0063] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for AR-assisted scale recovery of a 3D model, characterized in that, The method includes: Acquire target scene data, wherein the target scene data includes: a 3D model, an image, and camera parameters and pose information when the image was captured; After preprocessing the 3D model, it is overlaid and displayed on the image in a virtual camera, wherein the virtual camera is constructed based on the camera parameters; Based on the interactive commands input by the user through the AR editor, the 3D model and the image are aligned, and the first transformation matrix between the model coordinate system and the camera coordinate system is recorded. Based on the first transformation matrix, the position coordinates of the image in the model coordinate system are determined. Optimizing the position coordinates includes: Based on the first transformation matrix, calculate the 2D projection coordinates of each vertex in the 3D model in the image, and determine the association information between the 2D projection coordinates and each vertex. Here, one 2D projection coordinate corresponds to one vertex, and the correspondence between the 2D projection point and the model vertex is called the association information. The location coordinates are optimized based on the associated information, and the reliability of the optimized location coordinates is determined. Optimizing the location coordinates based on the associated information includes: Acquire multiple sets of related information, as well as multiple sets of images of the target scene taken from different angles; Feature points are extracted from the multiple sets of images, and the images are matched based on the feature points to obtain the 2D association relationship between the multiple sets of images; Based on the 2D association relationship and the position coordinates, triangulation is performed to obtain the 3D position of the feature point; Based on the 3D position and the parameters of the virtual camera, the pose of the image and the 3D points in the image are optimized using the SFM algorithm to obtain the optimized image position coordinates; The position coordinates are optimized, and based on the optimized position coordinates and the actual physical position of the object in the image, a second transformation matrix between the world coordinate system and the model coordinate system is determined, wherein the second transformation matrix contains scale scaling information; Based on the scale information, the 3D model is restored to its true size.

2. The method according to claim 1, characterized in that, Preprocessing the 3D model includes: Determine the direction of the gravity axis of the three-dimensional model, and rotate the three-dimensional model so that its gravity axis coincides with any coordinate axis of the three-dimensional coordinate system; Construct the bounding box of the 3D model and scale the shape of the 3D model to a size close to its real size.

3. The method according to claim 1, characterized in that, Aligning 3D models with images includes: The initialization process includes: aligning the top and bottom edges of the bounding box of the 3D model with the top and bottom edges of the field of view of the virtual camera, and rotating the 3D model so that its Z-axis is parallel to the gravity axis of the image coordinate system. The position alignment process includes: using the virtual camera, moving the 3D model in the front-back depth direction and the up-down direction to align its position with the image; The rotation alignment process includes: rotating the 3D model around the gravity axis using the virtual camera so that it completely coincides with the scene angle in the image.

4. The method according to claim 3, characterized in that: Based on at least one set of images, they are aligned with a 3D model in the virtual camera to obtain at least one set of alignment results.

5. The method according to claim 1, characterized in that, Determining the reliability of the optimized position coordinates includes: Based on the optimized position coordinates, the 2D projection coordinates of each vertex in the 3D model in the current image are recalculated; Determine whether the deviation between the recalculated 2D projection coordinates and the 2D projection coordinates before optimization is greater than a preset deviation threshold. If so, output an unreliable optimization result response. If not, based on the actual physical location of the object in the image and the location coordinates, determine the second transformation matrix between the world coordinate system and the model coordinate system.

6. A system for AR-assisted scale recovery of a 3D model, characterized in that, The system includes: an acquisition module, a preprocessing module, an alignment module, and a reconstruction and recovery module, wherein, The acquisition module is used to acquire target scene data, wherein the target scene data includes: a 3D model, an image, and camera parameters and pose information when the image was captured; The preprocessing module is used to preprocess the 3D model and then overlay it with the image in a virtual camera, wherein the virtual camera is constructed based on the camera parameters; The alignment module is used to align the 3D model with the image according to the interactive commands input by the user through the AR editor, and record the first transformation matrix between the model coordinate system and the camera coordinate system, and determine the position coordinates of the image in the model coordinate system according to the first transformation matrix; and optimize the position coordinates, including: Based on the first transformation matrix, the 2D projection coordinates of each vertex in the 3D model in the image are calculated, and the association information between the 2D projection coordinates and each vertex is determined; wherein, one 2D projection coordinate corresponds to one vertex, and the correspondence between the 2D projection point and the model vertex is called the association information; The location coordinates are optimized based on the associated information, and the reliability of the optimized location coordinates is determined. Optimizing the location coordinates based on the associated information includes: Acquire multiple sets of related information, as well as multiple sets of images of the target scene taken from different angles; Feature points are extracted from the multiple sets of images, and the images are matched based on the feature points to obtain the 2D association relationship between the multiple sets of images; Based on the 2D association relationship and the position coordinates, triangulation is performed to obtain the 3D position of the feature point; Based on the 3D position and the parameters of the virtual camera, the pose of the image and the 3D points in the image are optimized using the SFM algorithm to obtain optimized image position coordinates. The reconstruction and restoration module is used to optimize the position coordinates, and based on the optimized position coordinates and the actual physical position of the object in the image, determine a second transformation matrix between the world coordinate system and the model coordinate system, wherein the second transformation matrix includes scale information, and... Based on the scale information, the 3D model is restored to its true size.

7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • PTAM improvement method based on ground characteristics of intelligent robot

    CN104732518A

  • Object size detection method, object size detection device and mobile terminal

    CN110276317A

  • Augmented reality image display method and device, equipment and storage medium

    CN113902520A