Three-dimensional reconstruction method and apparatus based on augmented reality
By displaying virtual closed auxiliary lines and a movable cursor on the image acquisition device, automated pose matching of the image acquisition device is achieved, solving the problem of inconsistent accuracy in 3D reconstruction caused by manual shooting by ordinary users, improving image acquisition quality and overlap rate, and promoting high-quality 3D reconstruction.
Patent Information
- Application Number
- PCT/CN2024/102296
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2026-01-02
AI Technical Summary
When ordinary users manually capture images using image acquisition devices such as mobile phone cameras or cameras, it is difficult to achieve the precise and uniform requirements of 3D reconstruction for the position, posture, and angle of the lens of the image acquisition device, resulting in insufficient image acquisition quality and overlap rate.
The image acquisition device displays a virtual closed auxiliary line and a movable cursor around the target object on the viewfinder. By judging the matching between the device pose and the panoramic reference point in real time, it automatically completes image acquisition and records camera pose information to achieve 3D reconstruction.
It significantly improves the accuracy and uniformity of image acquisition, enhances image quality and angular coverage, and is beneficial for high-quality 3D reconstruction.
Smart Images

Figure CN2024102296_02012026_PF_FP_ABST
Abstract
Description
Augmented reality based three-dimensional reconstruction method and device TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, in particular to an augmented reality based three-dimensional reconstruction method and device, a computing device, a computer readable storage medium and a computer program product. BACKGROUND
[0002] Three-dimensional (3D) reconstruction is a technology for generating 3D models from two-dimensional images or data. This technology is widely used in various fields, including film production, game development, architectural design, engineering and manufacturing, etc. 3D reconstruction technology can help people better understand and analyze the shape, size, position and pose of objects, etc. At the same time, it can also provide more realistic and lively three-dimensional models for virtual reality, augmented reality and mixed reality technologies, thereby improving user experience. There are many methods for 3D reconstruction, such as manual modeling, scanning modeling, CAD modeling, modeling based on multi-view photos and / or depth maps, etc. These methods can obtain the surface information of objects in different ways, and then use computer graphics and visual algorithms to convert this information into 3D models.
[0003] With the update iteration of mobile devices, many mobile devices have been configured with cameras or cameras for capturing or shooting images or videos, so it has become an effective way for 3D reconstruction to use image resources shot by image acquisition devices such as mobile phones. However, on the one hand, the images required for computer three-dimensional modeling should have high quality, high resolution, multi-angle coverage and reasonable overlap, which requires the lens of the image acquisition device to be accurately and uniformly positioned during image acquisition; on the other hand, it is difficult for ordinary users (even professional photographers) to manually shoot images using image capture or acquisition devices such as mobile phone cameras or cameras to meet the requirements of accurate and uniform positioning, posture and angle of the lens of the image acquisition device for 3D reconstruction. Therefore, how to use image acquisition devices to collect effective images for 3D reconstruction has become a problem to be solved.
[0004] SUMMARY
[0005] In view of this, the present application provides an augmented reality based three-dimensional reconstruction method and device, a computing device, a computer readable storage medium and a computer program product, which is expected to alleviate or overcome some or all of the above-mentioned defects and other possible defects.
[0006] According to a first aspect of the present application, a method for three-dimensional reconstruction based on augmented reality is provided, comprising: displaying at least one virtual closed auxiliary line and a movable cursor on a viewfinder of an image acquisition device, the at least one virtual closed auxiliary line surrounding a target object, the movable cursor representing a spatial position of the image acquisition device, each virtual closed auxiliary line comprising a plurality of loop-shot reference points; for each virtual closed auxiliary line, in response to the movable cursor moving along the virtual closed auxiliary line, determining in real time whether a real-time pose of the image acquisition device matches a loop-shot reference point adjacent to the movable cursor among the plurality of loop-shot reference points on the virtual closed auxiliary line; for each loop-shot reference point on the at least one virtual closed auxiliary line, in response to the loop-shot reference point being adjacent to the movable cursor and the real-time pose of the image acquisition device matching the loop-shot reference point, obtaining a target object image frame corresponding to the loop-shot reference point and the real-time pose of the image acquisition device to form image frame-pose pair data; and performing three-dimensional reconstruction of the target object based on the image frame-pose pair data corresponding to each loop-shot reference point on the at least one virtual closed auxiliary line.
[0007] In the method for three-dimensional reconstruction based on augmented reality according to some embodiments of the present application, before the displaying of the at least one virtual closed auxiliary line and the movable cursor on the viewfinder of the image acquisition device, the method further comprises: displaying a scalable and movable three-dimensional viewfinder frame on the viewfinder, the shape of the three-dimensional viewfinder frame comprising at least one of a cylinder and a prism; and outputting first prompt information through audio and / or text, the first prompt information comprising selecting the target object by zooming and / or moving the three-dimensional viewfinder frame to surround the target object.
[0008] In the method for three-dimensional reconstruction based on augmented reality according to some embodiments of the present application, the displaying of the at least one virtual closed auxiliary line and the movable cursor on the viewfinder of the image acquisition device comprises: in response to the target object being surrounded by the three-dimensional viewfinder frame, displaying three virtual closed auxiliary lines on the viewfinder, the three virtual closed auxiliary lines surrounding the target object and being located on mutually parallel top, bottom and central planes of the three-dimensional viewfinder frame respectively, each virtual closed auxiliary line comprising a plurality of uniformly distributed loop-shot reference points; and displaying the movable cursor representing the spatial position of the image acquisition device on the viewfinder.
[0009] In the method for three-dimensional reconstruction based on augmented reality according to some embodiments of the present application, the at least one virtual closed auxiliary line comprises at least one of a circular auxiliary line and a regular polygon auxiliary line, and the plurality of loop-shot reference points on each virtual closed auxiliary line are uniformly distributed on the auxiliary line.
[0010] In the augmented reality based three-dimensional reconstruction method according to some embodiments of the present application, when the at least one virtual closed auxiliary line comprises a plurality of virtual closed auxiliary lines, the plurality of virtual closed auxiliary lines are respectively located on mutually parallel planes, and the step of judging, for each virtual closed auxiliary line, whether the real-time pose of the image acquisition device matches the loop shot reference point adjacent to the movable cursor on the virtual closed auxiliary line in real time in response to the movable cursor moving along the virtual closed auxiliary line, comprises, for each virtual closed auxiliary line, in response to the movable cursor moving along the virtual closed auxiliary line, the following steps based on a three-dimensional space coordinate system: obtaining the space coordinates of each loop shot reference point on the virtual closed auxiliary line; obtaining the space coordinates of the real-time position of the image acquisition device in real time; judging whether the real-time pose of the image acquisition device matches the loop shot reference point adjacent to the movable cursor in position according to the space coordinates of each loop shot reference point on the virtual closed auxiliary line and the space coordinates of the real-time position of the image acquisition device; in response to the real-time pose of the image acquisition device matching the loop shot reference point adjacent to the movable cursor in position, judging whether the real-time pose of the image acquisition device matches the loop shot reference point adjacent to the movable cursor in pose according to the space coordinates of each loop shot reference point on the virtual closed auxiliary line and the space coordinates of the real-time position of the image acquisition device; and in response to the real-time pose of the image acquisition device matching the loop shot reference point adjacent to the movable cursor in pose, determining the real-time pose of the image acquisition device as matching the loop shot reference point adjacent to the movable cursor.
[0011] In the augmented reality-based three-dimensional reconstruction method according to some embodiments of the present application, in the three-dimensional space coordinate system, the origin O is the initial position of the image acquisition device in the ring shot, the y-axis is in the first direction, the x-axis and the z-axis are located in the same plane perpendicular to the first direction, the z-axis points to the opposite direction of the image acquisition device shooting, the first direction represents the outer normal direction of the plane on which the at least one virtual closed auxiliary line is located, and the determination of whether the real-time pose of the image acquisition device matches the position of the ring shot reference point adjacent to the movable cursor according to the space coordinates of each ring shot reference point on the virtual closed auxiliary line and the space coordinates of the real-time position of the image acquisition device includes: calculating the first coordinate difference of the real-time position of the image acquisition device and the ring shot reference point adjacent to the movable cursor in the y-axis direction, the second coordinate difference in the x-axis direction, and the third coordinate difference in the z-axis direction according to the space coordinates of the real-time position of the image acquisition device and the space coordinates of each ring shot reference point on the virtual closed auxiliary line; calculating the spatial distance between the real-time position of the image acquisition device and the ring shot reference point, the first distance in the y-axis direction, the second distance in the x-axis direction, and the third distance in the z-axis direction according to the first coordinate difference, the second coordinate difference, and the third coordinate difference; in response to at least one of the following conditions being met, the real-time pose of the image acquisition device is determined to match the position of the ring shot reference point adjacent to the movable cursor on the virtual closed auxiliary line: the first distance between the real-time position of the image acquisition device and the ring shot reference point does not exceed a first preset threshold; the second distance between the real-time position of the image acquisition device and the ring shot reference point does not exceed a second preset threshold; the third distance between the real-time position of the image acquisition device and the ring shot reference point does not exceed a third preset threshold; and the spatial distance between the real-time position of the image acquisition device and the ring shot reference point does not exceed a fourth preset threshold.
[0012] In the method for three-dimensional reconstruction based on augmented reality according to some embodiments of the present application, for each ring-beat reference point in the at least one virtual closed auxiliary line, in response to the ring-beat reference point being adjacent to the movable cursor and the real-time pose of the image acquisition device matching the ring-beat reference point, the target object image frame corresponding to the ring-beat reference point and the real-time pose of the image acquisition device are acquired to form the image frame-pose pair data, including: for each ring-beat reference point in the at least one virtual closed auxiliary line, in response to the ring-beat reference point being adjacent to the movable cursor and the real-time pose of the image acquisition device matching the ring-beat reference point, the following steps are performed: the real-time pose of the image acquisition device corresponding to the ring-beat reference point is acquired and the image of the target object is captured; in response to the first coordinate difference between the real-time position in the real-time pose of the image acquisition device and the first coordinate of the ring-beat reference point in the y-axis direction being other than zero, the clipping range of the image of the target object in the y-axis direction is calculated according to the first coordinate difference; the image of the target object is clipped according to the clipping range of the image of the target object in the y-axis direction to obtain the target object image frame corresponding to the ring-beat reference point; and the image frame-pose pair data corresponding to the ring-beat reference point is formed according to the real-time pose of the image acquisition device corresponding to the ring-beat reference point and the target object image frame.
[0013] In the method for three-dimensional reconstruction based on augmented reality according to some embodiments of the present application, in response to the first coordinate difference between the real-time position in the real-time pose of the image acquisition device and the first coordinate of the ring-beat reference point in the y-axis direction being other than zero, the clipping range of the image of the target object in the y-axis direction is calculated according to the first coordinate difference, including: in response to the first coordinate difference between the real-time position in the real-time pose of the image acquisition device and the first coordinate of the ring-beat reference point in the y-axis direction being greater than zero, the clipping range of the image of the target object in the y-axis direction is moved downward by the pixel value corresponding to the first coordinate difference; and in response to the first coordinate difference between the real-time position in the real-time pose of the image acquisition device and the first coordinate of the ring-beat reference point in the y-axis direction being less than zero, the clipping range of the image of the target object in the y-axis direction is moved upward by the pixel value corresponding to the first coordinate difference.
[0014] In the method for three-dimensional reconstruction based on augmented reality according to some embodiments of the present application, in the three-dimensional coordinate system, the origin O is the initial position of the image acquisition device, the y-axis is in the first direction, the x-axis and the z-axis are in the same plane perpendicular to the first direction, the z-axis points to the opposite direction of the image acquisition device, the first direction represents the direction of the outer normal of the plane where the at least one virtual closed auxiliary line is located, and the determination of whether the real-time pose of the image acquisition device matches the position of the ring shot reference point adjacent to the movable cursor according to the spatial coordinates of each ring shot reference point on the virtual closed auxiliary line and the spatial coordinates of the real-time position of the image acquisition device comprises: calculating the spatial coordinates of the center point of the planar figure formed by the virtual closed auxiliary line according to the spatial coordinates of each ring shot reference point on the virtual closed auxiliary line; calculating the fourth distance between the real-time position of the image acquisition device and the center point according to the spatial coordinates of the center point of the planar figure formed by the virtual closed auxiliary line and the spatial coordinates of the real-time position of the image acquisition device; calculating the fifth distance between the real-time position of the image acquisition device and the ring shot reference point adjacent to the movable cursor in the y-axis direction according to the spatial coordinates of the ring shot reference point adjacent to the movable cursor and the spatial coordinates of the real-time position of the image acquisition device; calculating the equation of the straight line passing through the center point and the ring shot reference point according to the spatial coordinates of the center point and the spatial coordinates of the ring shot reference point adjacent to the movable cursor; calculating the sixth distance between the projection point of the real-time position of the image acquisition device in the plane where the virtual closed auxiliary line is located and the straight line according to the spatial coordinates of the real-time position of the image acquisition device and the equation of the straight line; and determining that the real-time pose of the image acquisition device matches the position of the ring shot reference point adjacent to the movable cursor in response to the fourth distance being within a preset interval, the fifth distance being less than a fifth preset threshold, and the sixth distance being less than a sixth preset threshold.
[0015] In the method for three-dimensional reconstruction based on augmented reality according to some embodiments of the present application, in response to the real-time pose of the image acquisition device matching the position of the loop reference point adjacent to the movable cursor, the real-time pose of the image acquisition device is determined to be posture matched with the loop reference point adjacent to the movable cursor according to the spatial coordinates of each loop reference point on the virtual closed auxiliary line and the spatial coordinates of the real-time position of the image acquisition device, including: in response to the real-time pose of the image acquisition device matching the position of the loop reference point adjacent to the movable cursor, the spatial coordinates of the center point of the planar figure surrounded by the virtual closed auxiliary line are calculated according to the spatial coordinates of each loop reference point on the virtual closed auxiliary line, and a first vector is constructed by connecting the center point and the movable cursor; a second vector is constructed by connecting the center point and the loop reference point adjacent to the movable cursor; the angle between the first vector and the second vector is calculated according to the spatial coordinates of the real-time position of the image acquisition device, the spatial coordinates of the loop reference point adjacent to the movable cursor, and the spatial coordinates of the center point; and in response to the angle between the first vector and the second vector being less than or equal to a seventh preset threshold, the real-time pose of the image acquisition device is determined to be posture matched with the loop reference point adjacent to the movable cursor.
[0016] In the method for three-dimensional reconstruction based on augmented reality according to some embodiments of the present application, the three-dimensional reconstruction of the target object is performed according to at least the image frame-pose pair data corresponding to each loop reference point on the at least one virtual closed auxiliary line, including: for each image frame corresponding to each loop reference point on the at least one virtual closed auxiliary line, a depth map of the image frame is obtained; and the three-dimensional reconstruction of the target object is performed according to the image frame-pose pair data corresponding to each loop reference point on the at least one virtual closed auxiliary line and the depth map.
[0017] In the method for three-dimensional reconstruction based on augmented reality according to some embodiments of the present application, after displaying at least one virtual closed auxiliary line surrounding the target object and a movable cursor on the viewfinder screen of the image acquisition device, the method further includes: outputting second prompt information through audio and / or text, the second prompt information relating to moving the image acquisition device around and aiming at the target object so that the movable cursor moves along the at least one virtual closed auxiliary line, respectively.
[0018] In the method for three-dimensional reconstruction based on augmented reality according to some embodiments of the present application, the image frame-position pair data corresponding to each loop beat reference point in the at least one virtual closed auxiliary line is obtained in response to the loop beat reference point being adjacent to the movable cursor and the real-time position of the image capture device matching the loop beat reference point, including: for each loop beat reference point in the at least one virtual closed auxiliary line, in response to the loop beat reference point being adjacent to the movable cursor and the real-time position of the image capture device matching the loop beat reference point, the following steps are performed: obtaining the real-time position of the image capture device corresponding to the loop beat reference point; capturing an image of the target object to obtain an image frame corresponding to the loop beat reference point; labeling the loop beat reference point; for each virtual closed auxiliary line, in response to each loop beat reference point on the virtual closed auxiliary line being labeled, stopping capturing the image of the target object and outputting third prompt information about the end of the loop beat corresponding to the current virtual closed auxiliary line.
[0019] In the method for three-dimensional reconstruction based on augmented reality according to some embodiments of the present application, the three-dimensional reconstruction of the target object is performed according to at least the image frame-position pair data corresponding to each loop beat reference point in the at least one virtual closed auxiliary line, including: transmitting all the image frame-position pair data corresponding to each loop beat reference point in the at least one virtual auxiliary line to a server for performing the three-dimensional reconstruction of the target object and issuing the result of the three-dimensional reconstruction to a terminal device.
[0020] According to a second aspect of the present application, a device for three-dimensional reconstruction based on augmented reality is provided, including: an auxiliary display module configured to display at least one virtual closed auxiliary line surrounding a target object and a movable cursor on a viewfinder screen of an image capture device, the movable cursor representing the spatial position of the image capture device, and each virtual closed auxiliary line including a plurality of loop beat reference points; a position matching module configured to, for each virtual closed auxiliary line, in response to the movable cursor moving along the virtual closed auxiliary line, judge in real time whether the real-time position of the image capture device matches a loop beat reference point adjacent to the movable cursor among the plurality of loop beat reference points on the virtual closed auxiliary line; a data acquisition module configured to, for each loop beat reference point in the at least one virtual closed auxiliary line, in response to the loop beat reference point being adjacent to the movable cursor and the real-time position of the image capture device matching the loop beat reference point, obtain an image frame of the target object corresponding to the loop beat reference point and the real-time position of the image capture device to form image frame-position pair data; and a three-dimensional reconstruction module configured to perform the three-dimensional reconstruction of the target object according to at least the image frame-position pair data corresponding to each loop beat reference point in the at least one virtual closed auxiliary line.
[0021] According to a third aspect of the present application, there is provided a computing device comprising: a memory and a processor, wherein the memory has stored therein a computer program which, when executed by the processor, causes the processor to perform the method according to some embodiments of the present application.
[0022] According to a fourth aspect of the present application, there is provided a computer readable storage medium having stored thereon computer readable instructions which, when executed, implement the method according to some embodiments of the present application.
[0023] According to a fifth aspect of the present application, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the steps of the method according to some embodiments of the present application.
[0024] In the augmented reality based three-dimensional reconstruction method and apparatus according to some embodiments of the present application, firstly, an AR augmented reality technology is used to plan a collection auxiliary line for an image ring collection process of a target object and display a movable cursor representing a real-time position of a camera, to guide a user to move an image collection device to implement ring shooting; secondly, the image collection device is automatically controlled to shoot target images in the ring shooting process based on real-time matching of a pose of the image collection device and a ring shooting reference point on the auxiliary line (without manual shooting by the user); and finally, three-dimensional reconstruction is implemented based on the images collected by ring shooting and corresponding camera pose information. On the one hand, the display of the closed virtual auxiliary line and the movable cursor based on AR in the image collection device enables the ring shooting movement process of the collection device to be traceable, and can effectively guide the shooter to correctly move the image collection device in the ring shooting process, thereby significantly improving the accuracy and uniformity of the shooting pose. On the other hand, the real-time matching of the pose of the image collection device and the ring shooting reference point on the auxiliary line and the automatic shooting of the target object based on the matching result can avoid the shooting position, pose and angle errors caused by manual determination of the shooting position and manual collection or shooting of the images, thereby significantly improving the quality, angle coverage and overlap rate of the image collection, and thus facilitating subsequent high-quality three-dimensional reconstruction.
[0025] These and other advantages of the application will become apparent from the embodiments described herein after and with reference to the drawings. BRIEF DESCRIPTION OF DRAWINGS
[0026] Embodiments of the present application will now be described in more detail, by way of example only, and with reference to the drawings, in which:
[0027] Fig. 1 schematically illustrates an example application scenario or implementation environment of an augmented reality based three-dimensional reconstruction method according to some embodiments of the present application;
[0028] FIG. 2 illustrates an interaction flow of an example implementation environment or application scenario of the augmented reality based three-dimensional reconstruction method according to some embodiments of the present application;
[0029] FIG. 3 illustrates a flowchart of the augmented reality based three-dimensional reconstruction method according to some embodiments of the present application;
[0030] FIG. 4 schematically illustrates an example interface corresponding to a target selection step in the augmented reality based three-dimensional reconstruction method according to some embodiments of the present application;
[0031] FIGS. 5A and 5B respectively illustrate example interfaces corresponding to an auxiliary display step in the augmented reality based three-dimensional reconstruction method according to some embodiments of the present application;
[0032] FIG. 6 illustrates a schematic diagram of a real scene and a viewfinder screen corresponding to a ring shot guiding process in the augmented reality based three-dimensional reconstruction method according to some embodiments of the present application;
[0033] FIG. 7 illustrates an example flow of a pose matching step in the augmented reality based three-dimensional reconstruction method according to some embodiments of the present application;
[0034] FIG. 8A illustrates an example flow of a position matching in a pose matching process of the augmented reality based three-dimensional reconstruction method according to some embodiments of the present application;
[0035] FIG. 8B illustrates an example flow of a position matching in a pose matching process of the augmented reality based three-dimensional reconstruction method according to some other embodiments of the present application;
[0036] FIG. 9A illustrates an example flow of a pose matching in a pose matching process of the augmented reality based three-dimensional reconstruction method according to some embodiments of the present application;
[0037] FIG. 9B illustrates a schematic diagram of a pose matching in the augmented reality based three-dimensional reconstruction method according to some embodiments of the present application;
[0038] FIG. 10 illustrates an example flow of a data acquisition step in the augmented reality based three-dimensional reconstruction method according to some embodiments of the present application;
[0039] FIG. 11 illustrates an example structural block diagram of an augmented reality based three-dimensional reconstruction apparatus according to some embodiments of the present application;
[0040] FIG. 12 schematically illustrates an example block diagram of a computing device according to some embodiments of the present application. DETAILED DESCRIPTION
[0041] Example embodiments now will be described more fully hereinafter with reference to the accompanying drawings. Example embodiments, however, can be implemented in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the example embodiments to those skilled in the art. Like reference numerals refer to like elements throughout the several views.
[0042] Moreover, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the application. One skilled in the relevant art will recognize, however, that the
[0043] The block diagrams in the drawings show only the functionality of the embodiments and do not imply that the embodiments will take the form discussed in connection therewith. As illustrated in the drawings, the enclosed blocks can be functional blocks that can be implemented in software or with hardware such as a processor of a mobile terminal or similar device.
[0044] The flow diagrams depicted herein are examples of sequences of operations that can be performed. Such sequences can be embodied in software or code modules executed by a processor that are not necessarily limited to any specific
[0045] It should be understood that although the terms first, second, third, etc. can be used herein to describe various components, these components should not be limited by these terms. Thus, a first component discussed below could be termed a second component without departing from the teachings of the present disclosure. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0046] Those skilled in the art will understand that the modules or flowcharts depicted in the accompanying drawings are only schematic and the modules or flowcharts do not necessarily correspond to the actual physical structure of the application. The modules or flowcharts can be implemented in software or hardware or a combination thereof.
[0047] Before embodiments of the present application are described in detail, some related concepts are explained first for the sake of clarity.
[0048] 1. Three-Dimensional (3D) Reconstruction: refers to the process of creating a mathematical model of a three-dimensional object that is suitable for computer representation and processing. It is the foundation for processing, manipulating, and analyzing the properties of objects in a computer environment, and is a key technology for creating virtual reality that represents the real world in a computer. In computer vision, 3D reconstruction is the process of reconstructing three-dimensional information from a single view or multiple views of images. Since the information from a single view is incomplete, it is necessary to use empirical knowledge. Multi-view 3D reconstruction (similar to human binocular positioning) is relatively easy, and its method is to reconstruct three-dimensional information using information from multiple two-dimensional images.
[0049] 2. Augmented Reality (AR): is a technology that skillfully combines virtual information with the real world, widely using multimedia, three-dimensional modeling, real-time tracking and registration, intelligent interaction, sensing and other technical means. After simulating virtual information such as text, images, three-dimensional models, music, and videos generated by computers, it is applied to the real world, and the two kinds of information complement each other, thus achieving "augmentation" of the real world. Augmented reality (AR) is different from virtual reality (VR), which completely creates a virtual environment, while augmented reality adds virtual elements to the real environment. Augmented reality can be achieved through various devices, such as smartphones, tablets, head-mounted displays (such as HoloLens, Magic Leap, etc.), or transparent augmented reality glasses. AR technology has a wide range of applications in gaming, education, medicine, design, entertainment, and many other fields. For example, in the field of education, AR can be used to create virtual historical scenes or scientific experiments, making the learning process more lively and interesting; in the medical field, AR can be used for surgical simulation, rehabilitation training, or to provide detailed visual information about human anatomy.
[0050] 3. Pose: is a concept commonly used in the fields of robotics and computer vision, used to describe the attitude and / or position of an object or robot in three-dimensional space. Specifically, pose refers to the position and / or orientation of an object or robot relative to a reference coordinate system. It contains specific position information (i.e. position) of the object or robot in three-dimensional space and / or direction information (i.e. attitude) it faces. Pose can be described using Cartesian coordinates (position) and / or Euler angles (attitude). The Cartesian coordinate system is composed of three coordinate axes (x, y, z) and an origin, and the position of an object is determined by the numerical values on the coordinate axes. For example, the position of an object in three-dimensional space can be represented by a vector containing three numerical values (x, y, z). Euler angles include pitch, yaw, and roll, which describe the rotation angles of the object around the x-axis, y-axis, and z-axis, respectively. The combination of these three angles can completely describe the attitude of the object in three-dimensional space.
[0051] In the related art, there are many methods for three-dimensional (3D) reconstruction, for example, including:
[0052] 1. Manual modeling: using professional 3D modeling software such as Autodesk 3ds Max, Maya, Blender, SketchUp, etc. for three-dimensional design from scratch. In these software, designers build complex 3D models by creating basic shapes (such as cubes, spheres, etc.), stretching, rotating, Boolean operations, etc.
[0053] 2. Scanning modeling: using a 3D scanner to convert a physical object into a 3D digital model. This is suitable for quickly obtaining the real geometric information of an object, and then the scanned data can be edited and optimized later.
[0054] 3. CAD modeling: for engineering design field, CAD (Computer Aided Design) software such as AutoCAD, SolidWorks, CATIA, etc. is used to accurately establish 3D models of mechanical parts, building structures, etc. of industrial level.
[0055] 4. Image-based modeling: using multi-view photos or depth maps to reconstruct 3D models.
[0056] These methods can all obtain the surface information of an object in different ways, and then use computer graphics and visual algorithms to convert this information into a 3D model. Manual modeling requires professionals to use professional 3D modeling software such as Maya, which is impossible for ordinary users. Scanning modeling requires a professional 3D scanner to scan the target object, and then edit the model on the scanned data, but the 3D scanner is expensive. CAD modeling also requires professionals to do it.
[0057] And image-based modeling is the most cost-effective way for ordinary users to model. With the update iteration of mobile devices, many mobile devices are equipped with (depth) cameras, so using a mobile phone camera to collect image resources to generate a 3D model is the easiest way to achieve. The photos used for 3D modeling should have high quality, high resolution, multi-angle coverage, and reasonable overlap rate. However, when ordinary users or photographers (even professional photographers) manually take pictures using image acquisition devices such as mobile phone cameras or cameras, it is difficult to achieve the precise and unified requirements of the position, posture, and angle of the lens of the image acquisition device for 3D reconstruction due to the inevitable errors of manual control. Therefore, how to assist or guide the photographer or user to more accurately and effectively collect images for 3D reconstruction has become a problem to be solved in three-dimensional reconstruction.
[0058] The present application relates to a three-dimensional reconstruction method based on augmented reality, which uses AR augmented reality technology to guide the user to move the image acquisition device by planning the collection of auxiliary lines for the 360-degree ring shooting process of the target object, and automatically completes the collection of all ring shooting images through real-time matching of the pose of the acquisition device and the ring shooting reference points on the auxiliary lines (without manual collection or shooting by the user), while recording the pose information of the corresponding camera of each collected image, so as to realize three-dimensional reconstruction through algorithms. In this way, the display of the AR-based ring shooting virtual auxiliary line and the movable cursor makes the ring shooting movement process of the image acquisition device traceable, significantly improving the accuracy and uniformity of the shooting pose of the image acquisition device; the automatic completion of the shooting based on the real-time matching of the pose of the image acquisition device and the ring shooting reference points on the auxiliary lines can avoid the errors in the shooting position, attitude and angle caused by manual collection and shooting of the image, significantly improving the quality, angle coverage and overlap rate of the image acquisition, thereby facilitating subsequent high-quality three-dimensional reconstruction.
[0059] Fig. 1 schematically shows an example application scenario or implementation environment 100 of the three-dimensional reconstruction method based on augmented reality according to some embodiments of the present disclosure. As shown in Fig. 1, the application scenario 100 can include an image acquisition device 110 and a server 120, as well as a network 130 for connecting the terminal device 110 and the server 120. In some embodiments, the terminal device 110 can be used to implement the three-dimensional reconstruction method based on augmented reality according to some embodiments of the present application. For example, the terminal device 110 can be deployed with corresponding programs or instructions for executing various methods provided by the present application. Alternatively, the terminal device 110 can also implement various methods according to the present application together with the server 120.
[0060] In some embodiments, the terminal device 110 can be an image capture device, such as a camera, video camera, scanner, or other device with a camera function, for capturing, converting, and transmitting image information. The main components of an image capture device can include a light source (e.g., a flash) to provide sufficient light for image capture to ensure image quality, an image sensor to convert optical signals into electrical signals, a lens to focus light to form a clear image, and an image capture card to convert analog signals into digital signals so that a computer can process them. In some embodiments, the terminal device 110 can also include any type of mobile computing device having an image capture device (e.g., a camera), including a mobile computer (e.g., a personal digital assistant (PDA), a laptop computer, a notebook computer, a tablet computer, a netbook, etc.), a mobile phone (e.g., a smartphone, etc.) as shown in FIG. 1, a wearable computing device (e.g., a smartwatch, a head-mounted device including smart glasses, etc.), or other type of mobile device. In some embodiments, the terminal device 110 can also be a stationary computing device including a removable image capture device, such as a desktop computer, a game console, a smart television, etc. In some embodiments, various mobile terminal devices 110 having an image capture device can also be broadly referred to as image capture devices.
[0061] The server 120 can be a single server or a cluster of servers, or can be a cloud server or a cluster of cloud servers capable of providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, web services, cloud communications, middleware services, domain name services, security services, CDNs, and big data and artificial intelligence platforms, and other basic cloud computing services. It should be understood that the servers referred to herein are typically server computers with a large amount of memory and processor resources, but other embodiments are also possible. In addition, the server 120 is shown by way of example only, and in fact, other devices or combinations of devices with computing and storage capabilities can also be used to provide the corresponding services alternatively or additionally. Optionally, the server 120 can also be a general computing device, including a display and a host computer, etc.
[0062] Examples of the network 130 include a local area network (LAN), a wide area network (WAN), a personal area network (PAN), and / or a combination of communication networks such as the Internet. The server 120 and the terminal device 110 can include at least one communication interface (not shown) capable of communicating over the network 130. Such a communication interface can be one or more of the following: any type of network interface (e.g., a network interface card (NIC)), a wired or wireless (such as IEEE 802.11 wireless LAN (WLAN)) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.
[0063] As shown in FIG. 1, the terminal device 110 can include a display screen 111 and the terminal user 140 can interact with the terminal application 112 via the display screen 111. The terminal application 112 can be a local application program, a Web application program, or a LiteApp (e.g., a mobile phone applet, a WeChat applet) as a lightweight application. In the case where the terminal application is a local application program that needs to be installed, the terminal application 112 can be installed in the terminal device 110. In the case where the terminal application 112 is a Web application program, the terminal application 112 can be accessed through a browser. In the case where the terminal application 112 is a LiteApp, the terminal application 112 can be directly opened on the image acquisition device 110 without installing the terminal application 112 by searching for relevant information of the terminal application (such as the name of the terminal application, etc.), scanning a graphic code (such as a bar code, a two-dimensional code, etc.) of the terminal application, and the like. In this article, the terminal application 112 can be an application for starting and managing a camera to realize image acquisition, such as an application for taking photos and / or videos in a mobile phone. After the camera application is started, a real-time image within the coverage range of the camera and various shooting settings or prompts can be displayed on the display screen of the terminal device for the user to realize image shooting.
[0064] The example application scenario or implementation environment of FIG. 1 is merely illustrative, and the augmented reality-based three-dimensional reconstruction method according to the present application is not limited to the example application scenario shown. It should be understood that, although in this article, the server 120 and the terminal device 110 are shown and described as separate structures, they can also be different components of the same computing device.
[0065] FIG. 2 shows an example interaction flow diagram of the augmented reality-based three-dimensional reconstruction method according to some embodiments of the present application implemented in the example implementation environment or application scenario 100 shown in FIG. 1. The working principle of the data acquisition method according to some embodiments of the present application in the implementation environment or application scenario 100 is briefly introduced below with reference to the example interaction flow diagram shown in FIG. 2.
[0066] First, the terminal device 110 can be configured to display at least one virtual closed auxiliary line surrounding the target object and a movable cursor on the viewfinder screen of the image capturing device, the movable cursor representing the spatial position of the image capturing device, and each virtual closed auxiliary line including a plurality of loop-shooting reference points. For example, different virtual closed auxiliary lines can be located on different planes, so that the image capturing device can capture images of different parts of the target object in all directions; in order to accurately guide the photographer to move the image capturing device during the loop-shooting process, a plurality of loop-shooting reference points are marked on each displayed closed auxiliary line, so that the photographer can try to make the movable cursor close to the loop-shooting reference points during the loop-shooting process, so as to quickly match the device pose with the reference points, thereby improving the loop-shooting efficiency.
[0067] Second, optionally, the auxiliary line display step can be performed after the user 140 selects the target object. Therefore, as shown in FIG. 2, the terminal device can first output a prompt information to the user 140; then, after the user 140 receives the target object selection prompt, the target object is selected in the current image (for example, by surrounding with a movable and scalable three-dimensional viewfinder frame (for example, a three-dimensional viewfinder frame)); finally, in response to the user selecting the target object, the auxiliary line display step is performed. Specifically, as shown in FIG. 2, optionally, before the auxiliary line display step, the terminal device 110 can be configured to output a first prompt information about selecting the target object in the viewfinder screen of the display screen 111. For example, as shown in FIG. 2, the terminal device 110 can prompt the user to select the target object, i.e., the object that needs to be reconstructed in three dimensions, from the current image focused by the image capturing device (for example, a camera) in the real-time viewfinder screen or image displayed on the display screen 111. Optionally, the terminal device 110 can also automatically select the target object from the real-time image of the display screen based on the user's pre-setting.
[0068] Third, optionally, as shown in FIG. 2, the terminal device 110 can be configured to output a second prompt information by audio and / or text, the second prompt information involving moving the image capturing device around and aiming at the target object, so that the movable cursor moves along the at least one virtual closed auxiliary line, respectively. After the movable cursor and the closed auxiliary line are displayed, the user 140 needs to be prompted to loop-shoot the target object according to the auxiliary line and the movable cursor, i.e., to make the cursor move slowly along the auxiliary line during the movement, while the lens always aims at the target object.
[0069] Then, as shown in FIG. 2, after the user 140 receives the prompt to start the loop-shooting according to the auxiliary lines and the cursor, the loop-shooting process is started, and the terminal device 110 can sense the movement of the image acquisition device and can be configured to: for each virtual closed auxiliary line, in response to the movable cursor moving along the virtual closed auxiliary line, determine in real time whether the real-time pose of the image acquisition device matches the loop-shooting reference point adjacent to the movable cursor among the plurality of loop-shooting reference points on the virtual closed auxiliary line.
[0070] Subsequently, as shown in FIG. 2, the terminal device 110 can be configured to: for each loop-shooting reference point in the at least one virtual closed auxiliary line, in response to the loop-shooting reference point being adjacent to the movable cursor and the real-time pose of the image acquisition device matching the loop-shooting reference point, acquire the target object image frame corresponding to the loop-shooting reference point and the real-time pose of the image acquisition device to form image frame-pose pair data.
[0071] Finally, as shown in FIG. 2, after the image frame and real-time pose corresponding to each loop-shooting reference point on all virtual closed auxiliary lines are acquired, the terminal device can be configured to: at least according to the image frame-pose pair data corresponding to each loop-shooting reference point in the at least one virtual closed auxiliary line, perform three-dimensional reconstruction of the target object.
[0072] Alternatively, since the three-dimensional modeling algorithm can involve a relatively complex operation process (such as training of a deep neural network, etc.), the three-dimensional reconstruction process of the target object can also be implemented by the server 120. As shown in FIG. 2, after the terminal device 110 obtains the image frame and real-time pose corresponding to each loop-shooting reference point on all virtual closed auxiliary lines, it can be transmitted to the server 120, which can be configured to receive the image frame-pose pair data corresponding to each loop-shooting reference point on all virtual closed auxiliary lines, and at least the image frame-pose pair data corresponding to each loop-shooting reference point, perform three-dimensional reconstruction of the target object, and then send the reconstruction result to the terminal device 110.
[0073] As shown in FIG. 2, optionally, after the three-dimensional reconstruction of the target object is completed, the terminal device 110 can be further configured to: show the three-dimensional reconstruction model to the user on the display screen, so as to enable the user to quickly and timely obtain the three-dimensional model of the target object.
[0074] FIG. 3 schematically shows a flowchart of an augmented reality-based three-dimensional reconstruction method according to some embodiments of the present application. The three-dimensional reconstruction method shown in FIG. 3 can be implemented in the application scenario shown in FIG. 1, and the execution subject can be the terminal device 110 shown in FIG. 1 or optionally can also include the server 120.
[0075] As shown in FIG. 3, the method for three-dimensional reconstruction based on augmented reality according to some embodiments of the present application can comprise the following steps:
[0076] S310, an auxiliary display step;
[0077] S320, a pose matching step;
[0078] S330, a data acquisition step;
[0079] S340, a three-dimensional reconstruction step.
[0080] Optionally, as shown in the dashed box in FIG. 3, the method for three-dimensional reconstruction based on augmented reality according to some embodiments of the present application can comprise a target selection step S350 before the auxiliary display step S310, and can comprise a loop shot guiding step S360 between S310 and S320.
[0081] The execution process of each of the above steps S310-S360 will be described in detail below in the order of the figures.
[0082] In the optional step S350 (target selection step), the first prompt information about selecting a target object in the viewfinder screen is output to the user.
[0083] According to the concept of the present application, in order to collect effective images for three-dimensional reconstruction, it is necessary to guide the photographer or user to accurately grasp the position and orientation (or pose) of the image acquisition device (especially its lens) during the shooting of the target object by generating and displaying virtual auxiliary lines in the viewfinder screen. Before the virtual auxiliary lines are generated and displayed, it is necessary to determine the object of three-dimensional reconstruction, i.e. the target object to be shot, so as to form the virtual auxiliary lines around the target object. It should be noted that the image acquisition device in the present application can have augmented reality (AR) function, because both the target selection process (step S350) and the subsequent auxiliary display process (step S310) can be performed based on AR.
[0084] Generally, as described in S350, the selection of the target object can be completed by the photographer or user, so that the user can be prompted to select or select the target object before shooting, for example, by audio and / or video. For example, the user can be shown the first prompt information about selecting the target object in the viewfinder screen by means of sound output and / or text display.
[0085] In some embodiments, the first prompt information can be displayed on the display screen at the same time when the viewfinder screen of the image capturing device is displayed, to inform the user to select or choose the target object in the viewfinder screen. As to how to make the user select the target object in the viewfinder screen of the display screen, it can be realized based on the augmented reality technology (see FIG. 4 and the corresponding description). Alternatively, the target choosing process can also be realized by other ways, such as image recognition technology, etc.
[0086] In this article, the image capturing device can represent a special photographing device with photographing function, such as a camera, a video camera, a digital camera, or can also be broadly defined to include various mobile computing devices with image capturing function and image capturing module or component (such as a camera), such as a smart phone, a tablet computer, etc. In some embodiments, the display screen of the image capturing device (such as the display screen of a digital camera) generally allows the photographer or user to view the viewfinder screen or image in real time when shooting and observe the photographed object, and can also be used to realize various shooting parameters, such as exposure, white balance, focus, etc. By observing the image on the display screen, the photographer can adjust these parameters in real time to achieve the best shooting effect. After shooting, the camera display screen can be used to play back the photo. The photographer can check the details, composition and exposure of the photo, etc. for subsequent evaluation and improvement. The display screen of the image capturing device can also display relevant information of the photo, such as shooting date, time, exposure data, etc. These information helps the photographer better understand the shooting situation of the photo, so as to carry out subsequent processing and classification. In summary, the display screen of the image capturing device plays an important role in real-time preview, parameter adjustment, photo playback, information viewing, etc.
[0087] In some embodiments, the viewfinder screen in step S350 refers to the real-time screen of the real scene to be photographed, i.e. the real scene covered by the lens, which is displayed through the display screen of the image capturing device when the photographer (or the photographer) is shooting.
[0088] In some embodiments, the target choosing step S350 (prompting to choose the target object in the viewfinder section on the display screen of the image capturing device) can include: displaying a scalable and movable three-dimensional viewfinder frame on the viewfinder screen, the shape of the three-dimensional viewfinder frame including at least one of a cylinder and a prism; outputting the first prompt information through audio and / or text, the first prompt information including choosing the target object by zooming and / or moving the three-dimensional viewfinder frame to surround the target object.
[0089] Here, the introduction of the stereoscopic framing box facilitates the user to convert the processing of the possibly irregular target object into the processing of regular geometric bodies, thereby simplifying the selection and positioning of the target object and the subsequent generation and positioning of the virtual auxiliary line and its ring-shot reference point. Generally, the stereoscopic framing box can refer to a geometric body whose horizontal cross-section is a regular figure (e.g., a circle, a regular polygon, etc.), and such a geometric body can include a prism, i.e., a prism (including a cuboid and a cube) or a cylinder; optionally, it can also include a truncated pyramid, such as a truncated cone or a truncated prism. In order to facilitate the user to select the target object, the system creates a stereoscopic framing box (e.g., a cube) on the framing picture. The cube supports gesture pinch-in zooming and zooming out and moving, and the user can move and / or scale the cube according to the actual size and position of the target object, and finally ensure that the cube completely wraps (i.e., surrounds) the target object.
[0090] FIG. 4 schematically shows an example interface corresponding to the target selection step in the augmented reality-based three-dimensional reconstruction method in the display screen of the image acquisition device according to some embodiments of the present application.
[0091] As shown in FIG. 4, the framing picture currently displayed in the display screen of the image acquisition device includes a cup 410 as the target object, and a scalable and movable stereoscopic framing box (i.e., a virtual cuboid 420) surrounded by a plurality of dashed lines is also displayed in the framing picture for the user to select or select the target object (e.g., the cup 410). As shown in FIG. 4, below the framing picture, the display screen displays the first prompt information: “Please move and / or scale the virtual cuboid to surround the target object”. When the target object of the three-dimensional reconstruction desired by the user is the cup 410, the above prompt information can be used to inform the user or the photographer to surround the cup 410 by moving and scaling the virtual cuboid 420 to select the cup 410 as the target object. The term “surround” here means that the stereoscopic framing box (e.g., the virtual cuboid 420) completely wraps the target object (e.g., the cup 410) inside, i.e., the entire cup 410 is completely located inside the cuboid 420. As shown in FIG. 4, after the user completes the selection of the target object or the virtual cuboid 420 surrounds the target object 410, the “Confirm” button can be clicked to feed back to the device that the target object selection step is completed; alternatively, the user can click the “Cancel” button during the scaling and / or moving of the virtual cuboid to restore it to the initial state, thereby starting the selection process again.
[0092] It should be noted that, as shown in FIG. 4, since the viewfinder picture is a two-dimensional scene image, the virtual cuboid 420 can only move left and right or up and down in the plane of the viewfinder picture, and can not move forward and backward (i.e., depth direction movement). In this way, in the process of the user selecting the target object cup 410 by using the virtual cuboid 420, if it is necessary to move the virtual cuboid 420 forward and backward in the current picture, the viewfinder lens of the image capture device can be moved to the side (relative to the current direction) so that the original forward and backward movement becomes left and right movement, thereby achieving complete surrounding of the target object 410 by the virtual cuboid 420. Therefore, optionally, as shown in FIG. 4, the first prompt information below the viewfinder picture can further include "please move the lens direction to the side of the target object if necessary".
[0093] In step S310 (auxiliary display step), at least one virtual closed auxiliary line surrounding the target object and a movable cursor representing the spatial position of the image capture device are displayed on the viewfinder picture of the image capture device, and each virtual closed auxiliary line includes a plurality of loop shooting reference points.
[0094] According to the concept of the present application, in order to comprehensively and extensively capture multi-angle and multi-direction images of the target object, the image capture device can perform multi-level 360-degree loop shooting on the target object for various target objects of different shapes and sizes to meet the requirements of multi-angle coverage and reasonable overlap rate of the captured images for three-dimensional reconstruction. In order to avoid the error of shooting angle and direction caused by manual loop shooting by the user, the movable cursor and the virtual auxiliary line and the plurality of loop shooting reference points distributed thereon can be introduced in the viewfinder picture based on AR technology (step S310) to guide the loop shooting process of the user, and further, the real-time matching (step 320) of the pose of the image capture device and each loop shooting reference point is used to automatically realize the positioning (pose) of the shooting moment of the device and the image frame capture (step 330).
[0095] In the auxiliary display step S310, a movable cursor is introduced in the viewfinder frame based on the AR technology to represent the relative position of the real-time position of the image capture device (e.g. its optical center) in the real world in the viewfinder frame, which is used to guide the photographer to move the image capture device in the ring shooting process to move the movable cursor along the virtual closed auxiliary line. In this context, the virtual closed auxiliary line refers to a virtual closed curve surrounding the target object generated and displayed in the viewfinder frame based on the augmented reality technology, which is used to guide the photographer to move the image capture device in combination with the movable cursor to achieve 360-degree ring shooting for the target object. The virtual closed auxiliary line includes a plurality of (virtual) ring shooting reference points for automatically matching the real-time pose of the target image capture device to automatically achieve accurate shooting of the target object. It should be noted that, in this context, for the purpose of clear expression, the position expression of the virtual auxiliary line and the ring shooting reference points in the viewfinder frame and their related concepts (such as the center of the figure surrounded by the virtual auxiliary line, etc.) refers to their spatial positions in the real scene corresponding to the viewfinder frame. Generally, the processing and calculation of various data in this context are based on the spatial positions of various entities (such as image capture devices, virtual auxiliary lines, ring shooting reference points, etc.) in the real scene.
[0096] In some embodiments, in order to guide the user to perform multi-level 360-degree ring shooting, it is necessary to form a virtual closed auxiliary line surrounding the target object on multiple levels (i.e. on multiple mutually different planes, such as multiple mutually parallel (horizontal) planes), so that a total of multiple virtual closed auxiliary lines need to be generated for ring shooting guidance at each level. In some embodiments, for target objects with smaller size and less details, in order to balance efficiency and quality, only one virtual closed auxiliary line (e.g. on the horizontal section where the center is located) can be planned to guide the ring shooting movement of the image capture device.
[0097] FIGS. 5A and 5B respectively show example interfaces corresponding to the auxiliary display step in the augmented reality-based three-dimensional reconstruction method according to some embodiments of the present application.
[0098] In some embodiments, the auxiliary display step S310 can include: in response to the target object being surrounded by the stereoscopic framing frame, displaying three virtual closed auxiliary lines surrounding the target object on the viewfinder frame, the three virtual closed auxiliary lines being located on mutually parallel top, bottom and center planes of the stereoscopic framing frame respectively, each virtual closed auxiliary line including a plurality of uniformly distributed ring shooting reference points; displaying a movable cursor representing the spatial position of the image capture device on the viewfinder frame.
[0099] As shown in FIGS. 5A and 5B, after the cup 510 as the target object is selected, the AR technology can generate a movable (cross) cursor 520 on the viewfinder screen to represent the virtual position marker of the real-time position of the image capturing device in the viewfinder screen; then in response to the target object 510 being surrounded by the stereoscopic viewfinder frame (not shown), i.e. the target object is selected, the AR technology can generate three virtual closed auxiliary lines around the target object 510 on the viewfinder screen, i.e. 530, 540, 550 in FIG. 5A and 560, 570, 580 in FIG. 5B, which are respectively located in three mutually parallel planes.
[0100] As shown in FIG. 5A, in some embodiments, the virtual closed auxiliary lines 530, 540, 550 can be circles around the target object, wherein the centers of the three circles can be on the vertical central axis of the cup 510, and the radii of the three circles can be the same to ensure the accurate uniformity of the shooting distance of the image capturing device. Alternatively, as shown in FIG. 5A, the planes where the auxiliary lines 530, 540, 550 are located are equidistantly distributed, i.e. the distances between adjacent planes are the same. As shown in FIG. 5B, in some embodiments, the virtual closed auxiliary lines 560, 570, 580 can be regular polygon curves around the target object, which are regular direction curves in FIG. 5B, wherein the centers of the three regular directions can be on the vertical central axis of the cup 510, and the side lengths of the three regular directions can be the same to ensure the uniformity of the shooting distance of the image capturing device. Alternatively, as shown in FIG. 5B, the planes where the auxiliary lines 560, 570, 580 are located are equidistantly distributed, i.e. the distances between adjacent planes are the same.
[0101] As shown in FIG. 5A, each of the virtual closed auxiliary lines 530-550 can include a plurality of loop shot reference points evenly distributed on the auxiliary line. For example, in the circular auxiliary line 540, a plurality of loop shot reference points can be evenly distributed, such that the angle between the line connecting any two adjacent loop shot reference points, for example, 540a and 540b, and the center of the circle is fixed. For example, in FIG. 5A, the circular auxiliary line 540 has 16 loop shot reference points, and thus the angle between the line connecting any two adjacent loop shot reference points, for example, 540a and 540b, and the center of the circle is 22.5 degrees. As shown in FIG. 5B, in the square auxiliary line 570, a plurality of loop shot reference points can also be evenly distributed, and the angle between the line connecting any two adjacent loop shot reference points and the center of the square is not necessarily the same. In some embodiments, it can also be predetermined that the angle between the line connecting any two adjacent loop shot reference points on each of the virtual closed auxiliary lines 530-580 and the center of the image enclosed by the auxiliary line is the same, for example, a loop shot reference point is set every 10 degrees on a closed circle or regular polygon, so as to achieve the precise uniformity of loop shot pose guidance.
[0102] In an optional step S360 (loop shot guidance step), second prompt information about moving the image capture device according to the movable cursor and the at least one virtual closed auxiliary line and aligning the lens of the image capture device with the target object is output.
[0103] Based on the concept of the present application, after introducing the movable cursor and the virtual auxiliary line and the plurality of loop shot reference points distributed thereon into the viewfinder image by AR technology, it is necessary to prompt or guide the user to move the image capture device according to the movable cursor and the auxiliary line to implement the loop shot process, for example, by audio and / or text, so as to overcome the error of the shooting position and orientation caused by manually moving the image capture device by the user.
[0104] Similar to the target selection step S350, the loop shot guidance step S360 can also show the user the second prompt information about moving the image capture device according to the movable cursor and the at least one virtual closed auxiliary line and aligning the lens of the image capture device with the target object through audio output and / or text display.
[0105] In some embodiments, the step S360 of the ring shot guiding can comprise outputting a second prompt information in audio and / or text, the second prompt information relating to moving the image capturing device around and aligning with the target object so that the movable cursors move along the at least one virtual closed auxiliary line respectively. For example, the second prompt information can be displayed on the display screen of the image capturing device, the second prompt information including please control the image capturing device to align with the target object and move around the target object so that the movable cursors move along the at least one virtual closed auxiliary line respectively. Alternatively, the second prompt information can also be outputted to the user through an audio output component (e.g. a speaker, etc.) in the image capturing device, or both in audio and in text.
[0106] As shown in FIGS. 5A and 5B, while the movable (cross) cursor 520 and the virtual closed auxiliary lines 530-580 are generated and displayed in the viewfinder of the display screen of the image capturing device, a second prompt information “please align with the target object and move around to make the movable cursors move along each closed auxiliary line respectively” can be displayed below the viewfinder of the display screen. As shown in FIG. 5A, when the user starts the 360-degree ring shot for the target object cup 510 for 3D reconstruction, the above-mentioned second prompt information can be used to guide the user to move around the target object 510 (e.g. take a photo every 10 degrees), so that the movable cross cursor 520 (the virtual position of the image capturing device) on the viewfinder moves along each virtual auxiliary line 530-550 respectively for one round (i.e. 360 degrees), a total of 3 rounds. After the 3 rounds of moving ring shot are completed, as shown in FIG. 5A, the “complete” button can be clicked to inform the image capturing device that the ring shot is completed. Alternatively, as shown in FIGS. 5A and 5B, if the user deviates from the planned path during the ring shot process (i.e. the moving process of the image capturing device), the “cancel” button can be clicked to make it return to the initial state of the ring shot, so as to start the ring shot process again.
[0107] In some embodiments, the user can control the orbiting process of the image capturing device according to the second prompt information in the following way. Taking the virtual auxiliary line 530 in FIG. 5A as an example, first the user can move the image capturing device so that the movable cross cursor 520 in the viewfinder is positioned at the initial position of the orbiting on the auxiliary line 530 (which can be preset), and then slowly move around the target object 510 from the current position of the image capturing device (try to keep the distance unchanged and the lens aimed at the target object), so that the movable cross cursor 520 in the viewfinder always moves along the auxiliary line 530, until it moves a full circle (360 degrees) and returns to the initial position. In this way, the orbiting process guided by the auxiliary line 530 is completed. In a similar way, the orbiting processes guided by the auxiliary lines 540-550 shown in FIG. 5A can be completed one by one. Similarly, the orbiting processes guided by the auxiliary lines 560-580 in FIG. 5B can also be implemented by the user in the above-mentioned way.
[0108] FIG. 6 shows a schematic diagram of the real scene and the viewfinder during the orbiting guiding step in the augmented reality-based three-dimensional construction method according to some embodiments of the present application. As shown in FIG. 6, the target object in the real scene is a cup 610, and the image capturing device 620 (such as a mobile phone) performs a 360-degree orbiting around the cup 610 in the direction of the arrow (i.e. clockwise to the left or counterclockwise to the right) from the initial position in the middle. The circular curve 630 represents the orbiting motion trajectory of the image capturing device 620. As shown in FIG. 6, such an orbiting process is implemented under the guidance of the movable cursor 640 and the virtual auxiliary line 650 displayed in the viewfinder of the display screen of the image capturing device 620. As shown in FIG. 6, the movable cross cursor 640 in the display screen of the image capturing device 620 is the real-time virtual marker of the spatial position of the device in the real scene in the viewfinder. As shown in FIG. 6, as the image capturing device 620 moves to the left or to the right according to the real trajectory 630 in the real scene under the guidance of the movable cross cursor 640 and the auxiliary line 650, the movable cross cursor 640 in the viewfinder also moves along the virtual auxiliary line 650, which indicates that the current orbiting motion trajectory 630 substantially corresponds to the virtual auxiliary line 650, achieving a high-quality orbiting guidance.
[0109] In step S320 (pose matching step), for each virtual closed auxiliary line, in response to the movable cursor moving along the virtual closed auxiliary line, it is determined in real time whether the real-time pose of the image capturing device matches the orbiting reference point adjacent to the movable cursor among the plurality of orbiting reference points on the virtual closed auxiliary line.
[0110] Based on the concept of the present application, in the process of guiding the user to take a picture of the target object by the movable cursor and the virtual auxiliary line based on AR technology, in order to realize the accurate unification of the shooting position and action, the shooting moment can be automatically determined and the image shooting action can be automatically realized by matching the real-time pose of the image acquisition device with each ring shot reference point on the auxiliary line in real time, thereby overcoming the error in shooting position, posture and timing of manual shooting.
[0111] In the present application, the pose of the image acquisition device refers to at least one of the position and the attitude (or orientation) of the device in the three-dimensional space. Among them, the position can usually be described by three-dimensional coordinates in the Cartesian coordinate system (such as three-dimensional space rectangular coordinate system); and the attitude or orientation can be described by Euler angles (such as pitch, yaw and roll). The Cartesian coordinate system is composed of three coordinate axes (x, y, z) and an origin, and the position of an object in three-dimensional space is determined by a three-dimensional array on the coordinate axis. The Euler angles include pitch, yaw and roll, which respectively describe the rotation angles of the object around the x-axis, y-axis and z-axis. The combination of these three angles can completely describe the attitude of the object in three-dimensional space. Further, the real-time pose of the image acquisition device refers to the current position and / or the current attitude information of the image acquisition device (such as its optical center) obtained in real time by, for example, a positioning device and / or a gyroscope, wherein "real-time" can be understood as obtaining the position and / or attitude of the device at every certain time interval (such as several milliseconds).
[0112] With the movement of the image acquisition device around the target object, each real-time matching object of the real-time pose of the image acquisition device can select a ring shot reference point adjacent to the movable cursor, without the need to simultaneously match all ring shot reference points on the virtual closed auxiliary line in real time, which can significantly reduce the amount of calculation and improve work efficiency.
[0113] In some embodiments of the present application, the so-called "matching" of the real-time pose of the image capturing device and the loop reference point can include at least one of position matching and attitude matching. The position matching can be understood as the real-time position of the image completely coincides with the position of the corresponding loop reference point in three-dimensional space, or more broadly understood as the coordinates of at least one dimension coincide, the coordinate difference is small, or the distance between the two is zero or less than a preset threshold distance, etc. The attitude matching can be understood as the real-time attitude of the image capturing device coincides with or has a small difference with the standard attitude corresponding to the loop reference point, such as the matching or coincidence degree of the pitch angle, yaw angle, and roll angle. Alternatively, since the shooting process of the image capturing device is a 360-degree loop around the target object, the matching of the yaw angle of the outer normal direction of the plane on which the auxiliary line is located (such as the vertical upward direction in FIGS. 5A and 5B) can be mainly determined. For the specific process of real-time matching, please refer to FIG. 7 and the corresponding description.
[0114] In step S330 (data acquisition step), for each loop reference point in the at least one virtual closed auxiliary line, in response to the loop reference point being adjacent to the movable cursor and the real-time pose of the image capturing device matching the loop reference point, the target object image frame corresponding to the loop reference point and the real-time pose of the image capturing device are acquired to form image frame-pose pair data.
[0115] Based on the concept of the present application, in the case that the real-time pose of the image capturing device matches the loop reference point adjacent to the movable cursor, automatic real-time shooting or image capturing for the target object can be realized, thereby replacing the adverse effects of shooting parameters (such as focusing, exposure, aperture, etc.) caused by the user manually pressing the shutter or shooting button, and realizing high-quality and high-resolution image acquisition. Moreover, the positions of each automatic shooting are accurately positioned by the real-time matching process and the shooting angles are uniform, thereby realizing reasonable angle overlap rate and coverage rate.
[0116] In some embodiments, since three-dimensional reconstruction not only needs two-dimensional image frames or photos obtained by 360-degree loop shooting of the target object, but also needs to record the real-time position and attitude of the image capturing device when the image frames are captured, in the data acquisition step S330, in addition to obtaining the image frame corresponding to each loop reference point, the real-time position information (such as three-dimensional space coordinates) and attitude information (such as Euler angle) of the image capturing device (such as its optical center) corresponding to each image frame also need to be acquired. In fact, as described in step S320, the real-time position and attitude of the image capturing device can be acquired by, for example, a positioning device and a gyroscope, etc. Therefore, after obtaining the target object image frame corresponding to each loop reference point matched in position and the real-time pose of the image capturing device, the image frame-pose pair data corresponding to each loop reference point can be formed for subsequent three-dimensional reconstruction step (S340).
[0117] The image frames in step S330 are the final two-dimensional images of the target object for three-dimensional reconstruction, which can be real-time images corresponding to each ring shot reference point automatically taken or captured through real-time pose matching, or images after further processing of the real-time images. In some embodiments, in the data acquisition step S330, the images of the target object automatically taken based on real-time pose matching can be directly used as image frames for three-dimensional reconstruction, because the automatically captured images of the target object have overcome various errors of manual ring shots to a certain extent through the above-mentioned steps based on auxiliary line ring shot guidance and pose matching, and thus the images obtained are likely to meet the stringent requirements (such as high quality, high resolution, reasonable angle coverage and overlap rate, etc.) of three-dimensional reconstruction on image frames.
[0118] In some embodiments, although the images automatically captured based on real-time pose matching are significantly better than the images obtained by manual ring shots, the poses of the images taken by the user moving the image acquisition device cannot be guaranteed to completely match the ring shot reference points every time. Therefore, in order to compensate for the pose errors of the captured images (especially the errors of the coordinates of the image acquisition device in the vertical direction and the ring shot reference points) after the images are formed, a post-dynamic compensation can be performed to finally obtain the image frames for three-dimensional reconstruction. For example, the errors in the vertical direction can be compensated by performing a cropping operation on the images of the target object, that is, the coordinate difference in the vertical (i.e. y) direction is dynamically compensated according to the real-time position coordinates of the image acquisition device corresponding to the image to be compensated and the position coordinates of the corresponding ring shot reference point, to ensure that the image frames of the same layer have consistent image ranges.
[0119] In some embodiments, the data acquisition step S330 can include: for each ring shot reference point in the at least one virtual closed auxiliary line, in response to the ring shot reference point being adjacent to the movable cursor and the real-time pose of the image acquisition device matching the ring shot reference point, performing the following steps: acquiring the real-time pose of the image acquisition device corresponding to the ring shot reference point; capturing an image of the target object to obtain an image frame corresponding to the ring shot reference point; labeling the ring shot reference point; for each virtual closed auxiliary line, in response to each ring shot reference point on the virtual closed auxiliary line being labeled, stopping capturing the image of the target object and outputting third prompt information about the end of the ring shot corresponding to the current virtual closed auxiliary line.
[0120] According to the concept of the present application, for a virtual closed auxiliary line (e.g. a circle), the pose matching step S320 and the data acquisition step S330 are performed for each loop shot reference point one by one until all the loop shot reference points on the closed auxiliary line have completed the two steps (it is considered that the current loop shot is completed). Therefore, in order to more clearly indicate in the viewfinder screen the matching completion of the loop shot reference points and the progress of the loop shot and the end point of the loop shot to avoid repeated shooting (which may cause a decrease in work efficiency and waste of resources), the loop shot reference point can be marked after each completion of the pose matching and (image frame-pose pair) data acquisition for one loop shot reference point, to inform the user or the shooter that the image corresponding to the loop shot reference point has been completed. In some embodiments, the marking of the loop shot reference point can include connecting the loop shot reference point with the center of the figure (e.g. a circle) enclosed by the auxiliary line to form a line segment, to indicate that the loop shot reference point has completed the pose matching and image shooting steps. Alternatively, the line segment can be further colored to be different from the color of the auxiliary line.
[0121] In step S340 (three-dimensional reconstruction step), three-dimensional reconstruction of the target object is performed according to at least the image frame-pose pair data corresponding to each loop shot reference point on the at least one virtual closed auxiliary line.
[0122] According to the concept of the present application, after obtaining the 360-degree loop shot image frame set of the target object based on the multi-level auxiliary line guidance and the real-time pose matching of the device with the loop shot reference points, and obtaining the real-time position and attitude of the image acquisition device corresponding to each image frame or each loop shot reference point, the image frame-pose pair data corresponding to each loop shot reference point is input into a three-dimensional reconstruction model to realize three-dimensional reconstruction of the target object.
[0123] In some embodiments, a NERF (Neural Radiance Fields) algorithm can be used to perform three-dimensional reconstruction. The NERF algorithm is a neural network-based three-dimensional reconstruction method that can reconstruct a high-quality three-dimensional scene from a single or multiple images. It uses a neural network to fit a directional radiance flow density function representing each 3D point in the scene, thereby achieving accurate three-dimensional reconstruction. The main idea of NERF is to use a multi-layer perception (MLP) network to learn a five-dimensional function that maps the position (x, y, z) and observation direction Density (σ) and color (c) of the scene are mapped. Then, new views are synthesized from multiple perspectives by ray tracing and volume rendering techniques. Compared with traditional three-dimensional reconstruction methods, NERF has a huge advantage in fineness and realism. Traditional methods usually rely on rough geometric models or sparse point cloud data, while NERF can generate more realistic and detailed three-dimensional scenes. At the same time, NERF can also handle details such as lighting and shadows, providing more rich scene information.
[0124] In some embodiments, the three-dimensional reconstruction step of the target object can include the following processes:
[0125] First, design a neural network with a continuous representation, which accepts three-dimensional spatial coordinates as input; the network architecture of this neural network can be a fully connected network, used to learn and predict the color (RGB value) and volume density of a given spatial point;
[0126] Second, through iterative optimization, associate each pixel on the image frame with the corresponding sampling point in three-dimensional space;
[0127] Third, for each pixel in the image frame, emit a ray along the viewing cone direction, and take multiple discrete points on the ray;
[0128] Fourth, send the three-dimensional coordinates of these discrete points into the network to obtain the predicted color and density;
[0129] Fifth, use the ray integration algorithm (such as volume rendering technique) to accumulate the color and transparency of all points along the ray path to calculate the final pixel color;
[0130] Sixth, compare the actual image pixel color with the synthesized pixel color, calculate the loss function (i.e. color loss) and update the network weights by backpropagation until the model can accurately reproduce the input image.
[0131] Seventh, use the trained neural network to present the three-dimensional geometric shape by extracting the voxelized grid or point cloud.
[0132] To achieve more accurate and realistic three-dimensional reconstruction, a depth map corresponding to each image frame can be further acquired (by using an image acquisition device equipped with a depth camera function, such as a depth camera, a mobile device with a depth camera, etc.) before three-dimensional reconstruction, and three-dimensional reconstruction of the target object can be achieved based on each image frame, the corresponding depth map, and the real-time position and attitude of the corresponding device. Thus, in some embodiments, the three-dimensional reconstruction step S340 can include: for each image frame corresponding to each loop reference point on the at least one virtual closed auxiliary line, acquiring a depth map of the image frame; and according to the image frame-attitude pair data and the depth map corresponding to each loop reference point on the at least one virtual closed auxiliary line, performing three-dimensional reconstruction of the target object.
[0133] In some embodiments, the three-dimensional reconstruction step S340 can include: transmitting all image frame-position pair data corresponding to each loop reference point on the at least one virtual auxiliary line to a server for performing three-dimensional reconstruction of the target object and issuing the result of the three-dimensional reconstruction to the terminal device. As shown in FIG. 2, since the three-dimensional modeling algorithm can involve a relatively complex operation process (such as training of a deep neural network, etc.), the three-dimensional reconstruction process of the target object can also be implemented by the server 120. As shown in FIG. 2, after the terminal device 110 obtains the image frame-attitude pair data corresponding to each loop reference point on all virtual closed auxiliary lines, it can be transmitted to the server 120, which can be configured to receive the image frame-attitude pair data corresponding to each loop reference point on all virtual closed auxiliary lines, and at least the image frame-attitude pair data corresponding to each loop reference point, to perform three-dimensional reconstruction of the target object, and then send the reconstruction result to the terminal device 110.
[0134] In the method for three-dimensional reconstruction based on augmented reality according to some embodiments of the present application, firstly, the AR augmented reality technology is used to plan the collection of auxiliary lines for the image ring shooting process of the target object and display a movable cursor representing the real-time position of the camera to guide the user to move the image collection device to achieve the ring shooting. Secondly, the shooting of the target image in the ring shooting process is automatically completed through the real-time matching of the pose of the image collection device and the ring shooting reference point on the auxiliary line (without manual shooting by the user). Finally, the three-dimensional reconstruction is realized based on the images collected by the ring shooting and the corresponding camera pose information. On the one hand, the display of the closed virtual auxiliary line and the movable cursor based on AR in the image collection device makes the ring shooting movement process of the collection device traceable, which can effectively guide the shooter to move the image collection device correctly during the ring shooting process, significantly improving the accuracy and uniformity of the shooting pose. On the other hand, the real-time matching of the pose of the image collection device and the ring shooting reference point on the auxiliary line and the automatic completion of the image shooting of the target object according to the matching result can avoid the shooting position, pose, and angle errors caused by manual determination of the shooting position and manual collection or shooting of the image, significantly improving the quality, angle coverage, and overlap rate of the image collection, thereby facilitating the subsequent realization of high-quality three-dimensional reconstruction. FIG. 7 shows an example flow of the pose matching step in the method for three-dimensional reconstruction based on augmented reality according to some embodiments of the present application.
[0135] As shown in FIG. 7, the pose matching step S320 can include, for each of the at least one virtual closed auxiliary line, in response to the movable cursor moving along the virtual closed auxiliary line, performing the following steps based on a three-dimensional space coordinate system:
[0136] S321, obtaining the space coordinates of each ring shooting reference point on the virtual closed auxiliary line;
[0137] S322, obtaining the space coordinates of the real-time position of the image collection device in real time;
[0138] S323, judging whether the real-time pose of the image collection device is positionally matched with the ring shooting reference point adjacent to the movable cursor according to the space coordinates of each ring shooting reference point on the virtual closed auxiliary line and the space coordinates of the real-time position of the image collection device, if yes, turning to S324, otherwise, turning to S322 to restart the obtaining of the real-time position of the image collection device;
[0139] S324, in response to the real-time pose of the image collection device being positionally matched with the ring shooting reference point adjacent to the movable cursor, judging whether the real-time pose of the image collection device is pose-matched with the ring shooting reference point adjacent to the movable cursor according to the space coordinates of each ring shooting reference point on the virtual closed auxiliary line and the space coordinates of the real-time position of the image collection device, if yes, turning to S325, otherwise, turning to S322 to restart the obtaining of the real-time position of the image collection device;
[0140] S325, in response to the real-time pose of the image capturing device matching the pose of the ring shot reference point adjacent to the movable cursor, determining the real-time pose of the image capturing device as matching the ring shot reference point adjacent to the movable cursor;
[0141] S326, judging whether all the ring shot reference points on the virtual closed auxiliary line have completed the pose matching, if yes, going to S327 - ending the pose matching, otherwise going to S322 - starting the real-time position acquisition of the image capturing device again;
[0142] S327, in response to all the ring shot reference points on the virtual closed auxiliary line having completed the pose matching with the real-time pose of the image capturing device, ending the pose matching step.
[0143] Generally, before performing the pose matching, a three-dimensional space coordinate system needs to be established, so that the positions can be converted into three-dimensional coordinates to facilitate the calculation of the position and pose matching. In some embodiments, for each virtual closed auxiliary line, a three-dimensional space coordinate system O-xyz can be established with the optical center of the initial shooting position of the image capturing device as the origin, wherein the y-axis points to the first direction, the x-axis and the z-axis are in the same plane perpendicular to the first direction, the z-axis points to the opposite direction of the shooting direction of the image capturing device, and the xyz-axis constitutes a right-hand system, and the first direction represents the outward normal direction of the plane on which the virtual closed auxiliary line is located (e.g. the vertical upward direction). Generally, the three-dimensional space coordinate system O-xyz can be a spatial rectangular coordinate system. Generally, after the three-dimensional space coordinate system is established, the coordinate system remains fixed until the current ring shot is completed.
[0144] As shown in S321-S322, after the three-dimensional space coordinate system is established, the spatial coordinates of each ring shot reference point on the virtual closed auxiliary line can be acquired first. It should be noted that the spatial coordinates of the quality ring shot reference point refer to the spatial coordinates of the corresponding (real) position in the real space, rather than the coordinates of the virtual position in the virtual picture. Since the shape, size and position of the virtual closed auxiliary line and the positions of the various ring shot reference points thereon can be pre-set according to the application scenario before being generated and displayed (e.g. the radius of the circular closed auxiliary line is determined according to the optimal shooting distance of the image capturing device and the target object, so as to determine the positions of the ring shot reference points), the spatial coordinates of the various ring shot reference points in the established three-dimensional space coordinate system can be directly obtained. Secondly, the spatial coordinates of the real-time position of the image capturing device can be acquired in real time along with its movement during the ring shot (e.g. through positioning devices and gyroscopes, etc.).
[0145] Based on the concept of the present application, in the process of the loop shooting of the target object, the matching process of the real-time pose of the image acquisition device and the loop reference point can include at least one of the position matching and the pose matching. In some embodiments, in order to make the real scene space position corresponding to the shooting pose and the loop reference point match to a higher degree, the pose matching can be defined as satisfying both the position matching and the pose matching, so that a more accurate unified image or photo is obtained in the case of such matching pose.
[0146] As shown in steps S323-S324, first, it can be judged whether the real-time position of the image acquisition device matches the position in the real scene corresponding to the loop reference point adjacent to the movable cursor according to the space coordinates of each loop reference point on the virtual closed auxiliary line and the space coordinates of the real-time position of the image acquisition device. The specific judgment method for the position matching can include two kinds. One is to directly calculate the coordinate difference and / or distance between the real-time position and the actual position of the reference point for, for example, a circular auxiliary line, and determine whether the position matches by the size of the coordinate difference and / or distance. The other is to really, for example, a square (or regular polygon) auxiliary line, the shooting distance is as same as possible, the vertical direction (y direction) coordinate is as constant as possible, and the real-time position is as close as possible to the line connecting the loop point and the center. Secondly, in the case of position matching, it can be judged whether the real-time pose of the image acquisition device matches the loop reference point adjacent to the movable cursor in terms of pose according to the space coordinates of each loop reference point on the virtual closed auxiliary line and the space coordinates of the real-time position of the image acquisition device. The pose matching mainly judges the matching of the yaw angle (because the 360-degree loop shooting is performed around the y-axis) around the y-axis, that is, the matching degree is mainly judged by judging the degree of coincidence between the line connecting the real-time position of the image acquisition device and the center point of the virtual closed curve and the line connecting the corresponding loop reference point and the center point.
[0147] As shown in S325-S327, only in the case that the real-time pose of the image acquisition device matches the loop reference point adjacent to the movable cursor in both position and pose, it can be considered that the real-time pose of the image acquisition device matches the loop reference point adjacent to the movable cursor, otherwise, it does not match, and then returns to the step of acquiring the real-time position coordinates of the image acquisition device. Similarly, only after all the loop reference points on the closed auxiliary line have completed the matching with the real-time pose of the image acquisition device, the pose matching process can be ended, otherwise, it returns to the step of acquiring the real-time position S322.
[0148] FIG. 8A schematically shows an example flow of position matching in the pose matching process in the augmented reality-based three-dimensional reconstruction method according to some embodiments of the present application.
[0149] As described above, for a circular virtual closed auxiliary line, the position matching process can be simplified as a judgment of the degree of position coincidence between the real-time position of the image acquisition device and the corresponding (adjacent movable cursor) ring shot reference point. Therefore, in some embodiments, the degree of position coincidence can be judged according to the distance between the real-time position and the spatial position of the ring shot reference point corresponding to the real-time position, and / or the distance in the direction of the three coordinate axes respectively.
[0150] As shown in FIG. 8A, the position matching step S323 shown in FIG. 7 (judging whether the real-time pose of the image acquisition device is positionally matched with the ring shot reference point adjacent to the movable cursor according to the spatial coordinates of each ring shot reference point on the virtual closed auxiliary line and the spatial coordinates of the real-time position of the image acquisition device) can include:
[0151] S323a, calculating a first coordinate difference in the y-axis direction, a second coordinate difference in the x-axis direction, and a third coordinate difference in the z-axis direction between the real-time position of the image acquisition device and the ring shot reference point adjacent to the movable cursor according to the spatial coordinates of the real-time position of the image acquisition device and the spatial coordinates of each ring shot reference point on the virtual closed auxiliary line;
[0152] S323b, calculating a spatial distance between the real-time position of the image acquisition device and the ring shot reference point, a first distance in the y-axis direction, a second distance in the x-axis direction, and a third distance in the z-axis direction according to the first coordinate difference, the second coordinate difference, and the third coordinate difference between the real-time position of the image acquisition device and the ring shot reference point adjacent to the movable cursor;
[0153] S323c, determining the real-time pose of the image acquisition device as positionally matched with the ring shot reference point adjacent to the movable cursor on the virtual closed auxiliary line in response to at least one of the following conditions being met:
[0154] The first distance between the real-time position of the image acquisition device and the ring shot reference point does not exceed a first preset threshold value;
[0155] The second distance between the real-time position of the image acquisition device and the ring shot reference point does not exceed a second preset threshold value;
[0156] The third distance between the real-time position of the image acquisition device and the ring shot reference point does not exceed a third preset threshold value;
[0157] The spatial distance between the real-time position of the image acquisition device and the ring shot reference point does not exceed a fourth preset threshold value.
[0158] In step S323c, the position matching condition is relatively loose, i.e. as long as at least one of the above four conditions is satisfied, the position matching can be considered. Alternatively, a more stringent matching condition can be defined, such as the case where the above four conditions are all satisfied, and then the position matching is considered, so that the position of the image acquisition device and the corresponding ring shot reference point is relatively high.
[0159] On the other hand, for non-circular closed auxiliary lines (such as directional auxiliary lines), the position of the corresponding ring shot reference point can not be the optimal shooting position, because the distances from each reference point of the non-circular auxiliary line to the center of the figure enclosed by the auxiliary line (i.e. the target object) are not the same. Therefore, in this case, the position matching process can be realized by considering the following three distances: one is the fourth distance between the real-time position (of the image acquisition device) and the real-time position of the center point corresponding to the auxiliary line (i.e. the distance from the device to the target object); the second is the fifth distance between the real-time position of the device in the y direction and the reference point (i.e. the distance between the device in the vertical direction and the plane of the auxiliary line); the third is the sixth distance between the projection point of the real-time position of the device on the plane (i.e. the xz plane) of the auxiliary line and the straight line passing through the ring shot reference point and the center point.
[0160] Fig. 8B schematically shows an example flow of position matching in the pose matching process in the augmented reality-based three-dimensional reconstruction method according to some embodiments of the present application.
[0161] As shown in Fig. 8B, the position matching step S323 shown in Fig. 7 can include:
[0162] S323a', calculating the spatial coordinates of the center point of the planar figure enclosed by the virtual closed auxiliary line according to the spatial coordinates of each ring shot reference point on the virtual closed auxiliary line;
[0163] S323b', calculating the fourth distance between the real-time position of the image acquisition device and the center point according to the spatial coordinates of the center point of the planar figure enclosed by the virtual closed auxiliary line and the spatial coordinates of the real-time position of the image acquisition device;
[0164] S323c', calculating the fifth distance between the real-time position of the image acquisition device and the ring shot reference point in the y axis direction according to the spatial coordinates of the ring shot reference point adjacent to the movable cursor and the spatial coordinates of the real-time position of the image acquisition device;
[0165] S324d', calculating the equation of the straight line passing through the center point and the ring shot reference point according to the spatial coordinates of the center point of the planar figure enclosed by the virtual closed auxiliary line and the spatial coordinates of the ring shot reference point adjacent to the movable cursor;
[0166] S323e', according to the spatial coordinates of the real-time position of the image acquisition device and the equation of the straight line passing through the center point and the ring shot reference point adjacent to the movable cursor, calculating a sixth distance between the projection point of the real-time position of the image acquisition device in the plane where the virtual closed auxiliary line is located and the straight line;
[0167] S323f', in response to the fourth distance being within the preset interval, the fifth distance being less than the fifth preset threshold, and the sixth distance being less than the sixth preset threshold, determining the real-time pose of the image acquisition device as matching the ring shot reference point position adjacent to the movable cursor.
[0168] FIG. 9A schematically shows an example flow of pose matching in the pose matching process of the augmented reality-based three-dimensional reconstruction method according to some embodiments of the present application.
[0169] In some embodiments, pose matching mainly judges the matching of the yaw angle around the y-axis (because the 360-degree ring shot is performed around the y-axis), that is, mainly judges the matching degree of the yaw angle by judging the degree of coincidence (such as the size of the included angle) between the line connecting the real-time position of the image acquisition device and the center point of the virtual closed curve and the line connecting the corresponding ring shot reference point and the center point.
[0170] As shown in FIG. 9A, the pose matching step S324 shown in FIG. 7 can include:
[0171] S324a, in response to the real-time pose of the image acquisition device matching the ring shot reference point position adjacent to the movable cursor, calculating the spatial coordinates of the center point of the planar figure surrounded by the virtual closed auxiliary line according to the spatial coordinates of each ring shot reference point on the virtual closed auxiliary line,
[0172] S324b, connecting the center point and the movable cursor to construct a first vector;
[0173] S324c, connecting the center point and the ring shot reference point adjacent to the movable cursor to construct a second vector;
[0174] S324d, according to the spatial coordinates of the real-time position of the image acquisition device, the spatial coordinates of the ring shot reference point adjacent to the movable cursor, and the spatial coordinates of the center, calculating the included angle between the first vector and the second vector;
[0175] S324e, in response to the included angle between the first vector and the second vector not exceeding a seventh preset threshold, determining the real-time position of the image acquisition device as matching the pose of the ring shot reference point adjacent to the movable cursor.
[0176] FIG. 9B schematically shows a principle diagram of pose matching in the pose matching process of the augmented reality based 3D reconstruction method shown in FIG. 9A, according to some embodiments of the present application. FIG. 9B schematically shows a 3D space coordinate system O-xyz established with the optical center of the initial shooting position of the image capturing device as the origin, wherein the y-axis points to the first direction, the x-axis and the z-axis are in the same plane perpendicular to the first direction, the z-axis points to the opposite direction of the shooting direction of the image capturing device, and the xyz-axes form a right-handed system, and the first direction represents the outward normal direction of the plane on which the virtual closed auxiliary line lies (e.g., the vertically upward direction). Generally, the 3D space coordinate system O-xyz can be a spatial rectangular coordinate system.
[0177] For the purpose of illustration, only the movable cursor 910 and one virtual auxiliary line 920 in the middle are shown in FIG. 9B (while other details are omitted) compared with the viewfinder screen shown in FIG. 5A. In addition, as shown in FIG. 9B, the position of the movable cursor 910 is the real-time position of the image capturing device, and the ring shot reference point adjacent to the movable cursor 910 is 921, and the center of the figure enclosed by the virtual auxiliary line 920 is 930. As shown in S324a, in order to perform pose matching, it is necessary to first calculate the spatial coordinates of the center point 930 of the figure enclosed by the virtual auxiliary line 920 (i.e., the 3D space coordinates of the center point in the real scene); then, as shown in S324b and S324c and FIG. 9B, in the viewfinder screen, the first vector a is constructed by connecting the center point 930 with the movable cursor 910, and the second vector b is constructed by connecting the center point 930 with the ring shot reference point 921 adjacent to the movable cursor 910, so that the degree of the angle θ between a and b can be used to determine the degree of coincidence between the real-time deflection angle of the image capturing device and the yaw angle corresponding to the ring shot reference point. Since the image capturing device moves around the target object in the horizontal plane, its pose matching is mainly in the matching of the yaw angle, so the greater the degree of θ, the worse the degree of pose matching; the smaller the degree of θ, the better the degree of matching. As shown in S324d, the angle θ between the two vectors a and b can be calculated according to the 3D space coordinates of the center point 930, the 3D space coordinates of the real-time position of the image capturing device corresponding to the movable cursor 910, and the 3D space coordinates of the ring shot reference point 921, for example, by using the knowledge of spatial analytic geometry. Then, as shown in S324e, a seventh preset threshold value representing the acceptable angle θ of pose matching can be determined in advance, for example, 0-2 degrees, so that when the degree of the angle θ between the first vector a and the second vector b does not exceed the seventh preset threshold value, the real-time pose of the image capturing device is determined to be pose matched with the ring shot reference point adjacent to the movable cursor.
[0178] FIG. 10 schematically shows an example flow of the data acquisition step in the augmented reality based 3D reconstruction method, according to some embodiments of the present application.
[0179] Based on the concept of the present application, the accuracy and uniformity of the user's ring shot of the target object are significantly improved through the guidance of the virtual auxiliary line, and the quality and uniformity of the captured image and the overlap rate and coverage rate of the ring shot image are further improved through the automatic image shooting based on the automatic real-time matching of the image acquisition device and the ring shot reference point, avoiding various inevitable errors caused by manual ring shot. However, in the automatic pose matching process, it is impossible to determine that the matching condition is completely consistent with the ring shot reference point (because if so, the condition is too harsh, and the user cannot accurately match the reference point every time, resulting in the inability to complete the shooting), therefore, as described above, a threshold is used to limit the pose matching condition, that is, the error is small enough to be considered as matching. Therefore, after completing the pose matching and image capture, due to the existence of pose matching error, the captured image may not actually correspond to the image obtained at the corresponding ring shot reference point. At this time, the image of the target object taken at the non-accurate position can be further compensated to improve its accuracy. That is, the pose error of the captured image (especially the error of the image acquisition device in the vertical direction and the ring shot reference point) is dynamically compensated later to finally obtain the image frame for three-dimensional reconstruction.
[0180] As shown in FIG. 10, in some embodiments, the data acquisition step S330 can include: for each ring shot reference point of each virtual closed auxiliary line in the at least one virtual closed auxiliary line, in response to the ring shot reference point being adjacent to the movable cursor and the real-time pose of the image acquisition device matching the ring shot reference point, performing the following steps:
[0181] S331, acquiring the real-time pose of the image acquisition device corresponding to the ring shot reference point and capturing the image of the target object;
[0182] S332, in response to the real-time position in the real-time pose of the image acquisition device being different from the first coordinate of the ring shot reference point in the y-axis direction, calculating the screenshot range of the image of the target object in the y-axis direction according to the first coordinate difference;
[0183] S333, according to the screenshot range of the image of the target object in the y-axis direction, cutting the image of the target object to obtain the target object image frame corresponding to the ring shot reference point;
[0184] S334, forming the image frame-pose pair data corresponding to the ring shot reference point according to the real-time pose of the image acquisition device corresponding to the ring shot reference point and the target object image frame.
[0185] As shown in FIG. 10, first, as shown in S331, when the pose matching passes, the image of the target object can be automatically captured or captured at the current real-time position directly for further processing to obtain the image frame for three-dimensional reconstruction, while the real-time pose of the image acquisition device is obtained and recorded for subsequent three-dimensional reconstruction. Subsequently, as shown in S332, when there is an error in pose matching, especially when the real-time position of the image acquisition device and the ring shot reference point have a distance greater than zero in the first direction (i.e., the outer normal direction of the plane (e.g., the xOz plane) where the ring shot is located, i.e., the y-axis direction), i.e., the first coordinate difference in this direction (y-axis direction) is not equal to zero, based on the first coordinate difference, the corresponding screenshot or image taking range is dynamically determined to ensure that a more accurate image frame is formed based on such a screenshot range, and to ensure that the image taking range of each image frame obtained by the ring shot at the same layer is basically consistent. For example, if the y-coordinate of the real-time position of the image acquisition device > the y-coordinate of the corresponding ring shot reference point, the image taking range is lowered; otherwise, it is raised. And the specific up and down movement amplitude can be determined according to the absolute value of the specific coordinate difference in the first direction or y direction. Subsequently, as shown in S333, the original captured image is screenshot or imaged using the obtained screenshot range, thereby obtaining the required image frame. During the acquisition of the image frame for three-dimensional reconstruction, the post-compensation of the captured image is a further processing of the image of the target object automatically obtained based on the auxiliary line guidance and pose matching, aiming to properly compensate for possible pose matching errors (especially in the numerical direction), so that the obtained image frame is more in line with the requirements of three-dimensional reconstruction. Finally, for each ring shot reference point, an image frame-pose pair data is formed based on its corresponding image frame and the real-time pose of the image acquisition device for subsequent three-dimensional reconstruction process.
[0186] In some embodiments, the screenshot range determination step S332 (calculating the screenshot range of the image of the target object in the y-axis direction according to the first coordinate difference in response to the first coordinate difference between the real-time position in the real-time pose of the image acquisition device and the ring shot reference point in the y-axis direction not being equal to zero) can include: in response to the first coordinate difference between the real-time position in the real-time pose of the image acquisition device and the ring shot reference point in the y-axis direction being greater than zero, lowering the screenshot range of the image of the target object in the y-axis direction by a pixel value corresponding to the first coordinate difference; in response to the first coordinate difference between the real-time position in the real-time pose of the image acquisition device and the ring shot reference point in the y-axis direction being less than zero, raising the screenshot range of the image of the target object in the y-axis direction by a pixel value corresponding to the first coordinate difference.
[0187] As described above, in relation to the problem of dynamically determining the screenshot range by using the coordinate difference, the coordinate difference can be first converted into the corresponding image pixel value, at which time the corresponding relationship between the determination of the coordinate unit length of the established coordinate system and the image pixel size needs to be understood. For example, assuming that 1 unit coordinate length is equal to the side length of 10 pixels, according to the numerical proportional relationship of 1:10, the corresponding coordinate difference can be converted into the pixel value, so as to realize the determination of the screenshot range according to the size (up or down movement amplitude) of the pixel value obtained based on the coordinate difference and the movement direction (up or down screenshot) obtained based on the coordinate difference.
[0188] FIG. 11 is an exemplary structural block diagram of an augmented reality-based three-dimensional reconstruction apparatus 1100 according to some embodiments of the present application. As shown in FIG. 11, the augmented reality-based three-dimensional reconstruction apparatus 1100 can include an auxiliary display module 1110, a pose matching module 1120, a data acquisition module 1130, and a three-dimensional reconstruction module 1140.
[0189] The auxiliary display module 1110 can be configured to display at least one virtual closed auxiliary line surrounding a target object and a movable cursor on a viewfinder screen of an image acquisition device, the movable cursor representing a spatial position of the image acquisition device, each virtual closed auxiliary line including a plurality of loop shot reference points.
[0190] The pose matching module 1120 can be configured to, for each virtual closed auxiliary line, in response to the movable cursor moving along the virtual closed auxiliary line, determine in real time whether a real-time pose of the image acquisition device matches a loop shot reference point adjacent to the movable cursor among the plurality of loop shot reference points on the virtual closed auxiliary line.
[0191] The data acquisition module 1130 can be configured to, for each loop shot reference point in the at least one virtual closed auxiliary line, in response to the loop shot reference point being adjacent to the movable cursor and the real-time pose of the image acquisition device matching the loop shot reference point, acquire a target object image frame corresponding to the loop shot reference point and the real-time pose of the image acquisition device to form image frame-pose pair data.
[0192] The three-dimensional reconstruction module 1140 can be configured to perform three-dimensional reconstruction of the target object according to at least the image frame-pose pair data corresponding to each loop shot reference point in the at least one virtual closed auxiliary line.
[0193] It should be noted that the various modules described above can be implemented in software or hardware or a combination of the two. Multiple different modules can be implemented in the same software or hardware structure, or one module can be implemented by multiple different software or hardware structures.
[0194] In the augmented reality based three-dimensional reconstruction device according to some embodiments of the present application, firstly, an AR augmented reality technology is used to plan a collection auxiliary line for an image ring collection process of a target object and display a movable cursor representing a real-time position of a camera, so as to guide a user to move an image collection device to achieve ring shooting; secondly, the shooting of the target image in the ring shooting process is automatically completed through real-time matching of the pose of the image collection device and a ring shooting reference point on the auxiliary line (without manual shooting by the user); and finally, three-dimensional reconstruction is realized based on the images collected by ring shooting and the corresponding camera pose information. On the one hand, the display of the closed virtual auxiliary line and the movable cursor based on the AR in the image collection device makes the ring shooting movement process of the collection device traceable, can effectively guide the shooter to correctly move the image collection device in the ring shooting process, and significantly improves the accuracy and uniformity of the shooting pose; on the other hand, the real-time matching of the pose of the collection device and the ring shooting reference point on the auxiliary line and the automatic completion of the image shooting of the target object according to the matching result can avoid the shooting position, pose and angle errors caused by manual determination of the shooting position and manual collection or shooting of the image, and significantly improve the quality, angle coverage and overlap rate of the image collection, thereby being beneficial to subsequent high-quality three-dimensional reconstruction.
[0195] FIG. 12 schematically illustrates an example block diagram of a computing device 1200 according to some embodiments of the present application. The computing device 1200 can represent a device to implement various apparatuses or modules described herein and / or perform various methods described herein. The computing device 1200 can be, for example, a server, a desktop computer, a laptop computer, a tablet, a smart phone, a smart watch, a wearable device, or any other suitable computing device or computing system, which can include devices of various levels from full resource devices with a large amount of storage and processing resources to low resource devices with limited storage and / or processing resources. In some embodiments, the augmented reality based three-dimensional reconstruction device 1100 described above with respect to FIG. 11 can be implemented in one or more computing devices 1200, respectively.
[0196] As shown in FIG. 12, the example computing device 1200 includes a processing system 1201, one or more computer-readable media 1202, and one or more I / O interfaces 1203 communicatively coupled to each other. Although not shown, the computing device 1200 can also include a system bus or other data and command transfer system that couples the various components to each other. The system bus can include any one or combination of different bus structures, such as a memory bus or memory controller, a peripheral bus, a serial bus, a parallel bus, and / or a processor or local bus with any of various bus architectures. Alternatively, the system can include a system in which the components are interconnected by a common bus system, by I / O device integrated with a computer system, by point-to-point communications in a star network topology, by some combination thereof, or the like.
[0197] The processing system 1201 is representative of the functionality performed by hardware as an example. As such, the processing system 1201 is illustrated as including hardware elements 1204 that can be configured to perform a variety of functions as an example. As illustrated, the hardware elements 1204 are those that can be readily implemented as specialized computer chips or other logic devices formed using one or more transistors (e.g., electronic integrated circuits (ICs)). The hardware elements 1204 are not limited by the materials from which they are formed or the processing mechanisms employed therein. For example, a processor can be comprised of semiconductor(s) and / or transistors (e.g., electronic integrated circuits (ICs)). In such context, processor-executable instructions can be electronic signals.
[0198] The computer-readable medium 1202 is illustrated as including memory / storage 1205. The memory / storage 1205 represents the memory / storage associated with one or more computer-readable media. The memory / storage 1205 can include volatile media (such as random access memory (RAM)) and / or nonvolatile media (such as read only memory (ROM), floppy disks, etc.). The memory / storage 1205 can include fixed and removable media, where appropriate. The memory / storage 1205 can be used to store various data used in execution of embodiments. The computer-readable medium 1202 can be configured in a variety of other ways as further described below.
[0199] The input / output interface(s) 1203 are representative of a functional interface that allows a user to enter commands and information to computing device 1200, and also allows information to be presented to the user and / or other components or devices using various input / output devices. Examples of input devices include a keyboard, cursor control device (e.g., a mouse), microphone (e.g., for voice inputs), a scanner, touch functionality (e.g., capacitive or other sensors that are configured to detect physical touch), a camera (e.g., which can employ visible or non-visible wavelengths such as infrared frequencies to detect movement that does not involve touch as gestures), a
[0200] The computing device 1200 also includes a three-dimensional reconstruction policy 1206. The three-dimensional reconstruction policy 1206 can be stored in the memory / storage 1205 as computer program instructions as well as hardware or firmware. The three-dimensional reconstruction policy 1206 can implement all of the functionality of the various modules of the augmented reality-based three-dimensional reconstruction apparatus 1100 described with respect to FIG. 11 in conjunction with the processing system 1201 and / or the like.
[0201] Various techniques can be described in the general context of software, hardware, elements, or program modules. Generally, such modules include routines, programs, objects, elements, components, data structures, and the like that perform particular tasks or implement particular abstract data types. The terms "module," "functionality," and the like as used herein generally represent software, firmware, hardware, or a combination thereof. The features of the techniques described herein are platform-independent, meaning that the techniques can be implemented on a variety of computing platforms having a variety of processors.
[0202] Implementations of the described modules and techniques can be stored or transmitted across some form of computer-readable media. Computer-readable media can include various media that can be accessed by the computing device 1200. By way of example, and not limitation, computer-readable media can include "computer-readable storage media" and "computer-readable signal media."
[0203] In contrast to signal-bearing media, "computer-readable storage media" refers to media or means configured to hold information for a period of time. Accordingly, computer-readable storage media does not include signals per se. Computer-readable storage media includes volatile and non-volatile, removable and non-removable media implemented in a method or technology for storage of information such as computer readable instructions, data structures, program modules, logical elements / circuits, or other data. Examples of computer-readable storage media can include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, hard disks, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or other storage devices, tangible media, or articles of manufacture that are appropriate for storage of desired information and that can be accessed by a computer.
[0204] "Computer-readable signal media" refers to a signal-bearing medium that is configured to transmit instructions to the hardware of the computing device 1200, such as via network routing. Computer-readable signal media can further include a computer program product. Computer-readable signal media can include a data signal communicated in a modulated data signal format, such as carried in or on a carrier wave, a baseband signal, or other transport mechanism. Computer-readable signal media can include any information delivery media.
[0205] As previously described, hardware elements 1204 and computer-readable media 1202 are representative of instructions, modules, programmable device logic and / or fixed device logic implemented in a hardware form that can be employed in some embodiments to implement at least portions of the techniques described herein. Hardware elements can include components of an integrated circuit or
[0206] The foregoing combination of software and / or hardware modules can also be employed to implement various techniques and modules described herein. Accordingly, software, hardware, or program modules and other program modules can be implemented as one or more instructions and / or logic embodied on some form of computer-readable storage media and / or by one or more hardware elements 1204. The computing device 1200 can be configured to implement particular instructions and / or functions corresponding to the software and / or hardware modules. Accordingly, implementation of a module that is executable by the computing device 1200 as software can be achieved at least partially in hardware, e.g., through the use of computer-readable storage media and / or hardware elements 1204 of the processing system. The instructions and / or functions can be executable / operable by one or more computing devices 1200 and / or processing systems 1201 to cause performance of techniques, modules, and examples described herein.
[0207] The techniques described herein can be supported by these various configurations of the computing device 1200 and are not limited to the specific examples of the techniques described herein.
[0208] In particular, the processes described above with reference to the flow charts can be implemented as a computer program. For example, an embodiment of the application provides a computer program product comprising a computer program carried on a computer readable medium, the computer program comprising program code for performing at least one of the steps of the method embodiments of the application.
[0209] In some embodiments of the present application, one or more computer readable storage media have computer readable instructions stored thereon which, when executed, implement an augmented reality based three-dimensional reconstruction method according to some embodiments of the present application. The various steps of the augmented reality based three-dimensional reconstruction method according to some embodiments of the present application can be converted into computer readable instructions by programming, and thus stored in a computer readable storage medium. When such a computer readable storage medium is read or accessed by a computing device or computer, the computer readable instructions therein are executed by a processor on the computing device or computer to implement the method according to some embodiments of the present application.
[0210] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. The illustrative description of the above terms in the present specification does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any suitable manner in any one or more embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.
[0211] Any process or method descriptions or descriptions of the flow diagrams in the present specification can be understood as representing code modules, segments, or portions of code which include one or more executable instructions for implementing the specified logic functions (i.e., "application specific circuitry"). It should also be understood that the preferred embodiments of the present application can include other code modules that cause a processor or computer to perform various functions. The other code modules that can be desirably combined with those described in the present specification include program modules, a program data, program components, other related modules, etc. which may
[0212] The logic and / or steps represented in the flow diagrams or otherwise described herein, for example, can be considered a list of executable instructions for implementing the logic function, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, processor- based system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions, or a combination of the above. For purposes of this specification, a "computer-readable medium" can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.
[0213] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (FPGAs), field-programmable gate arrays (FPGAs), etc.
[0214] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by hardware associated with program instructions. The program can be stored in a computer-readable storage medium, and when executed, the program includes performing one or a combination of the steps of the method embodiments.
[0215] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
Claims
1. A method for three-dimensional reconstruction based on augmented reality, characterized in that, The method comprises: displaying at least one virtual closed auxiliary line surrounding the target object and a movable cursor on the viewfinder screen of the image acquisition device, the movable cursor representing the spatial position of the image acquisition device, each virtual closed auxiliary line comprising a plurality of loop shot reference points; for each virtual closed auxiliary line, in response to the movable cursor moving along the virtual closed auxiliary line, determining in real time whether the real-time pose of the image acquisition device matches the loop shot reference point adjacent to the movable cursor among the plurality of loop shot reference points on the virtual closed auxiliary line; for each loop shot reference point in the at least one virtual closed auxiliary line, in response to the loop shot reference point being adjacent to the movable cursor and the real-time pose of the image acquisition device matching the loop shot reference point, obtaining the target object image frame corresponding to the loop shot reference point and the real-time pose of the image acquisition device to form image frame-pose pair data; performing three-dimensional reconstruction of the target object according to the image frame-pose pair data corresponding to each loop shot reference point in the at least one virtual closed auxiliary line.
2. The method of claim 1, wherein, Before the displaying at least one virtual closed auxiliary line surrounding the target object and a movable cursor on the viewfinder screen of the image acquisition device, the method further comprises: displaying a scalable and movable three-dimensional viewfinder frame on the viewfinder screen, the shape of the three-dimensional viewfinder frame comprising at least one of a cylinder and a prism; outputting first prompt information through audio and / or text, the first prompt information comprising selecting the target object by zooming and / or moving the three-dimensional viewfinder frame to surround the target object.
3. The method of claim 2, wherein, The displaying at least one virtual closed auxiliary line surrounding the target object and a movable cursor on the viewfinder screen of the image acquisition device comprises: in response to the target object being surrounded by the three-dimensional viewfinder frame, displaying three virtual closed auxiliary lines surrounding the target object on the viewfinder screen, the three virtual closed auxiliary lines being located on mutually parallel top, bottom and central planes of the three-dimensional viewfinder frame respectively, each virtual closed auxiliary line comprising a plurality of uniformly distributed loop shot reference points; displaying a movable cursor representing the spatial position of the image acquisition device on the viewfinder screen.
4. The method of claim 1, wherein, The at least one virtual closed auxiliary line comprises at least one of a circular auxiliary line and a regular polygon auxiliary line, and the plurality of loop shot reference points on each virtual closed auxiliary line are uniformly distributed on the auxiliary line.
5. The method of claim 4, wherein, When the at least one virtual closed auxiliary line comprises a plurality of virtual closed auxiliary lines, the plurality of virtual closed auxiliary lines are located on mutually parallel planes, and The determining in real time whether the real-time pose of the image acquisition device matches the loop shot reference point adjacent to the movable cursor among the plurality of loop shot reference points on the virtual closed auxiliary line for each virtual closed auxiliary line comprises: for each virtual closed auxiliary line, in response to the movable cursor moving along the virtual closed auxiliary line, performing the following steps based on a three-dimensional spatial coordinate system: obtaining the spatial coordinates of each loop shot reference point on the virtual closed auxiliary line; obtaining the spatial coordinates of the real-time position of the image acquisition device in real time; determining whether the real-time pose of the image acquisition device matches the ring-shot reference point adjacent to the movable cursor according to the spatial coordinates of each ring-shot reference point on the virtual closed auxiliary line and the spatial coordinates of the real-time position of the image acquisition device; in response to the real-time pose of the image acquisition device matching the ring-shot reference point adjacent to the movable cursor, determining whether the real-time pose of the image acquisition device matches the ring-shot reference point adjacent to the movable cursor according to the spatial coordinates of each ring-shot reference point on the virtual closed auxiliary line and the spatial coordinates of the real-time position of the image acquisition device; in response to the real-time pose of the image acquisition device matching the ring-shot reference point adjacent to the movable cursor, determining the real-time pose of the image acquisition device as matching the ring-shot reference point adjacent to the movable cursor.
6. The method of claim 5, wherein, In the three-dimensional coordinate system, the origin O is the initial ring-shot position of the image acquisition device, the y-axis is in the first direction, the x-axis and the z-axis are in the same plane perpendicular to the first direction, the z-axis points to the opposite direction of the image acquisition device shooting, the first direction represents the outer normal direction of the plane on which the at least one virtual closed auxiliary line is located, and The determination whether the real-time pose of the image acquisition device matches the ring-shot reference point adjacent to the movable cursor according to the spatial coordinates of each ring-shot reference point on the virtual closed auxiliary line and the spatial coordinates of the real-time position of the image acquisition device includes: According to the spatial coordinates of the real-time position of the image acquisition device and the spatial coordinates of each ring-shot reference point on the virtual closed auxiliary line, calculating a first coordinate difference in the y-axis direction, a second coordinate difference in the x-axis direction, and a third coordinate difference in the z-axis direction between the real-time position of the image acquisition device and the ring-shot reference point adjacent to the movable cursor; According to the first coordinate difference, the second coordinate difference, and the third coordinate difference, respectively calculating a spatial distance between the real-time position of the image acquisition device and the ring-shot reference point, a first distance in the y-axis direction, a second distance in the x-axis direction, and a third distance in the z-axis direction; in response to at least one of the following conditions being met, determining the real-time pose of the image acquisition device as matching the ring-shot reference point adjacent to the movable cursor on the virtual closed auxiliary line: the first distance between the real-time position of the image acquisition device and the ring-shot reference point does not exceed a first preset threshold value; the second distance between the real-time position of the image acquisition device and the ring-shot reference point does not exceed a second preset threshold value; the third distance between the real-time position of the image acquisition device and the ring-shot reference point does not exceed a third preset threshold value; the spatial distance between the real-time position of the image acquisition device and the ring-shot reference point does not exceed a fourth preset threshold value.
7. The method of claim 6, wherein, The method further includes, for each ring-shot reference point on the at least one virtual closed auxiliary line, in response to the ring-shot reference point being adjacent to the movable cursor and the real-time pose of the image acquisition device matching the ring-shot reference point, acquiring a target object image frame corresponding to the ring-shot reference point and the real-time pose of the image acquisition device to form image frame-pose pair data, including: For each of the at least one virtual closed auxiliary line, in response to the beat reference point adjacent to the movable cursor and the real-time pose of the image acquisition device matching the beat reference point, the following steps are performed: Obtaining the real-time pose of the image acquisition device corresponding to the beat reference point and capturing the image of the target object; In response to the first coordinate difference between the real-time position in the real-time pose of the image acquisition device and the first coordinate of the beat reference point in the y-axis direction being not equal to zero, calculating the screenshot range of the image of the target object in the y-axis direction according to the first coordinate difference; According to the screenshot range of the image of the target object in the y-axis direction, the image of the target object is intercepted to obtain the target object image frame corresponding to the beat reference point; According to the real-time pose of the image acquisition device corresponding to the beat reference point and the target object image frame, the image frame-pose data pair corresponding to the beat reference point is formed.
8. The method of claim 7, wherein, In response to the first coordinate difference between the real-time position in the real-time pose of the image acquisition device and the first coordinate of the beat reference point in the y-axis direction being not equal to zero, calculating the screenshot range of the image of the target object in the y-axis direction according to the first coordinate difference, including: In response to the first coordinate difference between the real-time position in the real-time pose of the image acquisition device and the first coordinate of the beat reference point in the y-axis direction being greater than zero, the screenshot range of the image of the target object in the y-axis direction is moved down by the pixel value corresponding to the first coordinate difference; In response to the first coordinate difference between the real-time position in the real-time pose of the image acquisition device and the first coordinate of the beat reference point in the y-axis direction being less than zero, the screenshot range of the image of the target object in the y-axis direction is moved up by the pixel value corresponding to the first coordinate difference.
9. The method of claim 5, wherein, In the three-dimensional space coordinate system, the origin O is the initial position of the image acquisition device, the y-axis is in the first direction, the x-axis and the z-axis are in the same plane perpendicular to the first direction, the z-axis points to the opposite direction of the image acquisition device, the first direction represents the outer normal direction of the plane on which the at least one virtual closed auxiliary line is located, and The method according to the virtual closed auxiliary line, the space coordinates of each beat reference point on the virtual closed auxiliary line and the space coordinates of the real-time position of the image acquisition device are used to determine whether the real-time pose of the image acquisition device matches the position of the beat reference point adjacent to the movable cursor, including: The space coordinates of the center point of the planar figure surrounded by the virtual closed auxiliary line are calculated according to the space coordinates of each beat reference point on the virtual closed auxiliary line; The fourth distance between the real-time position of the image acquisition device and the center point is calculated according to the space coordinates of the center point of the planar figure surrounded by the virtual closed auxiliary line and the space coordinates of the real-time position of the image acquisition device; The fifth distance between the real-time position of the image acquisition device and the beat reference point adjacent to the movable cursor in the y-axis direction is calculated according to the space coordinates of the beat reference point adjacent to the movable cursor and the space coordinates of the real-time position of the image acquisition device; The equation of the straight line passing through the center point and the beat reference point is calculated according to the space coordinates of the center point and the space coordinates of the beat reference point adjacent to the movable cursor; According to the spatial coordinates of the real-time position of the image acquisition device and the equation of the straight line, a projection point of the real-time position of the image acquisition device in a plane where the virtual closed auxiliary line is located is calculated A sixth distance from the straight line In response to the fourth distance being within a preset interval, the fifth distance being less than a fifth preset threshold, and the sixth distance being less than a sixth preset threshold, the real-time pose of the image acquisition device is determined to match the loop reference point position adjacent to the movable cursor.
10. The method of claim 5, wherein, In response to the real-time pose of the image acquisition device matching the loop reference point position adjacent to the movable cursor, according to the spatial coordinates of each loop reference point on the virtual closed auxiliary line and the spatial coordinates of the real-time position of the image acquisition device, it is judged whether the real-time pose of the image acquisition device matches the loop reference point adjacent to the movable cursor in terms of posture, comprising: In response to the real-time pose of the image acquisition device matching the loop reference point position adjacent to the movable cursor, according to the spatial coordinates of each loop reference point on the virtual closed auxiliary line, the spatial coordinates of the center point of the planar figure surrounded by the virtual closed auxiliary line are calculated, Connecting the center point and the movable cursor to construct a first vector; Connecting the center point and the loop reference point adjacent to the movable cursor to construct a second vector; According to the spatial coordinates of the real-time position of the image acquisition device, the spatial coordinates of the loop reference point adjacent to the movable cursor, and the spatial coordinates of the center point, the included angle between the first vector and the second vector is calculated; In response to the included angle between the first vector and the second vector being less than a seventh preset threshold, the real-time pose of the image acquisition device is determined to match the loop reference point adjacent to the movable cursor in terms of posture.
11. The method of claim 1, wherein, The three-dimensional reconstruction of the target object is performed at least according to the image frame-pose pair data corresponding to each loop reference point on the at least one virtual closed auxiliary line, comprising: For each image frame corresponding to each loop reference point on the at least one virtual closed auxiliary line, a depth map of the image frame is obtained; According to the image frame-pose pair data and the depth map corresponding to each loop reference point on the at least one virtual closed auxiliary line, the three-dimensional reconstruction of the target object is performed.
12. The method of claim 1, wherein, After displaying at least one virtual closed auxiliary line surrounding the target object and a movable cursor on the image acquisition device, the method further comprises: Outputting second prompt information through audio and / or text, and the second prompt information relates to moving the image acquisition device around and aligning the target object, so that the movable cursor moves along the at least one virtual closed auxiliary line, respectively. The image frame corresponding to the loop reference point and the real-time pose of the image acquisition device are obtained to form image frame-pose pair data in response to the loop reference point adjacent to the movable cursor and the real-time pose of the image acquisition device matching the loop reference point for each loop reference point on the at least one virtual closed auxiliary line, comprising:
13. The method of claim 1, wherein, For each of the at least one virtual closed auxiliary line, in response to the beat reference point adjacent to the movable cursor and the real-time pose of the image capture device matching the beat reference point, the following steps are performed: obtaining the real-time pose of the image capture device corresponding to the beat reference point; capturing an image of the target object to obtain an image frame corresponding to the beat reference point; annotating the beat reference point; For each virtual closed auxiliary line, in response to each beat reference point on the virtual closed auxiliary line being annotated, stopping capturing the image of the target object and outputting third prompt information about the end of the beat corresponding to the current virtual closed auxiliary line.
14. The method of claim 1, wherein, The at least according to the image frame-pose pair data corresponding to each beat reference point in the at least one virtual closed auxiliary line, performing three-dimensional reconstruction for the target object, comprising: transmitting all image frame-pose pair data corresponding to each beat reference point in the at least one virtual auxiliary line to a server for performing three-dimensional reconstruction for the target object and issuing the result of the three-dimensional reconstruction to a terminal device.
15. An augmented reality-based three-dimensional reconstruction apparatus, characterized by comprising: comprising: an auxiliary display module configured to display at least one virtual closed auxiliary line surrounding the target object and a movable cursor on the viewfinder screen of the image capture device, the movable cursor representing the spatial position of the image capture device, each virtual closed auxiliary line including a plurality of beat reference points; a pose matching module configured to, for each virtual closed auxiliary line, in response to the movable cursor moving along the virtual closed auxiliary line, judge in real time whether the real-time pose of the image capture device matches the beat reference point adjacent to the movable cursor among the plurality of beat reference points on the virtual closed auxiliary line; a data acquisition module configured to, for each of the at least one virtual closed auxiliary line, in response to the beat reference point adjacent to the movable cursor and the real-time pose of the image capture device matching the beat reference point, obtain the target object image frame corresponding to the beat reference point and the real-time pose of the image capture device to form image frame-pose pair data; a three-dimensional reconstruction module configured to perform three-dimensional reconstruction for the target object according to at least the image frame-pose pair data corresponding to each beat reference point in the at least one virtual closed auxiliary line.
16. A computing device, comprising: a memory and a processor, wherein the memory has stored therein a computer program which, when executed by the processor, causes the processor to perform the method of any one of claims 1-14.
17. A computer-readable storage medium having stored thereon computer-readable instructions which, when executed, implement the method of any one of claims 1-14.
18. A computer program product comprising a computer program which, when executed by a processor, implements the steps of the method according to any one of claims 1-14.
Citation Information
Patent Citations
Modeling method, related electronic equipment and storage medium
CN115526925A
Three-dimensional reconstruction method and terminal equipment
CN116797713A
Tracking and 3D reconstruction of unknown objects
US20240169563A1
Cited By
Augmented reality interaction method, device and system and storage medium
CN122064254A
An augmented reality interaction method, device, system and storage medium
CN122064254B