Method for determining posture and method, device, equipment and medium for generating three-dimensional model
By determining and transforming the relative pose in 3D reconstruction to make it consistent with the measurement unit of the point cloud data, the problems of a single image being unable to represent the complete scene and inconsistent measurement units are solved, high-precision 3D model stitching is achieved, and the robustness of 3D reconstruction and user experience are improved.
Patent Information
- Application Number
- CN202210964457.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-09
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2042-08-09
AI Technical Summary
During the 3D reconstruction process, a single image cannot represent the complete scene information, and since the measurement unit of the depth estimation result is inconsistent with the measurement unit of the relative pose, it makes model splicing difficult, affecting the robustness and accuracy of the 3D reconstruction.
By determining the initial relative pose and transforming it, the measurement unit of the relative pose is made consistent with that of the point cloud data. Deep learning technology is used to generate point cloud data, and the relative pose is calculated in combination with epipolar constraint relationships to achieve 3D model splicing in different coordinate systems.
It improves the robustness and accuracy of 3D reconstruction technology, enhances the stitching effect of large-scene 3D models, and improves user experience.
Smart Images

Figure CN115375740B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, specifically to technical fields such as virtual reality, augmented reality, metaverse, computer vision, and deep learning, and in particular to a method for determining a posture and a method, device, equipment, and medium for generating a three-dimensional model. Background Art
[0002] With the development of computer and network technologies, deep learning techniques have been widely applied in numerous fields. For example, deep learning can be applied to 3D reconstruction scenarios to improve efficiency and reduce costs. 3D reconstruction technology aims to generate 3D models of objects or scenes based on captured images, which can be applied to technologies such as virtual reality, augmented reality, and the metaverse. Summary of the Invention
[0003] The present disclosure aims to provide a method for determining a posture and a method, apparatus, device, and medium for generating a three-dimensional model that are conducive to combining deep learning technology and three-dimensional reconstruction technology.
[0004] According to one aspect of the present disclosure, a method for determining a posture is provided, including determining an initial relative posture for two images based on multiple feature point pairs of two images captured at different postures; wherein the feature point pair includes two feature points belonging to the two images respectively, and the two feature points match each other; the initial relative posture is based on a first measurement unit; based on the initial relative posture, determining first position information of a spatial point corresponding to the feature point pair based on the first measurement unit; based on point cloud data corresponding to the two images, determining second position information of the spatial point based on a second measurement unit, and the point cloud data is based on the second measurement unit; and transforming the initial relative posture according to the first position information and the second position information to obtain a relative posture based on the second measurement unit.
[0005] According to another aspect of the present disclosure, a method for generating a three-dimensional model is provided, comprising: for at least two images of a target scene, generating a three-dimensional sub-model for the image based on the image; taking any one of the at least two images as a reference image, determining an image pair consisting of the reference image and other images of the at least two images to obtain at least one image pair; determining a relative pose for the at least one image pair; and splicing the three-dimensional sub-model for the other image and the three-dimensional sub-model for the reference image based on the relative pose to obtain a three-dimensional model for the target scene, wherein the relative pose is determined using the pose determination method provided in an embodiment of the present disclosure.
[0006] According to another aspect of the present disclosure, a posture determination device is provided, including: a posture determination module, used to determine the initial relative posture for two images based on multiple feature point pairs of two images collected at different postures; wherein the feature point pair includes two feature points belonging to the two images respectively, and the two feature points match each other; the initial relative posture is based on a first measurement unit; a first position determination module, used to determine the first position information of the spatial point corresponding to the feature point pair based on the first measurement unit based on the initial relative posture; a second position determination module, used to determine the second position information of the spatial point based on the second measurement unit based on the point cloud data corresponding to the two images, and the point cloud data is based on the second measurement unit; and a posture transformation module, used to transform the initial relative posture according to the first position information and the second position information to obtain a relative posture based on the second measurement unit.
[0007] According to another aspect of the present disclosure, a device for generating a three-dimensional model is provided, including: a model generation module for generating a three-dimensional sub-model for each image according to at least two images of a target scene; an image pair determination module for determining an image pair consisting of a reference image and other images of the at least two images with any one of the at least two images as a reference image, to obtain at least one image pair; a posture determination module for determining a relative posture for at least one image pair; and a model stitching module for stitching the three-dimensional sub-models for other images and the three-dimensional sub-model for the reference image according to the relative posture, to obtain a three-dimensional model for the target scene, wherein the relative posture is determined using the posture determination device provided by an embodiment of the present disclosure.
[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so as to enable the at least one processor to execute the method for determining a posture and / or the method for generating a three-dimensional model provided by the present disclosure.
[0009] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the method for determining a posture and / or the method for generating a three-dimensional model provided by the present disclosure.
[0010] According to another aspect of the present disclosure, a computer program product is provided, including a computer program / instruction, which, when executed by a processor, implements the method for determining a posture and / or the method for generating a three-dimensional model provided by the present disclosure.
[0011] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0013] Figure 1 Schematic diagram of an application scenario of the method for determining a posture and the method and device for generating a three-dimensional model according to an embodiment of the present disclosure;
[0014] Figure 2 is a flowchart of a method for determining a posture according to an embodiment of the present disclosure;
[0015] Figure 3 is a schematic diagram of the principle of determining point cloud data according to an embodiment of the present disclosure;
[0016] Figure 4 is a schematic diagram of the principle of determining the initial relative posture according to an embodiment of the present disclosure;
[0017] Figure 5 is a flowchart of a method for generating a three-dimensional model according to an embodiment of the present disclosure;
[0018] Figure 6 is a structural block diagram of a posture determination device according to an embodiment of the present disclosure;
[0019] Figure 7 is a structural block diagram of a device for generating a three-dimensional model according to an embodiment of the present disclosure; and
[0020] Figure 8 It is a block diagram of an electronic device used to implement the method for determining the posture and / or the method for generating a three-dimensional model according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0021] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0022] In 3D reconstruction technology, the 3D layout information of a scene can be generated based on a single image. However, due to the influence of perspective, when the scene is too large or there are too many obstructions, a single image cannot represent the complete scene information. In order to generate a complete 3D model, a 3D model of a local area can be generated for each image, and then the spatial relationship between the 3D models of multiple local areas can be determined by positioning data, and the 3D models of multiple local areas can be spliced according to the spatial relationship. Among them, when realizing the splicing of 3D models in two different coordinate systems, the 3D models in the two different coordinate systems can be first transformed into the same coordinate system according to the transformation relationship between the two different coordinate systems, and then the 3D models can be spliced.
[0023] In one embodiment, the transformation relationship between two different coordinate systems can be represented by the change in pose when an image acquisition device captures two images corresponding to two three-dimensional models. That is, the transformation relationship can be represented by the relative pose of the two images when the image acquisition device captures them. The relative pose can be calculated based on an epipolar constraint relationship or a homography constraint relationship.
[0024] The three-dimensional layout information (i.e., the three-dimensional model of the local area) is usually generated based on the point cloud data corresponding to a single image. The generation of point cloud data depends on the depth estimation result obtained for a single image (i.e., the monocular depth estimation result). In the process of realizing the concept of the present disclosure, the inventors found that there is inconsistency between the measurement unit of the point cloud data generated based on the monocular depth estimation result and the measurement unit of the relative posture. For example, the relative posture is expressed in terms of the distance between the two positions where the optical center of the image acquisition device is located when the two images are acquired (i.e., the optical center distance), and the point cloud data is usually expressed in commonly used length measurement units (e.g., 1m).
[0025] Based on this, the present disclosure provides a method for determining a relative pose, which can ensure that the measurement unit of the determined relative pose is consistent with the measurement unit of the point cloud data, so as to facilitate model splicing. Based on the relative pose determined by the posture determination method, the present disclosure also provides a method for generating a three-dimensional model.
[0026] The following is combined first Figure 1 , describes the application scenarios of the methods and devices provided by the present disclosure.
[0027] Figure 1 It is a schematic diagram of an application scenario of the method for determining posture and the method and device for generating a three-dimensional model according to an embodiment of the present disclosure.
[0028] like Figure 1As shown, the application scenario 100 of this embodiment may include an electronic device 110, which may be various electronic devices with processing functions, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, servers, etc.
[0029] The electronic device 110 may process the collected multiple images of the target scene to generate a three-dimensional model of the target scene. The multiple images of the target scene may include, for example, images 121 to 123 , and the generated three-dimensional model may be, for example, model 130 .
[0030] In one embodiment, the electronic device 110 may first generate point cloud data for each image, and use the generated point cloud data to represent the 3D model corresponding to each image. Subsequently, the 3D model is obtained by aggregating the point cloud data generated for multiple images.
[0031] When aggregating point cloud data, for example, image 121 can be used as a reference image to determine the relative poses of image 122 and image 123 relative to image 121. Based on these relative poses, the point cloud data generated for image 122 and image 123 are transformed so that all point cloud data for the three images are represented by the same coordinate system. Subsequently, the point cloud data is aggregated. The relative pose between the two images can, for example, represent the change in pose of the acquisition device when the two images were acquired.
[0032] Among them, when generating point cloud data, for example, the depth of the image can be estimated first. For example, a multi-eye depth estimation method or a monocular depth estimation method can be used to perform depth estimation. Among them, for the multi-eye depth estimation method, it is necessary to pre-calibrate the camera parameters and complete the depth estimation based on paired images or image sequences. For the monocular depth estimation method, the depth of each image can be estimated separately. For example, this embodiment can use a monocular depth estimation model 140 constructed based on deep learning technology to perform depth estimation. Specifically, each image can be used as the input of the monocular depth estimation model 140, and the monocular depth estimation model 140 generates a depth map of each image.
[0033] In one embodiment, if Figure 1 As shown, the application scenario 100 may further include a server 150, which may be, for example, a background management server supporting the operation of client applications in the electronic device 110. The electronic device 110 may be connected to the server 150 via a network, which may include a wired or wireless communication link.
[0034] For example, the server 150 may generate the monocular depth estimation model 140 in a supervised manner or a self-supervised manner. The server 150 may also feed back the monocular depth estimation model 140 to the electronic device 110 in response to a request from the electronic device 110 .
[0035] In one embodiment, the electronic device 110 may further send the received multiple images to the server 150 , and the server 150 may process the multiple images and generate a three-dimensional model.
[0036] It should be noted that the method for determining the posture and the method for generating the three-dimensional model provided in the present disclosure can be executed by the electronic device 110 or by the server 150. Accordingly, the apparatus for determining the posture and the apparatus for generating the three-dimensional model provided in the present disclosure can be set in the server 150 or in the electronic device 110.
[0037] It should be understood that Figure 1 The number and type of the electronic devices 110 and the server 150 in FIG. 1 are merely illustrative. Any number and type of electronic devices 110 and servers 150 may be provided as required.
[0038] The following will be combined Figures 2 to 4 The method for determining the posture provided by the present disclosure is described in detail.
[0039] Figure 2 4 is a flow chart of a method for determining a posture according to an embodiment of the present disclosure.
[0040] like Figure 2 As shown, the posture determination method 200 of this embodiment may include operations S210 to S240.
[0041] In operation S210 , an initial relative pose for two images is determined based on a plurality of feature point pairs of two images collected at different poses.
[0042] According to an embodiment of the present disclosure, when determining relative pose, for example, a feature extraction algorithm may be first used to extract feature points from each image, and two feature points representing the same object and the same position in the two images are combined to form a feature point pair. The feature extraction algorithm may include, for example, any one of the Scale-Invariant Feature Transform (SIFT) algorithm, the Oriented Fast and Rotated Brief (ORB) algorithm, and the Speeded Up Robust Feature (SURF) extraction algorithm.
[0043] This embodiment can solve the essential matrix E or the basic matrix F based on multiple feature point pairs, and then solve the rotation matrix R and the translation vector t representing the initial relative pose by singular value decomposition. For example, the essential matrix E can be estimated by using an epipolar constraint relationship or a homography constraint relationship. For example, a feature point pair is set to include a feature point p1 belonging to the first image of the two images and a feature point p2 belonging to the second image of the two images, and the camera intrinsic parameter is set to K. Then, the rotation matrix R and the translation vector t should satisfy the following constraint relationship formula (1):
[0044]
[0045] Among them, x1 and x2 represent the positions of two feature points p1 and p2 respectively, t^ is the antisymmetric matrix of t, the essential matrix E can be represented by (t^R), and the basic matrix F can be represented by (K -T EK -1 ). According to this constraint, given the coordinates of the feature points in the known feature point pairs and the camera intrinsic parameter K, the rotation matrix R and the translation vector t can be obtained.
[0046] It is understandable that the number of multiple feature point pairs can be set according to actual needs, for example, it can be 5, 8 or more, and this disclosure does not limit this. The rotation matrix R and translation vector t obtained by solution are expressed based on a first unit of measurement. The first unit of measurement can be determined according to the solution method, and this disclosure does not limit this. For example, when using the epipolar constraint relationship to determine the initial relative pose, the first unit of measurement can be the optical center distance.
[0047] In one embodiment, before determining the relative pose based on the feature point pairs, a random sample consensus (RANSAC) algorithm may be used to determine multiple feature point pairs.
[0048] In one embodiment, the captured image may be, for example, a panoramic image. Thanks to the all-round field of view of 360° horizontally and 180° vertically, the captured panoramic image can obtain image information of the scene to the maximum extent, thus providing great advantages in scene understanding and structural information acquisition.
[0049] In operation S220 , first position information of a spatial point corresponding to the feature point pair based on a first measurement unit is determined according to the initial relative pose.
[0050] This embodiment can also determine the first position and initial posture of the optical center when the image acquisition device captures one of the two images, and then obtain the second position and posture of the optical center when capturing the other image based on the initial relative posture. According to the pixel positions of the two feature points in each feature point pair, the positions of the two feature points in the coordinate system constructed based on the image acquisition device can be converted. Subsequently, the position of the intersection of the first position and the first line connecting the feature point belonging to one image of the two feature points and the second position and the second line connecting the feature point belonging to the other image of the two feature points can be used as the first position information of the spatial point corresponding to each feature point pair. Since the initial relative posture is expressed based on the first unit of measurement, the unit of measurement of the coordinate system constructed based on the image acquisition device is the first unit of measurement, and the first position information obtained is also expressed based on the first unit of measurement.
[0051] It is understandable that the above method for determining the first position information is only used as an example to facilitate understanding of the present disclosure, and the present disclosure may also adopt other relevant methods to determine the first position information.
[0052] In operation S230 , second position information of the spatial point based on a second measurement unit is determined according to the point cloud data corresponding to the two images.
[0053] According to an embodiment of the present disclosure, the point cloud data corresponding to each image can be obtained using a point cloud generation network. For each feature point pair comprising two feature points, this embodiment can first use the point cloud data corresponding to the two feature points as the second position information based on the spatial position included in the point cloud data. Typically, the point cloud data corresponding to the image is expressed based on a commonly used unit of measurement (e.g., a second unit of measurement such as 1m), and the obtained second position information is based on the second unit of measurement.
[0054] It is understandable that the point cloud data corresponding to each point is a vector in a three-dimensional coordinate system. The vector includes three-dimensional coordinates representing a spatial position and may also include at least one of color information and reflection intensity information.
[0055] It is understandable that operation S230 may be executed synchronously with operations S210 to S220, and operation S230 may also be executed before operation S210 or operation S220. The present disclosure does not limit the execution order.
[0056] In operation S240 , the initial relative position is transformed according to the first position information and the second position information to obtain a relative position expressed based on a second measurement unit.
[0057] According to an embodiment of the present disclosure, a transformation relationship between a first measurement unit and a second measurement unit can be determined based on the first position information and the second position information of the spatial points corresponding to the plurality of feature point pairs. Subsequently, the initial relative pose is transformed based on the transformation relationship to obtain a relative pose expressed based on the second measurement unit.
[0058] For example, multiple first position information of spatial points corresponding to multiple feature point pairs can be aggregated into one information group, and multiple second position information of spatial points corresponding to multiple feature point pairs can be aggregated into one information group. The linear relationship between the two information groups determined by fitting is used as the transformation relationship between the first measurement unit and the second measurement unit. Alternatively, a transformation relationship can be determined for each feature point pair based on the first position information and the second position information of the corresponding spatial point, and a total of multiple transformation relationships can be determined. This embodiment can average the multiple transformation relationships to obtain a transformation relationship between the first measurement unit and the second measurement unit.
[0059] The embodiment of the present disclosure estimates the first position information of the spatial point based on the relative pose, obtains the second position information of the spatial point based on the point cloud data, and then transforms the initial relative pose based on the transformation relationship between the two position information, so that the relative pose obtained by the transformation has the same measurement unit as the point cloud data, so as to facilitate the splicing of the three-dimensional model generated according to the two images. Through the relative pose determination method of this embodiment, a method for determining the relative pose and a method for determining the point cloud data based on different measurement units can be adopted, which is conducive to improving the robustness of the three-dimensional reconstruction technology. For example, based on the method of the embodiment of the present disclosure, in the three-dimensional reconstruction technology, a monocular depth estimation model constructed based on deep learning can be used to estimate the depth, and point cloud data can be generated based on the depth. At the same time, the relative pose can be determined by using the epipolar constraint relationship.
[0060] According to an embodiment of the present disclosure, before determining the second position information, the method for determining the relative posture may, for example, first determine the point cloud data corresponding to the image to provide conditions for determining the second position information.
[0061] The following will be combined Figure 3 An exemplary principle of determining point cloud data is described.
[0062] Figure 3 It is a schematic diagram of the principle of determining point cloud data according to an embodiment of the present disclosure.
[0063] According to embodiments of the present disclosure, a point cloud generation network can be used to generate point cloud data corresponding to a single image. For example, a sparse point cloud network can be first used to generate sparse point cloud data based on a single image, and then the sparse point cloud data can be input into a dense model (Dense Module) to generate a dense point cloud. The sparse point cloud network can include an encoder and a decoder. The encoder is composed of a convolutional network, and the decoder is composed of a deconvolutional network and a convolutional network. The dense model can process the sparse point cloud data through feature extraction and feature expansion operations.
[0064] According to the embodiments of the present disclosure, Figure 3 As shown, when determining point cloud data, this embodiment 300 may also first use a monocular depth estimation model 320 to process a single image 310 to obtain a depth map 330 of the single image. Subsequently, based on the single image 310 and the depth map 330, the point cloud data 340 corresponding to the single image 310 is determined.
[0065] Among them, the monocular depth estimation model 320 is a generative model, the input is an image, and the output is an image containing depth information. The monocular depth estimation model can be a model built based on deep learning technology. The monocular depth estimation model can be a supervised monocular depth estimation model or a self-supervised monocular depth estimation model. Among them, the supervised monocular depth estimation model needs to be supervised by a real depth map and needs to rely on a high-precision depth sensor to capture real depth information. The self-supervised monocular depth estimation model can use constraints between consecutive frames to predict depth information. The self-supervised monocular depth estimation model can, for example, include a model framework MLDA-Net (Mult-Level Dual Attention-BasedNetwork for Self-Supervised Monocular Depth Estimation). Among them, the model framework MLDA-Net takes a low-resolution color image as input and can estimate the corresponding depth information in a self-supervised manner; the framework uses a multi-level feature fusion (MLFE) strategy to extract rich hierarchical representations from different receptive fields for high-quality depth prediction. This model framework can obtain effective features using a dual attention strategy, which enhances global and local structural information by combining global and local attention modules. This model framework uses a re-weighted strategy to calculate the loss function, re-weighting the depth information of the output of different layers, thereby effectively supervising the depth information of the final output.
[0066] The overall structure of this model framework consists of input data at multiple scales, with the scale being an optional parameter called scales. This input data is processed through two convolutional networks to extract features, which are then integrated and further extracted using an attention network (GA). The convolutional and attention networks form the encoding network structure. After feature extraction, the model framework feeds these features into a second network structure, which primarily performs feature extraction and upsampling based on two attention modules. The final output is depth maps of different scales corresponding to the input images of different scales.
[0067] It is understood that the model framework of the above-mentioned monocular depth estimation model is only used as an example to facilitate understanding of the present disclosure, and the present disclosure does not limit it. For example, the monocular depth estimation model can also be constructed based on Markov random fields, or based on the Monodepth algorithm or the SVS (Single View Stereo Matching) algorithm.
[0068] After obtaining the depth map 330, this embodiment can map the depth map 330 and the single image 310 into three-dimensional point cloud data based on three-dimensional geometric principles. The point cloud data may include three-dimensional coordinates, color information, and reflection intensity information. The color information can be represented, for example, by the RGB values of the corresponding pixels in the single image 310. Using a ray tracing algorithm and the RGB values of the corresponding pixels in the single image 310, the reflection intensity information included in each point cloud data corresponding to each pixel can be calculated.
[0069] This embodiment uses a monocular depth estimation model to generate a depth map corresponding to each image, and then determines point cloud data based on the depth map, which can improve the accuracy and efficiency of determining point cloud data.
[0070] According to an embodiment of the present disclosure, the operation of determining the initial relative posture described above will be expanded and supplemented below.
[0071] Figure 4 2 is a schematic diagram of the principle of determining the initial relative posture according to an embodiment of the present disclosure.
[0072] In embodiment 400, when determining the initial relative pose, a feature extraction algorithm may be first used to extract feature points of each of the two images captured at two different locations to obtain two feature point groups. For example, feature points of a first image 411 of the two images may be extracted to obtain a plurality of first feature points 421, which form a first feature point group. Feature points of a second image 412 of the two images may be extracted to obtain a plurality of second feature points 422, which form a second feature point group.
[0073] Subsequently, this embodiment can use a feature point matching algorithm to determine the matching of multiple first feature points 421 with multiple second feature points 422 to obtain second feature points that match the first feature points. The first feature point and the matched second feature point can form a feature point pair, and a total of multiple feature point pairs 430 can be obtained. It is understood that the number of multiple feature point pairs is less than or equal to the number of multiple first feature points. Among them, the feature point matching algorithm can include, for example, a Brute-Force Matcher (BFM) algorithm, a K-Nearest Neighbor (KNN) algorithm, etc.
[0074] After obtaining a plurality of feature point pairs 430, the initial relative pose 440 can be determined based on the plurality of feature point pairs 430 and the epipolar constraint relationship. For example, the N-point method can be used to solve the essential matrix described above, where N is 8, 16, 24, etc., and the present disclosure does not limit this. The essential matrix can reflect the relationship between a spatial point and the pixel points in the image captured by the acquisition device at different viewing angles. After solving the essential matrix, the essential matrix can be subjected to singular value decomposition to obtain a decomposition result, which includes the rotation matrix s and the translation vector t representing the initial relative pose.
[0075] In one embodiment, after obtaining feature point pairs using a feature point matching algorithm, a RANSAC algorithm or a minimum median search (LMedS) algorithm, for example, can be used to filter the multiple feature point pairs and eliminate incorrectly matched feature point pairs. Accordingly, this embodiment can determine the initial relative pose based on the filtered feature point pairs and the epipolar constraint relationship. In this way, the accuracy of the determined initial relative pose can be improved.
[0076] The disclosed embodiment uses epipolar constraints and multiple feature point pairs to determine the initial relative pose, expressed in optical center distance. Based on this, the disclosed embodiment's method for determining relative pose transforms the initial relative pose based on two pieces of position information, ensuring that the resulting relative pose and point cloud data use the same unit of measurement. This facilitates the splicing of 3D models and improves the accuracy of 3D models of larger scenes.
[0077] According to an embodiment of the present disclosure, after obtaining the rotation matrix R and the translation vector t representing the initial relative pose, the embodiment can use triangulation to restore the spatial position of the spatial point corresponding to the two matched feature points based on the coordinates of the two feature points in each feature point pair and the initial relative pose, and obtain the first position information representing the spatial position. For example, the embodiment can unify the coordinates of the two feature points into the same coordinate system based on the coordinates of the two feature points, and calculate the distance between the two feature points. According to the initial relative pose, the distance between the two optical centers can be obtained. When the focal length of the image acquisition device is known, according to the characteristics of similar triangles, the two depth values of the spatial points corresponding to the two feature points for the two images can be calculated. According to the two depth values and the distance between the two feature points, a spatial point can be uniquely determined, and the spatial point is the spatial point corresponding to the feature point pair, and the coordinate value of the spatial point is the first position information based on the first unit of measurement.
[0078] This embodiment uses a triangulation method to determine the first position information, which can simplify the principle of determining the first position information and improve the efficiency of determining the first position information.
[0079] Based on the posture determination method provided by the present disclosure, the present disclosure also provides a method for generating a three-dimensional model. Figure 5 The method is described in detail.
[0080] Figure 5 It is a flowchart of a method for generating a three-dimensional model according to an embodiment of the present disclosure.
[0081] like Figure 5 As shown, the method 500 for generating a three-dimensional model in this embodiment may include operations S510 to S540.
[0082] In operation S510 , for at least two images of a target scene, a three-dimensional sub-model for each image is generated according to the images.
[0083] According to an embodiment of the present disclosure, a single-view reconstruction method can be used to generate a 3D sub-model for each image. Single-view reconstruction methods are mainly divided into two categories: one is a data-driven reconstruction method, and the other is a constraint-based reconstruction method.
[0084] Constraint-based reconstruction methods mainly use geometric constraints to extract the structural features of objects from images to perform 3D reconstruction. These geometric constraints are established based on prior knowledge and mainly include vanishing points, parallelism, coplanarity, perpendicularity, and perspective relationships.
[0085] Data-driven reconstruction methods use machine learning and other methods to build a target training dataset. From this large amount of target data, they extract the mapping relationship between semantic labels and scene geometry, i.e., the data model. After training the model based on the dataset, data-driven methods analyze the input image and match it to the existing feature model to obtain segmentation results, semantic annotation results, and depth information for the image region. Based on this information, point cloud data for each image can be generated, and the point cloud data of each image can be used to represent the 3D sub-model of each image.
[0086] In operation S520, any one of the at least two images is used as a reference image, and an image pair consisting of the reference image and the other image of the at least two images is determined to obtain at least one image pair.
[0087] In this embodiment, the image captured earliest of at least two images may be used as a reference image. This image may be combined with each of the other images to obtain at least one image pair. Alternatively, in this embodiment, the image captured earlier of two adjacent images may be used as a reference image, and these two adjacent images may be combined to obtain at least one image pair. It will be appreciated that, depending on actual needs, any image may be used as a reference image and combined with the other images to obtain at least one image pair.
[0088] In operation S530 , a relative pose for at least one image pair is determined.
[0089] This operation can use the two images in each image pair as the two images captured at different poses involved in the method for determining relative pose described above. The relative pose of the two images in each image pair can be obtained through the method for determining relative pose described above as the relative pose for each image pair.
[0090] In operation S540 , the 3D sub-models for the other images and the 3D sub-model for the reference image are spliced together according to the relative poses to obtain a 3D model for the target scene.
[0091] In this embodiment, the point cloud data representing the 3D sub-model of the other image can be transformed based on the relative pose so that the 3D sub-model of the other image and the 3D sub-model of the reference image are represented in the same coordinate system. In this embodiment, the transformed point cloud data of the 3D sub-model of the other image can be aggregated with the point cloud data representing the 3D sub-model of the reference image to achieve splicing of the two 3D sub-models.
[0092] It is understood that when the earlier of two adjacent images is used as the reference image, this embodiment can perform operation S540 for each image pair to complete the splicing of the 3D sub-models of the two images in the image pair. Subsequently, based on the relative poses of the two reference images in the two image pairs, the two spliced 3D sub-models of the two image pairs are spliced together to ultimately obtain a 3D model of the target scene.
[0093] It is understandable that the target scene can be selected according to actual needs, for example, it can be a virtual scene of a game, or an indoor architectural scene, etc., and the present disclosure does not limit this.
[0094] This embodiment determines the relative pose between other images and the reference image by adopting the method for determining relative pose described above, and splices three-dimensional sub-models for multiple images based on the relative pose, which can improve the accuracy and realism of the generated three-dimensional model and help improve the user experience.
[0095] Based on the posture determination method provided by the present disclosure, the present disclosure also provides a posture determination device. Figure 6 The device is described in detail.
[0096] Figure 6 4 is a structural block diagram of a posture determination device according to an embodiment of the present disclosure.
[0097] like Figure 6 As shown, the posture determination device 600 of this embodiment includes a posture determination module 610 , a first position determination module 620 , a second position determination module 630 and a posture transformation module 640 .
[0098] Pose determination module 610 is configured to determine an initial relative pose for the two images based on a plurality of feature point pairs from the two images captured at different poses. A feature point pair includes two feature points belonging to each of the two images, and the two feature points match each other. The initial relative pose is expressed based on a first unit of measure. In one embodiment, pose determination module 610 may be configured to perform operation S210 described above, and will not be further described herein.
[0099] The first position determination module 620 is used to determine the first position information of the spatial point corresponding to the feature point pair based on the first measurement unit according to the initial relative position. In one embodiment, the first position determination module 620 can be used to perform the operation S220 described above, which will not be repeated here.
[0100] The second position determination module 630 is configured to determine second position information of the spatial point based on a second measurement unit based on the point cloud data corresponding to the two images, where the point cloud data is represented based on the second measurement unit. In one embodiment, the second position determination module 630 can be configured to perform operation S230 described above, which will not be further described herein.
[0101] The posture transformation module 640 is used to transform the initial relative posture according to the first position information and the second position information to obtain a relative posture represented by a second measurement unit. In one embodiment, the posture transformation module 640 can be used to perform the operation S240 described above, which will not be repeated here.
[0102] According to an embodiment of the present disclosure, the apparatus 600 may further include a depth map generation module and a point cloud data determination module. The depth map generation module is configured to process two images using a monocular depth estimation model to obtain depth maps corresponding to the images. The point cloud data determination module is configured to determine the point cloud data corresponding to the images based on the images and the depth maps.
[0103] According to an embodiment of the present disclosure, the pose determination module 610 may include a matching submodule and a pose determination submodule. The matching submodule is configured to use a feature point matching algorithm to determine, based on multiple first feature points in one of the two images, multiple second feature points in the other image that match the multiple first feature points, thereby obtaining multiple feature point pairs; each feature point pair consists of a first feature point and a matched second feature point. The pose determination submodule is configured to determine an initial relative pose based on the epipolar constraint relationship and the multiple feature point pairs.
[0104] According to an embodiment of the present disclosure, the first position determination module 620 is specifically configured to determine the first position information of the spatial point corresponding to the initial relative pose feature point pair by using a triangulation method.
[0105] According to an embodiment of the present disclosure, the posture transformation module 640 may include a transformation relationship determination submodule and a transformation submodule. The transformation relationship determination submodule is configured to determine a transformation relationship between a first measurement unit and a second measurement unit based on first and second position information of spatial points corresponding to a plurality of feature point pairs. The transformation submodule is configured to transform the initial relative posture based on the transformation relationship to obtain a relative posture based on the second measurement unit.
[0106] Based on the method for generating a three-dimensional model provided by the present disclosure, the present disclosure also provides a device for generating a three-dimensional model. Figure 7 The device is described in detail.
[0107] Figure 7 It is a structural block diagram of a device for generating a three-dimensional model according to an embodiment of the present disclosure.
[0108] like Figure 7 As shown, the three-dimensional model generation device 700 of this embodiment may include a model generation module 710 , an image pair determination module 720 , a posture determination module 730 and a model stitching module 740 .
[0109] The model generation module 710 is used to generate a 3D sub-model for the at least two images of the target scene according to the images. In one embodiment, the model generation module 710 can be used to perform the operation S510 described above, which will not be repeated here.
[0110] The image pair determination module 720 is configured to use any one of the at least two images as a reference image and determine an image pair consisting of the reference image and the other image in the at least two images, thereby obtaining at least one image pair. In one embodiment, the image pair determination module 720 may be configured to perform operation S520 described above, which will not be further described herein.
[0111] The pose determination module 730 is configured to determine a relative pose for at least one image pair. The relative pose may be determined using the apparatus for determining relative pose described above. In one embodiment, the pose determination module 730 may be configured to perform operation S530 described above, which will not be described in detail herein.
[0112] The model stitching module 740 is used to stitch the 3D sub-models for the other images and the 3D sub-model for the reference image according to the relative pose to obtain a 3D model for the target scene. In one embodiment, the model stitching module 740 can be used to perform the operation S540 described above, which will not be repeated here.
[0113] It should be noted that the collection, storage, use, processing, transmission, provision, disclosure, and application of user personal information in the technical solutions disclosed herein comply with relevant laws and regulations, employ necessary confidentiality measures, and do not violate public order and good morals. In the technical solutions disclosed herein, user authorization or consent is obtained before obtaining or collecting user personal information.
[0114] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0115] Figure 8A schematic block diagram of an example electronic device 800 that can be used to implement the method for determining the pose and / or the method for generating a three-dimensional model of an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0116] like Figure 8 As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0117] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0118] The computing unit 801 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the method for determining the pose and / or the method for generating a three-dimensional model. For example, in some embodiments, the method for determining the pose and / or the method for generating a three-dimensional model can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the method for determining the pose and / or the method for generating a three-dimensional model described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the method for determining the position and / or the method for generating the three-dimensional model in any other appropriate manner (for example, by means of firmware).
[0119] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0120] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0121] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0122] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0123] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0124] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system. It solves the problems of difficult management and poor business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server in a distributed system, or a server integrated with blockchain.
[0125] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0126] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for determining a posture, comprising: Determining an initial relative pose for two images based on a plurality of feature point pairs of two images captured at different poses, wherein the feature point pairs include two feature points belonging to the two images, respectively, and the two feature points match each other; and the initial relative pose is expressed based on a first unit of measurement; Determining, according to the initial relative pose, first position information of the spatial point corresponding to the feature point pair based on the first measurement unit; determining, based on the point cloud data corresponding to the two images, second position information of the spatial point based on a second measurement unit, wherein the point cloud data is represented based on the second measurement unit; determining a transformation relationship between the first measurement unit and the second measurement unit according to the first position information and the second position information of the spatial points corresponding to the plurality of feature point pairs; and The initial relative posture is transformed according to the transformation relationship to obtain a relative posture based on the second measurement unit.
2. The method according to claim 1, further comprising: For the image, use a monocular depth estimation model to process the image to obtain a depth map corresponding to the image; as well as Determine point cloud data corresponding to the image based on the image and the depth map.
3. The method according to claim 1, wherein Determining the initial relative poses of the two images based on a plurality of feature point pairs of the two images collected at different poses includes: For the plurality of first feature points of one of the two images, a feature point matching algorithm is used to determine a plurality of second feature points of the other of the two images that match the plurality of first feature points one-to-one, thereby obtaining a plurality of feature point pairs; the feature point pairs are composed of the first feature points and the matched second feature points; and The initial relative pose is determined according to the epipolar constraint relationship and the plurality of feature point pairs.
4. The method according to claim 1, wherein The determining, based on the initial relative pose, first position information of the spatial point corresponding to the feature point pair based on the first measurement unit includes: According to the initial relative position and the feature point pair, a triangulation method is used to determine first position information of a spatial point corresponding to the feature point pair.
5. A method for generating a three-dimensional model, comprising: For at least two images of a target scene, generating three-dimensional sub-models for the images according to the images; Taking any one of the at least two images as a reference image, determining an image pair consisting of the reference image and the other of the at least two images, to obtain at least one image pair; determining a relative pose for the at least one image pair; as well as splicing the three-dimensional sub-model for the other image and the three-dimensional sub-model for the reference image according to the relative pose to obtain a three-dimensional model for the target scene, Wherein, the relative posture is determined using the method described in any one of claims 1 to 4.
6. A device for determining a posture, comprising: a pose determination module configured to determine an initial relative pose for two images based on a plurality of feature point pairs of two images captured at different poses, wherein the feature point pairs include two feature points belonging to the two images, respectively, and the two feature points match each other; and the initial relative pose is expressed based on a first unit of measurement; A first position determination module is configured to determine, based on the initial relative posture, first position information of the spatial point corresponding to the feature point pair in the first measurement unit; a second position determination module, configured to determine second position information of the spatial point based on a second measurement unit based on point cloud data corresponding to the two images, wherein the point cloud data is represented based on the second measurement unit; a transformation relationship determination submodule, configured to determine a transformation relationship between the first measurement unit and the second measurement unit based on the first position information and the second position information of the spatial points corresponding to the plurality of feature point pairs; and A transformation submodule is used to transform the initial relative posture according to the transformation relationship to obtain a relative posture based on the second measurement unit.
7. The apparatus according to claim 6, further comprising: a depth map generation module, configured to process the image using a monocular depth estimation model to obtain a depth map corresponding to the image; and The point cloud data determination module is used to determine the point cloud data corresponding to the image based on the image and the depth map.
8. The device according to claim 6, wherein The posture determination module includes: a matching submodule, configured to determine, based on the plurality of first feature points of one of the two images, a plurality of second feature points of the other of the two images that match the plurality of first feature points one-to-one with the plurality of first feature points using a feature point matching algorithm, to obtain a plurality of feature point pairs; the feature point pairs consisting of the first feature points and the matched second feature points; and The posture determination submodule is used to determine the initial relative posture according to the epipolar constraint relationship and the plurality of feature point pairs.
9. The device according to claim 6, wherein The first location determination module is configured to: According to the initial relative position and the feature point pair, a triangulation method is used to determine first position information of a spatial point corresponding to the feature point pair.
10. A device for generating a three-dimensional model, comprising: A model generation module, configured to generate, for at least two images of a target scene, three-dimensional sub-models for the images according to the images; an image pair determination module, configured to use any one of the at least two images as a reference image and determine an image pair consisting of the reference image and the other of the at least two images, thereby obtaining at least one image pair; a pose determination module for determining a relative pose for the at least one image pair; as well as A model splicing module is used to splice the three-dimensional sub-model for the other image and the three-dimensional sub-model for the reference image according to the relative posture to obtain a three-dimensional model for the target scene. Wherein, the relative posture is determined using the device according to any one of claims 6 to 9.
11. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 5.
13. A computer program product, comprising a computer program / instructions, wherein the computer program / instructions are stored on at least one of a readable storage medium and an electronic device, and when the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Image positioning method and device based on ray model three-dimensional reconstruction
CN105844696A
A global motion initialization method and system for aerial image three-dimensional reconstruction
CN109493415A