A point cloud model construction method and device

By acquiring multiple 2D image data and performing feature extraction and filtering, the problem of low accuracy of point cloud models in monocular 3D detection technology is solved, achieving accurate restoration of object point cloud models and memory optimization.

CN117253051BActive Publication Date: 2026-04-21MIGU COMIC CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
MIGU COMIC CO LTD
Filing Date
2023-10-18
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In existing technologies, point cloud models constructed by monocular 3D detection technology have low accuracy, especially when the objects are unevenly distributed, resulting in gaps in the point cloud model of the objects and making it impossible to accurately restore the three-dimensional geometry of the objects.

Method used

By acquiring multiple two-dimensional image data of the target object, the target point cloud is extracted using camera parameters and pose, and multi-view feature extraction and feature point set filtering are performed to construct a point cloud model, ensuring that the point cloud is distributed on the object surface and improving the accuracy of the point cloud model.

Benefits of technology

Without increasing the number of point clouds in 3D space, the accuracy of the point cloud model is effectively improved, memory usage is reduced, and the point cloud model of the object can completely restore its 3D geometry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117253051B_ABST
    Figure CN117253051B_ABST
Patent Text Reader

Abstract

This application provides a point cloud model construction method and apparatus. The method includes: acquiring two-dimensional image data of a target object, the two-dimensional image data including multiple two-dimensional images taken from different perspectives, and camera parameters and poses corresponding to each two-dimensional image; extracting target point clouds of each two-dimensional image based on the camera parameters and poses corresponding to each two-dimensional image, the target point clouds being distributed on at least three point clouds on the surface of the target object; performing multi-view feature extraction on each two-dimensional image based on the target point clouds of each two-dimensional image to obtain a set of feature points for each two-dimensional image; acquiring a set of feature points of a subset of the multiple two-dimensional images, and filtering the set of feature points of the subset of the two-dimensional images to obtain target feature points of each two-dimensional image in the subset of the two-dimensional images; and constructing a point cloud model of the target object based on the target point clouds, the set of feature points, and the target feature points of the multiple two-dimensional images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of stereoscopic imaging technology, specifically to a method and apparatus for constructing point cloud models. Background Technology

[0002] Monocular 3D detection technology is a commonly used technique for constructing point cloud models in 3D space. In monocular 3D detection, a point cloud model in 3D space is constructed from images viewed from multiple perspectives. In existing technologies, the points in the constructed point cloud model are uniformly distributed in 3D space. However, there are instances where the point cloud representation of objects in 3D space is insufficient, leading to gaps in the object's point cloud model and preventing the creation of a complete object model. This results in low accuracy in constructing the object's point cloud model.

[0003] It is evident that existing technologies suffer from low accuracy in creating point cloud models of objects. Summary of the Invention

[0004] This application provides a point cloud model construction method and apparatus to solve the problem of low accuracy in building point cloud models of objects in the prior art.

[0005] To solve the above problems, this application is implemented as follows:

[0006] In a first aspect, embodiments of this application provide a point cloud model construction method, including:

[0007] Acquire two-dimensional image data of the target object, wherein the two-dimensional image data includes multiple two-dimensional images taken from different perspectives, as well as the camera parameters and pose corresponding to each two-dimensional image;

[0008] Based on the camera parameters and pose corresponding to each two-dimensional image, the target point cloud of each two-dimensional image is extracted, and the target point cloud is distributed on the surface of the target object by at least three point clouds;

[0009] Based on the target point cloud of each two-dimensional image, multi-view feature extraction is performed on each two-dimensional image to obtain the feature point set of each two-dimensional image;

[0010] Obtain feature point sets of a subset of the multiple two-dimensional images, and filter the feature point sets of the subset of two-dimensional images to obtain the target feature points of each two-dimensional image in the subset of two-dimensional images;

[0011] Based on the target point cloud, the set of feature points, and the target feature points of the multiple two-dimensional images, a point cloud model of the target object is constructed.

[0012] Secondly, embodiments of this application also provide a point cloud model construction apparatus, comprising:

[0013] The acquisition module is used to acquire two-dimensional image data of the target object. The two-dimensional image data includes multiple two-dimensional images with different shooting angles, as well as the camera parameters and pose corresponding to each two-dimensional image.

[0014] The first processing module is used to extract the target point cloud of each two-dimensional image based on the camera parameters and pose corresponding to each two-dimensional image, wherein the target point cloud is distributed on the surface of the target object by at least three point clouds.

[0015] The second processing module is used to perform multi-view feature extraction on each two-dimensional image based on the target point cloud of each two-dimensional image to obtain a feature point set of each two-dimensional image.

[0016] The third processing module is used to obtain a set of feature points of a portion of the multiple two-dimensional images, and filter the set of feature points of the portion of the two-dimensional images to obtain the target feature points of each two-dimensional image in the portion of the two-dimensional images.

[0017] A construction module is used to construct a point cloud model of the target object based on the target point cloud, the set of feature points, and the target feature points of the multiple two-dimensional images.

[0018] Thirdly, embodiments of this application also provide an electronic device, including a transceiver and a processor.

[0019] The transceiver is used to acquire two-dimensional image data of the target object. The two-dimensional image data includes multiple two-dimensional images with different shooting angles, as well as camera parameters and poses corresponding to each two-dimensional image.

[0020] The processor is used to extract the target point cloud of each two-dimensional image based on the camera parameters and pose corresponding to each two-dimensional image, wherein the target point cloud is distributed on the surface of the target object by at least three point clouds.

[0021] The processor is further configured to perform multi-view feature extraction on each two-dimensional image based on the target point cloud of each two-dimensional image, thereby obtaining a feature point set for each two-dimensional image;

[0022] The processor is further configured to acquire a set of feature points of a portion of the plurality of two-dimensional images, and filter the set of feature points of the portion of the two-dimensional images to obtain the target feature points of each two-dimensional image in the portion of the two-dimensional images;

[0023] The processor is further configured to construct a point cloud model of the target object based on the target point cloud, the set of feature points, and the target feature points of the plurality of two-dimensional images.

[0024] Fourthly, embodiments of this application also provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of the point cloud model construction method as described in the first aspect above.

[0025] Fifthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the point cloud model construction method described in the first aspect above.

[0026] In this embodiment, two-dimensional image data of the target object is acquired. This data includes multiple two-dimensional images taken from different perspectives, as well as camera parameters and poses corresponding to each image. Target point clouds are extracted from each two-dimensional image based on its camera parameters and poses. These target point clouds are distributed across at least three points on the surface of the target object. Multi-view feature extraction is performed on each two-dimensional image based on its target point cloud, resulting in a feature point set for each image. The feature points in this set are then distributed across the surface of the target object. A feature point set is obtained from a subset of the multiple two-dimensional images, and this set is filtered to obtain target feature points for each image. Constraints are applied to the feature points in the feature point set using the target point clouds and target feature points from the multiple two-dimensional images. A point cloud model of the target object is then constructed, ensuring that all points in the point cloud model are distributed across the surface of the target object. This increases the number of points in the point cloud model and thus improves its accuracy. Simultaneously, the point cloud model does not increase the number of points outside the target object in three-dimensional space, effectively reducing the memory footprint of the point cloud model. Attached Figure Description

[0027] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 This is a structural diagram of a point cloud model in the prior art provided in the embodiments of this application;

[0029] Figure 2 This is a flowchart of a point cloud model construction method provided in an embodiment of this application;

[0030] Figure 3 This is a distribution map of the target point cloud and target feature points provided in the embodiments of this application;

[0031] Figure 4 This is a schematic diagram of the three-dimensional space provided in the embodiments of this application;

[0032] Figure 5 This is a flowchart of constructing a point cloud model provided in an embodiment of this application;

[0033] Figure 6 This is a structural diagram of a point cloud model construction device provided in an embodiment of this application;

[0034] Figure 7 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0035] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0036] In existing technologies, point cloud models are constructed using monocular 3D detection. Specifically, this involves aggregating uniformly divided 3D meshes as anchor points (e.g., dividing the 3D space into an n*n mesh, with each mesh's feature points used as anchor points). By using the 3D mesh as anchor points, visual features are accumulated, and the 3D geometry is preserved. However, in these related technologies, due to the small field of view, large objects cannot be captured, resulting in a limited number of point clouds on the objects after construction, making it impossible to build a complete point cloud model. Figure 1 As shown, the reconstructed object shape is incomplete and cannot be accurately reproduced. It should be understood that in existing technologies, to increase the number of point clouds on an object, the 3D space needs to be divided into smaller grids, which results in the point cloud model consuming a large amount of memory.

[0037] In this embodiment, the point cloud model of the object is constructed by using the point cloud distributed on the surface of the object as anchor points. In this way, the point cloud model of the object can be effectively established without increasing the number of point clouds, so as to accurately restore the three-dimensional geometric shape of the object. See the following embodiments for details.

[0038] Please see Figure 2 , Figure 2 This is a flowchart of a point cloud model construction method provided in an embodiment of this application, such as... Figure 2 As shown, it includes the following steps:

[0039] Step 201: Obtain two-dimensional image data of the target object. The two-dimensional image data includes multiple two-dimensional images with different shooting angles, as well as the camera parameters and pose corresponding to each two-dimensional image.

[0040] The target object mentioned above is an object in three-dimensional space, and the point cloud model to be constructed in this application is the point cloud model of the object. The two-dimensional image mentioned above is a two-dimensional image obtained by the camera taking pictures of the object in three-dimensional space from different perspectives, that is, a three-dimensional red-green-blue (RGB) image. The camera parameters mentioned above are the shooting parameters of the camera when taking two-dimensional images, such as focal length, resolution, etc. The pose mentioned above is the pose of the camera when taking two-dimensional images, used to characterize the shooting angle when shooting the target object.

[0041] It should be understood that monocular depth estimation can be performed using the two-dimensional image data of the target object to obtain the depth estimation result of each two-dimensional image. Then, the specific position of the object in the two-dimensional image can be determined by the depth estimation result. Point cloud data can be collected on the surface of the object as anchor points to construct a point cloud model of the target object.

[0042] Step 202: Extract the target point cloud of each two-dimensional image based on the camera parameters and pose corresponding to each two-dimensional image. The target point cloud is distributed on the surface of the target object by at least three point clouds.

[0043] The aforementioned target point cloud consists of at least three point clouds distributed on the target object. The features of each of the at least three point clouds are the features of the object's surface. The point cloud of the target object's surface is collected as the target point cloud. Then, the feature that is similar to the target point cloud is extracted from each two-dimensional image. The point cloud model of the target object is constructed by using the corresponding positions of similar features as point clouds.

[0044] In this model, the target point cloud serves as the anchor point. By leveraging similar features between the target point cloud and the 2D image data, a point cloud model of the target object can be constructed. It should be understood that the constructed point cloud model is a 3D point cloud model, with the target point cloud acting as the anchor point. During the point cloud acquisition process, at least three point clouds need to be collected so that the point cloud model constructed from the target point cloud can reflect the geometric shape characteristics of the target object in 3D space.

[0045] It should be understood that the target point cloud of each two-dimensional image is extracted based on the camera parameters and pose corresponding to each two-dimensional image. Specifically, the target point cloud of an image is extracted based on the camera parameters and position of that image, as detailed in the following embodiments.

[0046] Step 203: Based on the target point cloud of each two-dimensional image, perform multi-view feature extraction on each two-dimensional image to obtain the feature point set of each two-dimensional image.

[0047] The aforementioned feature point set includes multiple feature points, each representing a location point in a 2D image. Each feature point characterizes the features of that location point in the 2D image, and the features corresponding to each feature point are similar to those in the target point cloud. The aforementioned multi-view feature extraction is performed on each 2D image based on the target point cloud of each image, resulting in a feature point set for each 2D image. Specifically, based on the target point cloud, camera parameters, and pose of each 2D image, multi-view feature extraction is performed on each image to obtain a feature point set for each 2D image.

[0048] It should be understood that by performing multi-view feature extraction on each two-dimensional image, a set of feature points is obtained. The features corresponding to the feature points in the feature point set are similar to the features of the target point cloud. It can be assumed that the feature points in the feature point set are distributed on the surface of the target object. Then, a point cloud model of the target object can be established through the feature point set. At this time, only the point cloud of the target object is added to establish a complete point cloud model of the target object. At the same time, no point cloud is added for locations in three-dimensional space where the target object does not exist, thus reducing the memory usage.

[0049] The above-mentioned multi-view feature extraction is performed on each two-dimensional image based on the target point cloud of each image, resulting in a feature point set for each two-dimensional image. In other words, by using the target point cloud of an image, multi-view feature extraction is performed on that image to obtain the image's feature points, which can be one or more. By performing multi-view feature extraction on each two-dimensional image based on the target point cloud of each image, the feature point set for each two-dimensional image can be obtained.

[0050] Step 204: Obtain the feature point set of a portion of the multiple two-dimensional images, and filter the feature point set of the portion of the two-dimensional images to obtain the target feature point of each two-dimensional image in the portion of the two-dimensional images.

[0051] The aforementioned target feature point is a feature point in a partial feature point set. This target feature point is used to characterize the average feature distribution at different locations of the target object. By using the target feature point, each feature point in the feature point set can be constrained, the position of each feature point in the feature point set can be optimized, and the position of some discrete feature points can be corrected to the surface of the target object, thereby improving the accuracy of the established point cloud model.

[0052] Furthermore, the feature point set of the aforementioned partial two-dimensional image is the set of feature points of each two-dimensional image in the partial two-dimensional image. By filtering the feature point set of the partial two-dimensional image, the target feature points are obtained. Compared with filtering the feature point set of the entire two-dimensional image, the amount of data that needs to be processed is effectively reduced, and the filtering efficiency is improved.

[0053] Step 205: Based on the target point cloud, the set of feature points, and the target feature points of the multiple two-dimensional images, construct a point cloud model of the target object.

[0054] The aforementioned target point cloud consists of at least three point clouds distributed on the target object. The aforementioned target feature points are feature points that represent the average features of the target object. Each feature point in the feature point set is constrained by the target point cloud and the target feature points, so that each constrained feature point can represent the features of different positions of the target object. Then, a point cloud model of the target object is constructed by each constrained feature point. This point cloud model is a complete model of the target object, and the geometric shape of the target object in three-dimensional space can be restored through the point cloud model.

[0055] For example, such as Figure 3 As shown, both the target point cloud and the target feature points are distributed on the surface of the target object. By constraining the feature points in the feature point set through the target point cloud and the target feature points, some discrete feature points in the feature point set (i.e. feature points not distributed on the target object) are corrected to be distributed on the surface of the target object, thereby improving the accuracy of the established point cloud model.

[0056] In this embodiment, two-dimensional image data of the target object is acquired. This data includes multiple two-dimensional images taken from different perspectives, as well as camera parameters and poses corresponding to each image. Target point clouds are extracted from each two-dimensional image based on its camera parameters and poses. These target point clouds are distributed across at least three points on the surface of the target object. Multi-view feature extraction is performed on each two-dimensional image based on its target point cloud, resulting in a feature point set for each image. The feature points in this set are then distributed across the surface of the target object. A feature point set is obtained from a subset of the multiple two-dimensional images, and this set is filtered to obtain target feature points for each image. Constraints are applied to the feature points in the feature point set using the target point clouds and target feature points from the multiple two-dimensional images. A point cloud model of the target object is then constructed, ensuring that all points in the point cloud model are distributed across the surface of the target object. This increases the number of points in the point cloud model and thus improves its accuracy. Simultaneously, the point cloud model does not increase the number of points outside the target object in three-dimensional space, effectively reducing the memory footprint of the point cloud model.

[0057] In one embodiment, extracting the target point cloud of each two-dimensional image based on the camera parameters and pose corresponding to each two-dimensional image includes:

[0058] Based on the camera parameters and pose corresponding to each two-dimensional image, monocular depth estimation is performed on each two-dimensional image to obtain the depth estimation result of each two-dimensional image;

[0059] The depth estimation results of each two-dimensional image are back-projected, and the target point cloud of each two-dimensional image is obtained by filtering and sampling based on the back-projection results.

[0060] The depth estimation results described above are used to characterize the distances from different locations in a 2D image to the camera. It should be understood that when a target object exists in 3D space, the distance from the target object to the camera differs from the distances from other objects in 3D space to the camera. The depth estimation results can identify the location of the target object in the 2D image, so as to sample the target point cloud at the location of the target object.

[0061] The depth estimation results of each two-dimensional image are back-projected so that the back-projected two-dimensional image can represent the depth at different locations, thereby realizing the acquisition of depth at different locations in at least three point clouds of the target object as the target point cloud.

[0062] For example, such as Figure 4 As shown, only the target object and the background wall exist in the 3D space. The target object is the table in the figure, which is obtained through monocular depth estimation. Figure 4 The depth estimation results show that the distance from the table to the camera is closer, while the distance from the background wall to the camera is farther. First, let's... Figure 4 The 2D image shown is segmented to obtain multiple regions. A point cloud is set for each region, resulting in multiple point clouds. Based on the depth estimation results, at least three point clouds are collected from the multiple point clouds. The depth of at least three point clouds is close to the depth of the table in the depth estimation results. It can be assumed that these at least three point clouds are distributed on the surface of the table in the 2D image.

[0063] In the backprojection process, the depth obtained from monocular depth estimation of each 2D image is backprojected into a 3D point cloud. To maintain a manageable amount of 2D image data, duplicate points that are too close to existing points need to be discarded; that is, the features of the three resulting point clouds must be distinct. For example, several hypothetical point clouds are sampled along the depth direction from each pixel of the 2D image, increasing the number of target point clouds several times. During backprojection, random sampling is performed on the surface of the target object to obtain at least three point clouds, ensuring that the features of these at least three sampled point clouds are significantly different.

[0064] In this embodiment, monocular depth estimation is performed on each two-dimensional image based on the camera parameters and pose corresponding to each two-dimensional image to obtain the depth estimation result of each two-dimensional image. Then, the depth estimation result of each two-dimensional image is back-projected, and the target point cloud of each two-dimensional image is obtained by filtering and sampling based on the back-projection result. The depth corresponding to the target point cloud is similar to the depth of the target object in the two-dimensional image, thereby determining the point cloud of the surface of the target object in the acquired two-dimensional image.

[0065] In one embodiment, the step of performing monocular depth estimation on each two-dimensional image based on the camera parameters and pose corresponding to each two-dimensional image to obtain the depth estimation result of each two-dimensional image includes:

[0066] The camera parameters and pose corresponding to each two-dimensional image, as well as each two-dimensional image, are input into an ordinal regression function for calculation to obtain multiple initial depths for each two-dimensional image. Each initial depth is the depth of a different region in the two-dimensional image.

[0067] Calculate the target loss for each region;

[0068] Redundant depths are removed from the plurality of initial depths to obtain the depth estimation result, wherein the target loss of the region corresponding to the redundant depth is less than a set threshold.

[0069] It should be understood that monocular depth estimation can calculate depth using multiple regression functions, with sequential regression typically chosen to calculate depth in two-dimensional images. However, target objects are continuous in three-dimensional space, and the depth calculated using sequential regression is discrete, which does not match the actual target object (the depth of different regions on the target object is continuous), resulting in inaccurate depth estimation results for certain objects. In this embodiment, an ordinal regression function is used instead of the sequential regression function. The depth output by the ordinal regression function is continuous, thus improving the accuracy of the depth estimation results.

[0070] Furthermore, after calculating the initial depth of different regions using ordinal regression functions, the target loss for each region is calculated, and the reliability of the calculated initial depth is measured by the target loss. Specifically, if the target loss of a region is greater than or equal to a set threshold, the depth of the region is considered reliable; if the target loss of a region is less than the set threshold, the depth of the region is considered unreliable and needs to be deleted to ensure the reliability of the depth estimation results.

[0071] The target loss mentioned above can be an ordinal loss, or it can be a loss optimized by using real image features and real depth. The specific optimization process is described in subsequent embodiments.

[0072] In this embodiment, by inputting the camera parameters and pose corresponding to each 2D image, and each 2D image into an ordinal regression function for calculation, multiple initial depths are obtained for each 2D image. Each initial depth represents the depth of a different region in the 2D image, making the initial depths continuous. This effectively represents the distance from the target object to the camera, improving the accuracy of the depth estimation results. Simultaneously, by calculating the target loss for each region and deleting redundant depths from the multiple initial depths, the depth estimation results are obtained. The target loss of the region corresponding to the redundant depth is less than a set threshold, thereby filtering out regions with target losses less than the set threshold and retaining regions with target losses greater than or equal to the set threshold, improving the reliability of the depth estimation results.

[0073] In one embodiment, calculating the target loss for each region includes:

[0074] Obtain the image features and true depth map of each region in each two-dimensional image;

[0075] Convolve the image features and initial depth of each region to obtain the positional residual;

[0076] The ordinal loss corresponding to each two-dimensional image is calculated.

[0077] The sum of the initial depth and position residuals of each region, the norm of the difference between the actual depth and the sum of the ordinal loss corresponding to each two-dimensional image are set as the target loss for each region.

[0078] The aforementioned image features are the features of a region in a two-dimensional image, and the aforementioned true depth map is the true depth map corresponding to the two-dimensional image, used to characterize the true depth of the two-dimensional image in different regions. It should be understood that this application employs the calculation of target loss through the positional residual of depth to improve the accuracy of the target loss. This requires determining the positional residual between the initial depth and the image features and true depth map of each region, and then calculating the target loss using the positional residual. Compared to directly using the loss function corresponding to ordinal regression, the embodiments of this application, by introducing real image features and real depth, can further improve the accuracy of the calculated target loss.

[0079] The positional residual is calculated using the following formula:

[0080] Formula 1: D' = Convs(D, F)

[0081] In Formula 1, D' represents the positional residual, Convs() is used to characterize the convolution calculation, D is the initial depth of the region, and F is the image feature.

[0082] The ordinal loss described above represents the loss for each region when calculating different regional values ​​based on the ordinal regression function. The target loss obtained through the ordinal loss is expressed by the following formula:

[0083] Formula 2: L depth =L ordinal +||D+D'-D * ||

[0084] In Formula 2, L ordinal D represents the average ordinal loss of each region in the entire two-dimensional image. * L represents the true depth of the region. depth This is the target loss. By incorporating realistic image features and depth, the accuracy of the calculated target loss can be further improved.

[0085] In this embodiment, the image features and true depth map of each region in each two-dimensional image are obtained; the image features and initial depth of each region are convolved to obtain the position residual; then the ordinal loss corresponding to each two-dimensional image is calculated; finally, the sum of the initial depth and position residual of each region, the norm of the difference between the actual depth and the actual depth, and the sum of the ordinal loss corresponding to each two-dimensional image are set as the target loss of each region. This realizes the correction of the ordinal loss of the two-dimensional image by the actual image features and actual depth, which can further improve the accuracy of the calculated target loss.

[0086] In one embodiment, obtaining the feature point set of a portion of the plurality of two-dimensional images includes the following:

[0087] Lightweight data is acquired, and multi-view feature extraction is performed on each two-dimensional image in the partial images based on the target point cloud of the partial images to obtain the feature point set of the partial two-dimensional images. The lightweight data includes partial images in the multiple two-dimensional images.

[0088] The feature point set of each two-dimensional image is filtered to obtain the feature point set of the partial two-dimensional image.

[0089] The aforementioned lightweight data is extracted from multiple two-dimensional images. The two-dimensional images in the lightweight data also include the target object, and the features of the two-dimensional images in the lightweight data are similar to the features in the multiple two-dimensional images. The target feature points can be determined through lightweight data, reducing the number of feature points, thereby reducing the computational burden and improving computational efficiency.

[0090] The flowchart for constructing the point cloud model is as follows: Figure 5As shown, all 2D image data are input into the 2D main chain. After monocular depth estimation, back projection and point cloud sampling are performed to obtain the target point cloud. Based on the 2D data and the target point cloud, multi-view feature extraction is performed on all 2D images to obtain a set of feature points. In the lightweight 2D main chain, lightweight data is input, and multi-view feature extraction is performed on each 2D image in the partial images based on the target point cloud of the partial images to obtain a set of feature points for the partial 2D images. Then, the feature point set of the partial 2D images is filtered to obtain the target feature points.

[0091] It should be understood that since performing multi-view feature extraction on all 2D images is the same process as performing multi-view feature extraction on 2D images in lightweight data, it is also possible to directly extract the feature point set corresponding to a portion of the 2D image from the feature point set, and then filter it to obtain the target feature points. This reduces one multi-view feature extraction step, further reducing the computational burden and improving computational efficiency.

[0092] In this embodiment, lightweight data is acquired, and multi-view feature extraction is performed on each two-dimensional image in the partial images based on the target point cloud of the partial images to obtain a feature point set of the partial two-dimensional images. The lightweight data includes partial images from multiple two-dimensional images, thereby reducing the number of feature points that need to be filtered, thus reducing the computational burden and improving computational efficiency. Alternatively, by filtering the feature point set of each two-dimensional image, a feature point set of the partial two-dimensional images can be obtained, eliminating one multi-view feature extraction process and further reducing the computational burden and improving computational efficiency.

[0093] In one embodiment, filtering the feature point set of the partial two-dimensional image to obtain the target feature points of each two-dimensional image in the partial two-dimensional image includes:

[0094] Obtain the projection position of each feature point in the feature point set of the partial two-dimensional image in each two-dimensional image;

[0095] Feature extraction is performed on each projection location to obtain the features of each projection location;

[0096] The features of each projection location are aggregated to obtain the aggregated features;

[0097] The target feature points are obtained by filtering the feature point set of the partial two-dimensional image based on the aggregated features.

[0098] The aforementioned aggregated features are the features obtained by aggregating the features of each feature point in each two-dimensional image. These features can be used to filter the features of a target object in three-dimensional space. It should be understood that for a point in three-dimensional space, its position and features will differ depending on the two-dimensional image taken from different viewing angles. In this embodiment, feature points are enhanced by multi-view image features, that is, target feature points are obtained by filtering through aggregated features. The target feature points can characterize the average features of the target object at different viewing angles and positions. Furthermore, each feature point is constrained by the target feature points to correct each feature point.

[0099] Specifically, the projected position of each feature point in each 2D image is obtained, and the feature at that projected position is also obtained. For example, the projected positions of feature point X in N 2D images are obtained, denoted as P, where P = [p0, p1, ..., p2]. N-1 ]∈R N×2 Where p0, p1, ..., p N-1 The projection position of each feature point in each 2D image is defined. The feature at each projection position is denoted as F, where F = [f0, f1, ..., f2]. N-1 ]∈R N×C Where, f0, f1, ..., f N-1 For each projection location, the features are defined, where C is the number of feature points in the feature point set. The features at each projection location are then aggregated to obtain aggregated features.

[0100] Furthermore, aggregated features can be obtained in the following way:

[0101] The features at each projection location are averaged and aggregated to obtain the aggregated features;

[0102] The variance of the features at each projection location is aggregated to obtain the aggregated features.

[0103] The above method is represented by the following formula three or formula four:

[0104] Formula 3:

[0105] Formula 4:

[0106] Formula 3 is used to calculate the mean, and Formula 4 is used to calculate the variance. In the formula, η(M) represents the number of projection positions (i.e., projection positions within the two-dimensional images) of the i-th feature point in the N two-dimensional images, and M is the mask of the projection positions, which is a pre-set value.

[0107] It should be understood that after calculating the mean and variance of the features, the target feature point is obtained by filtering the C feature points using the mean and variance of the features. This filtering method can either select the feature point from the C feature points that is closest to the mean, select the feature point with the smallest variance as the target feature point, or combine the mean and variance to filter the C feature points to obtain the target feature point. For example, 50% can be filtered out first using the mean, and then the feature point with the smallest variance can be selected from the remaining 50% as the target feature point.

[0108] In this embodiment, the projection position of each feature point in the feature point set of a partial two-dimensional image in each two-dimensional image is obtained, and then feature extraction is performed on each projection position to obtain the feature of each projection position; the features of each projection position are aggregated to obtain aggregated features; based on the aggregated features, the feature point set of the partial two-dimensional image is filtered to obtain target feature points, which can characterize the average feature of the target object at different viewpoints and positions, so that each feature point can be constrained by the target feature points to correct each feature point and improve the accuracy of the established point cloud model.

[0109] In one embodiment, the step of extracting features from each projection location to obtain the features of each projection location includes:

[0110] When the target projection position is inside the two-dimensional image, feature extraction is performed on the target projection to obtain the features of the target projection position;

[0111] When the target projection position is outside the two-dimensional image, the features corresponding to the target image are interpolated to obtain the features of the target projection position. The pose of the target image is adjacent to the pose of the two-dimensional image corresponding to the target projection position. The target image is one of the multiple two-dimensional images.

[0112] It should be understood that a two-dimensional image is an image of a target object captured by a camera in three-dimensional space at different times. Due to the limitations of the shooting angle, a feature point on the target object may or may not have a projected position in different two-dimensional images. When a projected position exists in a two-dimensional image (i.e., the projected position is within the two-dimensional image), feature extraction can be performed directly on that feature point to obtain its features at the projected position. However, when a projected position does not exist (i.e., the projected position is outside the two-dimensional image), it is necessary to determine the corresponding features in the two-dimensional image through adjacent images.

[0113] Specifically, during the process of capturing two-dimensional images, the target object is photographed from different angles. For each two-dimensional image, there are multiple two-dimensional images with adjacent shooting perspectives. Since the shooting perspective changes relatively little, the features of two-dimensional images without projection positions can be determined by using the two-dimensional images with adjacent shooting perspectives. The acquired target images are at least two images with projection positions among the multiple two-dimensional images with adjacent shooting perspectives. The features of the projection positions in the target images are then extracted, and interpolation is performed based on these features to obtain the features of the two-dimensional images without projection positions. It should be understood that the interpolation can be linear interpolation or bilinear interpolation.

[0114] In this embodiment, when the target projection position is inside the two-dimensional image, feature extraction is performed on the target projection to obtain the feature of the target projection position; when the target projection position is outside the two-dimensional image, interpolation is performed on the feature corresponding to the target image to obtain the feature of the target projection position. The pose corresponding to the target image is adjacent to the pose of the two-dimensional image corresponding to the target projection position. The target image is one of multiple two-dimensional images, thereby obtaining the feature of the projection position of the feature point in each two-dimensional image to realize subsequent calculation of aggregated features.

[0115] In one embodiment, constructing a point cloud model of the target object based on the target point cloud of the plurality of two-dimensional images, the set of feature points, and the target feature points includes:

[0116] Each feature point in the feature point set is averaged with the target feature point to obtain the sampling result of each feature point;

[0117] The sampling results of each feature point are aggregated with the target point cloud point by point to obtain the point cloud model of the target object.

[0118] The above sampling results are the average sampling results of each feature point and the target feature point, that is, the average value of each feature point and the target feature point. By averaging each feature point and the target feature point, each feature point is constrained. This effectively corrects some discrete feature points in the feature point set, making the corrected discrete feature points continuous with other feature points and reducing the weight of discrete feature points. Thus, it can be confirmed that the average sampling results are distributed on the surface of the target object, thereby improving the accuracy of the point cloud model.

[0119] The above-mentioned method aggregates the sampling results of each feature point with the target point cloud point by point. Specifically, it aggregates the sampling results of each feature point with at least three point clouds in the target point cloud point cloud point by point to obtain the point cloud model of the target object. Since the sampling results are all distributed on the surface of the target object, and the target point cloud is also distributed on the surface of the target object, the point cloud obtained by aggregating the sampling results of each feature point with the target point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud point cloud model point cloud model point cloud model point cloud model point cloud model point cloud model point cloud model point cloud model of target object ...

[0120] In this embodiment, by averaging each feature point in the feature point set with the target feature point, the sampling result of each feature point is obtained, which reduces the weight of discrete feature points in the feature point set and improves the effectiveness of the point cloud. Then, the sampling result of each feature point is aggregated with the target point cloud point by point to obtain the point cloud model of the target object. The number of point clouds corresponding to the target object is effectively increased compared with the prior art, thereby improving the accuracy of the point cloud model of the target object.

[0121] Please see Figure 6 , Figure 6 This is a structural diagram of a point cloud model building device provided in an embodiment of this application, such as... Figure 6 As shown, the point cloud model building device 600 includes:

[0122] The acquisition module 601 is used to acquire two-dimensional image data of the target object. The two-dimensional image data includes multiple two-dimensional images with different shooting angles, as well as camera parameters and poses corresponding to each two-dimensional image.

[0123] The first processing module 602 is used to extract the target point cloud of each two-dimensional image based on the camera parameters and pose corresponding to each two-dimensional image, wherein the target point cloud is distributed on the surface of the target object by at least three point clouds.

[0124] The second processing module 603 is used to perform multi-view feature extraction on each two-dimensional image based on the target point cloud of each two-dimensional image to obtain a feature point set of each two-dimensional image.

[0125] The third processing module 604 is used to obtain a set of feature points of a portion of the multiple two-dimensional images, and filter the set of feature points of the portion of the two-dimensional images to obtain the target feature points of each two-dimensional image in the portion of the two-dimensional images.

[0126] Construction module 605 is used to construct a point cloud model of the target object based on the target point cloud, the feature point set, and the target feature points of the multiple two-dimensional images. Optionally,

[0127] In one embodiment, the first processing module 602 includes:

[0128] The first processing submodule is used to perform monocular depth estimation on each two-dimensional image based on the camera parameters and pose corresponding to each two-dimensional image, and obtain the depth estimation result of each two-dimensional image;

[0129] The second processing submodule is used to back-project the depth estimation results of each two-dimensional image, and to filter and sample based on the back-projection results to obtain the target point cloud of each two-dimensional image.

[0130] In one embodiment, the first processing submodule includes:

[0131] The estimation unit is used to input the camera parameters and pose corresponding to each two-dimensional image, as well as each two-dimensional image, into an ordinal regression function for calculation to obtain multiple initial depths for each two-dimensional image, where each initial depth is the depth of a different region in the two-dimensional image;

[0132] The calculation unit is used to calculate the target loss for each region.

[0133] The deletion unit is used to delete redundant depths from the plurality of initial depths to obtain the depth estimation result, wherein the target loss of the region corresponding to the redundant depth is less than a set threshold.

[0134] In one embodiment, the computing unit includes:

[0135] A sub-unit is used to acquire the image features and true depth map of each region in each two-dimensional image;

[0136] A convolutional subunit is used to convolve the image features and initial depth of each region to obtain the positional residual.

[0137] The computational subunit is used to calculate the ordinal loss corresponding to each two-dimensional image.

[0138] The processing subunit is configured to set the sum of the initial depth and position residuals of each region, the norm of the difference between the sum and the true depth, and the sum of the ordinal loss corresponding to each two-dimensional image as the target loss for each region.

[0139] In one embodiment, the third processing module 604 includes the following:

[0140] The first acquisition submodule is used to acquire lightweight data and perform multi-view feature extraction on each two-dimensional image in the partial images based on the target point cloud of the partial images to obtain the feature point set of the partial two-dimensional images. The lightweight data includes partial images in the multiple two-dimensional images.

[0141] The first filtering submodule is used to filter the feature point set of each two-dimensional image to obtain the feature point set of the partial two-dimensional image.

[0142] In one embodiment, the third processing module 604 includes:

[0143] The second acquisition submodule is used to acquire the projection position of each feature point in the feature point set of the partial two-dimensional image in each two-dimensional image;

[0144] The extraction submodule is used to extract features from each projection position to obtain the features of each projection position;

[0145] An aggregation submodule is used to aggregate the features of each projection position to obtain the aggregated features;

[0146] The second filtering submodule is used to filter the feature point set of the partial two-dimensional image based on the aggregated features to obtain the target feature points.

[0147] In one embodiment, the aggregation submodule includes the following:

[0148] The first aggregation unit is used to perform average aggregation of the features at each projection location to obtain the aggregated features;

[0149] The second aggregation unit is used to perform variance aggregation on the features of each projection location to obtain the aggregated features.

[0150] In one embodiment, the extraction submodule includes:

[0151] The extraction unit is used to extract features from the target projection when the target projection position is inside the two-dimensional image, so as to obtain the features of the target projection position;

[0152] An interpolation unit is used to interpolate the features corresponding to the target image when the target projection position is outside the two-dimensional image to obtain the features of the target projection position. The pose corresponding to the target image is adjacent to the pose of the two-dimensional image corresponding to the target projection position. The target image is one of the plurality of two-dimensional images.

[0153] In one embodiment, the building module 605 includes:

[0154] A submodule is used to perform average sampling between each feature point in the feature point set and the target feature point to obtain the sampling result of each feature point;

[0155] A submodule is constructed to aggregate the sampling results of each feature point with the target point cloud point by point to obtain the point cloud model of the target object.

[0156] The point cloud model construction apparatus provided in this application embodiment can realize each process of the above-described point cloud model construction method, with one-to-one correspondence of technical features and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0157] It should be noted that the point cloud model building device in the embodiments of this application can be a device, or it can be a component, integrated circuit, or chip in an electronic device.

[0158] This application also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the various processes of the above-described point cloud model construction method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0159] For details, see Figure 7 This application also provides an electronic device, including a bus 701, a transceiver 702, an antenna 703, a bus interface 704, a processor 705, and a memory 706.

[0160] The transceiver 702 is used to acquire two-dimensional image data of the target object. The two-dimensional image data includes multiple two-dimensional images with different shooting angles, as well as camera parameters and poses corresponding to each two-dimensional image.

[0161] The processor 705 is used to extract the target point cloud of each two-dimensional image based on the camera parameters and pose corresponding to each two-dimensional image, wherein the target point cloud is distributed on the surface of the target object by at least three point clouds.

[0162] The processor 705 is further configured to perform multi-view feature extraction on each two-dimensional image based on the target point cloud of each two-dimensional image to obtain a feature point set of each two-dimensional image;

[0163] The processor 705 is further configured to acquire a set of feature points of a portion of the plurality of two-dimensional images, and filter the set of feature points of the portion of the two-dimensional images to obtain the target feature points of each two-dimensional image in the portion of the two-dimensional images.

[0164] The processor 705 is further configured to construct a point cloud model of the target object based on the target point cloud, the set of feature points, and the target feature points of the plurality of two-dimensional images.

[0165] In one embodiment, the processor 705 is further configured to perform monocular depth estimation on each two-dimensional image based on the camera parameters and pose corresponding to each two-dimensional image, and obtain the depth estimation result of each two-dimensional image;

[0166] The processor 705 is further configured to back-project the depth estimation results of each two-dimensional image, and perform filtering sampling based on the back-projection results to obtain the target point cloud of each two-dimensional image.

[0167] In one embodiment, the processor 705 is further configured to input the camera parameters and pose corresponding to each two-dimensional image, as well as each two-dimensional image, into an ordinal regression function for calculation to obtain multiple initial depths for each two-dimensional image, wherein each initial depth is the depth of a different region in the two-dimensional image;

[0168] The processor 705 is also used to calculate the target loss for each region;

[0169] The processor 705 is further configured to delete redundant depths from the plurality of initial depths to obtain the depth estimation result, wherein the target loss of the region corresponding to the redundant depth is less than a set threshold.

[0170] In one embodiment, the transceiver 702 is further configured to acquire image features and true depth maps of each region in each two-dimensional image;

[0171] The processor 705 is further configured to convolve the image features and initial depth of each region to obtain a positional residual;

[0172] The processor 705 is further configured to calculate the ordinal loss corresponding to each two-dimensional image for each two-dimensional image;

[0173] The processor 705 is further configured to set the sum of the initial depth and position residuals of each region, the norm of the difference between the actual depth and the sum of the ordinal loss corresponding to each two-dimensional image, as the target loss of each region.

[0174] In one embodiment, the processor 705 is further configured to acquire lightweight data and perform multi-view feature extraction on each two-dimensional image in the partial images based on the target point cloud of the partial images to obtain a feature point set of the partial two-dimensional images, wherein the lightweight data includes partial images in the plurality of two-dimensional images;

[0175] The processor 705 is further configured to filter the feature point set of each two-dimensional image to obtain the feature point set of the partial two-dimensional image.

[0176] In one embodiment, the transceiver 702 is further configured to acquire the projection position of each feature point in the feature point set of the partial two-dimensional image in each two-dimensional image;

[0177] The processor 705 is also used to extract features from each projection position to obtain features of each projection position;

[0178] The processor 705 is further configured to aggregate the features of each projection position to obtain the aggregated features;

[0179] The processor 705 is further configured to filter the feature point set of the partial two-dimensional image based on the aggregated features to obtain the target feature points.

[0180] In one embodiment, the processor 705 is further configured to perform average aggregation of the features at each projection location to obtain the aggregated features;

[0181] The processor 705 is further configured to perform variance aggregation on the features of each projection location to obtain the aggregated features.

[0182] In one embodiment, the processor 705 is further configured to extract features from the target projection when the target projection position is inside a two-dimensional image, thereby obtaining features of the target projection position.

[0183] The processor 705 is further configured to interpolate the features corresponding to the target image when the target projection position is outside the two-dimensional image, to obtain the features of the target projection position, wherein the pose corresponding to the target image is adjacent to the pose of the two-dimensional image corresponding to the target projection position, and the target image is one of the plurality of two-dimensional images.

[0184] In one embodiment, the processor 705 is further configured to perform average sampling on each feature point in the feature point set and the target feature point to obtain a sampling result for each feature point;

[0185] The processor 705 is further configured to aggregate the sampling results of each feature point with the target point cloud point by point to obtain the point cloud model of the target object.

[0186] exist Figure 7In this document, a bus architecture (represented by bus 701) is used. Bus 701 can include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 705 and memory represented by memory 706. Bus 701 can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 704 provides an interface between bus 701 and transceiver 702. Transceiver 702 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 705 is transmitted over a wireless medium via antenna 703, which further receives data and transmits data to processor 705.

[0187] Processor 705 manages bus 701 and general processing, and also provides various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. Memory 706 can be used to store data used by processor 705 during operation.

[0188] Optionally, the processor 705 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD).

[0189] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described point cloud model construction method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0190] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0191] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.

[0192] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for constructing a point cloud model, characterized in that, include: Acquire two-dimensional image data of the target object, wherein the two-dimensional image data includes multiple two-dimensional images taken from different perspectives, as well as the camera parameters and pose corresponding to each two-dimensional image; Based on the camera parameters and pose corresponding to each two-dimensional image, the target point cloud of each two-dimensional image is extracted, and the target point cloud is distributed on the surface of the target object by at least three point clouds; Based on the target point cloud of each two-dimensional image, multi-view feature extraction is performed on each two-dimensional image to obtain the feature point set of each two-dimensional image; Obtain feature point sets of a subset of the multiple two-dimensional images, and filter the feature point sets of the subset of two-dimensional images to obtain the target feature points of each two-dimensional image in the subset of two-dimensional images; Based on the target point cloud, the set of feature points, and the target feature points of the multiple two-dimensional images, a point cloud model of the target object is constructed. The filtering of the feature point set of the partial two-dimensional image to obtain the target feature points of each two-dimensional image in the partial two-dimensional image includes: Obtain the projection position of each feature point in the feature point set of the partial two-dimensional image in each two-dimensional image; Feature extraction is performed on each projection location to obtain the features of each projection location; The features of each projection location are aggregated to obtain aggregated features; Based on the aggregated features, the feature point set of the partial two-dimensional image is filtered to obtain the target feature points; The construction of a point cloud model of the target object based on the target point cloud, the feature point set, and the target feature points of the multiple two-dimensional images includes: Each feature point in the feature point set is averaged with the target feature point to obtain the sampling result of each feature point; The sampling results of each feature point are aggregated with the target point cloud point by point to obtain the point cloud model of the target object.

2. The method according to claim 1, characterized in that, The step of extracting the target point cloud of each two-dimensional image based on the camera parameters and pose corresponding to each two-dimensional image includes: Based on the camera parameters and pose corresponding to each two-dimensional image, monocular depth estimation is performed on each two-dimensional image to obtain the depth estimation result of each two-dimensional image; The depth estimation results of each two-dimensional image are back-projected, and the target point cloud of each two-dimensional image is obtained by filtering and sampling based on the back-projection results.

3. The method according to claim 2, characterized in that, The process of performing monocular depth estimation on each two-dimensional image based on the camera parameters and pose corresponding to each two-dimensional image to obtain the depth estimation result of each two-dimensional image includes: The camera parameters and pose corresponding to each two-dimensional image, as well as each two-dimensional image, are input into an ordinal regression function for calculation to obtain multiple initial depths for each two-dimensional image. Each initial depth is the depth of a different region in the two-dimensional image. Calculate the target loss for each region; Redundant depths are removed from the plurality of initial depths to obtain the depth estimation result, wherein the target loss of the region corresponding to the redundant depth is less than a set threshold.

4. The method according to claim 3, characterized in that, The calculation of the target loss for each region includes: Obtain the image features and true depth map of each region in each two-dimensional image; Convolve the image features and initial depth of each region to obtain the positional residual; The ordinal loss corresponding to each two-dimensional image is calculated. The sum of the initial depth and position residuals of each region, the norm of the difference between the actual depth and the sum of the ordinal loss corresponding to each two-dimensional image are set as the target loss for each region.

5. The method according to any one of claims 1 to 4, characterized in that, The step of obtaining a feature point set of a subset of the multiple two-dimensional images includes the following: Lightweight data is acquired, and multi-view feature extraction is performed on each of the two-dimensional images in the partial two-dimensional images based on the target point cloud of the partial two-dimensional images to obtain the feature point set of the partial two-dimensional images. The lightweight data includes partial two-dimensional images in the multiple two-dimensional images. The feature point set of each two-dimensional image is filtered to obtain the feature point set of the partial two-dimensional image.

6. The method according to claim 1, characterized in that, The aggregation of features at each projection location to obtain the aggregated features includes the following: The aggregated features are obtained by averaging and aggregating the features at each projection location. The variance of the features at each projection location is aggregated to obtain the aggregated features.

7. The method according to claim 1, characterized in that, The step of extracting features from each projection location to obtain the features of each projection location includes: When the target projection position is inside the two-dimensional image, feature extraction is performed on the target projection to obtain the features of the target projection position; When the target projection position is outside the two-dimensional image, the features corresponding to the target image are interpolated to obtain the features of the target projection position. The pose corresponding to the target image is adjacent to the pose of the two-dimensional image corresponding to the target projection position. The target image is one of the multiple two-dimensional images.

8. A point cloud model construction device, characterized in that, include: The acquisition module is used to acquire two-dimensional image data of the target object. The two-dimensional image data includes multiple two-dimensional images with different shooting angles, as well as the camera parameters and pose corresponding to each two-dimensional image. The first processing module is used to extract the target point cloud of each two-dimensional image based on the camera parameters and pose corresponding to each two-dimensional image, wherein the target point cloud is distributed on the surface of the target object by at least three point clouds. The second processing module is used to perform multi-view feature extraction on each two-dimensional image based on the target point cloud of each two-dimensional image to obtain a feature point set of each two-dimensional image. The third processing module is used to obtain a set of feature points of a portion of the multiple two-dimensional images, and filter the set of feature points of the portion of the two-dimensional images to obtain the target feature points of each two-dimensional image in the portion of the two-dimensional images. The construction module is used to construct a point cloud model of the target object based on the target point cloud, the feature point set, and the target feature points of the multiple two-dimensional images; The third processing module includes: The second acquisition submodule is used to acquire the projection position of each feature point in the feature point set of the partial two-dimensional image in each two-dimensional image; The extraction submodule is used to extract features from each projection position to obtain the features of each projection position; The aggregation submodule is used to aggregate the features of each projection position to obtain aggregated features; The second filtering submodule is used to filter the feature point set of the partial two-dimensional image based on the aggregated features to obtain the target feature points; The building module includes: A submodule is used to perform average sampling between each feature point in the feature point set and the target feature point to obtain the sampling result of each feature point; A submodule is constructed to aggregate the sampling results of each feature point with the target point cloud point by point to obtain the point cloud model of the target object.

Citation Information

Patent Citations

  • Heterogenous data fusion method and device, and storage medium

    CN112836734A

  • Synchronous positioning and mapping method integrating vision, IMU (Inertial Measurement Unit) and sonar

    CN113744337A