Three-dimensional modeling method and apparatus
Through deep learning technology combining point cloud segmentation and image segmentation, the geometric structure parameters of the target object are automatically extracted, solving the low automation rate and high cost problems caused by manual point selection in existing three-dimensional modeling, and achieving more efficient and accurate modeling effects.
Patent Information
- Application Number
- PCT/CN2024/107335
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-07-24
- Publication Date
- 2025-07-03
AI Technical Summary
The existing three-dimensional modeling technology relies on manual point selection, resulting in low automation rate, long time, high cost of modeling operations, and poor reconstruction effect in weak texture areas.
Using deep learning methods, point cloud features and mask feature maps of target objects are automatically extracted for registration through a combination of point cloud segmentation and image segmentation, geometric structure parameters are obtained, and manual participation is reduced.
Improves the automation rate of modeling, reduces modeling time and cost, and achieves more accurate reconstruction in weak textured areas, reducing point cloud noise.
Smart Images

Figure CN2024107335_03072025_PF_FP_ABST
Abstract
Description
Three-dimensional modeling method and device
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of the People's Republic of China on December 28, 2023, with application number 202311870535.X and invention name "A three-dimensional modeling method and device", the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of modeling technology, and in particular to a three-dimensional modeling method and device. Background Art
[0004] The execution process of scenario-based 3D modeling can be divided into four steps: data acquisition, data processing, 3D modeling, and output. 3D modeling solutions require different acquisition equipment, acquisition solutions, and data processing methods based on the needs of different modeling scenarios.
[0005] The current panoramic three-mode modeling solution uses a panoramic camera to collect all-round data of the scene to be modeled, and then processes the collected video or image data. Then, according to the drawings, the outer contour of the modeling object is manually selected to extract the modeling parameters and perform manual modeling. However, in the step of extracting modeling parameters, the manual modeling solution requires the modeling operator's prior knowledge and the information manually understood in the drawings to manually select points to identify the outer contour of the modeling object in the point cloud of the scene, and obtain the manually modeled contour according to the point selection operation. However, the accuracy of the manual point selection operation depends on the operator's ability, and after the point selection is completed, it also takes a long time to manually adjust the contour range. The automation rate of the modeling operation is low, the modeling time is long, and the modeling cost is high.
[0006] Summary of the Invention
[0007] The embodiments of the present application provide a three-dimensional modeling method and device to realize automated scene three-dimensional modeling tasks, which can reduce modeling time and modeling costs.
[0008] In a first aspect, an embodiment of the present application provides a three-dimensional modeling method. The method can be executed by a modeling device. The method includes: acquiring an image sequence, the image sequence including multiple images, the image sequence being obtained by performing image acquisition on a scene to be modeled; processing the image sequence into three-dimensional point cloud data of the scene to be modeled; obtaining point cloud features of a target object in the scene to be modeled based on the three-dimensional point cloud data; obtaining a mask feature map of the target object based on the image sequence; aligning the point cloud features and the mask feature map of the target object in three-dimensional space to obtain geometric structure parameters of the target object; and completing the modeling process of the target object based on the geometric structure parameters of the target object.
[0009] In the embodiment of the present application, the method of manually selecting points in the point cloud features is no longer adopted. Instead, more accurate geometric structure parameters of the target object are obtained by aligning the point cloud features of the target object through point cloud segmentation and the mask feature map of the target object through image segmentation. No human intervention is required, the automation rate of the modeling operation is improved, the modeling time is reduced, and the labor cost consumed by the modeling is reduced.
[0010] In one possible implementation, processing the image sequence into three-dimensional point cloud data of the scene to be modeled includes:
[0011] An MVS network using deep learning (such as a convolutional neural network) is used to process the image sequence into three-dimensional point cloud data of the scene to be modeled.
[0012] In one possible implementation, processing the image sequence into three-dimensional point cloud data of the scene to be modeled includes:
[0013] Obtain a depth feature map for each image in an image sequence, wherein the depth feature map of a first image is obtained by performing feature extraction based on the first image and at least one image in the neighborhood of the first image; perform a differentiable homography transformation operation on the depth feature map of each image in the image sequence to transform the depth feature map of each image into a corresponding feature body; and perform three-dimensional feature fusion on the feature bodies of the image sequence to obtain three-dimensional point cloud data of the scene to be modeled.
[0014] In the above implementation, by referring to one or more images in the neighborhood when calculating the depth of each image, and using a differentiable homography transformation operation, the reconstruction of weak texture areas can be ensured, the integrity of the effective area can be improved, and point cloud holes can be reduced.
[0015] In one possible implementation, three-dimensional feature fusion is performed on the feature bodies of the image sequence, including: using a fusion loss function to perform three-dimensional feature fusion on the feature bodies of the image sequence; wherein the fusion loss function is obtained by weighted fusion of multiple loss functions, different loss functions among the multiple loss functions correspond to different parameters of the feature bodies, and the parameters include at least one of depth, scale, gradient and intensity.
[0016] In the above implementation, the number of noise points in the point cloud data can be further reduced by adopting the loss function fusion method of differential measurement.
[0017] In one possible implementation, point cloud features of a target object in the scene to be modeled are obtained based on the three-dimensional point cloud data, including: using multi-size convolution kernels to extract features from the three-dimensional point cloud data to obtain multi-scale features; and fusing the multi-scale features to obtain point cloud features of the target object.
[0018] In the above implementation, multi-scale convolution kernels are used to extract features of 3D point clouds at different scales, and multi-scale features are fused across scales in the residual network to solve the problem of multi-scale object recognition. The input is the complete point cloud of the scene, and the output is the point cloud feature result annotated according to the recognition result.
[0019] In a possible implementation, obtaining the mask feature map of the target object according to the image sequence includes:
[0020] Stitching the image sequence to obtain a panoramic image;
[0021] The panoramic image is segmented using a first segmentation network to obtain a mask feature map of the target object.
[0022] For example, the target object includes a cabinet.
[0023] In one possible implementation, the first segmentation network includes a mask attention network and a first self-attention network;
[0024] Using a first segmentation network to perform segmentation processing on the panoramic image to obtain a mask feature map of the cabinet, including:
[0025] Using a mask attention network to mark the cabinet in the panoramic image;
[0026] Use the first self-attention network to extract features of the marked cabinet to obtain a mask feature map of the cabinet.
[0027] In one possible implementation, registering the point cloud features and the mask feature map of the target object in three-dimensional space to obtain geometric structure parameters of the target object includes:
[0028] Mapping the mask feature map of the cabinet into a three-dimensional space to obtain a three-dimensional feature of the cabinet;
[0029] Fusing the point cloud features of the cabinet with the three-dimensional features of the cabinet to obtain an initial bounding box of the cabinet;
[0030] The initial bounding box of the cabinet is aligned with the edge features in the point cloud features of the cabinet to obtain the optimized bounding box of the cabinet and the geometric structure parameters of the cabinet.
[0031] In the above implementation, after fusing the point cloud features of the cabinet with the three-dimensional features of the cabinet, a bounding box of the cabinet with the largest range is obtained, and then further alignment operation is performed to optimize the bounding box of the cabinet, which can improve the accuracy of cabinet modeling.
[0032] In a possible implementation, the target object also includes a wall, and the method further includes: using a second segmentation network on the panoramic image to obtain a mask feature map of the wall; and obtaining structural parameters of the wall based on the mask feature map of the wall.
[0033] In one possible implementation, the second segmentation network includes a semantic segmentation network and a second self-attention network; the second segmentation network is used on the panoramic image to obtain a mask feature map of the wall, including: using the semantic segmentation network to segment the panoramic image to obtain a segmentation result of the wall; and using the second self-attention network to perform feature extraction on the segmentation result of the wall to obtain a mask feature map of the wall. In the above implementation, a multi-scale feature extraction method is used for the main structure, such as a wall, and then a mask feature map is further obtained by combining the semantic segmentation network with the self-attention network. This can more effectively restore the occluded area and reduce the problem of point cloud holes when the main structure is heavily occluded.
[0034] In one possible implementation, the structural parameters of the wall are obtained according to the mask feature map of the wall, including: determining at least one candidate spatial plane where the wall is located based on the spatial plane parameters of the scene to be modeled and the mask feature map of the wall; and determining the target spatial plane of the wall and the structural parameters of the wall from the at least one candidate spatial plane according to a loss function.
[0035] In the above implementation, after determining the mask feature map, multiple candidate spatial planes are selected through spatial plane parameters, and then the loss function is combined to filter and obtain a unique target spatial plane, thereby improving the accuracy of wall determination.
[0036] In the second aspect, a modeling device is provided. The modeling device can be used to implement the functions described in the method described in the first aspect. In an optional implementation, the communication device includes a processing unit (sometimes also referred to as a processing module) and a transceiver unit (sometimes also referred to as a transceiver module). The transceiver unit can implement a sending function and a receiving function. When the transceiver unit implements the sending function, it can be called a sending unit (sometimes also referred to as a sending module). When the transceiver unit implements the receiving function, it can be called a receiving unit (sometimes also referred to as a receiving module). The sending unit and the receiving unit can be the same functional module, which is called a transceiver unit, and the functional module can implement a sending function and a receiving function; or, the sending unit and the receiving unit can be different functional modules, and the transceiver unit is a general term for these functional modules.
[0037] In a third aspect, an embodiment of the present application provides a modeling device, comprising a processor and a memory. The memory is configured to store program code. The processor is configured to read and execute the program code stored in the memory to implement the method described in the first aspect or any design of the first aspect.
[0038] In a fourth aspect, embodiments of the present application further provide a computer storage medium storing a software program that, when read and executed by one or more processors, can implement any of the methods provided in the first aspect.
[0039] In a fifth aspect, an embodiment of the present application provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute any of the methods provided by the design of the first aspect above.
[0040] In a sixth aspect, an embodiment of the present application provides a chip, the chip including a processor. The processor is configured to execute any one of the methods provided in the first aspect.
[0041] In one possible design, the chip further includes a communication interface coupled to the processor.
[0042] In one possible design, the chip is connected to a memory and is used to read and execute a software program stored in the memory to implement the method provided by any one of the designs of the first aspect, or to implement the method provided by any one of the designs of the second aspect.
[0043] Based on the implementations provided in the above aspects, this application can also be further combined to provide more implementations. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] FIG1 is a schematic diagram of the architecture of a 3D modeling system provided in an embodiment of the present application;
[0045] FIG2 is a schematic diagram of a three-dimensional modeling method according to an embodiment of the present application;
[0046] FIG3 is a schematic diagram of the data collection and processing flow provided in an embodiment of the present application;
[0047] FIG4 is a schematic diagram of a dense point cloud generation process according to an embodiment of the present application;
[0048] FIG5 is a schematic diagram of a modeling parameter extraction process according to an embodiment of the present application;
[0049] FIG6 is a schematic diagram of the modeling results provided in an embodiment of the present application;
[0050] FIG7 is a schematic diagram of the effect of the three-dimensional modeling solution provided in an embodiment of the present application;
[0051] FIG8 is a schematic diagram of a modeling device provided in an embodiment of the present application;
[0052] FIG9 is a schematic diagram of another modeling device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] The embodiments of the present application can be applied to three-dimensional modeling scenarios, such as scenario-based three-dimensional modeling, such as the creation of building information modeling (BIM).
[0054] To facilitate understanding of the embodiments of this application, the following first introduces the concepts of the technical terms involved:
[0055] (1) Three-dimensional modeling refers to the process of using computer methods and / or mathematical methods to build a three-dimensional mathematical model suitable for computer representation and processing based on the data of three-dimensional objects in the geographical world.
[0056] (2) Point cloud data refers to a set of vectors in a three-dimensional coordinate system, used to represent the surface shape of an object.
[0057] (3) Scene: This refers to a specific area in the geographical world. The scene in this application is a three-dimensional scene, such as an indoor machine room, a machine room under a tower, etc.
[0058] (4) Neural networks, used to implement machine learning, deep learning, search, reasoning, decision-making, etc. The neural networks mentioned in this application may include various types, such as deep neural networks (DNN), convolutional neural networks (CNN), recurrent neural networks (RNN), residual networks, attention networks, self-attention networks, neural networks using transformer models, or other neural networks. The following is an exemplary introduction to some neural networks.
[0059] The work of each layer in a deep neural network can be expressed mathematically as To describe: From a physical perspective, the work of each layer in a deep neural network can be understood as completing the transformation from input space to output space (i.e., from the row space to the column space of a matrix) through five operations on the input space (a set of input vectors). These five operations include: 1. Dimensionality increase / decrease; 2. Zoom in / out; 3. Rotation; 4. Translation; 5. "Bending". Operations 1, 2, and 3 are represented by Completed, operation 4 is completed by +b, and operation 5 is performed by a(). The word "space" is used here because the object being classified is not a single thing, but a category of things, and space refers to the collection of all individuals in this category. W is the weight vector, and each value in this vector represents the weight of a neuron in that layer of the neural network. This vector W determines the spatial transformation from input space to output space described above. In other words, the weight W of each layer controls how the space is transformed.
[0060] The purpose of training a neural network is to ultimately obtain the weight matrices of all layers of the trained neural network (a weight matrix formed by many layers of vectors W). Therefore, the training process of a neural network is essentially learning how to control spatial transformations, and more specifically, learning the weight matrices.
[0061] (5) An attention model is a neural network that uses an attention mechanism. In deep learning, the attention mechanism can be broadly defined as a weight vector that describes the importance of an element: this weight vector is used to predict or infer an element. For example, for a pixel in an image or a word in a sentence, the attention vector can be used to quantitatively estimate the correlation between the target element and other elements, and the weighted sum of the attention vectors can be used as an approximation of the target.
[0062] The attention mechanism in deep learning simulates the attention mechanism of the human brain. For example, when a person looks at a painting, although their eyes can see the entire painting, when they look closely, their eyes actually focus on only a small part of the pattern. At this time, the human brain focuses primarily on this small part. In other words, when a person observes an image carefully, the human brain's attention is not evenly distributed across the entire image; instead, it is weighted differently. This is the core idea of the attention mechanism.
[0063] Simply put, the human visual processing system tends to selectively focus on certain parts of an image while ignoring other irrelevant information, thereby facilitating human perception. Similarly, in deep learning's attention mechanism, in some problems involving language, speech, or vision, certain parts of the input may be more relevant than others. Therefore, through the attention mechanism in the attention model, the attention model can dynamically focus on only the parts of the input that help effectively perform the task at hand.
[0064] (6) Self-attention network is a neural network that applies the self-attention mechanism. The self-attention mechanism is an extension of the attention mechanism. The self-attention mechanism is actually an attention mechanism that associates different positions of a single sequence to calculate the representation of the same sequence. The self-attention mechanism can play a key role in machine reading, abstract summary or image description generation. Taking the application of self-attention network in natural language processing as an example, the self-attention network processes input data of arbitrary length and generates a new feature expression of the input data, and then converts the feature expression into the target word. The self-attention network layer in the self-attention network uses the attention mechanism to obtain the relationship between all other words, thereby generating a new feature expression for each word. The advantage of the self-attention network is that the attention mechanism can directly capture the relationship between all words in a sentence without considering the position of the words.
[0065] (7) Transformer model:
[0066] A neural network using the transformer model can include several encoders (also called blocks). Each encoder can include an attention layer and a feedforward layer. The attention layer can use a multi-head self-attention mechanism. The feedforward layer can use a feedforward neural network (FNN). In a feedforward neural network, neurons are arranged in layers, with each neuron connected only to neurons in the previous layer. The layers receive the output of the previous layer and output it to the next layer, with no feedback between layers. The encoder converts the input corpus into a feature vector. The multi-head self-attention layer calculates the data input to the encoder using calculations between three matrices: the query matrix Q (query), the key matrix K (key), and the value matrix V (value). When encoding the word at the current position in the sequence, the multi-head self-attention layer considers the various interdependencies between the word at the current position and words at other positions in the sequence. The feedforward layer is a linear transformation layer that performs a linear transformation on the representation of each word.
[0067] In the description of this application, unless otherwise specified, "plurality" means two or more than two. Furthermore, " / " indicates that the objects associated with each other are in an "or" relationship. For example, A / B can mean A or B. "And / or" in this application merely describes an association relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. A and B can be singular or plural. Furthermore, to facilitate the clear description of the technical solutions of the embodiments of this application, the words "first" and "second" are used in the embodiments of this application to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity or order of execution, and do not necessarily define differences. It should also be noted that, unless otherwise specified, the specific description of certain technical features in one embodiment can also be used to explain the corresponding technical features mentioned in other embodiments.
[0068] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0069] 1 , which is a schematic diagram of a 3D modeling system architecture provided by an embodiment of the present application, can include an acquisition device 100 and a modeling apparatus 200 .
[0070] The acquisition device 100 is used to capture the scene to be modeled to obtain an image sequence, i.e., a video or multiple images. For example, the acquisition device 100 may be a video camera, a mobile phone with a camera function, a panoramic camera, etc. The acquisition device 100 may capture images along a planned acquisition path within the scene to be modeled to obtain a video or image sequence.
[0071] Modeling device 200 is used to acquire image sequences from acquisition device 100 and perform 3D modeling of the scene to be modeled. Modeling device 200 can be a hardware device, such as a server or terminal computing device, or a software device, such as a software program running on hardware.
[0072] The acquisition device 100 and the modeling apparatus 200 may be connected via a network, which may be a wide area network (WAN) or a local area network (LAN), or a combination of the two.
[0073] In one possible implementation scenario, the 3D modeling system architecture may further include a storage unit. The storage unit may be independent of the modeling device 200 or deployed within the modeling device 200. For example, the storage unit may be a cloud storage server or a physical storage server. After the acquisition device 100 acquires the image sequence, the image sequence may be sent to the storage unit for storage. The modeling device 200 may obtain the image sequence from the storage unit. The storage unit has an independent data matching and transmission interface for transmitting the image sequence to the modeling device 200. The modeling device may have a data receiving interface for receiving the image sequence.
[0074] In one possible implementation scenario, the 3D modeling system architecture may also include a visualization device. The modeling device 200 may establish a communication connection with the visualization device, for example, via a network, which may be a wide area network (WAN), a local area network (LAN), or a combination of both. The modeling device 200 may transmit the scene modeling results to the visualization device, which then displays the model based on the results. For example, the 3D modeling system architecture may be a digital twin 3D modeling system.
[0075] In some embodiments, using a computer room modeling scenario as an example, the modeling device 200 sends the modeling parameters of a cabinet to a visualization device, so that the visualization device displays the model and annotates the cabinet's geometric structure parameters based on the cabinet's modeling parameters. In other embodiments, also using a computer room modeling scenario as an example, the modeling parameters of the cabinet can be sent to a storage unit, so that the visualization device can retrieve the cabinet's modeling parameters from the storage unit and display the model and annotate the cabinet's geometric structure parameters based on the cabinet's modeling parameters.
[0076] Exemplarily, when the modeling device 200 is a software device, in the three-dimensional modeling system, the modeling device 200 can run on a cloud computing device system (which may include at least one cloud computing device, such as a server, etc.), or on a physical server, or on an edge computing device system (which may include at least one edge computing device, such as a server, a desktop computer, etc.), or on various terminal computing devices, such as laptops, personal desktop computers, mobile phones, tablet computers, ultra-mobile personal computers (UMPC), netbooks, personal digital assistants (PDA), etc.
[0077] Exemplarily, the modeling apparatus 200 can also be logically composed of various parts, and the various components of the modeling apparatus 200 can run in different systems or servers. The various parts of the modeling apparatus 200 can run in any two of the cloud computing device system, the edge computing device system, and the terminal computing device. The cloud computing device system, the edge computing device system, and the terminal computing device are connected by a communication path and can communicate and transmit data with each other. Exemplarily, the modeling apparatus 200 runs on an edge computer device system and a cloud computing device system, the edge computing device system includes one or more edge computing devices, and the cloud computing device system includes one or more cloud computing devices.
[0078] Exemplarily, when the modeling apparatus 200 runs on an edge computing device system, the visualization device may be a mobile visualization device; when the modeling apparatus 200 runs on a cloud computing device system, the visualization device may be a central visualization device or a mobile visualization device; when the modeling apparatus 200 runs on a cloud computing device system and an edge computing device system, the visualization device may be a central visualization device or a mobile visualization device.
[0079] For example, the visualization device may be a laptop computer, a personal desktop computer, a mobile phone, a tablet computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and the like.
[0080] In some embodiments, the modeling apparatus 200 and the visualization device may also be deployed on the same device, such as a personal desktop computer, a laptop computer, and the like.
[0081] It should be noted that the interface involved in the embodiments of the present application can be an offline interface or an online interface, and the embodiments of the present application do not make specific limitations on this.
[0082] As we can see from the background, manual modeling requires the operator to use their prior knowledge and manually interpret the information in the blueprint to identify the outer contour of the modeled object in the scene's point cloud. This process then generates a manually modeled contour based on the selected points. However, the accuracy of manual point selection depends on the operator's ability, and even after point selection, manual adjustment of the contour range is time-consuming. This results in a low automation rate, long modeling time, and high modeling costs.
[0083] Based on this, the embodiments of the present application provide a three-dimensional modeling method and device, which adopts an automated modeling method and does not require human participation, thereby reducing modeling time and modeling costs.
[0084] 2 is a flow chart of a three-dimensional modeling method according to an embodiment of the present application. The three-dimensional modeling method may include the following steps S201 to S204. The three-dimensional modeling method may be executed by a modeling device, such as the modeling device 200 in FIG1 .
[0085] S201: Acquire an image sequence. This image sequence is obtained by capturing images of the scene to be modeled. In some embodiments, the scene to be modeled may be captured through video along a predetermined planned route to obtain video data, which may include the image sequence. In other embodiments, the scene to be modeled may be captured through video and then frame-by-frame extracted to obtain the image sequence. In still other embodiments, a panoramic camera may be used to capture the scene to be modeled from multiple perspectives to obtain the image sequence.
[0086] In some embodiments, pre-processing may be performed before subsequent processing of the image sequence, such as generating camera distortion parameters, deduplicating fisheye image gradients, converting IMU images to rotation matrices, cropping / framing fisheye images, converting gyroscope data to IMU images, or stitching fisheye images.
[0087] Optionally, the preprocessing may be performed by an acquisition device or a modeling device, and this embodiment of the present application does not limit this.
[0088] The formats of the obtained images or videos may include but are not limited to:
[0089] 1). Video pair with insv suffix.
[0090] 2) Images extracted from insv videos, such as images with a circular field of view and an aspect ratio of 1:1.
[0091] 3) Multiple images with the insp suffix. For example, two circular field of view images arranged side by side with an aspect ratio of 2:1.
[0092] 4). Multiple images with jpg suffix, aspect ratio 2:1. Or,
[0093] 5). Videos with .insv suffix and multiple images with .jpg suffix.
[0094] S202: Process the image sequence into three-dimensional point cloud data of the scene to be modeled. The three-dimensional point cloud data may be dense point cloud data.
[0095] For example, a neural network-based multi-view stereo (MVS) algorithm can be used to process the image sequence into three-dimensional point cloud data of the scene to be modeled. The neural network can be a deep neural network or a convolutional neural network.
[0096] Exemplarily, processing the image sequence into three-dimensional point cloud data of the scene to be modeled may include the following:
[0097] A1. Obtain a depth feature map (also referred to as a feature map) of each image in an image sequence, wherein the depth feature map of a first image is obtained by performing feature extraction based on the first image and at least one image in the neighborhood of the first image.
[0098] For example, the image sequence can be preprocessed first. The purpose of preprocessing is to obtain image matching pairs, that is, pairs of reference images and neighborhood images. One reference image corresponds to several neighborhood images. The neighborhood images can be called source images. That is, a reference image and at least one corresponding source image are obtained. It can be understood that each image corresponds to an image matching pair, or each image corresponds to a group of images. Each group of images includes a reference image and N s A source image. A reference image can be represented as I ref , N s Zhang Yuan image is represented as I source ={i∈Ω ref |i=1,2,...,N s}. ref The representative space and the reference image I ref A set of images taken by cameras that have a spatially adjacent relationship with each other. It can be determined based on the camera pose (R, t).
[0099] N s +1 feature map of the image It can be extracted by an encoder with shared weights, such as a convolutional neural network, for example, a ResNet / DenseNet encoding network. W×H represents the size of the feature map. C represents the matching image pair.
[0100] A2: performing a differentiable homography transformation operation on the depth feature map of the image sequence to transform the depth feature map of the image sequence into a corresponding feature volume.
[0101] After obtaining the feature map, a randomly sampled depth hypothesis {d k |k=1,2,..,N d}, as N d Using the differentiable homography transformation operation, the reference feature map F ref Warp transform to its adjacent N s Zhang Yuan characteristic map {F i ,i=1,2,...,N s}, for F ref A feature pixel p in the depth hypothesis plane is d k When it is projected to the source feature image F, it can be projected to the source feature image F by the following formula i At depth d k In the case of pixel p of the reference image and the corresponding pixel p of the i-th source image i The definition of the distortion between is shown in formula (1).
[0102] Where K, R and t represent the camera parameters. K represents the camera intrinsic parameter matrix. t 0,i represents the camera displacement matrix of the i-th source image relative to the reference image, R 0,i Represents the camera rotation matrix of the i-th source image relative to the reference image.
[0103] Through the above differentiable homography transformation, the source feature image F source Warp to the reference feature map F ref To construct the 3D cost volume C = {C i |i=1,2,..,N s According to the similarity of features, for the source feature image F i At depth d k The cost volume under is generated using the bi-norm of the feature difference. For example, the feature difference satisfies the following formula (2).
[0104] c i (p,d k )=‖F i [p i (p,d k )]-F ref (p)‖2 Formula (2).
[0105] Finally, N s 3D cost volumes are aggregated and adaptive weights w(ci (d k )) Perform cost body aggregation of different views. For example, the cost body aggregation result satisfies the following formula (3).
[0106] C(d k ) indicates that at depth d k In this way, pixels with key context information will be assigned larger weights to improve the problems caused by occlusion and different lighting conditions of non-Lambertian surfaces.
[0107] Perform softmax on the cost body to generate the probability body U. U(d k )=softmax(C(d k )). The depth hypothesis is weighted averaged using the probability volume to obtain the depth map. The depth map satisfies formula (4).
[0108] D(p) represents the depth of feature pixel p. U(p, d k ) represents the depth d k The probability of feature pixel p under .
[0109] The depth map can also be understood as a feature body, or the depth map is used to describe the feature body.
[0110] After obtaining the depth maps of all images through the deep learning-based MVS network in the embodiment of the present application, all the depth maps are filtered and fused to obtain dense point cloud data of the scene.
[0111] A3: Perform three-dimensional feature fusion on the feature bodies of the image sequence to obtain three-dimensional point cloud data of the scene to be modeled.
[0112] In the above scheme of the embodiment of the present application, a differentiable homography transformation operation is used to implicitly encode the feature map (which can also be understood in terms of geometric structure) obtained in the feature extraction network, construct a universal learning feature for the weak texture area in the scene, and use the learned features to restore the weak texture area at the missing position in space, thereby solving the point cloud hole problem.
[0113] In some possible implementations, in A3, when performing three-dimensional feature fusion on the feature bodies of the image sequence, a loss function of a differential metric corresponding to different depths, scales, gradients or intensities can be introduced according to the depth features of the image, and the loss functions are fused according to the weight values obtained by training. The depth map fusion result is given according to the fused loss function (which can be called a fusion loss function or other names). The depth map fusion result is dense point cloud data. Specifically, the fusion loss function is used to perform three-dimensional feature fusion on the feature bodies of the image sequence to obtain dense point cloud data (i.e., three-dimensional point cloud data of the scene to be modeled). The fusion loss function is obtained by weighted fusion of multiple loss functions, and different loss functions in the multiple loss functions correspond to different parameters of the feature body, and the parameters include at least one of depth, scale, gradient and intensity. For example, depth, scale, gradient and intensity each correspond to a loss function, and different loss functions correspond to a weight. By adjusting the weights corresponding to each parameter, the difference between the fused result and the feature body of the reference image is made smaller. The weights corresponding to each parameter can be obtained by training using a neural network.
[0114] For example, the input of the neural network is the panoramic image and camera parameters. The panoramic image is encoded and decoded layer by layer according to the neural network. The camera parameters are used as relevant parameters in the fusion loss function. According to the aerial triangulation error, the weight of the fusion loss function is reversely transmitted and fine-tuned to finally meet the aerial triangulation accuracy and obtain the weight of the loss function.
[0115] By adopting the loss function fusion method of differential measurement, the number of noise points in the point cloud data can be further reduced.
[0116] S203: Obtain point cloud features of a target object in the scene to be modeled according to the three-dimensional point cloud data.
[0117] A feature extraction network may be used to obtain point cloud features of a target object in the scene to be modeled based on the three-dimensional point cloud data. In one possible embodiment, the feature extraction network may include a convolutional neural network and a residual network. When executing S203, a convolutional neural network may be used to extract features from the three-dimensional point cloud data to obtain multi-scale features, and then the multi-scale features are fused to obtain point cloud features of the target object.
[0118] For example, convolutional neural networks include multi-scale convolution kernels. These kernels are used to extract features of 3D point cloud data at different scales, generating multi-scale features. Residual networks can be used to fuse these multi-scale features. The residual network fuses these multi-scale features across scales, generating point cloud features of target objects at various scales. This can solve the problem of multi-scale object recognition. For example, a feature extraction network composed of multi-scale convolution kernels and a residual network can take 3D point cloud data as input and output point cloud features of each target object labeled according to the recognition results.
[0119] The feature extraction network may also adopt other network structures, such as a multi-scale convolution kernel + encoder-decoder network. The embodiment of the present application does not specifically limit the structure of the feature extraction network.
[0120] S204: Obtain a mask feature map of the target object based on the image sequence. This can be understood as performing segmentation processing on the target object based on the image sequence to obtain a mask feature map of the target object. This step will be described in detail later and will not be repeated here.
[0121] S205 , registering the point cloud features of the target object and the mask feature map in three-dimensional space to obtain geometric structure parameters of the target object.
[0122] S206: Complete the modeling process of the target object according to the geometric structure parameters of the target object.
[0123] In the embodiment of the present application, the method of manually selecting points in the point cloud features is no longer adopted. Instead, a more accurate bounding box and geometric structure parameters of the target object are obtained by aligning the point cloud features of the target object of point cloud segmentation and the mask feature map of the target object of image segmentation. No human intervention is required, the automation rate of the modeling operation is improved, the modeling time is reduced, and the labor cost consumed by the modeling is reduced.
[0124] The present invention utilizes a deep learning-based MVS network, such as the backbone network of the CasMVSNet basic network structure, cost volume enhancement for weak texture areas, and feature extraction results from the MVS deep learning network. A stricter fusion judgment criterion is used to ensure accurate reconstruction results and effectively suppress noise. Universal feature learning ensures successful reconstruction of weak texture areas, and strict depth map fusion ensures accurate reconstruction results and effectively suppresses noise, reducing point cloud noise.
[0125] Exemplarily, when executing S205, the mask feature map may be first mapped into a three-dimensional space to obtain the three-dimensional features of the target object. This three-dimensional space is identical to the three-dimensional space corresponding to the point cloud features. It can be understood that the mask feature map is mapped into the three-dimensional space where the point cloud features reside. The mapping results may be aligned based on the depth information of the point cloud features. Furthermore, the point cloud features corresponding to the target object and the three-dimensional features corresponding to the target object are aligned to obtain the geometric structure parameters of the target object.
[0126] In one possible implementation, when obtaining the mask feature map of the target object, the image sequence can be spliced into a panoramic image, and then the panoramic image can be segmented using a segmentation network to obtain the mask feature map of the target object. In one possible example, the segmentation network can be used to segment a variety of target objects. In another possible example, different segmentation networks can be used for different types of target objects. The segmentation network can use a semantic segmentation network or other segmentation networks, which is not limited in the embodiments of the present application.
[0127] Exemplarily, after stitching the image sequence to obtain a panoramic image, the panoramic image can be segmented using a first segmentation network to obtain a mask feature map of the target object. As an example, the first segmentation network can include a masked attention network and a feature extraction network. The target object in the panoramic image is identified (or marked) by the masked attention network, and then the feature extraction network is used to extract features from the marked target object to obtain a mask feature map of the target object (also referred to as a mask). The feature extraction network can use a self-attention network, such as a transformer or a pixel decoder. In some embodiments, the feature extraction network can also use a cross-attention network. In some embodiments, the feature extraction network can also use two or three of the transformer, pixel decoder, or cross-attention network. For example, if the target object is a cabinet, the feature extraction network can use a first self-attention network, and the first self-attention network can use a transformer or a pixel decoder. After the cabinet is marked in the panoramic image using the masked attention mechanism, the mask feature map corresponding to the cabinet is generated according to the transformer mechanism.
[0128] After obtaining the mask feature map of the target object, the registration operation is performed in combination with the point cloud features of the target object obtained in S203 to improve the accuracy of target object segmentation and obtain the image segmentation results of each target object in the panoramic image. As an example, when performing the registration, the mask feature map of the target object can be mapped to a three-dimensional space to obtain the three-dimensional features of the target object. The point cloud features of the target object and the three-dimensional features of the target object are then fused to obtain the initial bounding box (or edge box) of the target object with the largest range. The target object can be a cabinet, a pipe or other regular geometric bodies and irregular geometric bodies, etc. Furthermore, the initial bounding box of the target object can be aligned with the edge point features of the point cloud features of the target object to obtain the optimized bounding box of the target object and the geometric structure parameters of the target object. For example, when aligning the initial bounding box of the target object with the edge point features of the point cloud features of the target object, an iterative closest point (ICP) algorithm, a random sample consensus (RANSAC) algorithm, a four-point matching algorithm (4PCS), or a registration algorithm based on geometric feature description can be used.
[0129] Taking a cabinet as an example, the cabinet's geometric structural parameters may include the cabinet's height, depth, and width. The cabinet's geometric structural parameters may also include the cabinet's name, center point coordinates, cabinet front center point coordinates, cabinet vertex coordinates, and so on. After obtaining these geometric structural parameters, the modeling device may save them in a box.json file. Of course, other file formats may also be used, and this embodiment of the application does not limit this. When modeling is required later, these geometric structural parameters may be read to obtain the modeling results.
[0130] In some application scenarios, the target object may also include a wall. Unlike the above-mentioned cabinets, pipes, etc., the wall does not require a bounding box. The segmentation network used for wall segmentation may be different from the segmentation network used for cabinets, etc. For ease of distinction, the segmentation network used for the wall is referred to as the second segmentation network. After obtaining the panoramic image, the second segmentation network can be used to segment the panoramic image to obtain a mask feature map of the wall, and then the structural parameters of the wall are obtained based on the mask feature map of the wall. For example, the second segmentation network may include a semantic segmentation network and a second self-attention network. The semantic segmentation network is used to segment the panoramic image to obtain the segmentation result of the wall; the second self-attention network is used to perform feature extraction on the segmentation result of the wall to obtain the mask feature map of the wall. The semantic segmentation network is used to perform regional estimation of the self-attention mechanism on the preliminary result of the wall segmentation to obtain the wall segmentation result of the occluded area and the wall mask feature map. Furthermore, based on the spatial plane parameters of the scene to be modeled and the mask feature map of the wall, at least one candidate spatial plane where the wall is located can be determined; and a target spatial plane of the wall and structural parameters of the wall can be determined from the at least one candidate spatial plane according to a loss function. For example, the second self-attention network can employ a Transformer network, etc.
[0131] The embodiment of the present application does not specifically limit the form of the loss function. For example, the loss function may be a weighted loss function, mean squared error (MSE), cross entropy (CE), or Smooth L1 loss.
[0132] The following examples illustrate the solutions provided in the embodiments of the present application in combination with specific application scenarios.
[0133] Scenario 1: Indoor equipment room, including cabinets. The modeling process may include:
[0134] 1. Data collection and processing.
[0135] As shown in Figure 3, scene modeling information can be collected using a panoramic camera along a planned acquisition route to produce the video or images required for data processing. For example, video acquisition time can be 2-4 minutes. For example, the images captured by the panoramic camera are fisheye images. An image sequence can be generated by extracting one frame from every consecutive frame. For example, a fisheye image frame is extracted every 10 consecutive frames and stitched together to produce a panoramic image. For example, the size of the panoramic image can be 6080*3040.
[0136] 2. Dense point cloud data generation: See S202 for details.
[0137] For example, see Figure 4, S401, feature extraction and mapping. Specifically, it may include: combining differentiable homography transformation operations and multi-scale feature extraction to reconstruct and optimize weak texture areas. It can solve the problems of low point cloud completeness and point cloud holes. S402, low-noise deep fusion. According to the depth features of the image, a loss function of differential metrics corresponding to different depths, scales, gradients or intensities is introduced, and the loss function is fused according to the weight values obtained by training. The depth map fusion result is given according to the fused loss function (which can be called a fusion loss function or other names).
[0138] 3. Modeling parameter extraction.
[0139] For example, please refer to Figure 5. S501, cabinet detection based on point cloud segmentation. Please refer to S203. Multi-size convolution kernels are used to extract features of point clouds at different scales, and multi-scale features are fused across scales in the residual network to solve the problem of multi-scale object recognition. The input is the complete point cloud of the scene, and the output can be the point cloud features of each cabinet annotated according to the recognition results. S502, cabinet segmentation processing is performed based on the image sequence to obtain the two-dimensional features of the cabinet (i.e., mask feature map). Please refer to S204, which will not be repeated here. S503, feature fusion of 2D-3D segmentation. 2D refers to the mask feature map. 3D refers to the point cloud feature. S504, registration and extraction of geometric structure parameters. For S503 and S504, please refer to the relevant description of S205, which will not be repeated here.
[0140] In this implementation, the MVS method, based on convolutional neural network feature extraction, replaces the traditional MVS dense point cloud reconstruction algorithm. This includes the backbone of the CasMVSNet basic network structure, an optimized image stitching module for panoramic scenes, and a cost volume for weakly textured areas. Based on the feature extraction results of the MVS AI deep learning network, a stricter fusion judgment criteria is used to ensure accurate reconstruction results and effectively suppress noise. Universal feature learning ensures successful reconstruction of weakly textured areas, and strict depth map fusion ensures accurate reconstruction results and effectively suppresses noise, reducing point cloud noise.
[0141] For example, the results of cabinet modeling based on the obtained geometric structural parameters are shown in Figure 6. The solution provided by this application can reduce manual operations, increase automation, and improve modeling efficiency. For example, Table 1 exemplifies the technical effects achieved by this application. It should be understood that the technical effects described below are related to the experimental environment and experimental data and are only used as examples. Even better results can be achieved under certain possible application environments and data.
[0142] Table 1
[0143] Scenario 2: Data center computer room. The modeling process may include:
[0144] 1. Data collection and processing: See the description of scenario 1 and will not be repeated here.
[0145] 2. Dense point cloud data generation. See the description of scenario 1 and will not be repeated here.
[0146] 3. Modeling parameter extraction. The modeling parameters for scenario 2 include the cabinet geometry and wall parameters. Data center computer room cabinets are large, and the main structure (i.e., walls) provides significant occlusion. Therefore, a multi-scale feature extraction approach can be used to superimpose an encoder-decoder mechanism to extract point cloud features of the target object.
[0147] The solution provided by this application can reduce manual operations, increase automation rate, and improve modeling efficiency. For example, Table 2 exemplifies the technical effects achieved by this application. It should be understood that the technical effects described below are related to the experimental environment and experimental data and are only used as examples. Even better effects can be achieved under certain possible application environments and data.
[0148] Table 2
[0149] As shown in Figure 7, the existing solution uses traditional MVS reconstruction technology after data collection and preprocessing, and then generates a point cloud and obtains modeling parameters through manual point selection. Furthermore, when completing dense point cloud matching, the traditional point cloud generation algorithm causes point cloud holes at corresponding locations in three-dimensional space due to the presence of a large number of textureless or weakly textured areas in indoor panoramic images. This affects the quality of the generated point cloud, which in turn affects operations such as identifying outer contours and manually selecting contours, ultimately directly affecting modeling accuracy and efficiency. As shown in Figure 7, the embodiment of the present application can achieve high-precision dense point cloud reconstruction based on multi-view panoramic images and automatic extraction of main structure and cabinet parameters based on semantic and fused segmentation features, without the need for manual point selection, thereby improving efficiency. The solution provided by this application can also achieve automatic cabinet placement based on semantic and fused segmentation feature recognition. Based on the automatic effect achieved by this method, the manual participation rate of cabinet placement is reduced from 100% to →15%. It can also achieve indoor cabinet modeling based on rule-based target modeling parameter extraction, reducing the modeling time of a single large-scale computer room by approximately 30%. Through 2D-3D fusion and registration, it can be ensured that weak texture areas can be successfully reconstructed. Strict depth map fusion ensures accurate reconstruction results and effectively suppresses noise, reducing point cloud noise.
[0150] FIG8 is a block diagram of a modeling device provided in an embodiment of the present application. The device can be implemented as part or all of the device through software. The device provided in an embodiment of the present application can implement the process described in FIG2 of the embodiment of the present application. The device includes: an acquisition module 810, a point cloud data generation module 820, and a modeling parameter extraction module 830, wherein:
[0151] The acquisition module 810 is used to acquire an image sequence, where the image sequence includes a plurality of images and is obtained by collecting images of a scene to be modeled.
[0152] The point cloud data generating module 820 is configured to process the image sequence into three-dimensional point cloud data of the scene to be modeled.
[0153] The modeling parameter extraction module 830 is used to obtain the point cloud features of the target object in the scene to be modeled based on the three-dimensional point cloud data; obtain the mask feature map of the target object based on the image sequence; and align the point cloud features and the mask feature map of the target object in three-dimensional space to obtain the geometric structure parameters of the target object.
[0154] In some embodiments, a model generation module may be further included, which is used to complete the modeling process of the target object according to the geometric structure parameters of the target object.
[0155] In one possible embodiment, the point cloud data generation module 820, when processing the image sequence into three-dimensional point cloud data of the scene to be modeled, is specifically used to: obtain a depth feature map of each image in the image sequence, wherein the depth feature map of the first image is obtained by feature extraction based on the first image and at least one image in the neighborhood of the first image; perform a differentiable homography transformation operation on the depth feature map of each image in the image sequence to transform the depth feature map of each image into a corresponding feature body; and perform three-dimensional feature fusion on the feature bodies of the image sequence to obtain three-dimensional point cloud data of the scene to be modeled.
[0156] In one possible embodiment, the modeling parameter extraction module 830 performs three-dimensional feature fusion on the feature bodies of the image sequence, including: using a fusion loss function to perform three-dimensional feature body fusion on the feature bodies of the image sequence; wherein the fusion loss function is obtained by weighted fusion of multiple loss functions, and different loss functions among the multiple loss functions correspond to different parameters of the feature bodies, and the parameters include at least one of depth, scale, gradient and intensity.
[0157] In one possible implementation, the modeling parameter extraction module 830 obtains point cloud features of the target object in the scene to be modeled based on the three-dimensional point cloud data, including: using multi-size convolution kernels to extract features from the three-dimensional point cloud data to obtain multi-scale features; and fusing the multi-scale features to obtain point cloud features of the target object.
[0158] In one possible embodiment, the target object includes a cabinet, and the modeling parameter extraction module 830 executes to obtain a mask feature map of the target object based on the image sequence, including: splicing the image sequence to obtain a panoramic image; and using a first segmentation network to segment the panoramic image to obtain a mask feature map of the cabinet.
[0159] In one possible embodiment, the first segmentation network includes a mask attention network and a first self-attention network; the modeling parameter extraction module 830 performs segmentation processing on the panoramic image using the first segmentation network to obtain a mask feature map of the cabinet, including: using the mask attention network to mark the cabinet in the panoramic image; using the first self-attention network to extract features of the marked cabinet to obtain a mask feature map of the cabinet.
[0160] In one possible embodiment, the modeling parameter extraction module 830 performs registration of the point cloud features and the mask feature map of the target object in three-dimensional space to obtain the geometric structure parameters of the target object, including: mapping the mask feature map of the cabinet into three-dimensional space to obtain the three-dimensional features of the cabinet; fusing the point cloud features of the cabinet with the three-dimensional features of the cabinet to obtain an initial bounding box of the cabinet; and aligning the initial bounding box of the cabinet with the edge features in the point cloud features of the cabinet to obtain an optimized bounding box of the cabinet and the geometric structure parameters of the cabinet.
[0161] In a possible embodiment, the target object also includes a wall, and the modeling parameter extraction module 830 is further used to use a second segmentation network on the panoramic image to obtain a mask feature map of the wall; and obtain the structural parameters of the wall based on the mask feature map of the wall.
[0162] In one possible embodiment, the second segmentation network includes a semantic segmentation network and a second self-attention network; the modeling parameter extraction module 830 executes the second segmentation network on the panoramic image to obtain the mask feature map of the wall, including: using the semantic segmentation network to segment the panoramic image to obtain the segmentation result of the wall; using the second self-attention network to perform feature extraction on the segmentation result of the wall to obtain the mask feature map of the wall.
[0163] In a possible embodiment, the modeling parameter extraction module 830 executes to obtain the structural parameters of the wall according to the mask feature map of the wall, including: determining at least one candidate spatial plane where the wall is located based on the spatial plane parameters of the scene to be modeled and the mask feature map of the wall; determining the target spatial plane of the wall and the structural parameters of the wall from the at least one candidate spatial plane according to the loss function.
[0164] The division of modules in the embodiments of the present application is illustrative and is merely a logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in the various embodiments of the present application may be integrated into a single processor, or may exist physically separately, or two or more modules may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or software functional modules.
[0165] If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions for enabling a terminal device (which can be a personal computer, or a network device, etc.) or a processor to execute all or part of the steps of the method of each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0166] The present application also provides a modeling device for 3D reconstruction of a scene. FIG9 exemplarily provides a possible architecture diagram of the modeling device.
[0167] The modeling device includes a memory 901 , a processor 902 , a communication interface 903 and a bus 904 . The memory 901 , the processor 902 and the communication interface 903 are connected to each other via the bus 904 .
[0168] Memory 901 may be a ROM, static storage device, dynamic storage device, or RAM. Memory 901 may store programs. When the program stored in memory 901 is executed by processor 902, processor 902 and communication interface 903 are used to execute the 3D modeling method for the scene shown in FIG. 2 , or to implement the functions of the apparatus shown in FIG. 8 . Memory 901 may also store image sequences.
[0169] The processor 902 may be a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU), or one or more integrated circuits.
[0170] The processor 902 may also be an integrated circuit chip with signal processing capabilities. During implementation, some or all of the functions of the modeling device of the present application may be performed by hardware integrated logic circuits or software instructions in the processor 902. The processor 902 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit, a field programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The methods, steps, and logic block diagrams disclosed in the above embodiments of the present application may be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present application may be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or the like.
[0171] The communication interface 903 uses a transceiver module such as, but not limited to, a transceiver to implement communication between the modeling device and other devices or a communication network. For example, point cloud data can be obtained through the communication interface 903.
[0172] The bus 904 may include a path for transmitting information between various components of the modeling device (eg, the memory 901 , the processor 902 , and the communication interface 903 ).
[0173] The descriptions of the processes corresponding to the above figures have different focuses. For parts that are not described in detail in a certain process, please refer to the relevant descriptions of other processes.
[0174] In the above embodiments, they can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a server or terminal, the process or function described in the embodiment of the present application is generated in whole or in part. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a server or terminal or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, and a tape, etc.), an optical medium (e.g., a digital video disk (DVD), etc.), or a semiconductor medium (e.g., a solid-state hard disk, etc.).
[0175] The present application is described with reference to the flow chart and / or block diagram of the method, device (system), and computer program product according to the embodiment of the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one flow chart flow or multiple flows and / or one box or multiple boxes of the block diagram.
[0176] Obviously, those skilled in the art may make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is intended to include these modifications and variations.
Claims
1. A three-dimensional modeling method, characterized in that, Including: Obtain an image sequence, where the image sequence includes multiple images, and the image sequence is obtained by collecting images for the scene to be modeled; Process the image sequence into three-dimensional point cloud data of the scene to be modeled; Obtain the point cloud features of the target object in the scene to be modeled according to the three-dimensional point cloud data; Obtain the mask feature map of the target object according to the image sequence; Register the point cloud features and the mask feature map of the target object in three-dimensional space to obtain the geometric structure parameters of the target object; Complete the modeling process of the target object according to the geometric structure parameters of the target object.
2. The method according to claim 1, wherein Processing the image sequence into three-dimensional point cloud data of the scene to be modeled includes: Obtain the depth feature map of each image in the image sequence, where the depth feature map of the first image is obtained by performing feature extraction according to the first image and at least one image in the neighborhood of the first image; Perform a differentiable homography transformation operation on the depth feature map of each image in the image sequence to transform the depth feature map of each image into a corresponding feature volume; Perform three-dimensional feature fusion on the feature volumes of the image sequence to obtain the three-dimensional point cloud data of the scene to be modeled.
3. The method according to claim 2, wherein Performing three-dimensional feature fusion on the feature volumes of the image sequence includes: Use a fusion loss function to perform three-dimensional feature volume fusion on the feature volumes of the image sequence; Wherein, the fusion loss function is obtained by weighted fusion of multiple loss functions, and the parameters of the feature volumes corresponding to different loss functions in the multiple loss functions are different, and the parameters include at least one of depth, scale, gradient, and intensity.
4. The method according to any one of claims 1 to 3, characterized in that, Obtaining the point cloud features of the target object in the scene to be modeled according to the three-dimensional point cloud data includes: Use convolutional kernels of multiple sizes to perform feature extraction on the three-dimensional point cloud data to obtain multi-scale features; Fuse the multi-scale features to obtain the point cloud features of the target object.
5. The method according to any one of claims 1 to 4, characterized in that, The target object includes a cabinet, and obtaining the mask feature map of the target object according to the image sequence includes: Stitch the image sequence to obtain a panoramic image; Use a first segmentation network to perform segmentation processing on the panoramic image to obtain the mask feature map of the cabinet.
6. The method according to claim 5, characterized in that, The first segmentation network includes a mask attention network and a first self-attention network; Using the first segmentation network to perform segmentation processing on the panoramic image to obtain the mask feature map of the cabinet includes: Use the mask attention network to mark the cabinet in the panoramic image; Use the first self-attention network to perform feature extraction on the marked cabinet to obtain the mask feature map of the cabinet.
7. The method according to claim 5 or 6, characterized in that, Registering the point cloud features and the mask feature map of the target object in three-dimensional space to obtain the geometric structure parameters of the target object includes: Map the mask feature map of the cabinet into three-dimensional space to obtain the three-dimensional features of the cabinet; Fuse the point cloud features of the cabinet with the three-dimensional features of the cabinet to obtain the initial bounding box of the cabinet; Register the bounding box of the initial cabinet with the edge features in the point cloud features of the cabinet to obtain the optimized bounding box of the cabinet and the geometric structure parameters of the cabinet.
8. The method according to any one of claims 5 to 7, characterized in that, The target object further includes a wall, and the method further includes: Use a second segmentation network on the panoramic image to obtain a mask feature map of the wall; Obtain the structure parameters of the wall according to the mask feature map of the wall.
9. The method according to claim 8, characterized in that, The second segmentation network includes a semantic segmentation network and a second self-attention network; Using the second segmentation network on the panoramic image to obtain a mask feature map of the wall includes: Use the semantic segmentation network to segment the wall from the panoramic image to obtain a segmentation result of the wall; Use the second self-attention network to extract features from the segmentation result of the wall to obtain a mask feature map of the wall.
10. The method according to claim 9, characterized in that, Obtaining the structure parameters of the wall according to the mask feature map of the wall includes: Based on the spatial plane parameters of the scene to be modeled and the mask feature map of the wall, determine at least one candidate spatial plane where the wall is located; Determine the target spatial plane of the wall and the structure parameters of the wall from the at least one candidate spatial plane according to a loss function.
11. A three-dimensional modeling system, characterized in that, Including a collection device and a modeling device; wherein, The collection device is used to capture the scene to be modeled and collect an image sequence; The modeling device is used to obtain the image sequence from the collection device, the image sequence includes multiple images, and the image sequence is obtained by image capture for the scene to be modeled; process the image sequence into three-dimensional point cloud data of the scene to be modeled; obtain the point cloud features of the target object in the scene to be modeled according to the three-dimensional point cloud data; obtain the mask feature map of the target object according to the image sequence; register the point cloud features and the mask feature map of the target object in three-dimensional space to obtain the geometric structure parameters of the target object; complete the modeling process of the target object according to the geometric structure parameters of the target object.
12. The system according to claim 11, wherein When the modeling device is used to process the image sequence into three-dimensional point cloud data of the scene to be modeled, it specifically includes: Obtain the depth feature map of each image in the image sequence, where the depth feature map of the first image is obtained by feature extraction according to the first image and at least one image in the neighborhood of the first image; perform a differentiable homography transformation operation on the depth feature map of each image in the image sequence to transform the depth feature map of each image to a corresponding feature volume; perform three-dimensional feature fusion on the feature volume of the image sequence to obtain the three-dimensional point cloud data of the scene to be modeled.
13. The system according to claim 12, wherein When the modeling device is used to perform three-dimensional feature fusion on the feature volume of the image sequence, it specifically includes: Use a fusion loss function to perform three-dimensional feature volume fusion on the feature volume of the image sequence; wherein, the fusion loss function is obtained by weighted fusion of multiple loss functions, and different loss functions in the multiple loss functions correspond to different parameters of the feature volume, and the parameters include at least one of depth, scale, gradient, and intensity.
14. The system according to any one of claims 11-13, characterized in that, When the modeling device is used to obtain the point cloud features of the target object in the to-be-modeled scene according to the three-dimensional point cloud data, it specifically includes: Performing feature extraction on the three-dimensional point cloud data using convolutional kernels of multiple sizes to obtain multi-scale features; fusing the multi-scale features to obtain the point cloud features of the target object.
15. The system according to any one of claims 11-14, characterized in that, The target object includes a cabinet. When the modeling device is used to obtain the mask feature map of the target object according to the image sequence, it specifically includes: Stitching the image sequence to obtain a panoramic image; using a first segmentation network to perform segmentation processing on the panoramic image to obtain the mask feature map of the cabinet.
16. The system according to any one of claims 11-15, characterized in that, The system further includes: a visualization device; wherein, The visualization device is configured to receive the modeling result of the modeling process from the modeling device; perform model display according to the modeling result. Display.
17. A modeling device, characterized in that, Includes: A memory for storing program instructions; A processor for being coupled to the memory, calling the program instructions in the memory, and executing the method according to any one of claims 1-10.
18. A chip, characterized in that, The chip is connected to the memory and is configured to read and execute the program code stored in the memory to implement the method according to any one of claims 1-10.
19. A computer storage medium, characterized in that, A computer program is stored in the computer storage medium. When the computer program is executed by the computer, the computer is caused to execute the method according to any one of claims 1-10.
20. A computer program product comprising instructions, characterized in that, When the computer program product runs on the computer, the computer is caused to execute the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Heterogenous data fusion method and device, and storage medium
CN112836734A
Model construction method, pose estimation method and object picking device
CN113034575A
Three-dimensional model generation method and device based on image sequence, equipment and medium
CN114758093A
Three-dimensional reconstruction method and system based on improved MVSNet
CN116912405A
Robotic control based on 3D bounding shape, for an object, generated using edge-depth values for the object
US20200376675A1
Cited By
Multi-level feature fusion-based cross-domain building extraction method and related device
CN120599482A
A cross-domain building extraction method and related device based on multi-level feature fusion
CN120599482B
Rural light storage and charging station intelligent management method and system combined with point cloud data
CN120764859A
Pipeline three-dimensional reconstruction and intelligent detection method based on panoramic stereoscopic vision
CN120833446A
Robust three-dimensional surface reconstruction method based on multi-scale fusion
CN120976446A