Three-dimensional modeling method and device
Through deep learning, the image sequence is processed to generate three-dimensional point cloud data, and the mask feature maps of point cloud segmentation and image segmentation are registered, solving the low automation rate and high cost problems caused by manual point selection in existing three-dimensional modeling, achieving more efficient and accurate modeling effects.
Patent Information
- Application Number
- CN202311870535.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-01
AI Technical Summary
The existing three-dimensional modeling technology relies on manual point selection, resulting in low automation rate, long time and high cost of modeling operations, and poor reconstruction effect in weak texture areas.
Using deep learning methods, the image sequence is processed through the MVS network to generate three-dimensional point cloud data, and the mask feature maps of point cloud segmentation and image segmentation are registered, and the geometric structural parameters of the target object are automatically extracted to reduce manual intervention.
Improves the automation rate of modeling, reduces modeling time and cost, and achieves more precise reconstruction in weak textured areas, reducing point cloud noise.
Smart Images

Figure CN120236003A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of modeling technologies, and particularly to a three-dimensional modeling method and apparatus. Background Art
[0002] The execution process of a scenario-based three-dimensional modeling service can be divided into four steps, namely data acquisition, data processing, 3D modeling, and result output. The three-dimensional modeling solution needs to select different acquisition devices, acquisition schemes, and data processing methods according to the requirements of different modeling scenarios.
[0003] Currently, for the panoramic three-mode modeling solution, a panoramic camera is used to perform omnidirectional data acquisition on the scene to be modeled, and then the acquired video or image data is processed. Then, according to the drawings, the external contour of the modeling object is manually selected to extract the modeling parameters and perform manual modeling. However, in the step of extracting the modeling parameters in the manual modeling solution, it is necessary to rely on the prior knowledge of the modeling operator, understand the information in the drawings manually, and use the method of manually selecting points to identify the external contour of the modeling object in the point cloud of the scene, and obtain the manual modeling contour according to the point selection operation. However, the accuracy of the manual point selection operation depends on the ability of the operator, and after the point selection is completed, it also takes a long time to manually adjust the contour range. The automation rate of the modeling operation is low, the modeling time is long, and the modeling cost is high. Summary of the Invention
[0004] Embodiments of this application provide a three-dimensional modeling method and apparatus to achieve an automated scene three-dimensional modeling task, which can reduce the modeling duration and the modeling cost.
[0005] In a first aspect, embodiments of this application provide a three-dimensional modeling method. This method can be executed by a modeling apparatus. The method includes: obtaining an image sequence, where the image sequence includes multiple images, and the image sequence is obtained by performing image acquisition on the scene to be modeled; processing the image sequence into three-dimensional point cloud data of the scene to be modeled; obtaining point cloud features of a target object in the scene to be modeled according to the three-dimensional point cloud data; obtaining a mask feature map of the target object according to the image sequence; registering the point cloud features of the target object and the mask feature map in a three-dimensional space to obtain geometric structure parameters of the target object; and completing the modeling process of the target object according to the geometric structure parameters of the target object.
[0006] In the embodiments of this application, instead of using the method of manually selecting points in the point cloud features, the geometric structure parameters of the target object are obtained more accurately through the registration of the point cloud features of the target object obtained by point cloud segmentation and the mask feature map of the target object obtained by image segmentation, without the participation of manual labor, improving the automation rate of the modeling operation, reducing the modeling duration, and reducing the labor cost consumed by the modeling.
[0007] In a possible implementation, processing the image sequence into the three-dimensional point cloud data of the to-be-modeled scene includes:
[0008] Processing the image sequence into the three-dimensional point cloud data of the to-be-modeled scene by using an MVS network of deep learning (such as a convolutional neural network).
[0009] In a possible implementation, processing the image sequence into the three-dimensional point cloud data of the to-be-modeled scene includes:
[0010] Obtaining the depth feature map of each image in the image sequence, where the depth feature map of the first image is obtained by performing feature extraction based on the first image and at least one image in the neighborhood of the first image; performing a differentiable homography transformation operation on the depth feature map of each image in the image sequence to transform the depth feature map of each image to the corresponding feature volume; performing three-dimensional feature fusion on the feature volume of the image sequence to obtain the three-dimensional point cloud data of the to-be-modeled scene.
[0011] In the above implementation, by calculating the depth of each image with reference to one or more images in the neighborhood and using a differentiable homography transformation operation, the reconstruction of the weak texture area can be ensured, the integrity of the effective area can be improved, and the point cloud holes can be reduced.
[0012] In a possible implementation, performing three-dimensional feature fusion on the feature volume of the image sequence includes: using a fusion loss function to perform three-dimensional feature volume fusion on the feature volume of the image sequence; where the fusion loss function is obtained by performing weighted fusion on multiple loss functions, and different loss functions in the multiple loss functions correspond to different parameters of the feature volume, and the parameters include at least one of depth, scale, gradient, and intensity.
[0013] In the above implementation, by adopting the method of fusing loss functions with differential metrics, the number of noise points in the point cloud data can be further reduced.
[0014] In a possible implementation, obtaining the point cloud features of the target object in the to-be-modeled scene based on the three-dimensional point cloud data includes: using convolutional kernels of multiple sizes to perform feature extraction on the three-dimensional point cloud data to obtain multi-scale features; performing fusion on the multi-scale features to obtain the point cloud features of the target object.
[0015] In the above implementation, convolutional kernels of multiple sizes are used to extract features of the 3D point cloud at different scales, and cross-scale fusion of the multi-scale features is performed in the residual network to solve the problem of identifying multi-scale objects. The input is the complete point cloud of the scene, and the output is the point cloud feature result labeled according to the recognition result.
[0016] In a possible implementation, obtaining the mask feature map of the target object according to the image sequence includes:
[0017] Stitch the image sequence to obtain a panoramic image;
[0018] Use a first segmentation network to perform segmentation processing on the panoramic image to obtain the mask feature map of the target object.
[0019] For example, the target object includes a cabinet.
[0020] In a possible implementation, the first segmentation network includes a mask attention network and a first self-attention network;
[0021] Using the first segmentation network to perform segmentation processing on the panoramic image to obtain the mask feature map of the cabinet includes:
[0022] Use the mask attention network to mark the cabinet in the panoramic image;
[0023] Use the first self-attention network to extract features from the marked cabinet to obtain the mask feature map of the cabinet.
[0024] In a possible implementation, registering the point cloud feature and the mask feature map of the target object in three-dimensional space to obtain the geometric structure parameters of the target object includes:
[0025] Map the mask feature map of the cabinet into three-dimensional space to obtain the three-dimensional feature of the cabinet;
[0026] Fuse the point cloud feature of the cabinet with the three-dimensional feature of the cabinet to obtain the initial bounding box of the cabinet;
[0027] Register the initial bounding box of the cabinet with the edge feature in the point cloud feature of the cabinet to obtain the optimized bounding box of the cabinet and the geometric structure parameters of the cabinet.
[0028] In the above implementation, after fusing the point cloud feature of the cabinet with the three-dimensional feature of the cabinet, the bounding box of the cabinet with the largest range is obtained, and then further registration operations are performed to optimize the bounding box of the cabinet, which can improve the accuracy of cabinet modeling.
[0029] In a possible implementation, the target object further includes a wall, and the method further includes: using a second segmentation network for the panoramic image to obtain the mask feature map of the wall; obtaining the structure parameters of the wall according to the mask feature map of the wall.
[0030] In a possible implementation, the second segmentation network includes a semantic segmentation network and a second self-attention network; using the second segmentation network for the panoramic image to obtain the mask feature map of the wall surface, including: using the semantic segmentation network to segment the wall surface from the panoramic image to obtain the segmentation result of the wall surface; using the second self-attention network to perform feature extraction on the segmentation result of the wall surface to obtain the mask feature map of the wall surface. In the above implementation, for the main structure, such as a wall, a multi-scale feature extraction method is adopted, and then the mask feature map is obtained by further combining the semantic segmentation network with the self-attention network. In the case where the main structure is more occluded, the occluded area can be more effectively restored, reducing the point cloud hole problem.
[0031] In a possible implementation, obtaining the structural parameters of the wall surface according to the mask feature map of the wall surface includes: based on the spatial plane parameters of the scene to be modeled and the mask feature map of the wall surface, determining at least one candidate spatial plane where the wall surface is located; determining the target spatial plane of the wall surface and the structural parameters of the wall surface from the at least one candidate spatial plane according to the loss function.
[0032] In the above implementation, after determining the mask feature map, multiple candidate spatial planes are selected through the spatial plane parameters, and then the unique target spatial plane is filtered out by combining the loss function, improving the accuracy of wall surface determination.
[0033] In a second aspect, a modeling device is provided. The modeling device can be used to implement the functions described in the method of the first aspect. In an optional implementation, the communication device includes a processing unit (sometimes also referred to as a processing module) and a transceiver unit (sometimes also referred to as a transceiver module). The transceiver unit can implement the sending function and the receiving function. When the transceiver unit implements the sending function, it can be referred to as a sending unit (sometimes also referred to as a sending module). When the transceiver unit implements the receiving function, it can be referred to as a receiving unit (sometimes also referred to as a receiving module). The sending unit and the receiving unit can be the same functional module, and this functional module is called the transceiver unit, which can implement the sending function and the receiving function; or, the sending unit and the receiving unit can be different functional modules, and the transceiver unit is a general term for these functional modules.
[0034] In a third aspect, an embodiment of the present application provides a modeling device, and the device includes a processor and a memory. The memory is used to store program code. The processor is used to read and execute the program code stored in the memory to implement the method described in the first aspect or any design of the first aspect.
[0035] Fourthly, an embodiment of the present application further provides a computer storage medium. A software program is stored in the storage medium, and when the software program is read and executed by one or more processors, the method provided by any design of the first aspect can be implemented.
[0036] Fifthly, an embodiment of the present application provides a computer program product containing instructions. When it runs on a computer, it enables the computer to execute the method provided by any design of the first aspect above.
[0037] Sixthly, an embodiment of the present application provides a chip, and the chip includes a processor. The processor is used to execute the method provided by any design of the first aspect.
[0038] In a possible design, the chip further includes a communication interface, and the communication interface is coupled to the processor.
[0039] In a possible design, the chip is connected to a memory and is used to read and execute the software program stored in the memory to implement the method provided by any design of the first aspect, or to implement the method provided by any design of the second aspect.
[0040] Based on the implementations provided in the above aspects of the present application, further combinations can be made to provide more implementations. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a schematic diagram of the three-dimensional modeling system architecture provided by an embodiment of the present application;
[0042] Figure 2 It is a schematic diagram of the three-dimensional modeling method flow provided by an embodiment of the present application;
[0043] Figure 3 It is a schematic diagram of the data acquisition and processing flow provided by an embodiment of the present application;
[0044] Figure 4 It is a schematic diagram of the dense point cloud generation process provided by an embodiment of the present application;
[0045] Figure 5 It is a schematic diagram of the modeling parameter extraction process provided by an embodiment of the present application;
[0046] Figure 6 It is a schematic diagram of the modeling result provided by an embodiment of the present application;
[0047] Figure 7 It is a schematic diagram of the effect of the three-dimensional modeling solution provided by an embodiment of the present application;
[0048] Figure 8 It is a schematic diagram of a modeling device provided by an embodiment of the present application;
[0049] Figure 9 Another schematic diagram of the modeling device provided by the embodiment of the present application. Detailed implementation manners
[0050] The embodiment of the present application can be applied to a 3D modeling scenario, such as in scenario-based 3D modeling, for example, the creation of a building information modeling (BIM).
[0051] To facilitate the understanding of the embodiment of the present application, the concepts of the technical terms involved are first introduced below:
[0052] (1) 3D modeling refers to the process of using computer methods and / or mathematical methods to establish a 3D mathematical model suitable for computer representation and processing based on the data of 3D objects collected in the geographical world.
[0053] (2) Point cloud data refers to a set of vectors in a three-dimensional coordinate system, which is used to represent the surface shape of an object.
[0054] (3) A scene represents a specific area in the geographical world. The scene in the present application is a 3D scene, for example, an indoor computer room, a computer room under a tower, etc.
[0055] (4) A neural network is used to implement machine learning, deep learning, search, reasoning, decision-making, etc. The neural networks mentioned in the present application can include various types, such as deep neural networks (DNNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), residual networks, attention networks, self-attention networks, neural networks using the transformer model, or other neural networks, etc. Some neural networks are introduced exemplarily below.
[0056] The work of each layer in a deep neural network can be described by the mathematical expression : From a physical level, the work of each layer in a deep neural network can be understood as completing the transformation from the input space (a set of input vectors) to the output space (i.e., from the row space to the column space of a matrix) through five operations on the input space. These five operations include: 1. Dimensionality increase / dimensionality reduction; 2. Enlargement / shrinkage; 3. Rotation; 4. Translation; 5. "Bending". Among them, the operations of 1, 2, and 3 are performed by Completion: The operation of 4 is completed by +b, and the operation of 5 is implemented by a(). The reason for using the word "space" here is that the objects to be classified are not individual things, but a class of things. Space refers to the set of all individuals of this class of things. Among them, W is the weight vector, and each value in this vector represents the weight value of a neuron in this layer of the neural network. This vector W determines the space transformation from the input space to the output space described above, that is, the weight W of each layer controls how to transform the space.
[0057] The purpose of training a neural network, that is, ultimately obtaining the weight matrices of all layers of the trained neural network (the weight matrix formed by many layers of vector W). Therefore, the training process of a neural network is essentially a process of learning the way to control space transformation, and more specifically, learning the weight matrix.
[0058] (5) The attention model is a neural network that applies the attention mechanism. In deep learning, the attention mechanism can be generally defined as a weight vector that describes importance: through this weight vector to predict or infer an element. For example, for a certain pixel in an image or a certain word in a sentence, an attention vector can be used to quantitatively estimate the correlation between the target element and other elements, and the weighted sum of the attention vector is used as an approximation of the target.
[0059] The attention mechanism in deep learning mimics the attention mechanism of the human brain. For example, when a human being observes a painting, although the human eye can see the whole picture, when the human observes it carefully, in fact, the eye only focuses on a part of the pattern in the whole picture. At this time, the human brain mainly focuses on this small piece of pattern. That is to say, when a human observes an image carefully, the attention of the human brain to the whole image is not balanced, and there is a certain weight distinction, which is the core idea of the attention mechanism.
[0060] Simply put, the human visual processing system often selectively focuses on some parts of the image and ignores other irrelevant information, which helps the human brain's perception. Similarly, in the attention mechanism of deep learning, in some problems involving language, speech or vision, some parts of the input may be more relevant than other parts. Therefore, through the attention mechanism in the attention model, the attention model can only dynamically focus on the part of the input that helps to effectively execute the task at hand.
[0061] (6) The self-attention network is a neural network that applies the self-attention mechanism. The self-attention mechanism is an extension of the attention mechanism. In fact, the self-attention mechanism is an attention mechanism that associates different positions of a single sequence to calculate the representation of the same sequence. The self-attention mechanism can play a key role in machine reading, abstract summarization, or image description generation. Taking the application of the self-attention network in natural language processing as an example, the self-attention network processes input data of any length and generates a new feature representation of the input data, and then converts the feature representation into the target word. The self-attention network layer in the self-attention network uses the attention mechanism to obtain the relationships between all other words, thereby generating a new feature representation for each word. The advantage of the self-attention network is that the attention mechanism can directly capture the relationships between all words in a sentence without considering the word positions.
[0062] (7) Transformer model:
[0063] The neural network using the Transformer model can include several encoders (which can also be called blocks). Each encoder can include a self-attention layer and a feed-forward layer. The self-attention layer can adopt the multi-head self-attention mechanism. The feed-forward layer can adopt a feedforward neural network (FNN). The neurons in the feedforward neural network are arranged in layers, and each neuron is only connected to the neurons in the previous layer. It receives the output of the previous layer and outputs to the next layer, and there is no feedback between layers. The encoder is used to convert the input corpus into a feature vector. The multi-head self-attention layer calculates the data input to the encoder using the calculations between three matrices. The three matrices include the query matrix Q, the key matrix K, and the value matrix V. When encoding the word at the current position in the sequence, the multi-head self-attention layer refers to various interdependent relationships between the word at the current position and the words at other positions in the sequence. The feed-forward layer is a linear transformation layer, which is used to perform a linear transformation on the representation of each word.
[0064] Among them, in the description of this application, unless otherwise specified, "a plurality of" means two or more than two. In addition, " / " means that the objects associated before and after are in an "or" relationship. For example, A / B can mean A or B. The "and / or" in this application is just a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Also, in order to clearly describe the technical solutions of the embodiments of this application, in the embodiments of this application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and terms such as "first" and "second" do not necessarily limit being different. It should also be noted that unless otherwise specified, the specific description of some technical features in one embodiment can also be applied to explain the corresponding technical features mentioned in other embodiments.
[0065] Next, the technical solutions in the embodiments of this application will be described with reference to the accompanying drawings in the embodiments of this application.
[0066] See Figure 1 As shown, it is a schematic diagram of a three-dimensional modeling system architecture provided by an embodiment of this application. The three-dimensional modeling system architecture may include a collection device 100 and a modeling device 200.
[0067] Among them, the collection device 100 is used to photograph the scene to be modeled to obtain an image sequence, that is, a video or multiple images. Exemplarily, the collection device 100 may be a camera, a mobile phone with a camera function, a panoramic camera, etc. The collection device 100 may collect according to the collection path planned in the scene to be modeled to obtain the collected video or image.
[0068] The modeling device 200 is used to obtain the image sequence from the collection device 100 and perform three-dimensional modeling on the scene to be modeled. The modeling device 200 can be either a hardware device, such as a server, a terminal computing device, etc., or a software device, such as a set of software programs running on the hardware.
[0069] The collection device 100 and the modeling device 200 can be connected through a network. The network can be a wide area network, a local area network, or a combination of the two.
[0070] In a possible implementation scenario, a storage unit may also be included in the three-dimensional modeling system architecture. The storage unit may be independent of the modeling device 200 or may be deployed inside the modeling device 200. For example, the storage unit may be a cloud storage server or an entity storage server. After the acquisition device 100 acquires an image sequence, it may send the image sequence to the storage unit for storage. The modeling device 200 may obtain the image sequence from the storage unit. The storage unit has an independent data matching and transmission interface for transmitting the image sequence to the modeling device 200. The modeling device may have a data receiving interface for receiving the image sequence.
[0071] In a possible implementation scenario, a visualization device may also be included in the three-dimensional modeling system architecture. The modeling device 200 may establish a communication connection with the visualization device, such as through a network, which may be a wide area network, a local area network, or a combination of the two. The modeling device 200 may send the modeling result of the scene to the visualization device, and the visualization device may perform model display according to the result. Exemplarily, the three-dimensional modeling system architecture may be a digital twin three-dimensional modeling system.
[0072] In some embodiments, taking the modeling scenario of a computer room as an example, the modeling device 200 sends the modeling parameters of the cabinet, etc. to the visualization device, so that the visualization device performs model display and annotation of the geometric structure parameters of the cabinet according to the modeling parameters of the cabinet. In some other embodiments, still taking the modeling scenario of a computer room as an example, the modeling parameters of the cabinet, etc. may be sent to the storage unit, so that the visualization device can retrieve the modeling parameters of the cabinet from the storage unit and perform model display and annotation of the geometric structure parameters of the cabinet according to the modeling parameters of the cabinet.
[0073] Exemplarily, when the modeling device 200 is a software device, in the three-dimensional modeling system, the modeling device 200 may run in a cloud computing device system (which may include at least one cloud computing device, such as a server, etc.), may also run on an entity server, may also run in an edge computing device system (which may include at least one edge computing device, such as a server, a desktop computer, etc.), and may also run on various terminal computing devices, such as a laptop computer, a personal desktop computer, a mobile phone, a tablet computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and so on.
[0074] Exemplarily, the modeling device 200 may also logically be a device composed of various parts, and the various components in the modeling device 200 may run on different systems or servers respectively. The various parts of the modeling device 200 may run on any two of the cloud computing device system, the edge computing device system, and the terminal computing device system. The cloud computing device system, the edge computing device system, and the terminal computing device system are connected by a communication path and can communicate with each other and transfer data. Exemplarily, the modeling device 200 runs on the edge computer device system and the cloud computing device system. The edge computing device system includes one or more edge computing devices, and the cloud computing device system includes one or more cloud computing devices.
[0075] Exemplarily, when the modeling device 200 runs on the edge computing device system, the visualization device may be a mobile visualization device; when the modeling device 200 runs on the cloud computing device system, the visualization device may be a central visualization device or a mobile visualization device; when the modeling device 200 runs on the cloud computing device system and the edge computing device system, the visualization device may be a central visualization device or a mobile visualization device.
[0076] Exemplarily, the visualization device may be a laptop computer, a personal desktop computer, a mobile phone, a tablet computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and so on.
[0077] In some embodiments, the modeling device 200 and the visualization device may also be deployed on the same device, such as a personal desktop computer, a laptop computer, and so on.
[0078] It should be noted that the interfaces involved in the embodiments of the present application may be offline interfaces or online interfaces, and the embodiments of the present application do not make specific limitations thereon.
[0079] As can be seen from the background art, in the step of extracting modeling parameters in the manual modeling solution, it is necessary to rely on the prior knowledge of the modeling operator, and according to the information manually understood in the drawing, use the method of manual point selection to identify the outer contour of the modeling object in the point cloud of the scene, and obtain the manual modeling contour according to the point selection operation. However, the accuracy of the manual point selection operation depends on the ability of the operator, and after the point selection is completed, it also takes a long time to manually adjust the contour range. The automation rate of the modeling operation is low, the modeling time is long, and the modeling cost is high.
[0080] Based on this, the embodiments of the present application provide a three-dimensional modeling method and device, which adopt an automated modeling method and do not require manual participation, and can reduce the modeling time and reduce the modeling cost.
[0081] See Figure 2 As shown, it is a schematic flowchart of a 3D modeling method provided by an embodiment of the present application. The 3D modeling method may include the following S201 to S204. The 3D modeling method may be executed by a modeling device, such as Figure 1 the modeling device 200 in
[0082] S201. Obtain an image sequence. The image sequence is obtained by collecting images of the scene to be modeled. In some embodiments, the scene to be modeled may be video-collected according to a set planned route to obtain video data, and the video data includes an image sequence. In other embodiments, the scene to be modeled may be video-collected and then the above image sequence is obtained by frame extraction. In still other embodiments, an omnidirectional camera may be used to capture the scene to be modeled from multiple perspectives to obtain an image sequence.
[0083] In some embodiments, before performing subsequent processing on the image sequence, preprocessing may be performed. For example, camera distortion parameter generation, fisheye image gradient de-duplication, IMU to rotation matrix conversion, fisheye image cutting / frame extraction, gyroscope data to IMU conversion, or fisheye image stitching, etc.
[0084] Optionally, the preprocessing may be performed by the acquisition device or by the modeling device, and the embodiments of the present application do not limit this.
[0085] The formats for obtaining images or videos may include but are not limited to:
[0086] 1). insv suffix video pair.
[0087] 2). Images after frame extraction of insv video. For example, images with a circular field of view and an aspect ratio of 1:1.
[0088] 3). Multiple images with the insp suffix. For example, two images with circular fields of view arranged left and right and an aspect ratio of 2:1.
[0089] 4). Multiple images with the jpg suffix and an aspect ratio of 2:1. Or,[[]]
[0090] 5). insv suffix video and multiple images with the jpg suffix.
[0091] S202. Process the image sequence into 3D point cloud data of the scene to be modeled. The 3D point cloud data may be dense point cloud data.
[0092] Exemplarily, a multi-view stereo (MVS) algorithm based on a neural network may be used to process the image sequence into 3D point cloud data of the scene to be modeled. The neural network may use a deep neural network or a convolutional neural network, etc.
[0093] Exemplarily, processing the image sequence into 3D point cloud data of the scene to be modeled may include the following:
[0094] A1. Obtain the depth feature map (which may be simply referred to as the feature map) of each image in the image sequence, where the depth feature map of the first image is obtained by performing feature extraction based on the first image and at least one image in the neighborhood of the first image.
[0095] For example, the image sequence may be preprocessed first. The purpose of the preprocessing is to obtain image matching pairs, that is, the pairing of a reference image and neighborhood images. One reference image corresponds to several neighborhood images, and the neighborhood images may be referred to as source images, that is, obtain a reference image and at least one corresponding source image. It can be understood that each image corresponds to an image matching pair, or each image corresponds to a set of images. Each set of images includes one reference image and N s source images. One reference image may be denoted as I ref , and N s source images are denoted as I source ={i∈Ω ref |i = 1, 2,..., N s}. Where Ω ref represents the set of images captured by cameras that have a spatially adjacent relationship with the camera corresponding to the reference image I ref . It can be determined according to the camera pose (R, t).
[0096] N s +1 feature maps of the images can be extracted by an encoder with shared weights, such as a convolutional neural network. For example, the ResNet / DenseNet encoding network can be used. W×H represents the size of the feature map. C represents the matching image pair.
[0097] A2. Perform a differentiable homography transformation operation on the depth feature maps of the image sequence to transform the depth feature maps of the image sequence into corresponding feature volumes.
[0098] After obtaining the feature maps, generate randomly sampled depth hypotheses {d k |k = 1, 2,.., N d} for each feature pixel as N d hypothetical depth planes. Use the differentiable homography transformation operation to warp and transform the reference feature map F ref to the N s adjacent source feature maps {F i , i = 1, 2,..., N s}. For a certain feature pixel p in F ref , on the depth hypothesis plane dk can be projected onto the source feature image F through the following formula i When the depth hypothesis value is d k In the case of, the distortion between the pixel p of the reference image and the pixel p of the corresponding i-th source image i is defined with reference to the formula (1) shown below
[0099]
[0100] where K, R and t represent camera parameters. Among them, K represents the camera intrinsic parameter matrix. t 0,i represents the camera displacement matrix of the i-th source image relative to the reference image, and R 0,i represents the camera rotation matrix of the i-th source image relative to the reference image
[0101] Through the above differentiable homography transformation, the source feature image F source is warped to the reference feature map F ref to construct a 3D cost volume C = {C i | i = 1, 2,.., N s}. According to the similarity of features, for the source feature image F i at the depth hypothesis d k the cost volume is generated using the L2 norm of the feature difference. For example, the feature difference satisfies the following formula (2)
[0102] c i (p, d k ) = ||F i [p i (p, d k )] - F ref (p)||2 Formula (2)
[0103] Finally, the N s 3D cost volumes need to be aggregated, and an adaptive weight w(c i (d k )) is used for cost volume aggregation of different views. For example, the aggregation result of the cost volume satisfies the following formula (3)
[0104]
[0105] C(d k ) represents the aggregation result of the cost volume at depth d k In this way, pixels with key context information will be assigned greater weights to improve problems caused by occlusion and different lighting conditions of non-Lambertian surfaces
[0106] Perform softmax on the cost volume to generate a probability volume U. U(d k) = softmax(C(d k ))。The depth map is obtained by performing weighted averaging on the depth hypotheses using the probability volume. The depth map satisfies Equation (4).
[0107]
[0108] D(p) represents the depth of the feature pixel p. U(p, d k ) represents the probability of the feature pixel p at depth d k .
[0109] The depth map can also be understood as a feature volume, or rather, the depth map is used to describe the feature volume.
[0110] After obtaining the depth maps of all the images through the deep learning-based MVS network in the embodiments of the present application, all the depth maps are filtered and fused to obtain the dense point cloud data of the scene.
[0111] A3. Perform three-dimensional feature fusion on the feature volume of the image sequence to obtain the three-dimensional point cloud data of the scene to be modeled.
[0112] In the above solution of the embodiments of the present application, a differentiable homography transformation operation is used to implicitly encode the feature map obtained in the feature extraction network (which can also be understood geometrically), construct a general learning feature for the weakly textured area in the scene, and restore the weakly textured area using the learned feature at the spatially missing position to solve the point cloud hole problem.
[0113] In some possible implementation manners, when performing three-dimensional feature fusion on the feature volume of the image sequence in A3, a loss function with differential metrics corresponding to different depths, scales, gradients, or intensities, etc. can be introduced according to the depth features of the images, the loss functions are fused according to the weight values obtained by training, and the depth map fusion result is given according to the fused loss function (which can be called the fusion loss function or can also be called otherwise). The depth map fusion result is the dense point cloud data. Specifically, the three-dimensional feature volume of the image sequence is fused using the fusion loss function to obtain the dense point cloud data (i.e., the three-dimensional point cloud data of the scene to be modeled). Among them, the fusion loss function is obtained by performing weighted fusion on multiple loss functions, and different loss functions in the multiple loss functions correspond to different parameters of the feature volume, and the parameters include at least one of depth, scale, gradient, and intensity. For example, depth, scale, gradient, and intensity respectively correspond to a loss function, and different loss functions correspond to a weight. By adjusting the weights corresponding to each parameter respectively, the difference between the fused result and the feature volume of the reference image is made smaller. The weights corresponding to each parameter respectively can be obtained by training using a neural network.
[0114] For example, the input to the neural network is a panoramic image and camera parameters. The panoramic image is encoded and decoded layer by layer by the neural network, and the camera parameters are used as relevant parameters in the fusion loss function. According to the resection and intersection synthesis error, the weights of the fusion loss function are backpropagated and fine-tuned, ultimately meeting the resection and intersection synthesis accuracy to obtain the weights of the loss function.
[0115] By adopting the method of fusing loss functions with differential metrics, the number of noise points in the point cloud data can be further reduced.
[0116] S203. Obtain the point cloud features of the target object in the to-be-modeled scene according to the three-dimensional point cloud data.
[0117] To obtain the point cloud features of the target object in the to-be-modeled scene according to the three-dimensional point cloud data, a feature extraction network can be used. In a possible implementation, the feature extraction network may include a convolutional neural network and a residual network. When executing S203, the convolutional neural network can perform feature extraction on the three-dimensional point cloud data to obtain multi-scale features, and then fuse the multi-scale features to obtain the point cloud features of the target object.
[0118] For example, the convolutional neural network includes convolutional kernels of multiple scales. The convolutional kernels of multiple scales are used to extract features of the three-dimensional point cloud data at different scales to obtain multi-scale features. When fusing the multi-scale features, a residual network can be used. The residual network performs cross-scale fusion on the multi-scale features to obtain the point cloud features of the target object at each scale, which can solve the problem of identifying multi-scale objects. For example, the convolutional kernels of multiple scales and the residual network form a feature extraction network, with the input being the three-dimensional point cloud data and the output being the point cloud features of each target object labeled according to the recognition results.
[0119] The feature extraction network can also adopt other network structures, such as convolutional kernels of multiple scales + encoder-decoder network. The structure of the feature extraction network is not specifically limited in the embodiments of the present application.
[0120] S204. Obtain the mask feature map of the target object according to the image sequence. It can be understood that the target object is segmented based on the image sequence to obtain the mask feature map of the target object. This step will be described in detail later and will not be elaborated here.
[0121] S205. Register the point cloud features of the target object and the mask feature map in three-dimensional space to obtain the geometric structure parameters of the target object.
[0122] S206. Complete the modeling process of the target object according to the geometric structure parameters of the target object.
[0123] In the embodiments of the present application, instead of manually selecting points in the point cloud features, the bounding box and geometric structure parameters of the target object are obtained through the registration of the point cloud features of the target object obtained by point cloud segmentation and the mask feature map of the target object obtained by image segmentation, without the participation of manual labor, improving the automation rate of the modeling operation, reducing the modeling duration, and reducing the labor cost consumed by the modeling.
[0124] In the embodiments of the present application, an MVS network based on deep learning is adopted, such as a backbone network (backbone) including the basic network structure of CasMVSNet, enhancement of the cost volume in weak texture regions, and the feature extraction results of the deep learning network based on MVS. A more stringent fusion determination criterion is adopted to ensure the accuracy of the reconstruction result and effectively suppress noise. General feature learning ensures that weak texture regions can be successfully reconstructed, and strict depth map fusion ensures the accuracy of the reconstruction result and effectively suppresses noise, reducing point cloud noise.
[0125] Exemplarily, when executing S205, the mask feature map can be first mapped into a three-dimensional space to obtain the three-dimensional features of the target object. This three-dimensional space is the same as the three-dimensional space corresponding to the point cloud features. It can be understood that the mask feature map is mapped into the three-dimensional space where the point cloud features are located. The registration of the mapping result can be performed according to the depth information of the point cloud features. Further, the point cloud features corresponding to the target object and the three-dimensional features corresponding to the target object are registered to obtain the geometric structure parameters of the target object.
[0126] In a possible implementation manner, when obtaining the mask feature map of the target object, the image sequence can be stitched into a panoramic image, and then the panoramic image is segmented by a segmentation network to obtain the mask feature map of the target object. In a possible example, the segmentation network can segment multiple target objects. In another possible example, different segmentation networks can be used for different types of target objects. The segmentation network can adopt a semantic segmentation network or other segmentation networks, and the embodiments of the present application do not limit this.
[0127] Exemplarily, after stitching an image sequence to obtain a panoramic image, the first segmentation network can be used to perform segmentation processing on the panoramic image, so as to obtain the mask feature map of the target object. As an example, the first segmentation network may include a masked attention network and a feature extraction network. The masked attention network is used to identify (or mark) the target object in the panoramic image, and then the feature extraction network is used to extract features of the marked target object to obtain the mask feature map of the target object (which can also be simply referred to as a mask). The feature extraction network can adopt a self-attention network, such as a transformer or a Pixel Decoder. In some embodiments, the feature extraction network can also adopt a Cross-Attention network. In some embodiments, the feature extraction network can also adopt two or three of the transformer, Pixel Decoder, or Cross-Attention networks. For example, if the target object is a cabinet, the feature extraction network can adopt a first self-attention network, and the first self-attention network can adopt a transformer or a Pixel Decoder. After marking the cabinet in the panoramic image using the masked attention mechanism, the mask feature map corresponding to the cabinet is generated according to the transformer mechanism.
[0128] After obtaining the mask feature map of the target object, combined with the point cloud feature of the target object obtained in S203, a registration operation is performed to improve the accuracy of target object segmentation, and the image segmentation results of each target object in the panoramic image are obtained. As an example, when performing registration, the mask feature map of the target object can be mapped into three-dimensional space to obtain the three-dimensional feature of the target object. Then, the point cloud feature of the target object and the three-dimensional feature of the target object are fused to obtain the initial bounding box (or edge box) of the target object with the largest range. The target object can be a cabinet, a pipeline, or other regular geometric bodies and irregular geometric bodies, etc. Further, the initial bounding box of the target object can be registered with the edge point feature of the point cloud feature of the target object to obtain the optimized bounding box of the target object and the geometric structure parameters of the target object. Exemplarily, when registering the initial bounding box of the target object with the edge point feature of the point cloud feature of the target object, an iterative closest point (ICP) algorithm, a random sample consensus (RANSAC) algorithm, a four-points congruent sets (4PCS) algorithm, or a registration algorithm based on geometric feature description can be used.
[0129] Taking a cabinet as an example, the geometric structure parameters of the cabinet may include the height, depth, and width of the cabinet. The geometric structure parameters of the cabinet may also include the name of the cabinet, the coordinates of the center point, the coordinates of the center point on the front of the cabinet, the vertex coordinates of the cabinet, and so on. After obtaining these geometric structure parameters, the modeling device can save them in a box.json file. Of course, other file formats can also be used, and the embodiments of the present application do not limit this. When modeling is required later, these geometric structure parameters can be read to obtain the modeling result.
[0130] In some application scenarios, the target object may also include a wall surface, which is different from the above-mentioned cabinets, pipelines, etc. and does not require a bounding box. The segmentation network used for segmenting the wall surface may be different from the segmentation network used for cabinets, etc. For the sake of easy distinction, the segmentation network used for the wall surface is called the second segmentation network. Then, after obtaining the panoramic image, the second segmentation network can be used to segment the panoramic image to obtain the mask feature map of the wall surface, and then the structural parameters of the wall surface can be obtained according to the mask feature map of the wall surface. For example, the second segmentation network may include a semantic segmentation network and a second self-attention network. The semantic segmentation network is used to segment the wall surface from the panoramic image to obtain the segmentation result of the wall surface; the second self-attention network is used to extract features from the segmentation result of the wall surface to obtain the mask feature map of the wall surface. When using the semantic segmentation network to perform regional estimation of the self-attention mechanism on the preliminary result of wall surface segmentation, the segmentation result of the wall surface in the occlusion area is obtained, and the mask feature map of the wall surface is obtained. Further, based on the spatial plane parameters of the to-be-modeled scene and the mask feature map of the wall surface, at least one candidate spatial plane where the wall surface is located can be determined; according to the loss function, the target spatial plane of the wall surface and the structural parameters of the wall surface are determined from the at least one candidate spatial plane. Exemplarily, the second self-attention network may adopt a Transformer network, etc.
[0131] The embodiments of the present application do not specifically limit the form of the loss function. For example, the loss function can adopt a weighted loss function, mean squared error (MSE), cross entropy (CE), or Smooth L1 loss, etc.
[0132] The solutions provided by the embodiments of the present application are illustrated below in combination with specific application scenarios.
[0133] Scenario 1: An indoor computer room, including cabinets. The modeling process may include:
[0134] 1. Data collection and processing.
[0135] See Figure 3As shown, the scene modeling information can be collected by a panoramic camera according to the acquisition route based on the scene plan to obtain the acquisition video or image required for data processing. For example, the video acquisition time is 2 - 4 minutes. For example, the image collected by the panoramic camera is a fisheye image. The image sequence can be obtained by extracting one frame from every consecutive multiple frames. For example, one fisheye image is extracted from every consecutive 10 frames, and the panoramic image is obtained after stitching. For example, the size of the panoramic image is 6080*3040.
[0136] 2. Generation of dense point cloud data. The method of S202 can be referred to.
[0137] For example, refer to Figure 4 As shown, S401, Feature extraction and mapping. Specifically, it can include: combining the differentiable homography transformation operation and multi-scale feature extraction to reconstruct and optimize the weak texture area. The problems of low point cloud integrity rate and point cloud holes can be solved. S402, Low-noise depth fusion. According to the depth features of the image, loss functions corresponding to different depths, scales, gradients or intensities, etc. are introduced, and the loss functions are fused according to the weight values obtained by training. The depth map fusion result is given according to the fused loss function (which can be called the fusion loss function or other names).
[0138] 3. Extraction of modeling parameters.
[0139] For example, it can be referred to Figure 5 As shown. S501, Cabinet detection based on point cloud segmentation. S203 can be referred to. Multi-size convolutional kernels are used to extract the features of the point cloud at different scales, and the multi-scale features are cross-scale fused in the residual network to solve the problem of identifying multi-scale objects. The input is the complete point cloud of the scene, and the output can be the point cloud features of each cabinet labeled according to the recognition result. S502, Cabinet segmentation processing based on the image sequence to obtain the two-dimensional features of the cabinet (i.e., the mask feature map). S204 can be referred to, and will not be elaborated here. S503, Feature fusion of 2D-3D segmentation. 2D refers to the mask feature map. 3D refers to the point cloud feature. S504, Registration and extraction of geometric structure parameters. For S503 and S504, the relevant descriptions of S205 can be referred to, and will not be elaborated here.
[0140] In this embodiment, an MVS method based on a convolutional neural network for feature extraction is adopted to replace the traditional MVS dense point cloud reconstruction algorithm, including the backbone of the CasMVSNet basic network structure, the optimization of the image stitching module for panoramic scenes, and the cost volume for weak texture regions. Based on the feature extraction results of the AI deep learning network based on MVS, a more stringent fusion determination criterion is adopted to ensure the accuracy of the reconstruction results and effectively suppress noise. General feature learning ensures that weak texture regions can be successfully reconstructed, and strict depth map fusion ensures the accuracy of the reconstruction results and effectively suppresses noise, reducing point cloud noise.
[0141] Exemplarily, the result of modeling the cabinet based on the obtained geometric structure parameters is shown in Figure 6 the figure shown. Through the solution provided by this application, manual operations can be reduced, the automation rate can be improved, and the modeling efficiency can be improved. For example, Table 1 exemplarily describes the technical effects achieved by this application. It should be understood that the technical effects described below are related to the experimental environment and experimental data and are only for example. Better effects can be achieved in some possible application environments and data.
[0142] Table 1
[0143]
[0144] Scenario 2: The computer room in the data center scenario. The modeling process may include:
[0145] 1. Data acquisition and processing. Refer to the description of Scenario 1, which will not be elaborated here.
[0146] 2. Generation of dense point cloud data. Refer to the description of Scenario 1, which will not be elaborated here.
[0147] 3. Extraction of modeling parameters. The modeling parameters in Scenario 2 include the geometric structure parameters of the cabinet and the wall structure parameters. The cabinets in the computer room of the data center are large in size and have many obstructions in the main structure (i.e., the wall). A multi-scale feature extraction method can be used to stack the encoder-decoder mechanism to extract the point cloud features of the target object.
[0148] Through the solution provided by this application, manual operations can be reduced, the automation rate can be improved, and the modeling efficiency can be improved. For example, Table 2 exemplarily describes the technical effects achieved by this application. It should be understood that the technical effects described below are related to the experimental environment and experimental data and are only for example. Better effects can be achieved in some possible application environments and data.
[0149] Table 2
[0150] Performance Item <![CDATA[Data center < 200m 2 > <![CDATA[400m 2 >Data center > 200m 2 > Accuracy Dimensions < 5 cm, angle < 5° Size < 10 cm, angle < 5° Number of Cabinets 171 432 Mapping Rate 94.7% 96.9% Number of Successes 162 419 Existing Time Consumption 16h 24h Latest Time Consumption 11h 16h Efficiency Improvement 31.2% 33.3%
[0151] Refer to Figure 7As shown in the figure, in the existing solution, after data acquisition and preprocessing, traditional MVS reconstruction technology is adopted, and then a manual point selection scheme is used to generate point clouds and obtain modeling parameters. When the traditional point cloud generation algorithm completes dense point cloud matching, due to the existence of many textureless or weakly textured areas in indoor panoramic images, point cloud holes appear in the corresponding positions in the three-dimensional space, affecting the quality of the generated point clouds, and further affecting operations such as identifying the outer contour and manually selecting the contour, ultimately directly affecting the modeling accuracy and efficiency. See Figure 7 As shown in the figure, the embodiments of the present application can achieve high-precision dense point cloud reconstruction based on multi-view panoramic images and automatic extraction of the main structure and cabinet parameters based on semantic and fusion segmentation features, without manual point selection, which can improve efficiency. The solution provided by the present application can also achieve automatic placement of cabinets based on semantic and fusion segmentation feature recognition. Based on this method, the manual participation rate in cabinet placement drops from 100% to →15%. It can achieve indoor cabinet modeling based on rule-based target modeling parameter extraction, and the modeling time for a single large-scale computer room is reduced by about 30%. Through the method of 2D-3D fusion and registration, it can ensure that weakly textured areas can be successfully reconstructed. Strict depth map fusion ensures the accuracy of the reconstruction result and effectively suppresses noise, reducing point cloud noise.
[0152] Figure 8 It is the structural diagram of the modeling device provided by the embodiments of the present application. This device can be implemented as part or all of the device through software. The device provided by the embodiments of the present application can implement the Figure 2 process described in the embodiments of the present application. The device includes: an acquisition module 810, a point cloud data generation module 820, and a modeling parameter extraction module 830, where:
[0153] The acquisition module 810 is used to acquire an image sequence, where the image sequence includes multiple images, and the image sequence is obtained by collecting images for the scene to be modeled.
[0154] The point cloud data generation module 820 is used to process the image sequence into three-dimensional point cloud data of the scene to be modeled.
[0155] The modeling parameter extraction module 830 is used to obtain the point cloud features of the target object in the scene to be modeled according to the three-dimensional point cloud data; obtain the mask feature map of the target object according to the image sequence; register the point cloud features and the mask feature map of the target object in the three-dimensional space to obtain the geometric structure parameters of the target object.
[0156] In some embodiments, it may further include a model generation module, which is used to complete the modeling process of the target object according to the geometric structure parameters of the target object.
[0157] In a possible implementation, when the point cloud data generation module 820 processes the image sequence into the three-dimensional point cloud data of the to-be-modeled scene, it is specifically configured to: obtain the depth feature map of each image in the image sequence, where the depth feature map of the first image is obtained by performing feature extraction on the first image and at least one image in the neighborhood of the first image; perform a differentiable homography transformation operation on the depth feature map of each image in the image sequence to transform the depth feature map of each image into a corresponding feature volume; perform three-dimensional feature fusion on the feature volumes of the image sequence to obtain the three-dimensional point cloud data of the to-be-modeled scene.
[0158] In a possible implementation, the modeling parameter extraction module 830 performs three-dimensional feature fusion on the feature volumes of the image sequence, including: using a fusion loss function to perform three-dimensional feature volume fusion on the feature volumes of the image sequence; where the fusion loss function is obtained by performing weighted fusion on multiple loss functions, and different loss functions in the multiple loss functions correspond to different parameters of the feature volumes, and the parameters include at least one of depth, scale, gradient, and intensity.
[0159] In a possible implementation, the modeling parameter extraction module 830 performs obtaining the point cloud features of the target object in the to-be-modeled scene according to the three-dimensional point cloud data, including: using convolutional kernels of multiple sizes to perform feature extraction on the three-dimensional point cloud data to obtain multi-scale features; performing fusion on the multi-scale features to obtain the point cloud features of the target object.
[0160] In a possible implementation, the target object includes a cabinet, and the modeling parameter extraction module 830 performs obtaining the mask feature map of the target object according to the image sequence, including: splicing the image sequence to obtain a panoramic image; using a first segmentation network to perform segmentation processing on the panoramic image to obtain the mask feature map of the cabinet.
[0161] In a possible implementation, the first segmentation network includes a mask attention network and a first self-attention network; the modeling parameter extraction module 830 performs using the first segmentation network to perform segmentation processing on the panoramic image to obtain the mask feature map of the cabinet, including: using the mask attention network to mark the cabinet in the panoramic image; using the first self-attention network to perform feature extraction on the marked cabinet to obtain the mask feature map of the cabinet.
[0162] In a possible implementation, the modeling parameter extraction module 830 performs registering the point cloud features and the mask feature map of the target object in a three-dimensional space to obtain the geometric structure parameters of the target object, including: mapping the mask feature map of the cabinet to the three-dimensional space to obtain the three-dimensional features of the cabinet; fusing the point cloud features of the cabinet with the three-dimensional features of the cabinet to obtain the initial bounding box of the cabinet; registering the initial bounding box of the cabinet with the edge features in the point cloud features of the cabinet to obtain the optimized bounding box of the cabinet and the geometric structure parameters of the cabinet.
[0163] In a possible implementation, the target object further includes a wall surface, and the modeling parameter extraction module 830 is further configured to use a second segmentation network for the panoramic image to obtain the mask feature map of the wall surface; obtain the structure parameters of the wall surface according to the mask feature map of the wall surface.
[0164] In a possible implementation, the second segmentation network includes a semantic segmentation network and a second self-attention network; the modeling parameter extraction module 830 performs using the second segmentation network for the panoramic image to obtain the mask feature map of the wall surface, including: using the semantic segmentation network to segment the wall surface from the panoramic image to obtain a segmentation result of the wall surface; using the second self-attention network to perform feature extraction on the segmentation result of the wall surface to obtain the mask feature map of the wall surface.
[0165] In a possible implementation, the modeling parameter extraction module 830 performs obtaining the structure parameters of the wall surface according to the mask feature map of the wall surface, including: determining at least one candidate space plane where the wall surface is located based on the spatial plane parameters of the to-be-modeled scene and the mask feature map of the wall surface; determining the target space plane of the wall surface and the structure parameters of the wall surface from the at least one candidate space plane according to a loss function.
[0166] The division of modules in the embodiments of the present application is illustrative. It is only a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present application, each functional module may be integrated in one processor, or may exist physically alone, or two or more modules may be integrated into one module. The above integrated modules may be implemented in the form of hardware or in the form of software functional modules.
[0167] When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a terminal device (which can be a personal computer, a network device, etc.) or a processor to execute all or part of the steps of the method in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0168] An embodiment of this application also provides a modeling device for three-dimensional reconstruction of a scene. Figure 9 An exemplary possible architecture diagram of the modeling device is provided.
[0169] The modeling device includes a memory 901, a processor 902, a communication interface 903, and a bus 904. Among them, the memory 901, the processor 902, and the communication interface 903 are communicatively connected to each other through the bus 904.
[0170] The memory 901 can be a ROM, a static storage device, a dynamic storage device, or a RAM. The memory 901 can store a program. When the program stored in the memory 901 is executed by the processor 902, the processor 902 and the communication interface 903 are used to execute the three-dimensional modeling method of the aforementioned Figure 2 shown scene, or implement the functions of the aforementioned Figure 8 shown device. The memory 901 can also store an image sequence.
[0171] The processor 902 can be a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU), or one or more integrated circuits.
[0172] The processor 902 can also be an integrated circuit chip with the ability to process signals. In the implementation process, some or all of the functions of the modeling device in this application can be completed by the integrated logic circuit in the hardware of the processor 902 or instructions in the form of software. The above-mentioned processor 902 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit, a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the above embodiments of this application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of this application can be directly embodied as being executed by the hardware decoding processor, or executed by a combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc.
[0173] The communication interface 903 uses a transceiver module such as, but not limited to, a transceiver to implement the communication between the modeling device and other devices or communication networks. For example, point cloud data and the like can be obtained through the communication interface 903.
[0174] The bus 904 can include a path for transmitting information between various components of the modeling device (for example, the memory 901, the processor 902, the communication interface 903).
[0175] The descriptions of the processes corresponding to the above respective drawings each have their own focuses. For parts not detailed in a certain process, reference can be made to the relevant descriptions of other processes.
[0176] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a server or a terminal, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial optical cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by the server or the terminal, or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, and a magnetic tape, etc.), an optical medium (such as a digital video disk (DVD), etc.), or a semiconductor medium (such as a solid-state drive, etc.).
[0177] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0178] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.
Claims
1. A three-dimensional modeling method, characterized in that, Including: Obtain an image sequence, where the image sequence includes a plurality of images, and the image sequence is obtained by collecting images for the scene to be modeled; Process the image sequence into three-dimensional point cloud data of the scene to be modeled; Obtain the point cloud features of the target object in the scene to be modeled according to the three-dimensional point cloud data; Obtain the mask feature map of the target object according to the image sequence; Register the point cloud features and the mask feature map of the target object in three-dimensional space to obtain the geometric structure parameters of the target object; Complete the modeling process of the target object according to the geometric structure parameters of the target object.
2. The method according to claim 1, wherein Processing the image sequence into three-dimensional point cloud data of the scene to be modeled includes: Obtain the depth feature map of each image in the image sequence, where the depth feature map of the first image is obtained by performing feature extraction on the first image and at least one image in the neighborhood of the first image; Perform a differentiable homography transformation operation on the depth feature map of each image in the image sequence to transform the depth feature map of each image into a corresponding feature volume; Perform three-dimensional feature fusion on the feature volumes of the image sequence to obtain the three-dimensional point cloud data of the scene to be modeled.
3. The method according to claim 2, wherein Performing three-dimensional feature fusion on the feature volumes of the image sequence includes Using a fusion loss function to perform three-dimensional feature volume fusion on the feature volumes of the image sequence; Wherein, the fusion loss function is obtained by performing weighted fusion on a plurality of loss functions, and the parameters of the feature volumes corresponding to different loss functions in the plurality of loss functions are different, and the parameters include at least one of depth, scale, gradient, and intensity.
4. The method according to any one of claims 1-3, characterized in that, Obtaining the point cloud features of the target object in the scene to be modeled according to the three-dimensional point cloud data includes: Use convolutional kernels of multiple sizes to perform feature extraction on the three-dimensional point cloud data to obtain multi-scale features; Fuse the multi-scale features to obtain the point cloud features of the target object.
5. The method according to any one of claims 1 to 4, characterized in that, The target object includes a cabinet, and obtaining the mask feature map of the target object according to the image sequence includes: Stitch the image sequence to obtain a panoramic image; Use a first segmentation network to perform segmentation processing on the panoramic image to obtain the mask feature map of the cabinet.
6. The method according to claim 5, characterized in that, The first segmentation network includes a mask attention network and a first self-attention network; Using the first segmentation network to perform segmentation processing on the panoramic image to obtain the mask feature map of the cabinet includes: Use the mask attention network to mark the cabinet in the panoramic image; Use the first self-attention network to perform feature extraction on the marked cabinet to obtain the mask feature map of the cabinet.
7. The method according to claim 5 or 6, characterized in that, Registering the point cloud features and the mask feature map of the target object in three-dimensional space to obtain the geometric structure parameters of the target object includes: Map the mask feature map of the cabinet into three-dimensional space to obtain the three-dimensional features of the cabinet; Fuse the point cloud features of the cabinet with the three-dimensional features of the cabinet to obtain the initial bounding box of the cabinet; Register the bounding box of the initial cabinet with the edge features in the point cloud features of the cabinet to obtain the optimized bounding box of the cabinet and the geometric structure parameters of the cabinet.
8. The method according to any one of claims 5 to 7, characterized in that The target object further includes a wall surface, and the method further includes: Use a second segmentation network for the panoramic image to obtain a mask feature map of the wall surface; Obtain the structural parameters of the wall surface according to the mask feature map of the wall surface.
9. The method according to claim 8, characterized in that, The second segmentation network includes a semantic segmentation network and a second self-attention network; Using the second segmentation network for the panoramic image to obtain a mask feature map of the wall surface includes: Use the semantic segmentation network to segment the wall surface from the panoramic image to obtain a segmentation result of the wall surface; Use the second self-attention network to extract features from the segmentation result of the wall surface to obtain a mask feature map of the wall surface.
10. The method according to claim 9, wherein Obtaining the structural parameters of the wall surface according to the mask feature map of the wall surface includes: Based on the spatial plane parameters of the scene to be modeled and the mask feature map of the wall surface, determine at least one candidate spatial plane where the wall surface is located; Determine the target spatial plane of the wall surface and the structural parameters of the wall surface from the at least one candidate spatial plane according to a loss function.
11. A modeling device, characterized in that, Includes: A memory for storing program instructions; A processor for being coupled with the memory, calling the program instructions in the memory, and executing the method according to any one of claims 1-10.
12. A chip, characterized in that, The chip is connected to the memory and is used to read and execute the program code stored in the memory to implement the method according to any one of claims 1-10.
13. A computer storage medium, characterized in that, A computer program is stored in the computer storage medium, and when the computer program is executed by the computer, the computer is caused to execute the method according to any one of claims 1-10.
14. A computer program product containing instructions, characterized in that, When the computer program product runs on a computer, the computer is caused to execute the method according to any one of claims 1 to 10.