Method for generating three-dimensional space structure and electronic equipment

Through single-image 3D reconstruction and panoramic segmentation technology, combined with semantic plane detection and constraint filtering, the problem of accurately restoring the 3D structure of the home space in a single image is solved, and efficient and accurate 3D space construction and interactive operations are achieved.

CN120655815APending Publication Date: 2025-09-16TAOBAO CHINA SOFTWARE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510517994.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

When reconstructing the 3D scene of a user's home space, existing technologies have difficulty accurately restoring the 3D structure of planes such as the ground and walls in a single image. In particular, when the bottom line of the wall is obscured, recognition errors or extensions may occur, resulting in the inability to accurately restore the 3D structure of each plane in the space.

Method used

Using single-image 3D reconstruction technology and panoramic segmentation technology, a 3D point cloud is extracted from a single 2D image. The 3D parameters of multiple semantic plane areas are obtained through semantic plane detection to generate a mesh space, including the 3D parameters of the ground, ceiling and wall. Combined with normal vector, height and position constraint filtering, the 3D spatial structure of the target space is generated.

Benefits of technology

It achieves efficient and accurate construction of 3D spatial structure in a single image, can accurately extract semantic plane areas of ceilings, floors and walls, supports users' interactive operations of adding objects in 3D space, and improves the practicality and operability of 3D space reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655815A_ABST
    Figure CN120655815A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a method for generating a three-dimensional space structure and electronic equipment. The method comprises the steps that a single two-dimensional picture about a target space place is acquired; predicting a three-dimensional point cloud corresponding to the two-dimensional picture by using a single-image three-dimensional reconstruction technology, and generating a three-dimensional background point cloud corresponding to a plurality of semantic plane regions with spatial background semantics by using a panoramic segmentation technology; semantic plane detection is carried out based on the three-dimensional background point clouds to obtain three-dimensional parameters of the multiple semantic plane areas, and the three-dimensional parameters comprise a plane equation and in-plane point clouds; and generating a mesh space according to the three-dimensional parameters of the plurality of semantic plane regions so as to generate a three-dimensional space structure about the target space place. According to the embodiment of the invention, the construction of the three-dimensional space structure can be completed more efficiently and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of three-dimensional interaction technology, and in particular to a method and electronic device for generating a three-dimensional spatial structure. Background Art

[0002] In traditional home improvement and home furnishing product purchasing scenarios, users often want to place furniture products in their own homes to check out the effects of coordination and size. In the past, this demand could only be met by transporting the furniture home offline, which was extremely costly for users due to the high experience and return costs. However, 3D (three-dimensional) e-commerce scenarios allow users to virtually place furniture products in their own homes to check out the effects of coordination and size. This is an application scenario based on technologies such as augmented reality (AR) and 3D modeling. 3D e-commerce scenarios for the home improvement and home furnishing industry meet users' needs for a highly personalized and immersive shopping experience at a very low cost and in a convenient manner. It also provides e-commerce platforms and brands with opportunities to improve conversion rates and user satisfaction.

[0003] However, recreating a user's home in 3D is a significant challenge. To allow users to place and move furniture in a 3D space that is sized and aligns with their real-world space, the primary challenge is reconstructing the 3D structure of the room's floor and walls. Initially, reconstructing the 3D scene of a user's home typically required specialized acquisition equipment and multiple photos of the space from various angles, which was time-consuming and labor-intensive. However, for consumer-oriented products, the goal is to complete 3D reconstruction of a room and restore the 3D structure of planar surfaces like the floor and walls with just a single image. Some algorithms, based on deep learning models, can detect the baseline of walls from 2D images and then predict the 3D structure of these surfaces. However, these algorithms may fail when the baseline is extensively obscured. Walls may be misidentified or extended incorrectly, making it impossible to accurately restore the 3D structure of the space's planes. For example, if the baseline of a TV wall is largely obscured by a TV cabinet, the detected baseline may be higher than the actual location. Summary of the Invention

[0004] The present application provides a method and electronic device for generating a three-dimensional spatial structure, which can more efficiently and accurately complete the construction of the three-dimensional spatial structure.

[0005] This application provides the following solutions:

[0006] A method for generating a three-dimensional spatial structure, comprising:

[0007] Obtain a single two-dimensional picture of the target spatial location;

[0008] Utilize a single Figure 3 The three-dimensional reconstruction technology is used to predict the three-dimensional point cloud corresponding to the two-dimensional image, and the panoptic segmentation technology is used to generate the three-dimensional background point cloud corresponding to multiple semantic plane areas with spatial background semantics;

[0009] Performing semantic plane detection based on the three-dimensional background point cloud to obtain three-dimensional parameters of the multiple semantic plane areas, the three-dimensional parameters including plane equations and point clouds within the planes;

[0010] A mesh space is generated according to the three-dimensional parameters of the multiple semantic plane areas so as to generate a three-dimensional spatial structure about the target spatial place.

[0011] Wherein, the semantic plane area includes: a ground plane;

[0012] The performing plane detection based on the three-dimensional background point cloud to obtain three-dimensional parameters of the multiple semantic plane areas includes:

[0013] By executing the random sampling consistency algorithm once, a plane that meets the ground saliency conditions is detected from the 3D background point cloud corresponding to the ground plane, and the plane is used as the detected ground plane, and the corresponding 3D parameters are output.

[0014] Among them, also include:

[0015] The extracted ground plane is rotated to be parallel to the horizontal plane in the standard coordinate system.

[0016] Wherein, the semantic plane area includes: ceiling plane;

[0017] The performing plane detection based on the three-dimensional background point cloud to obtain three-dimensional parameters of the multiple semantic plane areas includes:

[0018] By executing the plane detection algorithm once, a plane that meets the ceiling saliency conditions is detected from the 3D background point cloud corresponding to the ceiling plane, and this plane is used as the detected ceiling, and the corresponding 3D parameters are output.

[0019] Wherein, the semantic plane area includes: a wall;

[0020] The performing semantic plane detection based on the three-dimensional background point cloud to obtain three-dimensional parameters of the multiple semantic plane areas includes:

[0021] By iteratively executing the plane detection algorithm, multiple semantic plane detections are performed from the 3D background point cloud corresponding to the wall to obtain multiple planes that meet the wall saliency requirements.

[0022] The multiple planes are filtered according to preset constraints, the remaining planes are used as main wall surfaces, and the three-dimensional parameters of each main wall surface are output respectively.

[0023] The constraints include one or more of the following:

[0024] Normal vector constraint is used to filter out tilted planes and retain planes perpendicular to the ground plane. After the ground plane is rotated to be parallel to the horizontal plane in the standard coordinate system, the normal vector constraint is determined by determining whether the Y-axis component of the wall equation's normal vector is less than a threshold.

[0025] Height constraint to filter out planes whose wall height is lower than a height threshold; the wall height is determined by calculating the maximum distance from the wall point cloud to the ground, and the height threshold is related to the space height value;

[0026] Position constraint to filter out planes whose lowest position on the wall is higher than a position threshold. The lowest position of the wall is calculated by calculating the minimum distance from the wall point cloud to the ground. The position threshold is related to the space height value.

[0027] Deduplication constraints are used to filter out repeated planes in the same area, so that only the core planes in the same area are retained as the main walls. The core degree of the candidate plane in the area is judged by the length of its projection on the ground.

[0028] The step of generating a polygonal mesh space according to the three-dimensional parameters of the plurality of semantic plane regions includes:

[0029] approximating the ground plane and the ceiling plane in the plurality of semantic plane regions into polygons, and approximating the wall plane into a quadrilateral;

[0030] Determining a spatial height value of the target spatial location;

[0031] Determining, based on the three-dimensional parameters of the multiple semantic plane areas and the spatial height value, vertex position coordinates of the polygons corresponding to the ground plane and the ceiling plane, and vertex position coordinates of the quadrilateral corresponding to the wall surface;

[0032] According to the vertex position coordinates of the polygons and quadrilaterals, the connection relationship between the vertices is determined, and according to the connection relationship between the vertices, they are connected into a triangle or polygon mesh space.

[0033] The vertex position coordinates of the polygon corresponding to the ground plane are determined in the following way:

[0034] All three-dimensional point clouds corresponding to the target space are projected onto the ground plane, and the vertex coordinates of the convex polygons of the ground plane are approximately obtained through graphics convex hull optimization.

[0035] The vertex position coordinates of the polygon corresponding to the ceiling plane are determined in the following way:

[0036] The three-dimensional background point cloud corresponding to the ceiling plane is projected onto the ground plane, and the vertex coordinates of the convex polygon of the ceiling plane are approximated through graphics convex hull optimization, and the Y coordinates of the vertex coordinates are restored to the spatial height value.

[0037] The vertex position coordinates of the quadrilateral corresponding to the wall are determined in the following way:

[0038] Project the plane point corresponding to the wall onto the ground, use the line segment endpoint obtained by the projection as the bottom endpoint of the wall quadrilateral in the horizontal direction, and extend the bottom endpoint upward to the space height value as the top endpoint.

[0039] The spatial height value is determined by:

[0040] If the three-dimensional background point cloud includes a ceiling point cloud, the ceiling point cloud is subjected to a rotation and translation transformation using a homogeneous rotation and translation matrix, and an average distance from the ceiling point cloud to the ground is calculated, and the average distance is determined as the spatial height value;

[0041] If the three-dimensional background point cloud does not include the ceiling point cloud, the three-dimensional point cloud image is subjected to a rotation and translation transformation using a homogeneous rotation and translation matrix, and the average distance from the three-dimensional point cloud to the ground is calculated, and the average distance is determined as the space height.

[0042] Among them, also include:

[0043] The Mesh space is aligned with the pixel space of the two-dimensional image to generate a three-dimensional spatial structure about the target spatial location, so as to perform joint interaction based on the Mesh space and the pixel space of the two-dimensional image.

[0044] The aligning of the Mesh space with the two-dimensional image pixel space includes:

[0045] Based on the three-dimensional point cloud map of the target space and the camera parameter information extracted from the two-dimensional image, a projection change calculation is performed to determine the coordinates of the multiple vertices of the Mesh space in the normalized pixel space corresponding to the two-dimensional image.

[0046] The joint interaction based on the Mesh space and the pixel space of the two-dimensional image includes:

[0047] In response to the user's interactive operation of adding an object to the three-dimensional spatial structure of the target space, the three-dimensional position of the added object in the Mesh space is determined to determine the display size and angle of the object, and according to the alignment relationship between the Mesh space and the two-dimensional image pixel space, the operation result of the corresponding physical rule is displayed in the two-dimensional image pixel space.

[0048] Among them, also include:

[0049] Determine the plane where the added items are located;

[0050] In response to the user's interactive operation of dragging the added object to change its position, the display size and angle of the object in the two-dimensional image pixel space are adjusted according to the change of the three-dimensional position of the added object in the Mesh space during the dragging process, and the object is adsorbed within the plane and moved.

[0051] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of any of the aforementioned methods.

[0052] An electronic device, comprising:

[0053] one or more processors; and

[0054] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of any of the aforementioned methods.

[0055] A computer program product comprises a computer program / computer executable instructions, wherein the computer program / computer executable instructions are capable of implementing the steps of any of the aforementioned methods when executed by a processor in an electronic device.

[0056] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0057] Through the embodiment of this application, through a single Figure 33D reconstruction and panoramic segmentation extract 3D background point clouds that conform to the semantics of ceilings, floors, walls, etc., and then perform plane detection semantically. This can efficiently and accurately determine the 3D parameters of each plane, and then construct a Mesh space based on the 3D parameters of each semantic plane area to generate a 3D spatial structure about the target space. In this way, the user only needs to input a single 2D image to parse the 3D spatial structure of the target space, and can parametrically express the 3D spatial structure to facilitate interactions such as adding objects based on the reconstructed 3D spatial structure. Among them, in the process of parametrically expressing the 3D spatial structure, it no longer relies on the wall bottom line detection. Therefore, even if part of the wall bottom line is blocked, the 3D parameters of multiple semantic plane areas with spatial background semantics can still be more accurately extracted, thereby completing the construction of the 3D spatial structure more efficiently and accurately.

[0058] In a preferred implementation, the constructed Mesh space can be aligned with the pixel space of the original 2D image, allowing for joint interaction between the 3D Mesh space and the 2D image pixel space. Specifically, the 3D Mesh space can be used to determine the position, size, angle, etc. of specific objects in the 3D space, and the results of the relative physical rules can be reflected in the pixel space, allowing users to intuitively view the matching effects of specific objects in the soft and hard furnishings of the space, etc., allowing users to easily restore the room and interact in 3D on the front end, improving the practicality and operability of 3D space structure reconstruction and bringing users a more intuitive and convenient experience.

[0059] Among them, when performing wall detection, an iterative plane detection method is provided. In addition, a wall constraint filtering method based on normal vector constraints, height constraints, position constraints, wall deduplication, etc. is designed. The combination of the two can improve the robustness of wall detection.

[0060] Of course, any product implementing the present application does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0062] Figure 1 It is a schematic diagram of the system architecture provided by the embodiment of the present application;

[0063] Figure 2 is a flow chart of the method provided in an embodiment of the present application;

[0064] Figure 3 This is a single embodiment provided by the present application Figure 3 Schematic diagram of the dimensional reconstruction results;

[0065] Figure 4 is a schematic diagram of a background point cloud provided in an embodiment of the present application;

[0066] Figure 5 is a schematic diagram of the wall detection results provided in an embodiment of the present application;

[0067] Figure 6 This is a schematic diagram of Mesh reconstruction provided by an embodiment of the present application;

[0068] Figure 7 is a schematic diagram of spatial alignment provided by an embodiment of the present application;

[0069] Figure 8 Schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0070] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0071] In the embodiment of the present application, a 3D spatial structure can be generated from a single 2D image. In addition, during the generation process, the detection of the bottom line of the wall is not relied on. Instead, a single 2D image can be used to generate the 3D spatial structure. Figure 3 3D reconstruction technology reconstructs the basic 3D point cloud representation of the spatial scene based on a single 2D picture of the target space; then, a panoramic segmentation algorithm is used to obtain plane areas that conform to the semantics of the ceiling, ground, wall, etc. (which can be referred to as "semantic plane areas"), from which 3D background point clouds corresponding to multiple semantic plane areas are extracted; semantic plane detection is then performed on the 3D background point clouds that conform to the semantics of the ceiling, ground, wall, etc. to obtain 3D parametric expressions of multiple semantic plane areas that conform to the semantics of the ceiling, ground, wall, etc. Then, a Mesh space can be generated based on the parametric expression of multiple semantic plane areas and aligned with the original 2D picture to generate a complete 3D spatial structure. Subsequently, based on the 3D spatial structure, specific interactions such as adding items to the target space can be completed.

[0072] To facilitate understanding of the specific implementation solutions provided by the embodiments of the present application, several related concepts are briefly introduced below.

[0073] 1. 3D Reconstruction: 3D reconstruction involves inferring the three-dimensional structure and shape of an object or scene from 2D images or other sensor data. Common representations of 3D structures include point clouds, voxels, Gaussian splatting, and meshes. Traditional 3D reconstruction is achieved through multi-view geometry techniques, which utilize multiple 2D images from different angles to reconstruct the 3D structure of a scene. This process utilizes methods such as feature matching, parallax calculation, and camera pose estimation to gradually derive a 3D model of the scene.

[0074] 2. Single Figure 3 D reconstruction: single Figure 3 3D reconstruction relies only on a single 2D image. Directly predicting the 3D representation corresponding to each pixel from a single image requires a combination of deep learning and computer vision technologies. Common paradigms include monocular depth estimation (such as Depth Anything and Metric3D models) and monocular point cloud estimation (such as DUSt3R and MoGe models).

[0075] It should be noted here that based on the Figure 3 After 3D reconstruction technology completes 3D reconstruction, although a 3D structure of a space can be obtained, there is no parameterized information or semantic information, that is, it does not know which plane is the ceiling, which plane is the wall or the ground, and where their respective positions and boundaries are, etc. Therefore, if only a single Figure 3 After 3D reconstruction, the resulting 3D structure is usually in a non-interactive and non-editable state, and can usually only provide a preview function. In scenarios such as home improvement and home furnishing product purchases, users usually need to place the 3D model of a certain product in the 3D structure corresponding to their home space to view the matching effects with the home's hard decoration environment and other furniture. Therefore, it is necessary to interact with the 3D structure of the space, and the 3D model of the specific product needs to be placed on a plane with a certain specific semantics (for example, the sofa needs to be placed on the ground, the mural needs to be attached to the wall, etc.). When the user drags the product to change its position, it usually needs to be attached to the semantic plane and move, and it cannot float in the air or appear on an inappropriate plane. This requires obtaining the semantic information of the specific plane in the 3D structure and determining the corresponding position, boundary, etc. Obviously, if only a single Figure 3 Therefore, in the embodiment of the present application, a single Figure 3 3D reconstruction is only the first step in the specific implementation process, and subsequent Figure 3 Based on the 3D reconstruction, the specific 3D structure is parametrically expressed to obtain the semantic information of the specific spatial plane, as well as the corresponding position, boundary, etc., to support specific interactions such as adding objects to the space.

[0076] 3. Panoramic segmentation: Panoramic segmentation is a computer vision task that aims to classify each pixel in the image and distinguish different target instances at the same time to achieve a complete understanding of the scene. Panoramic segmentation is relative to instance segmentation. Instance segmentation usually only requires the classification of pixels related to foreground objects in the scene. For example, it is sufficient to identify sofas, TV cabinets, etc. in the space, while the walls and floors in the background will be ignored. Panoramic segmentation requires the classification of pixels corresponding to foreground objects and backgrounds in the scene. That is, for the background part, it is also necessary to determine whether it belongs to the wall, floor, ceiling, etc. In the embodiment of the present application, after completing a single Figure 3 After 3D reconstruction, panoptic segmentation techniques can be used to generate 3D background point clouds corresponding to multiple semantic planar regions with spatial background semantics (i.e., the aforementioned ceiling, floor, and wall). Most modern panoptic segmentation methods are based on deep learning, with popular panoptic segmentation models such as Oneformer and Segment Anything (SAM).

[0077] 4. Plane Detection: Plane detection is a key concept in computer vision, robotics, and 3D reconstruction. It refers to the process of identifying and extracting planar regions from point cloud data or images. The output of plane detection is typically a set of detected planes, the geometric parameters of each plane (such as its normal vector and position), and the points in the point cloud that belong to each plane.

[0078] 5. Semantic plane detection: It refers to understanding the semantics of each pixel in the scene, retaining the plane area that conforms to the semantics of the spatial background, and then combining the single pixel with the plane area. Figure 3 The result of D reconstruction is used to perform plane detection on the specified semantic plane area.

[0079] From the perspective of system architecture, the embodiments of the present application can provide a system for generating 3D spatial structures based on 2D images, see Figure 1 ,The system can be mainly divided into three modules: scene pre-analysis, plane detection and Mesh model construction.

[0080] First, in the scene pre-analysis module, the input single 2D image is processed through deep learning and computer vision technology to generate the preliminary analysis results of the scene. Figure 3 3D reconstruction technology is used to predict the corresponding 3D point cloud. Simultaneously, panoramic segmentation technology is used to extract mask information for semantic planar regions within the 2D image that correspond to the semantics of the ceiling, walls, and floor. The mask information is then combined with the 3D point cloud to generate a 3D background point cloud containing only semantic planar regions such as the ceiling, walls, and floor.

[0081] In the plane detection module, semantic plane detection technology is used to perform a series of visual processing and geometric calculations on the 3D background point cloud, resolving the 3D parameters of semantic plane areas such as ceilings, walls, and floors. (Specifically, these can be the 3D coordinates of the bounding polygons corresponding to the semantic plane areas. Bounding polygons are composed of many points, so the specific 3D parameters can be the 3D coordinates of the many points that make up the boundary polygons.) These parameters provide key geometric information for subsequent 3D spatial structure reconstruction.

[0082] Finally, in the Mesh model construction module, mesh models of each semantic plane area can be constructed based on the 3D parameters analyzed by the plane detection module. (Mesh modeling is a fundamental technology in computer graphics and 3D digitization. It involves using a series of polygons (triangles or quadrilaterals) to represent the surface of a 3D object. Its core lies in creating and editing the mesh structure composed of these polygons to accurately and efficiently describe the object's shape, texture, and surface properties.) In a preferred manner, it can also be aligned with the pixel space of the original image (the original 2D picture) to generate a more complete 3D spatial structure representation.

[0083] This process enables efficient reconstruction from a single 2D image into a structured, parameterized 3D space. This approach to semantic plane detection based on 3D semantic point clouds no longer focuses on the wall baseline. Therefore, as long as there is a visible area of ​​a semantic plane such as a wall, floor, or ceiling in the image (even if the wall baseline is partially obscured, as long as a portion is visible), the largest plane parameter containing this plane can be determined, thereby extracting and accurately restoring the 3D structure of semantic plane areas such as the ceiling, floor, and wall.

[0084] The specific implementation scheme provided in the embodiments of this application is introduced in detail below.

[0085] First, the present invention provides a method for generating a three-dimensional space structure. Figure 2 , the method may include the following steps:

[0086] S201: Acquire a single two-dimensional picture of the target space location.

[0087] This step S201 is the starting point of the solution. When it is implemented specifically, it can provide users (which can be consumer users, buyer users, etc. in the product information service system) with relevant function entrances for generating 3D spatial structure representations for target space places. After the user enters the relevant interface, he can upload a pre-prepared 2D picture. Among them, the target space place can specifically be the user's home place, or workplace, etc. The specific 2D picture can be a photo of a room taken by the user through an image acquisition device such as a mobile phone, and the single 2D picture can be used as the input picture of the process. Among them, the specific single 2D picture usually needs to include at least the ground and walls, and the ceiling part may or may not be included.

[0088] In addition, in the specific implementation, some pre-processing can be performed on the 2D images uploaded by the user. For example, the image resolution can be normalized so that the resolution of the 2D image input into the process is fixed at a certain value.

[0089] S202: Utilize a single Figure 3 The three-dimensional reconstruction technology is used to predict the three-dimensional point cloud corresponding to the two-dimensional image, and the panoramic segmentation technology is used to generate a three-dimensional background point cloud corresponding to multiple semantic plane areas with spatial background semantics.

[0090] After receiving a single 2D picture, you can Figure 3 The 3D reconstruction model predicts the 3D coordinates of each pixel in the image corresponding to the spatial point. The 3D coordinates are affine invariant and single Figure 3 The 3D reconstruction model can use a point cloud image as a 3D representation output. A point cloud image is a 3D representation with dimensions of H*W*3, where H*W represents the image resolution and 3 represents the XYZ coordinates of the point cloud.

[0091] Specifically in single Figure 3 When reconstructing in D, you can use pre-trained models such as MoGe to complete single Figure 3 D reconstruction. For example, in the specific implementation, after inputting a single 2D picture, the image encoder can be used to extract the features of the single image first, and then the convolution decoder can be used to output the predicted point cloud map. The point cloud map predicted by the model is affine invariant and can restore a 3D space with relatively consistent structure. In the reconstructed space, the ceiling remains parallel to the ground, the walls maintain good verticality, the overall structural error is small, and the geometric relationship is accurately preserved. In addition, the model can simultaneously predict the intrinsic parameters of the camera, which also solves the problem of subsequent projection transformation, and can align the reconstructed space with the pixel space without knowing the intrinsic parameters of the camera in advance. In actual use, the reconstruction model can be inferred in 512 dimensions, and the output point cloud map dimension is 512, which ensures a reconstruction speed of seconds while maintaining accuracy. The specific 3D reconstruction effect can be as follows Figure 3 As shown, Figure 3 (A) shows the original 2D image, and (B) shows the visualization result converted from the point cloud image.

[0092] In completing the order Figure 3 After D reconstruction, some segmentation models can be used to predict the panoramic segmentation results. For example, Oneformer is an advanced panoramic segmentation model that can handle semantic segmentation and instance segmentation tasks simultaneously under the same network framework. In an embodiment of the present application, segmentation models such as Oneformer can be fine-tuned based on 3D rendered self-made indoor home photo panoramic segmentation data, so that the fine-tuned segmentation model maintains higher accuracy in home scenes. The results of panoramic segmentation include segmentation results of semantic plane areas such as walls, floors, and ceilings. In actual use, the segmentation model performs reasoning in 512 dimensions and can complete segmentation prediction within half a second.

[0093] The output of panoptic segmentation includes component masks for semantic plane areas such as walls, floors, and ceilings. Applying each component mask to the point cloud (H*W*3) and performing matrix multiplication, and then recombining it into a point cloud (N*3), we can obtain the 3D background point cloud for semantic plane areas such as walls, floors, and ceilings. For example, for Figure 3 The original 2D image shown in (A) and the 3D background point cloud obtained here can be Figure 4 As shown, the range shown by 41 is the ceiling point cloud, the range shown by 42 is the wall point cloud, and the range shown by 43 is the ground point cloud.

[0094] In an optional approach, to eliminate the influence of outliers on the 3D background point cloud, the 3D background point cloud can be averaged downsampled and / or statistically removed to obtain a filtered 3D background point cloud. Finally, the sign of the point cloud coordinates is corrected so that the camera coordinate system has the Y axis pointing upward, the X axis pointing to the right, and the Z axis pointing backward.

[0095] Among them, the so-called average downsampling is used to reduce the density of the point cloud while retaining its overall structural characteristics. In actual operation, the points in each voxel can be sampled according to the side length of the downsampled voxel (for example, 0.03), and usually one representative point is retained. Statistical outlier removal is used to remove noise points or abnormal points in the point cloud. This method is based on the principle of statistical analysis and determines which points are outliers by calculating the distance distribution between each point and its neighboring points. In actual operation, the number of neighboring points can be set to 50 (or other values), and the standard deviation threshold can be set to 3 times (or other values) the average neighborhood distance.

[0096] S203: Perform semantic plane detection based on the three-dimensional background point cloud to obtain three-dimensional parameters of the multiple semantic plane areas, where the three-dimensional parameters include a plane equation and a point cloud within the plane.

[0097] After obtaining the 3D background point cloud corresponding to the semantic plane areas such as walls, floors, and ceilings, semantic plane detection can be performed based on this 3D background point cloud to obtain the 3D parameters of the multiple semantic plane areas, where the specific 3D parameters can include the plane equation and the point cloud within the plane for each semantic plane area.

[0098] Among them, for planes with different semantics, the specific semantic plane detection methods may be different. First, the prior based on which the ground and ceiling detection is based may be that a room usually only needs one plane to represent the ground and one plane to represent the ceiling. Therefore, by inputting the ground point cloud, the plane equation and inner points of the ground can be output through the plane detection algorithm. Similarly, by inputting the ceiling point cloud, the plane equation and inner points of the ceiling can be output through the plane detection algorithm. Specifically, for the ground plane, the plane detection algorithm can be executed once to detect the plane that meets the ground saliency conditions from the 3D background point cloud corresponding to the ground plane, and the plane is used as the detected ground plane, and the corresponding 3D parameters are output. For the ceiling plane, the plane detection algorithm can be executed once to detect the plane that meets the ceiling saliency conditions from the 3D background point cloud corresponding to the ceiling plane, and the plane is used as the detected ceiling, and the corresponding 3D parameters are output.

[0099] There are many specific plane detection algorithms. For example, one of them may be the RANSAC (Random Sample Consensus) algorithm, which is used to fit a plane from point cloud data and extract plane parameters. The workflow is as follows:

[0100] a. Randomly select a minimum number of points (for plane fitting, at least 3 points are required) and calculate a candidate plane equation based on these points;

[0101] b. Calculate the distance from all points to the plane and count the number of points whose distance is less than the threshold (inliers);

[0102] c. Repeat the above process multiple times and select the plane with the largest number of internal points as the final result.

[0103] Specifically, when using the aforementioned RANSAC algorithm to detect ground planes, based on the prior assumption that a room's ground is represented by a single plane, a single RANSAC algorithm run can be used to extract the most prominent plane in the ground point cloud and output its plane equation and a point cloud of points within that plane. Similarly, for ceiling planes, a single RANSAC algorithm run is required to extract the most prominent plane in the ceiling point cloud and output its plane equation and a point cloud of points within that plane. If the user's uploaded 2D image does not capture the ceiling (i.e., does not have a ceiling point cloud), the ceiling detection step can be omitted.

[0104] In addition, in 3D space reconstruction, the ground is usually a key reference plane. However, due to reasons such as the tilt of the camera angle, single Figure 3 The ground plane reconstructed by D may not be parallel to the spatial coordinate system, which will make subsequent spatial modeling and alignment complicated. Therefore, in an optional way, in order to simplify the complexity of the subsequent algorithm, a ground plane axis alignment process can be performed. The goal of this process is to rotate the extracted ground plane to be parallel to the XOZ plane in the standard coordinate system (the Y axis of the camera coordinate system is upward). This operation can reduce the difficulty of subsequent geometric calculations and spatial reconstruction, while improving the robustness and efficiency of the algorithm. In specific implementation, the rotation axis and rotation angle can be first calculated based on the ground plane equation detected in the previous steps and the normal vector [0, 1, 0] of the XOZ plane equation in the standard coordinate system, and then the rotation matrix R is constructed using the rotation formula. Find any point on the ground and obtain the translation vector by solving the plane equation, and finally obtain the homogeneous rotation and translation matrix T. Then, the ground plane axis alignment process can be completed by performing a rotation and translation transformation on T.

[0105] The above describes the plane detection of the floor and ceiling. Next, we will introduce the plane detection of the wall. Regarding the wall, unlike the floor and ceiling, there is usually more than one wall in a room. Therefore, wall detection can be an iterative plane detection process. Each time a plane is detected, the points in the plane are removed, and the planes of the remaining point cloud are detected. The process is repeated until the termination condition is met. In other words, by executing the plane detection algorithm multiple times, multiple plane detections can be performed from the 3D background point cloud corresponding to the wall to obtain multiple planes whose wall significance meets the conditions. Then, the multiple detected planes can be filtered according to the preset constraints, and the remaining planes are used as the main wall surfaces, and the 3D parameters of each main wall surface are output separately.

[0106] In specific implementation, the RANSAC algorithm can also be used for wall detection. Specifically, first input the wall point cloud, use the RANSAC algorithm to extract the most significant plane of the current point cloud, then remove the internal points of the current plane to obtain the remaining wall point cloud, and continue plane detection. Among them, in order to prevent the plane detection algorithm from classifying irrelevant point clouds as internal points of the wall, clustering and filtering can also be performed on the internal points of the wall. Specifically, when performing clustering and filtering, you can input the initial detected plane internal points, use algorithms such as DBSCAN to implement density-based connected domain clustering, and only take the largest connected domain point cloud as the internal point of the current most significant plane to maintain the single connectivity of the wall.

[0107] Iterative plane detection of walls requires setting a termination condition. This setting can avoid detecting small, nonfunctional, and minor walls while improving computational efficiency. In practice, the present embodiment can control the termination condition by setting two parameters: the number of iterations, for example, setting a maximum of five planes to be detected for a wall point cloud; and the internal point threshold for the current plane, for example, set to 200. If the number of remaining point clouds within a plane is less than 200, then the wall plane detection for that plane is terminated.

[0108] After iterative plane detection, the plane belonging to the main wall can be detected from the 3D background point cloud corresponding to the wall, and the corresponding 3D parameters can be output. Figure 4 The wall point cloud part in the 3D background point cloud can be decomposed into Figure 5 The two planar point clouds shown at 51 and 52 represent the two main wall planes.

[0109] Furthermore, due to the complexity of the user's room environment, walls may appear incomplete due to occlusion. Complex hard structures (such as doors, windows, and decorative objects) can also interfere with detection, leading to false or missed detections. Furthermore, wall detection typically involves multi-plane detection, which is susceptible to the effects of segmentation results. For example, the introduction of point cloud data unrelated to the wall further increases the risk of false detection. Therefore, wall detection can also introduce additional constraints and filtering conditions to improve the accuracy, reliability, and robustness of detection results and mitigate the adverse effects of environmental complexity.

[0110] Specifically, a series of wall equations and interior points have been obtained through the aforementioned wall detection. To facilitate subsequent calculations, the wall equations and interior points can first be transformed using the homogeneous rotation and translation matrix T described above to obtain the wall equations and interior points after rotation and translation. Wall constraint filtering can then be performed based on the transformed parameters. Specific constraints may include one or more of the following:

[0111] a. Normal vector constraint:

[0112] The wall and the ground are perpendicular to each other, so after the constraint transformation, the wall normal vector is parallel to the XOZ plane in the standard coordinate system. Specifically, if the Y-axis component of the wall normal vector is less than a threshold (e.g., 0.1), the normal constraint is considered met. This ensures that the detected wall plane is perpendicular to the ground, and severely tilted planes can be filtered out.

[0113] b. Height Constraint

[0114] The height of the wall from the ground is typically similar to the room height. The wall height can be calculated as the maximum distance from the wall point cloud to the ground plane, which can be converted to the maximum Y coordinate of the wall point cloud. The height of the target space (referred to as the room height) can also be calculated. The wall height constraint filters out walls above a height threshold, thereby filtering out very short walls. In practice, the height threshold can be set to 0.8 times the room height, for example.

[0115] c. Position constraints

[0116] When restoring the room structure, you can also ignore hard-mounted ceilings, which may be identified as a small portion of the top wall. Position constraints can be used to filter out these top ceiling walls. Specifically, the minimum distance from the wall point cloud to the ground plane can be calculated, which serves as the lowest point on the wall. This can then be converted to the minimum Y coordinate for calculating the wall point cloud. A threshold can be set, such as 0.8 times the room height. If the lowest point on the wall exceeds the threshold, it can be filtered out.

[0117] The concept of room height is involved here. The room height parameter not only provides guidance for wall constraints but also serves as an important basis for the subsequent determination of wall vertices. Therefore, the present embodiment also provides a method for calculating room height. Specifically, the room height can be calculated based on both the presence and absence of a ceiling in the 2D image.

[0118] Case 1: There is a ceiling in the picture

[0119] The ceiling point cloud can be rotated and translated using matrix T to calculate the average distance from the ceiling point cloud to the ground plane. Since the ground plane was aligned with the standard coordinate system in the previous step, making it overlap with the XOZ plane, the average distance from the ceiling to the ground plane can be calculated as the average Y coordinate of the ceiling point cloud. In this case, the average distance d can be used as the room height.

[0120] Case 2: There is no ceiling in the picture

[0121] The reconstructed scene point cloud is rotated and translated using matrix T to calculate the maximum distance between the scene point cloud and the ground plane. Similarly, when the ground plane is aligned with the standard coordinate system, since the ground plane overlaps with the XOZ plane, the maximum distance between the scene point cloud and the ground plane is the maximum Y coordinate of the scene point cloud. In this case, the maximum distance d can be used as the room height.

[0122] d. Wall deduplication constraint

[0123] When interfering planes exist on a wall, the wall can be deduplicated, filtering out repeated planes in the same area and retaining only the most core planes in the same area as the main wall. The core of the deduplication constraint is to retain the most core planes in the same area. The core degree of the plane in the area can be judged by the projection length of the candidate plane on the ground plane. The projection length is abstracted as the two endpoints of the projection line segment. In actual operation, hash table storage and similarity checks can be used to retain unique wall planes. For example, in a specific implementation, the specific process of deduplication processing may include:

[0124] The first step is to sort the walls in descending order, giving priority to larger surfaces.

[0125] The second step is to traverse each wall plane and calculate its hash value based on the plane equation. Check whether the current plane is similar to the stored plane (based on the plane equation). If so, calculate the projection range of the plane's projected line segment on the plane. If the ranges intersect, a plane with a larger range can be formed. If the ranges do not intersect, proceed to the next plane.

[0126] The third step is to convert the unique planes in the hash table into a list and return it.

[0127] The above describes the specific implementation of semantic plane detection for semantic plane areas such as ceilings, floors, and walls. Through this step, the plane equations corresponding to semantic plane areas such as ceilings, floors, and walls, as well as the point cloud within the planes, can be obtained, thereby determining the position and boundaries of each semantic plane.

[0128] S204: Generate a mesh space according to the three-dimensional parameters of the multiple semantic plane areas, so as to generate a three-dimensional spatial structure about the target spatial place.

[0129] After obtaining the 3D parameters of multiple semantic plane areas (that is, the plane equation of each plane and the point cloud within the plane), the corresponding Mesh space can be generated based on these 3D parameters. When generating the Mesh space, the vertex coordinates of each plane can be calculated first. Specifically, the ground and ceiling can be approximated as polygons, and the walls can be approximated as quadrilaterals. Then, based on the 3D parameters of multiple semantic plane areas and the calculated room height value, the vertex coordinates of each polygon and quadrilateral can be determined, and then based on the vertex position coordinates of the polygons and quadrilaterals, the connection relationship between the vertices can be determined, and based on the connection relationship between the vertices, they can be connected into a triangular or polygonal mesh Mesh space.

[0130] When determining the vertex position coordinates of the polygon corresponding to the ground plane, since the user's room is usually filled with various items such as furniture, it is not enough to restore the complete ground polygon by only segmenting the ground point cloud. Based on this situation, in the embodiment of the present application, the vertex position coordinates of the polygon corresponding to the ground plane can be determined in the following way: First, the vertex position coordinates of the polygon corresponding to the ground plane can be determined by the single point cloud. Figure 3 The scene point cloud obtained by 3D reconstruction is projected onto the ground plane. Then, through graphics convex hull optimization, the vertices of the convex polygon of the ground are approximated and the vertex coordinates are determined. In an optimal method, after obtaining the convex polygon vertices, the optimized ground plane vertices can be obtained by clipping the key points of the wall.

[0131] When determining the vertex coordinates of the polygon corresponding to the ceiling plane, since the ceiling is usually segmented to obtain all visible parts of the ceiling in the image, and the lamps will not affect the integrity of the ceiling, the ceiling point cloud can be projected onto the ground. Through graphics convex hull optimization, the polygon vertices of the ceiling can be approximated. The Y coordinates of the vertex coordinates are then restored to the room height, thus forming the polygon vertices of the ceiling.

[0132] When determining the vertex position coordinates of the quadrilateral corresponding to the wall, you can first project the plane point corresponding to the wall to the ground, and use the projected line segment endpoint as the bottom endpoint of the wall quadrilateral in the horizontal direction, and then extend the bottom endpoint upward to the spatial height value as the top endpoint to obtain the vertex coordinates of the quadrilateral, thereby obtaining the coordinates of the four vertices of the wall quadrilateral.

[0133] After obtaining the plane equations for each semantic plane region and the corresponding polygon or quadrilateral vertex coordinates, we can further determine the connectivity between the vertices and then generate a triangular or polygonal mesh based on this connectivity. In practice, since the convex hull polygon vertices corresponding to the floor and ceiling and the quadrilateral vertices corresponding to the wall are known to be connected in sequence, we can simply generate triangle indices in that order and ultimately construct the mesh's vertex list and triangle index list.

[0134] The schematic diagram of plane Mesh reconstruction can be as follows Figure 6 As shown in the figure, compared to the original 2D image, the ground plane, ceiling plane, and two walls have been reconstructed. Since the walls represent the area where the user's soft furnishings are placed, and the hard ceiling cannot accommodate wall furnishings such as hanging paintings, the hard ceiling is ignored during the reconstruction process, resulting in a slight gap between the visible wall and ceiling planes.

[0135] After constructing the mesh space, the relevant information about the 3D space corresponding to the specific location is already known, including parameterized information such as the positions of various semantic planes. Theoretically, this mesh space can be used to interact with items within the space, including adding a new product model to the space. However, to allow users to more intuitively see how the added product works with the existing soft and hard furnishings in the space, the mesh space can be further aligned with the pixel space of the original 2D image. This allows the specific product model to be placed based on the mesh space, including determining the product model's position, size, and rotation angle in the space, while the user is presented with the results of the relative physical rules in the pixel space. This allows users to intuitively see how the product will fit in with the other soft and hard furnishings in the room after being placed. This allows users to easily restore the room and interact with it in 3D on the front end, significantly improving the practicality and operability of 3D room structure reconstruction and providing a more intuitive and convenient user experience.

[0136] Specifically, there are several ways to align the mesh space with the pixel space of the original 2D image. For example, one approach involves performing a projection change calculation based on the 3D point cloud of the target space and the camera parameter information extracted from the 2D image to determine the UV (normalized pixel space) coordinates corresponding to multiple vertices of the mesh space in the 2D image. In other words, if the original image is imaged on the camera's focal plane, the UV coordinates of the mesh space are aligned one-to-one with the original image, and the entire 3D space of the mesh space is aligned with the pixel space of the original image.

[0137] The diagram of the alignment between the mesh space and the pixel space of the 2D image can be found in Figure 7 , where each closed wireframe represents a plane, indicating that the established 3D space is fully aligned with the pixel space.

[0138] After completing the alignment between the mesh space and the pixel space of the original image, you can perform any 3D editing in the 3D space of the mesh space, such as adding objects to the space, translating and rotating the added objects, etc., which can all reflect the results of the relative physical rules in the pixel space.

[0139] Specifically, when the user performs an interactive operation of adding an object to the 3D spatial structure of the target space, the 3D position of the added object in the Mesh space can be determined first to determine the display size and angle of the object, and then the display effect of placing the object in the 2D image pixel space can be displayed based on the alignment relationship between the Mesh space and the 2D image pixel space.

[0140] In addition, the plane where the added items are located can also be determined. For example, the plane to which they belong can be determined based on the category of the specific item, the location where they are placed, etc. When the user performs interactive operations such as translation and rotation on the added items, the display size and angle of the items in the 2D image can be adjusted according to the changes in the 3D position of the added items in the Mesh space during the translation and rotation process, and the items can be adsorbed within the plane and moved. In other words, if the user adds a sofa, then during the translation or rotation process, the 3D model of the sofa will move attached to the ground plane and will not float in the air.

[0141] It can be seen that through the embodiment of the present application, by a single Figure 3 3D reconstruction and panoramic segmentation extract 3D background point clouds that conform to the semantics of ceilings, floors, walls, etc., and then perform plane detection semantically. This can efficiently and accurately determine the 3D parameters of each plane, and then construct a Mesh space based on the 3D parameters of each semantic plane area to generate a 3D spatial structure about the target space. In this way, the user only needs to input a single 2D image to parse the 3D spatial structure of the target space, and can parametrically express the 3D spatial structure to facilitate interactions such as adding objects based on the reconstructed 3D spatial structure. Among them, in the process of parametrically expressing the 3D spatial structure, it no longer relies on the wall bottom line detection. Therefore, even if part of the wall bottom line is blocked, the 3D parameters of multiple semantic plane areas with spatial background semantics can still be more accurately extracted, thereby completing the construction of the 3D spatial structure more efficiently and accurately.

[0142] In a preferred implementation, the constructed Mesh space can be aligned with the pixel space of the original 2D image, allowing for joint interaction between the 3D Mesh space and the 2D image pixel space. Specifically, the 3D Mesh space can be used to determine the position, size, angle, etc. of specific objects in the 3D space, and the results of the relative physical rules can be reflected in the pixel space, allowing users to intuitively view the matching effects of specific objects in the soft and hard furnishings of the space, etc., allowing users to easily restore the room and interact in 3D on the front end, improving the practicality and operability of 3D space structure reconstruction and bringing users a more intuitive and convenient experience.

[0143] Among them, when performing wall detection, an iterative plane detection method is provided. In addition, a wall constraint filtering method based on normal vector constraints, height constraints, position constraints, wall deduplication, etc. is designed. The combination of the two can improve the robustness of wall detection.

[0144] It should be noted that the embodiments of the present application may involve the use of user data. In actual applications, user-specific personal data can be used in the scheme described herein within the scope permitted by applicable laws and regulations, subject to the requirements of applicable laws and regulations of the country where the user is located (for example, with the user's explicit consent, effective notification to the user, etc.).

[0145] Corresponding to the aforementioned method embodiment, an embodiment of the present application further provides a device for generating a three-dimensional spatial structure, which may include:

[0146] A two-dimensional image acquisition unit, configured to acquire a single two-dimensional image of a target space;

[0147] one Figure 3 dimensional reconstruction unit for utilizing single Figure 3 The three-dimensional reconstruction technology is used to predict the three-dimensional point cloud corresponding to the two-dimensional image, and the panoptic segmentation technology is used to generate the three-dimensional background point cloud corresponding to multiple semantic plane areas with spatial background semantics;

[0148] a plane detection unit, configured to perform semantic plane detection based on the three-dimensional background point cloud to obtain three-dimensional parameters of the plurality of semantic plane regions, the three-dimensional parameters including a plane equation and a point cloud of points within the plane;

[0149] A mesh space generating unit is used to generate a mesh space according to the three-dimensional parameters of the multiple semantic plane areas, so as to generate a three-dimensional space structure about the target space place.

[0150] Wherein, the semantic plane area includes: a ground plane;

[0151] At this time, the plane detection unit can be specifically used for:

[0152] By executing the random sampling consistency algorithm once, a plane that meets the ground saliency conditions is detected from the 3D background point cloud corresponding to the ground plane, and the plane is used as the detected ground plane, and the corresponding 3D parameters are output.

[0153] At this time, the device may further include:

[0154] The ground plane alignment unit is used to rotate the extracted ground plane to be parallel to the horizontal plane in the standard coordinate system.

[0155] Additionally, the semantic plane region includes: a ceiling plane;

[0156] At this time, the plane detection unit can be specifically used for:

[0157] By executing the plane detection algorithm once, a plane that meets the ceiling saliency conditions is detected from the 3D background point cloud corresponding to the ceiling plane, and this plane is used as the detected ceiling, and the corresponding 3D parameters are output.

[0158] The semantic plane area may also include: a wall;

[0159] At this time, the plane detection unit can be specifically used for:

[0160] By iteratively executing the plane detection algorithm, multiple semantic plane detections are performed from the 3D background point cloud corresponding to the wall to obtain multiple planes that meet the wall saliency requirements.

[0161] The multiple planes are filtered according to preset constraints, the remaining planes are used as main wall surfaces, and the three-dimensional parameters of each main wall surface are output respectively.

[0162] Specifically, the constraint conditions may include one or more of the following:

[0163] Normal vector constraint is used to filter out tilted planes and retain planes perpendicular to the ground plane. After the ground plane is rotated to be parallel to the horizontal plane in the standard coordinate system, the normal vector constraint is determined by determining whether the Y-axis component of the wall equation's normal vector is less than a threshold.

[0164] Height constraint to filter out planes whose wall height is lower than a height threshold; the wall height is determined by calculating the maximum distance from the wall point cloud to the ground, and the height threshold is related to the space height value;

[0165] Position constraint to filter out planes whose lowest position on the wall is higher than a position threshold. The lowest position of the wall is calculated by calculating the minimum distance from the wall point cloud to the ground. The position threshold is related to the space height value.

[0166] Deduplication constraints are used to filter out repeated planes in the same area, so that only the core planes in the same area are retained as the main walls. The core degree of the candidate plane in the area is judged by the length of its projection on the ground.

[0167] In specific implementation, the Mesh space generation unit can be used for:

[0168] approximating the ground plane and the ceiling plane in the plurality of semantic plane regions into polygons, and approximating the wall plane into a quadrilateral;

[0169] Determining a spatial height value of the target spatial location;

[0170] Determining, based on the three-dimensional parameters of the multiple semantic plane areas and the spatial height value, vertex position coordinates of the polygons corresponding to the ground plane and the ceiling plane, and vertex position coordinates of the quadrilateral corresponding to the wall surface;

[0171] According to the vertex position coordinates of the polygons and quadrilaterals, the connection relationship between the vertices is determined, and according to the connection relationship between the vertices, they are connected into a triangle or polygon mesh space.

[0172] Specifically, the vertex position coordinates of the polygon corresponding to the ground plane can be determined in the following way:

[0173] All three-dimensional point clouds corresponding to the target space are projected onto the ground plane, and the vertex coordinates of the convex polygons of the ground plane are approximately obtained through graphics convex hull optimization.

[0174] Determine the vertex position coordinates of the polygon corresponding to the ceiling plane in the following way:

[0175] The three-dimensional background point cloud corresponding to the ceiling plane is projected onto the ground plane, and the vertex coordinates of the convex polygon of the ceiling plane are approximated through graphics convex hull optimization, and the Y coordinates of the vertex coordinates are restored to the spatial height value.

[0176] Determine the vertex coordinates of the quadrilateral corresponding to the wall in the following way:

[0177] Project the plane point corresponding to the wall onto the ground, use the line segment endpoint obtained by the projection as the bottom endpoint of the wall quadrilateral in the horizontal direction, and extend the bottom endpoint upward to the space height value as the top endpoint.

[0178] The space height value is determined by:

[0179] If the three-dimensional background point cloud includes a ceiling point cloud, the ceiling point cloud is subjected to a rotation and translation transformation using a homogeneous rotation and translation matrix, and an average distance from the ceiling point cloud to the ground is calculated, and the average distance is determined as the spatial height value;

[0180] If the three-dimensional background point cloud does not include the ceiling point cloud, the three-dimensional point cloud image is subjected to a rotation and translation transformation using a homogeneous rotation and translation matrix, and the average distance from the three-dimensional point cloud to the ground is calculated, and the average distance is determined as the space height.

[0181] In addition, the device may further include:

[0182] A spatial alignment unit is configured to align the Mesh space with the pixel space of the two-dimensional image to generate a three-dimensional spatial structure about the target spatial location, so as to perform joint interaction based on the Mesh space and the pixel space of the two-dimensional image.

[0183] Specifically, the spatial alignment unit can be used to:

[0184] Based on the three-dimensional point cloud map of the target space and the camera parameter information extracted from the two-dimensional image, a projection change calculation is performed to determine the coordinates of the multiple vertices of the Mesh space in the normalized pixel space corresponding to the two-dimensional image.

[0185] Specifically, when performing joint interaction based on the Mesh space and the pixel space of the two-dimensional image, the spatial alignment unit can be used to:

[0186] In response to the user's interactive operation of adding an object to the three-dimensional spatial structure of the target space, the three-dimensional position of the added object in the Mesh space is determined to determine the display size and angle of the object, and according to the alignment relationship between the Mesh space and the two-dimensional image pixel space, the operation result of the corresponding physical rule is displayed in the two-dimensional image pixel space.

[0187] At this time, the device may further include:

[0188] A plane determining unit, used to determine the plane where the added item is located;

[0189] The interactive operation response unit is used to respond to the user's interactive operation of dragging the added object to change its position, adjust the display size and angle of the object in the two-dimensional image pixel space according to the change of the three-dimensional position of the added object in the Mesh space during the dragging process, and make the object adsorbed in the plane for movement.

[0190] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.

[0191] And an electronic device comprising:

[0192] one or more processors; and

[0193] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the aforementioned method embodiments.

[0194] A computer program product includes a computer program / computer executable instructions, which implement the steps of the method described in the above method embodiment when executed by a processor in an electronic device.

[0195] in, Figure 8 The electronic device architecture is shown as an example, and may include a processor 810, a video display adapter 811, a disk drive 812, an input / output interface 813, a network interface 814, and a memory 820. The processor 810, the video display adapter 811, the disk drive 812, the input / output interface 813, the network interface 814, and the memory 820 may be communicatively connected via a communication bus 830.

[0196] Among them, the processor 810 can be implemented by a general-purpose CPU (Central Processing Unit, processor), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in this application.

[0197] The memory 820 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 820 can store an operating system 821 for controlling the operation of the electronic device 800, and a basic input and output system (BIOS) for controlling the low-level operations of the electronic device 800. In addition, a web browser 823, a data storage management system 824, and a three-dimensional space structure generation processing system 825, etc. can also be stored. The above-mentioned three-dimensional space structure generation processing system 825 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided by the present application is implemented by software or firmware, the relevant program code is stored in the memory 820 and is called and executed by the processor 810.

[0198] The input / output interface 813 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.

[0199] The network interface 814 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).

[0200] The bus 830 comprises a pathway for transmitting information between the various components of the device (eg, the processor 810 , the video display adapter 811 , the disk drive 812 , the input / output interface 813 , the network interface 814 , and the memory 820 ).

[0201] It should be noted that although the above device only shows the processor 810, video display adapter 811, disk drive 812, input / output interface 813, network interface 814, memory 820, bus 830, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may also include only the components necessary to implement the solution of the present application, and does not necessarily include all the components shown in the figure.

[0202] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.

[0203] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0204] The above is a detailed introduction to the method and electronic device for generating a three-dimensional spatial structure provided by this application. Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is only intended to help understand the method and core ideas of this application. At the same time, for those skilled in the art, based on the ideas of this application, there may be changes in the specific implementation methods and application scope. In summary, the contents of this specification should not be understood as limiting this application.

Claims

1. A method for generating a three-dimensional spatial structure, characterized in that: include: Obtain a single two-dimensional picture of the target spatial location; Using a single-image 3D reconstruction technique to predict a 3D point cloud corresponding to the 2D image, and using a panoptic segmentation technique to generate a 3D background point cloud corresponding to multiple semantic plane regions having spatial background semantics; Performing semantic plane detection based on the three-dimensional background point cloud to obtain three-dimensional parameters of the multiple semantic plane areas, the three-dimensional parameters including plane equations and point clouds within the planes; A mesh space is generated according to the three-dimensional parameters of the multiple semantic plane areas so as to generate a three-dimensional spatial structure about the target spatial place.

2. The method according to claim 1, characterized in that The semantic plane area includes: a ground plane; The performing plane detection based on the three-dimensional background point cloud to obtain three-dimensional parameters of the multiple semantic plane areas includes: By executing the random sampling consistency algorithm once, a plane that meets the ground saliency conditions is detected from the 3D background point cloud corresponding to the ground plane, and the plane is used as the detected ground plane, and the corresponding 3D parameters are output.

3. The method according to claim 2, characterized in that Also includes: The extracted ground plane is rotated to be parallel to the horizontal plane in the standard coordinate system.

4. The method according to claim 1, wherein The semantic plane area includes: a ceiling plane; The performing plane detection based on the three-dimensional background point cloud to obtain three-dimensional parameters of the multiple semantic plane areas includes: By executing the plane detection algorithm once, a plane that meets the ceiling saliency conditions is detected from the 3D background point cloud corresponding to the ceiling plane, and this plane is used as the detected ceiling, and the corresponding 3D parameters are output.

5. The method according to claim 1, characterized in that The semantic plane area includes: a wall; The performing semantic plane detection based on the three-dimensional background point cloud to obtain three-dimensional parameters of the multiple semantic plane areas includes: By iteratively executing the plane detection algorithm, multiple semantic plane detections are performed from the 3D background point cloud corresponding to the wall to obtain multiple planes that meet the wall saliency requirements. The multiple planes are filtered according to preset constraints, the remaining planes are used as main wall surfaces, and the three-dimensional parameters of each main wall surface are output respectively.

6. The method according to claim 5, characterized in that The constraints include one or more of the following: Normal vector constraint is used to filter out tilted planes and retain planes perpendicular to the ground plane. After the ground plane is rotated to be parallel to the horizontal plane in the standard coordinate system, the normal vector constraint is determined by determining whether the Y-axis component of the wall equation's normal vector is less than a threshold. Height constraint to filter out planes whose wall height is lower than a height threshold; the wall height is determined by calculating the maximum distance from the wall point cloud to the ground, and the height threshold is related to the space height value; Position constraint to filter out planes whose lowest position on the wall is higher than a position threshold. The lowest position of the wall is calculated by calculating the minimum distance from the wall point cloud to the ground. The position threshold is related to the space height value. Deduplication constraints are used to filter out repeated planes in the same area, so that only the core planes in the same area are retained as the main walls. The core degree of the candidate plane in the area is judged by the length of its projection on the ground.

7. The method according to claim 1, characterized in that Generating a polygonal mesh space according to the three-dimensional parameters of the plurality of semantic plane areas includes: approximating the ground plane and the ceiling plane in the plurality of semantic plane regions into polygons, and approximating the wall plane into a quadrilateral; Determining a spatial height value of the target spatial location; Determining, based on the three-dimensional parameters of the multiple semantic plane areas and the spatial height value, vertex position coordinates of the polygons corresponding to the ground plane and the ceiling plane, and vertex position coordinates of the quadrilateral corresponding to the wall surface; According to the vertex position coordinates of the polygons and quadrilaterals, the connection relationship between the vertices is determined, and according to the connection relationship between the vertices, they are connected into a triangle or polygon mesh space.

8. The method according to claim 7, characterized in that Determine the vertex position coordinates of the polygon corresponding to the ground plane in the following way: All three-dimensional point clouds corresponding to the target space are projected onto the ground plane, and the vertex coordinates of the convex polygons of the ground plane are approximately obtained through graphics convex hull optimization.

9. The method according to claim 7, characterized in that Determine the vertex position coordinates of the polygon corresponding to the ceiling plane in the following way: The three-dimensional background point cloud corresponding to the ceiling plane is projected onto the ground plane, and the vertex coordinates of the convex polygon of the ceiling plane are approximated through graphics convex hull optimization, and the Y coordinates of the vertex coordinates are restored to the spatial height value.

10. The method according to claim 7, characterized in that Determine the vertex coordinates of the quadrilateral corresponding to the wall in the following way: Project the plane point corresponding to the wall onto the ground, use the line segment endpoint obtained by the projection as the bottom endpoint of the wall quadrilateral in the horizontal direction, and extend the bottom endpoint upward to the space height value as the top endpoint.

11. The method according to claim 6 or 7, characterized in that The space height value is determined by: If the three-dimensional background point cloud includes a ceiling point cloud, the ceiling point cloud is subjected to a rotation and translation transformation using a homogeneous rotation and translation matrix, and an average distance from the ceiling point cloud to the ground is calculated, and the average distance is determined as the spatial height value; If the three-dimensional background point cloud does not include the ceiling point cloud, the three-dimensional point cloud image is subjected to a rotation and translation transformation using a homogeneous rotation and translation matrix, and the average distance from the three-dimensional point cloud to the ground is calculated, and the average distance is determined as the space height.

12. The method according to claim 1, characterized in that Also includes: The Mesh space is aligned with the pixel space of the two-dimensional image to generate a three-dimensional spatial structure about the target spatial location, so as to perform joint interaction based on the Mesh space and the pixel space of the two-dimensional image.

13. The method according to claim 12, characterized in that The aligning the Mesh space with the two-dimensional image pixel space includes: Based on the three-dimensional point cloud map of the target space and the camera parameter information extracted from the two-dimensional image, a projection change calculation is performed to determine the coordinates of the multiple vertices of the Mesh space in the normalized pixel space corresponding to the two-dimensional image.

14. The method according to claim 12, characterized in that The joint interaction based on the Mesh space and the pixel space of the two-dimensional image includes: In response to the user's interactive operation of adding an object to the three-dimensional spatial structure of the target space, the three-dimensional position of the added object in the Mesh space is determined to determine the display size and angle of the object, and according to the alignment relationship between the Mesh space and the two-dimensional image pixel space, the operation result of the corresponding physical rule is displayed in the two-dimensional image pixel space.

15. The method according to claim 14, characterized in that Also includes: Determine the plane where the added items are located; In response to the user's interactive operation of dragging the added object to change its position, the display size and angle of the object in the two-dimensional image pixel space are adjusted according to the change of the three-dimensional position of the added object in the Mesh space during the dragging process, and the object is adsorbed within the plane and moved.

16. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 15 are implemented.

17. An electronic device, characterized in that: include: one or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, perform the steps of the method according to any one of claims 1 to 15.

18. A computer program product comprising a computer program / computer executable instructions, characterized in that When the computer program / computer executable instructions are executed by a processor in an electronic device, the steps of the method according to any one of claims 1 to 15 are implemented.