Spatial three-dimensional layout acquisition method and device, electronic equipment and storage medium

By extracting the pixel coordinates of key points and coordinate information in the three-dimensional spatial model from the target image, the three-dimensional spatial model is optimized to obtain the real-scale spatial three-dimensional layout model, and the problem of relying on high-cost hardware equipment to obtain spatial three-dimensional layout information in the existing technology is solved, and low-cost and high-precision spatial three-dimensional layout information is achieved.

CN120070720APending Publication Date: 2025-05-30GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311606155.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, obtaining spatial three-dimensional layout information requires relying on high-cost hardware equipment such as depth sensors, and the accuracy depends on the accuracy of the hardware equipment, resulting in higher costs.

Method used

By obtaining the target image taken for the preset spatial three-dimensional layout, using the pixel coordinates of key points in the image and coordinate information in the three-dimensional spatial model, the three-dimensional spatial model is optimized to obtain the real-scale spatial three-dimensional layout model, and avoiding relying on hardware devices such as depth sensors.

Benefits of technology

It realizes that the spatial three-dimensional layout information can be obtained without relying on hardware devices such as depth sensors, reducing the cost of obtaining spatial three-dimensional layout information, and the model accuracy does not depend on the accuracy of the hardware device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070720A_ABST
    Figure CN120070720A_ABST
Patent Text Reader

Abstract

The invention discloses a spatial three-dimensional layout acquisition method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring a target image shot for a preset spatial three-dimensional layout, and taking pixel coordinates of key points in the target image as first two-dimensional coordinates; obtaining a second two-dimensional coordinate of the key point in a three-dimensional space model based on the three-dimensional space model matched with the spatial three-dimensional layout and a pose model and camera parameters of a camera used for collecting the target image under the three-dimensional space model; optimizing the three-dimensional space model according to the first two-dimensional coordinates and the second two-dimensional coordinates; and performing scale recovery on the optimized three-dimensional space model to obtain a real-scale space three-dimensional layout model. According to the spatial three-dimensional layout acquisition method provided by the embodiment of the invention, the spatial three-dimensional layout model can truly display the spatial three-dimensional layout information without depending on hardware equipment for providing spatial depth information, and the precision of the spatial three-dimensional layout acquisition method does not depend on the precision of the hardware equipment, so that the three-dimensional layout information acquisition cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image processing, and particularly to a method, apparatus, electronic device and storage medium for obtaining a three-dimensional spatial layout. Background Art

[0002] Obtaining three-dimensional spatial layout information is very important for augmented reality (AR), virtual reality (VR), and some tasks that require the use of three-dimensional spatial layout information, such as depth-based camera autofocus, initial parameter estimation of a microphone audio array based on spatial perception, etc.

[0003] However, the images obtained by a general camera for photographing a three-dimensional space are two-dimensional signals generated by perspective projection of the three-dimensional space, and the three-dimensional layout information of the space cannot be directly obtained therefrom. In related technologies, obtaining three-dimensional spatial layout information often requires using spatial depth information obtained by depth sensors such as TOF and structured light sensors. This way of obtaining three-dimensional spatial layout information highly depends on the accuracy of hardware devices, and the cost of hardware devices with higher accuracy is higher.

[0004] The above statements are only used to provide background technical information related to the present application, and do not necessarily constitute prior art. Summary of the Invention

[0005] Embodiments of the present application provide a method, apparatus, electronic device and storage medium for obtaining a three-dimensional spatial layout, so as to solve the problem of how to obtain three-dimensional spatial layout information without relying on high-cost hardware devices such as depth sensors.

[0006] To provide a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary part is not a general review, nor is it intended to identify key / important constituent elements or delineate the protection scope of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the subsequent detailed description.

[0007] According to one aspect of the embodiments of the present application, a method for obtaining a three-dimensional spatial layout is provided, including:

[0008] Obtaining a target image captured for a preset three-dimensional spatial layout, and obtaining the pixel coordinates of key points in the target image as first two-dimensional coordinates; the key points are preset position points in the preset three-dimensional spatial layout;

[0009] Based on a three-dimensional spatial model matching the preset three-dimensional spatial layout, a pose model of a camera for collecting the target image in the three-dimensional spatial model, and the camera parameters of the camera, obtaining second two-dimensional coordinates of the key points in the three-dimensional spatial model;

[0010] Optimize the three-dimensional space model according to the first two-dimensional coordinates and the second two-dimensional coordinates to obtain an optimized three-dimensional space model;

[0011] Perform scale recovery on the optimized three-dimensional space model to obtain a spatial three-dimensional layout model with the true scale.

[0012] In some embodiments of the present application, the obtaining the target image captured for the preset spatial three-dimensional layout and obtaining the pixel coordinates of the key points in the target image as the first two-dimensional coordinates includes:

[0013] Determine a preset two-dimensional spatial layout estimation model;

[0014] Select position points representing the geometric spatial structure of the spatial three-dimensional layout from the preset spatial three-dimensional layout as key points;

[0015] Input the target image captured for the preset spatial three-dimensional layout into the two-dimensional spatial layout estimation model, and output the pixel coordinates of the key points in the target image as the first two-dimensional coordinates.

[0016] In some embodiments of the present application, the obtaining the second two-dimensional coordinates of the key points in the three-dimensional space model based on the three-dimensional space model matching the preset spatial three-dimensional layout, the pose model of the camera used to capture the target image in the three-dimensional space model, and the camera parameters of the camera includes:

[0017] Construct a three-dimensional space model matching the preset spatial three-dimensional layout;

[0018] Obtain the camera parameters of the camera used to capture the target image;

[0019] Determine the pose model of the camera in the three-dimensional space model;

[0020] Based on the camera parameters, the pose model, and the spatial position information of the key points in the three-dimensional space model, calculate the second two-dimensional coordinates of the key points in the three-dimensional space model.

[0021] In some embodiments of the present application, the constructing a three-dimensional space model matching the preset spatial three-dimensional layout includes:

[0022] Obtain the spatial position information of the key points in the preset spatial three-dimensional layout;

[0023] Determine the geometric spatial structure of the preset spatial three-dimensional layout;

[0024] Construct a three-dimensional space model that matches the preset three-dimensional spatial layout based on the spatial position information of the key points and the geometric spatial structure.

[0025] In some embodiments of the present application, the determining the pose model of the camera in the three-dimensional space model includes:

[0026] Determine the plane in which the camera is located in the three-dimensional space model;

[0027] Construct the pose of the camera relative to the world coordinate system based on the spatial position relationship between the plane where the camera is located and the key points as the pose model;

[0028] The calculating the second two-dimensional coordinates of the key points in the three-dimensional space model based on the camera parameters, the pose model, and the spatial position information of the key points in the three-dimensional space model includes:

[0029] Obtain the spatial position information of the key points in the three-dimensional space model;

[0030] Use the camera parameters and the pose model to perform coordinate transformation on the spatial position information of the key points to obtain the second two-dimensional coordinates of the key points in the three-dimensional space model.

[0031] In some embodiments of the present application, the optimizing the three-dimensional space model according to the first two-dimensional coordinates and the second two-dimensional coordinates to obtain an optimized three-dimensional space model includes:

[0032] Construct a loss function according to the first two-dimensional coordinates and the second two-dimensional coordinates;

[0033] Optimize the three-dimensional space model according to the loss function to obtain an optimized three-dimensional space model.

[0034] In some embodiments of the present application, the performing scale recovery on the optimized three-dimensional space model to obtain a spatial three-dimensional layout model with a true scale includes:

[0035] Determine a preset depth estimation model;

[0036] Input the target image into the depth estimation model, and output the depth map of the target image;

[0037] Based on the depth information corresponding to the key points in the depth map and the camera position information of the key points in the camera coordinate system of the camera, obtain a scale factor;

[0038] Use the scale factor to perform scale recovery on the optimized three-dimensional space model to obtain a spatial three-dimensional layout model with a true scale.

[0039] According to another aspect of the embodiments of the present application, a three-dimensional spatial layout acquisition device is provided, including:

[0040] A first acquisition module, configured to acquire a target image captured for a preset three-dimensional spatial layout, and obtain pixel coordinates of key points in the target image as first two-dimensional coordinates; the key points are preset position points in the preset three-dimensional spatial layout;

[0041] A second acquisition module, configured to obtain second two-dimensional coordinates of the key points in the three-dimensional spatial model based on a three-dimensional spatial model matching the preset three-dimensional spatial layout, a pose model of a camera for acquiring the target image in the three-dimensional spatial model, and camera parameters of the camera;

[0042] An optimization module, configured to optimize the three-dimensional spatial model according to the first two-dimensional coordinates and the second two-dimensional coordinates to obtain an optimized three-dimensional spatial model;

[0043] A scale recovery module, configured to perform scale recovery on the optimized three-dimensional spatial model to obtain a spatial three-dimensional layout model with a true scale.

[0044] According to another aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the three-dimensional spatial layout acquisition method according to any embodiment of the present application.

[0045] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and the computer program is executed by a processor to implement the three-dimensional spatial layout acquisition method according to any embodiment of the present application.

[0046] The technical solution provided by one aspect of the embodiments of the present application may include the following beneficial effects:

[0047] The method for obtaining a three-dimensional spatial layout provided by an embodiment of the present application includes: obtaining a target image captured for a preset three-dimensional spatial layout, and obtaining the pixel coordinates of key points in the target image as first two-dimensional coordinates; the key points are preset position points in the preset three-dimensional spatial layout; based on a three-dimensional spatial model matching the preset three-dimensional spatial layout, the pose model of the camera used to capture the target image in the three-dimensional spatial model, and the camera parameters of the camera, obtaining the second two-dimensional coordinates of the key points in the three-dimensional spatial model; optimizing the three-dimensional spatial model according to the first two-dimensional coordinates and the second two-dimensional coordinates to obtain an optimized three-dimensional spatial model; performing scale recovery on the optimized three-dimensional spatial model to obtain a spatial three-dimensional layout model with a real scale. This spatial three-dimensional layout model can truly display the spatial three-dimensional layout information corresponding to the preset three-dimensional spatial layout, without relying on hardware devices such as depth sensors to provide spatial depth information. Moreover, this spatial three-dimensional layout model is optimized based on the first two-dimensional coordinates and the second two-dimensional coordinates of the preset position points (i.e., key points) in the three-dimensional spatial layout, and the accuracy of its model does not need to rely on the accuracy of hardware devices, thereby reducing the cost of obtaining spatial three-dimensional layout information. Furthermore, since the target image is captured for the preset three-dimensional spatial layout, and the key points in the target image correspond to the preset position points in the preset three-dimensional spatial layout, in the process of applying this method, if different spatial three-dimensional layout models are to be obtained, the target images under different spatial three-dimensional layouts can be collected and the key point information in these target images can be used to optimize the three-dimensional spatial model matching this spatial three-dimensional layout, which is convenient, fast, and has strong universality.

[0048] The above description is only an overview of the technical solutions of the embodiments of the present application. In order to be able to understand the technical means of the embodiments of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the embodiments of the present application more obvious and understandable, the following specifically gives the specific implementation manners of the present application. Description of the Drawings

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0050] Figure 1 Shows the flowchart of the method for obtaining a three-dimensional spatial layout according to an embodiment of the present application.

[0051] Figure 2 Shows the target image in an embodiment of the present application.

[0052] Figure 3 Shows a target image in an embodiment of the present application with the connection lines between key points marked.

[0053] Figure 4 Shows a schematic diagram of a three-dimensional space cuboid model in the world coordinate system in an embodiment of the present application.

[0054] Figure 5 Shows a schematic diagram of converting a point in the world coordinate system to the camera coordinate system in an embodiment of the present application.

[0055] Figure 6 Shows a schematic diagram of a three-dimensional space model after scale recovery in a specific example of the present application.

[0056] Figure 7 Shows a structural block diagram of a spatial three-dimensional layout acquisition device in an embodiment of the present application.

[0057] Figure 8 Shows a structural block diagram of an electronic device in an embodiment of the present application.

[0058] Figure 9 Shows a schematic diagram of a computer-readable storage medium in an embodiment of the present application. Detailed implementation manners

[0059] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0060] Those skilled in the art can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the technical field to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as here.

[0061] In the related art, it is necessary to utilize the spatial depth information obtained by hardware devices such as depth sensors for providing spatial depth information or the RGB color information obtained by a color camera to obtain the spatial three-dimensional layout information. The dependence on hardware devices is relatively high, and the accuracy of the obtained spatial three-dimensional layout information also depends on the accuracy of the hardware devices for providing spatial depth information. Since the technical solution for obtaining the spatial three-dimensional layout information in the related art has a relatively high dependence on hardware devices, and the cost of hardware devices such as depth sensors and color cameras used is relatively high, the cost of obtaining the spatial three-dimensional layout information is relatively high.

[0062] In view of the technical problems existing in the related art, an embodiment of the present application provides a method for obtaining a spatial three-dimensional layout. A target image captured for a preset spatial three-dimensional layout is obtained, and the pixel coordinates of key points in the target image are used as the first two-dimensional coordinates; the key points are preset position points in the preset spatial three-dimensional layout; based on a three-dimensional space model matching the preset spatial three-dimensional layout, the pose model of the camera used for capturing the target image in the three-dimensional space model, and the camera parameters of the camera, the second two-dimensional coordinates of the key points in the three-dimensional space model are obtained; the three-dimensional space model is optimized according to the first two-dimensional coordinates and the second two-dimensional coordinates to obtain an optimized three-dimensional space model; the optimized three-dimensional space model is subjected to scale recovery to obtain a spatial three-dimensional layout model with a real scale. This spatial three-dimensional layout model can truly display the spatial three-dimensional layout information corresponding to the preset spatial three-dimensional layout, without relying on hardware devices such as depth sensors for providing spatial depth information. Moreover, this spatial three-dimensional layout model is optimized based on the first two-dimensional coordinates and the second two-dimensional coordinates of the preset position points (i.e., key points) in the spatial three-dimensional layout, and the accuracy of its model does not need to depend on the accuracy of hardware devices, thereby reducing the cost of obtaining the spatial three-dimensional layout information. Furthermore, since the target image is captured for a preset spatial three-dimensional layout, and the key points in the target image correspond to the preset position points in the preset spatial three-dimensional layout, in the process of applying this method, if different spatial three-dimensional layout models are to be obtained, the target images under different spatial three-dimensional layouts can be collected and the key point information in these target images can be used to optimize the three-dimensional space model matching this spatial three-dimensional layout, which is convenient, fast, and has strong universality.

[0063] Next, a method, device, electronic device, and storage medium for obtaining a spatial three-dimensional layout proposed according to an embodiment of the present application will be described with reference to the accompanying drawings.

[0064] Refer to Figure 1 As shown, an embodiment of the present application provides a method for obtaining a spatial three-dimensional layout, which may include steps S10 to S40:

[0065] S10. Obtain a target image captured for a preset three-dimensional spatial layout, and obtain the pixel coordinates of key points in the target image as the first two-dimensional coordinates.

[0066] In this embodiment, the target image may be an image collected by a camera for a preset three-dimensional spatial layout. The preset three-dimensional spatial layout may be a scene pre-selected according to task requirements, such as indoor scenes like meeting rooms, offices, exhibition rooms, etc. The embodiments of the present application do not make specific limitations thereto.

[0067] Among them, the key points are preset position points in the preset three-dimensional spatial layout. For example, any position point belonging to the three-dimensional spatial layout can be pre-selected as a key point according to task requirements in the preset three-dimensional spatial layout. The selection of key points can refer to the structural characteristics of the geometric spatial structure of the three-dimensional spatial layout (such as length, width, height, vertices, intersections, perpendicular points, etc.). The geometric spatial structure may be a three-dimensional structure composed of multiple planes, such as a cuboid and other three-dimensional structures. It can be understood that there is more than one key point in the three-dimensional spatial layout in the embodiments of the present application, and there may be multiple.

[0068] In a specific implementation manner, S10 may include the following steps:

[0069] S101. Determine a preset two-dimensional spatial layout estimation model.

[0070] The two-dimensional spatial layout estimation model may be, for example, a RoomNet model, a MonoLayout model, etc.

[0071] S102. Select position points representing the geometric spatial structure of the three-dimensional spatial layout from the preset three-dimensional spatial layout as key points.

[0072] In S102, the preset three-dimensional spatial layout is a spatial scene pre-selected according to task requirements. This spatial scene generally has a regular geometric spatial structure, such as a cuboid, a cube, etc. Therefore, the preset three-dimensional spatial layout in this embodiment also has a regular geometric spatial structure, and these geometric spatial structures have multiple vertices. These vertices are position points that can represent the main geometric features of the geometric spatial structure, and the vertices can contain rich position information in the three-dimensional spatial layout. In one example, the vertices of the spatial structure where the preset three-dimensional spatial layout is located can be selected from the preset three-dimensional spatial layout as position points representing the geometric spatial structure of the three-dimensional spatial layout as key points.

[0073] For example, the target image may be an indoor image of a room. Usually, the overall outline of the room is in the shape of a cuboid. Therefore, the key points generally include all the vertices of the cuboid. The cuboid can be determined through all the vertices of the cuboid. Refer to Figure 2As shown, in one example, the target image is an image of a conference room scene captured by a camera, in which the main planes for spatial layout estimation in the image are divided. Each plane is the plane where each wall in the conference room is located, and each vertex in the spatial structure of the conference room can be used as a key point. As Figure 3 shown, Figure 3 In [the figure], the connections between the key points are further marked. Each connection is the intersection line between the planes where two adjacent walls are located.

[0074] In another example, key points for a preset three-dimensional spatial layout can be selected through neural network / deep learning models such as object detection models and feature recognition models commonly used in the field of images. When using the model to select key points, it is necessary to ensure that the key points meet the prerequisite conditions of geometric spatial structure features that can represent the three-dimensional spatial layout, so that the selected key points have rich geometric structure features, which is beneficial to constructing a three-dimensional space model that matches the preset three-dimensional spatial layout next.

[0075] S103: Input the target image captured for the preset three-dimensional spatial layout into the two-dimensional spatial layout estimation model, and output the pixel coordinates of the key points in the target image as the first two-dimensional coordinates.

[0076] In S103, the key points can be pre-identified in the target image, so that when the target image is input into the two-dimensional spatial layout estimation model, the pixel coordinates of the key points in the target image can be output conveniently and quickly as the first two-dimensional coordinates.

[0077] In a specific example, referring to Figure 4 shown, Figure 4 is a three-dimensional space model that matches the preset three-dimensional spatial layout. This three-dimensional space model is a three-dimensional space cuboid model constructed based on the world coordinate system. In this example, by processing the target image I through a preset two-dimensional spatial layout estimation model (such as the RoomNet model or the MonoLayout model), the two-dimensional spatial layout coordinates (i.e., the first two-dimensional coordinates) pi(xi, yi) of the key points in the pixel coordinate system can be obtained, where i ∈ [0, 7]. Figure 4 The key points in [the figure] include the vertices of the cuboid, namely P0, P1, P2, P3, P4, P5, P6, and P7.

[0078] In some examples, the region of interest (ROI) identification algorithm in the field of images can be used to identify the key points of the target image, or a pre-trained image recognition model or semantic segmentation model can be used to identify the key points of the target image. The embodiments of the present application do not make specific limitations on this.

[0079] S20. Based on the three-dimensional space model that matches the preset three-dimensional space layout, the pose model of the camera for collecting the target image in the three-dimensional space model, and the camera parameters of the camera, obtain the second two-dimensional coordinates of the key points in the three-dimensional space model.

[0080] Among them, the three-dimensional space model that matches the preset three-dimensional space layout is the three-dimensional space model that matches the target image.

[0081] In one embodiment, S20 may include the following specific steps:

[0082] S201. Construct a three-dimensional space model that matches the preset three-dimensional space layout.

[0083] In a specific implementation manner, the spatial position information of the key points in the preset three-dimensional space layout may be obtained; the geometric space structure of the preset three-dimensional space layout is determined; and a three-dimensional space model that matches the preset three-dimensional space layout is constructed based on the spatial position information of the key points and the geometric space structure.

[0084] The spatial position information of the key points in the preset three-dimensional space layout is the three-dimensional coordinates of the key points in the world coordinate system. For example, in the Figure 4 corresponding three-dimensional space layout, the key points are the vertices of a cuboid, and the spatial position information of the key points is the three-dimensional coordinates of each vertex in the world coordinate system.

[0085] The geometric space structure of the three-dimensional space layout may be, for example, a three-dimensional structure such as a cuboid structure or a cube structure. In indoor scenes such as meeting rooms, offices, and exhibition rooms, the geometric space structure of the three-dimensional space layout is mainly a cuboid structure. For example, in the Figure 4 corresponding geometric space structure is a cuboid structure.

[0086] A neural network model or tools such as CAD software may be used to construct a three-dimensional space model that matches the preset three-dimensional space layout based on the three-dimensional coordinates of the key points in the world coordinate system and the geometric space structure. In the example shown in Figure 4 the three-dimensional space model that matches the preset three-dimensional space layout is a cuboid model formed by the vertices P1, P2, P3, P4, P5, P6, P7, and P8.

[0087] S202. Obtain the camera parameters of the camera for collecting the target image.

[0088] The camera internal parameters describe some properties inherent in the camera itself, such as focal length, pixel pitch, etc. Specifically, the camera parameters of the camera for collecting the target image can be represented by the internal parameter matrix of the camera.

[0089] The internal parameter matrix K of the camera can be expressed as K =

[0090]

[0091] Among them, \(f_x\) is the component length of the camera's focal length in the x-axis direction in the pixel coordinate system, \(f_y\) is the component length of the camera's focal length in the y-axis direction in the pixel coordinate system, \(c_x\) represents the principal point in the x-axis direction, and \(c_y\) represents the principal point in the y-axis direction.

[0092] S203. Determine the pose model of the camera in the three-dimensional space model.

[0093] The pose model of the camera in the three-dimensional space model can be represented as an external parameter matrix. The pose model is used to realize the transformation, rotation, and translation from the world coordinate system to the camera coordinate system. In a specific implementation, the plane where the camera is located in the three-dimensional space model can be determined first, and then the pose of the camera relative to the world coordinate system can be constructed based on the spatial position relationship between the plane where the camera is located and the key points as the pose model. It can be understood that there is more than one key point in the embodiments of the present application, and there can be multiple.

[0094] As Figure 4 shown in the example, the plane where the camera C is located in the three-dimensional space model is the plane formed by the vertices P4, P5, P6, and P7. The pose of the camera relative to the world coordinate system is constructed based on the spatial position relationship between the plane where the camera C is located and the key points as the pose model.

[0095] Specifically, in one form of expression, the pose model can be represented as the external parameter matrix of the camera. The external parameter matrix can be represented as a matrix T = [R|t] composed of a rotation matrix R and a translation matrix t (the external parameter matrix can also be called a pose transformation matrix). Among them, R represents the rotation matrix, \(R\in(3\times3)\), t represents the translation matrix, \(t=(x,y,z)\), \(t\in(3\times1)\). The initial values of R and the initial value of t can be obtained by random initialization or set according to actual application experience.

[0096] The rotation matrix R contains 4 rotation variables, and the translation matrix t contains 3 position variables. The 4 rotation variables can be represented by the quaternion representation method, so as to obtain the rotation matrix R represented by the quaternion. Among them, according to the Rodrigue's formula, the rotation variable q represented by the quaternion and the rotation matrix R can be mutually converted, that is, \(q = Rodrigue(R)\), \(R = Rodrigue\_inv(q)\).

[0097] For example, if the rotation variable q represented by the quaternion is \(q=(x,y,z,w)\), then the rotation matrix R can be represented as

[0098]

[0099] That is to say, the external parameter matrix contains 4 rotation variables and 3 position variables, a total of 7 variables. These 7 variables are variables to be optimized, and subsequently, these 7 variables can be optimized according to a preset loss function.

[0100] Figure 5 It is a schematic diagram for converting points in the world coordinate system to the camera coordinate system through the external parameter matrix. Refer to Figure 5 As shown, specifically, the point Pi(X, Y, Z) in the world coordinate system is converted to the camera coordinate system, and the obtained coordinates in the camera coordinate system are Pi’(X’, Y’, Z’). The conversion formula is

[0101]

[0102] Among them, R represents the rotation matrix, R ∈ (3×3), t represents the translation matrix, t = (x, y, z), and t ∈ (3×1).

[0103] S204. Calculate the second two-dimensional coordinates of the key point in the three-dimensional space model based on the camera parameters, the pose model, and the spatial position information of the key point in the three-dimensional space model.

[0104] In one implementation, the spatial position information of the key point in the three-dimensional space model can be obtained;

[0105] The spatial position information of the key point is subjected to coordinate transformation using the camera parameters and the pose model to obtain the second two-dimensional coordinates of the key point in the three-dimensional space model.

[0106] Specifically, in an example, refer to Figure 6 , the spatial position information of the key point in the three-dimensional space model is the three-dimensional coordinates Pi of the key point (the vertex in the cuboid model), i ∈ [0, 7]. The spatial position information of the key point is subjected to coordinate transformation using the internal parameter matrix K of the camera and the external parameter matrix T represented by the pose model to obtain the second two-dimensional coordinates, that is, the second two-dimensional coordinates pi’ = K * T * Pi, i ∈ [0, 7].

[0107] Projecting the spatial position information Pi(X, Y, Z) of the key point using the internal parameter matrix K and the external parameter matrix T can obtain the corresponding second two-dimensional coordinates pi’(u, v).

[0108] Pi(X, Y, Z) and pi’(u, v) satisfy the following relationship:

[0109]

[0110] Among them, K is the internal parameter matrix, and [R|t] is the external parameter matrix.

[0111] S30. Optimize the three-dimensional space model according to the first two-dimensional coordinates and the second two-dimensional coordinates to obtain an optimized three-dimensional space model.

[0112] Optimizing the three-dimensional space model according to the first two-dimensional coordinates and the second two-dimensional coordinates to obtain an optimized three-dimensional space model may include: constructing a loss function according to the first two-dimensional coordinates and the second two-dimensional coordinates; optimizing the three-dimensional space model according to the loss function to obtain an optimized three-dimensional space model.

[0113] Continuing with the above specific example, the first two-dimensional coordinate is pi', and the second two-dimensional coordinate is pi; constructing a loss function according to the first two-dimensional coordinates and the second two-dimensional coordinates may include: using pi' and pi to construct a loss term loss_pixel through the L2 loss function, constructing the slope of the vanishing line as a loss term loss_slope using L2 as well, and constructing a loss function as loss = loss_pixel + loss_slope using the loss term loss_pixel and the loss term loss_slope. The vanishing line is the connection line of key points, which is the edges of the cuboid in this example.

[0114] Optimizing the three-dimensional space model according to the loss function to obtain an optimized three-dimensional space model may include: using an optimization algorithm such as the simulated annealing algorithm or the differential evolution algorithm to optimize the length, width, and height of the three-dimensional space cuboid until the value of the loss function reaches a preset value, obtaining optimized parameters. During the optimization process, the loss function is minimized. At this time, the length, width, and height of the three-dimensional space cuboid and the camera pose model reach the required optimized values, and at the same time, the length, width, and height of the three-dimensional space cuboid and the camera pose model that reach this optimized value are determined.

[0115] S40. Perform scale recovery on the optimized three-dimensional space model to obtain a spatial three-dimensional layout model with the true scale.

[0116] Taking the three-dimensional space model as shown in Figure 4 as an example, after the steps of S10 to S30, the 8 vertex coordinates Pi (i ∈ [0, 7]) of the three-dimensional space cuboid, the length, width, and height of the three-dimensional space cuboid, and the camera pose model are all known, and an optimized three-dimensional space model is obtained. However, at this time, since the scale parameters of the three-dimensional space cuboid sought are all constructed according to normalized parameters and the depth information is all 1, which does not represent the actual depth distance value, that is, it is only a relative scale at this time. Therefore, it is necessary to perform scale recovery on the optimized three-dimensional space model to obtain the true scale (for example, restoring the cuboid model in centimeters to the actual cuboid model in meters).

[0117] In one embodiment, S40 may include the following specific steps:

[0118] S401. Determine a preset depth estimation model.

[0119] Specifically, the depth estimation model can be, for example, a monocular depth estimation model.

[0120] S402. Input the target image into the depth estimation model, and output the depth map of the target image.

[0121] S403. Based on the depth information corresponding to the key points in the depth map and the camera position information of the key points in the camera coordinate system of the camera, obtain a scale factor.

[0122] The depth information corresponding to the key points can be found in the obtained depth map, and the coordinates of the key points in the camera coordinate system of the camera are obtained as the camera position information. The scale factor is calculated according to the preset projection relationship between the camera position information and the depth information of the key points.

[0123] S404. Use the scale factor to perform scale recovery on the optimized three-dimensional space model to obtain a spatial three-dimensional layout model with the true scale.

[0124] In a specific example, monocular depth maps can be used for scale recovery. In this example, the size of the depth map D is the same as that of the target image I. The size of the target image I is width W and height H, and the size of the depth map D is also width W and height H. The depth map D is obtained by inputting the target image I into the monocular depth estimation model for processing. For example, the depth D[pi] corresponding to the key point pi can be found in the depth map D. According to the principle of rigid body motion, the corresponding coordinates of the key point Pi in the world coordinate system in the camera coordinate system are Pi_c = T * Pi, where the coordinates of Pi_c are expressed as (Xi_c, Yi_c, Zi_c). Then D[pi] = Xi_c * S, where S is the scale factor, and S = D[pi] / Xi_c. Use the scale factor S to perform scale recovery on the optimized three-dimensional space model to obtain a scaled 3D layout cube with length, width, and height of (Length, Width, Height) = (b_l * S, b_w * S, 1 * S).

[0125] To deepen the understanding of the above embodiments, specific examples will be explained below.

[0126] In a specific example, a target image of a meeting room is obtained. The geometric space structure of the meeting room is a cuboid structure, and 8 vertices of the cuboid structure are selected as key points. The target image of the meeting room is input into a preset two-dimensional space layout estimation model, and the pixel coordinates of the key points in the target image are output as the first two-dimensional coordinates. A three-dimensional space model matching the meeting room is constructed, which is a cuboid model.

[0127] Set the length, width, and height of the three-dimensional cuboid model to \(b_l\), \(b_w\), and \(b_h\) respectively. Refer to Figure 4 As shown, taking the position of the lower left corner vertex \(P0\) in the cuboid model as the coordinate origin \((0, 0, 0)\), a world coordinate system is established. The \(x\)-axis of this world coordinate system is parallel to the side where \(P0\) and \(P4\) are located, the \(z\)-axis is parallel to the side where \(P0\) and \(P3\) are located, and the \(y\)-axis is parallel to the side where \(P0\) and \(P1\) are located. The key points include the 8 vertices of the cuboid model. The coordinates of these 8 vertices in the world coordinate system can be expressed as \(P_i (i\in[0,7])\). \(P0\) to \(P7\) can be respectively expressed as \(P0(0, 0, 0)\), \(P1(0, b_w, 0)\), \(P2(0, b_w, b_h)\), \(P3(0, 0, b_h)\), \(P4(b_l, 0, 0)\), \(P5(b_l, b_w, 0)\), \(P6(b_l, b_w, b_h)\), \(P7(b_l, 0, b_h)\); among them, \(b_l\), \(b_w\), and \(b_h\) are 3 variables to be optimized. Figure 4 The midpoint \(C\) represents the position point of the camera. The position point where the camera \(C\) is located is on the plane where \(P4\), \(P5\), \(P6\), and \(P7\) are located.

[0128] Since the target image is taken for a preset three-dimensional spatial layout, and the key points in the target image correspond to the preset position points in the preset three-dimensional spatial layout. In the process of applying this method, if different three-dimensional spatial layout models are to be obtained, different target images under different three-dimensional spatial layouts can be collected and the key point information in these target images can be used to optimize the three-dimensional spatial model matching the three-dimensional spatial layout, which is convenient, fast, and has strong universality.

[0129] Furthermore, obtain the camera parameters of the camera used to collect the target image, that is, the internal parameter matrix \(K\); and determine the pose model of the camera under this cuboid model. This pose model can be expressed as an external parameter matrix. The external parameter matrix can be expressed as a matrix \(T = [R|t]\) composed of a rotation matrix \(R\) and a translation matrix \(t\). Among them, \(R\) represents the rotation matrix, \(R\in(3\times3)\), \(t\) represents the translation matrix, \(t=(x,y,z)\), \(t\in(3\times1)\). The initial values of \(R\) and the initial values of \(t\) can be obtained through random initialization or set according to actual application experience. Using the internal parameter matrix \(K\) and the external parameter matrix \(T\) to perform coordinate transformation on \(P_i(X,Y,Z)\) can obtain the second two-dimensional coordinates \(p_i'(u, v)\), where \(p_i' = K*T*P_i\), \(i\in[0,7]\).

[0130] To facilitate the solution of the variables to be optimized in the pose model of the camera, some prior information can be introduced. For example, when the three-dimensional space model is a three-dimensional rectangular parallelepiped model, the installation position of the camera is default set on the opposite plane of the plane where the world coordinate system is located, that is, this opposite plane is perpendicular to the x-axis of the world coordinate system and parallel to the yoz plane formed by the origin, the y-axis, and the z-axis. And the origin of the world coordinate system is set at a certain vertex of the rectangular parallelepiped model, and the x, y, and z axes of this world coordinate system are flush with the three edges that diverge from this vertex (the origin) of the rectangular parallelepiped model and are perpendicular to each other in pairs. Refer to Figure 4 As shown, the position of the lower left corner vertex P0 in the rectangular parallelepiped model is the coordinate origin (0, 0, 0). A world coordinate system is established with this coordinate origin. The x-axis of this world coordinate system is parallel to the edge where P0 and P4 are located, the z-axis is parallel to the edge where P0 and P3 are located, and the y-axis is parallel to the edge where P0 and P1 are located. By default, the camera is located on the plane formed by the four vertices P4, P5, P6, and P7. Based on Figure 4 , it can be confirmed that there is a relatively definite translational position relationship between the position of the camera and the world coordinate system, that is, there is a relatively definite translational position relationship between the camera coordinate system and the world coordinate system (represented by the translation matrix t). Therefore, when solving each variable in the pose model of the camera, the coordinates Pi (i ∈ [0, 7]) of the 8 vertices can be successively substituted for solution. For the specific values of the length b_l, width b_w, and height b_h involved in the vertex coordinates Pi substituted for solution, the number of variables to be optimized can be reduced by means of assignment. For example, b_h can be set to 1 or 2 or 3, etc. The specific assigned value can be selected according to actual needs.

[0131] After the above steps, the 8 vertex coordinates Pi (i ∈ [0, 7]) of the cuboid in three-dimensional space, the length, width, and height of the cuboid in three-dimensional space, and the pose model of the camera are all known, and an optimized three-dimensional space model is obtained. However, at this time, since the scale parameter of the cuboid in the three-dimensional space to be obtained is a normalized parameter with reference to the depth information b_h = 1, it does not represent the actual distance value, that is, it is only a relative scale at this time, and scale recovery is required to obtain the true scale of the cuboid (for example, a cuboid in meters). Find the depth D[pi] corresponding to the key point pi in the depth map D. The depth map D is obtained by inputting the target image I into a monocular depth estimation model for processing. According to the principle of rigid body motion, the corresponding coordinates of the key point Pi in the world coordinate system in the camera coordinate system are Pi_c = T * Pi, where the coordinates of Pi_c are expressed as (Xi_c, Yi_c, Zi_c), then D[pi] = Xi_c * S, where S is the scale factor, and S = D[pi] / Xi_c. Use the scale factor S to perform scale recovery on the optimized three-dimensional space model, and obtain the length, width, and height of the 3D layout cube with scale as (Length, Width, Height) = (b_l * S, b_w * S, S). The three-dimensional space model after scale recovery is as Figure 6 shown.

[0132] The method for obtaining the three-dimensional spatial layout in this example applies the principles of visual geometry, two-dimensional layout estimation, and monocular depth estimation. Only relying on the images obtained by the camera can obtain the 3D layout of the indoor space and the pose of the camera, and can truly display the three-dimensional spatial layout information corresponding to the preset three-dimensional spatial layout. It does not rely on hardware devices such as depth sensors to provide spatial depth information, and the three-dimensional spatial layout model is optimized based on the first two-dimensional coordinates and the second two-dimensional coordinates of the key points in the three-dimensional spatial layout. The accuracy of its model does not depend on the accuracy of the hardware device, thus reducing the cost of obtaining the three-dimensional spatial layout information.

[0133] Refer to Figure 7 shown. Another embodiment of the present application provides a device for obtaining a three-dimensional spatial layout, which may include:

[0134] A first acquisition module, configured to acquire a target image taken for a preset three-dimensional spatial layout, and obtain the pixel coordinates of the key points in the target image as the first two-dimensional coordinates; the key points are preset position points in the preset three-dimensional spatial layout;

[0135] A second acquisition module, configured to obtain the second two-dimensional coordinates of the key points in the three-dimensional spatial model based on a three-dimensional spatial model matching the preset three-dimensional spatial layout, the pose model of the camera for acquiring the target image in the three-dimensional spatial model, and the camera parameters of the camera;

[0136] An optimization module, configured to optimize the three-dimensional space model according to the first two-dimensional coordinates and the second two-dimensional coordinates, so as to obtain an optimized three-dimensional space model;

[0137] A scale recovery module, configured to perform scale recovery on the optimized three-dimensional space model to obtain a spatial three-dimensional layout model with a true scale.

[0138] In some embodiments, the first acquisition module may include:

[0139] A determination unit, configured to determine a preset two-dimensional space layout estimation model;

[0140] A selection unit, configured to select position points representing the geometric space structure of the spatial three-dimensional layout from the preset spatial three-dimensional layout as key points;

[0141] A two-dimensional space layout estimation unit, configured to input a target image captured for the preset spatial three-dimensional layout into the two-dimensional space layout estimation model, and output the pixel coordinates of the key points in the target image as the first two-dimensional coordinates.

[0142] In some embodiments, the second acquisition module may include:

[0143] A construction unit, configured to construct a three-dimensional space model matching the preset spatial three-dimensional layout;

[0144] A camera parameter acquisition unit, configured to acquire the camera parameters of the camera used to capture the target image;

[0145] A pose model determination unit, configured to determine the pose model of the camera under the three-dimensional space model;

[0146] A calculation unit, configured to calculate the second two-dimensional coordinates of the key points in the three-dimensional space model based on the camera parameters, the pose model, and the spatial position information of the key points in the three-dimensional space model.

[0147] In some embodiments, the construction unit may be further specifically configured to: acquire the spatial position information of the key points in the preset spatial three-dimensional layout; determine the geometric space structure of the preset spatial three-dimensional layout; and construct a three-dimensional space model matching the preset spatial three-dimensional layout based on the spatial position information of the key points and the geometric space structure.

[0148] The pose model determination unit may be further specifically configured to: determine the plane in which the camera is located in the three-dimensional space model; and construct the pose of the camera relative to the world coordinate system as the pose model based on the spatial position relationship between the plane in which the camera is located and the key points.

[0149] The calculation unit can be further used to: obtain the spatial position information of the key points in the three-dimensional space model; perform coordinate transformation on the spatial position information of the key points by using the camera parameters and the pose model to obtain the second two-dimensional coordinates of the key points in the three-dimensional space model.

[0150] In some embodiments, the optimization module can be further specifically used to: construct a loss function according to the first two-dimensional coordinates and the second two-dimensional coordinates; optimize the three-dimensional space model according to the loss function to obtain an optimized three-dimensional space model.

[0151] In some embodiments, the scale recovery module can be further specifically used to: determine a preset depth estimation model; input the target image into the depth estimation model to output the depth map of the target image; obtain a scale factor based on the depth information corresponding to the key points in the depth map and the camera position information of the key points in the camera coordinate system of the camera; perform scale recovery on the optimized three-dimensional space model by using the scale factor to obtain a spatial three-dimensional layout model with a real scale.

[0152] The spatial three-dimensional layout acquisition device according to the embodiments of the present application, the spatial three-dimensional layout model thereof can truly display the spatial three-dimensional layout information corresponding to the preset spatial three-dimensional layout, without relying on hardware devices such as depth sensors to provide spatial depth information, and the spatial three-dimensional layout model is optimized based on the first two-dimensional coordinates and the second two-dimensional coordinates of the preset position points (i.e., key points) in the spatial three-dimensional layout, and the accuracy of the model does not need to depend on the accuracy of the hardware device, thereby reducing the cost of obtaining the spatial three-dimensional layout information.

[0153] Another embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the spatial three-dimensional layout acquisition method according to any one of the above embodiments.

[0154] Reference Figure 8 As shown, the electronic device 10 may include: a processor 100, a memory 101, a bus 102, and a communication interface 103. The processor 100, the communication interface 103, and the memory 101 are connected through the bus 102; a computer program executable on the processor 100 is stored in the memory 101, and when the processor 100 runs the computer program, it executes the method provided in any one of the foregoing embodiments of the present application.

[0155] Among them, the memory 101 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the device network element and at least one other network element is realized through at least one communication interface 103 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0156] The bus 102 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 101 is used to store a program. After receiving an execution instruction, the processor 100 executes the program. Any implementation manner of the method disclosed in any embodiment of the present application can be applied to the processor 100 or implemented by the processor 100.

[0157] The processor 100 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 100 or by instructions in software form. The above-mentioned processor 100 can be a general-purpose processor, which can include a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor, or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or by a combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and combines its hardware to complete the steps of the above method.

[0158] The electronic device provided by the embodiments of the present application and the method provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run, or implemented by it.

[0159] Another embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the method for obtaining a three-dimensional spatial layout in any of the above embodiments. Refer to Figure 9 As shown, the computer-readable storage medium shown is an optical disc 20, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will execute the method provided in any of the foregoing embodiments.

[0160] It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated here one by one.

[0161] The computer-readable storage medium provided in the above embodiments of the present application and the method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.

[0162] It should be noted that:

[0163] The term "module" is not intended to be limited to a specific physical form. Depending on the specific application, a module can be implemented as hardware, firmware, software, and / or a combination thereof. In addition, different modules can share common components or even be implemented by the same components. There may or may not be clear boundaries between different modules.

[0164] The algorithms and displays provided herein are not inherently related to any particular computer, virtual device, or other equipment. Various general-purpose devices can also be used in conjunction with the examples based herein. Based on the above description, the structure required to construct such a device is obvious. In addition, the present application is not directed to any particular programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the description of a specific language above is to disclose the best implementation mode of the present application.

[0165] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially according to the indication of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this document, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0166] The above embodiments only express the implementation manners of the present application, and the description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for obtaining a three-dimensional spatial layout, characterized in that, it includes: Obtain a target image taken for a preset three-dimensional spatial layout, and obtain the pixel coordinates of key points in the target image as the first two-dimensional coordinates; The key points are preset position points in the preset three-dimensional spatial layout; Based on a three-dimensional spatial model matching the preset three-dimensional spatial layout, the pose model of the camera used to collect the target image in the three-dimensional spatial model, and the camera parameters of the camera, obtain the second two-dimensional coordinates of the key points in the three-dimensional spatial model; Optimize the three-dimensional spatial model according to the first two-dimensional coordinates and the second two-dimensional coordinates to obtain an optimized three-dimensional spatial model; Perform scale recovery on the optimized three-dimensional spatial model to obtain a spatial three-dimensional layout model with real scale.

2. The method according to claim 1, characterized in that, The obtaining of the target image taken for the preset three-dimensional spatial layout and obtaining the pixel coordinates of the key points in the target image as the first two-dimensional coordinates includes: Determine a preset two-dimensional spatial layout estimation model; Select position points representing the geometric spatial structure of the three-dimensional spatial layout from the preset three-dimensional spatial layout as key points; Input the target image taken for the preset three-dimensional spatial layout into the two-dimensional spatial layout estimation model, and output the pixel coordinates of the key points in the target image as the first two-dimensional coordinates.

3. The method according to claim 1, characterized in that, The obtaining of the second two-dimensional coordinates of the key points in the three-dimensional spatial model based on a three-dimensional spatial model matching the preset three-dimensional spatial layout, the pose model of the camera used to collect the target image in the three-dimensional spatial model, and the camera parameters of the camera includes: Construct a three-dimensional spatial model matching the preset three-dimensional spatial layout; Obtain the camera parameters of the camera used to collect the target image; Determine the pose model of the camera in the three-dimensional spatial model; Based on the camera parameters, the pose model, and the spatial position information of the key points in the three-dimensional spatial model, calculate the second two-dimensional coordinates of the key points in the three-dimensional spatial model.

4. The method according to claim 3, characterized in that, The constructing of the three-dimensional spatial model matching the preset three-dimensional spatial layout includes: Obtain the spatial position information of the key points in the preset three-dimensional spatial layout; Determine the geometric spatial structure of the preset three-dimensional spatial layout; Based on the spatial position information of the key points and the geometric spatial structure, construct a three-dimensional spatial model matching the preset three-dimensional spatial layout.

5. The method according to claim 3, characterized in that, The determining of the pose model of the camera in the three-dimensional spatial model includes: Determine the plane where the camera is located in the three-dimensional spatial model; Based on the spatial position relationship between the plane where the camera is located and the key points, construct the pose of the camera relative to the world coordinate system as the pose model; Calculating the second two-dimensional coordinates of the key point in the three-dimensional space model based on the camera parameters, the pose model, and the spatial position information of the key point in the three-dimensional space model includes: Obtaining the spatial position information of the key point in the three-dimensional space model; Performing coordinate transformation on the spatial position information of the key point by using the camera parameters and the pose model to obtain the second two-dimensional coordinates of the key point in the three-dimensional space model.

6. The method according to any one of claims 1-5, wherein, Optimizing the three-dimensional space model according to the first two-dimensional coordinates and the second two-dimensional coordinates to obtain an optimized three-dimensional space model includes: Constructing a loss function according to the first two-dimensional coordinates and the second two-dimensional coordinates; Optimizing the three-dimensional space model according to the loss function to obtain an optimized three-dimensional space model.

7. The method according to any one of claims 1-5, wherein, Performing scale recovery on the optimized three-dimensional space model to obtain a spatial three-dimensional layout model with a true scale includes: Determining a preset depth estimation model; Inputting the target image into the depth estimation model and outputting a depth map of the target image; Obtaining a scale factor based on the depth information corresponding to the key point in the depth map and the camera position information of the key point in the camera coordinate system of the camera; Performing scale recovery on the optimized three-dimensional space model by using the scale factor to obtain a spatial three-dimensional layout model with a true scale.

8. A spatial three-dimensional layout acquisition device, wherein, includes: A first acquisition module, configured to acquire a target image taken for a preset spatial three-dimensional layout, and obtain the pixel coordinates of the key points in the target image as the first two-dimensional coordinates; the key points are preset position points in the preset spatial three-dimensional layout; A second acquisition module, configured to obtain the second two-dimensional coordinates of the key points in the three-dimensional space model based on a three-dimensional space model matching the preset spatial three-dimensional layout, a pose model of the camera used to acquire the target image in the three-dimensional space model, and the camera parameters of the camera; An optimization module, configured to optimize the three-dimensional space model according to the first two-dimensional coordinates and the second two-dimensional coordinates to obtain an optimized three-dimensional space model; A scale recovery module, configured to perform scale recovery on the optimized three-dimensional space model to obtain a spatial three-dimensional layout model with a true scale.

9. An electronic device, wherein, includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the spatial three-dimensional layout acquisition method according to any one of claims 1-7.

10. A computer-readable storage medium, on which a computer program is stored, wherein, the computer program is executed by a processor to implement the spatial three-dimensional layout acquisition method according to any one of claims 1-7.