Scene modeling method and device, equipment, medium and program product

By combining SLAM modeling and room type matching technology, whole-house VR modeling is carried out based on panoramic video or panoramic images, the problem of data acquisition time in the existing technology is solved, efficient and accurate scene modeling is achieved, and user experience is improved.

CN120147547APending Publication Date: 2025-06-13KE COM (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510296755.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing scenario modeling technology takes a long time in the data acquisition and modeling process, resulting in poor user experience and affecting the wide range of applications.

Method used

The real-time positioning and map construction (SLAM) modeling method and room type matching technology are used to realize whole-house VR modeling based on panoramic video or panoramic maps. By obtaining real and estimating layout data, the rotation matrix and displacement vector are determined, and a three-dimensional model that conforms to the real layout is generated.

Benefits of technology

It greatly reduces the cost of data acquisition, improves the efficiency and accuracy of scenario modeling, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147547A_ABST
    Figure CN120147547A_ABST
Patent Text Reader

Abstract

The invention relates to a scene modeling method and device, equipment, a medium and a program product. The method comprises the steps that real layout data and estimated layout data for a target scene are acquired based on a panorama of the target scene, the real layout data comprise a real boundary of the target scene, and the estimated layout data comprise an estimated boundary of the target scene; a target rotation matrix and a target displacement vector are determined, and the target rotation matrix and the target displacement vector are used for processing the estimated layout data, so that the overlapping degree of the processed estimated boundary and the real boundary meets a preset condition; and generating a three-dimensional model of the target scene based on the target rotation matrix and the target displacement vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a method for scene modeling, a device for scene modeling, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Scene modeling has a wide range of applications in life (for example, indoor VR modeling of houses is widely used in house sales, rentals, and decoration businesses), which helps to bring an immersive viewing experience to users. Limited by the data collection and modeling methods, in practical applications, the data collection time required by the operator may be long or the effect may be poor, affecting the user experience.

[0003] The methods described in this section are not necessarily methods that have been previously envisioned or adopted. Unless otherwise specified, no method described in this section should be considered prior art solely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention

[0004] It would be advantageous to provide a mechanism that alleviates, mitigates, or even eliminates one or more of the above problems.

[0005] According to one aspect of the present disclosure, there is provided a method for scene modeling, including: obtaining real layout data and estimated layout data for a target scene based on a panoramic view of the target scene, where the real layout data includes real boundaries of the target scene, and where the estimated layout data includes estimated boundaries of the target scene; determining a target rotation matrix and a target displacement vector, where the target rotation matrix and the target displacement vector are used to process the estimated layout data so that the overlap degree between the processed estimated boundaries and the real boundaries meets a preset condition; and generating a three-dimensional model of the target scene based on the target rotation matrix and the target displacement vector.

[0006] According to one aspect of the present disclosure, there is provided an apparatus for scene modeling, including: an acquisition unit configured to acquire real layout data and estimated layout data for a target scene based on a panoramic view of the target scene, wherein the real layout data includes real boundaries of the target scene, and wherein the estimated layout data includes estimated boundaries of the target scene; a determination unit configured to determine a target rotation matrix and a target displacement vector, wherein the target rotation matrix and the target displacement vector are used to process the estimated layout data such that an overlap degree between the processed estimated boundaries and the real boundaries meets a preset condition; and a generation unit configured to generate a three-dimensional model of the target scene based on the target rotation matrix and the target displacement vector.

[0007] According to one aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and at least one memory storing a computer program thereon, wherein when the computer program is executed by the at least one processor, the at least one processor is caused to execute the above method.

[0008] According to one aspect of the present disclosure, there is provided a computer-readable storage medium storing a computer program thereon, and when the computer program is executed by a processor, the processor is caused to execute the above method.

[0009] According to one aspect of the present disclosure, there is provided a computer program product including a computer program, and when the computer program is executed by a processor, the processor is caused to execute the above method.

[0010] These and other aspects of the present disclosure will be apparent from and will be elucidated with reference to the embodiments described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings exemplarily illustrate embodiments and constitute a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements. In the following description of exemplary embodiments with reference to the drawings, more details, features, and advantages of the present disclosure are disclosed. In the drawings:

[0012] Figure 1 is a schematic diagram illustrating an example system in which various methods described herein may be implemented according to an exemplary embodiment;

[0013] Figure 2 is an exemplary flowchart illustrating a method of scene modeling according to some exemplary embodiments;

[0014] Figures 3A to 3C is a schematic diagram showing an overlap degree calculation process according to some exemplary embodiments;

[0015] Figures 4A to 4B is a schematic diagram showing a whole-house layout estimation process according to some exemplary embodiments;

[0016] Figures 5A to 5C is a schematic diagram showing estimated layout data and actual layout data of multiple functional rooms according to some exemplary embodiments;

[0017] Figures 6A to 6B is a schematic diagram showing a whole-house layout matching process according to some exemplary embodiments;

[0018] Figure 7 is a schematic diagram showing a height calculation process according to some exemplary embodiments;

[0019] Figures 8A to 8B is a schematic diagram showing a gravity correction process according to some exemplary embodiments;

[0020] Figure 9 is a schematic block diagram showing a method for scene modeling according to some exemplary embodiments;

[0021] Figure 10 is a schematic block diagram showing an apparatus for scene modeling according to some exemplary embodiments;

[0022] Figure 11 is a block diagram showing an exemplary computer device that can be applied to exemplary embodiments. Detailed implementation manners

[0023] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the description of the context, they may also refer to different instances.

[0024] In the description of various examples in the present disclosure, the terms used are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. As used herein, the term "a plurality" means two or more, and the term "based on" should be interpreted as "at least partially based on". In addition, the terms "and / or" and "at least one of..." cover any one of the listed items and all possible combinations.

[0025] In the technical solution of the present disclosure, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of user information complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0026] Through scene modeling technologies such as indoor VR modeling of a house, an immersive viewing experience (e.g., an online viewing experience) can be provided for users. During the VR modeling process, data can be collected at various points within the target scene based on a panoramic camera, and finally, reconstruction between multiple points is performed to obtain a VR three-dimensional model. However, this collection and modeling method requires the operator to perform discrete collection separately between multiple points in each room, which takes a long time for collection, resulting in poor application in many business scenarios and affecting the user's viewing experience.

[0027] Based on this, the present disclosure proposes a lightweight collection and modeling solution for application in various scene modeling processes such as indoor VR modeling. It can utilize the Simultaneous Localization and Mapping (SLAM) modeling method and house type matching technology to achieve whole-house VR modeling based on panoramic videos or panoramic images, greatly reducing the data collection cost.

[0028] For clarity, some embodiments of the present disclosure are described by taking indoor VR modeling of a house as an example. It should be understood that in addition to indoor VR modeling of a house, the method proposed by the present disclosure can also be applied to the modeling process in any other field.

[0029] The exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0030] Figure 1 FIG. is a schematic diagram illustrating an example system 100 in which various methods described herein can be implemented according to an exemplary embodiment.

[0031] Refer to Figure 1 FIG., the system 100 includes a client device 110, a server 120, and a network 130 communicatively coupling the client device 110 and the server 120.

[0032] The client device 110 includes a display 114 and a client application (APP) 112 that can be displayed via the display 114. The client application 112 can be an application program that needs to be downloaded and installed before running or a mini-program (liteapp) as a lightweight application program. In the case where the client application 112 is an application program that needs to be downloaded and installed before running, the client application 112 can be pre-installed on the client device 110 and activated. In the case where the client application 112 is a mini-program, the user 102 can directly run the client application 112 on the client device 110 without installing the client application 112 by searching for the client application 112 in the host application (e.g., by the name of the client application 112, etc.) or scanning the graphical code of the client application 112 (e.g., barcode, QR code, etc.). In some embodiments, the client device 110 can be any type of mobile computer device, including a mobile computer, a mobile phone, a wearable computer device (such as a smart watch, a head-mounted device, including smart glasses, etc.) or other types of mobile devices. In some embodiments, the client device 110 can alternatively be a stationary computer device, such as a desktop computer, a server computer or other types of stationary computer devices.

[0033] The server 120 is typically a server deployed by an Internet service provider (ISP) or an Internet content provider (ICP). The server 120 can represent a single server, a cluster of multiple servers, a distributed system, or a cloud server that provides basic cloud services (such as cloud databases, cloud computing, cloud storage, cloud communications). It will be understood that although Figure 1 the server 120 is shown communicating with only one client device 110, the server 120 can provide background services for multiple client devices simultaneously.

[0034] Examples of the network 130 include a local area network (LAN), a wide area network (WAN), a personal area network (PAN), and / or a combination of communication networks such as the Internet. The network 130 can be a wired or wireless network. In some embodiments, technologies and / or formats including HyperText Markup Language (HTML), Extensible Markup Language (XML), etc. are used to process the data exchanged through the network 130. In addition, encryption technologies such as Secure Sockets Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some of the links. In some embodiments, custom and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.

[0035] For the purposes of the embodiments of the present disclosure, in Figure 1In the example, the client application 112 can be an application for displaying a model of a target scene. Correspondingly, the server 120 can be a server used with an application for displaying a 3D model of a target scene. The server 120 can run a scene modeling method and provide the modeled model data to the client application 112 running in the client device 110.

[0036] Figure 2 is an exemplary flowchart illustrating a method for scene modeling according to some exemplary embodiments.

[0037] The method 200 can be executed at a server (e.g., the server 120 shown in Figure 1 ), that is, the execution entity of each step of the method 200 can be the server 120 shown in Figure 1 In some embodiments, the method 200 can be executed at a client device (e.g., the client device 110 shown in Figure 1 ). In some embodiments, the method 200 can be executed in combination by a client device (e.g., the client device 110) and a server (e.g., the server 120). Hereinafter, taking the execution entity as the server 120 as an example, each step of the method 200 will be described in detail.

[0038] Referring to Figure 2 , wherein the method 200 includes steps S202 to S206.

[0039] Step S202, based on the panoramic view of the target scene, obtain the real layout data and the estimated layout data for the target scene, wherein the real layout data includes the real boundary of the target scene, and wherein the estimated layout data includes the estimated boundary of the target scene.

[0040] In some examples, the target scene can be one or more functional rooms in a house, the real layout data of the target scene can be the house type data, and the real boundary of the target scene can be the boundary formed by the wall positions of the house type.

[0041] In some examples, the estimated layout data can be the estimated house type data obtained from the captured panoramic video or panoramic view (e.g., one or more frames in the panoramic video), and the estimated boundary of the target scene can be the boundary formed by the estimated wall positions.

[0042] Step S204, determine the target rotation matrix and the target displacement vector, wherein the target rotation matrix and the target displacement vector are used to process the estimated layout data so that the overlap degree between the processed estimated boundary and the real boundary meets a preset condition.

[0043] It can be understood that the estimated layout data can be two-dimensional matrix data represented in any suitable form. The rotation and translation of the estimated boundary can be achieved by applying the rotation matrix and displacement vector to the estimated layout data.

[0044] In some embodiments, making the overlap degree between the processed estimated boundary and the true boundary meet a preset condition may include: the overlap degree is the maximum overlap degree in multiple calculations, and the overlap degree is greater than or equal to a predetermined threshold.

[0045] Step S206, generate a three-dimensional model of the target scene based on the target rotation matrix and the target displacement vector.

[0046] Thus, by matching the estimated boundary with the true boundary, the registration between the estimated data and the true layout data is achieved, so that the finally generated three-mode model conforms to the true layout data, and thus the three-dimensional modeling task can be completed efficiently and accurately.

[0047] In some embodiments, the overlap degree between the estimated boundary and the true boundary can be determined by the overlapping area between the area enclosed by the estimated boundary and the area enclosed by the true boundary. In some embodiments, the overlap degree can also be determined by the distance between the geometric center point of the estimated boundary and the geometric center point of the true boundary. The following will be combined with Figures 3A to 3C Further introduce various methods for determining the overlap degree.

[0048] Figures 3A to 3C FIG. is a schematic diagram showing a process for calculating the overlap degree according to some exemplary embodiments. Exemplarily, the dashed boundary can be the estimated boundary of the target scene in the estimated layout data, and the solid boundary can be the true boundary of the target scene in the true layout data.

[0049] Figure 3A The gray part in FIG. shows the overlapping area between the area enclosed by the estimated boundary and the area enclosed by the true boundary. In some embodiments, the overlap degree can be the ratio of the overlapping area to the area of the region enclosed by the true boundary, or the ratio of the overlapping area to the area of the region enclosed by the estimated boundary. In some embodiments, the overlap degree can be the ratio of the overlapping area to the total area of the region enclosed by the true boundary and the estimated boundary.

[0050] Figures 3B to 3CFurther shown is a polar coordinate-based overlap degree calculation method according to some exemplary embodiments. For each of at least one preset direction, the following operations are performed: calculating a first radial vector between a preset point (e.g., point O_lt) within the estimated boundary and the intersection point of the estimated boundary along the preset direction, and a second radial vector between the preset point and the intersection point of the true boundary along the preset direction; determining the smaller value and the larger value among the lengths of the first radial vector and the second radial vector; and calculating the ratio between the square of the smaller value and the square of the larger value.

[0051] By summing up the ratios calculated for each of at least one preset direction, the overlap degree between the estimated boundary and the true boundary is obtained.

[0052] In some embodiments, the preset point may be the geometric center point of the estimated boundary. In some other embodiments, the preset point may be a point close to the geometric center point of the estimated boundary (e.g., a point within a predetermined range from the geometric center point).

[0053] Exemplarily, Figure 3B shows an example of the smaller value among the lengths of the first radial vector and the second radial vector, while Figure 3C shows an example of the larger value among the lengths of the first radial vector and the second radial vector.

[0054] In some embodiments, at least one preset direction may be a plurality of angles (e.g., N angles) evenly divided along [0, 2pi]. For the sake of clarity, Figures 3B to 3C only a partial number of the angles are shown. The lengths of the radial vectors between the preset point and the estimated boundary along the plurality of angles can be represented as r 1i i, while the lengths of the radial vectors between the preset point and the true boundary can be represented as r 2i i, where i indicates the corresponding angle.

[0055] As described above, the ratio between the square of the smaller value and the square of the larger value can be calculated, and the overlap degree can be obtained by summing up the ratios calculated for each of at least one preset direction, that is, the overlap degree is calculated through the following formula:

[0056]

[0057] In some embodiments, the overlap degree can also be calculated through the following formula to achieve a balance between accuracy and computational efficiency:

[0058]

[0059] Thus, through the overlap degree calculation method based on polar coordinates, the overlap degree can be quickly estimated to evaluate whether the estimated boundary matches the true boundary. Generally, the larger the overlap degree, the better the matching result.

[0060] For indoor VR modeling, estimated layout data can be obtained from panoramic videos or panoramic images. Figures 4A to 4B FIG. shows a schematic diagram of a whole-house layout estimation process according to some exemplary embodiments.

[0061] As Figure 4A shown, through a deep network or computer vision technology, the main room structures in the interior (e.g., Figure 4A the intersection lines of the ceiling, walls, and floor shown in ) can be estimated in one or more frames (also referred to as "key frames") of the panoramic video or in the panoramic image, thereby obtaining a three-dimensional spatial structure as shown in Figure 4B The three-dimensional structure includes height information, and the two-dimensional projection of the three-dimensional structure can be used as the estimated layout data for matching with the true layout data.

[0062] During the indoor VR modeling process, a house may involve one or more functional rooms. Figures 5A to 5C FIG. shows a schematic diagram of the estimated layout data and true layout data of multiple functional rooms according to some exemplary embodiments.

[0063] As Figure 5A shown, the estimated layout data of each functional room, including the estimated boundary of the functional room, can be obtained through the panoramic video or panoramic image of each functional room. The relative positions of the functional rooms can be obtained through the panoramic video or panoramic image at the link positions of the functional rooms, obtaining the estimated layout data of the house as shown in Figure 5B for comparison with the true layout data of the house as shown in Figure 5C It can be understood that various computer learning technologies and computer vision technologies can be used to implement the process of obtaining the estimated layout data from the panoramic image.

[0064] For a house including multiple functional rooms, the ratio between the square of the smaller value and the square of the larger value between the first radial vector and the second radial vector can be obtained for each functional room, and the ratios of multiple functional rooms are summed to obtain the overlap degree of multiple functional rooms. In some embodiments, the ratios of multiple functional rooms can be weighted and summed to obtain the final overlap degree, and the weights in the weighted sum can be preset based on the type of the functional room, such as presetting different weights based on whether the functional room belongs to a bedroom, a kitchen, or a bathroom, or can be related to the area or shape of the functional room.

[0065] In some embodiments, the maximum overlap degree can be obtained through iterative calculations to achieve the matching of the whole-house layout. Figures 6A to 6B FIG. shows a schematic diagram of the whole-house layout matching process according to some exemplary embodiments, including a rotation process and a displacement process.

[0066] As Figure 6A shown, there may be an angular difference between the estimated layout data ( Figure 6A the part shown in light gray in ) and the real layout data ( Figure 6A the part shown in dark gray in ). The estimated layout data can be processed by a rotation matrix to match the angle of the real layout data.

[0067] In some embodiments, the panoramic image can be corrected (e.g., gravity correction) so that the camera direction is perpendicular to a certain wall. Thus, when the walls of the real house type data are parallel to the coordinate axes, the angular difference between the estimated layout data and the real layout data is close to one of 0 degrees, 90 degrees, 180 degrees, and 270 degrees. The rotation matrices corresponding to 0 degrees, 90 degrees, 180 degrees, and 270 degrees can be applied to the estimated layout data, and the difference between the processed estimated layout data and the real layout data is calculated. The result with the smallest difference (e.g., the result with the highest overlap degree) is the correct rotation direction.

[0068] As Figure 6B shown, there may be a displacement difference between the estimated layout data and the real layout data. The estimated layout data can be processed by a displacement vector to match the position of the real layout data.

[0069] In some embodiments, due to possible errors in the estimated layout data, some structures may be incorrect or missing. For example, as Figure 6B shown, the estimated layout data shown by the dashed line is missing the corner area in the lower right corner. Using traditional corner matching algorithms may result in large errors, or even incorrect matching results as shown in the Figure 6B middle picture. In addition, using traditional discrete-point-based matching methods for global optimal registration may cause large position deviations between some functions and cannot obtain the correct matching results as shown in the Figure 6B right picture.

[0070] According to some embodiments, through an iterative algorithm, the target rotation matrix and the target displacement vector can be determined to match the estimated layout data and the real layout data. Specifically, steps S204 in the above method 200, determining the target rotation matrix and the target displacement vector, may include:

[0071] For each preset rotation matrix in at least one preset rotation matrix, obtain an initial displacement vector, and iteratively perform the following operations until a preset criterion is met: obtain the displaced estimated boundary by applying the displacement vector to the estimated layout data; calculate the overlap degree between the displaced estimated boundary and the true boundary; and update the displacement vector based on the error vector of the displaced estimated boundary relative to the true boundary.

[0072] Based on the preset rotation matrix and displacement vector corresponding to the highest overlap degree among the overlap degrees between the displaced estimated boundary and the true boundary, determine the target rotation matrix and the target displacement vector.

[0073] Thus, through iterative calculation, the target rotation matrix and the target displacement vector under the optimal matching result (i.e., the highest overlap degree) can be determined, so as to be used for processing the three-dimensional model of the target scene, making the three-mode model match the true layout data.

[0074] Exemplarily, the preset rotation matrix can be a matrix that can rotate the two-dimensional layout data by a predetermined angle. For example, the two-dimensional layout data is rotated by θ degrees through the following rotation matrix R:

[0075]

[0076] In some embodiments, the above-described polar coordinate-based overlap degree calculation method can be used in the iterative algorithm to calculate the overlap degree. In some embodiments, the preset criterion can be that the error vector is less than a given threshold, or the number of iterations reaches a given threshold. Figures 3B to 3C Based on the error vector of the displaced estimated boundary relative to the true boundary, updating the displacement vector includes: for each preset direction in at least one preset direction, calculate the difference vector between the intersection point of a preset point on the displaced estimated boundary along the preset direction and the displaced estimated boundary and the intersection point with the true boundary; determine the error vector based on the difference vector with the largest length; and add the error vector to the displacement vector to update the displacement vector.

[0077] In some embodiments, the difference vector can be the difference vector between the radial vector of the preset point along the plurality of angles relative to the estimated boundary and the radial vector relative to the true boundary as described above. Thus, the error vector represents the direction with the largest radial difference, and the displacement direction can be iteratively optimized through the error vector.

[0078] In some embodiments, the difference vector can be as described above in combination with Figures 3B to 3C the radial vector between the preset point along the plurality of angles relative to the estimated boundary and the radial vector relative to the true boundary. Thus, the error vector represents the direction with the largest radial difference, and the displacement direction can be iteratively optimized through the error vector.

[0079] According to some embodiments, the target scene corresponds to one of multiple functional spaces, and based on the difference vector with the largest length in the difference vectors, determining the error vector includes: performing a weighted sum of the difference vectors with the largest length in the difference vectors of each functional space among the multiple functional spaces as the error vector, where the weights for the weighted sum are based on the area of the corresponding functional space. Thus, iterative calculation for the multi-functional space can be achieved, fully considering the importance of different functional spaces and their impact on the modeling result.

[0080] In some embodiments, the weight can also be preset based on the type of the functional space.

[0081] Exemplarily, in some embodiments, the following algorithm can be adopted:

[0082] Input:

[0083] Layout set (e.g., the above-mentioned estimated layout data), {L_i}, i = [0, N - 1], N represents the number of panoramic images, and L_i represents the corresponding estimated boundary

[0084] Room set (e.g., the above-mentioned true layout data), {H_j}, j = [0, M - 1], M represents the number of rooms, and H_j represents the corresponding true boundary

[0085] Output:

[0086] R: Rotation (e.g., the above-mentioned target rotation matrix)

[0087] t: translation (e.g., the above-mentioned target shift vector)

[0088] Algorithm:

[0089] · FOR Ri corresponding to [0, 90, 180, 270]:

[0090] a) Calculate the outer frames of {L_i} and {H_j} respectively, BBox_L and BBbox_H

[0091] b) Calculate the initial displacement vector translation between the two through BBox_L and BBbox_H

[0092] c) Based on translation, calculate the initial matching relationship between {L_i} and {H_j}, and determine the belonging functional space by the landing point of the preset point (e.g., the geometric center point or close to the geometric center point) of L_i

[0093] d) According to translation, calculate the overlap degree of {L_i} and {H_j}

[0094] e) Weight-average the radial vector errors to obtain delta_xy, where the weights are from the areas between the corresponding functions.

[0095] f) translation += delta_xy. When |delta_xy| is greater than a certain threshold, repeat steps d - f.

[0096] g) Retain the maximum overlap value, as well as the corresponding translation and Ri, {Ri, translation, overlap}.

[0097] · The rotation matrix and displacement vector corresponding to the result with the maximum overlap are the target rotation matrix and target displacement vector.

[0098] In addition to the rotation matrix and displacement vector, in some embodiments, other parameters are also considered, such as a scaling parameter.

[0099] According to some embodiments, method 200 may further include: obtaining a height estimate for the target scene based on the panoramic view; determining a scaling parameter based on the height estimate and the true height of the target scene; and applying the scaling parameter to the estimated layout data. Since there is a strong correlation between the estimated layout data and the floor height of the target scene, by constraining with the true height, the scaling parameter relative to the true scene can be determined. In some embodiments, a constraint of consistent floor height is also imposed on all the estimated layout data, such that the scaling factors of all the estimated layout data are kept consistent.

[0100] Figure 7 FIG. shows a schematic diagram of a height calculation process according to some exemplary embodiments. In the Figure 7 shown shooting scene, the shooting angle of the intersection line between the ceiling and the wall is θ c , and the shooting angle of the intersection line between the floor and the wall is θ f .

[0101] The height h of the camera shooting position from the ceiling and the height h c from the floor can be calculated through the following formulas f for use in estimating the scaling of the layout data and model reconstruction:

[0102]

[0103] h c + h f = true floor height

[0104] According to some embodiments, method 200 may further include: correcting the panoramic view based on the vanishing point information in the panoramic view such that the gravity direction of the panoramic view is the vertical direction and the shooting direction of the panoramic view is perpendicular to the wall of the target scene.

[0105] In some embodiments, the vanishing point information of the entire image is calculated based on the orientations of line segments in the image (such as the intersection lines between the ceiling and the wall, and between the ceiling and the floor), and the vanishing point information of each line segment in the corrected panoramic image is consistent with each other.

[0106] During the data acquisition process, the camera may face any direction. Through gravity correction, the distortion in the image can be reduced, and the angular difference between the estimated layout data and the real layout data can be close to one of 0 degrees, 90 degrees, 180 degrees, and 270 degrees, reducing the number of calculations for the rotation matrix.

[0107] Figures 8A to 8B FIG. is a schematic diagram showing the gravity correction process according to some exemplary embodiments. Among them, Figure 8A shows the panoramic image before gravity correction. It can be seen that its gravity direction is not vertical. Figure 8B shows the panoramic image after gravity correction, and its gravity direction is vertical.

[0108] In some embodiments, one or more frames (also referred to as key frames) of the panoramic video can be corrected, and the average value of the correction parameters can be obtained. The entire model reconstruction data is corrected by this average value.

[0109] According to some embodiments, step S206 in method 200, generating a three-dimensional model of the target scene based on the target rotation matrix and the target displacement vector, may include: obtaining the reconstructed pose information of the target scene via the Simultaneous Localization and Mapping (SLAM) modeling method based on the panoramic video of the target scene; applying the target rotation matrix and the target displacement vector to the reconstructed pose information to obtain the real pose information of the target scene; and generating a three-dimensional model of the target scene based on the real pose information.

[0110] Thus, by combining the SLAM modeling method, the real pose information of one or more frames in the panoramic video relative to the real house type can be obtained to improve the accuracy and display effect of the final three-dimensional model.

[0111] Figure 9The figure is a schematic diagram showing a method for scene modeling according to some exemplary embodiments. Among them, first, a panoramic video of the target scene is collected and SLAM reconstruction is performed to obtain the pose information of one or more frames. Apply the above-mentioned gravity correction method to correct one or more frames of panoramic images, and based on the mean value of the correction parameters of the one or more frames, correct the entire SLAM reconstruction result. Based on the corrected panoramic image, estimated layout data can be obtained, and further, a whole-house layout estimate including one or more functional rooms can be obtained. Based on the real layout data, the above-mentioned whole-house layout matching method can be applied to obtain the target rotation matrix and the target displacement vector. Finally, the target rotation matrix and the target displacement vector are applied to the reconstructed pose information to obtain the real pose information of the target scene and generate a three-dimensional VR model of the target scene.

[0112] Exemplarily, in some embodiments, scene modeling can be implemented based on the following specific steps:

[0113] Step 1: Video acquisition through a panoramic camera. Starting from the position of the entrance door, walk uniformly along each room indoors to ensure video acquisition of each room, and the video finally returns to the initial position.

[0114] Step 2: Perform whole-house reconstruction on the collected panoramic video through SLAM to obtain the point cloud of the whole house and the pose information of one or more frames.

[0115] Step 3: Correct the SLAM result in the gravity direction

[0116] 1. Based on the correction principle of the panoramic image, correct the gravity direction of one or more frames, and calculate the relative rotation matrix of each frame with respect to the gravity axis direction;

[0117] 2. For all frames in one or more frames, calculate the mean matrix of the relative rotation matrices as the relative rotation of the SLAM reconstruction result with respect to the gravity axis direction;

[0118] 3. Apply this mean matrix to the SLAM result, and correct the SLAM reconstruction result (i.e., the pose information of each frame)

[0119] to be perpendicular to the gravity axis.

[0120] Step 4: Estimate the whole-house layout based on one or more frames of the panoramic video

[0121] 1. Perform layout estimation on one or more frames;

[0122] 2. Based on the relative pose information between one or more frames, filter out frames with too large errors (for example, frames with pose information that is too different from other frames);

[0123] 3. Vertically project the whole-house layout result onto the horizontal plane, and this result can be regarded as the estimated layout data.

[0124] Step 5: Match the estimated layout data with the real layout data to obtain the true pose information of the SLAM result relative to the real layout data;

[0125] Step 6: Output the VR modeling result

[0126] 1. Generate a geometric model of the whole house based on the real layout data;

[0127] 2. Each functional room selects points under the functional room according to certain rules, and the user can select the points to preview the 3D VR model at the points;

[0128] 3. Each functional room generates the texture map of the functional room based on one point according to the panoramic view and pose information of the point;

[0129] 4. Finally, output the geometric model with texture, as well as the panoramic views and pose information of all points;

[0130] Step 7: VR effect visualization. Through the VR front-end tool, the VR modeling result is visualized.

[0131] Figure 10 FIG. is a schematic block diagram showing a device 1000 for scene modeling according to some exemplary embodiments. The device 1000 includes:

[0132] An acquisition unit 1010 configured to acquire real layout data and estimated layout data for a target scene based on a panoramic view of the target scene, wherein the real layout data includes the real boundary of the target scene, and wherein the estimated layout data includes the estimated boundary of the target scene;

[0133] A determination unit 1020 configured to determine a target rotation matrix and a target displacement vector, wherein the target rotation matrix and the target displacement vector are used to process the estimated layout data so that the overlap degree between the processed estimated boundary and the real boundary meets a preset condition; and

[0134] A generation unit 1030 configured to generate a three-dimensional model of the target scene based on the target rotation matrix and the target displacement vector.

[0135] It should be understood that Figure 10 each module of the device 1000 shown in Figure 2 may correspond to each step in the method 200 described with reference to

[0136] Although specific functions have been discussed above with reference to specific modules, it should be noted that the functions of the various modules discussed herein can be divided into multiple modules, and / or at least some of the functions of multiple modules can be combined into a single module. The actions performed by the specific modules discussed herein include the specific module itself performing the action, or alternatively the specific module invoking or otherwise accessing another component or module that performs the action (or performs the action in combination with the specific module). Thus, a specific module that performs an action can include the specific module itself that performs the action and / or another module that the specific module invokes or otherwise accesses and that performs the action. For example, the acquisition unit 1010 and the determination unit 1020 can be combined into a single module in some embodiments. As another example, the acquisition unit 1010 can include the determination unit 1020 in some embodiments. As used herein, the phrase "entity A initiates action B" can mean that entity A issues an instruction to perform action B, but entity A itself does not necessarily perform action B.

[0137] It should also be understood that the various techniques herein can be described in the general context of software-hardware elements or program modules. The various modules described above Figure 10 can be implemented in hardware or in hardware in combination with software and / or firmware. For example, these modules can be implemented as computer program code / instructions that are configured to execute in one or more processors and stored in a computer-readable storage medium. Alternatively, these modules can be implemented as hardware logic / circuits. For example, in some embodiments, one or more of the acquisition unit 1010, the determination unit 1020, and the generation unit 1030 can be implemented together in a system on chip (SoC). The SoC can include an integrated circuit chip (which includes one or more components such as a processor (e.g., a central processing unit (CPU), a microcontroller, a microprocessor, a digital signal processor (DSP), etc.), a memory, one or more communication interfaces, and / or other circuits), and can optionally execute the received program code and / or include embedded firmware to perform functions.

[0138] According to one aspect of the present disclosure, there is provided an electronic device, at least one processor; and at least one memory having stored thereon a computer program, which when executed by the at least one processor, causes the at least one processor to perform the steps of any of the method embodiments described above.

[0139] According to one aspect of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program, which when executed by a processor, implements the steps of any of the method embodiments described above.

[0140] According to one aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the steps of any of the method embodiments described above.

[0141] In the following, illustrative examples of such electronic devices, non-transitory computer-readable storage media, and computer program products will be described in conjunction with Figure 11 description.

[0142] Figure 11 An example configuration of a computer device 1100 that can be used to implement the methods described herein is shown. By way of example, Figure 1 the server 120 and / or client device 110 shown in may include an architecture similar to that of the computer device 1100. The above-described apparatus 1000 may also be implemented in whole or at least in part by the computer device 1100 or a similar device or system.

[0143] The computer device 1100 may be of various different types. Examples of the computer device 1100 include, but are not limited to: desktop computers, server computers, laptop or netbook computers, mobile devices (e.g., tablets, cellular or other wireless telephones (e.g., smartphones), notepad computers, mobile stations), wearable devices (e.g., glasses, watches), entertainment devices (e.g., entertainment appliances, set-top boxes communicatively coupled to a display device, gaming consoles), televisions or other display devices, automotive computers, and the like.

[0144] The computer device 1100 may include at least one processor 1102, a memory 1104, (one or more) communication interfaces 1106, a display device 1108, other input / output (I / O) devices 1110, and one or more mass storage devices 1112 that are capable of communicating with each other, such as via a system bus 1114 or other suitable connection.

[0145] The processor 1102 may be a single processing unit or multiple processing units, and all processing units may include a single or multiple computing units or multiple cores. The processor 1102 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any device that manipulates signals based on operation instructions. Among other capabilities, the processor 1102 may be configured to obtain and execute computer-readable instructions stored in the memory 1104, the mass storage device 1112, or other computer-readable media, such as program code of an operating system 1116, program code of an application 1118, program code of other programs 1120, and the like.

[0146] Memory 1104 and mass storage device 1112 are examples of computer-readable storage media for storing instructions that are executed by processor 1102 to implement the various functions described above. For example, memory 1104 generally can include both volatile and non-volatile memory (e.g., RAM, ROM, etc.). In addition, mass storage device 1112 generally can include a hard disk drive, a solid state drive, removable media, including external and removable drives, memory cards, flash memory, floppy disks, optical discs (e.g., CD, DVD), storage arrays, network attached storage, storage area networks, and the like. Memory 1104 and mass storage device 1112 can both be collectively referred to herein as memory or computer-readable storage media and can be non-transitory media capable of storing computer-readable, processor-executable program instructions as computer program code that can be executed by processor 1102 as a particular machine configured to implement the operations and functions described in the examples herein.

[0147] Multiple programs can be stored on mass storage device 1112. These programs include operating system 1116, one or more application programs 1118, other programs 1120, and program data 1122, and they can be loaded into memory 1104 for execution. Examples of such application programs or program modules can include, for example, computer program logic (e.g., computer program code or instructions) for implementing the following components / functions: client application 112, method 200 (including any suitable steps of method 200) and / or additional embodiments described herein.

[0148] Although illustrated as being stored in memory 1104 of computer device 1100 in Figure 11 , module 1116, 1118, 1120, and 1122 or portions thereof can be implemented using any form of computer-readable medium accessible by computer device 1100. As used herein, "computer-readable medium" includes at least two types of computer-readable media, namely computer-readable storage media and communication media.

[0149] A computer-readable storage medium includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVDs), or other optical storage devices, magnetic cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible to a computing device. In contrast, a communication medium can embody computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism. The computer-readable storage media as defined herein does not include a communication medium.

[0150] One or more communication interfaces 1106 are used to exchange data with other devices, such as via a network, a direct connection, and the like. Such communication interfaces can be one or more of the following: any type of network interface (e.g., network interface card (NIC)), wired or wireless (such as IEEE 802.11 wireless local area network (WLAN)) wireless interface, Worldwide Interoperability for Microwave Access (WiMAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, BluetoothTM interface, Near Field Communication (NFC) interface, etc. The communication interface 1106 can facilitate communication within a variety of network and protocol types, including wired networks (e.g., LAN, cable, etc.) and wireless networks (e.g., WLAN, cellular, satellite, etc.), the Internet, and the like. The communication interface 1106 can also provide communication with external storage devices (not shown) such as in storage arrays, network-attached storage, storage area networks, and the like.

[0151] In some examples, a display device 1108, such as a monitor, can be included for displaying information and images to a user. Other I / O devices 1110 can be devices that receive various inputs from a user and provide various outputs to the user, and can include touch input devices, gesture input devices, cameras, keyboards, remote controls, mice, printers, audio input / output devices, and the like.

[0152] The techniques described herein may be supported by these various configurations of the computer device 1100 and are not limited to the specific examples of the techniques described herein. For example, the functionality may also be implemented in whole or in part using a distributed system on a "cloud". The cloud comprises and / or represents a platform for resources. The platform abstracts the underlying functionality of the hardware (e.g., servers) and software resources of the cloud. The resources may include applications and / or data that may be used when performing computing processing on servers remote from the computer device 1100. The resources may also include services provided over the Internet and / or over a subscriber network such as a cellular or Wi-Fi network. The platform may abstract the resources and functionality to connect the computer device 1100 with other computer devices. Accordingly, the implementation of the functionality described herein may be distributed across the entire cloud. For example, the functionality may be implemented in part on the computer device 1100 and in part through a platform that abstracts the functionality of the cloud.

[0153] Although the present disclosure has been illustrated and described in detail in the accompanying drawings and foregoing description, such illustration and description are to be considered illustrative and exemplary, and not restrictive; the present disclosure is not limited to the disclosed embodiments. Variations of the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed subject matter, from a study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps not listed, the indefinite article "a" or "an" does not exclude a plurality, the term "plural" means two or more, and the term "based on" shall be construed as "at least partially based on". The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.

Claims

1. A method for scene modeling, comprising: Based on a panoramic image of a target scene, acquiring real layout data and estimated layout data for the target scene, wherein the real layout data includes a real boundary of the target scene, and wherein the estimated layout data includes an estimated boundary of the target scene; Determining a target rotation matrix and a target displacement vector, wherein the target rotation matrix and the target displacement vector are used to process the estimated layout data so that the overlap between the processed estimated boundary and the real boundary satisfies a preset condition; and A three-dimensional model of the target scene is generated based on the target rotation matrix and the target displacement vector.

2. The method according to claim 1, further comprising: For each of at least one preset direction, do the following: Calculating a first radial vector between a preset point within the estimated boundary and an intersection point of the estimated boundary along the preset direction, and a second radial vector between the preset point and an intersection point of the real boundary along the preset direction; determining a smaller value and a larger value of a length of the first radial vector and a length of the second radial vector; as well as calculating a ratio between the square of the smaller value and the square of the larger value; as well as The overlap degree between the estimated boundary and the actual boundary is obtained by summing the ratios calculated for each of the at least one preset direction.

3. The method according to claim 1 or 2, wherein: Determining the target rotation matrix and the target displacement vector includes: For each preset rotation matrix in at least one preset rotation matrix, an initial displacement vector is obtained, and the following operations are iteratively performed until a preset criterion is met: Obtaining a displaced estimated boundary by applying the displacement vector to the estimated layout data; Calculating the degree of overlap between the shifted estimated boundary and the actual boundary; and updating the displacement vector based on an error vector of the displaced estimated boundary relative to the true boundary; and The target rotation matrix and the target displacement vector are determined based on a preset rotation matrix and a displacement vector corresponding to a highest degree of overlap among the degrees of overlap between the displaced estimated boundary and the real boundary.

4. The method according to claim 3, wherein: Based on an error vector of the shifted estimated boundary relative to the true boundary, updating the displacement vector comprises: For each preset direction of at least one preset direction, calculating a difference vector between an intersection point of a preset point in the displaced estimated boundary along the preset direction with the displaced estimated boundary and an intersection point with the real boundary; determining the error vector based on the difference vector with the largest length among the difference vectors; and The error vector is added to the displacement vector to update the displacement vector.

5. The method according to claim 4, wherein: The target scenario corresponds to a functional room among a plurality of functional rooms, and wherein the determining the error vector based on the difference vector with the largest length among the difference vectors comprises: The difference vector with the longest length among the difference vectors between each of the multiple functional rooms is weightedly summed to obtain the error vector, wherein the weight of the weighted sum is based on the area of ​​the corresponding functional room.

6. The method according to claim 1 or 2, further comprising: Based on the panoramic image, obtaining a height estimate for the target scene; determining a scaling parameter based on the height estimate and a true height of the target scene; as well as The scaling parameter is applied to the estimated layout data.

7. The method according to claim 1 or 2, wherein: Generating a three-dimensional model of the target scene based on the target rotation matrix and the target displacement vector includes: Based on the panoramic video of the target scene, the reconstructed pose information of the target scene is obtained through a real-time positioning and map construction SLAM modeling method; Applying the target rotation matrix and the target displacement vector to the reconstructed pose information to obtain the real pose information of the target scene; and Based on the real pose information, the three-dimensional model of the target scene is generated.

8. The method according to claim 1 or 2, further comprising: The panoramic image is corrected based on the vanishing point information in the panoramic image, so that the gravity direction of the panoramic image is a vertical direction, and the shooting direction of the panoramic image is perpendicular to the wall of the target scene.

9. A scene modeling device, comprising: an acquisition unit configured to acquire real layout data and estimated layout data for a target scene based on a panoramic image of the target scene, wherein the real layout data includes a real boundary of the target scene, and wherein the estimated layout data includes an estimated boundary of the target scene; a determining unit configured to determine a target rotation matrix and a target displacement vector, wherein the target rotation matrix and the target displacement vector are used to process the estimated layout data so that an overlap between the processed estimated boundary and the real boundary satisfies a preset condition; and A generating unit is configured to generate a three-dimensional model of the target scene based on the target rotation matrix and the target displacement vector.

10. An electronic device, comprising: at least one processor; as well as at least one memory having a computer program stored thereon, When the computer program is executed by the at least one processor, the at least one processor executes the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program, which, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 8.