Model training method and device, three-dimensional model establishing method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202480004604.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-07
- Filing Date
- 2024-05-14
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art lacks universality when establishing three-dimensional models, making it difficult to adapt to different types of cameras and tasks, resulting in high algorithm complexity and poor scalability.
A method based on sequence prediction is proposed, and a three-dimensional model is established by determining the target prediction sequence using the initial model and image feature sequence. The method includes two main steps: first, use the initial model and the image feature sequence of the sample three-dimensional scene to adjust the initial model to obtain the target model; second, use the target model and the image feature sequence of the three-dimensional scene to be modeled to determine the target prediction sequence to establish the three-dimensional model.
A general three-dimensional scene modeling algorithm framework is implemented, which can be applied to different types of cameras and modeling tasks, reducing the complexity and resource requirements of three-dimensional modeling, and improving modeling efficiency and accuracy.
Smart Images

Figure CN120153400A_ABST
Abstract
Description
Model training method and device, method and device for establishing three-dimensional model, electronic device and storage medium Technical Field
[0001] The present disclosure relates to the field of computer technology, in particular to the fields of computer vision, three-dimensional reconstruction, and three-dimensional generation, and provides a model training method and device, a method and device for establishing a three-dimensional model, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] 3D reconstruction refers to the process of creating a 3D model of a 3D scene (or 3D object) from a 2D image of that scene (or 3D object). 3D reconstruction is a long-standing research topic in computer vision.
[0003] Summary of the Invention
[0004] A first aspect of the present disclosure provides a model training method, including:
[0005] Determining a target prediction sequence corresponding to the sample three-dimensional scene using the initial model and the image feature sequence corresponding to the sample three-dimensional scene; wherein the target prediction sequence is used to establish a three-dimensional model corresponding to the sample three-dimensional scene, and the image feature sequence is determined based on a two-dimensional image corresponding to the sample three-dimensional scene; and
[0006] The target prediction sequence and the target label sequence of the sample 3D scene are used to adjust the initial model to obtain the target model.
[0007] A second aspect of the embodiments of the present disclosure provides a method for establishing a three-dimensional model, comprising:
[0008] Determining a target prediction sequence corresponding to the three-dimensional scene to be modeled using the target model and an image feature sequence corresponding to the three-dimensional scene to be modeled; wherein the image feature sequence is determined based on a two-dimensional image corresponding to the three-dimensional scene to be modeled; and
[0009] According to the target prediction sequence, a three-dimensional model corresponding to the three-dimensional scene to be modeled is established; wherein the target model is trained using the model training method provided by the first aspect of the embodiment of the present disclosure.
[0010] A third aspect of the present disclosure provides a model training device, including:
[0011] The first determination module is configured to use the initial model and the image feature sequence corresponding to the sample three-dimensional scene to determine the target prediction sequence corresponding to the sample three-dimensional scene; wherein the target prediction sequence is used to establish a three-dimensional model corresponding to the sample three-dimensional scene, and the image feature sequence is determined based on the two-dimensional image corresponding to the sample three-dimensional scene; and the adjustment module is configured to use the target prediction sequence and the target label sequence of the sample three-dimensional scene to adjust the initial model to obtain the target model.
[0012] A fourth aspect of the embodiments of the present disclosure provides an apparatus for establishing a three-dimensional model, including:
[0013] a third determining module configured to determine a target prediction sequence corresponding to the three-dimensional scene to be modeled using the target model and an image feature sequence corresponding to the three-dimensional scene to be modeled; wherein the image feature sequence is determined based on a two-dimensional image corresponding to the three-dimensional scene to be modeled; and
[0014] The modeling module is configured to establish a three-dimensional model corresponding to the three-dimensional scene to be modeled according to the target prediction sequence; wherein the target model is trained using the model training device provided by the third aspect of the embodiment of the present disclosure.
[0015] A fifth aspect of the present disclosure provides a model training method, including:
[0016] Determining a feature sequence corresponding to the sample three-dimensional scene based on prompt information corresponding to the sample three-dimensional scene;
[0017] Determining a target prediction sequence corresponding to the sample three-dimensional scene using the initial model and the feature sequence; wherein the target prediction sequence is used to establish a three-dimensional model corresponding to the sample three-dimensional scene; and
[0018] The target prediction sequence and the target label sequence of the sample 3D scene are used to adjust the initial model to obtain the target model.
[0019] A sixth aspect of the embodiments of the present disclosure provides a method for establishing a three-dimensional model, comprising:
[0020] Determining a feature sequence corresponding to the three-dimensional scene to be modeled according to prompt information corresponding to the three-dimensional scene to be modeled;
[0021] Using the target model and the feature sequence, a target prediction sequence corresponding to the three-dimensional scene to be modeled is determined; and based on the target prediction sequence, a three-dimensional model corresponding to the three-dimensional scene to be modeled is established; wherein the target model is trained using the model training method provided in the fifth aspect of the embodiment of the present disclosure.
[0022] A seventh aspect of the embodiments of the present disclosure provides a model training device, including:
[0023] A first acquisition module is configured to determine a feature sequence corresponding to the sample three-dimensional scene based on prompt information corresponding to the sample three-dimensional scene;
[0024] The first determination module is configured to use the initial model and the feature sequence to determine the target prediction sequence corresponding to the sample three-dimensional scene; wherein the target prediction sequence is used to establish a three-dimensional model corresponding to the sample three-dimensional scene; and the adjustment module is configured to use the target prediction sequence and the target label sequence of the sample three-dimensional scene to adjust the initial model to obtain the target model.
[0025] An eighth aspect of the embodiments of the present disclosure provides an apparatus for establishing a three-dimensional model, including:
[0026] a second acquisition module configured to determine a feature sequence corresponding to the three-dimensional scene to be modeled based on prompt information corresponding to the three-dimensional scene to be modeled;
[0027] a third determining module configured to determine a target prediction sequence corresponding to the three-dimensional scene to be modeled using the target model and the feature sequence; and
[0028] The modeling module is configured to establish a three-dimensional model corresponding to the three-dimensional scene to be modeled according to the target prediction sequence; wherein the target model is obtained by training using the model training device provided by the seventh aspect of the embodiment of the present disclosure.
[0029] A ninth aspect of the present disclosure provides an electronic device, including:
[0030] at least one processor; and
[0031] a memory communicatively connected to the at least one processor; wherein,
[0032] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method provided by any aspect of the embodiments of the present disclosure.
[0033] A tenth aspect of the embodiments of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute the method provided by any aspect of the embodiments of the present disclosure.
[0034] An eleventh aspect of the embodiments of the present disclosure provides a computer program product, including a computer program, which implements the method provided by any aspect of the embodiments of the present disclosure when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the multiple drawings represent the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments provided in accordance with the present disclosure and should not be regarded as limiting the scope of the present disclosure.
[0036] FIG1 is a schematic diagram of an application scenario according to an embodiment of the present disclosure.
[0037] FIG2 is a flowchart of an implementation of a model training method provided in the first aspect of an embodiment of the present disclosure.
[0038] FIG3A is a first schematic diagram of a two-dimensional image according to an embodiment of the present disclosure.
[0039] FIG3B is a second schematic diagram of a two-dimensional image according to an embodiment of the present disclosure.
[0040] FIG4 is a schematic diagram of determining a target tag sequence corresponding to a sample three-dimensional scene according to an embodiment of the present disclosure.
[0041] FIG5 is a schematic block diagram of an initial model according to an embodiment of the present disclosure.
[0042] FIG6A is a flowchart illustrating an implementation of a method for establishing a three-dimensional model provided in the second aspect of an embodiment of the present disclosure.
[0043] FIG6B is a first schematic diagram of a three-dimensional model obtained based on a two-dimensional image according to an embodiment of the present disclosure.
[0044] FIG6C is a second schematic diagram of a three-dimensional model obtained based on a two-dimensional image according to an embodiment of the present disclosure.
[0045] FIG6D is a third schematic diagram of a three-dimensional model obtained based on a two-dimensional image according to an embodiment of the present disclosure.
[0046] FIG6E is a fourth schematic diagram of a three-dimensional model obtained based on a two-dimensional image according to an embodiment of the present disclosure.
[0047] FIG7 is a schematic structural diagram of a model training device provided in the third aspect of an embodiment of the present disclosure.
[0048] FIG8 is another structural diagram of the model training device provided in the third aspect of an embodiment of the present disclosure.
[0049] FIG9 is a schematic structural diagram of an apparatus for establishing a three-dimensional model provided in the fourth aspect of an embodiment of the present disclosure.
[0050] FIG10 is a flowchart of an implementation of a model training method provided in the fifth aspect of an embodiment of the present disclosure.
[0051] FIG11 is a schematic diagram of obtaining a feature sequence based on prompt information according to an embodiment of the present disclosure.
[0052] FIG12 is a schematic diagram of an initial model according to an embodiment of the present disclosure.
[0053] FIG13 is a flowchart of an implementation of a method for establishing a three-dimensional model provided in the sixth aspect of an embodiment of the present disclosure.
[0054] FIG14A is a first schematic diagram of a three-dimensional model obtained based on prompt information according to the sixth aspect of an embodiment of the present disclosure.
[0055] FIG14B is a second schematic diagram of a three-dimensional model obtained based on prompt information according to the sixth aspect of an embodiment of the present disclosure.
[0056] FIG14C is a third schematic diagram of a three-dimensional model obtained based on prompt information according to the sixth aspect of an embodiment of the present disclosure.
[0057] FIG14D is a fourth schematic diagram of a three-dimensional model obtained based on prompt information according to the sixth aspect of an embodiment of the present disclosure.
[0058] FIG14E is a fifth schematic diagram of a three-dimensional model obtained based on prompt information according to the sixth aspect of an embodiment of the present disclosure.
[0059] FIG14F is a sixth schematic diagram of a three-dimensional model obtained based on prompt information according to the sixth aspect of an embodiment of the present disclosure.
[0060] FIG15 is a schematic structural diagram of a model training device provided in the seventh aspect of an embodiment of the present disclosure.
[0061] FIG16 is a schematic structural diagram of an apparatus for establishing a three-dimensional model provided in the eighth aspect of an embodiment of the present disclosure.
[0062] FIG17 is a structural block diagram of an electronic device provided in the ninth aspect of an embodiment of the present disclosure. DETAILED DESCRIPTION
[0063] The present disclosure will be described in further detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.
[0064] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, circuits, etc. well known to those skilled in the art are not described in detail in order to highlight the main purpose of the present disclosure.
[0065] 3D reconstruction refers to the process of creating a 3D model of a 3D scene using a 2D image of the scene. Specifically, taking an indoor 3D scene as an example, the existing 3D modeling process is as follows:
[0066] Step 1: Structure estimation and layout estimation of indoor 3D scenes.
[0067] In some embodiments, the structure of the indoor 3D scene can be estimated based on the 2D image corresponding to the indoor 3D scene (e.g., estimating room outline information of the indoor 3D scene). Specifically, if the 2D image corresponding to the indoor 3D scene is a perspective image, then the structure of the indoor 3D scene can be estimated using an object detection algorithm; or, if the 2D image corresponding to the indoor 3D scene is a panoramic image, then the structure of the indoor 3D scene can be estimated using a regression algorithm.
[0068] Furthermore, the furniture layout of the 3D indoor scene can be estimated based on the 2D image corresponding to the 3D indoor scene. For example, the furniture contained in the 3D indoor scene and its attribute information (such as location information, size information, type information, etc.) can be estimated based on the 2D image corresponding to the 3D indoor scene.
[0069] Step 2: Based on the structure estimation results and layout estimation results of the indoor three-dimensional scene, a three-dimensional model corresponding to the indoor three-dimensional scene is established.
[0070] Specifically, the structure estimation result and the layout estimation result of the indoor three-dimensional scene can be directly combined to establish a three-dimensional model corresponding to the indoor three-dimensional scene.
[0071] However, the current methods for building three-dimensional models design independent algorithms for images captured by different types of cameras (such as pinhole cameras, panoramic cameras, etc.) and different types of tasks (such as room structure estimation, furniture layout estimation). These algorithms lack versatility and are difficult to expand to other tasks.
[0072] Specifically, most algorithms for building indoor three-dimensional scenes based on two-dimensional images decompose the problem into room structure estimation and furniture layout restoration, and then solve these two sub-problems separately.
[0073] Room structure estimation usually uses different algorithms depending on the image type: (1) When the input image is a perspective image, an object detection algorithm is usually used to achieve room structure estimation; (2) When the input image is a panoramic image, room structure estimation can usually be achieved by regressing room corner points and wall contours.
[0074] For the sub-problem of restoring furniture layout, related technologies first detect furniture from images, and then use multimodal retrieval or 3D reconstruction to restore individual furniture.
[0075] Therefore, to address the above-mentioned issues, the first aspect of the present disclosure proposes a model training method. Based on the sequence prediction method, a universal three-dimensional scene modeling algorithm framework is implemented. The framework is applicable to images captured by different types of cameras (such as pinhole cameras, panoramic cameras, etc.) and to different modeling targets (such as room structures and furniture layouts). FIG1 is a schematic diagram of an application scenario according to an embodiment of the present disclosure. As shown in FIG1 , the application scenario includes: a server 110 and a terminal device 120.
[0076] The server 110 can be used to train the initial model to obtain a target model; and deploy the target model to the terminal device 120. The terminal device 120 can use the internally deployed target model and the image feature sequence corresponding to the three-dimensional scene to be modeled to determine the target prediction sequence corresponding to the three-dimensional scene to be modeled; and use the target prediction sequence to determine the three-dimensional model corresponding to the three-dimensional scene to be modeled. The server 110 and the terminal device 120 can be connected via any type of wired or wireless network. In addition, in some embodiments, the server 110 can include an independent physical server, a server cluster consisting of multiple physical servers, a distributed system, and a distribution network (Content Delivery Network, CDN) that can provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, and security services; the terminal device 120 can include an electronic device used by the user, such as a personal computer, mobile phone, tablet computer, notebook, and e-book reader, etc., a computer device with certain computing capabilities.
[0077] It should be noted that the above application scenarios are only provided to facilitate understanding of the spirit and principles of the present disclosure, and the embodiments of the present disclosure are not limited in this respect. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.
[0078] Figure 2 is a flowchart of the implementation of the training method of the model provided by the first aspect of the embodiment of the present disclosure. The method can be applied to the training device of the model. For example, the device can be deployed in a terminal or server or other processing device in a single machine, multi-machine or cluster system. Among them, the terminal can be a user equipment (UE, User Equipment), a mobile device, a personal digital assistant (PDA, Personal Digital Assistant) handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementations, the method can also be implemented by a processor calling a computer-readable instruction stored in a memory. As shown in Figure 2, the training method of the model includes S210 and S220.
[0079] In S210, the target prediction sequence corresponding to the sample three-dimensional scene is determined using the initial model and the image feature sequence corresponding to the sample three-dimensional scene; wherein the target prediction sequence is used to establish a three-dimensional model corresponding to the sample three-dimensional scene, and the image feature sequence is determined based on the two-dimensional image corresponding to the sample three-dimensional scene.
[0080] In S220 , the target prediction sequence and the target label sequence of the sample three-dimensional scene are used to adjust the initial model to obtain a target model.
[0081] Figures 3A and 3B are schematic diagrams of a two-dimensional image according to an embodiment of the present disclosure. As shown in Figures 3A and 3B, the two-dimensional image may include a two-dimensional view corresponding to a sample three-dimensional scene.
[0082] In one example, an image feature sequence corresponding to a sample 3D scene (i.e., an image feature sequence corresponding to the 2D image) can be determined based on a pre-trained model and a 2D image corresponding to the sample 3D scene. Pre-trained models include a Contrastive Language-Image Pre-training (CLIP) model, a Masked AutoEncoder (MAE), and a DINO model. Furthermore, the sample 3D scene can include an indoor 3D scene. The initial model can include a Sequence-to-Sequence (Seq2Seq)-based model, i.e., a model used for 3D modeling before or during training. The target model is a trained model for 3D modeling obtained after training the initial model.
[0083] Based on the model training method proposed in the embodiments of the present disclosure, the trained target model can determine a target prediction sequence corresponding to the 3D scene based on the image feature sequence corresponding to the 3D scene, and use this target prediction sequence to build a 3D model corresponding to the 3D scene. In other words, a sequence prediction method can be used to predict the target sequence based on the image feature sequence, thereby achieving 3D reconstruction (or 3D modeling), and thus is independent of the type of image or the type of object to be modeled.
[0084] It can be understood that since the purpose of three-dimensional modeling can be achieved based on the sequence prediction method, before training the initial model, the target label sequence corresponding to the sample three-dimensional scene can also be obtained to achieve the training of the initial model. That is, before training the initial model, the target label sequence can also be determined based on the sample three-dimensional scene, wherein the target label sequence includes a real sequence corresponding to the sample three-dimensional scene, and the real sequence is used to establish a real three-dimensional model corresponding to the sample three-dimensional scene.
[0085] Specifically, determining the target label sequence based on the sample three-dimensional scene can include: determining an initial label sequence based on the sample three-dimensional scene; wherein the initial label sequence includes a first initial label sequence and / or a second initial label sequence; and, sequentially performing discretization processing and feature embedding processing on the initial label sequence to obtain a target label sequence.
[0086] Specifically, the initial label sequence can be used to represent all and / or some features of the sample 3D scene. In one example, a target label sequence corresponding to the sample 3D scene can be determined based on the sample 3D scene and a domain-specific language (DSL) corresponding to the sample 3D scene. Specifically, taking an indoor 3D scene as an example, an initial label sequence corresponding to the indoor 3D scene can be determined based on the indoor 3D scene and the DSL corresponding to the indoor 3D scene.
[0087] The DSL may include a language dedicated to describing a specific domain. Compared with a common cross-domain general language (such as Java), the DSL can only be applied to a specific domain, but the DSL can accurately express the characteristic information of the specific domain.
[0088] Therefore, the initial label sequence determined based on the sample 3D scene and the DSL corresponding to the sample 3D scene can reflect most of the feature information contained in the sample 3D scene, thereby improving the accuracy of the subsequent 3D model obtained based on the target model.
[0089] It should be noted that since DSL may include a custom language focusing on a specific field, before determining the initial label sequence based on the sample 3D scene and the DSL corresponding to the sample 3D scene, the DSL corresponding to the sample 3D scene may also be determined.
[0090] Typically, the DSL for a certain field may include all and / or part of the feature information used to describe the field. For example, as shown in FIG3A and FIG3B , taking the sample three-dimensional scene as an indoor three-dimensional scene as an example, the structural information (such as contour information) and layout information (such as furniture information) corresponding to the indoor three-dimensional scene may be used to describe the indoor three-dimensional scene. Then, it can be determined that the DSL corresponding to the indoor three-dimensional scene may include information for describing the structure (such as contour) of the indoor three-dimensional scene and / or information for describing the layout (such as furniture) of the indoor three-dimensional scene. For this reason, taking the sample three-dimensional scene as an indoor three-dimensional scene as an example, based on the sample three-dimensional scene and the sample three-dimensional scene DSL, the determined initial label sequence may include a sequence for describing the structure (such as contour) of the indoor three-dimensional scene; and / or, a sequence for describing the layout (such as furniture) corresponding to the indoor three-dimensional scene. That is, the initial label sequence can be expressed by formula (1): S = (g, q) Formula (1)
[0091] Among them, g can represent a sequence used to describe the structure (such as outline) of the indoor three-dimensional scene; q can represent a sequence used to describe the layout (such as furniture) of the indoor three-dimensional scene; S can represent the initial label sequence corresponding to the indoor three-dimensional scene.
[0092] After determining the DSL corresponding to the sample three-dimensional scene and the content contained in the initial label sequence (i.e., the sequence used to describe the structure of the indoor three-dimensional scene and / or the sequence used to describe the layout of the indoor three-dimensional scene), the specific content contained in the sequence used to describe the structure of the indoor three-dimensional scene and / or the sequence used to describe the layout of the indoor three-dimensional scene (for example, the various elements contained in the sequence used to describe the layout of the indoor three-dimensional scene) can also be determined to determine the initial label sequence.
[0093] For this reason, determining the initial label sequence based on the sample three-dimensional scene may include: determining the M vertices contained in the sample three-dimensional scene, and determining a first initial label sequence corresponding to the sample three-dimensional scene based on the first coordinate information of the M vertices; M is a positive integer; and / or determining the N target objects contained in the sample three-dimensional scene, and determining the target object feature sequences corresponding to the N target objects respectively; determining a second initial label sequence corresponding to the sample three-dimensional scene based on the N target object feature sequences corresponding to the target objects; N is a positive integer.
[0094] Taking the example that the sample three-dimensional scene includes an indoor three-dimensional scene, the first initial label sequence may include a sequence (i.e., g) used to describe the structure (such as the outline) of the indoor three-dimensional scene; the second initial label sequence may include a sequence (i.e., q) used to describe the layout (such as furniture) of the indoor three-dimensional scene; the M vertices contained in the sample three-dimensional scene may include the corners of the sample three-dimensional scene; and the N target objects contained in the sample three-dimensional scene may include the furniture contained in the sample three-dimensional scene.
[0095] According to the embodiments of the present disclosure, components contained in a three-dimensional scene (such as vertices and target objects) can be represented in the form of a sequence, which is beneficial for reducing the training complexity of the target model and improving the accuracy of the three-dimensional model obtained based on the target model.
[0096] In addition, in order to better illustrate the specific steps of using the sample three-dimensional scene and the DSL corresponding to the sample three-dimensional scene to determine the initial label sequence (such as the first initial label sequence and / or the second initial label sequence) corresponding to the sample three-dimensional scene, the following content will combine Figure 4 to explain in detail how to use the sample three-dimensional scene and the DSL corresponding to the sample three-dimensional scene to determine the initial label sequence corresponding to the sample three-dimensional scene, that is, how to use the sample three-dimensional scene to determine the first initial label sequence and the second initial label sequence.
[0097] Figure 4 is a schematic diagram of determining a target label sequence corresponding to a sample three-dimensional scene according to an embodiment of the present disclosure. As shown in Figure 4, the following steps 1.1 and 1.2 may be used to determine a first initial label sequence corresponding to the sample three-dimensional scene.
[0098] In step 1.1, the M vertices contained in the sample three-dimensional scene are determined.
[0099] In step 1.2, a first initial label sequence corresponding to the sample three-dimensional scene is determined based on the first coordinate information of the M vertices.
[0100] Specifically, determining a first initial label sequence corresponding to the sample three-dimensional scene based on the first coordinate information of the M vertices may include: using the first coordinate information of the M vertices to determine a first arrangement order corresponding to the M vertices; and determining the first initial label sequence based on the first coordinate information of the M vertices and the first arrangement order. The first arrangement order corresponding to the M vertices may be used to represent positions of the first coordinate information of the M vertices in the first initial label sequence.
[0101] Taking the sample 3D scene as an indoor 3D scene as an example, the outline of the indoor 3D scene can be determined based on a polygonal wall. When the outline of the indoor 3D scene is a polygonal wall, multiple vertices (such as corners) included in the indoor 3D scene are determined, and a first initial label sequence is determined based on the first coordinate information of the multiple vertices (such as corners). The first coordinate information of the vertex (such as the corner) can be determined based on the camera shooting position and orientation of the 2D image corresponding to the sample 3D scene.
[0102] Specifically, the first initial tag sequence can be determined based on formula (2-1): g=(g1, g2, ..., g i ,…,g M ) Formula (2-1)
[0103] Wherein, g can represent the first initial label sequence corresponding to the indoor three-dimensional scene; g i It can represent the first coordinate information of the i-th vertex (such as a corner) contained in the indoor three-dimensional scene.
[0104] Furthermore, if a right-handed coordinate system is established with the optical center of the camera of the two-dimensional image corresponding to the sample three-dimensional scene as the coordinate origin, the camera optical axis direction as the longitudinal axis (Y axis), and the negative direction of gravity as the vertical axis (Z axis), then in this case, (x i ,y i ,z i ) represents the first coordinate information of the i-th vertex in the two-dimensional image, where x i Can represent the x-axis coordinate of the i-th vertex; y i Can represent the vertical coordinate of the i-th vertex; z i It can represent the z-axis coordinate of the i-th vertex.
[0105] It should be noted that, in the training of the model used for three-dimensional modeling, the z-axis coordinates of each vertex (such as a corner) have little influence on the training of the three-dimensional model. Therefore, in order to reduce the training complexity and training resources required for the model used for three-dimensional modeling, the first coordinate information of each vertex can optionally include only the x-axis coordinates and y-axis coordinates of each vertex, that is, the first coordinate information of the i-th vertex in the two-dimensional image can be expressed by formula (2-2): g i =(x i ,y i ) Formula (2-2)
[0106] In addition, the first arrangement order corresponding to the M vertices (such as corners) included in the sample three-dimensional scene may be determined according to the first coordinate information of the M vertices.
[0107] For example, assuming a sample three-dimensional scene includes three corners, namely, a first corner G1, a second corner G2, and a third corner G3; and in a right-handed coordinate system established with the optical center of the camera of the two-dimensional image corresponding to the sample three-dimensional scene as the coordinate origin, the camera optical axis as the longitudinal axis (Y axis), and the negative gravity direction as the vertical axis (Z axis), the first coordinate information of corner G1 is g1, the first coordinate information of corner G2 is g2, and the first coordinate information of corner G3 is g3. Corners G1, G2, and G3 can first be sorted clockwise in this coordinate system to obtain a clockwise order corresponding to corners G1, G2, and G3, such as (G1, G2, G3); then, the first corner corresponding to the left edge of the camera can be determined based on the camera center position and / or orientation, and the initial arrangement order can be adjusted based on the first corner corresponding to the left edge of the camera to obtain a first arrangement order. For example, if the first corner corresponding to the left boundary of the camera is determined to be G2 based on the camera center position and / or orientation, then the first arrangement order corresponding to corner G2 can be determined as: the first coordinate information of corner G2 is in the first position in the first initial label sequence; and based on the positional relationship between corner G2 and corner G1 and corner G3 in the clockwise sequence, the first arrangement order corresponding to corner G1 and corner G3 is determined, so that the first initial label sequence can be determined based on the first arrangement order corresponding to corner G1, corner G2 and corner G3, such as (g2, g3, g1).
[0108] Using the coordinate information of each vertex to determine the first initial label sequence can, to a certain extent, unify the method of determining the first initial label sequence corresponding to multiple sample three-dimensional scene images, thereby avoiding the increased complexity of the target model training process due to the different methods of determining the first initial label sequences corresponding to multiple sample three-dimensional scene images.
[0109] In addition, not only the first initial label sequence corresponding to the sample 3D scene can be determined, but also the second initial label sequence corresponding to the sample 3D scene can be determined. Still as shown in FIG4 , the following steps 2.1, 2.2, and 2.3 can be used to determine the second initial label sequence corresponding to the sample 3D scene.
[0110] In step 2.1, N target objects contained in the sample three-dimensional scene are determined.
[0111] In step 2.2, target object feature sequences corresponding to the N target objects are determined respectively.
[0112] Specifically, determining a target object feature sequence corresponding to a target object may include: determining attribute information of the target object; wherein the attribute information includes at least one of the second coordinate information, category, size information, angle information, and identification information of the target object, and the identification information is used to characterize the target object; and, determining the target object feature sequence using the attribute information of the target object.
[0113] Each target object may correspond to unique identification information, and the identification information may be used to indicate a three-dimensional model corresponding to the target object.
[0114] Taking the sample 3D scene as an indoor 3D scene as an example, the target object feature sequence corresponding to the target object in the indoor 3D scene can be determined based on the attribute information of the target object (such as furniture). For example, if the attribute information can include at least one of the second coordinate information p, the category c, the size information l, the angle information r, and the identification information f of the target object, then formula (3) can be used to express the determination of the target object Q in the sample 3D scene. j Corresponding target object feature sequence: q j =(c j ,l j ,p j ,r j ,f j ) Formula (3)
[0115] Among them, q j It can represent the j-th target object Q in the indoor three-dimensional scene j The corresponding target object feature sequence; c j It can represent the j-th target object Q in the indoor three-dimensional scene j Category of j It can represent the j-th target object Q in the indoor three-dimensional scene j Size information of p j It can represent the j-th target object Q in the indoor three-dimensional scene j The second coordinate information of r j It can represent the j-th target object Q in the indoor three-dimensional scene j Angle information of f j It can represent the j-th target object Q in the indoor three-dimensional scene j identification information.
[0116] Furthermore, if a right-handed coordinate system is established with the optical center of the camera of the two-dimensional image corresponding to the sample three-dimensional scene as the coordinate origin, the camera optical axis direction as the longitudinal axis (Y axis), and the negative direction of gravity as the vertical axis (Z axis), then in this case, the j-th target object Qj The second coordinate information p j You can use (x Q ,y Q ,z Q ) means (where x Q Can represent the target object Q j The x-axis coordinate of Q Can represent the target object Q j The y-axis coordinate of Q It can represent the z-axis coordinate of the target object Q). In addition, the j-th target object Q j Size information j The target object Q j At least one of the length, height and width of the j-th target object Q j Angle information r j Can include the j-th target object Q j The rotation angle relative to the horizontal plane (such as the ground); the j-th target object Q j Category c j It can include doors, windows, beds, wardrobes, lamps, bedside tables, curtains, decorative paintings, carpets, desks, dressing tables, chairs and / or TV cabinets, etc.; the jth target object Q j Identification information f j Can be used to characterize the target object Q j The corresponding three-dimensional model.
[0117] Using the target object's attribute information to determine the target object feature sequence corresponding to the target object can, to a certain extent, avoid the negative impact on the subsequent training of the target three-dimensional model caused by the target object feature sequence containing too little information.
[0118] In step 2.3, a second initial label sequence corresponding to the sample three-dimensional scene is determined based on the N target object feature sequences corresponding to the target objects.
[0119] Specifically, determining the second initial label sequence based on N target object feature sequences corresponding to the target objects may include: determining a second arrangement order corresponding to the N target objects using at least one of second coordinate information, categories, and occurrence frequencies of the N target objects in the sample three-dimensional scene; and determining the second initial label sequence based on the second arrangement order and the N target object feature sequences corresponding to the target objects. The second arrangement order may be used to characterize the position of the target feature sequence corresponding to the target object in the second initial label sequence.
[0120] Taking the sample 3D scene as an indoor 3D scene as an example, the layout information of the indoor 3D scene can be determined based on multiple target objects (such as furniture) in the indoor 3D scene. Then, the second initial label sequence can be determined based on the multiple target objects (such as furniture) contained in the indoor 3D scene and the target object feature sequences corresponding to each target object. For example, if the indoor 3D scene includes N pieces of furniture, the second initial label sequence can be expressed by formula (4): q = (q1, q2, ..., q j ,…,q N ) Formula (4)
[0121] Wherein, q may represent the second initial label sequence corresponding to the indoor three-dimensional scene; q j It can represent the j-th target object Q in the indoor three-dimensional scene j The corresponding target object feature sequence.
[0122] In addition, the second arrangement order corresponding to each target object (such as furniture) included in the sample three-dimensional scene may be determined based on at least one of the second coordinate information, category, and occurrence frequency of each target object.
[0123] For example, if the sample three-dimensional scene contains furniture Q1, furniture Q2, and furniture Q3, where furniture Q1 belongs to the category of bed, furniture Q2 belongs to the category of door, and furniture Q3 belongs to the category of carpet. Then, if the predetermined second arrangement order corresponding to each target object includes: in the second initial label sequence, the target object feature sequence corresponding to the furniture belonging to the category of bed is before the target object feature sequence corresponding to the furniture belonging to the category of door, the target object feature sequence corresponding to the furniture belonging to the category of bed is before the target object feature sequence corresponding to the furniture belonging to the category of carpet, and the target object feature sequence corresponding to the furniture belonging to the category of door is before the target object feature sequence corresponding to the furniture belonging to the category of carpet, then the second initial label sequence can be determined to be (q1, q2, q3).
[0124] Alternatively, for another example, the sample three-dimensional scene includes furniture Q1, furniture Q2, and furniture Q3, where the frequency of occurrence corresponding to furniture Q1 is w1, the frequency of occurrence corresponding to furniture Q2 is w2, and the frequency of occurrence corresponding to furniture Q3 is w3; then, by comparing the frequencies of occurrence corresponding to furniture Q1, furniture Q2, and furniture Q3, the second arrangement order corresponding to each piece of furniture can be determined, and based on the second arrangement order corresponding to each piece of furniture, the second initial label sequence can be determined. Specifically, the second arrangement order corresponding to each piece of furniture Q1, furniture Q2, and furniture Q3 can be determined based on the ascending (or descending) order of the frequencies of occurrence corresponding to furniture Q1, furniture Q2, and furniture Q3. For example, if the ascending order of the occurrence frequencies corresponding to the furniture Q1, furniture Q2, and furniture Q3 is "w1, w2, w3", then it can be determined that the second arrangement order corresponding to the furniture Q1, furniture Q2, and furniture Q3 may include: in the second initial label sequence, the target object feature sequence corresponding to furniture Q1 is before the target object feature sequence corresponding to furniture Q2, the target object feature sequence corresponding to furniture Q2 is before the target object feature sequence corresponding to furniture Q3, and the target object feature sequence corresponding to furniture Q1 is before the target object feature sequence corresponding to furniture Q3, then the second initial label sequence can be determined to be (q1, q2, q3).
[0125] Or, for another example, the sample three-dimensional scene includes furniture Q1, furniture Q2 and furniture Q3, and when a right-handed coordinate system is established with the camera optical center of the two-dimensional image corresponding to the sample three-dimensional scene as the coordinate origin, the camera optical axis direction as the longitudinal axis (Y axis), and the negative direction of gravity as the vertical axis (Z axis), the second coordinate information p1 of the furniture Q1 is (x1, y1, z1) (that is, the x-axis coordinate of the furniture Q1 is x1), the second coordinate information p2 of the furniture Q2 is (x2, y2, z2) (that is, the x-axis coordinate of the furniture Q2 is x2), and the second coordinate information p3 of the furniture Q3 is (x3, y3, z3) (that is, the x-axis coordinate of the furniture Q3 is x3). Then, the second arrangement order corresponding to each piece of furniture can be determined by comparing the x-axis coordinates of furniture Q1, furniture Q2, and furniture Q3 (or by comparing the y-axis coordinates of furniture Q1, furniture Q2, and furniture Q3, or by comparing the z-axis coordinates of furniture Q1, furniture Q2, and furniture Q3), and the second initial label sequence can be determined based on the second arrangement order corresponding to each piece of furniture. Specifically, the position of the target object feature sequence corresponding to each piece of furniture in the second initial label sequence can be determined based on the ascending order (or descending order) of the x-axis coordinates of furniture Q1, furniture Q2, and furniture Q3. For example, if the x-axis coordinates of furniture Q1, furniture Q2, and furniture Q3 are arranged in ascending order as "x1, x2, x3", then it can be determined that the second arrangement order corresponding to furniture Q1, furniture Q2, and furniture Q3 may include: in the second initial label sequence, the target object feature sequence corresponding to furniture Q1 is before the target object feature sequence corresponding to furniture Q2, the target object feature sequence corresponding to furniture Q2 is before the target object feature sequence corresponding to furniture Q3, and the target object feature sequence corresponding to furniture Q1 is before the target object feature sequence corresponding to furniture Q3. Then, the second initial label sequence can be determined to be (q1, q2, q3).
[0126] The second initial label sequence is determined based on at least one of the second coordinate information, category and occurrence frequency of each target object, which can achieve the purpose of sorting the feature sequences of each target object in the second initial label sequence, thereby reducing the complexity, required resources and time required for training the model used for three-dimensional modeling.
[0127] It should be noted that the above order of determining the first initial tag sequence and the second initial tag sequence is only an example, and the present disclosure does not limit the specific order of determining the first initial tag sequence and the second initial tag sequence.
[0128] It should be noted that, in some embodiments, after the initial tag sequence is determined, the target tag sequence may be determined using the initial tag sequence.
[0129] Specifically, still as shown in FIG4 , the following steps 3.1 and 3.2 may be used to determine the target tag sequence.
[0130] In step 3.1, the initial tag sequence S (for example, the first initial tag sequence or the second initial tag sequence) is discretized to obtain a tag sequence S′ after discretization. The discretized sequence S′ contains each element s′ i or s′ j All are discrete values.
[0131] Specifically, we can first determine the elements s contained in the initial tag sequence S i or s j Whether it is a discrete value; the initial label sequence S contains the element s i or s j Can be divided into the following six categories: Elements s i g is included in the first initial tag sequence i ; element s j c is included in the second initial tag sequence j , the second initial tag sequence contains l j , the second initial tag sequence contains p j , the second initial tag sequence contains r j , and / or the second initial tag sequence contains f j wait.
[0132] Then, if any element s contained in the initial tag sequence S i or s j is a discrete value, then the element s can be directly i or s j Determine the element s' in the sequence S' i or s′ j , and the position of the element in the initial label sequence S is the same as that in the sequence S′; or, if any element s contained in the initial label sequence S i or s j is a continuous value, then the element s i or s j Perform discretization to obtain the discretized element s′ i or s′ j , and the discretized element s′ i or s′ j Save to sequence S', and element s i or s j The position of S in the initial label sequence corresponds to the discretized element s' i or s′ jThe positions in the sequence S' are the same.
[0133] In step 3.2, feature embedding is performed on the discretized label sequence S′ to obtain the target label sequence E(S′).
[0134] Specifically, each element s′ in the discretely processed tag sequence S′ can be converted into i or s′ j Perform word embedding processing to obtain each element s′ i or s′ j The corresponding feature embedding E(s′ i ) or E(s′ j ), and use formula (5-2) to determine the target tag sequence E(S′): E(s′ i )=E value (s′ i )+E type (s′ i )+E pos (s′ i ) Formula (5-1-1) E(s′) j )=E value (s′ j )+E type (s′ j )+E pos (s′ j ) Formula (5-1-2) E(S′)=[E(s′1),…,E(s′ i )…,E(s′ M ),E(s′1),…,E(s′ j )…,E(s′ N )] Formula (5-2)
[0135] Among them, E value (s′ i ) can represent element s′ i The corresponding value embedding, E type (s′ i ) can represent element s′ i Type embedding, E pos (s′ i ) represents the element s′ i The relative position in the first initial tag sequence; E value (s′ j ) can represent element s′ j The corresponding value embedding, E type (s′ j ) can represent element s′j Type embedding, E pos (s′ j ) represents the element s′ j The relative position in the second initial tag sequence.
[0136] It should be noted that, when the initial label sequence includes the first initial label sequence and the second initial label sequence, the above formula (5-2) is used to determine the target label sequence E(S′); when the initial label sequence includes only one of the first initial label sequence and the second initial label sequence, the part of the above formula (5-2) corresponding to the other of the first initial label sequence and the second initial label sequence can be empty.
[0137] In one example, in element s′ i g is included in the first initial tag sequence i , or s′ j c is included in the second initial tag sequence j , the second initial tag sequence contains l j , the second initial tag sequence contains p j Or the second initial tag sequence contains r j When, we can use the element s′ i or s′ j The corresponding eigenvector determines E value (s′ i ) or E value (s′ j ); in the element s′ j f is included in the second initial tag sequence j When the element s′ j The corresponding image feature determines E value (s′ j ). Get the element s′ j The corresponding image features may include: inputting a 2D rendering of the 3D model of the target object (e.g., furniture) into a pre-trained model, such as a Contrastive Language-Image Pre-training (CLIP) model, and having the CLIP model output a value that corresponds to the element s′. j The corresponding image features can be obtained by comparing with the element s′ i or s′ j The corresponding six categories determine E type (s′ i ) or E type (s′ j ), that is, the first initial tag sequence contains g i , the second initial tag sequence contains c j, the second initial tag sequence contains l j , the second initial tag sequence contains p j , the second initial tag sequence contains r j and the second initial tag sequence contains f j These six categories can be classified according to the element s′ i The position in the first initial tag sequence is determined by E pos (s′ i ), can be calculated based on the element s′ j The position in the second initial tag sequence is determined by E pos (s′ j ).
[0138] The above describes how to determine the target label sequence and image feature sequence corresponding to the sample 3D scene. It should be noted that after determining the target label sequence and image feature sequence, the target label sequence and image feature sequence can also be used to train the initial 3D model to obtain the target 3D model.
[0139] Therefore, how to train the initial 3D model based on the target label sequence and the image feature sequence is also a problem considered in the embodiments of the present disclosure. Specifically, the following will explain in detail how to train the initial 3D model using the target label sequence and the image feature sequence.
[0140] FIG5 is a schematic block diagram of an initial model according to an embodiment of the present disclosure. As shown in FIG5 , the initial model may include a transformer encoding layer and a transformer decoding layer.
[0141] The Transformer encoding layer is configured to encode the image feature sequence to obtain semantic features corresponding to the image feature sequence.
[0142] The Transformer decoding layer is configured to use semantic features and historical prediction sequences corresponding to the sample three-dimensional scene to determine a first prediction sequence; wherein the first prediction sequence is used to determine a target prediction sequence.
[0143] The Transformer encoding layer and Transformer decoding layer included in this initial model can both be composed of 6 Transformer modules.
[0144] In one example, the target prediction sequence may include the prediction sequence of the initial model for the sample three-dimensional scene at the first moment, and the historical prediction sequence may include the prediction sequence of the initial model for the sample three-dimensional scene at the first time series; wherein the first moment is after the first time series.
[0145] For example, the first time series includes moments within the time period [tk, t), for example, the first time series includes moments tk, t-k+1, t-k+2, ..., t-1, and the first moment may include moment t.
[0146] Based on this, if the image feature sequence is input into the initial model, the decoding layer of the initial model can receive the semantic features of the image feature sequence output by the encoding layer at time t, as well as the output sequence (i.e., historical prediction sequence) E(s) of the decoding layer at time tk, t-k+1, t-k+2, ..., and t-1 for the image feature sequence. <t ); The decoding layer of the initial model receives the semantic features of the image feature sequence output by the encoding layer at time t and the output sequence (i.e., historical prediction sequence) E(s) of the encoding layer at the previous k moments (i.e., time tk, t-k+1, t-k+2, ..., and t-1) for the image feature sequence. <t ) can be used to obtain the semantic features of the image feature sequence at time t and the output sequence E(s) of the previous k moments. <t ), determine the first prediction sequence corresponding to the sample three-dimensional scene, that is, the hidden feature h corresponding to the sample three-dimensional scene t Thus, the initial model can further determine the target prediction sequence based on the first prediction sequence.
[0147] Transformer has the ability to model long-range dependencies, enabling the initial model to simultaneously consider global information. Therefore, when the initial model includes a Transformer encoding layer and a Transformer decoding layer, the efficiency of training the initial model can be improved, and the initial model has advantages such as strong expressive power, parallel computing, long-range dependency modeling, multi-head attention mechanism, and interpretability, thereby improving the performance and accuracy of the initial model. In addition, because the target prediction sequence is determined using multiple intermediate values predicted by the initial model (i.e., the historical prediction sequence), it can reduce the error caused by a single prediction and improve the prediction accuracy of the target model.
[0148] Furthermore, as still shown in FIG5 , the initial model may also include a linear layer and a normalization layer.
[0149] The linear layer is configured to adjust the feature dimension of the first prediction sequence to obtain an adjusted first prediction sequence.
[0150] The normalization layer is configured to perform normalization processing on the adjusted first prediction sequence according to the image feature sequence corresponding to the sample three-dimensional scene to obtain a second prediction sequence; wherein the second prediction sequence is used to determine the target prediction sequence.
[0151] For example, still taking the first time series including time tk, t-k+1, t-k+2, ..., and t-1, and the first time including time t as an example, after obtaining the first prediction sequence corresponding to the sample three-dimensional scene, that is, the hidden feature h corresponding to the sample three-dimensional scene t Afterwards, the second prediction sequence corresponding to the sample three-dimensional scene can be determined using formula (6): p(s t |s <t ,I)=p(E(s t )|E(s <t ),F)=softmax(linear(h t )) Formula (6)
[0152] Where I represents the two-dimensional image corresponding to the sample three-dimensional scene, F can represent the image feature sequence corresponding to the two-dimensional image I (that is, the image feature sequence corresponding to the sample three-dimensional scene); p(s t |s <t ,I) can represent the second prediction sequence corresponding to the sample three-dimensional scene.
[0153] That is to say, we can first use the linear layer of the initial model to adjust the hidden feature h t to obtain the adjusted hidden features; then, the normalization layer of the initial model can be used to process the hidden features after the dimension adjustment to obtain a second prediction sequence corresponding to the sample three-dimensional scene and an occurrence probability corresponding to the second prediction sequence.
[0154] Furthermore, after determining the second prediction sequence corresponding to the sample three-dimensional scene, the target prediction sequence can be determined using formula (7): p(S′|I)=∏ t p(s t |s <t ,I)=∏ t p(E(s t )|E(s <t ),F) Formula (7)
[0155] Wherein, I represents the two-dimensional image corresponding to the sample three-dimensional scene, F represents the image feature sequence corresponding to the two-dimensional image I; p(S′|I) can represent the target prediction sequence corresponding to the sample three-dimensional scene.
[0156] It should be noted that after determining the target prediction sequence corresponding to the sample three-dimensional scene, the target prediction sequence and the target label sequence of the sample three-dimensional scene can be used to adjust the initial model to obtain the target model.
[0157] Specifically, the initial model can be adjusted based on the cross-entropy loss function according to the target prediction sequence and the target label sequence corresponding to the sample three-dimensional scene to obtain the target model.
[0158] Among them, the value of the loss function is related to the gap between the target prediction sequence and the target label sequence. The larger the gap between the two, the larger the absolute value of the corresponding loss function. Determining the value of the loss function in this way can accelerate the convergence of the model and improve the training speed of the model.
[0159] The above method for determining the loss function is only an example, and other methods for determining the loss function may also be used. The embodiments of the present disclosure do not limit the method for determining the loss function.
[0160] The second aspect of the embodiment of the present disclosure proposes a method for establishing a three-dimensional model. FIG6A is a flowchart of the implementation of the method for establishing a three-dimensional model provided by the second aspect of the embodiment of the present disclosure, including: S610 and S620.
[0161] In S610, a target prediction sequence corresponding to the three-dimensional scene to be modeled is determined using the target model and the image feature sequence corresponding to the three-dimensional scene to be modeled; wherein the image feature sequence is determined based on the two-dimensional image corresponding to the three-dimensional scene to be modeled, and the target model is trained using the model training method provided by the first aspect of the embodiment of the present disclosure.
[0162] In S620 , a three-dimensional model corresponding to the three-dimensional scene to be modeled is established according to the target prediction sequence.
[0163] The two-dimensional image may include a panoramic view and / or a perspective view. As shown in Figures 6B and 6C, when the two-dimensional image includes a panoramic view, the image feature sequence corresponding to the three-dimensional scene to be modeled is input into the target model, and the target model can output a target prediction sequence for the three-dimensional scene to be modeled, and use the target prediction sequence to determine the three-dimensional model corresponding to the indoor three-dimensional scene reconstruction task; and, as shown in Figures 6D and 6E, when the two-dimensional image includes a perspective view, the image feature sequence corresponding to the three-dimensional scene to be modeled is input into the target model, and the target model can output a target prediction sequence for the three-dimensional scene to be modeled, and use the target prediction sequence to determine the three-dimensional model corresponding to the indoor three-dimensional scene reconstruction task.
[0164] The method for establishing a three-dimensional model proposed in the embodiment of the present disclosure can achieve the purpose of three-dimensional prediction based on the sequence prediction method, thereby reducing the complexity of three-dimensional modeling, improving the efficiency of three-dimensional modeling, and improving the accuracy of the three-dimensional model obtained by modeling.
[0165] The target model can include Transformer encoding layers and Transformer decoding layers.
[0166] The Transformer encoding layer is configured to process the image feature sequence to obtain semantic features corresponding to the image feature sequence.
[0167] The Transformer decoding layer is configured to determine a first prediction sequence using semantic features and a historical prediction sequence corresponding to the three-dimensional scene to be modeled; wherein the first prediction sequence is used to determine a target prediction sequence.
[0168] The Transformer's ability to model long-range dependencies enables the target model to simultaneously consider global information. Therefore, since the target model includes both the Transformer encoding and decoding layers, it improves the target model's prediction efficiency and offers advantages such as powerful expressiveness, parallel computing, long-range dependency modeling, multi-head attention mechanisms, and interpretability, thereby improving the target model's performance and accuracy.
[0169] The target model can also include linear layers and normalization layers.
[0170] The linear layer is configured to adjust the feature dimension of the first prediction sequence to obtain an adjusted first prediction sequence.
[0171] The normalization layer is configured to perform normalization processing on the adjusted first prediction sequence to obtain a second prediction sequence; wherein the second prediction sequence is used to determine the target prediction sequence.
[0172] In one example, the target prediction sequence may include a prediction sequence of the target model for the three-dimensional scene to be modeled at a second moment, and the historical prediction sequence may include a prediction sequence of the target model for the three-dimensional scene to be modeled at a second time series; wherein the second moment is after the second time series.
[0173] For example, the second time series includes moments within the time period [tk, t]. For example, the second time series may include moments tk, t-k+1, t-k+2, ..., t-1, and the second moment may include moment t.
[0174] It should be noted that the specific implementation method of using the target model and the image feature sequence corresponding to the three-dimensional scene to be modeled to determine the target prediction sequence corresponding to the three-dimensional scene to be modeled is similar to the specific implementation method of the model training method provided in the first aspect of the embodiment of the present disclosure using the initial model and the image feature sequence corresponding to the sample three-dimensional scene to determine the target prediction sequence corresponding to the sample three-dimensional scene. The embodiment of the present disclosure will not be repeated here.
[0175] The third aspect of the embodiment of the present disclosure also proposes a model training device. Figure 7 is a structural diagram of the model training device provided by the third aspect of the embodiment of the present disclosure. The model training device 700 includes: a first determination module 710 and an adjustment module 720.
[0176] The first determination module 710 is configured to use the initial model and the image feature sequence corresponding to the sample three-dimensional scene to determine the target prediction sequence corresponding to the sample three-dimensional scene; wherein the target prediction sequence is used to establish a three-dimensional model corresponding to the sample three-dimensional scene; and the image feature sequence is determined based on the two-dimensional image corresponding to the sample three-dimensional scene.
[0177] The adjustment module 720 is configured to adjust the initial model using the target prediction sequence and the target label sequence of the sample three-dimensional scene to obtain a target model.
[0178] FIG8 is another schematic diagram of the structure of the model training apparatus provided by the third aspect of the embodiments of the present disclosure. As shown in FIG8 , the model training apparatus 800 may further include: a second determination module 830 configured to determine a target label sequence based on the sample 3D scene, wherein the target label sequence includes a real sequence corresponding to the sample 3D scene, and the real sequence is used to establish a real 3D model corresponding to the sample 3D scene.
[0179] In some embodiments, the second determination module 830 is configured to determine an initial label sequence based on the sample three-dimensional scene; the initial label sequence includes a first initial label sequence and / or a second initial label sequence; and the initial label sequence is discretized and feature embedded in sequence to obtain a target label sequence.
[0180] In some embodiments, the second determination module 830 may include: a first determination submodule 831 and / or a second determination submodule 832 .
[0181] The first determination submodule 831 is configured to determine M vertices included in the sample three-dimensional scene, and determine a first initial label sequence corresponding to the sample three-dimensional scene according to first coordinate information of the M vertices; M is a positive integer.
[0182] The second determination submodule 832 is configured to determine the N target objects contained in the sample three-dimensional scene, and respectively determine the target object feature sequences corresponding to the N target objects, and determine the second initial label sequence corresponding to the sample three-dimensional scene based on the N target object feature sequences corresponding to the target objects; N is a positive integer.
[0183] In some embodiments, the first determination submodule 831 is configured to: use the first coordinate information of M vertices to determine the position of the first coordinate information of the M vertices in the first initial label sequence; and determine the first initial label sequence based on the first coordinate information of the M vertices and the position of the first coordinate information of the M vertices in the first initial label sequence.
[0184] In some embodiments, the second determination submodule 832 is configured to: determine the positions of N target object feature sequences corresponding to the target objects in the second initial label sequence using at least one of the second coordinate information, category, and occurrence frequency of the N target objects in the sample three-dimensional scene; and determine the second initial label sequence based on the N target object feature sequences and the positions of the N target object feature sequences in the second initial label sequence.
[0185] In some embodiments, the second determination submodule 832 is configured to: determine attribute information of the target object; wherein the attribute information includes at least one of the second coordinate information, category, size information, angle information and identification information of the target object, and the identification information corresponds to the target object; and, using the attribute information of the target object, determine a target object feature sequence corresponding to the target object.
[0186] In some embodiments, the initial model may include a Transformer encoding layer and a Transformer decoding layer.
[0187] The Transformer encoding layer is configured to encode the image feature sequence to obtain semantic features corresponding to the image feature sequence.
[0188] The Transformer decoding layer is configured to use semantic features and historical prediction sequences corresponding to the sample three-dimensional scene to determine a first prediction sequence; wherein the first prediction sequence is used to determine a target prediction sequence.
[0189] In some implementations, the initial model may further include a linear layer and a normalization layer.
[0190] The linear layer is configured to adjust the feature dimension of the first prediction sequence to obtain an adjusted first prediction sequence.
[0191] The normalization layer is configured to perform normalization processing on the adjusted first prediction sequence to obtain a second prediction sequence; wherein the second prediction sequence is used to determine the target prediction sequence.
[0192] In some embodiments, the target prediction sequence may include a prediction sequence of the initial model for the sample three-dimensional scene at a first moment, and the historical prediction sequence may include a prediction sequence of the initial model for the sample three-dimensional scene at a first time series; wherein the first moment is after the first time series.
[0193] For the specific functions and exemplary descriptions of the modules and sub-modules of the model training device provided in the third aspect of the embodiment of the present disclosure, please refer to the relevant descriptions of the corresponding steps in the model training method provided in the first aspect of the embodiment of the present disclosure, and the embodiment of the present disclosure will not be repeated here.
[0194] The fourth aspect of the embodiment of the present disclosure also proposes a device for establishing a three-dimensional model. Figure 9 is a structural schematic diagram of the device 900 for establishing a three-dimensional model provided by the fourth aspect of the embodiment of the present disclosure. The device 900 for establishing a three-dimensional model includes: a third determination module 910 and a modeling module 920.
[0195] The third determination module 910 is configured to determine a target prediction sequence corresponding to the three-dimensional scene to be modeled using the target model and the image feature sequence corresponding to the three-dimensional scene to be modeled; wherein the image feature sequence is determined based on the two-dimensional image corresponding to the three-dimensional scene to be modeled, and the target model is obtained by training using the model training device provided by the third aspect of the embodiment of the present disclosure.
[0196] The modeling module 920 is configured to establish a three-dimensional model corresponding to the three-dimensional scene to be modeled according to the target prediction sequence.
[0197] In some embodiments, the target model may include a Transformer encoding layer and a Transformer decoding layer.
[0198] The Transformer encoding layer is configured to process the image feature sequence to obtain semantic features corresponding to the image feature sequence.
[0199] The Transformer decoding layer is configured to determine a first prediction sequence using semantic features and a historical prediction sequence corresponding to the three-dimensional scene to be modeled; wherein the first prediction sequence is used to determine a target prediction sequence.
[0200] In some embodiments, the target model may further include a linear layer and a normalization layer.
[0201] The linear layer is configured to adjust the feature dimension of the first prediction sequence to obtain an adjusted first prediction sequence.
[0202] The normalization layer is configured to perform normalization processing on the adjusted first prediction sequence to obtain a second prediction sequence; wherein the second prediction sequence is used to determine the target prediction sequence.
[0203] In some embodiments, the target prediction sequence may include a prediction sequence of the target model for the three-dimensional scene to be modeled at a second moment, and the historical prediction sequence may include a prediction sequence of the target model for the three-dimensional scene to be modeled at a second time series; wherein the second moment is after the second time series.
[0204] For the specific functions and exemplary descriptions of the modules and submodules of the apparatus for establishing a three-dimensional model provided in the fourth aspect of the embodiment of the present disclosure, please refer to the relevant descriptions of the corresponding steps in the method for establishing a three-dimensional model provided in the second aspect of the embodiment of the present disclosure, and the embodiment of the present disclosure will not be repeated here.
[0205] Indoor three-dimensional scene modeling algorithms in related technologies are implemented using different algorithms for different applications.
[0206] Most applications for indoor 3D scene reconstruction use 3D object detection algorithms. These algorithms first identify the target objects in the scene through 3D object detection. Based on the detection results, they then use different types of 3D reconstruction representation algorithms to model the 3D structure of each target object, thereby reconstructing the indoor 3D scene.
[0207] Applications for indoor 3D scene generation are mostly based on Seq2Seq models. Seq2Seq models are typically used to map from one sequence to another. In scene generation applications, Seq2Seq models can effectively map input sequences (such as scene descriptions) to output sequences (such as scene representations), thereby enabling the generation of indoor 3D scenes.
[0208] Indoor 3D scene modeling algorithms are designed to solve different applications independently, which poses a significant challenge to their implementation. Because different applications may require different feature extraction, data processing, and / or model optimization, existing indoor 3D scene modeling algorithms often struggle to adapt to diverse needs. This wastes resources and time, and limits the algorithm's flexibility and scalability.
[0209] In order to solve the above technical problems, based on the same inventive concept as the first to fourth aspects of the embodiments of the present disclosure, the fifth aspect of the embodiments of the present disclosure proposes a model training method, which regards any prior knowledge about indoor three-dimensional scenes, such as two-dimensional images (for example, for three-dimensional scene reconstruction tasks), floor plans (for example, for three-dimensional scene generation tasks), three-dimensional models of partial target objects and / or bounding box information of partial target objects (for example, for scene completion tasks), or other forms of information (for example, text description information, or point cloud information), as prompts. These prompts are first represented by a series of word units, and then these word units are input into the Seq2Seq model to generate a unified representation of the indoor three-dimensional scene, so that the tasks of understanding and generating various indoor three-dimensional scenes can be unified into a single model.
[0210] FIG10 is a flowchart of an implementation of a model training method provided in the fifth aspect of an embodiment of the present disclosure. The model training method includes S1001 , S1002 and S1003 .
[0211] In S1001 , a feature sequence corresponding to the sample three-dimensional scene is determined according to prompt information corresponding to the sample three-dimensional scene.
[0212] In one embodiment, the prompt information may include a two-dimensional image corresponding to the sample three-dimensional scene, a floor plan corresponding to the sample three-dimensional scene, a three-dimensional model of the first target object contained in the sample three-dimensional scene, bounding box information of the second target object contained in the sample three-dimensional scene, text description information, point cloud information, or at least one of voxel information.
[0213] In one embodiment, as shown in FIG11 , S1001 determines a feature sequence based on prompt information, which may include at least one of the following.
[0214] In a case where the prompt information includes a two-dimensional image, feature extraction is performed on the two-dimensional image based on the first encoder to obtain a first word-gram list.
[0215] In one example, the first encoder may be an image encoder, specifically a DINOv2 (ViT-L / 14) image encoder. The DINOv2 (ViT-L / 14) image encoder may be a self-supervised visual pre-training model for extracting features suitable for visual tasks (such as image classification, instance retrieval, video understanding, depth estimation, and semantic segmentation). The DINOv2 (ViT-L / 14) image encoder may use the ViT (Vision Transformer) network architecture. ViT is a Transformer framework specifically designed for visual tasks. It first decomposes the image into a series of image patches, then converts the image patch sequence into a set of feature vectors, and finally processes this set of vectors through the Transformer.
[0216] DINOv2 (ViT-L / 14) image encoder takes the input two-dimensional image I∈R H×W×C (where H represents the height of the two-dimensional image I, W represents the width of the two-dimensional image I, and C represents the number of channels of the two-dimensional image I. For example, the number of channels of an RGB image is usually 3) is processed to output a visual word list (i.e., the first word list) Among them, N I represents the number of image blocks into which the DINOv2 (ViT-L / 14) image encoder decomposes the two-dimensional image I, and D represents the feature dimension. In addition, in the embodiment of the present disclosure, R represents a real number.
[0217] In a case where the prompt information includes a floor plan, the floor plan is converted into a binary mask, and feature extraction is performed on the binary mask based on the second encoder to obtain a second word list.
[0218] Floor plans are a common input for indoor 3D scene generation tasks, which specify the structure of a room (i.e., walls, doors, and windows). Floor plans can be represented using a binary mask, where 1 represents being inside the room outline and 0 represents being outside the room outline (or vice versa).
[0219] In one example, the second encoder can be a ResNet-34 model. The ResNet-34 model can be used to encode the binary mask of the floor plan, and the features before the ResNet pooling layer can be selected and the features before the ResNet pooling layer can be expanded in the spatial dimension to obtain the second word list. Among them, N F It represents the number of image blocks that the ResNet model decomposes the floor plan into, and D represents the feature dimension.
[0220] When the prompt information includes a three-dimensional model of the first target object, the three-dimensional model of the first target object is rendered to obtain a rendering image of the first target object, and features are extracted from the rendering image based on the third encoder to obtain a third word list.
[0221] The 3D model of the target object typically includes an original 3D mesh representing the shape of the 3D model, texture, material and / or other attribute information. The 3D model of the target object may also be associated with text description information such as category, style or function.
[0222] In one example, the third encoder may be a CLIP model. The rendering of the three-dimensional model of the first target object may be input into the CLIP model with the ViT-L / 14 image encoder, and the feature vector corresponding to the classification header (Class Token) of the CLIP model is selected as the word element corresponding to the three-dimensional model of the first target object (that is, the feature vector corresponding to the three-dimensional model of the first target object), thereby obtaining the third word element list V O ∈R D , where D represents the feature dimension.
[0223] In a case where the prompt information includes the bounding box information of the second target object, the values in the bounding box information of the second target object are encoded based on a set encoding rule to obtain a fourth word-gram list.
[0224] The bounding box information of the second target object can be expressed as B=(c, p, l, r), where c represents the category of the second target object, p represents the coordinate information of the second target object, l represents the size information of the second target object, and r represents the angle information of the second target object.
[0225] Furthermore, the coordinate information p of the second target object can be expressed as (x, y, z) (wherein, x can represent the x-axis coordinate of the second target object (for example, the center point of the second target object); y can represent the y-axis coordinate of the second target object; and z can represent the z-axis coordinate of the second target object). In addition, the size information l of the second target object may include at least one of the length, height, and width of the second target object; the angle information r of the second target object may include the rotation angle of the second target object relative to the horizontal plane (such as the ground); the category c to which the second target object belongs may include, for example, doors, windows, beds, wardrobes, lamps, bedside tables, curtains, decorative paintings, carpets, desks, dressing tables, chairs and / or TV cabinets, etc., and the embodiments of the present disclosure do not impose any limitations on this.
[0226] In one example, each value in the bounding box information B = (c, p, l, r) of the second target object can be encoded based on a set encoding rule. For example, the coordinate value of the coordinate information p of the second target object and the size information l of the second target object can be quantized into 8-bit integers; the angle interval [0, 2π) is divided into 24 intervals, each interval corresponds to an angle span of 15 degrees, and the angle information r of the second target object is mapped to the corresponding interval; and a numerical value is assigned to the category c to which the second target object belongs. Thus, a fourth word list V corresponding to the bounding box information of the second target object can be obtained. B =[v c ,v p ,v l ,v r ]∈R 8×D , where D represents the feature dimension.
[0227] It should be noted that the first target object and the second target object can be the target objects to be completed in the scene completion task, and the second target object and the first target object can be the same target object or different target objects, which is not limited in this embodiment of the present disclosure.
[0228] In a case where the prompt information includes text description information, feature extraction is performed on the text description information based on the fourth encoder to obtain a fifth word-gram list.
[0229] Optionally, the text description information can be a text description of the sample three-dimensional scene; optionally, the text description information can also be a text description of the target object in the sample three-dimensional scene, and the target object can also be the target object to be completed in the scene completion task, and can be the same target object as the first target object and / or the second target object, or can be a different target object.
[0230] In one example, the fourth encoder may be a text encoder of CLIP. Specifically, the text encoder of CLIP may be used to extract features from the text description information, and the feature vector corresponding to the classification head (Class Token) of the text encoder of CLIP may be selected as the word element corresponding to the text description information, thereby obtaining the fifth word element list V T ∈R D , where D represents the feature dimension.
[0231] In a case where the prompt information includes point cloud information, feature extraction is performed on the point cloud information based on the fifth encoder to obtain a sixth word list.
[0232] Optionally, the point cloud information may be point cloud information of a sample three-dimensional scene; optionally, the point cloud information may also be point cloud information of a target object in the sample three-dimensional scene, and the target object may also be a target object to be completed in the scene completion task, and may be the same target object as the first target object and / or the second target object, or may be a different target object.
[0233] In one example, the fifth encoder may be a Swin3D model encoder. Specifically, the Swin3D model encoder may be used to perform feature extraction on the point cloud information to obtain the sixth word list. Among them, N P Represents the number of points in the point cloud information, and D represents the feature dimension.
[0234] In a case where the prompt information includes voxel information, feature extraction is performed on the voxel information based on the sixth encoder to obtain a seventh word list.
[0235] Optionally, the voxel information may be voxel information of a sample three-dimensional scene; optionally, the voxel information may also be voxel information of a target object in the sample three-dimensional scene, and the target object may also be a target object to be completed in a scene completion task, and may be the same target object as the first target object and / or the second target object, or may be a different target object.
[0236] In one example, the sixth encoder may be a Swin3D model encoder. Specifically, the Swin3D model encoder may be used to perform feature extraction on the voxel information to obtain the seventh word list Among them, N VG represents the number of voxels in the voxel information, and D represents the feature dimension.
[0237] By implementing the above S1001, a feature sequence corresponding to the sample three-dimensional scene can be obtained based on the prompt information corresponding to the sample three-dimensional scene, wherein the feature sequence may include at least one of the above-mentioned first word list, second word list, third word list, fourth word list, fifth word list, sixth word list or seventh word list.
[0238] In S1002, the target prediction sequence corresponding to the sample three-dimensional scene is determined using the initial model and the feature sequence; wherein the target prediction sequence is used to establish a three-dimensional model corresponding to the sample three-dimensional scene.
[0239] As shown in FIG12 , it is a schematic diagram of the initial model according to an embodiment of the present disclosure. I , the second word list V F 、Third word list V O 、The fourth word list VB 、The fifth word list V T or the sixth lexical list V P At least one of them) is input into the initial model to obtain a target prediction sequence for establishing a three-dimensional model corresponding to the sample three-dimensional scene.
[0240] In the target prediction sequence, the sample 3D scene can be represented as a set of ordered objects Each object can be represented by Z i ={O i ,B i}, where O i is the unique model ID of the i-th object in the target object database (i.e., the 3D model database), B i Represents the 3D bounding box of the i-th object in the sample 3D scene. The 3D bounding box can be represented as B i =(c i ,l i ,p i ,r i ), including object category c i ∈R, the three-dimensional coordinates of the center point of the object p i ∈R 3 , the three-dimensional size of the object l i ∈R 3 , the object's rotation angle r around the vertical direction i ∈R.
[0241] Here, since almost all indoor objects are aligned with the horizontal plane, a full three-axis rotation can be omitted. A pseudo-perspective-centered right-handed coordinate system can be used, with the origin O located at the center of the camera, the X-axis pointing to the right, the Y-axis pointing forward, and the Z-axis pointing upward. The Z-axis can be aligned in the opposite direction of gravity, and the XOY plane is parallel to the ground. Thus, the rotation of the object is the rotation of the object in the -Y direction on the XOY plane relative to the initial orientation (angle).
[0242] The initial model can include Transformer encoding layers and Transformer decoding layers.
[0243] The Transformer encoding layer is configured to encode the feature sequence to obtain semantic features corresponding to the feature sequence.
[0244] The Transformer decoding layer is configured to use semantic features and historical prediction sequences corresponding to the sample three-dimensional scene to determine a first prediction sequence; wherein the first prediction sequence is used to determine a target prediction sequence.
[0245] In one embodiment, the Transformer encoding layer and the Transformer decoding layer included in the initial model can both be composed of 6 Transformer modules.
[0246] In one example, the target prediction sequence may include the prediction sequence of the initial model for the sample three-dimensional scene at the first moment, and the historical prediction sequence may include the prediction sequence of the initial model for the sample three-dimensional scene at the first time series; wherein the first moment is after the first time series.
[0247] For example, the first time series includes moments within the time period [tk, t), for example, the first time series includes moments tk, t-k+1, t-k+2, ..., t-1, and the first moment may include moment t.
[0248] Based on this, if the feature sequence is input into the initial model, the decoding layer of the initial model can receive the semantic features of the feature sequence output by the encoding layer at time t, as well as the output sequence (i.e., historical prediction sequence) E(s) of the decoding layer at time tk, t-k+1, t-k+2, ..., and t-1 for the feature sequence. <t ); The decoding layer of the initial model receives the semantic features of the feature sequence output by the encoding layer at time t and the output sequence (i.e., historical prediction sequence) E(s) of the feature sequence output by the encoding layer at the previous k moments (i.e., time tk, t-k+1, t-k+2, ..., and t-1) <t ) After that, we can use the semantic features of the feature sequence at time t and the output sequence E(s) of the previous k moments <t ), determine the first prediction sequence corresponding to the sample three-dimensional scene, that is, the hidden feature h corresponding to the sample three-dimensional scene t Thus, the initial model can further determine the target prediction sequence based on the first prediction sequence.
[0249] The initial model may further include a linear layer and a normalization layer. The linear layer may be configured to adjust the feature dimensions of the first prediction sequence to obtain an adjusted first prediction sequence. The normalization layer may be configured to normalize the adjusted first prediction sequence based on a feature sequence corresponding to the sample three-dimensional scene to obtain a second prediction sequence; wherein the second prediction sequence is used to determine the target prediction sequence.
[0250] For example, still taking the first time series including time tk, t-k+1, t-k+2, ..., and t-1, and the first time including time t as an example, after obtaining the first prediction sequence corresponding to the sample three-dimensional scene, that is, the hidden feature h corresponding to the sample three-dimensional scene t Afterwards, the second prediction sequence corresponding to the sample three-dimensional scene can be determined using formula (8): p(st |s <t ,P)=p(E(s t )|E(s <t ),F)=softmax(linear(h t )) Formula (8)
[0251] Wherein, P represents the prompt information corresponding to the sample three-dimensional scene, and F can represent the feature sequence corresponding to the prompt information P (that is, the feature sequence corresponding to the sample three-dimensional scene); p(s t |s <t ,P) can represent the second prediction sequence corresponding to the sample three-dimensional scene.
[0252] That is to say, we can first use the linear layer of the initial model to adjust the hidden feature h t to obtain the adjusted hidden features; then, the normalization layer of the initial model can be used to process the hidden features after the dimension adjustment to obtain a second prediction sequence corresponding to the sample three-dimensional scene and an occurrence probability corresponding to the second prediction sequence.
[0253] Furthermore, after determining the second prediction sequence corresponding to the sample three-dimensional scene, the target prediction sequence can be determined using formula (9): p(S′|P)=∏ t p(s t |s <t ,P)=∏ t p(E(s t )|E(s <t ),F) Formula (9)
[0254] Among them, P represents the prompt information corresponding to the sample three-dimensional scene, F represents the feature sequence corresponding to the prompt information P; p(S′|P) can represent the target prediction sequence corresponding to the sample three-dimensional scene.
[0255] In S1003 , the target prediction sequence and the target label sequence of the sample three-dimensional scene are used to adjust the initial model to obtain a target model.
[0256] In one embodiment, the model training method may further include: determining a target label sequence based on a sample three-dimensional scene; wherein the target label sequence includes a real sequence corresponding to the sample three-dimensional scene, and the real sequence is used to establish a real three-dimensional model corresponding to the sample three-dimensional scene.
[0257] In one embodiment, determining a target label sequence based on a sample three-dimensional scene may include: determining an initial label sequence based on the sample three-dimensional scene; the initial label sequence includes a first initial label sequence and / or a second initial label sequence; and, sequentially performing discretization processing and feature embedding processing on the initial label sequence to obtain a target label sequence.
[0258] In one embodiment, determining an initial label sequence based on a sample three-dimensional scene may include: determining M vertices contained in the sample three-dimensional scene, and determining a first initial label sequence corresponding to the sample three-dimensional scene based on first coordinate information of the M vertices; M is a positive integer; and / or determining N third target objects contained in the sample three-dimensional scene, and respectively determining target object feature sequences corresponding to the N third target objects; determining a second initial label sequence corresponding to the sample three-dimensional scene based on the N target object feature sequences corresponding to the third target objects; N is a positive integer.
[0259] In one embodiment, determining a first initial label sequence based on the first coordinate information of M vertices may include: using the first coordinate information of the M vertices to determine a first arrangement order corresponding to the M vertices; and determining the first initial label sequence based on the first coordinate information and the first arrangement order of the M vertices.
[0260] In one embodiment, determining a second initial label sequence based on N target object feature sequences corresponding to the third target object may include: using at least one of the second coordinate information, category, and occurrence frequency of the N third target objects in the sample three-dimensional scene to determine a second arrangement order corresponding to the N third target objects; and determining a second initial label sequence based on the second arrangement order and the N target object feature sequences corresponding to the third target object.
[0261] In one embodiment, determining a target object feature sequence corresponding to a third target object may include: determining attribute information of the third target object; wherein the attribute information includes at least one of second coordinate information, category, size information, angle information, and identification information of the third target object, and the identification information is used to characterize the third target object; and, using the attribute information of the third target object, determining the target object feature sequence.
[0262] It should be noted that the specific implementation method of obtaining the target label sequence of the sample three-dimensional scene can refer to the relevant description of the model training method provided in the first aspect of the embodiment of the present disclosure, and the embodiment of the present disclosure will not be repeated here.
[0263] In addition, the specific implementation of adjusting the initial model using the target prediction sequence and the target label sequence of the sample three-dimensional scene to obtain the target model can also refer to the relevant description of the model training method provided in the first aspect of the embodiment of the present disclosure, and the embodiment of the present disclosure will not be repeated here.
[0264] The fifth aspect of the embodiments of the present disclosure provides a method for unifying various indoor three-dimensional scene modeling tasks (such as indoor three-dimensional scene analysis and indoor three-dimensional scene synthesis) through prompts. By inputting different applications in the form of prompts, various applications of indoor three-dimensional scene modeling are modeled as the problem of how to generate sequences, so that a unified algorithm framework can be used to implement various related applications of indoor three-dimensional scene modeling, including but not limited to indoor three-dimensional scene reconstruction, indoor three-dimensional scene generation, and / or scene completion.
[0265] The sixth aspect of the embodiment of the present disclosure proposes a method for establishing a three-dimensional model. Figure 13 is an implementation flow chart of the method for establishing a three-dimensional model provided by the sixth aspect of the embodiment of the present disclosure, including: S1301, S1302 and S1303.
[0266] In S1301 , a feature sequence corresponding to the three-dimensional scene to be modeled is determined according to prompt information corresponding to the three-dimensional scene to be modeled.
[0267] In one embodiment, the prompt information may include a two-dimensional image corresponding to the three-dimensional scene to be modeled, a floor plan corresponding to the three-dimensional scene to be modeled, a three-dimensional model of the first target object contained in the three-dimensional scene to be modeled, bounding box information of the second target object contained in the three-dimensional scene to be modeled, text description information, point cloud information, or at least one of voxel information.
[0268] In one embodiment, S1301 determines a feature sequence based on the prompt information, including at least one of the following:
[0269] In a case where the prompt information includes a two-dimensional image, feature extraction is performed on the two-dimensional image based on the first encoder to obtain a first word-gram list.
[0270] In one example, the first encoder may be an image encoder, specifically a DINOv2 (ViT-L / 14) image encoder. The DINOv2 (ViT-L / 14) image encoder may be a self-supervised visual pre-training model for extracting features suitable for visual tasks (such as image classification, instance retrieval, video understanding, depth estimation, and semantic segmentation). The DINOv2 (ViT-L / 14) image encoder may use the ViT (Vision Transformer) network architecture. ViT is a Transformer framework specifically designed for visual tasks. It first decomposes the image into a series of image patches, then converts the image patch sequence into a set of feature vectors, and finally processes this set of vectors through the Transformer.
[0271] DINOv2 (ViT-L / 14) image encoder takes the input two-dimensional image I∈R H×W×C (where H represents the height of the two-dimensional image I, W represents the width of the two-dimensional image I, and C represents the number of channels of the two-dimensional image I. For example, the number of channels of an RGB image is usually 3) is processed to output a visual word list (i.e., the first word list) Among them, N I It represents the number of image blocks into which the DINOv2 (ViT-L / 14) image encoder decomposes the two-dimensional image I, and D represents the feature dimension.
[0272] In a case where the prompt information includes a floor plan, the floor plan is converted into a binary mask, and feature extraction is performed on the binary mask based on the second encoder to obtain a second word list.
[0273] Floor plans are a common input for indoor 3D scene generation tasks, which specify the structure of a room (i.e., walls, doors, and windows). Floor plans can be represented using a binary mask, where 1 represents being inside the room outline and 0 represents being outside the room outline (or vice versa).
[0274] In one example, the second encoder can be a ResNet-34 model. The ResNet-34 model can be used to encode the binary mask of the floor plan, and the features before the ResNet pooling layer can be selected and the features before the ResNet pooling layer can be expanded in the spatial dimension to obtain the second word list. Among them, N F It represents the number of image blocks that the ResNet model decomposes the floor plan into, and D represents the feature dimension.
[0275] When the prompt information includes a three-dimensional model of the first target object, the three-dimensional model of the first target object is rendered to obtain a rendering image of the first target object, and features are extracted from the rendering image based on the third encoder to obtain a third word list.
[0276] The 3D model of the target object typically includes an original 3D mesh representing the shape of the 3D model, texture, material and / or other attribute information. The 3D model of the target object may also be associated with text description information such as category, style or function.
[0277] In one example, the third encoder may be a CLIP model. The rendering of the first target object may be input into the CLIP model with the ViT-L / 14 image encoder, and the feature vector corresponding to the classification header (Class Token) of the CLIP model is selected as the word element corresponding to the three-dimensional model of the first target object (that is, the feature vector corresponding to the three-dimensional model of the first target object), thereby obtaining the third word element list V O ∈R D , where D represents the feature dimension.
[0278] In a case where the prompt information includes the bounding box information of the second target object, the values in the bounding box information of the second target object are encoded based on a set encoding rule to obtain a fourth word-gram list.
[0279] The bounding box information of the second target object can be expressed as B=(c, p, l, r), where c represents the category of the second target object, p represents the coordinate information of the second target object, l represents the size information of the second target object, and r represents the angle information of the second target object.
[0280] Furthermore, the coordinate information p of the second target object can be expressed as (x, y, z) (wherein, x can represent the x-axis coordinate of the second target object (for example, the center point of the second target object); y can represent the y-axis coordinate of the second target object; and z can represent the z-axis coordinate of the second target object). In addition, the size information l of the second target object may include at least one of the length, height, and width of the second target object; the angle information r of the second target object may include the rotation angle of the second target object relative to the horizontal plane (such as the ground); the category c to which the second target object belongs may include, for example, doors, windows, beds, wardrobes, lamps, bedside tables, curtains, decorative paintings, carpets, desks, dressing tables, chairs and / or TV cabinets, etc., and the embodiments of the present disclosure do not impose any restrictions on this.
[0281] In one example, each value in the bounding box information B = (c, p, l, r) of the second target object can be encoded based on a set encoding rule. For example, the coordinate value of the coordinate information p of the second target object and the size information l of the second target object can be quantized into 8-bit integers; the angle interval [0, 2π) is divided into 24 intervals, each interval corresponds to an angle span of 15 degrees, and the angle information r of the second target object is mapped to the corresponding interval; and a numerical value is assigned to the category c to which the second target object belongs. Thus, a fourth word list V corresponding to the bounding box information of the second target object can be obtained. B =[v c ,v p ,v l ,v r ]∈R 8×D , where D represents the feature dimension.
[0282] It should be noted that the first target object and the second target object can be the target objects to be completed in the scene completion task, and the second target object and the first target object can be the same target object or different target objects, which is not limited in this embodiment of the present disclosure.
[0283] In a case where the prompt information includes text description information, feature extraction is performed on the text description information based on the fourth encoder to obtain a fifth word-gram list.
[0284] Optionally, the text description information can be a text description of the three-dimensional scene to be modeled; optionally, the text description information can also be a text description of the target object in the three-dimensional scene to be modeled, and the target object can also be the target object to be completed in the scene completion task, and can be the same target object as the first target object and / or the second target object, or can be a different target object.
[0285] In one example, the fourth encoder may be a text encoder of CLIP. Specifically, the text encoder of CLIP may be used to extract features from the text description information, and the feature vector corresponding to the classification head (Class Token) of the text encoder of CLIP may be selected as the word element corresponding to the text description information, thereby obtaining the fifth word element list V T ∈R D , where D represents the feature dimension.
[0286] In a case where the prompt information includes point cloud information, feature extraction is performed on the point cloud information based on the fifth encoder to obtain a sixth word list.
[0287] Optionally, the point cloud information may be point cloud information of the three-dimensional scene to be modeled; optionally, the point cloud information may also be point cloud information of a target object in the three-dimensional scene to be modeled, and the target object may also be the target object to be completed in the scene completion task, and may be the same target object as the first target object and / or the second target object, or may be a different target object.
[0288] In one example, the fifth encoder may be a Swin3D model encoder. Specifically, the Swin3D model encoder may be used to perform feature extraction on the point cloud information to obtain the sixth word list. Among them, N P Represents the number of points in the point cloud information, and D represents the feature dimension.
[0289] In a case where the prompt information includes voxel information, feature extraction is performed on the voxel information based on the sixth encoder to obtain a seventh word list.
[0290] Optionally, the voxel information may be voxel information of the three-dimensional scene to be modeled; optionally, the voxel information may also be voxel information of a target object in the three-dimensional scene to be modeled, and the target object may also be the target object to be completed in the scene completion task, and may be the same target object as the first target object and / or the second target object, or may be a different target object.
[0291] In one example, the sixth encoder may be a Swin3D model encoder. Specifically, the Swin3D model encoder may be used to perform feature extraction on the voxel information to obtain the seventh word list Among them, N VG represents the number of voxels in the voxel information, and D represents the feature dimension.
[0292] By implementing the above S1301, a feature sequence corresponding to the three-dimensional scene to be modeled can be obtained based on the prompt information corresponding to the three-dimensional scene to be modeled, wherein the feature sequence may include at least one of the first word list, the second word list, the third word list, the fourth word list, the fifth word list, the sixth word list or the seventh word list.
[0293] In S1302, a target prediction sequence corresponding to the three-dimensional scene to be modeled is determined using the target model and the feature sequence, wherein the target model is trained using the model training method provided by the fifth aspect of the embodiment of the present disclosure.
[0294] In the target prediction sequence, the 3D scene to be modeled can be represented as a set of ordered objects Each object can be represented by Z i ={O i,B i}, where O i is the unique model ID of the i-th object in the target object database (i.e., the 3D model database), B i Represents the 3D bounding box of the i-th object in the 3D scene to be modeled. The 3D bounding box can be represented as B i =(c i ,l i ,p i ,r i ), including object category c i ∈R, the three-dimensional coordinates of the center point of the object p i ∈R 3 , the three-dimensional size of the object l i ∈R 3 , the object's rotation angle r around the vertical direction i ∈R.
[0295] The target model can include Transformer encoding layers and Transformer decoding layers.
[0296] The Transformer encoding layer is configured to process the feature sequence to obtain semantic features corresponding to the feature sequence.
[0297] The Transformer decoding layer is configured to determine a first prediction sequence using semantic features and a historical prediction sequence corresponding to the three-dimensional scene to be modeled; wherein the first prediction sequence is used to determine a target prediction sequence.
[0298] The Transformer's ability to model long-range dependencies enables the target model to simultaneously consider global information. Therefore, since the target model includes both the Transformer encoding and decoding layers, it improves the target model's prediction efficiency and offers advantages such as powerful expressiveness, parallel computing, long-range dependency modeling, multi-head attention mechanisms, and interpretability, thereby improving the target model's performance and accuracy.
[0299] The target model may further include a linear layer and a normalization layer. The linear layer is configured to adjust the feature dimensions of the first prediction sequence to obtain an adjusted first prediction sequence. The normalization layer is configured to normalize the adjusted first prediction sequence to obtain a second prediction sequence; wherein the second prediction sequence is used to determine the target prediction sequence.
[0300] In one example, the target prediction sequence may include a prediction sequence of the target model for the 3D scene to be modeled at a second moment, and the historical prediction sequence may include a prediction sequence of the target model for the 3D scene to be modeled at a second time series; wherein the second moment is after the second time series. For example, the second time series includes moments within the time period [tk, t], for example, the second time series may include moments tk, t-k+1, t-k+2, ..., t-1, and the second moment may include moment t.
[0301] It should be noted that the specific implementation method of using the target model and feature sequence to determine the target prediction sequence corresponding to the three-dimensional scene to be modeled is similar to the specific implementation method of using the initial model and feature sequence to determine the target prediction sequence corresponding to the sample three-dimensional scene in the model training method provided in the fifth aspect of the embodiment of the present disclosure. The embodiment of the present disclosure will not be repeated here.
[0302] In S1303 , a three-dimensional model corresponding to the three-dimensional scene to be modeled is established according to the target prediction sequence.
[0303] As shown in Figures 14A and 14B, when the prompt information includes a two-dimensional image, the feature sequence corresponding to the two-dimensional image is input into the target model. The target model can output a target prediction sequence for the three-dimensional scene to be modeled, and use the target prediction sequence to establish a three-dimensional model corresponding to the two-dimensional image, that is, a three-dimensional model corresponding to the three-dimensional scene to be modeled.
[0304] As shown in Figures 14C and 14D, when the prompt information includes a floor plan, the feature sequence corresponding to the floor plan is input into the target model. The target model can output a target prediction sequence for the three-dimensional scene to be modeled, and use the target prediction sequence to generate a three-dimensional model corresponding to the floor plan, that is, a three-dimensional model corresponding to the three-dimensional scene to be modeled.
[0305] As shown in Figures 14E and 14F, when the prompt information includes a three-dimensional model of the first target object (the target object to be completed), the feature sequence corresponding to the three-dimensional model of the first target object is input into the target model, and the target model can output a target prediction sequence for the three-dimensional scene to be modeled, and use the target prediction sequence to complete the three-dimensional model of the three-dimensional scene to be modeled.
[0306] The sixth aspect of the embodiments of the present disclosure provides a method for unifying various indoor three-dimensional scene modeling tasks (such as indoor three-dimensional scene analysis and indoor three-dimensional scene synthesis) through prompts. By inputting different applications in the form of prompts, various applications of indoor three-dimensional scene modeling are modeled as the problem of how to generate sequences, so that a unified algorithm framework can be used to implement various related applications of indoor three-dimensional scene modeling, including but not limited to indoor three-dimensional scene reconstruction, indoor three-dimensional scene generation, and / or scene completion.
[0307] The seventh aspect of the embodiment of the present disclosure also proposes a model training device. Figure 15 is a structural schematic diagram of the model training device 1500 provided by the seventh aspect of the embodiment of the present disclosure, including: a first acquisition module 1501, a first determination module 1502 and an adjustment module 1503.
[0308] The first acquisition module 1501 is configured to determine a feature sequence corresponding to the sample three-dimensional scene according to prompt information corresponding to the sample three-dimensional scene.
[0309] The first determination module 1502 is configured to determine a target prediction sequence corresponding to the sample three-dimensional scene using the initial model and the feature sequence; wherein the target prediction sequence is used to establish a three-dimensional model corresponding to the sample three-dimensional scene.
[0310] The adjustment module 1503 is configured to adjust the initial model using the target prediction sequence and the target label sequence of the sample three-dimensional scene to obtain a target model.
[0311] In one embodiment, the prompt information includes at least one of a two-dimensional image corresponding to the sample three-dimensional scene, a floor plan corresponding to the sample three-dimensional scene, a three-dimensional model of the first target object contained in the sample three-dimensional scene, bounding box information of the second target object contained in the sample three-dimensional scene, text description information, point cloud information, or voxel information.
[0312] In one embodiment, the first acquisition module 1501 may be configured to perform at least one of the following:
[0313] When the prompt information includes a two-dimensional image, performing feature extraction on the two-dimensional image based on the first encoder to obtain a first word-gram list;
[0314] In a case where the prompt information includes a floor plan, converting the floor plan into a binary mask, and performing feature extraction on the binary mask based on a second encoder to obtain a second word list;
[0315] When the prompt information includes a three-dimensional model of the first target object, rendering the three-dimensional model of the first target object to obtain a rendered image of the first target object, and performing feature extraction on the rendered image based on a third encoder to obtain a third word list;
[0316] In a case where the prompt information includes bounding box information of the second target object, encoding the values in the bounding box information of the second target object based on a set encoding rule to obtain a fourth word-gram list;
[0317] When the prompt information includes text description information, performing feature extraction on the text description information based on the fourth encoder to obtain a fifth word list;
[0318] When the prompt information includes point cloud information, performing feature extraction on the point cloud information based on the fifth encoder to obtain a sixth word list; or
[0319] When the prompt information includes voxel information, performing feature extraction on the voxel information based on the sixth encoder to obtain a seventh word list;
[0320] The feature sequence includes at least one of a first word list, a second word list, a third word list, a fourth word list, a fifth word list, a sixth word list, or a seventh word list.
[0321] In one embodiment, the model training device may further include:
[0322] The second determination module 1504 is configured to determine a target tag sequence according to the sample three-dimensional scene;
[0323] The target label sequence includes a real sequence corresponding to the sample three-dimensional scene, and the real sequence is used to establish a real three-dimensional model corresponding to the sample three-dimensional scene.
[0324] In one implementation, the second determining module 1504 may be configured to:
[0325] Determine an initial label sequence according to the sample three-dimensional scene; the initial label sequence includes a first initial label sequence and / or a second initial label sequence; and
[0326] The initial label sequence is discretized and feature embedded in turn to obtain the target label sequence.
[0327] For the specific functions and exemplary descriptions of the modules and sub-modules of the model training device provided in the seventh aspect of the embodiment of the present disclosure, please refer to the relevant descriptions of the corresponding steps in the model training method provided in the fifth aspect of the embodiment of the present disclosure, and the embodiment of the present disclosure will not be repeated here.
[0328] The eighth aspect of the embodiment of the present disclosure also proposes a device for establishing a three-dimensional model. Figure 16 is a structural schematic diagram of the device 1600 for establishing a three-dimensional model provided by the eighth aspect of the embodiment of the present disclosure, including: a second acquisition module 1601, a third determination module 1602 and a modeling module 1603.
[0329] The second acquisition module 1601 is configured to determine a feature sequence corresponding to the three-dimensional scene to be modeled according to the prompt information corresponding to the three-dimensional scene to be modeled.
[0330] The third determination module 1602 is configured to use the target model and the feature sequence to determine the target prediction sequence corresponding to the three-dimensional scene to be modeled, wherein the target model is trained using the model training device provided by the seventh aspect of the embodiment of the present disclosure.
[0331] The modeling module 1603 is configured to establish a three-dimensional model corresponding to the three-dimensional scene to be modeled according to the target prediction sequence.
[0332] For the specific functions and exemplary descriptions of the modules and submodules of the device for establishing a three-dimensional model provided in the eighth aspect of the embodiment of the present disclosure, please refer to the relevant descriptions of the corresponding steps in the method for establishing a three-dimensional model provided in the sixth aspect of the embodiment of the present disclosure, and the embodiment of the present disclosure will not be repeated here.
[0333] The seventh and eighth aspects of the embodiments of the present disclosure provide a device for unifying various indoor three-dimensional scene modeling tasks (such as indoor three-dimensional scene analysis and indoor three-dimensional scene synthesis) through prompts. By inputting different applications in the form of prompts, various applications of indoor three-dimensional scene modeling are modeled as the problem of how to generate sequences, so that a unified algorithm framework can be used to implement various related applications of indoor three-dimensional scene modeling, including but not limited to indoor three-dimensional scene reconstruction, indoor three-dimensional scene generation, and / or scene completion.
[0334] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0335] Figure 17 is a block diagram of the structure of an electronic device according to an embodiment of the present disclosure. As shown in Figure 17, the electronic device includes: a memory 1710 and a processor 1720, and the memory 1710 stores a computer program that can be run on the processor 1720. The number of memories 1710 and processors 1720 can be one or more. The memory 1710 can store one or more computer programs. When the one or more computer programs are executed by the electronic device, the electronic device executes any of the methods provided in the embodiments of the present disclosure. The electronic device may also include: a communication interface 1730 for communicating with external devices and performing data exchange and transmission.
[0336] If the memory 1710, processor 1720, and communication interface 1730 are implemented independently, the memory 1710, processor 1720, and communication interface 1730 can be interconnected via a bus and communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, FIG17 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.
[0337] Optionally, in a specific implementation, if the memory 1710, the processor 1720 and the communication interface 1730 are integrated on a chip, the memory 1710, the processor 1720 and the communication interface 1730 can communicate with each other through an internal interface.
[0338] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the Advanced RISC Machines (ARM) architecture.
[0339] Furthermore, optionally, the above-mentioned memory may include a read-only memory and a random access memory, and may also include a non-volatile random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct memory bus random access memory (DR RAM).
[0340] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the process or function according to the embodiment of the present disclosure is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (for example: coaxial cable, optical fiber, data subscriber line (Digital Subscriber Line, DSL)) or wireless (for example: infrared, Bluetooth, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more available media integrated. Available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital versatile discs (DVDs)), or semiconductor media (e.g., solid-state drives (SSDs)). It is worth noting that the computer-readable storage media mentioned in the present disclosure may be non-volatile storage media, in other words, non-transitory storage media.
[0341] Those skilled in the art will understand that all or part of the steps of implementing the above embodiments may be accomplished by hardware, or by programs instructing related hardware to accomplish the steps. The programs may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk, or an optical disk, etc.
[0342] In the description of the embodiments of the present disclosure, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.
[0343] In the description of the embodiments of the present disclosure, unless otherwise specified, " / " means or. For example, A / B can mean A or B. "And / or" in this document is only a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0344] In the description of the embodiments of the present disclosure, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present disclosure, unless otherwise specified, "plurality" means two or more.
[0345] The above are merely exemplary embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present disclosure shall be included in the scope of protection of the present disclosure.
Claims
1. A model training method, comprising: Determining a target prediction sequence corresponding to the sample three-dimensional scene using an initial model and an image feature sequence corresponding to the sample three-dimensional scene; wherein the target prediction sequence is used to establish a three-dimensional model corresponding to the sample three-dimensional scene, and the image feature sequence is determined based on a two-dimensional image corresponding to the sample three-dimensional scene; and The target prediction sequence and the target label sequence of the sample three-dimensional scene are used to adjust the initial model to obtain a target model.
2. The training method of the model according to claim 1, further comprising: Determining the target tag sequence according to the sample three-dimensional scene; The target label sequence includes a real sequence corresponding to the sample three-dimensional scene, and the real sequence is used to establish a real three-dimensional model corresponding to the sample three-dimensional scene.
3. The training method of the model according to claim 2, wherein: Determining the target tag sequence according to the sample three-dimensional scene includes: Determine an initial label sequence according to the sample three-dimensional scene; the initial label sequence includes a first initial label sequence and / or a second initial label sequence; and The initial label sequence is sequentially subjected to discretization processing and feature embedding processing to obtain the target label sequence.
4. The training method of the model according to claim 3, wherein: Determining the initial label sequence according to the sample three-dimensional scene includes: Determine M vertices included in the sample three-dimensional scene, and determine a first initial label sequence corresponding to the sample three-dimensional scene according to first coordinate information of the M vertices; M is a positive integer; and / or, Determine N target objects contained in the sample three-dimensional scene, and respectively determine target object feature sequences corresponding to the N target objects; determine a second initial label sequence corresponding to the sample three-dimensional scene based on the N target object feature sequences corresponding to the target objects; N is a positive integer.
5. The training method of the model according to claim 4, wherein: Determining the first initial label sequence according to the first coordinate information of the M vertices includes: Using the first coordinate information of the M vertices, determining a first arrangement order corresponding to the M vertices; and The first initial label sequence is determined according to the first coordinate information of the M vertices and the first arrangement order.
6. The training method of the model according to claim 4, wherein: Determining the second initial label sequence according to the N target object feature sequences corresponding to the target objects includes: Determine a second arrangement order corresponding to the N target objects by using at least one of the second coordinate information, the category and the occurrence frequency of the N target objects in the sample three-dimensional scene; and The second initial label sequence is determined according to the second arrangement order and the N target object feature sequences corresponding to the target objects.
7. The training method of the model according to claim 6, wherein: Determining a target object feature sequence corresponding to the target object includes: Determine attribute information of the target object; wherein the attribute information includes at least one of second coordinate information, category, size information, angle information and identification information of the target object, and the identification information is used to characterize the target object; and The target object feature sequence is determined using the attribute information of the target object.
8. The method for training a model according to claim 1, wherein: The initial model includes a transformer encoding layer and a transformer decoding layer; wherein, The Transformer encoding layer is configured to process the image feature sequence to obtain semantic features corresponding to the image feature sequence; and The Transformer decoding layer is configured to determine a first prediction sequence using the semantic features and a historical prediction sequence corresponding to the sample three-dimensional scene; wherein the first prediction sequence is used to determine the target prediction sequence.
9. The training method of the model according to claim 8, wherein: The initial model also includes a linear layer and a normalization layer; wherein, The linear layer is configured to adjust the feature dimension of the first prediction sequence to obtain an adjusted first prediction sequence; and The normalization layer is configured to perform normalization processing on the adjusted first prediction sequence to obtain a second prediction sequence; wherein the second prediction sequence is used to determine the target prediction sequence.
10. The training method of the model according to claim 9, wherein: The target prediction sequence includes: a prediction sequence of the initial model for the sample three-dimensional scene at a first moment; and The historical prediction sequence includes: a prediction sequence of the initial model for the sample three-dimensional scene in a first time sequence; The first moment is after the first time series.
11. A method for establishing a three-dimensional model, comprising: Determining a target prediction sequence corresponding to the three-dimensional scene to be modeled using a target model and an image feature sequence corresponding to the three-dimensional scene to be modeled; wherein the image feature sequence is determined based on a two-dimensional image corresponding to the three-dimensional scene to be modeled; and According to the target prediction sequence, a three-dimensional model corresponding to the three-dimensional scene to be modeled is established; wherein the target model is trained using the model training method according to any one of claims 1-10.
12. A model training device, comprising: A first determination module is configured to determine a target prediction sequence corresponding to the sample three-dimensional scene using an initial model and an image feature sequence corresponding to the sample three-dimensional scene; wherein the target prediction sequence is used to establish a three-dimensional model corresponding to the sample three-dimensional scene, and the image feature sequence is determined based on a two-dimensional image corresponding to the sample three-dimensional scene; and The adjustment module is configured to adjust the initial model using the target prediction sequence and the target label sequence of the sample three-dimensional scene to obtain a target model.
13. The training device of the model according to claim 12, wherein: Also includes: A second determination module is configured to determine the target tag sequence according to the sample three-dimensional scene; The target label sequence includes a real sequence corresponding to the sample three-dimensional scene, and the real sequence is used to establish a real three-dimensional model corresponding to the sample three-dimensional scene.
14. The training device of the model according to claim 13, wherein: The second determining module is configured to: Determine an initial label sequence according to the sample three-dimensional scene; the initial label sequence includes a first initial label sequence and / or a second initial label sequence; and The initial label sequence is sequentially subjected to discretization processing and feature embedding processing to obtain the target label sequence.
15. A device for building a three-dimensional model, comprising: a third determination module configured to determine a target prediction sequence corresponding to the three-dimensional scene to be modeled using the target model and an image feature sequence corresponding to the three-dimensional scene to be modeled; wherein the image feature sequence is determined based on a two-dimensional image corresponding to the three-dimensional scene to be modeled; and A modeling module is configured to establish a three-dimensional model corresponding to the three-dimensional scene to be modeled according to the target prediction sequence; wherein the target model is trained using a model training device according to any one of claims 12-14.
16. A model training method comprising: Determining a feature sequence corresponding to the sample three-dimensional scene according to prompt information corresponding to the sample three-dimensional scene; Determining a target prediction sequence corresponding to the sample three-dimensional scene using the initial model and the feature sequence; wherein the target prediction sequence is used to establish a three-dimensional model corresponding to the sample three-dimensional scene; and The target prediction sequence and the target label sequence of the sample three-dimensional scene are used to adjust the initial model to obtain a target model.
17. The training method of the model according to claim 16, wherein: The prompt information includes at least one of a two-dimensional image corresponding to the sample three-dimensional scene, a floor plan corresponding to the sample three-dimensional scene, a three-dimensional model of a first target object contained in the sample three-dimensional scene, bounding box information of a second target object contained in the sample three-dimensional scene, text description information, point cloud information, or voxel information.
18. The model training method according to claim 17, wherein: Determining the feature sequence according to the prompt information includes at least one of the following: In a case where the prompt information includes the two-dimensional image, extracting features from the two-dimensional image based on a first encoder to obtain a first word list; In a case where the prompt information includes the floor plan, converting the floor plan into a binary mask, and performing feature extraction on the binary mask based on a second encoder to obtain a second word list; In a case where the prompt information includes a three-dimensional model of the first target object, rendering the three-dimensional model of the first target object to obtain a rendering of the first target object, and performing feature extraction on the rendering based on a third encoder to obtain a third word list; In a case where the prompt information includes the bounding box information of the second target object, encoding the value in the bounding box information of the second target object based on a set encoding rule to obtain a fourth word list; In a case where the prompt information includes the text description information, performing feature extraction on the text description information based on a fourth encoder to obtain a fifth word-gram list; When the prompt information includes the point cloud information, performing feature extraction on the point cloud information based on a fifth encoder to obtain a sixth word list; or In a case where the prompt information includes the voxel information, performing feature extraction on the voxel information based on a sixth encoder to obtain a seventh word list; Wherein, the feature sequence includes at least one of the first word-gram list, the second word-gram list, the third word-gram list, the fourth word-gram list, the fifth word-gram list, the sixth word-gram list or the seventh word-gram list.
19. The model training method according to claim 16, further comprising: Determining the target tag sequence according to the sample three-dimensional scene; The target label sequence includes a real sequence corresponding to the sample three-dimensional scene, and the real sequence is used to establish a real three-dimensional model corresponding to the sample three-dimensional scene.
20. The model training method according to claim 19, wherein: Determining the target tag sequence according to the sample three-dimensional scene includes: Determine an initial label sequence according to the sample three-dimensional scene; the initial label sequence includes a first initial label sequence and / or a second initial label sequence; and The initial label sequence is sequentially subjected to discretization processing and feature embedding processing to obtain the target label sequence.
21. The method for training a model according to claim 20, wherein: Determining the initial label sequence according to the sample three-dimensional scene includes: Determine M vertices included in the sample three-dimensional scene, and determine a first initial label sequence corresponding to the sample three-dimensional scene according to first coordinate information of the M vertices; M is a positive integer; and / or, Determine N third target objects contained in the sample three-dimensional scene, and respectively determine the target object feature sequences corresponding to the N third target objects; determine a second initial label sequence corresponding to the sample three-dimensional scene based on the N target object feature sequences corresponding to the third target objects; N is a positive integer.
22. The training method of the model according to claim 21, wherein: Determining the first initial label sequence according to the first coordinate information of the M vertices includes: Using the first coordinate information of the M vertices, determining a first arrangement order corresponding to the M vertices; and The first initial label sequence is determined according to the first coordinate information of the M vertices and the first arrangement order.
23. The method for training a model according to claim 21, wherein: Determining the second initial label sequence according to the N target object feature sequences corresponding to the third target object includes: Determine a second arrangement order corresponding to the N third target objects by using at least one of the second coordinate information, the category and the occurrence frequency of the N third target objects in the sample three-dimensional scene; and The second initial label sequence is determined according to the second arrangement order and the N target object feature sequences corresponding to the third target object.
24. The method for training a model according to claim 23, wherein: Determining a target object feature sequence corresponding to the third target object includes: Determine the attribute information of the third target object; wherein the attribute information includes the third target object at least one of the second coordinate information, the category, the size information, the angle information and the identification information of the third target object, and the identification information is used to characterize the third target object; and The target object feature sequence is determined using the attribute information of the third target object.
25. The method for training a model according to claim 16, wherein: The initial model includes a transformer encoding layer and a transformer decoding layer; wherein, The Transformer encoding layer is configured to process the feature sequence to obtain semantic features corresponding to the feature sequence; and The Transformer decoding layer is configured to determine a first prediction sequence using the semantic features and a historical prediction sequence corresponding to the sample three-dimensional scene; wherein the first prediction sequence is used to determine the target prediction sequence.
26. The method for training a model according to claim 25, wherein: The initial model also includes a linear layer and a normalization layer; wherein, The linear layer is configured to adjust the feature dimension of the first prediction sequence to obtain an adjusted first prediction sequence; and The normalization layer is configured to perform normalization processing on the adjusted first prediction sequence to obtain a second prediction sequence; wherein the second prediction sequence is used to determine the target prediction sequence.
27. The method for training a model according to claim 26, wherein: The target prediction sequence includes: a prediction sequence of the initial model for the sample three-dimensional scene at a first moment; and The historical prediction sequence includes: a prediction sequence of the initial model for the sample three-dimensional scene in a first time sequence; The first moment is after the first time series.
28. A method for establishing a three-dimensional model, comprising: Determining a feature sequence corresponding to the three-dimensional scene to be modeled according to prompt information corresponding to the three-dimensional scene to be modeled; Determining a target prediction sequence corresponding to the three-dimensional scene to be modeled using the target model and the feature sequence; and According to the target prediction sequence, a three-dimensional model corresponding to the three-dimensional scene to be modeled is established; wherein the target model is trained using the model training method according to any one of claims 16-27.
29. The method for establishing a three-dimensional model according to claim 28, wherein: The prompt information includes at least one of a two-dimensional image corresponding to the three-dimensional scene to be modeled, a floor plan corresponding to the three-dimensional scene to be modeled, a three-dimensional model of a first target object contained in the three-dimensional scene to be modeled, bounding box information of a second target object contained in the three-dimensional scene to be modeled, text description information, point cloud information, or voxel information.
30. The method for establishing a three-dimensional model according to claim 29, wherein: Determining the feature sequence according to the prompt information includes at least one of the following: In a case where the prompt information includes the two-dimensional image, extracting features from the two-dimensional image based on a first encoder to obtain a first word list; In a case where the prompt information includes the floor plan, converting the floor plan into a binary mask, and performing feature extraction on the binary mask based on a second encoder to obtain a second word list; In a case where the prompt information includes a three-dimensional model of the first target object, rendering the three-dimensional model of the first target object to obtain a rendering of the first target object, and performing feature extraction on the rendering based on a third encoder to obtain a third word list; In a case where the prompt information includes the bounding box information of the second target object, encoding the value in the bounding box information of the second target object based on a set encoding rule to obtain a fourth word list; In a case where the prompt information includes the text description information, performing feature extraction on the text description information based on a fourth encoder to obtain a fifth word-gram list; When the prompt information includes the point cloud information, performing feature extraction on the point cloud information based on a fifth encoder to obtain a sixth word list; or In a case where the prompt information includes the voxel information, performing feature extraction on the voxel information based on a sixth encoder to obtain a seventh word list; The feature sequence includes the first word list, the second word list, the third word list, at least one of the first word list, the fourth word list, the fifth word list, the sixth word list, or the seventh word list.
31. A model training device, comprising: A first acquisition module is configured to determine a feature sequence corresponding to the sample three-dimensional scene according to prompt information corresponding to the sample three-dimensional scene; A first determination module is configured to determine a target prediction sequence corresponding to the sample three-dimensional scene using an initial model and the feature sequence; wherein the target prediction sequence is used to establish a three-dimensional model corresponding to the sample three-dimensional scene; as well as The adjustment module is configured to adjust the initial model using the target prediction sequence and the target label sequence of the sample three-dimensional scene to obtain a target model.
32. A training device for a model according to claim 31, wherein: The prompt information includes at least one of a two-dimensional image corresponding to the sample three-dimensional scene, a floor plan corresponding to the sample three-dimensional scene, a three-dimensional model of a first target object contained in the sample three-dimensional scene, bounding box information of a second target object contained in the sample three-dimensional scene, text description information, point cloud information, or voxel information.
33. A training device for a model according to claim 32, wherein: The first acquisition module is configured to perform at least one of the following: In a case where the prompt information includes the two-dimensional image, extracting features from the two-dimensional image based on a first encoder to obtain a first word list; In a case where the prompt information includes the floor plan, converting the floor plan into a binary mask, and performing feature extraction on the binary mask based on a second encoder to obtain a second word list; In a case where the prompt information includes a three-dimensional model of the first target object, rendering the three-dimensional model of the first target object to obtain a rendering of the first target object, and performing feature extraction on the rendering based on a third encoder to obtain a third word list; In a case where the prompt information includes the bounding box information of the second target object, encoding the value in the bounding box information of the second target object based on a set encoding rule to obtain a fourth word list; In a case where the prompt information includes the text description information, performing feature extraction on the text description information based on a fourth encoder to obtain a fifth word-gram list; When the prompt information includes the point cloud information, performing feature extraction on the point cloud information based on a fifth encoder to obtain a sixth word list; or In a case where the prompt information includes the voxel information, performing feature extraction on the voxel information based on a sixth encoder to obtain a seventh word list; Wherein, the feature sequence includes at least one of the first word-gram list, the second word-gram list, the third word-gram list, the fourth word-gram list, the fifth word-gram list, the sixth word-gram list or the seventh word-gram list.
34. The training device of the model according to claim 31 further comprises: A second determination module is configured to determine the target tag sequence according to the sample three-dimensional scene; The target label sequence includes a real sequence corresponding to the sample three-dimensional scene, and the real sequence is used to establish a real three-dimensional model corresponding to the sample three-dimensional scene.
35. A training device for a model according to claim 34, wherein: The second determining module is configured to: Determine an initial label sequence according to the sample three-dimensional scene; the initial label sequence includes a first initial label sequence and / or a second initial label sequence; and The initial label sequence is sequentially subjected to discretization processing and feature embedding processing to obtain the target label sequence.
36. A device for building a three-dimensional model, comprising: A second acquisition module is configured to determine a feature sequence corresponding to the three-dimensional scene to be modeled according to prompt information corresponding to the three-dimensional scene to be modeled; A third determination module is configured to determine a target prediction sequence corresponding to the three-dimensional scene to be modeled using the target model and the feature sequence; as well as A modeling module is configured to establish a three-dimensional model corresponding to the three-dimensional scene to be modeled according to the target prediction sequence; wherein the target model is a model according to any one of claims 31-35 Obtained through training with a training device.
37. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1-11 and 16-30.
38. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause a computer to execute the method according to any one of claims 1-11 and 16-30.
39. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-11 and 16-30.