Three-dimensional scene reconstruction method and apparatus, and cluster, product and storage medium
Patent Information
- Application Number
- PCT/CN2026/076711
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-28
- Filing Date
- 2026-02-03
- Publication Date
- 2026-09-03
Smart Images

Figure CN2026076711_03092026_PF_FP_ABST
Abstract
Description
A method, apparatus, cluster, product, and storage medium for reconstructing three-dimensional scenes.
[0001] This application claims priority to Chinese Patent Application No. 202510249321.3, filed with the State Intellectual Property Office of China on February 28, 2025, entitled “A method, apparatus, cluster, product and storage medium for three-dimensional scene reconstruction”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of computer technology, and in particular to a method, apparatus, cluster, product, and storage medium for reconstructing three-dimensional scenes. Background Technology
[0003] 3D scene reconstruction involves identifying a computer-aided design (CAD) model that closely resembles objects in a real 3D scene, and then arranging the CAD model according to the orientation of the objects in the real 3D scene to create a virtual 3D scene with the same layout as the real 3D scene.
[0004] Common methods for reconstructing 3D scenes include: manually labeling the categories of objects in the real 3D scene; selecting a fixed-category detector (e.g., 3D-SIS) that can identify objects of that category based on the labeled categories; using this fixed-category detector to identify the objects in the real 3D scene; then, using a graph neural network to aggregate and extract the geometric features of the identified objects; and based on these geometric features, retrieving the target CAD model with the highest similarity to the object from multiple CAD models (referred to as a small-scale CAD asset library) belonging to the categories specified by the user; finally, reconstructing the virtual 3D scene based on this target CAD model.
[0005] However, the above-mentioned three-dimensional scene reconstruction method requires manual specification of the small-scale CAD asset library for objects in the real three-dimensional scene, which reduces the efficiency of reconstructing virtual three-dimensional scenes. Summary of the Invention
[0006] This application provides a method, apparatus, cluster, product, and storage medium for reconstructing three-dimensional scenes, which can improve the efficiency of reconstructing virtual three-dimensional scenes.
[0007] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:
[0008] In a first aspect, embodiments of this application provide a method for reconstructing a three-dimensional scene. The method includes: acquiring a global feature vector and a textual feature vector of an object in a real three-dimensional scene; wherein the global feature vector is generated based on multi-view images of the object and is used to characterize the overall features of the object; the textual feature vector of the object is generated based on the semantic text of the object and is used to characterize the overall or local features of the object; based on the global feature vector and the textual feature vector of the object, determining a target CAD model from a CAD asset library that satisfies the similarity condition to the object; wherein the CAD asset library includes multiple CAD models, which are CAD models corresponding to various categories of objects; and generating a virtual three-dimensional scene corresponding to the real three-dimensional scene based on the target CAD model; the virtual three-dimensional scene includes virtual objects corresponding to the objects.
[0009] This application provides a method for reconstructing a 3D scene. This method determines a target CAD model that meets a certain similarity condition from multiple CAD models (i.e., a large number of rich CAD models) corresponding to objects of various categories, based on the global feature vector and text feature vector of an object in a real 3D scene. Then, it generates a virtual 3D scene corresponding to the real 3D scene based on the target CAD model. Since the global feature vector of the object is used to represent the overall features of the object, and the global feature vector is generated based on the object's multi-view images, and the text feature vector is generated based on the object's semantic text, and the semantic text is used to represent the object's overall or local features, this application determines the target CAD model from the large number of rich CAD models based on the multimodal features of the object, without requiring user intervention. Therefore, it improves the efficiency of reconstructing a virtual 3D scene.
[0010] Furthermore, compared to traditional techniques that retrieve target CAD models based on the geometric features of objects, this application determines the target CAD model from the massive CAD models based on the multimodal features of the aforementioned objects. These multimodal features integrate visual information from multiple perspective images of the object and the object's semantic text (e.g., descriptive information of the semantic text). The visual information and the semantic text complement each other in the process of retrieving the target CAD model (i.e., retrieving the target CAD model based on both visual information and semantic text modal features). Therefore, it can more quickly determine the target CAD model from the rich and massive CAD models, thus improving the efficiency of retrieving the target CAD model.
[0011] In one possible implementation, the above-mentioned determination of a target CAD model from a CAD asset library that meets the similarity condition of the object based on the object's global feature vector and text feature vector includes: calculating the similarity between the object's global feature vector and the feature vector of each of the multiple CAD models, and determining multiple candidate CAD models from the multiple CAD models; obtaining the object's target features; wherein the object's target features include at least one of the following: visual features of each viewpoint image in the object's multi-view images, or geometric features of the object; visual features of one viewpoint image in the object's multi-view images used to represent the overall features of the viewpoint image; geometric features of the object used to describe the object's shape; and determining a target CAD model from the multiple candidate CAD models based on the object's target features.
[0012] The above embodiments calculate the similarity between the object's global feature vector and textual feature vector and the feature vector of each of the multiple CAD models, respectively, to roughly retrieve (referred to as "coarse retrieval") multiple candidate CAD models with high similarity from a massive number of CAD models. Then, based on the visual features of the object's image from each viewpoint and / or the object's geometric features, the target CAD model with the highest similarity is precisely retrieved (referred to as "fine retrieval") from these multiple candidate CAD models, thereby improving the accuracy of the target CAD model.
[0013] In one possible implementation, the similarity calculation between the object's global feature vector and text feature vector and the feature vector of each of the multiple CAD models is performed to determine multiple candidate CAD models from these multiple CAD models. This includes: calculating the similarity between the object's global feature vector and the feature vector of each of the multiple CAD models to determine a first model set; the first model set includes M CAD models, which are selected from those with the highest similarity between the object's global feature vector and the feature vector of each of the multiple CAD models, in descending order of similarity. The CAD models corresponding to the top M similarity scores are sorted by size; M is greater than or equal to 1. Based on the similarity calculation between the text feature vector of the object and the feature vector of each of the multiple CAD models, a second model set is determined. The second model set includes N CAD models, which are the CAD models corresponding to the top N similarity scores between the text feature vector of the object and the feature vector of each of the multiple CAD models, sorted from largest to smallest; N is greater than or equal to 1. The union of the first model set and the second model set is obtained. The CAD models in the union are the candidate CAD models.
[0014] In one possible implementation, based on the target features of the aforementioned object, a target CAD model is determined from the plurality of candidate CAD models, including: determining P candidate CAD models based on the visual features of the object's multi-view images and the visual features of the multi-view images of each candidate CAD model in the plurality of candidate CAD models through a similarity voting strategy; where P is greater than or equal to 2; wherein, the visual features of the object's multi-view images refer to the visual features of each view in the object's multi-view images; the similarity voting strategy is used to indicate that the candidate CAD models among the plurality of candidate CAD models whose number of target visual features meets the condition are selected as candidate CAD models; the target visual feature is the visual feature with the highest similarity to a visual feature of the object among the visual features of the multi-view images of the plurality of candidate CAD models; and determining the target CAD model from the P candidate CAD models, wherein the similarity between the geometric features of the target CAD model and the geometric features of the object meets the condition.
[0015] This embodiment first determines P candidate CAD models from a plurality of candidate CAD models based on the visual features of the object's image from each viewpoint; this is equivalent to further narrowing down the scope of the target CAD model search based on the candidate CAD models. Then, based on the similarity between the geometric features of each of the P candidate CAD models and the geometric features of the object, the target CAD model is determined from the P candidate CAD models after the search scope has been narrowed down. This avoids the problem of low accuracy in determining the target CAD model due to ignoring the differences between the image of the object in the real 3D scene and the CAD model in the CAD asset library from different viewpoints, as well as the differences between the pose parameters of the object and the pose of the CAD module, thus improving the accuracy of the target CAD model.
[0016] In one possible implementation, determining the target CAD model from P candidate CAD models includes: aligning the target parameters of each of the P candidate CAD models with the target parameters of the object, and obtaining the 3D point cloud of each candidate CAD model after the target parameters are aligned; the target parameters include pose parameters, or the target parameters include pose parameters and relative scale; determining the chamfer distance and relative scale difference between the 3D point cloud of each candidate CAD model after the target parameters are aligned and the 3D point cloud of the object; and determining the target CAD model corresponding to the 3D point cloud of the model whose sum of chamfer distance and relative scale difference satisfies the condition from the P candidate CAD models.
[0017] This embodiment determines P candidate CAD models from a plurality of candidate CAD models based on the visual features of the object's image from each viewpoint. Then, the chamfer distance and relative scale difference between the 3D point cloud of each candidate CAD model and the 3D point cloud of the object after target parameter alignment are determined respectively, and the target CAD model is determined based on the sum of the chamfer distance and relative scale difference. It can be seen that the above embodiment improves the accuracy of the target CAD model by combining the chamfer distance and relative scale difference to determine the target CAD model.
[0018] In one possible implementation, before obtaining the global feature vector and text feature vector of the object in the real 3D scene, the method further includes: obtaining a multi-view image of the real 3D scene; inputting the multi-view image of the real 3D scene into a 3D scene instance segmentation model to obtain a multi-view image of the object; inputting the multi-view image of the object into a language vision model to obtain the semantic text of the object.
[0019] The above embodiments use a 3D scene instance segmentation model to segment each view image in a multi-view image of a real 3D scene, obtaining multi-view images of objects in the real 3D scene; and inputting the multi-view images of the objects into a language vision model to obtain semantic text of the objects. Since this 3D scene instance segmentation model supports segmentation of open-category objects, it can identify and segment multi-view images of any category of objects in a real 3D scene without requiring manual labeling of the objects, thus improving the efficiency of virtual 3D scene reconstruction.
[0020] In one possible implementation, the above-mentioned generation of a virtual 3D scene corresponding to the real 3D scene based on the target CAD model includes: configuring the target CAD model and generating a virtual object according to the pose parameters of the object and the relative scale of the object.
[0021] The above embodiments configure a target CAD model based on the object's pose parameters and relative scale to generate a virtual object, ensuring that the position, orientation, and relative scale (i.e., 9-DOF pose) of the target CAD model are identical to the object. This more accurately recreates objects in a real 3D scene, thus improving the accuracy of the virtual 3D scene. In one possible implementation, the virtual 3D scene includes multiple target CAD models corresponding to multiple objects. The method further includes: obtaining a first CAD model and a second CAD model with a support relationship from the multiple target CAD models; the support surface of the first CAD model is used to support the bottom surface of the second CAD model; if the height difference between the bottom surface of the second CAD model and the support surface is greater than 0, the second CAD model is moved downwards by the height difference.
[0022] The above embodiment determines whether the second CAD model is suspended relative to the first CAD model by the height difference between the bottom surface of the second CAD model, which has a supporting relationship with the first CAD model in the virtual 3D scene; and if it is suspended, the second CAD model is moved downward by the height difference; thereby solving the suspension problem in the virtual 3D scene and improving the accuracy of the virtual 3D scene.
[0023] In one possible implementation, obtaining the first CAD model and the second CAD model with a support relationship from multiple target CAD models includes: obtaining a third CAD model located above the support surface of the first CAD model and without any other CAD models between it and the first CAD model; wherein the multiple target CAD models include the third CAD model; determining the overlapping area of the projection of the bottom surface of the third CAD model onto the horizontal plane and the projection of the support surface onto the horizontal plane; determining the second CAD model from the third CAD model based on the overlapping area corresponding to the third CAD model; wherein the ratio of the overlapping area corresponding to the second CAD model to the area of the bottom surface of the second CAD model is greater than a preset ratio.
[0024] The above embodiment first obtains a third CAD model located above the support surface of the first CAD model, and without any other CAD models between it and the first CAD model. Then, based on the overlapping area of the projection of the bottom surface of the third CAD model onto the horizontal plane and the projection of the support surface of the first CAD model onto the horizontal plane, a second CAD model is determined from the third CAD model; the ratio of the overlapping area of the second CAD model to the area of its bottom surface is greater than a preset ratio. This avoids identifying a CAD model above the first CAD model and staggered from it on the horizontal plane as the second CAD model; therefore, the accuracy of the second CAD model is improved.
[0025] In one possible implementation, the method further includes: obtaining an overlapping model pair from a virtual 3D scene; the overlapping model pair consists of two CAD models with an overlapping space volume greater than 0; the overlapping model pair includes a fourth CAD model and a fifth CAD model; determining a CAD model to be moved from the fourth CAD model and the fifth CAD model; moving the CAD model to be moved; and the moved fourth CAD model and the fifth CAD model do not have an overlapping volume.
[0026] In one possible implementation, the above-mentioned moving of the CAD model to be moved includes: determining the target distance by the sum of the side length of the side perpendicular to the maximum plane in the three-dimensional bounding box of the overlapping space of the fourth CAD model and the fifth CAD model and a preset length; the maximum plane is the maximum plane in the three-dimensional bounding box of the overlapping space; and moving the CAD model to be moved by the target distance based on the pointing direction; the pointing direction is the direction from the non-moving CAD model to the CAD model to be moved.
[0027] The above embodiment determines the CAD model to be moved from an overlapping model pair in a virtual 3D scene. This overlapping model pair includes a fourth CAD model and a fifth CAD model. Then, the target distance is determined by the sum of the length of the side perpendicular to the largest face of the 3D bounding box of the overlapping space of the fourth and fifth CAD models and a preset length. Finally, the CAD model to be moved is moved by this target distance in the direction that the non-moving CAD model is pointing towards; this separates the overlapping fourth and fifth CAD models, thereby improving the accuracy of the virtual 3D scene.
[0028] In one possible implementation, the above-mentioned method of moving the CAD model to be moved by a target distance based on the pointing direction includes: determining the vertical direction of the largest plane as the direction to be moved; the direction to be moved includes: a positive direction and a negative direction; the positive direction and the negative direction are opposite in direction; determining the target direction from the positive direction and the negative direction based on the pointing direction, the angle between the target direction and the pointing direction is smaller than the angle between the non-moving direction and the pointing direction; and moving the CAD model to be moved by a target distance in the target direction.
[0029] The above embodiment determines the direction to be moved by the vertical direction of the largest plane in the three-dimensional bounding box of the overlapping space; then, the positive and negative directions in the direction to be moved, respectively, with the smaller angle between them and the pointing direction, are determined as the target directions; thereby avoiding the problem that the CAD model to be moved will move to the wrong oblique direction because the fourth CAD model is on the oblique side of the fifth CAD model (e.g., left oblique, right oblique, upper oblique, or lower oblique); thus improving the accuracy of the virtual three-dimensional scene.
[0030] Secondly, embodiments of this application provide a three-dimensional scene reconstruction device, which includes: an acquisition module and a processing module; the acquisition module is used to acquire global feature vectors and text feature vectors of objects in a real three-dimensional scene; wherein, the global feature vector is generated based on multi-view images of the object and is used to characterize the overall features of the object; the text feature vector of the object is generated based on the semantic text of the object and is used to characterize the overall or local features of the object; the processing module is used to determine a target CAD model that meets the similarity condition of the object from a CAD asset library based on the global feature vector and the text feature vector of the object; wherein, the CAD asset library includes multiple CAD models, which are CAD models corresponding to various categories of objects; the processing module is also used to generate a virtual three-dimensional scene corresponding to the real three-dimensional scene based on the target CAD model; the virtual three-dimensional scene includes virtual objects corresponding to the objects.
[0031] In one possible implementation, the processing module is used to calculate the similarity between the object's global feature vector and text feature vector and the feature vector of each of the multiple CAD models, respectively, to determine multiple candidate CAD models from the multiple CAD models; the acquisition module is used to acquire the object's target features; wherein, the object's target features include at least one of the following: visual features of each view image in the object's multi-view images, or geometric features of the object; the visual features of one view image in the object's multi-view images are used to represent the overall features of the view image; the geometric features of the object are used to describe the features of the object's shape; the processing module is used to determine the target CAD model from the multiple candidate CAD models based on the object's target features.
[0032] In one possible implementation, the processing module is used to calculate the similarity between the global feature vector of the object and the feature vector of each of the multiple CAD models to determine a first model set. This first model set includes M CAD models, which are the top M CAD models corresponding to the highest similarity scores between the global feature vector of the object and the feature vector of each of the multiple CAD models, sorted from largest to smallest; M is greater than or equal to 1. The processing module is also used to calculate the similarity between the text feature vector of the object and the feature vector of each of the multiple CAD models to determine a second model set. This second model set includes N CAD models, which are the top N CAD models corresponding to the highest similarity scores between the text feature vector of the object and the feature vector of each of the multiple CAD models, sorted from largest to smallest; N is greater than or equal to 1. The processing module is used to obtain the union of the first model set and the second model set; where the CAD models in the union are candidate CAD models.
[0033] In one possible implementation, the processing module is used to determine P candidate CAD models based on the visual features of multi-view images of an object and the visual features of multi-view images of each candidate CAD model among multiple candidate CAD models, using a similarity voting strategy; P is greater than or equal to 2; wherein, the visual features of the multi-view images of the object refer to the visual features of each view in the multi-view images of the object; the similarity voting strategy is used to indicate that the candidate CAD models among the multiple candidate CAD models whose number of target visual features meets the condition are selected as candidate CAD models; the target visual feature is the visual feature with the highest similarity to a visual feature of the object among the visual features of the multi-view images of the multiple candidate CAD models; the processing module is also used to determine a target CAD model from the P candidate CAD models, wherein the geometric features of the target CAD model satisfy the similarity condition with the geometric features of the object.
[0034] In one possible implementation, the processing module is used to align the target parameters of P candidate CAD models with the target parameters of the object, and obtain the 3D point cloud of each candidate CAD model after target parameter alignment; the target parameters include: pose parameters, or the target parameters include: pose parameters and relative scale; the processing module is used to determine the chamfer distance and relative scale difference between the 3D point cloud of each candidate CAD model after target parameter alignment and the 3D point cloud of the object; the processing module is used to determine the target CAD model corresponding to the 3D point cloud of the model whose sum of chamfer distance and relative scale difference satisfies the condition from the P candidate CAD models.
[0035] In one possible implementation, the acquisition module is used to acquire multi-view images of a real 3D scene; the processing module is used to input the multi-view images of the real 3D scene to a 3D scene instance segmentation model to obtain multi-view images of the object; and the processing module is used to input the multi-view images of the object to a language vision model to obtain the semantic text of the object.
[0036] In one possible implementation, the processing module is used to configure the target CAD model and generate a virtual object based on the pose parameters of the object and the relative scale of the object.
[0037] In one possible implementation, the processing module is used to obtain a first CAD model and a second CAD model with a support relationship from multiple target CAD models; the support surface of the first CAD model is used to support the bottom surface of the second CAD model; and the processing module is used to move the second CAD model downward by the height difference when the height difference between the bottom surface of the second CAD model and the support surface is greater than 0.
[0038] In one possible implementation, the acquisition module is used to acquire a third CAD model located above the support surface of the first CAD model and without any other CAD models between it and the first CAD model; wherein the third CAD model is included among multiple target CAD models; the processing module is used to determine the overlapping area of the projection of the bottom surface of the third CAD model onto the horizontal plane and the projection of the support surface onto the horizontal plane; the processing module is used to determine a second CAD model from the third CAD model based on the overlapping area corresponding to the third CAD model; the ratio of the overlapping area corresponding to the second CAD model to the bottom surface area of the second CAD model is greater than a preset ratio.
[0039] In one possible implementation, the acquisition module is used to acquire overlapping model pairs from the virtual 3D scene; the overlapping model pair is two CAD models with an overlapping space volume greater than 0; the overlapping model pair includes: a fourth CAD model and a fifth CAD model; the processing module is used to determine the CAD model to be moved from the fourth CAD model and the fifth CAD model; the processing module is used to move the CAD model to be moved; the fourth CAD model and the fifth CAD model after moving do not have overlapping volumes.
[0040] In one possible implementation, the processing module is used to determine the target distance by summing the side length of the edge perpendicular to the maximum plane in the three-dimensional bounding box of the overlapping space of the fourth CAD model and the fifth CAD model with a preset length; the maximum plane is the maximum plane in the three-dimensional bounding box of the overlapping space; the processing module is used to move the CAD model to be moved by the target distance based on the pointing direction; the pointing direction is the direction from which the non-moving CAD model points to the CAD model to be moved.
[0041] In one possible implementation, the processing module is used to determine the vertical direction of the maximum plane as the direction to be moved; the direction to be moved includes: a positive direction and a negative direction; the positive direction and the negative direction are opposite in direction; the processing module is used to determine the target direction from the positive direction and the negative direction based on the pointing direction, the angle between the target direction and the pointing direction is smaller than the angle between the non-moving direction and the pointing direction; the processing module is used to move the CAD model to be moved by a target distance in the target direction.
[0042] Thirdly, this application provides a computing device cluster including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, such that the computing device cluster performs the method described in the first aspect and any of its possible implementations.
[0043] Fourthly, this application provides a computer-readable storage medium having computer instructions stored thereon, which, when executed on a computing device, cause the computing device to perform the method described in any one of the first aspects and its possible implementations.
[0044] Fifthly, this application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the method described in any one of the first aspects and their possible implementations.
[0045] It should be understood that the beneficial effects of the technical solutions of the second to fifth aspects of this application and the corresponding possible implementations can be referred to the above-described technical effects of the first aspect and its corresponding possible implementations, and will not be repeated here. Attached Figure Description
[0046] Figure 1 is a schematic diagram of the architecture of a cloud platform provided in an embodiment of this application;
[0047] Figure 2 is a schematic diagram of a three-dimensional scene reconstruction system provided in an embodiment of this application;
[0048] Figure 3 is a flowchart illustrating a three-dimensional scene reconstruction method provided in an embodiment of this application;
[0049] Figure 4 is a schematic flowchart of a method for determining a target CAD model provided in an embodiment of this application;
[0050] Figure 5 is a schematic flowchart of a method for determining candidate CAD models provided in an embodiment of this application;
[0051] Figure 6 is a schematic flowchart of another method for determining a target CAD model provided in an embodiment of this application;
[0052] Figure 7 is a flowchart illustrating a method for handling suspended CAD models according to an embodiment of this application;
[0053] Figure 8 is a schematic flowchart of a method for determining a second CAD model provided in an embodiment of this application;
[0054] Figure 9 is a schematic diagram of a suspended structure provided in an embodiment of this application;
[0055] Figure 10 is a schematic diagram of a support relationship provided in an embodiment of this application;
[0056] Figure 11 is a schematic diagram of CAD model overlap processing provided in an embodiment of this application;
[0057] Figure 12 is a schematic flowchart of a CAD model overlap processing method provided in an embodiment of this application;
[0058] Figure 13 is a schematic flowchart of another CAD model overlap processing method provided in an embodiment of this application;
[0059] Figure 14 is a schematic diagram of a three-dimensional scene reconstruction device provided in an embodiment of this application;
[0060] Figure 15 is a hardware schematic diagram of a computing device provided in an embodiment of this application;
[0061] Figure 16 is a schematic diagram of a computing device cluster provided in an embodiment of this application;
[0062] Figure 17 is a schematic diagram of a network connection provided in an embodiment of this application. Detailed Implementation
[0063] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0064] The terms "first CAD model" and "second CAD model," etc., used in the specification and claims of this application are used to distinguish different CAD models, rather than to describe a specific order of CAD models. For example, "first model set" and "second model set," etc., are used to distinguish different model sets, rather than to describe a specific order of model sets.
[0065] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0066] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more. For example, multiple CAD models refer to two or more CAD models.
[0067] Open category objects: These are objects or objects belonging to categories that do not exist in the model's training data.
[0068] Similarity voting strategy: A similarity score is obtained by calculating the similarity between each object to be processed and a reference object. Then, a vote is held based on these similarity scores to determine which object will be selected for final processing.
[0069] Cosine similarity: also known as cosine similarity, it is used to evaluate the similarity between two vectors by calculating the cosine value of the angle between them; the larger the cosine value, the higher the similarity between the two vectors; the smaller the cosine value, the lower the similarity between the two vectors.
[0070] RGB-D (red breen blue depth) camera: An RGB-D camera is a device that can capture color images and depth information simultaneously.
[0071] 6-DOF pose: An object has six degrees of freedom in space, namely the degree of freedom of movement along the three rectangular coordinate axes x, y, and z, and the degree of freedom of rotation about these three coordinate axes.
[0072] Language vision model: This is a model that connects a visual encoder and a large language model to achieve general vision and language understanding.
[0073] Contrastive language-image pre-training (CLIP) model: A multimodal contrastive learning model that maps text and images to the same semantic space to achieve cross-modal understanding.
[0074] Large language and vision assistant (LLaVA) model: This is a multimodal large model that integrates vision and language to achieve visual reasoning and natural language interaction in complex scenarios.
[0075] Based on the background technology, the commonly used 3D scene reconstruction methods require manually specifying a small-scale CAD asset library for objects in the real 3D scene, resulting in low efficiency in reconstructing virtual 3D scenes.
[0076] Based on this, embodiments of this application provide a three-dimensional scene reconstruction method. This method determines a target CAD model that meets the similarity criteria to an object from multiple CAD models (i.e., a large number of rich CAD models) corresponding to objects of various categories, based on the global feature vector and text feature vector of an object in a real three-dimensional scene. Then, it generates a virtual three-dimensional scene corresponding to the real three-dimensional scene based on the target CAD model. This is because the global feature vector of the object is used to represent the overall features of the object, and the global feature vector is generated based on the object's multi-view images; and the text feature vector of the object is generated based on the object's semantic text, and the semantic text is used to represent the object's overall or local features.
[0077] As can be seen, this application determines the target CAD model from the rich CAD models based on the multimodal characteristics of the aforementioned objects, without requiring user participation, thus improving the efficiency of reconstructing virtual 3D scenes.
[0078] Furthermore, compared to traditional techniques that retrieve target CAD models based on the geometric features of objects, this application determines the target CAD model from the massive CAD models based on the multimodal features of the aforementioned objects. These multimodal features integrate visual information from multiple perspective images of the object and the object's semantic text (e.g., descriptive information of the semantic text). The visual information and the semantic text complement each other in the process of retrieving the target CAD model (i.e., retrieving the target CAD model based on both visual information and semantic text modal features). Therefore, it can more quickly determine the target CAD model from the rich and massive CAD models, thus improving the efficiency of retrieving the target CAD model.
[0079] The three-dimensional scene reconstruction method provided in this application can be applied to the cloud platform shown in Figure 1. Figure 1 is a schematic diagram of the architecture of a cloud platform provided in this application. As shown in Figure 1, the cloud platform 100 includes a computing server cluster 110, a storage server cluster 120, a management server cluster 130, a network device cluster 140, and tenant terminals 150. The computing server cluster 110, storage server cluster 120, and management server cluster 130 communicate with the tenant terminals 150 through the network device cluster 140.
[0080] The computing server cluster 110 serves as a computing resource within the cloud platform 100. The computing server cluster 110 includes one or more computing servers (computing server 111 and computing server 112 shown in Figure 1). The computing servers can be electronic devices with data processing capabilities, including but not limited to servers and desktop computers. The computing servers are used to generate and allocate computing resources based on virtualization technology and tenant needs. In terms of hardware, the computing servers are equipped with processors and memory (not shown in Figure 1), and the computing functions of the computing servers are implemented by the processor running programs in memory. The computing servers can also read / write data on various storage servers in the storage server cluster 120 according to tenant needs.
[0081] Storage server cluster 120 serves as a storage resource within cloud platform 100. Storage server cluster 120 includes one or more storage servers (storage server 121 and storage server 122 as shown in Figure 1). Storage servers can be electronic devices with data storage capabilities, including but not limited to: servers, desktop computers, or controllers and hard disk enclosures of storage arrays. Storage servers provide logical disk storage, unstructured data storage, and consolidated backup services for cloud virtual machines within cloud platform 100. In terms of hardware, storage servers are equipped with network interface cards (NICs), processors, and memory. The processor in the storage server processes data from outside the storage server. The NIC controls the access process to the memory, such as by controlling address signals, data signals, and various command signals, enabling the storage server to provide the memory as a storage resource to tenants. Memory is used to store data and may include RAM and / or hard disks. RAM refers to internal memory that directly exchanges data with the processor; RAM can quickly read and write data at any time, serving as temporary data storage for the operating system or other running programs. Unlike RAM, hard disks are slower to read and write data and are typically used for persistent data storage.
[0082] The management server cluster 130 is used to manage all computing services, shared storage, and network of the entire cloud platform 100, and to provide tenants or administrators with an API to manage the entire node. The management server cluster 130 includes one or more management servers (management server 131 and management server 132 as shown in Figure 1). Similarly, management servers can be electronic devices with data processing capabilities, including but not limited to servers, desktop computers, etc.
[0083] Network device cluster 140 includes one or more network devices. These network devices include, but are not limited to, switches and routers. For example, network device cluster 140 includes router 141, switch 142, switch 143, switch 144, and switch 145. Tenant terminal 150 is connected to router 141 via the Internet. Router 141 is connected to switch 143 via switch 142. Switch 143 is connected to each computing server in computing server cluster 110. Switch 144 is connected to each computing server in computing server cluster 110 and each storage server in storage server cluster 120. Switch 145 is connected to each computing server in computing server cluster 110, each storage server in storage server cluster 120, and each management server in management server cluster 130.
[0084] The number and type of network devices included in the network device cluster 140 can be adjusted according to the needs of the cloud platform 100. For example, switches 142, 143, 144, and 145 can be switches with different functions, such as switch 142 being the core switch, and switches 143, 144, and 145 being switches used to manage specific network segments (e.g., switch 143 can be an internal / external switching network segment switch, switch 144 can be a storage network segment switch, and switch 145 can be a management network segment switch).
[0085] Tenant terminal 150 includes one or more tenant terminals (tenant terminal 151 and tenant terminal 152 as shown in Figure 1). The tenant terminal contains the interfaces and applications required to access cloud platform 100.
[0086] It is worth noting that Figure 1 is merely a schematic diagram and should not be construed as limiting this application. The cloud platform 100 may include more or fewer devices, and this application does not limit this. For example, the computing server cluster 110 of the cloud platform 100 may include more or fewer computing servers. Similarly, the storage server cluster 120 of the cloud platform 100 may include more or fewer storage servers. Furthermore, the management server cluster 130 of the cloud platform 100 may include more or fewer management servers. Moreover, the network device cluster 140 of the cloud platform 100 may include more or fewer network devices. Finally, the tenant terminal 150 of the cloud platform 100 may include more or fewer tenant terminals. And, depending on the needs of actual applications, the cloud platform may also include other devices; this application does not limit the specific content included in the cloud platform 100.
[0087] Based on the equipment of the cloud platform 100 shown in Figure 1, the cloud platform 100 can implement the functions of service nodes based on IaaS, PSSS, and SaaS. The cloud platform 100 provides computing services, storage services, network services, and other services to tenants through these service nodes. Service nodes can be cloud nodes (such as control nodes, worker nodes, etc.) obtained by virtualizing the resources (such as computing resources and storage resources) of the cloud platform 100.
[0088] Figure 2 is a system architecture diagram of the three-dimensional scene reconstruction system provided in an embodiment of this application. The system includes a reconstruction system 201 and a display system 202.
[0089] In one implementation, the display system 202 can be a tenant terminal in the tenant terminal 150 shown in Figure 1; the reorganization system 201 can be a computing server in the computing server cluster 110 shown in Figure 1.
[0090] The recombination system 201 is used to extract features of objects in the acquired multi-view images of a real 3D scene, and based on these features, to retrieve a target CAD model from a CAD resource library containing a massive number of CAD models that meets the similarity criteria (e.g., the highest similarity). Finally, a virtual 3D scene that is the same as or similar to the real 3D scene is generated based on the target CAD model; the specific implementation process is described in S110-S170 below, and will not be repeated here.
[0091] The reconfiguration system 201 is a computing device with image processing and computing functions. Specifically, the reconfiguration system 201 can be a server (such as a cloud server or physical server), a computer, or a mobile phone.
[0092] It should be noted that the reassembly system 201 can obtain the multi-view images of the real three-dimensional scene from the storage server in the storage server cluster 120 shown in Figure 2, or it can obtain the multi-view images of the real three-dimensional scene acquired by the acquisition system. Specifically, this application embodiment does not limit the implementation method of the reassembly system 201 to obtain the multi-view images of the real three-dimensional scene.
[0093] In the case where the reconstruction system 201 acquires multi-view images of a real 3D scene from the acquisition system, the aforementioned 3D scene reconstruction system further includes an acquisition system 203. The acquisition system 203 is used to acquire images (referred to as multi-view images) of a real 3D scene (such as a classroom or a room) from different perspectives, and sends the multi-view images of the real 3D scene to the reconstruction system 201.
[0094] Among them, the acquisition system 203 is a system or device with image acquisition function; specifically, the acquisition system 203 can be an RGB-D camera that can simultaneously capture color images and depth information, or it can be a LiDAR and a regular camera.
[0095] The display system 202 is used to receive the virtual three-dimensional scene sent by the reconstruction system 201 and display the virtual three-dimensional scene.
[0096] Among them, the display system 202 is a device with display function; specifically, the display system 202 can be a mobile phone, a large screen, or a computer monitor, etc.
[0097] The three-dimensional scene reconstruction method provided in this application is applied to the reconstruction system 201 in the scene of generating a virtual three-dimensional scene from multi-view images of a real three-dimensional scene.
[0098] It should be noted that the system architecture and application scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0099] This application provides a three-dimensional scene reconstruction method, which is applied to a reconstruction system 201 (hereinafter referred to as the reconstruction system), as shown in Figure 3. The method includes: S110-S170.
[0100] S110: Obtain multi-view images of a real 3D scene.
[0101] In this application, objects in a real 3D scene include one or more objects in that real 3D scene.
[0102] In this application, S110 can be implemented by acquiring multi-view images of the real three-dimensional scene captured by the ordinary camera from an acquisition device or acquisition system (such as an ordinary camera); or by acquiring multi-view images of the real three-dimensional scene from other devices; or by acquiring multi-view images of the real three-dimensional scene locally from the reconstruction system; the specific implementation of S110 in this application embodiment does not limit the implementation of S110.
[0103] In this application, multi-view images include different images from multiple perspectives; that is, the aforementioned multi-view images of a real 3D scene include different images of the real 3D scene from different perspectives.
[0104] For example, multi-view images of a real 3D scene include: an image acquired at 0 degrees relative to the real 3D scene (referred to as: 0-degree image), a 60-degree image, a 120-degree image, a 180-degree image, a 240-degree image, and a 300-degree image.
[0105] S120. Input multi-view images of a real 3D scene into the 3D scene instance segmentation model to obtain multi-view images of objects.
[0106] The 3D scene instance segmentation model in this application is a model that supports the segmentation of open-category objects; this model is used to identify different objects in an image and segment images of different objects from the image. Here, open-category objects are objects or objects of categories that do not exist in the model's training data.
[0107] For example, the 3D scene instance segmentation model can be a semantically related general segmentation model (grounded-segment-anything model, Grounded-SAM).
[0108] In this application, S120 is implemented by inputting each view image from the multi-view images of the real three-dimensional scene into Grounded-SAM, which identifies the objects in each view image and segments the images of the identified objects from each view image, thereby obtaining multi-view images of the objects in the real three-dimensional scene.
[0109] For example, based on the example in S110, assuming the real 3D scene includes a table and a chair; then the Grounded-SAM can obtain the images of the table and chair from the 0-degree viewpoint by processing the aforementioned 0-degree image. Similarly, the Grounded-SAM can obtain the images of the table and chair from the 60-degree viewpoint by processing the aforementioned 60-degree image; this process is repeated until the Grounded-SAM has processed each viewpoint image in the multi-view images of the real 3D scene, thereby obtaining the 0-degree, 60-degree, 120-degree, 180-degree, 240-degree, and 300-degree images of the table (i.e., multi-view images of the table), and the 0-degree, 60-degree, 120-degree, 180-degree, 240-degree, and 300-degree images of the chair (i.e., multi-view images of the chair).
[0110] It should be noted that when multiple objects (such as tables and chairs) exist in a real 3D scene, after multi-view image segmentation of the real 3D scene based on Grounded-SAM, multiple segmentation results are obtained (i.e., images of multiple objects from different perspectives). Since these multiple segmentation results are scattered and clustered together, the reconstruction system needs to determine the multi-view images of the table and the chair from these multiple segmentation results; the specific determination methods include: Method 1 and / or Method 2 as follows.
[0111] Method 1: Extract image features (e.g., image feature vectors) from each segmentation result and calculate the similarity between the image features of each segmentation result; determine the segmentation results with similarity (e.g., distance between image feature vectors) greater than a first threshold as multi-view images of the same object (e.g., table), thereby obtaining multi-view images of different objects (e.g., table and chair).
[0112] Method 2: By projecting each segmentation result into 3D space, 3D point clouds of different segmentation results are obtained. Then, the overlap ratio of the 3D point clouds of different segmentation results is calculated, and the segmentation results with an overlap ratio greater than a second threshold are determined as multi-view images of the same object.
[0113] It should be understood that in the implementation of S110, where the reconstructing system acquires multi-view images of a real 3D scene from an RGB-D camera, the RGB-D camera is an acquisition device or system that acquires images of the real 3D scene from multiple perspectives by moving around it. Therefore, the images of different frames acquired by the RGB-D camera are images of the real 3D scene from different angles (i.e., multi-view images of the real 3D scene). Subsequently, Grounded-SAM processes (i.e., segments) each frame acquired by the RGB-D camera to obtain multi-view images of objects in the real 3D scene.
[0114] It should be noted that S110-S120 is one way to acquire multi-view images of objects in a real 3D scene; another way is that the reassembly system can directly acquire multi-view images of objects in the real 3D scene from other devices or locally; specifically, the embodiments of this application do not limit the way to acquire multi-view images of objects in a real 3D scene.
[0115] S130. Generate the global feature vector of the object based on the multi-view image of the object.
[0116] It should be noted that when there are multiple objects in the real 3D scene, the reconstruction system executes S130-S160 for each object and S170 for the target CAD model corresponding to each object.
[0117] In this application, the global feature vector of an object is used to characterize the overall features of the object. These overall features are extracted from the object's overall perspective (i.e., multiple perspectives). Specifically, the overall features of the object include: the object's overall texture, the object's overall edges, and the object's overall shape.
[0118] The global feature vector of the aforementioned object is generated based on multi-view images of the object; the generation process includes: inputting the multi-view images of the object into a visual language base model, and obtaining the global feature vector of the object output by the visual language base model. The viewpoint semantic model can be a CLIP model.
[0119] It should be understood that after acquiring multi-view images of an object (object A), the CLIP model encodes the multi-view images of object A into a high-dimensional feature vector; this feature vector integrates visual information extracted from multiple view images of the object as well as semantic category information.
[0120] It should be noted that the CLIP model is a model trained based on joint contrastive learning between semantic text and images. This CLIP model has the ability to extract features of open-category objects.
[0121] S140. Input multi-view images of the object and use a language vision model to obtain the semantic text of the object.
[0122] In this application, the language vision model connects a visual encoder and a large language model to achieve general visual and language understanding. Specifically, this language vision model is a model that generates semantic text describing an image. For example, this language vision model could be an LLaVA model.
[0123] In this application, S140 is implemented by: inputting multi-view images of the object into the LLaVA model; and then obtaining the semantic text of the object output by the LLaVA model.
[0124] S150. Based on the semantic text of objects, generate text feature vectors of objects.
[0125] In this application, the text feature vector of an object is generated based on the semantic text of the object (e.g., the semantic text description information of the object); wherein, the semantic text description information of the object is used to describe the overall or local features of the object, and the specific embodiments of this application do not limit the features of the object described by the semantic text description information of the object.
[0126] In this application, the method for generating the text feature vector of an object is similar to the method for generating the global feature vector of an object. For details on the method for generating the text feature vector of an object, please refer to the above description of the method for generating the global feature vector of an object, which will not be repeated here.
[0127] It should be understood that S130 can be executed before S140-S150, after S140-S150, or simultaneously with S140-S150; specifically, the execution order of S130 and S140-S150 is not limited in the embodiments of this application.
[0128] It should be noted that S130-S150 is one implementation method for obtaining the global feature vector and text feature vector of an object in a real 3D scene. Another implementation method could be that the reconstruction system directly obtains the global feature vector and text feature vector of the object in the real 3D scene from other devices. Specifically, this application does not limit the implementation method for obtaining the global feature vector and text feature vector of an object in a real 3D scene.
[0129] The above embodiments use a 3D scene instance segmentation model to segment each view image in a multi-view image of a real 3D scene, obtaining multi-view images of objects in the real 3D scene; and inputting the multi-view images of the objects into a language vision model to obtain semantic text of the objects. Since this 3D scene instance segmentation model supports segmentation of open-category objects, it can identify and segment multi-view images of any category of objects in a real 3D scene without requiring manual labeling of the objects, thus improving the efficiency of virtual 3D scene reconstruction.
[0130] S160. Based on the global feature vector and textual feature vector of the object, determine the target CAD model that meets the similarity condition of the object from the CAD asset library.
[0131] In this application, the CAD asset library includes multiple CAD models, which are CAD models corresponding to various categories of objects. That is, unlike the small-scale CAD asset libraries in traditional technologies that specify object categories, the CAD asset library in this application includes a rich and massive number of CAD models. For example, the CAD asset library in this application is a collection that integrates existing large-scale 3D CAD model storage libraries ShapeNet and CAD models from 3D Future.
[0132] It should be understood that all CAD models in this application are three-dimensional CAD models, which will not be described further hereafter.
[0133] In this application, S160 can be implemented as follows: a candidate CAD model is determined based on a coarse-grained retrieval method; then, a target CAD model is determined from the candidate CAD models based on a fine-grained retrieval method; as described in method 1 below; or the target CAD model can be directly determined based on a coarse-grained retrieval method; as described in method 2 below.
[0134] Method 1: Based on the global feature vector and text feature vector of the object, determine multiple candidate CAD models with high similarity to the object from the CAD asset library; then, determine the target CAD model from these multiple candidate CAD models; see S210-S230 below for specific implementation details, which will not be repeated here.
[0135] Method 2: Based on the global feature vector and text feature vector of the object, determine the target CAD model from the CAD asset library that meets the similarity condition (e.g., the highest similarity) with the object; that is, the implementation of S160 is based on the global feature vector and text feature vector of the object, and determine the CAD model with the highest similarity with the object from the CAD asset library as the target CAD model.
[0136] The specific implementation of identifying the target CAD model with the highest similarity to the object from the CAD asset library includes: identifying the CAD model with the highest similarity between its feature vector and the global feature vector and text feature vector of the object from among multiple CAD models as the target CAD model.
[0137] It should be understood that the similarity of feature vectors can be represented by the cosine similarity between feature vectors or by the distance between feature vectors; specific embodiments of this application are not limited.
[0138] It should be noted that the feature vector of the CAD model can be obtained by using a sphere sampling strategy to collect multi-view images of the CAD model; then, based on the visual language base model (such as the CLIP model) and the multi-view images of the CAD model, the feature vector of the CAD model is generated.
[0139] S170. Based on the target CAD model, generate a virtual 3D scene that corresponds to the real 3D scene.
[0140] In this application, S170 can be implemented by aligning the pose and relative dimensions of the target CAD model with the pose and relative dimensions of the object, as described in method 1 below; or by aligning the pose of the target CAD model with the pose of the object, as described in method 2 below.
[0141] Method 1: Configure the target CAD model based on the object's pose parameters and relative scale to generate a virtual object.
[0142] The virtual object mentioned above is the configured target CAD model; the configuration specifically includes: setting the pose parameters of the target CAD model to the pose parameters of the object corresponding to the target CAD model in the real 3D scene, and setting the relative scale of the target CAD model to the relative scale of the object, thereby generating a virtual 3D scene.
[0143] The pose parameters of an object include: the object's position information (i.e., three-dimensional coordinates) and the object's three-dimensional orientation (also known as rotation); that is, the pose parameters of an object are used to indicate the object's 6-DOF pose.
[0144] The pose parameters of an object are determined based on its 3D point cloud. For details on the determination method, please refer to relevant technologies, which will not be elaborated here.
[0145] The three-dimensional point cloud of the aforementioned object can be obtained based on multi-view images of the object, or it can be obtained from an acquisition device or acquisition system (such as a lidar). Specifically, the embodiments of this application do not limit the method of obtaining the three-dimensional point cloud of the object.
[0146] It should be understood that the implementation of obtaining the 3D point cloud of an object based on multi-view images of the object includes: projecting each view image in the multi-view images into a 3D space to obtain a 3D object; then, obtaining the 3D point cloud of the object based on the 3D object; for specific implementation, please refer to relevant technologies, which will not be elaborated here.
[0147] In this application, relative scales are used to indicate the length, width, and height of an object.
[0148] It should be noted that the virtual 3D scene includes virtual objects corresponding to objects in the real 3D scene; these virtual objects are the target CAD models configured as described above. When there are multiple objects in the real 3D scene, the virtual 3D scene is composed of the virtual objects corresponding to each of these objects. When there is only one object in the real 3D scene, the virtual 3D scene is obtained based on the virtual objects corresponding to that object.
[0149] The above embodiments configure the target CAD model based on the object's pose parameters and relative scale, and generate a virtual object so that the position, orientation, and relative scale (i.e., 9-DOF pose) of the target CAD model are the same as the object, thereby more accurately reproducing the object in the real 3D scene, thus improving the accuracy of the virtual 3D scene.
[0150] Method 2: Configure the target CAD model based on the object's pose parameters to generate a virtual object.
[0151] In this application, the specific implementation of method 2 includes: setting the pose parameters of the target CAD model to the pose parameters of the object corresponding to the target CAD model in the real three-dimensional scene, and generating a virtual three-dimensional scene.
[0152] It should be understood that after setting the pose parameters of the target CAD model to the pose parameters of the object corresponding to the target CAD model in the real 3D scene, the target CAD model and the object have the same position in the 3D coordinate system, and the pose of the target CAD model and the object is also the same; thus, the object in the real 3D scene is restored through the target CAD model.
[0153] This application provides a method for reconstructing a 3D scene. This method determines a target CAD model that meets a certain similarity condition from multiple CAD models (i.e., a large number of rich CAD models) corresponding to objects of various categories, based on the global feature vector and text feature vector of an object in a real 3D scene. Then, it generates a virtual 3D scene corresponding to the real 3D scene based on the target CAD model. Since the global feature vector of the object is used to represent the overall features of the object, and the global feature vector is generated based on the object's multi-view images, and the text feature vector is generated based on the object's semantic text, and the semantic text is used to represent the object's overall or local features, this application determines the target CAD model from the large number of rich CAD models based on the multimodal features of the object, without requiring user intervention. Therefore, it improves the efficiency of reconstructing a virtual 3D scene.
[0154] Furthermore, compared to traditional techniques that retrieve target CAD models based on the geometric features of objects, this application determines the target CAD model from the massive CAD models based on the multimodal features of the aforementioned objects. These multimodal features integrate visual information from multiple perspective images of the object and the object's semantic text (e.g., descriptive information of the semantic text). The visual information and the semantic text complement each other in the process of retrieving the target CAD model (i.e., retrieving the target CAD model based on both visual information and semantic text modal features). Therefore, it can more quickly determine the target CAD model from the rich and massive CAD models, thus improving the efficiency of retrieving the target CAD model.
[0155] The following section details the scheme for determining candidate CAD models using a coarse-grained retrieval method in S160, followed by determining the target CAD model from these candidate CAD models using a fine-grained retrieval method.
[0156] It should be noted that, since the images from multiple perspectives are encoded into a single feature vector (e.g., the global feature vector of the object in S160), the differences between the images of the object in the real 3D scene and the CAD model in the CAD asset library from different perspectives are ignored, as well as the differences between the pose parameters of the object and the pose of the CAD module. Therefore, the global feature vector and text feature vector of the object in S160 are insufficient to retrieve the target CAD model that is most visually similar to the object.
[0157] Based on this, and referring to Figure 3, this application embodiment provides an implementation method for S160, as shown in Figure 4, which includes S210-S230.
[0158] S210. Based on the global feature vector and text feature vector of the object, calculate the similarity between them and the feature vector of each CAD model in multiple CAD models, and determine multiple candidate CAD models from multiple CAD models.
[0159] In this application, the candidate CAD model is a CAD model among multiple CAD models in the aforementioned CAD asset library whose feature vector has a higher similarity (abbreviated as: overall feature vector similarity) with the global feature vector and text feature vector of the object than the overall feature vector similarity of other CAD models among the multiple CAD models.
[0160] The implementation of S210 in this application, as shown in Figure 5, includes: S210A-S210C.
[0161] S210A: Calculate the similarity between the global feature vector of the object and the feature vector of each CAD model in multiple CAD models to determine the first model set.
[0162] In this application, the first model set includes M CAD models; these M CAD models are the top M CAD models corresponding to the highest similarity scores between the global feature vector of the aforementioned object and the feature vectors of each of the multiple CAD models in the CAD asset library, sorted from largest to smallest. In other words, the CAD models in the first model set are the top M CAD models among the multiple CAD models whose feature vectors have a high similarity to the global feature vector of the object. Where M is greater than or equal to 1.
[0163] Method 1: When the above similarity is expressed by cosine similarity, the above S210A is based on the following formula (1) to obtain the first model set. C img =TopM(cos(f) img ,F CAD ),M) (1)
[0164] Among them, C img Let f represent the first model set. img F is the global feature vector of the aforementioned objects. CAD is the feature vector of the CAD model, M is the number of CAD models in the first model set, TopM function means selecting the M CAD models with the highest cosine similarity, and cos() is the cosine expression.
[0165] Based on this, the specific implementation of S210A includes: calculating the cosine similarity between the feature vector of each CAD model in the CAD asset library and the global feature vector of the object; and determining the CAD models corresponding to the top M cosine similarities with the largest cosine similarities as the CAD models in the first model set.
[0166] It should be understood that the closer the cosine value of two vectors is to 1, the higher the similarity between the two vectors; that is, the larger the cosine value, the higher the cosine similarity between the two vectors, and thus the higher the similarity between the two vectors.
[0167] For example, assume M is 5; further assume the CAD asset library contains 80,000 CAD models; among them, the top 5 CAD models with the highest cosine similarity to the global feature vectors of the aforementioned objects are: CAD model_1 (cosine similarity 0.95), CAD model_9 (cosine similarity 0.95), CAD model_20 (cosine similarity 0.9), CAD model_100 (cosine similarity 0.85), and CAD model_200 (cosine similarity 0.8). Then, these 5 CAD models—CAD model_1, CAD model_9, CAD model_20, CAD model_100, and CAD model_200—are determined as the CAD models in the first model set.
[0168] Method 2: In the case where the similarity is represented by the distance between vectors, the specific implementation of S210A includes: calculating the distance between the feature vector of each CAD model in multiple CAD models and the global feature vector of the above object, and determining the top M CAD models with the smallest distance as the CAD models in the first model set.
[0169] S210B: Calculate the similarity between the text feature vector of the object and the feature vector of each CAD model in multiple CAD models to determine the second model set.
[0170] In this application, the second model set includes N CAD models. These N CAD models are the top N CAD models whose text feature vectors are similar to the feature vectors of each of the multiple CAD models, ranked from largest to smallest. In other words, the CAD models in the second model set are the top N CAD models whose feature vectors have a high similarity to the text feature vectors of the object. Here, N is greater than or equal to 1.
[0171] It should be noted that the sizes of M and N can be the same or different, and the specific embodiments of this application do not limit the sizes of M and N.
[0172] Method 1: When the above similarity is expressed by cosine similarity, the above S210B is based on the following formula (2) to obtain the second model set. text =TopN(cos(f text ,F CAD ),N) (2)
[0173] Among them, C text Let f represent the second model set. text F is the text feature vector of the aforementioned object. CAD is the feature vector of the CAD model, N is the number of CAD models in the second model set, the TopN function selects the top N CAD objects with the highest cosine similarity, and cos() is the cosine expression.
[0174] Based on this, the specific implementation of S210B includes: calculating the cosine similarity between the feature vector of each CAD model in the CAD asset library and the text feature vector of the object; and determining the CAD models corresponding to the top N largest cosine similarities as the CAD models in the second model set.
[0175] For example, based on the example in S210A, assuming the value of N is 4; the top four CAD models with the highest cosine similarity to the text feature vectors of the aforementioned objects are: CAD model_1 (cosine similarity 0.8), CAD model_300 (cosine similarity 0.8), CAD model_100 (cosine similarity 0.8), and CAD model_20 (cosine similarity 0.75). Therefore, these four CAD models—CAD model_1, CAD model_300, CAD model_100, and CAD model_20—are determined as the CAD models in the second model set.
[0176] Method 2: In the case where the similarity is represented by the distance between vectors, the specific implementation of S210B includes: calculating the distance between the feature vector of each CAD model in multiple CAD models and the text feature vector of the above object, and determining the N CAD models with the smallest distance as the CAD models in the second model set.
[0177] S210C, Obtain the union of the first model set and the second model set; wherein, the CAD models in the union are candidate CAD models.
[0178] In this application, S210C is based on the following formula (3) to obtain multiple candidate CAD models. c =C img ∪C text (3)
[0179] Among them, Cand c ∪ represents a set of multiple candidate CAD models, where ∪ is the union symbol.
[0180] Based on this, the specific implementation of S210C includes: determining the CAD models in the union of the first model set and the second model set as candidate CAD models.
[0181] For example, based on the examples in S210A and S210B; the first model set includes: CAD model_1, CAD model_9, CAD model_20, CAD model_100, and CAD model_200. The second model set includes: CAD model_1, CAD model_300, CAD model_100, CAD model_20, and CAD model_200. Then, the union of the first model set and the second model set is determined as the candidate CAD model, which includes: CAD model_1, CAD model_9, CAD model_20, CAD model_100, CAD model_200, and CAD model_300.
[0182] S220. Obtain the target features of the object.
[0183] In this application, the target features of an object include at least one of the following: visual features of each viewpoint image in a multi-view image of the object, or geometric features of the object.
[0184] In this application, the visual features of one viewpoint image (e.g., a 60-degree viewpoint image) from a multi-view image of an object are used to represent the overall features of that viewpoint image (i.e., the 60-degree viewpoint image); that is, the visual features of one viewpoint image are used to represent the overall features of the object from that viewpoint. Since one viewpoint image of an object can only describe the local features of the object, the overall features of the viewpoint image are a detailed description of those local features. For example, the visual features of one viewpoint image of an object include: the outline of the object from that viewpoint, the size of the object from that viewpoint, and the color of the object from that viewpoint.
[0185] In this application, the visual features of an object from each viewpoint are obtained by extracting the visual features of the object from each viewpoint image using a visual feature extractor. For example, the visual feature extractor can be an unsupervised robust visual feature extractor (DINOv2).
[0186] In this application, the geometric features of an object are used to describe the shape of the object; specifically, the geometric features of the object are used to describe the shape of the object in three-dimensional space. The geometric features of the object include: the pose parameters of the object.
[0187] The method for obtaining the geometric features of the aforementioned object includes: obtaining the three-dimensional point cloud of the object; and then extracting the geometric features of the object from the three-dimensional point cloud of the object.
[0188] It should be noted that the three-dimensional point cloud of the above-mentioned object can be obtained directly from the lidar by the reconstruction system; or it can be obtained by projecting the multi-view image of the object onto a three-dimensional coordinate system (i.e., three-dimensional space); specifically, the embodiments of this application do not limit the implementation method of obtaining the three-dimensional point cloud of the object.
[0189] S230. Based on the target features of the object, determine the target CAD model from multiple candidate CAD models.
[0190] When the target features of the above-mentioned object include the visual features of each view image in the multi-view images of the object, the implementation method 1 of S230 includes: determining the candidate CAD model with the highest similarity to the multiple visual features of the object among multiple candidate CAD models as the target CAD model; the specific implementation method is similar to S230A below, and will not be repeated here.
[0191] When the target features of the object include the geometric features of the object, the second implementation of S230 includes: determining the candidate CAD model with the highest similarity to the geometric features of the object among multiple candidate CAD models as the target CAD model; the specific implementation is similar to S230B-S230D below, and will not be repeated here.
[0192] When the target features of an object include the visual features of each view in the multi-view images of the object and the geometric features of the object, the implementation method 3 of S230 can be to first determine the candidate CAD model from the candidate CAD models based on the geometric features of the object, and then determine the target CAD model from the candidate CAD models based on the visual features of the object; or it can be to first determine the candidate CAD model from the candidate CAD models based on the visual features of the object, and then determine the target CAD model from the candidate CAD models based on the geometric features of the object.
[0193] It should be noted that the embodiments of this application use the following example to illustrate the target features of an object: the visual features of each view image in the multi-view image of the object and the geometric features of the object. This will not be repeated hereafter.
[0194] The above embodiments calculate the similarity between the object's global feature vector and textual feature vector and the feature vector of each of the multiple CAD models, respectively, to roughly retrieve (referred to as "coarse retrieval") multiple candidate CAD models with high similarity from a massive number of CAD models. Then, based on the visual features of the object's image from each viewpoint and / or the object's geometric features, the target CAD model with the highest similarity is precisely retrieved (referred to as "fine retrieval") from these multiple candidate CAD models, thereby improving the accuracy of the target CAD model.
[0195] Method 31: The specific implementation of the above method, which first selects candidate CAD models from the candidate CAD models based on the geometric features of the object, and then determines the target CAD model from the candidate CAD models based on the visual features of the object, includes: first, identifying the top P candidate CAD models with high similarity to the geometric features of the object as candidate CAD models. Then, identifying the candidate CAD model with the highest similarity to multiple visual features of the object among these P candidate CAD models as the target CAD model; the specific implementation is similar to the implementation methods of S230A-S230D described below, and will not be repeated here.
[0196] Method 32: The specific implementation of first determining the candidate CAD model from the candidate CAD models based on the object's visual features, and then determining the target CAD model from the candidate CAD models based on the object's geometric features, includes: identifying the top P candidate CAD models with high similarity to multiple visual features of the object as candidate CAD models. Then, identifying the candidate CAD model with the highest similarity to the object's geometric features among the P candidate CAD models as the target CAD model; for specific implementation, see S230A-S230D below, which will not be elaborated here.
[0197] The following is a detailed explanation of method 32 above:
[0198] When the target features of the object include the visual features of each view in the multi-view image of the object and the geometric features of the object, the specific implementation method of S230 is shown in Figure 6, including: S230A-S230D.
[0199] S230A, based on the visual features of multi-view images of objects and the visual features of multi-view images of each candidate CAD model among multiple candidate CAD models, determines P candidate CAD models through a similarity voting strategy.
[0200] In this application, P is greater than or equal to 2 and less than the number of candidate CAD models.
[0201] It should be noted that the visual features of the multi-view images of each candidate CAD model mentioned above include: the visual features of each view image of each candidate CAD model. The method for obtaining the visual features of the multi-view images of the candidate CAD model includes: acquiring multi-view images of the candidate CAD model; and then extracting the visual features of each view image of the candidate CAD model based on a visual feature extractor.
[0202] In this application, the similarity voting strategy is used to instruct the determination of target visual features corresponding to the visual features of each viewpoint image of an object from the candidate visual feature set. The target visual feature corresponding to a visual feature of an object is the visual feature with the highest similarity to that visual feature in the candidate visual feature set; the candidate visual feature set includes the visual features of each viewpoint image of each candidate CAD model. Then, based on the target visual features, a candidate CAD model is determined from the candidate CAD models; the number of target visual features in the candidate CAD model is greater than the number of target visual features in other candidate CAD models. In other words, the similarity voting strategy instructs that the number of target visual features in the candidate CAD model is greater than the number of target visual features in the non-candidate CAD models; that is, the similarity voting strategy instructs that the candidate CAD model whose number of target visual features meets the condition among the multiple candidate CAD models be selected as the candidate CAD model.
[0203] For example, based on the example in S210, the number of candidate CAD models mentioned above is 6. These 6 candidate CAD models include: CAD model_1, CAD model_9, CAD model_20, CAD model_100, CAD model_200, and CAD model_300. Assuming the value of P is 3, each candidate CAD model includes 12 visual features corresponding to 12 viewpoint images, that is, the candidate visual feature set includes 72 (6x12) visual features. Further assuming that the multi-view images of the object are 12 viewpoint images, that is, the visual features of the multi-view images of the object include 12 visual features corresponding to 12 viewpoint images.
[0204] Then, for each of the 12 visual features corresponding to the object, the target visual feature with the highest similarity in the aforementioned candidate visual feature set is determined, thus obtaining 12 target visual features. Next, the candidate CAD models containing these 12 target visual features are determined; for example, CAD model_1, CAD model_9, CAD model_300, and CAD model_100. Among these, CAD model_9 contains 4 target visual features, CAD model_1 and CAD model_300 contain 3 target visual features, and CAD model_100 contains 2 target visual features. Finally, CAD model_9, CAD model_1, and CAD model_300 are selected as the candidate CAD models.
[0205] It should be noted that, in the case that the target features of the object in S230 include the visual features of each view in the multi-view images of the object, the reconstructing system will identify the candidate CAD model that contains the most of the above-mentioned target visual features as the target CAD model.
[0206] S230B: Align the target parameters of the P candidate CAD models with the target parameters of the object, and obtain the 3D point cloud of each candidate CAD model after the target parameters are aligned.
[0207] In this application, the target parameters include: pose parameters, or the target parameters include: pose parameters and relative scale.
[0208] In this application, the method for aligning the target parameters of the candidate CAD model with the target parameters of the object is as follows: the target parameters of the candidate CAD model are set as the target parameters of the object.
[0209] S230C: Determine the chamfer distance and relative scale difference between the three-dimensional point cloud of each candidate CAD model after target parameter alignment and the three-dimensional point cloud of the object.
[0210] In this application, S230C is implemented by determining the chamfer distance between the 3D point cloud of each candidate CAD model after target parameter alignment and the 3D point cloud of the object, based on the chamfer distance formula; and by determining the relative scale difference between the 3D point cloud of each candidate CAD model after target parameter alignment and the 3D point cloud of the object, based on the relative scale difference formula; for specific implementation, please refer to related technologies, which will not be elaborated here.
[0211] S230D: From P candidate CAD models, determine the target CAD model corresponding to the 3D point cloud of the model whose sum of chamfer distance and relative scale difference satisfies the condition.
[0212] It should be understood that the smaller the chamfer distance between two 3D point clouds, the higher the similarity between the two 3D point clouds; similarly, the smaller the relative scale difference between two 3D point clouds, the higher the similarity between the two 3D point clouds.
[0213] In one implementation, S230D is implemented by: identifying the candidate CAD model with the smallest sum of the chamfer distance and the relative scale difference among P candidate CAD models as the target CAD model.
[0214] For example, based on the example in S230A, the candidate CAD models include: CAD model_9, CAD model_1, and CAD model_300. Assume that the sum of the chamfer distance and relative scale difference between the 3D point cloud of CAD model_9 (after target parameter alignment) and the 3D point cloud of the aforementioned object is 3. The sum of the chamfer distance and relative scale difference between the 3D point cloud of CAD model_1 (after target parameter alignment) and the 3D point cloud of the aforementioned object is 4. The sum of the chamfer distance and relative scale difference between the 3D point cloud of CAD model_300 (after target parameter alignment) and the 3D point cloud of the aforementioned object is 1. Therefore, since the sum of the chamfer distance and relative scale difference between the 3D point cloud of CAD model_300 (after target parameter alignment) and the 3D point cloud of the aforementioned object is the smallest, CAD model_300 is selected as the target CAD model.
[0215] In another implementation, the candidate CAD model with the smallest sum of chamfer distance weighted value and relative scale difference weighted value among P candidate CAD models is determined as the target CAD model.
[0216] It should be understood that S230B-S230D are the specific implementations of the target CAD model (i.e., the similarity between the geometric features of the target CAD model and the geometric features of the object satisfies the condition) determined from the P candidate CAD models identified in S230A. Based on this, when the target features of the object in S230 include the geometric features of the object, the reorganization system determines the candidate CAD model with the highest similarity between its geometric features and the geometric features of the object as the target CAD model.
[0217] This embodiment first determines P candidate CAD models from a plurality of candidate CAD models based on the visual features of the object's image from each viewpoint; this is equivalent to further narrowing down the scope of the target CAD model search based on the candidate CAD models. Then, based on the similarity between the geometric features of each of the P candidate CAD models and the geometric features of the object, the target CAD model is determined from the P candidate CAD models after the search scope has been narrowed down. This avoids the problem of low accuracy in determining the target CAD model due to ignoring the differences between the image of the object in the real 3D scene and the CAD model in the CAD asset library from different viewpoints, as well as the differences between the pose parameters of the object and the pose of the CAD module, thus improving the accuracy of the target CAD model.
[0218] Furthermore, this embodiment determines P candidate CAD models from the plurality of candidate CAD models based on the visual features of the object's image from each viewpoint. Then, the chamfer distance and relative scale difference between the 3D point cloud of each candidate CAD model after target parameter alignment and the 3D point cloud of the object are determined respectively, and the target CAD model is determined based on the sum of the chamfer distance and relative scale difference. It can be seen that the above embodiment improves the accuracy of the target CAD model by combining the chamfer distance and relative scale difference to determine the target CAD model.
[0219] It should be noted that during the generation of the virtual 3D scene in S170, special factors (such as lighting, noise, incorrect gravity settings, and collision detection errors) may cause two target CAD models that should have a supporting relationship to appear suspended, or the two target CAD models to collide and overlap. For example, a vase on a table may appear suspended above the tabletop. Another example is the two tables shown in Figure 11(a) that overlap in the left-right direction.
[0220] The solutions for the suspension phenomenon and the collision overlap phenomenon are explained below:
[0221] Solutions to the problem of suspension
[0222] Referring to Figure 3, this application embodiment provides a specific implementation method, which, after S170, further includes S310-S330 as shown in Figure 7.
[0223] S310. Obtain a first CAD model and a second CAD model with a support relationship from multiple target CAD models in a virtual 3D scene.
[0224] The support surface of the first CAD model is used to support the bottom surface of the second CAD model.
[0225] For example, suppose the first CAD model is a table and the second CAD model is a bottle. When the tabletop is used to support the bottle, the table and the bottle have a supporting relationship.
[0226] The implementation of S310 in this application is shown in Figure 8, including: S310A-S310C.
[0227] S310A: Obtain a third CAD model located above the support surface of the first CAD model, and which is not connected to the first CAD model by any other CAD model.
[0228] In this application, "above the support surface of the first CAD model" means that the support surface of the first CAD model is above the support surface in the z-axis direction.
[0229] In this application, the third CAD model is included among the multiple target CAD models in the virtual 3D scene.
[0230] The supporting surface of the first CAD model is a plane among multiple planes on the first CAD model, where the angle between the normal vector of the plane and the z-axis (i.e., the axis used to indicate altitude) in the three-dimensional coordinate system is less than a preset degree.
[0231] S310B, Determine the overlapping area of the projection of the bottom surface of the third CAD model onto the horizontal plane and the projection of the support surface of the first CAD model onto the horizontal plane.
[0232] In this application, the horizontal plane is the plane formed by the x-axis (i.e., the axis used to indicate length) and the y-axis (i.e., the axis used to indicate width) in a three-dimensional coordinate system.
[0233] In this application, S310B can be implemented by determining the overlapping area based on the projection of the bottom surface of the third CAD model and the projection of the supporting surface of the first CAD model, as described in method 1 below; or by determining the overlapping area based on the projection of the bottom surface of the three-dimensional bounding box of the third CAD model and the projection of the supporting surface of the first CAD model, as described in method 2 below.
[0234] Method 1: Project the bottom surface of the third CAD model and the support surface of the first CAD model onto the horizontal plane respectively; then, based on the coordinates of the overlapping area between the projection of the bottom surface of the third CAD model and the projection of the support surface of the first CAD model, determine the area of the overlapping area, i.e., the overlapping area.
[0235] Method 2: Determine the overlapping area of the projection of the bottom face of the 3D bounding box of the third CAD model onto the horizontal plane and the projection of the support face of the first CAD model onto the horizontal plane.
[0236] The above embodiment determines the overlapping area based on the projection of the bottom surface of the three-dimensional bounding box of the third CAD model onto the horizontal plane and the projection of the supporting surface of the first CAD model onto the horizontal plane, thus avoiding the problem that the overlapping area cannot be determined because the bottom surface of the supported object is a point (e.g., the supported object is a sphere).
[0237] S310C: Based on the overlapping area corresponding to the third CAD model, determine the second CAD model from the third CAD model.
[0238] In this application, the ratio of the overlapping area of the second CAD model to the area of the bottom surface of the second CAD model is greater than a preset ratio.
[0239] It should be noted that when S412 is implemented in mode 2 of S412, the ratio of the overlapping area of the second CAD model to the area of the bottom face of the three-dimensional bounding box of the second CAD model is greater than the preset ratio.
[0240] The above embodiment first obtains a third CAD model located above the support surface of the first CAD model, and without any other CAD models between it and the first CAD model. Then, based on the overlapping area of the projection of the bottom surface of the third CAD model onto the horizontal plane and the projection of the support surface of the first CAD model onto the horizontal plane, a second CAD model is determined from the third CAD model; the ratio of the overlapping area of the second CAD model to the area of its bottom surface is greater than a preset ratio. This avoids identifying a CAD model above the first CAD model and staggered from it on the horizontal plane as the second CAD model; therefore, the accuracy of the second CAD model is improved.
[0241] S320. Determine whether the height difference between the bottom surface of the second CAD model and the support surface of the first CAD model is greater than 0.
[0242] If the height difference between the bottom surface of the second CAD model and the support surface of the first CAD model is greater than 0, execute the following S330.
[0243] If the height difference between the bottom surface of the second CAD model and the supporting surface of the first CAD model is less than or equal to 0, execute the termination action to end the current method flow.
[0244] In this application, S320 is implemented by obtaining the height of the bottom surface of the second CAD model in the virtual three-dimensional scene (i.e., the coordinate value on the z-axis) and the height of the support surface of the first CAD model in the virtual three-dimensional scene; then, the height difference between the bottom surface of the second CAD model and the support surface of the first CAD model is determined.
[0245] S330. Move the second CAD model downwards by the height difference between the bottom surface of the second CAD model and the support surface of the first CAD model.
[0246] For example, as shown in Figure 9, the first CAD model is a table, and the second CAD model is a bottle above the table; the bottle is suspended relative to the table; assuming the bottom surface of the bottle is 5 and the height of the table is 3; then, the height difference between the bottom surface of the bottle and the table is 2; then, the bottle is moved down 2 so that the bottle falls on the table, as shown in Figure 10.
[0247] The above embodiment determines whether the second CAD model is suspended relative to the first CAD model by the height difference between the bottom surface of the second CAD model, which has a supporting relationship with the first CAD model in the virtual 3D scene; and if it is suspended, the second CAD model is moved downward by the height difference; thereby solving the suspension problem in the virtual 3D scene and improving the accuracy of the virtual 3D scene.
[0248] Solutions to collision overlap phenomenon
[0249] Referring to Figure 3 or Figure 7, this application embodiment provides a specific implementation method, which, after S170 above, further includes S410-S440 as shown in Figure 12.
[0250] S410. Obtain overlapping model pairs from the virtual 3D scene.
[0251] In this application, an overlapping model pair refers to two CAD models whose overlapping space volume is greater than 0. Specifically, the overlapping model pair includes a fourth CAD model and a fifth CAD model.
[0252] The above-mentioned S410 is implemented based on the position information (i.e., three-dimensional coordinates) of the CAD models in the virtual three-dimensional scene to determine the two overlapping CAD models; then, if the volume of the overlapping space of the two CAD models is greater than 0, the two CAD models are determined as an overlapping model pair.
[0253] S420. Determine the CAD model to be moved from the fourth and fifth CAD models.
[0254] It should be noted that the CAD model to be moved can be either the fourth CAD model or the fifth CAD model. Alternatively, the CAD model to be moved can be determined based on the similarity between the fourth CAD model and the corresponding object in the real 3D scene (i.e., the object corresponding to the fourth CAD model in the real 3D scene) and the similarity between the fifth CAD model and the corresponding object in the real 3D scene. The specific implementation method of S420 is not limited in this application embodiment.
[0255] When determining the CAD model to be moved based on the similarity between the fourth CAD model and the corresponding object in the real 3D scene, and the similarity between the fifth CAD model and the corresponding object in the real 3D scene, the implementation of S420 includes: determining the similarity between the fourth CAD model and the corresponding object in the real 3D scene, and the similarity between the fifth CAD model and the corresponding object in the real 3D scene; and determining the CAD model with the lowest similarity as the CAD model to be moved.
[0256] It should be understood that the object corresponding to a CAD model (e.g., CAD model A) in a virtual 3D scene is the object with the highest similarity to CAD model A in the real 3D scene; that is, the target CAD model corresponding to the object is CAD model A.
[0257] The aforementioned similarity can be represented by the chamfer distance and / or relative scale between the CAD model and the corresponding object in the real 3D scene, or by the visual features of the multi-view images of the CAD model and the visual features of the multi-view images of the corresponding object in the real 3D scene. The specific embodiments of this application do not limit the way the similarity is represented.
[0258] S430. The sum of the side length of the side perpendicular to the maximum plane in the three-dimensional bounding box of the overlapping space of the fourth CAD model and the fifth CAD model and the preset length is determined as the target distance.
[0259] In this application, the preset length is greater than or equal to 0;
[0260] The aforementioned maximum plane is the maximum plane within the three-dimensional bounding box of the overlapping space of the fourth and fifth CAD models.
[0261] In this application, S430 is implemented by: determining the maximum face from the three-dimensional bounding box of the overlapping space, then determining the edge perpendicular to the maximum face from the three-dimensional bounding box; and finally determining the target distance as the sum of the side length of the edge and a preset length.
[0262] S440: Based on the pointing direction, move the CAD model to be moved by the target distance.
[0263] In this application, the pointing direction is the direction in which the non-movable CAD model points to the CAD model to be moved in the fourth and fifth CAD models.
[0264] In this application, S440 can be implemented by moving the CAD model to be moved based on the pointing direction, as described in method 1 below; or by moving the CAD model to be moved based on the target direction determined by the pointing direction, as described in method 2 below.
[0265] Method 1: Move the CAD model to be moved by the target distance in the pointing direction.
[0266] For example, as shown in Figure 11(a), the fourth CAD model is the right table where the overlap occurs, and the fifth CAD model is the left table where the overlap occurs. When the CAD model to be moved is the fourth CAD model, the pointing direction is to the right of the fifth CAD model pointing to the fourth CAD model; the fourth CAD model is moved to the right a target distance to separate it from the fifth CAD model, as shown in Figure 11(b).
[0267] The above embodiment determines the CAD model to be moved from an overlapping model pair in a virtual 3D scene. This overlapping model pair includes a fourth CAD model and a fifth CAD model. Then, the target distance is determined by the sum of the length of the side perpendicular to the largest face of the 3D bounding box of the overlapping space of the fourth and fifth CAD models and a preset length. Finally, the CAD model to be moved is moved by this target distance in the direction that the non-moving CAD model is pointing towards; this separates the overlapping fourth and fifth CAD models, thereby improving the accuracy of the virtual 3D scene.
[0268] Method 2: Move the CAD model to be moved based on the target direction determined by the pointing direction.
[0269] As shown in Figure 13, it includes: S440A-S440C.
[0270] S440A: Determine the vertical direction of the largest plane in the three-dimensional bounding box of the overlapping space as the direction to be moved.
[0271] In this application, the direction to be moved includes a positive direction and a negative direction; the positive direction is opposite to the negative direction.
[0272] For example, as shown in Figure 11(a), since the largest face in the three-dimensional bounding box is perpendicular to the x-axis, the movement to be made includes: to the right (i.e., the positive direction) and to the left (i.e., the negative direction).
[0273] S440B: Determines the target direction from the positive and negative directions based on the pointing direction.
[0274] In this application, the angle between the target direction and the pointing direction is smaller than the angle between the non-target direction and the pointing direction.
[0275] In this application, S542 is implemented by: determining the angles between the positive direction and the negative direction and the pointing direction, respectively. If the angle between the positive direction and the pointing direction is less than the angle between the negative direction and the pointing direction, the positive direction is determined as the target direction. If the angle between the positive direction and the pointing direction is greater than the angle between the negative direction and the pointing direction, the negative direction is determined as the target direction.
[0276] S440C: Move the CAD model to be moved a target distance in the target direction.
[0277] It should be noted that S430-S440 is one way to move the CAD model to be moved so that the fourth and fifth CAD models after the move do not have overlapping volumes; another way is to move the CAD model to be moved a preset distance in a preset direction, so that the fourth and fifth CAD models after the move do not have overlapping volumes.
[0278] The above embodiment determines the direction to be moved by the vertical direction of the largest plane in the three-dimensional bounding box of the overlapping space; then, the positive and negative directions in the direction to be moved, respectively, with the smaller angle between them and the pointing direction, are determined as the target directions; thereby avoiding the problem that the CAD model to be moved will move to the wrong oblique direction because the fourth CAD model is on the oblique side of the fifth CAD model (e.g., left oblique, right oblique, upper oblique, or lower oblique); thus improving the accuracy of the virtual three-dimensional scene.
[0279] The above primarily describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, the 3D scene reconstruction device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, in conjunction with the modules and algorithm steps of the various examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0280] This application embodiment can, according to the above method, exemplarily divide the 3D scene reconstruction device into functional modules. For example, the 3D scene reconstruction device may include functional units corresponding to each functional division, or two or more functions may be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division; in actual implementation, there may be other division methods.
[0281] With each functional unit divided according to its corresponding function, Figure 14 shows a possible structural schematic diagram of the three-dimensional scene reconstruction device involved in the above embodiments; the three-dimensional scene reconstruction device is deployed in the reconstruction system; the three-dimensional scene reconstruction device includes: an acquisition module 1301 and a processing module 1302.
[0282] The acquisition module 1301 is used to acquire the global feature vector and text feature vector of objects in a real 3D scene.
[0283] The processing module 1302 is used to determine the target CAD model that meets the similarity condition of the object from the CAD asset library based on the global feature vector and the text feature vector of the object; for example, it executes step S160 in the above method embodiment.
[0284] The processing module 1302 is also used to generate a virtual 3D scene corresponding to the real 3D scene based on the target CAD model; for example, by executing step S170 in the above method embodiment.
[0285] Optionally, the processing module 1302 is used to calculate the similarity between the global feature vector and the text feature vector of the object and the feature vector of each of the multiple CAD models, respectively, to determine multiple candidate CAD models from the multiple CAD models; for example, to perform step S210 in the above method embodiment.
[0286] The acquisition module 1301 is used to acquire the target features of the object; for example, by executing step S220 in the above method embodiment.
[0287] The processing module 1302 is used to determine the target CAD model from multiple candidate CAD models based on the target features of the object; for example, by performing step S230 in the above method embodiment.
[0288] Optionally, the acquisition module 1301 is used to calculate the similarity between the global feature vector of the object and the feature vector of each CAD model in the multiple CAD models to determine a first model set, and to calculate the similarity between the text feature vector of the object and the feature vector of each CAD model in the multiple CAD models to determine a second model set; for example, to execute steps S210A-S210B in the above method embodiment.
[0289] The processing module 1302 is used to obtain the union of the first model set and the second model set; wherein the CAD models in the union are candidate CAD models; for example, the step S210C in the above method embodiment is executed.
[0290] Optionally, the processing module 1302 is used to determine P candidate CAD models based on the visual features of the multi-view images of the object and the visual features of the multi-view images of each candidate CAD model among multiple candidate CAD models, through a similarity voting strategy; for example, by performing step S230A in the above method embodiment.
[0291] The processing module 1302 is used to determine the target CAD model from P candidate CAD models.
[0292] Optionally, the processing module 1302 is used to align the target parameters of the P candidate CAD models with the target parameters of the object, and obtain the 3D point cloud of each candidate CAD model after the target parameters are aligned; for example, by executing step S230B in the above method embodiment.
[0293] The processing module 1302 is used to determine the chamfer distance and relative scale difference between the three-dimensional point cloud of each candidate CAD model after target parameter alignment and the three-dimensional point cloud of the object; for example, it executes step S230C in the above method embodiment.
[0294] The processing module 1302 is used to determine the target CAD model corresponding to the 3D point cloud of the model that satisfies the condition of the sum of chamfer distance and relative scale difference from P candidate CAD models; for example, it executes step S230D in the above method embodiment.
[0295] Optionally, the acquisition module 1301 is used to acquire multi-view images of a real 3D scene; for example, by performing step S110 in the above method embodiment.
[0296] The processing module 1302 is used to input multi-view images of a real 3D scene into a 3D scene instance segmentation model to obtain multi-view images of objects; for example, to execute step S120 in the above method embodiment.
[0297] The processing module 1302 is used to input multi-view images of objects and use a language vision model to obtain semantic text of the objects; for example, it executes step S140 in the above method embodiment.
[0298] Optionally, the processing module 1302 is used to configure the target CAD model and generate a virtual object based on the object's pose parameters and relative scale.
[0299] Optionally, the acquisition module 1301 is used to acquire a first CAD model and a second CAD model with a support relationship from multiple target CAD models in a virtual 3D scene; for example, by performing step S310 in the above method embodiment.
[0300] The processing module 1302 is used to move the second CAD model downward by the height difference when the height difference between the bottom surface and the support surface of the second CAD model is greater than 0; for example, by executing steps S320-S330 in the above method embodiment.
[0301] Optionally, the acquisition module 1301 is used to acquire a third CAD model located above the support surface of the first CAD model and without any other CAD models between it and the first CAD model; for example, by executing step S310A in the above method embodiment.
[0302] The processing module 1302 is used to determine the overlapping area of the projection of the bottom surface of the third CAD model onto the horizontal plane and the projection of the support surface of the first CAD model onto the horizontal plane; for example, it executes step S310B in the above method embodiment.
[0303] The processing module 1302 is used to determine the second CAD model from the overlapping area corresponding to the third CAD model; for example, by performing step S310C in the above method embodiment.
[0304] Optionally, the acquisition module 1301 is used to acquire overlapping model pairs from the virtual 3D scene; for example, by performing step S410 in the above method embodiment.
[0305] The processing module 1302 is used to determine the CAD model to be moved from the fourth CAD model and the fifth CAD model; for example, by performing step S420 in the above method embodiment.
[0306] The processing module 1302 is used to move the CAD model to be moved.
[0307] Optionally, the processing module 1302 is used to determine the target distance by summing the side length of the side perpendicular to the maximum plane in the three-dimensional bounding box of the overlapping space of the fourth CAD model and the fifth CAD model with a preset length; for example, by performing step S430 in the above method embodiment.
[0308] The processing module 1302 is used to move the CAD model to be moved a target distance based on the pointing direction; for example, by executing step S440 in the above method embodiment.
[0309] Optionally, the processing module 1302 is used to determine the vertical direction of the largest plane in the three-dimensional bounding box of the overlapping space as the direction to be moved; for example, by performing step S440A in the above method embodiment.
[0310] The processing module 1302 is used to determine the target direction from the positive and negative directions based on the pointing direction; for example, by performing step S440B in the above method embodiment.
[0311] The processing module 1302 is used to move the CAD model to be moved a target distance in the target direction; for example, by executing step S440A in the above method embodiment.
[0312] Both the acquisition module 1301 and the processing module 1302 can be implemented in software or in hardware. For example, the implementation of the processing module 1302 will be described below. The implementation of the acquisition module 1301 can be referenced from the implementation of the processing module 1302.
[0313] As an example of a software functional unit, processing module 1302 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the aforementioned computing instance may be one or more. For example, processing module 1302 may include code running on multiple hosts / virtual machines / containers.
[0314] It should be noted that the multiple hosts / virtual machines / containers used to run this code can be distributed within the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run this code can be distributed within the same availability zone (AZ) or in different AZs, each AZ comprising one or more geographically proximate data centers. Typically, a region can include multiple AZs.
[0315] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0316] As an example of a hardware functional unit, the processing module 1302 may include at least one computing device, such as a server. Alternatively, the processing module 1302 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0317] The processing module 1302 includes multiple computing devices that can be distributed in the same region or in different regions. Similarly, the processing module 1302 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the processing module 1302 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0318] It should be noted that, in other embodiments, the processing module 1302 can be used to execute any step in the above-described 3D scene reconstruction method, and the acquisition module 1301 can also be used to execute any step in the above-described 3D scene reconstruction method. The steps implemented by the acquisition module 1301 and the processing module 1302 can be specified as needed. By implementing different steps in the above-described 3D scene reconstruction method through the acquisition module 1301 and the processing module 1302 respectively, all functions of the management device can be realized.
[0319] This application also provides a computing device, specifically the reconfiguration system 201 in Figure 2; as shown in Figure 15, the computing device may include a processor 301, a memory 302, and a communication interface 303. The processor 301, memory 302, and communication interface 303 can be connected via a bus 304 or other means. This computing device can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device.
[0320] Processor 301 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0321] The memory 302 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, or optical memory, disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer.
[0322] The memory 302 stores executable program code, and the processor 301 executes the executable program code to implement the functions of the aforementioned acquisition module 1301 and processing module 1302, thereby realizing the three-dimensional scene reconstruction method. That is, the memory 302 stores instructions for executing the three-dimensional scene reconstruction method.
[0323] Alternatively, the memory 302 stores executable code, which the processor 301 executes to implement the functions of the aforementioned 3D scene reconstruction device, thereby realizing the 3D scene reconstruction method. That is, the memory 302 stores instructions for executing the 3D scene reconstruction method.
[0324] In one possible implementation, the memory 302 can exist independently of the processor 301. The memory 302 can be connected to the processor 301 via a bus 304 and is used to store data, instructions, or program code. When the processor 301 calls and executes the instructions or program code stored in the memory 302, it can implement the relevant steps in the three-dimensional scene reconstruction method provided in the embodiments of this application.
[0325] In another possible implementation, the memory 302 can also be integrated with the processor 301.
[0326] The communication interface 303 can be an acquisition module used to communicate with other devices or communication networks, such as Ethernet, RAN, and wireless local area networks (WLAN). The communication interface 303 can receive commands, messages, or data. The acquisition module can be a transceiver or similar device.
[0327] Optionally, the communication interface 303 can also be a transceiver circuit located within the processor 301, used to implement signal input and signal output of the processor 301, such as acquiring multi-view images of a real 3D scene. The communication interface 303 can be a wired interface (port), such as a fiber distributed data interface (FDDI) or a gigabit Ethernet (GE) interface, or it can also be a wireless interface.
[0328] Bus 304 can be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus. This bus can be categorized as an address bus, data bus, control bus, etc. Buses can also be classified as serial buses and parallel buses. For ease of representation, only one thick line is used in Figure 15, but this does not indicate that there is only one bus or one type of bus.
[0329] It should be understood that the computing device in Figure 15 is merely an example of a computing device, which may have more or fewer components than those shown in Figure 15, may combine two or more components, or may have different component configurations. For example, the computing device may also include a smart network interface card, such as a data processing unit (DPU).
[0330] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0331] As shown in Figure 16, the computing device cluster includes at least one computing device 100. The memory 302 of one or more computing devices 100 in the computing device cluster may store the same instructions for executing the aforementioned three-dimensional scene reconstruction method.
[0332] In some possible implementations, the memory 302 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the aforementioned 3D scene reconstruction method. In other words, a combination of one or more computing devices 100 can jointly execute the instructions for executing the aforementioned 3D scene reconstruction method.
[0333] It should be understood that the processor and bus in the computing device 100 described above are the same as the processor 301 and bus 304 in FIG15. For a detailed description of the processor and bus in the computing device 100, please refer to the relevant description for FIG15, which will not be repeated here.
[0334] It should be noted that the memory 302 in different computing devices 100 within the computing device cluster can store different instructions, each used to execute a portion of the management device's functions. That is, the instructions stored in the memory 302 of different computing devices 100 can implement the functions of one or more modules in the acquisition module 1301 and processing module 1302.
[0335] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 17 illustrates one possible implementation. As shown in Figure 17, two computing devices 100A and 100B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 302 in computing device 100A stores instructions for executing the functions of processing module 1302. Simultaneously, the memory 302 in computing device 100B stores instructions for executing the functions of acquisition module 1301.
[0336] The connection method between the computing device clusters shown in Figure 17 can be considered as follows: taking into account that the three-dimensional scene reconstruction method provided in this application needs to frequently acquire and send data, the function implemented by the acquisition module 1301 is considered to be executed by the computing device 100B.
[0337] It should be understood that the functions of computing device 100A shown in Figure 17 can also be performed by multiple computing devices 100. Similarly, the functions of computing device 100B can also be performed by multiple computing devices 100.
[0338] This application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similar to the connection method of the computing device cluster described in Figures 16 and 17. The difference is that the memory 302 of one or more computing devices 100 in this computing device cluster can store the same instructions for executing the three-dimensional scene reconstruction method.
[0339] In some possible implementations, the memory 302 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the 3D scene reconstruction method. In other words, a combination of one or more computing devices 100 can jointly execute the instructions for executing the 3D scene reconstruction method.
[0340] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform a three-dimensional scene reconstruction method.
[0341] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct a computing device to perform a three-dimensional scene reconstruction method.
[0342] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for reconstructing a three-dimensional scene, characterized in that, The method includes: Obtain the global feature vector and the text feature vector of the object in a real 3D scene; wherein, the global feature vector is generated based on the multi-view image of the object and is used to represent the overall features of the object; the text feature vector of the object is generated based on the semantic text of the object and is used to represent the overall or local features of the object. Based on the global feature vector and the text feature vector of the object, a target CAD model that meets the similarity condition with the object is determined from the CAD asset library of computer-aided design files; wherein, the CAD asset library includes multiple CAD models, which are CAD models corresponding to various categories of objects; Based on the target CAD model, a virtual 3D scene corresponding to the real 3D scene is generated; the virtual 3D scene includes virtual objects corresponding to the object.
2. The method according to claim 1, characterized in that, The step of determining a target CAD model from the CAD asset library that meets the similarity criteria to the object, based on the object's global feature vector and text feature vector, includes: Based on the global feature vector and text feature vector of the object, the similarity between the object and the feature vector of each of the multiple CAD models is calculated to determine multiple candidate CAD models from the multiple CAD models. Obtain the target features of the object; wherein the target features of the object include at least one of the following: visual features of each view image in the multi-view images of the object, or geometric features of the object; the visual features of one view image in the multi-view images of the object are used to represent the overall features of the one view image; the geometric features of the object are used to describe the features of the object's shape. Based on the target features of the object, the target CAD model is determined from the plurality of candidate CAD models.
3. The method according to claim 2, characterized in that, The similarity calculation between the global feature vector and the text feature vector of the object and the feature vector of each of the multiple CAD models, respectively, to determine multiple candidate CAD models from the multiple CAD models, includes: Based on the similarity calculation between the global feature vector of the object and the feature vector of each of the multiple CAD models, a first model set is determined; the first model set includes M CAD models, which are the CAD models corresponding to the top M similarity scores between the global feature vector of the object and the feature vector of each of the multiple CAD models, sorted from largest to smallest; M is greater than or equal to 1; Based on the similarity calculation between the text feature vector of the object and the feature vector of each of the multiple CAD models, a second model set is determined; the second model set includes N CAD models, which are the CAD models corresponding to the top N similarity scores between the text feature vector of the object and the feature vector of each of the multiple CAD models, sorted from largest to smallest; N is greater than or equal to 1; Obtain the union of the first model set and the second model set; wherein, the CAD models in the union are the candidate CAD models.
4. The method according to claim 2 or 3, characterized in that, The step of determining the target CAD model from the plurality of candidate CAD models based on the target features of the object includes: Based on the visual features of the object's multi-view images and the visual features of each candidate CAD model's multi-view images among the multiple candidate CAD models, P candidate CAD models are determined through a similarity voting strategy; P is greater than or equal to 2. Wherein, the visual features of the multi-view images of the object refer to the visual features of each view image in the multi-view images of the object; the similarity voting strategy is used to indicate that the candidate CAD model whose number of target visual features meets the condition among the multiple candidate CAD models is the candidate CAD model; the target visual feature is the visual feature with the highest similarity to a visual feature of the object among the visual features of the multi-view images of the multiple candidate CAD models. From the P candidate CAD models, the target CAD model is determined, wherein the similarity between the geometric features of the target CAD model and the geometric features of the object satisfies the condition.
5. The method according to claim 4, characterized in that, The step of determining the target CAD model from the P candidate CAD models includes: Align the target parameters of the P candidate CAD models with the target parameters of the object, and obtain the 3D point cloud of each candidate CAD model after the target parameters are aligned; the target parameters include: pose parameters, or the target parameters include: pose parameters and relative scale; The chamfer distance and relative scale difference between the 3D point cloud of each candidate CAD model after alignment with the target parameters and the 3D point cloud of the object are determined respectively. From the P candidate CAD models, determine the target CAD model corresponding to the 3D point cloud of the model whose sum of chamfer distance and relative scale difference satisfies the condition.
6. The method according to any one of claims 1-5, characterized in that, Before obtaining the global feature vectors of objects in the real 3D scene and the text feature vectors of the objects, the method further includes: Acquire multi-view images of the real 3D scene; Input the multi-view images of the real 3D scene into the 3D scene instance segmentation model to obtain the multi-view images of the object; Input multi-view images of the object into a language vision model to obtain the semantic text of the object.
7. The method according to any one of claims 1-6, characterized in that, The step of generating a virtual 3D scene corresponding to the real 3D scene based on the target CAD model includes: Based on the pose parameters and relative scale of the object, the target CAD model is configured to generate the virtual object.
8. The method according to any one of claims 1-7, characterized in that, The virtual 3D scene includes multiple target CAD models corresponding to multiple objects; the method further includes: A first CAD model and a second CAD model with a support relationship are obtained from the plurality of target CAD models; the support surface of the first CAD model is used to support the bottom surface of the second CAD model; If the height difference between the bottom surface of the second CAD model and the supporting surface is greater than 0, the second CAD model is moved downward by the height difference.
9. The method according to claim 8, characterized in that, The step of obtaining a first CAD model and a second CAD model with a support relationship from the plurality of target CAD models includes: Obtain a third CAD model located above the support surface of the first CAD model, and where there are no other CAD models between the third CAD model and the first CAD model; wherein the third CAD model is included among the plurality of target CAD models; Determine the overlapping area of the projection of the bottom surface of the third CAD model onto the horizontal plane and the projection of the supporting surface onto the horizontal plane; Based on the overlapping area corresponding to the third CAD model, the second CAD model is determined from the third CAD model; the ratio of the overlapping area corresponding to the second CAD model to the bottom surface area of the second CAD model is greater than a preset ratio.
10. The method according to any one of claims 1-9, characterized in that, The method further includes: Obtain overlapping model pairs from the virtual 3D scene; the overlapping model pair consists of two CAD models with an overlapping space volume greater than 0; the overlapping model pair includes: a fourth CAD model and a fifth CAD model; The CAD model to be moved is determined from the fourth and fifth CAD models; The CAD model to be moved is moved; the fourth and fifth CAD models after the move do not have overlapping volumes.
11. The method according to claim 10, characterized in that, Moving the CAD model to be moved includes: The target distance is determined by the sum of the side length of the side perpendicular to the maximum plane in the three-dimensional bounding box of the overlapping space of the fourth CAD model and the fifth CAD model, and a preset length; the maximum plane is the maximum plane in the three-dimensional bounding box of the overlapping space. Based on the pointing direction, the CAD model to be moved is moved by the target distance; the pointing direction is the direction in which the non-moving CAD model points to the CAD model to be moved.
12. The method according to claim 11, characterized in that, The step of moving the CAD model to be moved by the target distance based on the pointing direction includes: The vertical direction of the largest plane is determined as the direction to be moved; the direction to be moved includes a positive direction and a negative direction; the positive direction is opposite to the negative direction. Based on the pointing direction, a target direction is determined from the positive direction and the negative direction, wherein the angle between the target direction and the pointing direction is smaller than the angle between the non-moving direction and the pointing direction; Move the CAD model to be moved by the target distance in the target direction.
13. A three-dimensional scene reconstruction device, characterized in that, The 3D scene reconstruction device includes: an acquisition module and a processing module; The acquisition module is used to acquire the global feature vector and the text feature vector of the object in the real 3D scene; wherein, the global feature vector is generated based on the multi-view image of the object and is used to represent the overall features of the object; the text feature vector of the object is generated based on the semantic text of the object and is used to represent the overall features or local features of the object. The processing module is used to determine a target CAD model that meets the similarity condition of the object from the CAD asset library based on the global feature vector and the text feature vector of the object; wherein, the CAD asset library includes multiple CAD models, which are CAD models corresponding to various categories of objects; The processing module is further configured to generate a virtual 3D scene corresponding to the real 3D scene based on the target CAD model; the virtual 3D scene includes virtual objects corresponding to the object.
14. The three-dimensional scene reconstruction device according to claim 13, characterized in that, The processing module is used to calculate the similarity between the global feature vector and the text feature vector of the object and the feature vector of each of the plurality of CAD models, respectively, and to determine a plurality of candidate CAD models from the plurality of CAD models. The acquisition module is used to acquire the target features of the object; wherein the target features of the object include at least one of the following: visual features of each view image in the multi-view images of the object, or geometric features of the object; the visual features of one view image in the multi-view images of the object are used to represent the overall features of the one view image; the geometric features of the object are used to describe the features of the object's shape. The processing module is used to determine the target CAD model from the plurality of candidate CAD models based on the target features of the object.
15. The three-dimensional scene reconstruction device according to claim 14, characterized in that, The processing module is used to determine a first model set based on the similarity calculation between the global feature vector of the object and the feature vector of each of the plurality of CAD models; the first model set includes M CAD models, which are the CAD models corresponding to the top M similarity scores between the global feature vector of the object and the feature vector of each of the plurality of CAD models, sorted from largest to smallest; M is greater than or equal to 1; The processing module is used to determine a second model set based on the similarity calculation between the text feature vector of the object and the feature vector of each of the plurality of CAD models; the second model set includes N CAD models, which are the CAD models corresponding to the top N similarity scores between the text feature vector of the object and the feature vector of each of the plurality of CAD models, sorted from largest to smallest; N is greater than or equal to 1; The processing module is used to obtain the union of the first model set and the second model set; wherein the CAD model in the union is the candidate CAD model.
16. The three-dimensional scene reconstruction device according to claim 14 or 15, characterized in that, The processing module is used to determine P candidate CAD models based on the visual features of the multi-view images of the object and the visual features of the multi-view images of each candidate CAD model among the multiple candidate CAD models, through a similarity voting strategy; P is greater than or equal to 2. Wherein, the visual features of the multi-view images of the object refer to the visual features of each view image in the multi-view images of the object; the similarity voting strategy is used to indicate that the candidate CAD model whose number of target visual features meets the condition among the multiple candidate CAD models is the candidate CAD model; the target visual feature is the visual feature with the highest similarity to a visual feature of the object among the visual features of the multi-view images of the multiple candidate CAD models. The processing module is used to determine the target CAD model from the P candidate CAD models, wherein the similarity between the geometric features of the target CAD model and the geometric features of the object satisfies a condition.
17. The three-dimensional scene reconstruction device according to claim 16, characterized in that, The processing module is used to align the target parameters of the P candidate CAD models with the target parameters of the object, and to obtain the 3D point cloud of each candidate CAD model after the target parameters are aligned. The target parameters include: pose parameters, or the target parameters include: pose parameters and relative scale; The processing module is used to determine the chamfer distance and relative scale difference between the 3D point cloud of each candidate CAD model after the target parameters are aligned and the 3D point cloud of the object. The processing module is used to determine the target CAD model corresponding to the 3D point cloud of the model that satisfies the condition of the sum of the chamfer distance and the relative scale difference from the P candidate CAD models.
18. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1-12.
19. A computer-readable storage medium, characterized in that, The device stores computer instructions that, when executed on a computing device, cause the computing device to perform the method as described in any one of claims 1-12.
20. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster causes the computing device cluster to perform the method as described in any one of claims 1-12.