Three-dimensional point cloud generation method and device, computer equipment and storage medium
The initial three-dimensional point cloud generation model is used to process the initial three-dimensional point cloud, which solves the problem of low resolution and achieves the improvement of the number of feature points and the improvement of resolution.
Patent Information
- Application Number
- CN202510231514.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, the resolution of three-dimensional point clouds is low, resulting in a small number of feature points, which cannot be edited or improved in resolution.
By pre-training the 3D point cloud generation model, the initial 3D point cloud is processed to generate high-resolution 3D point clouds. The model is based on sample description text and sample three-dimensional point cloud training generation, and outputs the target three-dimensional point cloud by inputting text features and spatial features.
The resolution improvement of low-resolution three-dimensional point clouds is achieved, and the number of feature points generated by the target three-dimensional point cloud is greater than that of the initial point cloud, meeting the needs of high resolution.
Smart Images

Figure CN120147535A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly relates to a three-dimensional point cloud generation method, apparatus, computer device, and storage medium. Background Art
[0002] The three-dimensional point cloud model is an important manifestation of the digitalization of real objects. Through computer technology, the corresponding three-dimensional point cloud of an object can be generated, and then through the three-dimensional point cloud, the development and testing of three-dimensional point cloud-related algorithms, as well as the development of robot interaction control methods related to the three-dimensional point cloud, can be realized.
[0003] In related technologies, the three-dimensional point cloud is generally obtained from the real world through three-dimensional scanning devices (such as lidar, depth cameras, etc.). However, due to equipment and cost limitations, the obtained three-dimensional point cloud has the problem of low resolution. The most obvious manifestation is the small number of feature points in the three-dimensional point cloud. In related technologies, the three-dimensional point cloud cannot be edited and changed, so the resolution corresponding to the three-dimensional point cloud cannot be improved. Summary of the Invention
[0004] Embodiments of this application provide a three-dimensional point cloud generation method, apparatus, computer device, and storage medium, which can process a low-resolution three-dimensional point cloud through a pre-trained point cloud generation model to generate a high-resolution three-dimensional point cloud.
[0005] To achieve the above objective, an embodiment of this application provides a three-dimensional point cloud generation method, including:
[0006] Obtain an initial three-dimensional point cloud, and determine the three-dimensional space coordinates corresponding to each feature point in the initial three-dimensional point cloud;
[0007] Generate the spatial feature corresponding to each feature point according to the three-dimensional space coordinates corresponding to each feature point;
[0008] Obtain the point cloud generation target description text corresponding to the initial three-dimensional point cloud, and determine the text feature corresponding to the point cloud generation target description text;
[0009] Input the text feature and the spatial feature corresponding to each feature point into a pre-trained three-dimensional point cloud generation model, and output a target three-dimensional point cloud;
[0010] Wherein, the number of feature points in the target three-dimensional point cloud is greater than the number of feature points in the initial three-dimensional point cloud, and the pre-trained three-dimensional point cloud generation model is trained and generated based on a sample description text and a sample three-dimensional point cloud.
[0011] To achieve the above objective, an embodiment of this application provides a three-dimensional point cloud generation apparatus, including:
[0012] The first acquisition module is used to acquire an initial three-dimensional point cloud and determine the three-dimensional spatial coordinates corresponding to each feature point in the initial three-dimensional point cloud;
[0013] The generation module is used to generate spatial features corresponding to each feature point according to the three-dimensional spatial coordinates corresponding to each feature point;
[0014] The second acquisition module is used to acquire a point cloud generation target description text corresponding to the initial three-dimensional point cloud and determine text features corresponding to the point cloud generation target description text;
[0015] The input module is used to input the text features and the spatial features corresponding to each feature point into a pre-trained three-dimensional point cloud generation model to output a target three-dimensional point cloud;
[0016] Wherein, the number of feature points in the target three-dimensional point cloud is greater than the number of feature points in the initial three-dimensional point cloud, and the pre-trained three-dimensional point cloud generation model is trained and generated based on a sample description text and a sample three-dimensional point cloud.
[0017] In some embodiments, the first acquisition module is used to:
[0018] Determine the initial three-dimensional spatial coordinates of each feature point in the initial three-dimensional point cloud;
[0019] Multiply the coordinate values in the initial three-dimensional spatial coordinates of each feature point by two to obtain the three-dimensional spatial coordinates corresponding to each feature point in the initial three-dimensional point cloud.
[0020] In some embodiments, the generation module is used to:
[0021] Obtain the resolution level corresponding to the initial three-dimensional point cloud;
[0022] Obtain the octree position number and octree code to which the three-dimensional spatial coordinates corresponding to each feature point belong in the octree, wherein the three-dimensional spatial coordinates corresponding to each feature point are obtained by octree segmentation based on the three-dimensional spatial coordinates of the corresponding parent feature point in the historical three-dimensional point cloud of the previous resolution level;
[0023] Generate spatial features corresponding to each feature point according to the resolution level, the octree position number, and the octree code.
[0024] In some embodiments, the three-dimensional point cloud generation device includes an update module, which is used to:
[0025] After inputting the text features and the spatial features corresponding to each feature point into the pre-trained three-dimensional point cloud generation model to output a target three-dimensional point cloud, determine whether the resolution level corresponding to the target three-dimensional point cloud reaches a preset resolution level;
[0026] When the resolution level corresponding to the target 3D point cloud does not reach the preset resolution level, determine the target 3D point cloud as the initial 3D point cloud, and return to execute the step of determining the 3D spatial coordinates corresponding to each feature point in the initial 3D point cloud until the target 3D point cloud output by the pre-trained 3D point cloud generation model reaches the preset resolution level.
[0027] In some embodiments, the pre-trained 3D point cloud generation model includes a first pre-trained 3D point cloud generation sub-model and a second pre-trained 3D point cloud generation sub-model; an input module, configured to:
[0028] Input the text feature and the spatial feature corresponding to each feature point into the first pre-trained 3D point cloud generation sub-model, and output a first 3D point cloud feature;
[0029] Perform upsampling processing on the first 3D point cloud feature to obtain a second 3D point cloud feature;
[0030] Input the second 3D point cloud feature and the text feature into the second pre-trained 3D point cloud generation sub-model, and output a third 3D point cloud feature;
[0031] Perform convolution processing on the third 3D point cloud feature to obtain the target 3D point cloud.
[0032] In some embodiments, the first pre-trained 3D point cloud generation sub-model includes a plurality of feature extraction modules; an input module, configured to:
[0033] Input the text feature and the spatial feature corresponding to each feature point into the first feature extraction module of the first pre-trained 3D point cloud generation sub-model, and output a point cloud extraction feature;
[0034] Input the point cloud extraction feature into the next feature extraction module of the first feature extraction module, and output an updated point cloud extraction feature;
[0035] Repeat inputting the point cloud extraction feature output by the previous feature extraction module and the text feature into the next feature extraction module, and output an updated point cloud extraction feature until the first 3D point cloud feature output by the last feature extraction module of the first pre-trained 3D point cloud generation sub-model is obtained.
[0036] In some embodiments, the feature extraction module includes a convolution sub-module, a cross-attention sub-module, a first fully-connected sub-module, a self-attention sub-module, and a second fully-connected sub-module connected in sequence; an input module, configured to:
[0037] Input the spatial feature corresponding to each feature point into the convolution sub-module, and output a convolution feature;
[0038] Add the convolutional feature and the spatial feature corresponding to each feature point to obtain a first input feature, and input the first input feature and the text feature into a cross-attention sub-module to output a first attention feature;
[0039] Add the first attention feature and the first input feature to obtain a second input feature, and input the second input feature into a first fully connected sub-module to output a first fully connected feature;
[0040] Add the first fully connected feature and the second input feature to obtain a third input feature, and input the third input feature into a self-attention sub-module to output a second attention feature;
[0041] Add the second attention feature and the third input feature to obtain a fourth input feature, and input the fourth input feature into a second fully connected sub-module to output a second fully connected feature;
[0042] Add the second fully connected feature and the fourth input feature to obtain a point cloud extraction feature.
[0043] In some embodiments, the second pre-trained three-dimensional point cloud generation sub-model includes a plurality of feature extraction modules; an input module for:
[0044] Input the text feature and the second three-dimensional point cloud feature into the first feature extraction module of the second pre-trained three-dimensional point cloud generation sub-model to output a target point cloud extraction feature;
[0045] Input the target point cloud extraction feature into the next feature extraction module of the first feature extraction module to output an updated target point cloud extraction feature;
[0046] Repeat the process of inputting the target point cloud extraction feature output by the previous feature extraction module and the text feature into the next feature extraction module to output an updated target point cloud extraction feature until the third three-dimensional point cloud feature output by the last feature extraction module of the second pre-trained three-dimensional point cloud generation sub-model is obtained.
[0047] In some embodiments, the pre-trained three-dimensional point cloud generation model includes a plurality of sequentially connected convolutional sub-models; an input module for:
[0048] Input the third three-dimensional point cloud feature into the first convolutional sub-model to output an intermediate convolutional feature;
[0049] Input the intermediate convolutional feature output by the first convolutional sub-model into the next convolutional sub-model to output an updated intermediate convolutional feature;
[0050] Repeatedly input the intermediate convolutional features output by the previous convolutional sub-model into the next convolutional sub-model, and output the updated intermediate convolutional features until the last convolutional sub-model outputs the target three-dimensional point cloud.
[0051] In some embodiments, the three-dimensional point cloud generation device further includes a training module for:
[0052] Before inputting the text features and the spatial features corresponding to each feature point into the pre-trained three-dimensional point cloud generation model to output the target three-dimensional point cloud, obtain a sample three-dimensional point cloud and the labeled three-dimensional point cloud corresponding to the sample three-dimensional point cloud;
[0053] Determine the sample three-dimensional spatial coordinates corresponding to each sample feature point in the sample three-dimensional point cloud;
[0054] Generate the sample spatial features corresponding to each sample feature point according to the sample three-dimensional spatial coordinates corresponding to each sample feature point;
[0055] Obtain the sample description text corresponding to the labeled three-dimensional point cloud, and determine the sample text features corresponding to the sample description text;
[0056] Input the sample text features and the sample spatial features corresponding to each sample feature point into the three-dimensional point cloud generation model to output a predicted three-dimensional point cloud;
[0057] Determine the difference between the predicted three-dimensional point cloud and the labeled three-dimensional point cloud. When the difference does not meet the preset iteration condition, iterate the three-dimensional point cloud generation model according to the difference, and return to execute the step of determining the sample three-dimensional spatial coordinates corresponding to each sample feature point in the sample three-dimensional point cloud until the difference meets the preset iteration condition to obtain the pre-trained three-dimensional point cloud model;
[0058] Wherein, the number of feature points of the predicted three-dimensional point cloud and the number of feature points of the labeled three-dimensional point cloud are both greater than the number of feature points of the sample three-dimensional point cloud.
[0059] To achieve the above object, an embodiment of the present application provides a computer-readable storage medium storing multiple instructions suitable for being loaded by a processor to execute the three-dimensional point cloud generation method provided by the embodiment of the present application.
[0060] To achieve the above object, an embodiment of the present application provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the three-dimensional point cloud generation method provided by the embodiment of the present application.
[0061] In an embodiment of the present application, an initial three-dimensional point cloud is obtained, and the three-dimensional spatial coordinates corresponding to each feature point in the initial three-dimensional point cloud are determined; spatial features corresponding to each feature point are generated according to the three-dimensional spatial coordinates corresponding to each feature point; a point cloud generation target description text corresponding to the initial three-dimensional point cloud is obtained, and text features corresponding to the point cloud generation target description text are determined; the text features and the spatial features corresponding to each feature point are input into a pre-trained three-dimensional point cloud generation model to output a target three-dimensional point cloud; wherein, the number of feature points in the target three-dimensional point cloud is greater than the number of feature points in the initial three-dimensional point cloud, and the pre-trained three-dimensional point cloud generation model is trained and generated based on a sample description text and a sample three-dimensional point cloud.
[0062] Thus, by obtaining the three-dimensional spatial coordinates of each feature point in the initial three-dimensional point cloud, the spatial features of each feature point are determined through the three-dimensional spatial coordinates of each feature point, and then the point cloud generation target description text corresponding to the initial three-dimensional point cloud and the text features corresponding to the point cloud generation target description text are determined. The spatial features of each feature point and the text features of the initial three-dimensional point cloud are input into the pre-trained three-dimensional point cloud model. The pre-trained three-dimensional point cloud model can perform feature point expansion and feature point screening on the initial three-dimensional point cloud based on the text features, so as to obtain a target three-dimensional point cloud with higher resolution. The number of feature points in the target three-dimensional point cloud is greater than the number of feature points in the initial three-dimensional point cloud. Compared with the related art where the resolution of a three-dimensional point cloud cannot be improved, in the present application, a pre-trained point cloud generation model can be used to process a low-resolution three-dimensional point cloud to generate a high-resolution three-dimensional point cloud.
[0063] Other features and advantages of the present application will be described in the subsequent specification, and, in part, will become apparent from the specification or will be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following described drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0065] Figure 1 It is a schematic diagram of the system framework corresponding to the three-dimensional point cloud generation method provided by the embodiment of the present application;
[0066] Figure 2 It is a schematic diagram of the scenario of the three-dimensional point cloud generation method provided by the embodiment of the present application;
[0067] Figure 3It is a schematic flowchart of the 3D point cloud generation method provided by an embodiment of the present application;
[0068] Figure 4 It is a schematic diagram of feature points in an octree provided by an embodiment of the present application;
[0069] Figure 5 It is a schematic structural diagram of a pre-trained 3D point cloud generation model provided by an embodiment of the present application;
[0070] Figure 6 It is a schematic structural diagram of a first pre-trained 3D point cloud generation sub-model provided by an embodiment of the present application;
[0071] Figure 7 It is a schematic structural diagram of a feature extraction module provided by an embodiment of the present application;
[0072] Figure 8 It is a schematic structural diagram of a second pre-trained 3D point cloud generation sub-model provided by an embodiment of the present application;
[0073] Figure 9 It is a schematic structural diagram of an output module provided by an embodiment of the present application;
[0074] Figure 10 It is another schematic flowchart of the 3D point cloud generation method provided by an embodiment of the present application;
[0075] Figure 11 It is a schematic structural diagram of a 3D point cloud generation device provided by an embodiment of the present application;
[0076] Figure 12 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0077] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts fall within the protection scope of the present application.
[0078] It should be noted that in the specific implementation manners of the present application, when it comes to data related to 3D point clouds, when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards.
[0079] In some processes described in the specification, claims, and the above-mentioned drawings, there are multiple steps that appear in a specific order. However, it should be clearly understood that these steps can be executed not in the order in which they appear herein or in parallel. The step numbers are only used to distinguish different steps, and the numbers themselves do not represent any execution order. In addition, descriptions such as "first", "second", or "target" in this text are used to distinguish similar objects and do not necessarily describe a specific order or sequence.
[0080] The 3D point cloud generation method provided by the embodiments of this application relates to the field of artificial intelligence technology. The 3D point cloud generation method provided by the embodiments of this application can be applied to a terminal, or to a server, or can be software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the 3D point cloud generation method, etc., but is not limited to the above forms.
[0081] The embodiments of this application provide a 3D point cloud generation method, apparatus, computer device, and storage medium. Specifically, the embodiments of this application will be described from the perspective of the 3D point cloud generation apparatus. The 3D point cloud generation apparatus can be specifically integrated in a computer device, which can be a server or a terminal device, etc. Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or can be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Among them, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, smart home appliances, in-vehicle terminals, intelligent voice interaction devices, aircraft, etc., but is not limited thereto. The embodiments of this application can be applied to various scenarios, including but not limited to cloud office, enterprise management, etc.
[0082] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are explained. The nouns and terms involved in the embodiments of this application are applicable to the following explanations:
[0083] 3D Point Cloud: A 3D point cloud is a set of a large number of discrete points in 3D space, used to represent the shape and surface information of an object or a scene. Each point usually contains 3D coordinate information (X, Y, Z). The X, Y, and Z axes can be defined according to a specific coordinate system. Generally speaking, X and Y represent planar positions, and Z represents height or depth.
[0084] Feature Point: A feature point is a point that constitutes a 3D point cloud, specifically each point on the 3D point cloud.
[0085] The above are the explanations of related terms. If other terms are involved in the following text, they will be explained in the following text.
[0086] First, describe the technical problems existing in the related art:
[0087] A 3D point cloud model is an important manifestation of the digitalization of real objects. Through computer technology, a 3D point cloud corresponding to an object can be generated. Then, through the 3D point cloud, the development and testing of 3D point cloud-related algorithms, as well as the development of robot interaction control methods related to 3D point clouds, can be realized.
[0088] In the related art, a 3D point cloud is generally obtained from the real world through 3D scanning devices (such as lidar, depth cameras, etc.). However, due to device and cost limitations, the obtained 3D point cloud has the problem of low resolution. The most obvious manifestation is the small number of feature points in the 3D point cloud. In the related art, the 3D point cloud cannot be edited and changed, so the resolution of the corresponding 3D point cloud cannot be improved.
[0089] To solve the above technical problems, the embodiments of the present application provide a pre-trained 3D point cloud generation model. By obtaining the 3D spatial coordinates of each feature point in the initial 3D point cloud, determining the spatial features of each feature point through the 3D spatial coordinates of each feature point, and then determining the point cloud generation target description text corresponding to the initial 3D point cloud and the text features corresponding to the point cloud generation target description text, the spatial features of each feature point and the text features of the initial 3D point cloud are input into the pre-trained 3D point cloud model. The pre-trained 3D point cloud model can expand the feature points and screen the feature points of the initial 3D point cloud based on the text features, so as to obtain a target 3D point cloud with a higher resolution. The number of feature points of the target 3D point cloud is greater than the number of feature points in the initial 3D point cloud. Compared with the related art where the resolution of a 3D point cloud cannot be improved, in the present application, a low-resolution 3D point cloud can be processed by the pre-trained point cloud generation model to generate a high-resolution 3D point cloud.
[0090] Please refer to Figure 1 , Figure 1It is a schematic diagram of the system framework corresponding to the 3D point cloud generation method provided by the embodiments of the present application. The 3D point cloud generation method provided by the embodiments of the present application can be applied to this system framework.
[0091] It includes a terminal 140, the Internet 130, a gateway 120, a server 110, etc.
[0092] The terminal 140 or the server 110 can be a device for executing the 3D point cloud generation method.
[0093] The terminal 140 includes but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc. The embodiments of the present application can be applied to various scenarios, including but not limited to 3D point cloud generation, 3D modeling, etc. In addition, it can be a single device or a collection of multiple devices combined. For example, multiple desktop computers are interconnected through a local area network and share a monitor, etc. to work collaboratively, jointly constituting a terminal 140. The terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data.
[0094] The server 110 refers to a computer system that can provide certain services to the terminal 140. Compared with ordinary terminals 140, the server 110 has higher requirements in terms of stability, security, performance, etc. The server 110 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.
[0095] The gateway 120 is also called an inter-network connector and protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a conversion function. Between two systems using different communication protocols, data formats or languages, and even completely different architectures, the gateway is a translator. At the same time, the gateway can also provide filtering and security functions. The message sent by the terminal 140 to the server 110 needs to be sent to the corresponding server 110 through the gateway 120. The message sent by the server 110 to the terminal 140 also needs to be sent to the corresponding terminal 140 through the gateway 120.
[0096] The 3D point cloud generation method in the embodiments of the present application can be applied to multiple scenarios, such as 3D modeling, 3D point cloud generation and other scenarios. The scenarios to which the 3D point cloud generation method in the present application is applied are not limited herein.
[0097] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the scenario of the 3D point cloud generation method provided by the embodiments of the present application.
[0098] Among them, the initial three-dimensional point cloud is a three-dimensional point cloud generation model with low resolution. For example, the number of feature points of the initial three-dimensional point cloud in three-dimensional space is 8.
[0099] Then, determine the three-dimensional space coordinates corresponding to each feature point in the initial three-dimensional point cloud. Then, generate the spatial feature corresponding to each feature point according to the three-dimensional space coordinates corresponding to each feature point. The spatial feature is the feature vector corresponding to each feature point.
[0100] The point cloud generation target description text corresponding to the initial three-dimensional point cloud can be obtained. For example, the point cloud generation target description text is "A chair with a back". The point cloud generation target description text can be encoded to generate the text feature corresponding to the point cloud generation target description text. The text feature can be understood as a vector.
[0101] Input the text feature and the spatial feature corresponding to each feature point into the pre-trained three-dimensional point cloud generation model, and output the target three-dimensional point cloud. Among them, the number of feature points in the target three-dimensional point cloud is greater than the number of feature points in the initial three-dimensional point cloud. For example, the number of feature points obtained in the target three-dimensional point cloud is 16, and the resolution of the target three-dimensional point cloud is higher than that of the initial three-dimensional point cloud. Among them, the pre-trained three-dimensional point cloud generation model is trained and generated based on the sample description text and the sample three-dimensional point cloud.
[0102] As can be seen from the above, in this application, by obtaining the three-dimensional space coordinates of each feature point in the initial three-dimensional point cloud, determining the spatial feature of each feature point through the three-dimensional space coordinates of each feature point, and then determining the point cloud generation target description text corresponding to the initial three-dimensional point cloud and the text feature corresponding to the point cloud generation target description text, the spatial feature of each feature point and the text feature of the initial three-dimensional point cloud are input into the pre-trained three-dimensional point cloud model. The pre-trained three-dimensional point cloud model can expand the feature points and screen the feature points of the initial three-dimensional point cloud based on the text feature, so as to obtain a target three-dimensional point cloud with higher resolution. The number of feature points in the target three-dimensional point cloud is greater than the number of feature points in the initial three-dimensional point cloud. Compared with the related art in which the resolution of the three-dimensional point cloud cannot be improved, in this application, the pre-trained point cloud generation model can be used to process the low-resolution three-dimensional point cloud to generate a high-resolution three-dimensional point cloud.
[0103] To understand the three-dimensional point cloud generation method provided in the embodiments of the present application in more detail, please refer to Figure 3 , Figure 3 which is a schematic flowchart of the three-dimensional point cloud generation method provided in the embodiments of the present application. The three-dimensional point cloud generation method may include the following steps:
[0104] Step 210: Obtain the initial 3D point cloud and determine the 3D spatial coordinates corresponding to each feature point in the initial 3D point cloud;
[0105] Step 220: Generate the spatial features corresponding to each feature point according to the 3D spatial coordinates corresponding to each feature point;
[0106] Step 230: Obtain the point cloud generation target description text corresponding to the initial 3D point cloud and determine the text features corresponding to the point cloud generation target description text;
[0107] Step 240: Input the text features and the spatial features corresponding to each feature point into the pre-trained 3D point cloud generation model to output the target 3D point cloud; wherein, the number of feature points in the target 3D point cloud is greater than the number of feature points in the initial 3D point cloud, and the pre-trained 3D point cloud generation model is trained and generated based on the sample description text and the sample 3D point cloud.
[0108] The following will describe Steps 210 to 240 in detail.
[0109] In Step 210, the initial 3D point cloud is obtained and the 3D spatial coordinates corresponding to each feature point in the initial 3D point cloud are determined.
[0110] Among them, the initial 3D point cloud is a low-resolution point cloud. For example, the number of feature points is small. In practical applications, a low-resolution point cloud may have too low a resolution to establish a high-precision visualization model or may cause failure to implement relevant point cloud algorithms. Therefore, it is necessary to improve the resolution of the initial 3D point cloud to generate a target 3D point cloud with a higher resolution.
[0111] The 3D spatial coordinates corresponding to each feature point in the initial 3D point cloud can be determined. For example, a 3D coordinate system can be established in 3D space, and in this 3D coordinate system, each feature point corresponds to a 3D spatial coordinate.
[0112] In some embodiments, determining the 3D spatial coordinates corresponding to each feature point in the initial 3D point cloud includes:
[0113] (1.1) Determine the initial 3D spatial coordinates of each feature point in the initial 3D point cloud;
[0114] (1.2) Multiply the coordinate values in the initial 3D spatial coordinates of each feature point by two to obtain the 3D spatial coordinates corresponding to each feature point in the initial 3D point cloud.
[0115] Among them, first, determine the initial three-dimensional spatial coordinates of each feature point in the initial three-dimensional point cloud. The initial three-dimensional spatial coordinates can be understood as the default coordinates in the three-dimensional coordinate system set in the three-dimensional space. However, considering the improvement of the resolution of the initial three-dimensional point cloud in this application, actually, new feature points are added based on the feature points of the initial three-dimensional point cloud to increase the number of feature points, so as to obtain the target three-dimensional point cloud with more feature points.
[0116] Therefore, it is necessary to multiply the coordinate values in the initial three-dimensional spatial coordinates of each feature point by two to obtain the three-dimensional spatial coordinates corresponding to each feature point in the initial three-dimensional point cloud. This can leave enough space for the subsequent newly added feature points, so as to achieve the improvement of the resolution of the initial three-dimensional point cloud.
[0117] For example, the coordinate range in the three-dimensional space of the initial three-dimensional point cloud is 0 - 63. After multiplying the coordinate values in the initial three-dimensional spatial coordinates of each feature point by two to obtain the three-dimensional spatial coordinates corresponding to each feature point in the initial three-dimensional point cloud, the coordinate range in the three-dimensional space corresponding to the initial three-dimensional point cloud is 0 - 126. In this application, after the subsequent upsampling process, the coordinate range in the three-dimensional space is 0 - 127, thus achieving a two-fold improvement in the resolution of the initial three-dimensional point cloud.
[0118] The advantage of doing this is to leave enough space for the subsequent expansion of feature points to ensure the improvement effect of the resolution of the initial three-dimensional point cloud.
[0119] In step 220, generate the spatial feature corresponding to each feature point according to the three-dimensional spatial coordinates corresponding to each feature point.
[0120] Among them, the spatial feature corresponding to each feature point can be generated according to the three-dimensional spatial coordinates corresponding to each feature point. For example, the three-dimensional spatial coordinates corresponding to each feature point can be input into the corresponding function to generate the feature vector corresponding to each feature point, and this feature vector is the spatial feature corresponding to each feature point.
[0121] In some embodiments, generating the spatial feature corresponding to each feature point according to the three-dimensional spatial coordinates corresponding to each feature point includes:
[0122] (1.1) Obtain the resolution level corresponding to the initial three-dimensional point cloud;
[0123] (1.2) Obtain the octree position number and octal code of each feature point in the octree corresponding to the three-dimensional spatial coordinates of each feature point, where the three-dimensional spatial coordinates of each feature point are obtained by octree segmentation based on the three-dimensional spatial coordinates of the corresponding parent feature point in the historical three-dimensional point cloud of the previous resolution level;
[0124] (1.3) Generate the spatial feature corresponding to each feature point according to the resolution level, octree position number, and octant code.
[0125] Among them, in the process of improving the resolution of the three-dimensional point cloud, it involves multiple expansions and improvements of the feature points, that is, multiple resolutions of the three-dimensional point cloud are improved. Each time the resolution of the three-dimensional point cloud is improved, there will be a corresponding resolution level. For example, each time the resolution is improved, the resolution level is increased by one level.
[0126] The initial three-dimensional point cloud can be a three-dimensional point cloud that has never had its resolution improved, or a three-dimensional point cloud with its resolution improved during a certain resolution improvement process. Therefore, the resolution level corresponding to the initial three-dimensional point cloud can be determined according to the number of times the resolution of the initial three-dimensional point cloud is improved. For example, if the initial three-dimensional point cloud is the three-dimensional point cloud with its resolution improved for the third time, the resolution level can be 3.
[0127] Then obtain the octree position number and octant code that the three-dimensional spatial coordinates corresponding to each feature point belong to in the octree. For example, in the process of expanding the feature points, the number of feature points is continuously increasing. The feature points in the previous three-dimensional point cloud are associated with the feature points in the next three-dimensional point cloud. For example, the feature points in the previous three-dimensional point cloud are parent feature points, and the multiple feature points in the next three-dimensional point cloud are expanded based on the parent feature points. These multiple feature points are called child feature points. During this process, there is a certain association relationship between the parent feature point and the multiple child feature points. The parent feature point is the previous node, and the multiple child feature points are the next nodes. The relationship can be specifically represented by an octree.
[0128] Please combine Figure 4 , Figure 4 is a schematic diagram of the feature points in the octree provided by the embodiment of the present application. Among them, each feature point corresponds to an octree position number and an octant code. For example, the octant code position numbers are 1, 2, 3, 4, 5, 6, 7, 8, etc. Each octant code position number corresponds to a feature point. Each feature point also corresponds to three-dimensional spatial coordinates and an octant code. The octant code is determined according to the actual octree position number.
[0129] Finally, generate the spatial feature corresponding to each feature point according to the resolution level, octree position number, and octant code. For example, the resolution level, octree position number, and octant code can be generated into an array, and this array is the spatial feature corresponding to the feature point. For example, the spatial feature is [level, octant, octree code], where level is the resolution level, octant is the octree position number, and octree code is the octant code. The spatial feature can be represented by this array.
[0130] In step 230, obtain the point cloud generation target description text corresponding to the initial three-dimensional point cloud, and determine the text features corresponding to the point cloud generation target description text.
[0131] Among them, the point cloud generation target description text corresponding to the initial three-dimensional point cloud can be obtained. This point cloud generation target description text is used to describe the relevant features of the three-dimensional model. For example, the point cloud generation target description text is: A chair with aback. The point cloud generation target description text can be encoded by a pre-trained text encoder to generate the text features corresponding to the point cloud generation target description text.
[0132] For example, input the point cloud generation target description text into the pre-trained text encoding model. The pre-trained text encoding model can encode different words in the point cloud generation target description text through the vocabulary to generate multiple tokens. Specifically, each word or phrase corresponds to an integer. For example, "A chair" corresponds to 9320, "with a" corresponds to 25609, and "back" corresponds to 12043. Thus, a token sequence like [9320, 25609, 12043] is obtained, and then this token sequence is encoded into a word embedding vector, that is, each token is mapped to a vector, and then position encoding is added to each vector, that is, position information is added to each vector, so as to obtain the text feature vector, that is, the text feature.
[0133] In step 240, input the text features and the spatial features corresponding to each feature point into the pre-trained three-dimensional point cloud generation model to output the target three-dimensional point cloud; among them, the number of feature points in the target three-dimensional point cloud is greater than the number of feature points in the initial three-dimensional point cloud, and the pre-trained three-dimensional point cloud generation model is trained and generated based on the sample description text and the sample three-dimensional point cloud.
[0134] Among them, please refer to Figure 5 , Figure 5 which is the structural schematic diagram of the pre-trained three-dimensional point cloud generation model provided by the embodiments of the present application. The pre-trained three-dimensional point cloud generation model includes a first pre-trained three-dimensional point cloud generation sub-model and a second pre-trained three-dimensional point cloud generation sub-model, and also includes an upsampling module and an output module.
[0135] The first pre-trained 3D point cloud generation sub-model is used to initially combine text features and spatial features to improve the resolution of the initial 3D point cloud. The upsampling module is used to perform upsampling processing on the output result of the first pre-trained 3D point cloud generation sub-model, and then input the upsampled result and text features into the second pre-trained 3D point cloud generation sub-model. The second pre-trained 3D point cloud generation sub-model further selects more relevant feature points from multiple feature points for the point cloud generation target description text and the initial 3D point cloud. Finally, the relevant features of these more relevant feature points are input into the output module, so that the output module outputs the target 3D point cloud.
[0136] In some embodiments, inputting text features and spatial features corresponding to each feature point into a pre-trained 3D point cloud generation model to output a target 3D point cloud includes:
[0137] (1.1) Inputting text features and spatial features corresponding to each feature point into the first pre-trained 3D point cloud generation sub-model to output the first 3D point cloud feature;
[0138] (1.2) Performing upsampling processing on the first 3D point cloud feature to obtain the second 3D point cloud feature;
[0139] (1.3) Inputting the second 3D point cloud feature and text features into the second pre-trained 3D point cloud generation sub-model to output the third 3D point cloud feature;
[0140] (1.4) Performing convolution processing on the third 3D point cloud feature to obtain the target 3D point cloud.
[0141] Among them, inputting text features and spatial features corresponding to each feature point into the first pre-trained 3D point cloud generation sub-model to output the first 3D point cloud feature, and the first 3D point cloud feature can be understood as the 3D point cloud feature obtained by processing the initial 3D point cloud and conforming to the content described in the point cloud generation target description text.
[0142] Then, perform upsampling processing on the first 3D point cloud feature to obtain the second 3D point cloud feature. In the process of this upsampling, the number of feature points of the point cloud can be expanded to obtain the second 3D point cloud feature.
[0143] Then, the second 3D point cloud feature and the text feature are input into the second pre-trained 3D point cloud generation sub-model, and the third 3D point cloud feature is output. In this process, the second pre-trained 3D point cloud generation sub-model can determine the point cloud features of the feature points related to the text of the target description for point cloud generation in the second 3D point cloud feature according to the text feature of the target description text for point cloud generation, so as to obtain the third 3D point cloud feature. The third 3D point cloud feature can be understood as the feature that more conforms to the object described in the target description text for point cloud generation. The third 3D point cloud feature has more feature information of feature points than the first 3D point cloud feature.
[0144] Finally, convolution processing is performed on the third 3D point cloud feature, the feature is dimensionally reduced, and the reduced feature values are screened to obtain the target 3D point cloud. Specifically, in the process of performing convolution processing on the third 3D point cloud feature, the feature corresponding to each point in the 3D space can be dimensionally reduced to one dimension, and then the feature value is mapped to the interval [0, 1]. Finally, the points with feature values greater than or equal to 0.5 are retained, and these retained points constitute the target 3D point cloud. The number of feature points of the target 3D point cloud is more than that of the initial 3D point cloud, so the target 3D point cloud has a higher resolution.
[0145] As can be seen from the above, in this application, through the spatial feature of each feature point in the initial 3D point cloud combined with the text feature of the target description text for point cloud generation of the initial 3D point cloud, the text feature can guide the pre-trained 3D point cloud generation model to generate the corresponding target 3D point cloud, so as to realize the conversion of the initial 3D point cloud with low resolution to generate the target 3D point cloud with high resolution.
[0146] To understand the first pre-trained 3D point cloud generation sub-model provided by the embodiments of this application in more detail, please refer to Figure 6 , Figure 6 which is a schematic structural diagram of the first pre-trained 3D point cloud generation sub-model provided by the embodiments of this application. Among them, the first pre-trained 3D point cloud generation sub-model includes multiple feature extraction modules, specifically N feature extraction modules, and the value of N can be 8.
[0147] In some embodiments, inputting the text feature and the spatial feature corresponding to each feature point into the first pre-trained 3D point cloud generation sub-model and outputting the first 3D point cloud feature includes:
[0148] (1.1.1) Input the text feature and the spatial feature corresponding to each feature point into the first feature extraction module of the first pre-trained 3D point cloud generation sub-model, and output the point cloud extraction feature;
[0149] (1.1.2) Input the point cloud extraction feature into the next feature extraction module of the first feature extraction module, and output the updated point cloud extraction feature;
[0150] (1.1.3) Repeatedly input the point cloud extraction features and text features output by the previous feature extraction module into the next feature extraction module, and output the updated point cloud extraction features until the first 3D point cloud features output by the last feature extraction module of the first pre-trained 3D point cloud generation sub-model are obtained.
[0151] Among them, in the first feature extraction module of the first pre-trained 3D point cloud generation sub-model, input the text features and the spatial features corresponding to each feature point. The first feature extraction module outputs the point cloud extraction features, and then input the point cloud extraction features into the next feature extraction module of the first feature extraction module, and output the updated point cloud extraction features, that is, the second feature extraction module outputs the updated point cloud extraction features.
[0152] And so on, repeatedly input the point cloud extraction features and text features output by the previous feature extraction module into the next feature extraction module, and output the updated point cloud extraction features until the first 3D point cloud features output by the last feature extraction module of the first pre-trained 3D point cloud generation sub-model are obtained.
[0153] As can be seen from the above, in the embodiments of the present application, by setting multiple feature extraction modules, the multiple feature extraction modules can more accurately extract the features of the feature points in the initial 3D point cloud features. At the same time, by introducing text features into each feature extraction module, it is possible to guide the feature extraction of the feature extraction module, so that the point cloud extraction features extracted by each feature extraction module are more in line with the description of the text features.
[0154] For a more detailed understanding of the feature extraction module provided in the embodiments of the present application, please refer to Figure 7 , Figure 7 which is the structural schematic diagram of the feature extraction module provided in the embodiments of the present application.
[0155] Among them, the feature extraction module includes a convolutional sub-module, a cross-attention sub-module, a first fully-connected sub-module, a self-attention sub-module, and a second fully-connected sub-module connected in sequence.
[0156] In some embodiments, inputting the text features and the spatial features corresponding to each feature point into the first feature extraction module of the first pre-trained 3D point cloud generation sub-model and outputting the point cloud extraction features includes:
[0157] (1.1.1.1) Input the spatial features corresponding to each feature point into the convolutional sub-module and output the convolutional features;
[0158] (1.1.1.2) Add the convolutional features and the spatial features corresponding to each feature point to obtain the first input feature, and input the first input feature and the text feature into the cross-attention sub-module to output the first attention feature;
[0159] (1.1.1.3) Add the first attention feature and the first input feature to obtain the second input feature, and input the second input feature into the first fully-connected sub-module to output the first fully-connected feature;
[0160] (1.1.1.4) Add the first fully-connected feature and the second input feature to obtain the third input feature, and input the third input feature into the self-attention sub-module to output the second attention feature;
[0161] (1.1.1.5) Add the second attention feature and the third input feature to obtain the fourth input feature, and input the fourth input feature into the second fully-connected sub-module to output the second fully-connected feature;
[0162] (1.1.1.6) Add the second fully-connected feature and the fourth input feature to obtain the point cloud extraction feature.
[0163] Among them, taking the first feature extraction module in the first pre-trained three-dimensional point cloud generation sub-model as an example, the input of the convolutional sub-module of the feature extraction module is the spatial feature of each feature point in the initial three-dimensional point cloud, and the convolutional sub-module outputs the convolutional feature.
[0164] Then, the convolution feature and the spatial feature corresponding to each feature point are added to obtain the first input feature, and the first input feature and the text feature are input into the cross-attention submodule to output the first attention feature. Among them, the fusion of the convolution feature and the spatial feature corresponding to each feature point can improve the stability of the point cloud generation model during the training process, so that the model can avoid the model non-convergence caused by gradient explosion or gradient disappearance during training on the basis of a deep network with multiple submodules superimposed. Then the text feature and the first input feature are input into the cross-attention submodule, so that the information of the two different modes of text and point cloud can communicate and associate with each other. The text describes the geometric structure information of the point cloud. Through the cross-attention mechanism, the spatial feature of each feature point can be focused on the key area mentioned in the text, and the text feature can also correspond to the specific content of the three-dimensional point cloud, thereby establishing a corresponding relationship between the two and better understanding the cross-modal information. In this way, a deeper feature extraction of the feature points in the initial three-dimensional point cloud in the three-dimensional space can be achieved more accurately, and at the same time, those feature parts that are more important in cross-modal understanding can be highlighted, and irrelevant or secondary information can be suppressed, such as suppressing feature points that are not related to the text description. At the same time, it can provide more targeted information for subsequent feature extraction and processing. For example, when combining text features and spatial features in the point cloud extraction task, the features processed by the cross-attention module can guide the subsequent sub-modules to learn in a direction that is more consistent with cross-modal semantics, helping the model to better capture and utilize the complementary information of the two modalities to achieve more accurate point cloud feature extraction.
[0165] Then, the first attention feature and the first input feature are added to obtain the second input feature, and the second input feature is input into the first fully connected submodule to output the first fully connected feature. The first fully connected submodule can linearly transform the features output by the previous cross-attention submodule. The feature dimensions and distribution output by different submodules may not be suitable for subsequent processing. The first fully connected submodule maps the features to a new feature space through weight matrix operations, so that the features are more in line with the requirements of subsequent self-attention submodules and other processing, which helps to explore the potential relationship between features.
[0166] The first fully connected feature and the second input feature are then added together to obtain the third input feature, and the third input feature is input into the self-attention submodule to output the second attention feature. The self-attention mechanism of the self-attention submodule enables the model to focus on each position in the entire third input feature, rather than just the local area, when processing sequences or features. For the third input feature, it can capture long-distance dependencies between different parts of the feature. For example, when processing point cloud-related features, even if some parts of the feature are far apart in space, there may be semantic or structural associations between them. The self-attention module can discover and utilize this relationship, which helps to more fully understand the overall structure and semantic information of the feature.
[0167] Then, the second attention feature and the third input feature are added to obtain a fourth input feature, and the fourth input feature is input into the second fully-connected sub-module to output a second fully-connected feature. The second fully-connected sub-module can first adjust the dimension of the input fourth input feature. After the self-attention module and the previous addition operation, the dimension of the feature may not meet the requirements of subsequent tasks (such as final point cloud extraction or classification, etc.). The fully-connected layer can map the feature from one dimension space to another through its weight matrix, so that the feature dimension matches the dimension of the final expected output. For example, if the subsequent task requires a lower-dimensional feature representation for efficient calculation or storage, the second fully-connected sub-module can play a role in dimensionality reduction; conversely, if a richer feature representation is needed to improve the expression ability of the model, it can also perform a dimensionality increase operation.
[0168] In addition, the fourth input feature is obtained by adding the second attention feature and the third input feature, and contains the information processed by the self-attention module and the previous original input information. The second fully-connected sub-module will perform a deeper fusion of this information. It uses its own weight parameters to synthesize these features from different sources, so that the fused second fully-connected feature can more effectively extract the information useful for the final task. This fusion method can enhance the semantic expression ability of the feature, better combine the important features highlighted by the self-attention module and other potentially useful information in the previous input features, and provide a more discriminative feature representation for subsequent tasks.
[0169] Finally, the second fully-connected feature and the fourth input feature are added to obtain a point cloud extraction feature. The fourth input feature contains the original information in the early stage and the feature combination after a series of processes (such as self-attention, etc.). The second fully-connected feature is obtained after being processed by the fully-connected layer, and it has refined and transformed some key information. Adding the two can retain both the useful information in the original input and the important features extracted by the fully-connected layer, playing a role in information enhancement.
[0170] As can be seen from the above, through the feature extraction module, the point cloud extraction feature of the input initial three-dimensional point cloud can be accurately extracted.
[0171] Please refer to Figure 8 , Figure 8It is a schematic structural diagram of the second pre-trained 3D point cloud generation sub-model provided by an embodiment of the present application. The second pre-trained 3D point cloud generation sub-model includes multiple feature extraction modules, specifically M feature extraction modules, and the value of M can be 2. The specific structures of the multiple feature extraction modules in the second pre-trained 3D point cloud generation sub-model are the same as those of the multiple feature extraction modules in the first pre-trained 3D point cloud generation sub-model, specifically as shown in Figure 7 shown, except that the inputs of the multiple feature extraction modules in the second pre-trained 3D point cloud generation sub-model are different.
[0172] In some embodiments, inputting the second 3D point cloud feature and the text feature into the second pre-trained 3D point cloud generation sub-model, and outputting the third 3D point cloud feature, including:
[0173] (1.3.1) Inputting the text feature and the second 3D point cloud feature into the first feature extraction module of the second pre-trained 3D point cloud generation sub-model, and outputting the target point cloud extraction feature;
[0174] (1.3.2) Inputting the target point cloud extraction feature into the next feature extraction module of the first feature extraction module, and outputting the updated target point cloud extraction feature;
[0175] (1.3.3) Repeatedly inputting the target point cloud extraction feature output by the previous feature extraction module and the text feature into the next feature extraction module, and outputting the updated target point cloud extraction feature until the third 3D point cloud feature output by the last feature extraction module of the second pre-trained 3D point cloud generation sub-model is obtained.
[0176] Among them, in the first feature extraction module of the second pre-trained 3D point cloud generation sub-model, inputting the text feature and the second 3D point cloud feature, the first feature extraction module outputs the target point cloud extraction feature, and then inputting the target point cloud extraction feature into the next feature extraction module of the first feature extraction module, and outputting the updated target point cloud extraction feature, that is, the second feature extraction module outputs the updated target point cloud extraction feature.
[0177] And so on, repeatedly inputting the target point cloud extraction feature output by the previous feature extraction module and the text feature into the next feature extraction module, and outputting the updated target point cloud extraction feature until the third 3D point cloud feature output by the last feature extraction module of the second pre-trained 3D point cloud generation sub-model is obtained.
[0178] As can be seen from the above, in the embodiment of the present application, by setting a plurality of feature extraction modules, the plurality of feature extraction modules can more accurately extract the relevant features of the feature points in the second three-dimensional point cloud feature. At the same time, by introducing text features into each feature extraction module, it is possible to guide the feature extraction of the feature extraction module, so that the target point cloud extraction features extracted by each feature extraction module are more in line with the description of the text features.
[0179] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of the output module provided by the embodiment of the present application. In some embodiments, the pre-trained three-dimensional point cloud generation model includes a plurality of sequentially connected convolutional sub-models, and these plurality of sequentially connected convolutional sub-models can be correspondingly included in the Figure 5 output module. The convolutional sub-model can be K, and the value of K can specifically be 3. Performing convolutional processing on the third three-dimensional point cloud feature to obtain the target three-dimensional point cloud includes:
[0180] (1.4.1) Input the third three-dimensional point cloud feature into the first convolutional sub-model to output an intermediate convolutional feature;
[0181] (1.4.2) Input the intermediate convolutional feature output by the first convolutional sub-model into the next convolutional sub-model to output an updated intermediate convolutional feature;
[0182] (1.4.3) Repeat inputting the intermediate convolutional feature output by the previous convolutional sub-model into the next convolutional sub-model to output an updated intermediate convolutional feature until the last convolutional sub-model outputs the target three-dimensional point cloud.
[0183] Among them, input the third three-dimensional point cloud feature into the first convolutional sub-model to output an intermediate convolutional feature, and then input the intermediate convolutional feature output by the first convolutional sub-model into the next convolutional sub-model to output an updated intermediate convolutional feature. And so on, it is possible to repeat inputting the intermediate convolutional feature output by the previous convolutional sub-model into the next convolutional sub-model to output an updated intermediate convolutional feature until the last convolutional sub-model outputs the target three-dimensional point cloud.
[0184] Among them, the convolutional sub-model extracts deeper features related to the input three-dimensional point cloud, thereby achieving dimensionality reduction of the features. Through convolutional operations, the local geometric structure and spatial features of the three-dimensional point cloud can be captured, such as the relative position relationship between points, local shape features, etc. Different convolutional sub-models can transform and fuse the features. During the processing of multiple consecutive convolutional sub-models, the features are continuously transformed into a new feature space, and the features of different channels are also fused with each other, so that the features are continuously compressed until each three-dimensional point corresponds to a one-dimensional feature value. Then, the feature values are mapped to the interval [0, 1]. Finally, the points with feature values greater than or equal to 0.5 are retained, and these retained points constitute the target three-dimensional point cloud.
[0185] In some embodiments, after inputting the text features and the spatial features corresponding to each feature point into the pre-trained three-dimensional point cloud generation model and outputting the target three-dimensional point cloud, it further includes:
[0186] (2.1) Determine whether the resolution level corresponding to the target three-dimensional point cloud reaches the preset resolution level;
[0187] (2.2) When the resolution level corresponding to the target three-dimensional point cloud does not reach the preset resolution level, determine the target three-dimensional point cloud as the initial three-dimensional point cloud, and return to execute the step of determining the three-dimensional spatial coordinates corresponding to each feature point in the initial three-dimensional point cloud until the target three-dimensional point cloud output by the pre-trained three-dimensional point cloud generation model reaches the preset resolution level.
[0188] Among them, for the resolution improvement of the three-dimensional point cloud, it can be improved multiple times to obtain the target three-dimensional point cloud at the preset resolution level. First, it can be determined whether the resolution level corresponding to the target three-dimensional point cloud reaches the preset resolution level. If it is determined that the resolution level corresponding to the target three-dimensional point cloud reaches the preset resolution level, then the target three-dimensional point cloud is used as the target three-dimensional point cloud at the finally output preset resolution level.
[0189] When the resolution level corresponding to the target three-dimensional point cloud does not reach the preset resolution level, determine the target three-dimensional point cloud as the initial three-dimensional point cloud, and return to execute the step of determining the three-dimensional spatial coordinates corresponding to each feature point in the initial three-dimensional point cloud until the target three-dimensional point cloud output by the pre-trained three-dimensional point cloud generation model reaches the preset resolution level.
[0190] In some embodiments, before inputting the text features and the spatial features corresponding to each feature point into the pre-trained three-dimensional point cloud generation model and outputting the target three-dimensional point cloud, it further includes:
[0191] (3.1) Obtain the sample three-dimensional point cloud and the labeled three-dimensional point cloud corresponding to the sample three-dimensional point cloud;
[0192] (3.2) Determine the sample three-dimensional spatial coordinates corresponding to each sample feature point in the sample three-dimensional point cloud;
[0193] (3.3) Generate the sample spatial features corresponding to each sample feature point according to the sample three-dimensional spatial coordinates corresponding to each sample feature point;
[0194] (3.4) Obtain the sample description text corresponding to the labeled three-dimensional point cloud, and determine the sample text features corresponding to the sample description text;
[0195] (3.5) Input the sample text features and the sample spatial features corresponding to each sample feature point into the three-dimensional point cloud generation model, and output the predicted three-dimensional point cloud;
[0196] (3.6) Determine the difference between the predicted three-dimensional point cloud and the labeled three-dimensional point cloud. When the difference does not meet the preset iteration condition, iterate the three-dimensional point cloud generation model according to the difference, and return to execute the step of determining the sample three-dimensional spatial coordinates corresponding to each sample feature point in the sample three-dimensional point cloud until the difference meets the preset iteration condition, and obtain the pre-trained three-dimensional point cloud model; wherein, the number of feature points of the predicted three-dimensional point cloud and the number of feature points of the labeled three-dimensional point cloud are both greater than the number of feature points of the sample three-dimensional point cloud.
[0197] In this application, the pre-trained three-dimensional point cloud generation model is trained and generated based on the sample description text and the sample three-dimensional point cloud. First, the sample three-dimensional point cloud and the labeled three-dimensional point cloud corresponding to the sample three-dimensional point cloud can be obtained. The sample three-dimensional point cloud is used as training data, and the labeled three-dimensional point cloud corresponding to the sample three-dimensional point cloud is used as label data, and the labeled three-dimensional point cloud is real data. For example, the sample three-dimensional point cloud is a low-resolution three-dimensional point cloud, and the labeled three-dimensional point cloud is a high-resolution labeled three-dimensional point cloud. The sample three-dimensional point cloud corresponds to a chair composed of feature points with a first number of feature points, and the labeled three-dimensional point cloud corresponds to a chair composed of feature points with a second number of feature points, and the second number of feature points is higher than the first number of feature points.
[0198] Then determine the sample three-dimensional spatial coordinates corresponding to each sample feature point in the sample three-dimensional point cloud. For example, in a preset three-dimensional coordinate system in three-dimensional space, determine the sample three-dimensional spatial coordinates corresponding to each sample feature point in the sample three-dimensional point cloud.
[0199] Then, generate the sample spatial feature corresponding to each sample feature point according to the sample three-dimensional space coordinates corresponding to each sample feature point. Specifically, the resolution level corresponding to the sample three-dimensional point cloud can be obtained; obtain the octree position number and octal code of the three-dimensional space coordinates corresponding to each feature point in the octree, where the octal code corresponding to each feature point is obtained by octree segmentation based on the three-dimensional space coordinates of the corresponding parent feature point in the historical three-dimensional point cloud of the previous resolution level; generate the spatial feature corresponding to each feature point according to the resolution level, octree position number, and octal code.
[0200] Then, obtain the sample description text corresponding to the labeled three-dimensional point cloud, and determine the sample text feature corresponding to the sample description text. Input the sample text feature and the sample spatial feature corresponding to each sample feature point into the three-dimensional point cloud generation model, and output the predicted three-dimensional point cloud. Finally, determine the difference between the predicted three-dimensional point cloud and the labeled three-dimensional point cloud. When the difference does not meet the preset iteration condition, iterate the three-dimensional point cloud generation model according to the difference, and return to execute the step of determining the three-dimensional space coordinates corresponding to each sample feature point in the sample three-dimensional point cloud until the difference meets the preset iteration condition, and obtain the pre-trained three-dimensional point cloud model; wherein, the number of feature points of the predicted three-dimensional point cloud and the number of feature points of the labeled three-dimensional point cloud are both greater than the number of feature points of the sample three-dimensional point cloud.
[0201] Among them, to determine the difference between the predicted three-dimensional point cloud and the labeled three-dimensional point cloud, the difference between the two can be determined by means of binary cross-entropy. For example, calculate the average value of the binary cross-entropy of the corresponding point feature values in the predicted point cloud and the labeled point cloud. Determine this average value as the difference between the two. The preset iteration condition can be that the difference between the two is less than the preset difference value. Or the preset iteration condition is that the difference between the two is stable in an interval and does not decrease as the iteration continues.
[0202] In the embodiment of the present application, by obtaining the initial three-dimensional point cloud and determining the three-dimensional space coordinates corresponding to each feature point in the initial three-dimensional point cloud; generating the spatial feature corresponding to each feature point according to the three-dimensional space coordinates corresponding to each feature point; obtaining the point cloud generation target description text corresponding to the initial three-dimensional point cloud, and determining the text feature corresponding to the point cloud generation target description text; inputting the text feature and the spatial feature corresponding to each feature point into the pre-trained three-dimensional point cloud generation model, and outputting the target three-dimensional point cloud; wherein, the number of feature points in the target three-dimensional point cloud is greater than the number of feature points in the initial three-dimensional point cloud, and the pre-trained three-dimensional point cloud generation model is trained and generated based on the sample description text and the sample three-dimensional point cloud.
[0203] Thus, by obtaining the three-dimensional spatial coordinates of each feature point in the initial three-dimensional point cloud, determining the spatial features of each feature point based on the three-dimensional spatial coordinates of each feature point, and then determining the point cloud generation target description text corresponding to the initial three-dimensional point cloud and the text features corresponding to the point cloud generation target description text, the spatial features of each feature point and the text features of the initial three-dimensional point cloud are input into a pre-trained three-dimensional point cloud model. The pre-trained three-dimensional point cloud model can expand the feature points and screen the feature points of the initial three-dimensional point cloud based on the text features, so as to obtain a target three-dimensional point cloud with higher resolution. The number of feature points of the target three-dimensional point cloud is greater than the number of feature points in the initial three-dimensional point cloud. Compared with the related art where the resolution of a three-dimensional point cloud cannot be improved, in this application, a pre-trained point cloud generation model can be used to process a low-resolution three-dimensional point cloud to generate a high-resolution three-dimensional point cloud.
[0204] Please refer to Figure 10 , Figure 10 which is another schematic flowchart of the three-dimensional point cloud generation method provided by the embodiments of this application. The three-dimensional point cloud generation method may include the following steps:
[0205] Step 301, obtain a sample three-dimensional point cloud and a labeled three-dimensional point cloud corresponding to the sample three-dimensional point cloud, and determine the sample three-dimensional spatial coordinates corresponding to each sample feature point in the sample three-dimensional point cloud;
[0206] Step 302, generate the sample spatial features corresponding to each sample feature point according to the sample three-dimensional spatial coordinates corresponding to each sample feature point;
[0207] Step 303, obtain the sample description text corresponding to the labeled three-dimensional point cloud, and determine the sample text features corresponding to the sample description text;
[0208] Step 304, input the sample text features and the sample spatial features corresponding to each sample feature point into the three-dimensional point cloud generation model, and output a predicted three-dimensional point cloud;
[0209] Step 305, determine the difference between the predicted three-dimensional point cloud and the labeled three-dimensional point cloud. When the difference does not meet the preset iteration condition, iterate the three-dimensional point cloud generation model according to the difference, and return to execute the step of determining the sample three-dimensional spatial coordinates corresponding to each sample feature point in the sample three-dimensional point cloud until the difference meets the preset iteration condition to obtain a pre-trained three-dimensional point cloud model; wherein, the number of feature points of the predicted three-dimensional point cloud and the number of feature points of the labeled three-dimensional point cloud are both greater than the number of feature points of the sample three-dimensional point cloud;
[0210] Step 306, determine the initial three-dimensional spatial coordinates of each feature point in the initial three-dimensional point cloud;
[0211] Step 307: Multiply the coordinate values in the initial three-dimensional space coordinates of each feature point by two to obtain the three-dimensional space coordinates corresponding to each feature point in the initial three-dimensional point cloud;
[0212] Step 308: Obtain the resolution level corresponding to the initial three-dimensional point cloud, and obtain the octree position number and octal code of each feature point in the octree where the three-dimensional space coordinates of each feature point are located. The three-dimensional space coordinates of each feature point are obtained by octree segmentation based on the three-dimensional space coordinates of the corresponding parent feature point in the historical three-dimensional point cloud of the previous resolution level;
[0213] Step 309: Generate the spatial feature corresponding to each feature point according to the resolution level, octree position number, and octal code;
[0214] Step 310: Obtain the point cloud generation target description text corresponding to the initial three-dimensional point cloud, and determine the text feature corresponding to the point cloud generation target description text;
[0215] Step 311: Input the text feature and the spatial feature corresponding to each feature point into the first pre-trained three-dimensional point cloud generation sub-model, output the first three-dimensional point cloud feature, and perform upsampling processing on the first three-dimensional point cloud feature to obtain the second three-dimensional point cloud feature;
[0216] Step 312: Input the second three-dimensional point cloud feature and the text feature into the second pre-trained three-dimensional point cloud generation sub-model, output the third three-dimensional point cloud feature, and perform convolution processing on the third three-dimensional point cloud feature to obtain the target three-dimensional point cloud.
[0217] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the detailed description of the three-dimensional point cloud generation method above, and details will not be repeated here.
[0218] Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of the three-dimensional point cloud generation device provided in the embodiments of the present application. This three-dimensional point cloud generation device is used to execute the above three-dimensional point cloud generation method.
[0219] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0220] The three-dimensional point cloud generation device 400 includes:
[0221] The first acquisition module 410 is configured to acquire an initial three-dimensional point cloud and determine the three-dimensional spatial coordinates corresponding to each feature point in the initial three-dimensional point cloud;
[0222] The generation module 420 is configured to generate spatial features corresponding to each feature point according to the three-dimensional spatial coordinates corresponding to each feature point;
[0223] The second acquisition module 430 is configured to acquire a point cloud generation target description text corresponding to the initial three-dimensional point cloud and determine text features corresponding to the point cloud generation target description text;
[0224] The input module 440 is configured to input the text features and the spatial features corresponding to each feature point into a pre-trained three-dimensional point cloud generation model and output a target three-dimensional point cloud;
[0225] Wherein, the number of feature points in the target three-dimensional point cloud is greater than the number of feature points in the initial three-dimensional point cloud, and the pre-trained three-dimensional point cloud generation model is trained and generated based on a sample description text and a sample three-dimensional point cloud.
[0226] In some embodiments, the first acquisition module 410 is configured to:
[0227] Determine the initial three-dimensional spatial coordinates of each feature point in the initial three-dimensional point cloud;
[0228] Multiply the coordinate values in the initial three-dimensional spatial coordinates of each feature point by two to obtain the three-dimensional spatial coordinates corresponding to each feature point in the initial three-dimensional point cloud.
[0229] In some embodiments, the generation module 420 is configured to:
[0230] Obtain the resolution level corresponding to the initial three-dimensional point cloud;
[0231] Obtain the octree position number and octal code of each feature point in the octree corresponding to the three-dimensional spatial coordinates of each feature point, wherein the three-dimensional spatial coordinates of each feature point are obtained by octree segmentation based on the three-dimensional spatial coordinates of the corresponding parent feature point in the historical three-dimensional point cloud of the previous resolution level;
[0232] Generate spatial features corresponding to each feature point according to the resolution level, octree position number, and octal code.
[0233] In some embodiments, the three-dimensional point cloud generation device 400 includes an update module, which is configured to:
[0234] After inputting the text features and the spatial features corresponding to each feature point into the pre-trained three-dimensional point cloud generation model and outputting the target three-dimensional point cloud, determine whether the resolution level corresponding to the target three-dimensional point cloud reaches a preset resolution level;
[0235] When the resolution level corresponding to the target 3D point cloud does not reach the preset resolution level, the target 3D point cloud is determined as the initial 3D point cloud, and the step of determining the 3D spatial coordinates corresponding to each feature point in the initial 3D point cloud is executed until the target 3D point cloud output by the pre-trained 3D point cloud generation model reaches the preset resolution level.
[0236] In some embodiments, the pre-trained 3D point cloud generation model includes a first pre-trained 3D point cloud generation sub-model and a second pre-trained 3D point cloud generation sub-model; an input module 440, configured to:
[0237] Input the text feature and the spatial feature corresponding to each feature point into the first pre-trained 3D point cloud generation sub-model, and output a first 3D point cloud feature;
[0238] Perform upsampling processing on the first 3D point cloud feature to obtain a second 3D point cloud feature;
[0239] Input the second 3D point cloud feature and the text feature into the second pre-trained 3D point cloud generation sub-model, and output a third 3D point cloud feature;
[0240] Perform convolution processing on the third 3D point cloud feature to obtain the target 3D point cloud.
[0241] In some embodiments, the first pre-trained 3D point cloud generation sub-model includes a plurality of feature extraction modules; an input module 440, configured to:
[0242] Input the text feature and the spatial feature corresponding to each feature point into the first feature extraction module of the first pre-trained 3D point cloud generation sub-model, and output a point cloud extraction feature;
[0243] Input the point cloud extraction feature into the next feature extraction module of the first feature extraction module, and output an updated point cloud extraction feature;
[0244] Repeat inputting the point cloud extraction feature output by the previous feature extraction module and the text feature into the next feature extraction module, and output an updated point cloud extraction feature until the first 3D point cloud feature output by the last feature extraction module of the first pre-trained 3D point cloud generation sub-model is obtained.
[0245] In some embodiments, the feature extraction module includes a convolutional sub-module, a cross-attention sub-module, a first fully connected sub-module, a self-attention sub-module, and a second fully connected sub-module connected in sequence; an input module 440, configured to:
[0246] Input the spatial feature corresponding to each feature point into the convolutional sub-module, and output a convolutional feature;
[0247] Add the convolutional features and the spatial features corresponding to each feature point to obtain the first input feature, and input the first input feature and the text feature into the cross-attention sub-module to output the first attention feature;
[0248] Add the first attention feature and the first input feature to obtain the second input feature, and input the second input feature into the first fully-connected sub-module to output the first fully-connected feature;
[0249] Add the first fully-connected feature and the second input feature to obtain the third input feature, and input the third input feature into the self-attention sub-module to output the second attention feature;
[0250] Add the second attention feature and the third input feature to obtain the fourth input feature, and input the fourth input feature into the second fully-connected sub-module to output the second fully-connected feature;
[0251] Add the second fully-connected feature and the fourth input feature to obtain the point cloud extraction feature.
[0252] In some embodiments, the second pre-trained three-dimensional point cloud generation sub-model includes multiple feature extraction modules; an input module 440 for:
[0253] Input the text feature and the second three-dimensional point cloud feature into the first feature extraction module of the second pre-trained three-dimensional point cloud generation sub-model to output the target point cloud extraction feature;
[0254] Input the target point cloud extraction feature into the next feature extraction module of the first feature extraction module to output the updated target point cloud extraction feature;
[0255] Repeat inputting the target point cloud extraction feature output by the previous feature extraction module and the text feature into the next feature extraction module to output the updated target point cloud extraction feature until the third three-dimensional point cloud feature output by the last feature extraction module of the second pre-trained three-dimensional point cloud generation sub-model is obtained.
[0256] In some embodiments, the pre-trained three-dimensional point cloud generation model includes multiple sequentially-connected convolutional sub-models; an input module 440 for:
[0257] Input the third three-dimensional point cloud feature into the first convolutional sub-model to output the intermediate convolutional feature;
[0258] Input the intermediate convolutional feature output by the first convolutional sub-model into the next convolutional sub-model to output the updated intermediate convolutional feature;
[0259] Repeatedly input the intermediate convolutional features output by the previous convolutional sub-model into the next convolutional sub-model, and output the updated intermediate convolutional features until the last convolutional sub-model outputs the target three-dimensional point cloud.
[0260] In some embodiments, the three-dimensional point cloud generation device 400 further includes a training module for:
[0261] Before inputting the text features and the spatial features corresponding to each feature point into the pre-trained three-dimensional point cloud generation model to output the target three-dimensional point cloud, obtain the sample three-dimensional point cloud and the labeled three-dimensional point cloud corresponding to the sample three-dimensional point cloud;
[0262] Determine the sample three-dimensional spatial coordinates corresponding to each sample feature point in the sample three-dimensional point cloud;
[0263] Generate the sample spatial features corresponding to each sample feature point according to the sample three-dimensional spatial coordinates corresponding to each sample feature point;
[0264] Obtain the sample description text corresponding to the labeled three-dimensional point cloud, and determine the sample text features corresponding to the sample description text;
[0265] Input the sample text features and the sample spatial features corresponding to each sample feature point into the three-dimensional point cloud generation model, and output the predicted three-dimensional point cloud;
[0266] Determine the difference between the predicted three-dimensional point cloud and the labeled three-dimensional point cloud. When the difference does not meet the preset iteration condition, iterate the three-dimensional point cloud generation model according to the difference, and return to execute the step of determining the sample three-dimensional spatial coordinates corresponding to each sample feature point in the sample three-dimensional point cloud until the difference meets the preset iteration condition to obtain the pre-trained three-dimensional point cloud model;
[0267] Wherein, the number of feature points of the predicted three-dimensional point cloud and the number of feature points of the labeled three-dimensional point cloud are both greater than the number of feature points of the sample three-dimensional point cloud.
[0268] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference may be made to the detailed description of the above three-dimensional point cloud generation method, which will not be elaborated here.
[0269] In an embodiment of the present application, an initial three-dimensional point cloud is obtained through a first acquisition module 410, and the three-dimensional spatial coordinates corresponding to each feature point in the initial three-dimensional point cloud are determined; a generation module 420 generates spatial features corresponding to each feature point according to the three-dimensional spatial coordinates corresponding to each feature point; a second acquisition module 430 acquires a point cloud generation target description text corresponding to the initial three-dimensional point cloud, and determines text features corresponding to the point cloud generation target description text; an input module 440 inputs the text features and the spatial features corresponding to each feature point into a pre-trained three-dimensional point cloud generation model, and outputs a target three-dimensional point cloud; wherein, the number of feature points in the target three-dimensional point cloud is greater than the number of feature points in the initial three-dimensional point cloud, and the pre-trained three-dimensional point cloud generation model is trained and generated based on a sample description text and a sample three-dimensional point cloud.
[0270] Thus, by obtaining the three-dimensional spatial coordinates of each feature point in the initial three-dimensional point cloud, determining the spatial features of each feature point through the three-dimensional spatial coordinates of each feature point, then determining the point cloud generation target description text corresponding to the initial three-dimensional point cloud and the text features corresponding to the point cloud generation target description text, inputting the spatial features of each feature point and the text features of the initial three-dimensional point cloud into the pre-trained three-dimensional point cloud model, the pre-trained three-dimensional point cloud model can perform feature point expansion and feature point screening on the initial three-dimensional point cloud based on the text features, so as to obtain a target three-dimensional point cloud with higher resolution, and the number of feature points in the target three-dimensional point cloud is greater than the number of feature points in the initial three-dimensional point cloud. Compared with the related art in which the resolution of a three-dimensional point cloud cannot be improved, in the present application, a pre-trained point cloud generation model can be used to process a low-resolution three-dimensional point cloud to generate a high-resolution three-dimensional point cloud.
[0271] An embodiment of the present application further provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above three-dimensional point cloud generation method is implemented. The computer device can be a server or a terminal.
[0272] Please refer to Figure 12 , Figure 12 which schematically shows the hardware structure of a computer device in another embodiment. The computer device includes:
[0273] A processor 501, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0274] The memory 502 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 502 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 502, and the processor 501 is used to call and execute the three-dimensional point cloud generation method of the embodiments of this application;
[0275] The input / output interface 503 is used to implement information input and output;
[0276] The communication interface 504 is used to implement communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0277] The bus 505 transmits information between various components of the device (such as the processor 501, the memory 502, the input / output interface 503, and the communication interface 504);
[0278] Among them, the processor 501, the memory 502, the input / output interface 503, and the communication interface 504 are communicatively connected to each other inside the device through the bus 505.
[0279] The embodiments of this application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned three-dimensional point cloud generation method is implemented.
[0280] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0281] The embodiments of the present application provide a three-dimensional point cloud generation method, apparatus, computer device, and storage medium. By obtaining an initial three-dimensional point cloud and determining the three-dimensional space coordinates corresponding to each feature point in the initial three-dimensional point cloud; generating a spatial feature corresponding to each feature point according to the three-dimensional space coordinates corresponding to each feature point; obtaining a point cloud generation target description text corresponding to the initial three-dimensional point cloud and determining a text feature corresponding to the point cloud generation target description text; inputting the text feature and the spatial feature corresponding to each feature point into a pre-trained three-dimensional point cloud generation model to output a target three-dimensional point cloud; wherein, the number of feature points in the target three-dimensional point cloud is greater than the number of feature points in the initial three-dimensional point cloud, and the pre-trained three-dimensional point cloud generation model is trained and generated based on a sample description text and a sample three-dimensional point cloud.
[0282] Thus, by obtaining the three-dimensional space coordinates of each feature point in the initial three-dimensional point cloud, determining the spatial feature of each feature point through the three-dimensional space coordinates of each feature point, then determining the point cloud generation target description text corresponding to the initial three-dimensional point cloud and the text feature corresponding to the point cloud generation target description text, and inputting the spatial feature of each feature point and the text feature of the initial three-dimensional point cloud into the pre-trained three-dimensional point cloud model, the pre-trained three-dimensional point cloud model can perform feature point expansion and feature point screening on the initial three-dimensional point cloud based on the text feature, so as to obtain a target three-dimensional point cloud with higher resolution, and the number of feature points in the target three-dimensional point cloud is greater than the number of feature points in the initial three-dimensional point cloud. Compared with the related art where the resolution of a three-dimensional point cloud cannot be improved, in the present application, a pre-trained point cloud generation model can be used to process a low-resolution three-dimensional point cloud to generate a high-resolution three-dimensional point cloud.
[0283] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0284] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown, or combine certain steps, or different steps.
[0285] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0286] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.
[0287] As used in the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0288] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the relationship between associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0289] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned unit division is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.
[0290] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0291] In addition, each functional unit in various embodiments of the present application may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0292] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs and other various media that can store programs.
[0293] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.
Claims
1. A three-dimensional point cloud generation method, characterized in that: include: Acquire an initial three-dimensional point cloud, and determine the three-dimensional space coordinates corresponding to each feature point in the initial three-dimensional point cloud; Generate a spatial feature corresponding to each feature point according to the three-dimensional spatial coordinates corresponding to each feature point; Obtaining a point cloud generation target description text corresponding to the initial three-dimensional point cloud, and determining text features corresponding to the point cloud generation target description text; Input the text features and the spatial features corresponding to each feature point into a pre-trained three-dimensional point cloud generation model, and output a target three-dimensional point cloud; The number of feature points in the target three-dimensional point cloud is greater than the number of feature points in the initial three-dimensional point cloud, and the pre-trained three-dimensional point cloud generation model is generated based on sample description text and sample three-dimensional point cloud training.
2. The three-dimensional point cloud generation method according to claim 1, characterized in that: The determining of the three-dimensional space coordinates corresponding to each feature point in the initial three-dimensional point cloud includes: Determining the initial three-dimensional spatial coordinates of each feature point in the initial three-dimensional point cloud; The coordinate value of the initial three-dimensional space coordinate of each feature point is multiplied by two to obtain the three-dimensional space coordinate corresponding to each feature point in the initial three-dimensional point cloud.
3. The three-dimensional point cloud generation method according to claim 1, characterized in that: Generating the spatial feature corresponding to each feature point according to the three-dimensional spatial coordinates corresponding to each feature point includes: Obtaining a resolution level corresponding to the initial three-dimensional point cloud; Obtaining the octree position number and octree code of the three-dimensional spatial coordinates corresponding to each feature point in the octree, wherein the three-dimensional spatial coordinates corresponding to each feature point are obtained by performing octree segmentation based on the three-dimensional spatial coordinates of the corresponding parent feature point in the historical three-dimensional point cloud at the previous resolution level; Generate a spatial feature corresponding to each feature point according to the resolution level, the octree position number and the octree code.
4. The three-dimensional point cloud generation method according to claim 1, characterized in that: After inputting the text features and the spatial features corresponding to each feature point into the pre-trained three-dimensional point cloud generation model and outputting the target three-dimensional point cloud, the method further includes: Determining whether the resolution level corresponding to the target three-dimensional point cloud reaches a preset resolution level; When the resolution level corresponding to the target three-dimensional point cloud does not reach the preset resolution level, the target three-dimensional point cloud is determined as the initial three-dimensional point cloud, and the step of determining the three-dimensional space coordinates corresponding to each feature point in the initial three-dimensional point cloud is returned to execute until the target three-dimensional point cloud output by the pre-trained three-dimensional point cloud generation model reaches the preset resolution level.
5. The three-dimensional point cloud generation method according to claim 1, characterized in that: The pre-trained 3D point cloud generation model includes a first pre-trained 3D point cloud generation sub-model and a second pre-trained 3D point cloud generation sub-model; The step of inputting the text feature and the spatial feature corresponding to each feature point into a pre-trained three-dimensional point cloud generation model and outputting a target three-dimensional point cloud comprises: Inputting the text feature and the spatial feature corresponding to each feature point into a first pre-trained three-dimensional point cloud generation sub-model, and outputting a first three-dimensional point cloud feature; Performing upsampling processing on the first three-dimensional point cloud features to obtain second three-dimensional point cloud features; Inputting the second three-dimensional point cloud feature and the text feature into a second pre-trained three-dimensional point cloud generation sub-model, and outputting a third three-dimensional point cloud feature; The third three-dimensional point cloud feature is subjected to convolution processing to obtain a target three-dimensional point cloud.
6. The three-dimensional point cloud generation method according to claim 5, characterized in that: The first pre-trained 3D point cloud generation sub-model includes a plurality of feature extraction modules; The step of inputting the text feature and the spatial feature corresponding to each feature point into a first pre-trained three-dimensional point cloud generation sub-model and outputting a first three-dimensional point cloud feature comprises: Inputting the text feature and the spatial feature corresponding to each feature point into a first feature extraction module of the first pre-trained 3D point cloud generation sub-model, and outputting the point cloud extraction feature; Inputting the point cloud extraction features into a next feature extraction module of the first feature extraction module, and outputting updated point cloud extraction features; Repeatedly input the point cloud extraction features and the text features output by the previous feature extraction module into the next feature extraction module, and output the updated point cloud extraction features, until the first three-dimensional point cloud features output by the last feature extraction module of the first pre-trained three-dimensional point cloud generation sub-model are obtained.
7. The three-dimensional point cloud generation method according to claim 6, characterized in that: The feature extraction module includes a convolution submodule, a cross attention submodule, a first fully connected submodule, a self-attention submodule and a second fully connected submodule connected in sequence; The step of inputting the text feature and the spatial feature corresponding to each feature point into a first feature extraction module of the first pre-trained 3D point cloud generation sub-model and outputting the point cloud extraction feature comprises: Input the spatial features corresponding to each feature point into the convolution submodule, and output the convolution features; Adding the convolution feature and the spatial feature corresponding to each feature point to obtain a first input feature, and inputting the first input feature and the text feature into a cross attention submodule to output a first attention feature; Adding the first attention feature and the first input feature to obtain a second input feature, and inputting the second input feature into a first fully connected submodule to output a first fully connected feature; Adding the first fully connected feature and the second input feature to obtain a third input feature, and inputting the third input feature into the self-attention submodule to output a second attention feature; Adding the second attention feature and the third input feature to obtain a fourth input feature, inputting the fourth input feature into a second fully connected submodule, and outputting a second fully connected feature; The second fully connected feature and the fourth input feature are added to obtain a point cloud extraction feature.
8. The three-dimensional point cloud generation method according to claim 5, characterized in that: The second pre-trained 3D point cloud generation sub-model includes a plurality of feature extraction modules; The step of inputting the second three-dimensional point cloud feature and the text feature into a second pre-trained three-dimensional point cloud generation sub-model and outputting a third three-dimensional point cloud feature comprises: Inputting the text features and the second three-dimensional point cloud features into a first feature extraction module of a second pre-trained three-dimensional point cloud generation sub-model, and outputting target point cloud extraction features; Inputting the target point cloud extraction features into a next feature extraction module of the first feature extraction module, and outputting updated target point cloud extraction features; Repeatedly input the target point cloud extraction features and the text features output by the previous feature extraction module into the next feature extraction module, and output the updated target point cloud extraction features, until the third three-dimensional point cloud features output by the last feature extraction module of the second pre-trained three-dimensional point cloud generation sub-model are obtained.
9. The three-dimensional point cloud generation method according to claim 5, characterized in that: The pre-trained 3D point cloud generation model includes a plurality of sequentially connected convolution sub-models; The step of performing convolution processing on the third three-dimensional point cloud feature to obtain a target three-dimensional point cloud comprises: Inputting the third three-dimensional point cloud feature into the first convolution sub-model, and outputting the intermediate convolution feature; Inputting the intermediate convolution features output by the first convolution sub-model into the next convolution sub-model, and outputting updated intermediate convolution features; Repeatedly input the intermediate convolution features output by the previous convolution sub-model into the next convolution sub-model, and output the updated intermediate convolution features until the last convolution sub-model outputs the target 3D point cloud.
10. The three-dimensional point cloud generation method according to claim 1, characterized in that: Before inputting the text features and the spatial features corresponding to each feature point into the pre-trained three-dimensional point cloud generation model and outputting the target three-dimensional point cloud, the method further includes: Obtaining a sample three-dimensional point cloud and a label three-dimensional point cloud corresponding to the sample three-dimensional point cloud; Determine the sample three-dimensional space coordinates corresponding to each sample feature point in the sample three-dimensional point cloud; Generate a sample space feature corresponding to each sample feature point according to the sample three-dimensional space coordinates corresponding to each sample feature point; Obtaining a sample description text corresponding to the labeled three-dimensional point cloud, and determining a sample text feature corresponding to the sample description text; Inputting the sample text features and the sample space features corresponding to each sample feature point into a three-dimensional point cloud generation model, and outputting a predicted three-dimensional point cloud; Determine the difference between the predicted three-dimensional point cloud and the labeled three-dimensional point cloud, and when the difference does not meet the preset iteration condition, iterate the three-dimensional point cloud generation model according to the difference, and return to the step of determining the sample three-dimensional space coordinates corresponding to each sample feature point in the sample three-dimensional point cloud until the difference meets the preset iteration condition, thereby obtaining a pre-trained three-dimensional point cloud model; The number of feature points of the predicted three-dimensional point cloud and the number of feature points of the label three-dimensional point cloud are both greater than the number of feature points of the sample three-dimensional point cloud.
11. A three-dimensional point cloud generation device, characterized in that: include: A first acquisition module is used to acquire an initial three-dimensional point cloud and determine the three-dimensional space coordinates corresponding to each feature point in the initial three-dimensional point cloud; A generating module, used for generating a spatial feature corresponding to each feature point according to the three-dimensional spatial coordinates corresponding to each feature point; A second acquisition module is used to acquire a point cloud generation target description text corresponding to the initial three-dimensional point cloud, and determine text features corresponding to the point cloud generation target description text; An input module, used for inputting the text features and the spatial features corresponding to each feature point into a pre-trained three-dimensional point cloud generation model, and outputting a target three-dimensional point cloud; The number of feature points in the target three-dimensional point cloud is greater than the number of feature points in the initial three-dimensional point cloud, and the pre-trained three-dimensional point cloud generation model is generated based on sample description text and sample three-dimensional point cloud training.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the three-dimensional point cloud generation method according to any one of claims 1 to 10.
13. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the three-dimensional point cloud generation method according to any one of claims 1 to 10 is implemented.