Three-dimensional model processing method and device, electronic equipment and storage medium
Through the three-dimensional model generation network combined with model coding and semantic feature extraction, the problems of high technical threshold and low efficiency of building three-dimensional models with rich details are solved, and the rapid and efficient generation of three-dimensional models with rich details are achieved.
Patent Information
- Application Number
- CN202510210195.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, the technical threshold for building three-dimensional models with rich details is high, and it takes a long time and is less efficient.
By obtaining the three-dimensional model to be optimized and the detailed description text, model coding processing and semantic feature extraction are performed, and a three-dimensional model generation network is used to generate a target three-dimensional model with rich details based on the text semantic features and model coding features.
It lowers the technical threshold for building three-dimensional models with rich details, improves construction efficiency, and allows non-professional modeling personnel to automatically generate three-dimensional models with rich details.
Smart Images

Figure CN120147522A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and more particularly, to a method, apparatus, electronic device, and storage medium for processing three-dimensional models. Background Art
[0002] With the development of computer technology, the application of three-dimensional modeling has become increasingly widespread. In related technologies, in order to construct a three-dimensional model with rich details, users need to have relatively high three-dimensional modeling capabilities to construct a three-dimensional model with rich details. The technical threshold is relatively high. Moreover, the time taken to construct a three-dimensional model with rich details is relatively long and the efficiency is relatively low. Summary of the Invention
[0003] In view of the above problems, embodiments of the present application propose a method, apparatus, electronic device, and storage medium for processing three-dimensional models to solve the problem of relatively high technical threshold for constructing a three-dimensional model with rich details in related technologies.
[0004] According to one aspect of the embodiments of the present application, there is provided a method for processing a three-dimensional model, including: obtaining a three-dimensional model to be optimized and a detail description text; performing encoding processing on the three-dimensional model to be optimized to obtain a model encoding feature; performing semantic feature extraction on the detail description text to obtain a text semantic feature; and generating, by a three-dimensional model generation network, a target three-dimensional model based on the text semantic feature and the model encoding feature, where the basic shape of the target three-dimensional model conforms to that of the three-dimensional model to be optimized, and the details conform to the detail description text.
[0005] According to one aspect of the embodiments of the present application, there is provided a device for processing a three-dimensional model, including: an obtaining module, configured to obtain a three-dimensional model to be optimized and a detail description text; an encoding module, configured to perform encoding processing on the three-dimensional model to be optimized to obtain a model encoding feature; a semantic feature extraction module, configured to perform semantic feature extraction on the detail description text to obtain a text semantic feature; and a three-dimensional model generation module, configured to generate, by a three-dimensional model generation network, a target three-dimensional model based on the text semantic feature and the model encoding feature, where the basic shape of the target three-dimensional model conforms to that of the three-dimensional model to be optimized, and the details conform to the detail description text.
[0006] In some embodiments, the three-dimensional model generation network includes a diffusion network and a model decoding network; the three-dimensional model generation module includes: a diffusion denoising processing unit, configured to use the diffusion network to perform T rounds of diffusion denoising processing on the model encoding features with the text semantic features as diffusion conditions, to obtain the target model features output by the T-th round of diffusion processing; wherein, in the T rounds of diffusion processing, the model features output by the i-th round of diffusion processing are used as the input of the (i + 1)-th round of diffusion processing, i and T are positive integers, and i + 1 ≤ T; a model decoding unit, configured to perform model decoding processing on the target model features by the model decoding network, to obtain a target three-dimensional model whose basic shape conforms to the three-dimensional model to be optimized and whose details conform to the detailed description text.
[0007] In some embodiments, the processing device for the three-dimensional model further includes: a first acquisition module, configured to acquire the sample encoding features of the first three-dimensional model; a second acquisition module, configured to acquire the sample text semantic features of the sample detailed description text of the first three-dimensional model; a first diffusion and noise addition module, configured to perform N rounds of diffusion and noise addition processing on the sample encoding features by the diffusion network, to obtain a diffusion and noise addition result; the diffusion and noise addition result includes the target noise-added model features obtained by the N-th round of diffusion and noise addition processing and the noise added in each round of the diffusion and noise addition processing; N is a positive integer not exceeding T; a first diffusion denoising module, configured to use the sample text semantic features as diffusion conditions by the diffusion network to perform N rounds of diffusion denoising processing on the target noise-added model features, to obtain a diffusion denoising result; the diffusion denoising result includes the noise removed in each round of the N rounds of diffusion denoising processing; a diffusion loss determination module, configured to determine a diffusion loss according to the noise added in each round of the diffusion and noise addition processing and the noise removed in each round of the diffusion denoising processing; a first adjustment module, configured to at least adjust the parameters of the diffusion network according to the diffusion loss until a first training end condition is reached.
[0008] In some embodiments, the diffusion denoising result further includes the target denoised model features output by the N-th round of diffusion denoising processing; the processing device for the three-dimensional model further includes: a first decoding module, configured to perform model decoding processing on the target denoised model features by the model decoding network, to obtain a second three-dimensional model; a model reconstruction loss determination module, configured to calculate a model reconstruction loss according to the second three-dimensional model and the first three-dimensional model; correspondingly, the first adjustment module is configured to: perform weighted processing on the diffusion loss and the model reconstruction loss to obtain a target loss; at least adjust the parameters of the diffusion network and the model decoding network according to the target loss until a first training end condition is reached.
[0009] In some embodiments, the processing device for the three-dimensional model further includes: an image acquisition module, configured to acquire two-dimensional model images of the first three-dimensional model from at least one perspective; a text generation module, configured to generate text descriptions for the two-dimensional model images from each of the perspectives to obtain image description texts corresponding to the two-dimensional model images; a first combination module, configured to combine the image description texts corresponding to the two-dimensional model images from the at least one perspective to obtain a sample detailed description text of the first three-dimensional model.
[0010] In some embodiments, the processing device for the three-dimensional model further includes: a collection module, configured to collect three-dimensional surface point cloud data of the first three-dimensional model; a first encoding module, configured to perform encoding processing on the three-dimensional surface point cloud data by a model encoding network to obtain a sample encoding feature of the first three-dimensional model.
[0011] In some embodiments, the processing device for the three-dimensional model further includes: a sampling module, configured to sample the three-dimensional surface point cloud data to obtain M sets of sub-surface point cloud data; M is a positive integer greater than 1; a second encoding module, configured to perform encoding processing on each set of sub-surface point cloud data in the M sets of sub-surface point cloud data by the model encoding network to obtain encoding features of each set of sub-surface point cloud data; a second combination module, configured to combine the encoding features of the M sets of sub-surface point cloud data to obtain a sample encoding feature of the first three-dimensional model.
[0012] In some embodiments, the processing device for the three-dimensional model further includes: a sample acquisition module, configured to acquire a plurality of target samples; the target samples include target point cloud data collected for a reference three-dimensional model and position labels of each point cloud in the target point cloud data; the target point cloud data includes internal point clouds, surface point clouds, and external point clouds located outside the reference three-dimensional model of the reference three-dimensional model; the position labels are used to indicate the relative position relationship of the corresponding point clouds with respect to the reference three-dimensional model; a third encoding module, configured to perform feature encoding on each point cloud in the target point cloud data by the model encoding network to obtain encoding features of each point cloud in the target point cloud data; a position decoding module, configured to perform point cloud position decoding according to the encoding features of each point cloud in the target point cloud data to obtain predicted position labels of each point cloud in the target point cloud data; a prediction loss calculation module, configured to calculate a prediction loss according to the position labels and the corresponding predicted position labels of each point cloud in the target point cloud data; a second adjustment module, configured to adjust parameters of the model encoding network according to the prediction loss until a second training end condition is reached.
[0013] In some embodiments, the processing device for the three-dimensional model further includes: a first acquisition module, configured to acquire internal point clouds located inside the reference three-dimensional model to obtain a first point cloud set; a second acquisition module, configured to acquire external point clouds located outside the reference three-dimensional model to obtain a second point cloud set; a third acquisition module, configured to acquire surface point clouds located on the surface of the reference three-dimensional model to obtain a third point cloud set; and a third combination module, configured to combine the first point cloud set, the second point cloud set, and the third point cloud set to obtain the target point cloud data.
[0014] In some embodiments, the processing device for the three-dimensional model further includes: a curvature information acquisition module, configured to acquire curvature information of each surface of the reference three-dimensional model; a sampling density determination module, configured to determine, according to the curvature information, the point cloud acquisition density corresponding to each surface, where the point cloud acquisition density refers to the number of point clouds acquired per unit area; the point cloud acquisition density is positively correlated with the curvature; and a third acquisition module, configured to: acquire surface point clouds located on the surface of the reference three-dimensional model according to the point cloud acquisition density corresponding to each surface to obtain a third point cloud set.
[0015] In some embodiments, the encoding module includes: a first acquisition unit, configured to acquire three-dimensional surface point cloud data of the three-dimensional model to be optimized; and an encoding processing unit, configured to perform encoding processing on the three-dimensional surface point cloud data of the three-dimensional model to be optimized by a model encoding network to obtain the model encoding feature.
[0016] According to one aspect of the embodiments of the present application, an electronic device is provided, including: a processor; and a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the above-described method for processing a three-dimensional model is implemented.
[0017] According to one aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the above-described method for processing a three-dimensional model is implemented.
[0018] According to one aspect of the embodiments of the present application, a computer program product is provided, including computer instructions, and when the computer instructions are executed by a processor, the above-described method for processing a three-dimensional model is implemented.
[0019] In the present application, by describing the text semantic features of the text in detail, a three-dimensional model generation network is guided to generate a target three-dimensional model based on the model coding features of the three-dimensional model to be optimized, where the basic shape conforms to the three-dimensional model to be optimized and the details conform to the detailed description text. Thus, the refinement processing of the three-dimensional model to be optimized using the detailed description text is realized, a refined target three-dimensional model with richer details is generated, the technical threshold for constructing a three-dimensional model with rich details is reduced, and the efficiency of constructing a three-dimensional model with rich details is improved. In this way, even non-professional modeling personnel can automatically generate the desired three-dimensional model with rich details according to the method of the present application. Brief Description of the Drawings
[0020] The drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0021] Figure 1A It is a schematic diagram of the application scenario of the present application shown according to an embodiment of the present application.
[0022] Figure 1B It is a schematic diagram of generating a target three-dimensional model shown according to an embodiment of the present application.
[0023] Figure 2 It is a flowchart of a method for processing a three-dimensional model shown according to an embodiment of the present application.
[0024] Figure 3 It is a flowchart of a method for processing a three-dimensional model shown according to an embodiment of the present application.
[0025] Figure 4 It is a flowchart of training a three-dimensional model generation network shown according to an embodiment of the present application.
[0026] Figure 5 It is a flowchart of training a three-dimensional model generation network shown according to another embodiment of the present application.
[0027] Figure 6 It is a schematic diagram of training a three-dimensional model generation network shown according to an embodiment of the present application.
[0028] Figure 7 It is a flowchart of training a model coding network shown according to an embodiment of the present application.
[0029] Figure 8 Exemplarily shows a schematic diagram of processing using a model coding network and a position label decoding network.
[0030] Figure 9 is a flowchart of a method for processing a three-dimensional model shown in another embodiment of the present application.
[0031] Figure 10 is a flowchart of a training model encoding network and a three-dimensional model generation network shown in an embodiment of the present application.
[0032] Figure 11 is a flowchart of a test model encoding network and a three-dimensional model generation network shown in an embodiment of the present application.
[0033] Figure 12 is a block diagram of a three-dimensional model processing device shown in an embodiment of the present application.
[0034] Figure 13 shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners
[0035] The following details the implementation manners of the present application. The examples of the implementation manners are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The implementation manners described below by referring to the accompanying drawings are exemplary only for explaining the present application and should not be construed as a limitation to the present application.
[0036] To enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0037] In the following description, the terms "first\second" etc. involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first\second" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0038] As used herein, "a plurality of" means two or more. "And / or" describes the relationship between associated objects and indicates that there can be three relationships. For example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. In the following description, when referring to "some embodiments or some implementation manners", it describes a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0039] Figure 1A is a schematic diagram of the application scenario of the present application shown according to an embodiment of the present application. As Figure 1A shown, this application scenario includes a terminal 110 and a server 120 communicatively connected to the terminal 110. The terminal 110 can send a three-dimensional model to be optimized and a detailed description text to the server 120. After that, the server 120 generates a three-dimensional model according to the method of the present application. The specific process is as follows: performing encoding processing on the three-dimensional model to be optimized to obtain a model encoding feature; extracting semantic features from the detailed description text to obtain a text semantic feature; and generating a network from the three-dimensional model to generate a target three-dimensional model according to the text semantic feature and the model encoding feature, so that the basic shape conforms to the three-dimensional model to be optimized, and the details conform to the target three-dimensional model of the detailed description text.
[0040] Figure 1B is a schematic diagram of generating a target three-dimensional model shown according to an embodiment of the present application. As Figure 1B shown, if the three-dimensional model to be optimized is a three-dimensional model of a sofa with a relatively simple shape, and the detailed description text is "a single-seat sofa with an arc-shaped handle", after obtaining the model encoding feature of the three-dimensional model to be optimized and the text semantic feature of the detailed description text, input the model encoding feature and the text semantic feature into the three-dimensional model generation network. After the three-dimensional model generation network processes them, it outputs a target three-dimensional model. From Figure 1B it can be seen that the output target three-dimensional model is a sofa with an arc-shaped handle and is a single-seat.
[0041] After that, the server 120 can return the target three-dimensional model to the terminal 110, and the target three-dimensional model can be displayed on the display interface of the terminal 110. The terminal 110 can be a desktop computer, a laptop computer, a smart phone, a vehicle-mounted terminal, a smart TV, a mixed reality device, etc., and is not specifically limited herein.
[0042] The server 120 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0043] In addition, the method of this application can also be executed by the terminal 110. The three-dimensional model generation network is deployed on the terminal 110, and the terminal 110 generates a three-dimensional model according to the method of this application.
[0044] The implementation details of the technical solution of the embodiments of this application are elaborated in detail below:
[0045] Figure 2 is a flowchart of a method for processing a three-dimensional model shown according to an embodiment of this application. This method can be executed by an electronic device. The electronic device can be a terminal, a server, etc., or implemented by the interaction between a server and a terminal. Refer to Figure 2 As shown, this method at least includes steps 210 to 240, which are introduced in detail as follows:
[0046] Step 210, obtain a three-dimensional model to be optimized and a detail description text.
[0047] The three-dimensional model to be optimized refers to a three-dimensional model to be refined. In some embodiments, the three-dimensional model to be optimized can be a three-dimensional model constructed by a user using three-dimensional geometric bodies provided in a modeling application. The three-dimensional geometric bodies are, for example, cylinders, cuboids, cubes, triangular prisms, triangular pyramids, spheres, cones, etc., which are not specifically limited herein. In other embodiments, the three-dimensional model to be optimized can be a three-dimensional model selected by a user from a three-dimensional model library, and the three-dimensional model library includes multiple basic three-dimensional models for the user to select.
[0048] The detail description text refers to a description text used to describe the model details of the to-be-generated three-dimensional refined model; this detail description text is used to indicate the user's requirements for the to-be-generated three-dimensional refined model, that is, to indicate the detail content that the to-be-generated three-dimensional refined model needs to present. It can also be understood as a description text used to indicate the detail optimization of the three-dimensional model to be optimized. The detail description text can be input by the user according to needs.
[0049] Step 220, perform an encoding process on the three-dimensional model to be optimized to obtain a model encoding feature.
[0050] In some embodiments, the three-dimensional model to be optimized can be described by the three-dimensional model data of the three-dimensional model to be optimized. The three-dimensional model data is, for example, the mesh model data of the three-dimensional model to be optimized. Therefore, the model encoding network can be used to encode the three-dimensional model data of the three-dimensional model to be optimized to obtain model encoding features, which are the feature representations of the three-dimensional model data of the three-dimensional model to be optimized in the latent feature space. In this embodiment, the model encoding network is a neural network model for encoding three-dimensional model data. The model encoding network can be a three-dimensional convolutional neural network, a mesh convolutional neural network (Mesh CNN), etc., which will not be specifically limited herein.
[0051] In some embodiments, considering that the three-dimensional model can also be described by the point cloud collected for the three-dimensional model, step 220 includes: collecting the three-dimensional surface point cloud data of the three-dimensional model to be optimized; and encoding the three-dimensional surface point cloud data of the three-dimensional model to be optimized by the model encoding network to obtain model encoding features.
[0052] In this embodiment, the model encoding network is a neural network model for encoding three-dimensional point cloud data. For example, it can be the encoding network in PointNet (point cloud processing network), Variational Auto-encoder (VAE), etc. The model encoding network can encode each point cloud (or called sampling point), and then fuse the features of all point clouds to obtain the global point cloud features representing the whole.
[0053] The point clouds in the three-dimensional surface point cloud data of the three-dimensional model to be optimized are the three-dimensional coordinate information of multiple surface point clouds collected on the surface of the three-dimensional model to be optimized. The surface point cloud refers to the point cloud representing the surface of the three-dimensional model. The three-dimensional surface point cloud data of the three-dimensional model to be optimized includes the surface point clouds respectively collected on each surface in the three-dimensional model to be optimized. Therefore, the three-dimensional surface point cloud data of the three-dimensional model to be optimized can accurately describe the shape of the three-dimensional model to be optimized.
[0054] In some embodiments, the three-dimensional surface point cloud data of the three-dimensional model to be optimized can be input into the model encoding network as a whole, and processed by the model encoding network to output model encoding features.
[0055] In some embodiments, considering that the number of surface point clouds in the three-dimensional surface point cloud data of the three-dimensional model to be optimized is large, inputting the entire three-dimensional surface point cloud data of the three-dimensional model to be optimized into the model encoding network for encoding processing involves a large amount of computation and more complex processing, which may lead to a longer time consumption. Therefore, the three-dimensional surface point cloud data of the three-dimensional model to be optimized can be encoded in batches. Alternatively, in some cases, the model encoding network has a length requirement for the input. If the number of surface point clouds in the three-dimensional surface point cloud data is large, it may cause the input to exceed the maximum input length of the model encoding network. Therefore, the three-dimensional surface point cloud data of the three-dimensional model to be optimized can be encoded in batches. In this case, step 220 may include the following steps A1 to A3:
[0056] Step A1: Sample the three-dimensional surface point cloud data of the three-dimensional model to be optimized to obtain Q sets of reference surface point cloud data; Q is a positive integer greater than 1.
[0057] Among them, each set of reference surface point cloud data includes multiple surface point clouds located on each surface of the three-dimensional model to be optimized. In this way, each set of reference surface point cloud data can be separately used to describe the shape of the three-dimensional model to be optimized. There are different surface point clouds in different sets of reference surface point cloud data. Of course, there may also be some identical surface point clouds.
[0058] In some embodiments, the three-dimensional surface point cloud data of the three-dimensional model to be optimized can be randomly sampled to avoid a high proportion of identical surface point clouds in different sets of reference surface point cloud data, resulting in excessive repeated processing by the model encoding network.
[0059] In some embodiments, Q can be set according to actual needs, and Q is at least 2. In some embodiments, to make full use of the information of all surface point clouds in the three-dimensional surface point cloud data of the three-dimensional model to be optimized, it can be set that the total number of surface point clouds in the Q sets of reference surface point cloud data is not less than the total number of surface point clouds in the three-dimensional surface point cloud data of the three-dimensional model to be optimized.
[0060] In some embodiments, the number of surface point clouds in each set of reference surface point cloud data (for the convenience of description, it is called the target number) can be preset in advance. In this case, Q can be determined according to the total number of surface point clouds (assumed to be the first number) in the three-dimensional surface point cloud data of the three-dimensional model to be optimized and the target number. For example, the integer part of the quotient of the first number and the target number (it can be rounded up or rounded down) can be used as Q; or, the integer part of the quotient of the first number and the target number plus a preset surplus amount can be used as Q, and the surplus amount is a positive integer.
[0061] After that, sampling is performed on the three-dimensional surface point cloud data of the three-dimensional model to be optimized according to the target quantity. On the one hand, it is ensured that the total number of surface point clouds in each group of reference surface point cloud data is the target quantity. On the other hand, it is ensured that each group of reference surface point cloud data covers the surface point clouds located on all surfaces of the three-dimensional model to be optimized, and it is ensured that the proportion of the same surface point clouds in different groups of reference surface point cloud data is lower than the proportion threshold.
[0062] Step A2: The model encoding network encodes each group of reference surface point cloud data in the Q groups of reference surface point cloud data respectively to obtain the encoding features of each group of reference surface point cloud data.
[0063] Each group of reference surface point cloud data is input into the model encoding network respectively, and the model encoding network processes it to output the encoding features of each group of reference surface point cloud data. It can be understood that if the total number of surface point clouds in a group of reference surface point cloud data is t, then the encoding features of the reference surface point cloud data are t in the dimension representing the number of point clouds. For example, if the model encoding network can map each point cloud to an implicit encoding with a length of 1024, then the dimension of the encoding features of a group of reference surface point cloud data is t×1024.
[0064] Step A3: Combine the encoding features of the Q groups of reference surface point cloud data to obtain the sample encoding features of the three-dimensional model to be optimized.
[0065] For example, if the dimension of the encoding features of a group of reference surface point cloud data is t×1024, the encoding features of the Q groups of reference surface point cloud data can be combined, and the dimension of the sample encoding features of the three-dimensional model to be optimized is Q×t×1024, where Q, t, and 1024 are the lengths in different dimensions respectively.
[0066] Step 230: Extract semantic features from the detailed description text to obtain text semantic features.
[0067] The text encoder can be used to extract semantic features from the detailed description text to obtain text semantic features.
[0068] Step 240: The three-dimensional model generation network generates a target three-dimensional model based on the text semantic features and the model encoding features, where the basic shape is consistent with the three-dimensional model to be optimized and the details conform to the detailed description text.
[0069] In other words, the target 3D model is a 3D model obtained by refining the details of the 3D model to be optimized according to the detail description text. Compared with the 3D model to be optimized, the target 3D model is a more refined 3D model. On the contrary, compared with the target 3D model, the 3D model to be optimized is a coarser 3D model. The model details of the target 3D model are richer than those of the 3D model to be optimized. Therefore, the solution of this application realizes the refinement of the 3D model to be optimized through the detail description text, and obtains a more refined target 3D model whose basic shape conforms to the 3D model to be optimized and whose details conform to the detail description text, thus realizing the refinement process of the 3D model.
[0070] The 3D model generation network is a neural network model for generating 3D models. In some embodiments, the 3D model generation network includes a diffusion network and a model decoding network; step 240 includes the following steps B1 to B2:
[0071] Step B1: The diffusion network uses the text semantic feature as the diffusion condition to perform T rounds of diffusion denoising processing on the model encoding feature, and obtains the target model feature output by the T - th round of diffusion processing; where, in the T rounds of diffusion processing, the model feature output by the i - th round of diffusion processing is used as the input of the (i + 1) - th round of diffusion processing, i and T are positive integers, and i + 1 ≤ T.
[0072] The diffusion network is used to perform diffusion denoising processing on the model encoding feature, the model decoding network is used to perform model decoding on the target model feature output by the diffusion network, and output a 3D model. The text semantic feature is used to guide the diffusion network to perform the diffusion denoising process on the model encoding feature, so as to inject the text semantic feature into the diffusion denoising process, so that the finally output target model feature is the detail information expressed by the injected text semantic feature.
[0073] In the process of each round of diffusion processing (taking the k - th round of diffusion denoising processing as an example, k is a positive integer greater than 1, k ≤ T), the diffusion network performs noise prediction based on the text semantic feature and the input model feature of the k - th round of diffusion denoising processing, and obtains the predicted noise of the k - th round of diffusion processing; then, the predicted noise of the k - th round of diffusion denoising processing is removed from the input model feature of the k - th round of diffusion denoising processing, and the model feature output by the k - th round of diffusion denoising processing is obtained. (Where, when k is greater than 1, the input model feature of the k - th round of diffusion denoising processing is the model feature output by the previous round of diffusion denoising processing, and when k = 1, the input model feature of the first round of diffusion denoising processing is the model encoding feature). T can be set according to actual needs, for example, T is 50, 100, 500, 900, 1000, etc.
[0074] In some embodiments, for the model encoding features, T rounds of diffusion denoising may be performed. First, the noise for each round of diffusion denoising process may be predicted, and then the noise predicted for the T rounds of diffusion denoising process may be denoised all at once to obtain the target model features.
[0075] Step B2: The model decoding network performs model decoding processing on the target model features to obtain a target 3D model whose basic shape conforms to the 3D model to be optimized and whose details conform to the detail description text.
[0076] The model decoding network is a neural network model used to decode model features into the 3D model space to output a 3D model. For example, the model decoding network may be a decoder in a 3D shape variational autoencoder, a U-shaped convolutional network, etc., which is not specifically limited herein.
[0077] Figure 3 It is a flowchart of a method for processing a 3D model according to an embodiment of the present application. As Figure 3 shown, after extracting the model encoding features of the 3D model to be optimized and the text semantic features of the detail description text, the model encoding features and the text semantic features are input into the diffusion network. Among them, the diffusion network uses the text semantic features as diffusion conditions to perform multiple rounds of diffusion denoising processing on the model encoding features. Figure 3 FIG. shows a schematic diagram of the first round of diffusion denoising processing. In the first round of diffusion denoising processing, the diffusion network may perform attention fusion on the text semantic features and the model encoding features. For example, according to the text semantic features and the model encoding features, a query vector (Query, Q), a key vector, and a value vector (Value, V) may be determined, and then the dot product attention may be calculated to predict the noise for the first round of diffusion denoising processing, and accordingly, the model encoding features may be denoised according to the predicted noise. Then, subsequent diffusion denoising processes are performed in a similar manner. After obtaining the target model features output by the T-th round of diffusion denoising processing, the model decoding network decodes the target model features to output the target 3D model.
[0078] In the present application, through the text semantic features of the detail description text, the 3D model generation network is guided to generate a target 3D model whose basic shape conforms to the 3D model to be optimized and whose details conform to the detail description text according to the model encoding features of the 3D model to be optimized. Thereby, the refinement processing of the 3D model to be optimized using the detail description text is realized, and a refined target 3D model with richer details is generated. For users, the technical threshold for constructing a 3D model with rich details is reduced, and the efficiency of constructing a 3D model with rich details is improved. In this way, even non-professional modelers can automatically generate a 3D model with rich details desired by users according to the method of the present application.
[0079] To ensure the accuracy of the 3D model generated by the 3D model generation network, it is necessary to pre-train the 3D model generation network. In some embodiments, as Figure 4 shown, before step 240, the method further includes:
[0080] Step 410, obtaining the sample encoding feature of the first 3D model.
[0081] The sample encoding feature refers to the model encoding feature of the first 3D model, which is the vectorized representation of the first 3D model in the feature space. For ease of distinction, the 3D model used to train the 3D model generation network is referred to as the first 3D model, and the first 3D models involved in different first training samples may be different. In some embodiments, the first 3D model may be a 3D model with relatively rich details to fully train the 3D model generation network.
[0082] In some embodiments, the sample encoding feature of the first 3D model may be generated according to the following process 1)-2): 1) Collecting the 3D surface point cloud data of the first 3D model.
[0083] 2) The model encoding network performs encoding processing according to the 3D surface point cloud data to obtain the sample encoding feature of the first 3D model.
[0084] In some embodiments, the entire 3D surface point cloud data of the first 3D model may be input into the model encoding network, which is processed by the model encoding network and outputs the sample encoding feature of the first 3D model.
[0085] In other embodiments, the step of the model encoding network performing encoding processing according to the 3D surface point cloud data to obtain the sample encoding feature of the first 3D model may include: sampling the 3D surface point cloud data to obtain M groups of sub-surface point cloud data; M is a positive integer greater than 1; the model encoding network respectively performs encoding processing on each group of sub-surface point cloud data in the M groups of sub-surface point cloud data to obtain the encoding features of each group of sub-surface point cloud data; combining the encoding features of the M groups of sub-surface point cloud data to obtain the sample encoding feature of the first 3D model.
[0086] M may be the same as or different from Q in the above text, and can be determined according to the actual situation. The method for determining the value of M may be similar to the method for determining the value of Q in the above text, which will not be elaborated here. The implementation details of determining the sample encoding feature of the first 3D model as above are similar to the implementation details of steps A1-A3 in the above text, and specific details can be referred to the above description, which will not be elaborated here.
[0087] Step 420, obtaining the sample text semantic feature of the sample detailed description text of the first 3D model.
[0088] The sample detailed description text of the first 3D model refers to the description text that describes the model details of the first 3D model. The sample detailed description text of the first 3D model can describe one or more details of the first 3D model. In some embodiments, the sample detailed description text of the first 3D model can be written manually based on the first 3D model. In some embodiments, a neural network model with the ability to generate the description text of a 3D model can be used. The first 3D model is input into the model, and the model performs feature encoding on the first 3D model, and then outputs the sample detailed description text of the first 3D model.
[0089] In some other embodiments, the sample detailed description text of the first 3D model can be generated according to the following process ①-③: ① Obtain the 2D model images of the first 3D model from at least one perspective.
[0090] Multiple perspectives can be selected to face the first 3D model for image acquisition to obtain 2D model images from multiple perspectives. Or the perspective that can best reflect the appearance characteristics of the first 3D model can be selected, and then the first 3D model is imaged from this perspective. The 2D model image from each perspective can present the external characteristics of the first 3D model from the corresponding perspective.
[0091] In some embodiments, the first 3D model can be imaged from 6 different perspectives to obtain 2D model images of the first 3D model from 6 perspectives, which are: front view, left view, right view, rear view, top view, and bottom view. Of course, one or more perspectives can also be selected from these 6 perspectives.
[0092] ② Generate text descriptions for the 2D model images from each perspective to obtain the image description text corresponding to each 2D model image.
[0093] Each 2D model image from each perspective can be input into the image-to-text model respectively, and the image-to-text model generates the image description text and outputs the image description text of the 2D model image. It can be understood that since the object presented in the 2D model image is the first 3D model, the object actually described by the image description text corresponding to the 2D model image is also the first 3D model. Therefore, the image description text corresponding to the 2D model image is equivalent to the description text of the first 3D model or the detailed description text of the first 3D model. The image-to-text model can be a neural network model such as GPT (Generative Pre-trained Transformer, the pre-trained model base for multi-round dialogue scenarios) that has the ability to generate text according to images. In addition, prompt text can also be input into the image-to-text model to instruct the image-to-text model to generate the description text of the input 2D model image.
[0094] ③ Combine the image description texts corresponding to the two-dimensional model images at at least one perspective to obtain the sample detailed description text of the first three-dimensional model.
[0095] If only the two-dimensional model images of one perspective of the first three-dimensional model are collected, the image description text corresponding to the two-dimensional model images of this perspective can be correspondingly used as the sample detailed description text of the first three-dimensional model; if multiple perspectives are involved, the image description texts corresponding to the two-dimensional model images of multiple perspectives can be combined, and the combined text obtained is used as the sample detailed description text of the first three-dimensional model. The number of perspectives can be set according to actual needs. The more the number of perspectives, the more and more comprehensive the external characteristics of the first three-dimensional model presented in the two-dimensional model images, which can ensure that the content expressed in the subsequent sample detailed description text of the first three-dimensional model is richer and more comprehensive.
[0096] Step 430: The diffusion network performs N rounds of diffusion and noise addition processing on the sample encoded features to obtain a diffusion and noise addition result; the diffusion and noise addition result includes the target noise-added model features obtained from the Nth round of diffusion and noise addition processing and the noise added in each round of diffusion and noise addition processing; N is a positive integer not exceeding T.
[0097] In each round of diffusion and noise addition processing, noise is gradually added to the sample encoded features. Similarly, the features output from the previous round of diffusion and noise addition processing are used as the input for the next round of diffusion and noise addition processing. In some embodiments, noise sampling can be performed in a Gaussian distribution to add noise that follows a Gaussian distribution in each round of diffusion and noise addition processing.
[0098] In some embodiments, to ensure the full training of the three-dimensional model generation network, different Ns can be set for different first training samples, that is, there are first training samples with different N values among multiple first training samples, and the N corresponding to different first training samples does not exceed T. The N set for the first training sample is used as the total number of rounds of diffusion and noise addition processing during training and as the total number of rounds of diffusion and noise removal processing during training.
[0099] Step 440: The diffusion network performs N rounds of diffusion and noise removal processing on the target noise-added model features with the sample text semantic features as the diffusion condition to obtain a diffusion and noise removal result; the diffusion and noise removal result includes the noise removed in each round of diffusion and noise removal processing during the N rounds of diffusion and noise removal processing.
[0100] Similarly, in the process of each round of diffusion denoising (taking the j-th round of diffusion denoising as an example, where j is a positive integer greater than 1 and k ≤ N), the diffusion network performs noise prediction based on the semantic features of the sample text and the input model features of the j-th round of diffusion denoising, and obtains the predicted noise of the j-th round of diffusion denoising. Then, the predicted noise of the j-th round of diffusion denoising is removed from the input model features of the j-th round of diffusion denoising to obtain the model features output by the j-th round of diffusion denoising. Among them, when j is greater than 1, the input model features of the j-th round of diffusion denoising are the model features output by the previous round of diffusion denoising. When j = 1, the input model features of the first round of diffusion denoising are the target noise-added model features.
[0101] Step 450: Determine the diffusion loss according to the noise added in each round of diffusion noise addition and the noise removed in each round of diffusion denoising.
[0102] The diffusion loss is used to reflect the difference between the noise added in the N rounds of diffusion noise addition and the noise removed in the N rounds of diffusion denoising. It can be understood that the closer the noise removed in the N rounds of diffusion denoising is to the noise added in the N rounds of diffusion denoising, the better the recovery effect of the diffusion network for adding noise to the sample coding features and then removing the noise, indicating that the diffusion denoising effect of the diffusion network is better.
[0103] In some embodiments, the mean square error between the noise added in each round of diffusion noise addition and the noise removed in each round of diffusion denoising can be calculated as the diffusion loss. In other embodiments, the negative log-likelihood loss between the noise added in each round of diffusion noise addition and the noise removed in each round of diffusion denoising can be calculated as the diffusion loss.
[0104] Step 460: Adjust at least the parameters of the diffusion network according to the diffusion loss until the first training end condition is reached.
[0105] The first training end condition can be that the diffusion loss converges, or the number of iterations of the diffusion network reaches the first number threshold, which is not specifically limited here. In some embodiments, the parameters of the diffusion network can be adjusted in a gradient descent manner according to the diffusion loss until the first training end condition is reached. In other embodiments, the model decoding network and the diffusion network can be trained together. Correspondingly, in step 460, the parameters of the diffusion network and the model encoding network can be adjusted in a gradient descent manner according to the diffusion loss until the first training end condition is reached.
[0106] After training the diffusion network through the above process, the diffusion network can accurately combine the text semantic features of the detailed description text, guiding the process of diffusion denoising of the model encoding features of the three-dimensional model, so that the model features output after diffusion denoising fully integrate the detailed content expressed by the text semantic features, facilitating the subsequent decoding of the three-dimensional model whose details conform to the detailed description text using the model features after diffusion denoising.
[0107] In some other embodiments, the diffusion denoising result further includes the target denoised model features output in the Nth round of diffusion denoising processing; the first training sample further includes a first three-dimensional model; as Figure 5 shown, compared with Figure 4 the training process, in this embodiment, before step 460, the method further includes:
[0108] Step 510, performing model decoding processing on the target denoised model features by a model decoding network to obtain a second three-dimensional model.
[0109] Step 520, calculating a model reconstruction loss according to the second three-dimensional model and the first three-dimensional model.
[0110] The model reconstruction loss is used to reflect the difference between the first three-dimensional model and the second three-dimensional model. The larger the model reconstruction loss, the greater the difference between the first three-dimensional model and the second three-dimensional model. On the contrary, the smaller the model reconstruction loss, the smaller the difference between the first three-dimensional model and the second three-dimensional model.
[0111] It can be understood that the closer the noise predicted by the diffusion network in each round of diffusion denoising processing during the diffusion denoising process is to the noise added during the diffusion noise addition process, that is, the more accurate the noise predicted by the diffusion network during the diffusion denoising process, and the higher the model decoding accuracy of the model decoding network, the smaller the difference between the second three-dimensional model and the first three-dimensional model. Therefore, the model reconstruction loss is affected by the noise prediction accuracy of the diffusion network and the model decoding accuracy of the model decoding network. Of course, if the model encoding network and the three-dimensional model generation network are trained together, the model reconstruction loss is also affected by the encoding accuracy of the model encoding network.
[0112] In some embodiments, the model reconstruction loss between the second three-dimensional model and the first three-dimensional model can be calculated through loss functions such as the mean square error loss function, the cross-entropy loss function, and the absolute value loss function.
[0113] Correspondingly, step 460 includes: step 461, performing weighted processing on the diffusion loss and the model reconstruction loss to obtain a target loss. During the weighted processing, the weighting coefficients of the diffusion loss and the model reconstruction loss can be the same or different, which can be specifically set according to needs.
[0114] Step 462: Adjust at least the parameters of the diffusion network and the model decoding network according to the target loss until the first training end condition is reached.
[0115] Similarly, if the model encoding network has been pre-trained before training the three-dimensional model generation network, in Step 462, the parameters of the diffusion network and the model decoding network can be adjusted according to the target loss in the manner of gradient descent until the first training end condition is reached. If the three-dimensional model generation network and the model encoding network are trained synchronously, in Step 462, the parameters of the model encoding network, the diffusion network, and the model decoding network can be adjusted according to the target loss in the manner of gradient descent until the first training end condition is reached.
[0116] Figure 6 It is a schematic diagram showing the training of a three-dimensional model generation network according to an embodiment of the present application. As Figure 6 shown, the first three-dimensional model is feature-encoded by the model encoding network to obtain sample encoding features. Then, the diffusion network performs multiple rounds of diffusion and noise addition processing on the target encoding features to output target noisy model features. Then, the sample text semantic features of the sample detailed description text of the first three-dimensional model are obtained. The diffusion network uses the sample text semantic features as diffusion conditions and performs multiple rounds of diffusion and denoising processing on the target noisy model features to obtain target denoised model features. Thereafter, the model decoding network performs model decoding on the target denoised model features to obtain a second three-dimensional model.
[0117] In the above embodiment, the diffusion network and the model decoding network are trained together, so that the diffusion denoising accuracy of the trained diffusion network is higher, and the decoding accuracy of the model decoding network is higher, which is convenient for accurately generating a three-dimensional model that meets the requirements by using the three-dimensional model generation network subsequently.
[0118] In some embodiments, the model encoding network and the three-dimensional model generation network can be separately trained, that is, the model encoding network is trained first, and then the trained model encoding network is used to generate sample encoding features of the first three-dimensional model for training the three-dimensional model generation network. Among them, the process of training the model encoding network can be as Figure 7 shown, including:
[0119] Step 710: Obtain a plurality of target samples; the target samples include target point cloud data collected for a reference three-dimensional model and position labels of each point cloud in the target point cloud data; the target point cloud data includes internal point clouds, surface point clouds, and external point clouds located outside the reference three-dimensional model of the reference three-dimensional model; the position labels are used to indicate the relative position relationship of the corresponding point cloud with respect to the reference three-dimensional model.
[0120] In this application, for the convenience of distinction, the three-dimensional model used to train the model encoding network is referred to as the reference three-dimensional model. The internal point cloud of the reference three-dimensional model refers to the point cloud collected inside the reference three-dimensional model (also called sampling points); the surface point cloud of the reference three-dimensional model refers to the point cloud collected on the surface of the reference three-dimensional model; the external point cloud of the reference three-dimensional model refers to the point cloud collected outside the reference three-dimensional model representing the external environment where the reference three-dimensional model is located.
[0121] In some embodiments, the relative position relationship of the point cloud with respect to the reference three-dimensional model can be used to indicate whether the point cloud belongs to the reference three-dimensional model; in this case, the position label can include a first position label for indicating that the point cloud belongs to the reference three-dimensional model and a second position label for indicating that the point cloud does not belong to the reference three-dimensional model. That is to say, if a point cloud is collected inside the reference three-dimensional model or on the surface of the reference three-dimensional model, the position label of the point cloud is the first position label; if a point cloud is collected outside the reference three-dimensional model, the position label of the point cloud is the second position label. In some embodiments, the first position label can be represented as 1, and the second position label can be represented as -1 or 0.
[0122] In some embodiments, the relative position relationship of the point cloud with respect to the reference three-dimensional model can be used to indicate whether the point cloud is located inside, on the surface, or outside the reference three-dimensional model. In this case, the position label can include: a surface position label for indicating that the point cloud is located on the surface of the reference three-dimensional model, an internal position label for indicating that the point cloud is located inside the reference three-dimensional model, and an external position label for indicating that the point cloud is located outside the reference three-dimensional model; for example, the surface position label can be represented as 1, the internal position label can be represented as 0, and the external position label can be represented as -1.
[0123] Step 720: The model encoding network performs feature encoding on each point cloud in the target point cloud data to obtain the encoded features of each point cloud in the target point cloud data.
[0124] Similar to the above, if the number of point clouds in the target point cloud data is large, the target point cloud data can be grouped into multiple point cloud groups, and then each point cloud group is input into the model encoding network for encoding processing respectively.
[0125] Step 730: Perform point cloud position decoding according to the encoded features of each point cloud in the target point cloud data to obtain the predicted position labels of each point cloud in the target point cloud data.
[0126] In some embodiments, if the position tags involved in multiple target samples are the first position tag and the second position tag, correspondingly, in step 730, the point cloud position is decoded according to the encoding features of each point cloud, that is, according to the encoding features of each point cloud, the position tag to which the point cloud belongs is predicted from the first position tag and the second position tag, and the predicted position tag becomes the predicted position tag of the point cloud.
[0127] In other embodiments, if the position tags involved in multiple target samples are the internal position tag, the surface position tag, and the external position tag, in step 730, according to the encoding features of each point cloud, the predicted position tag to which the point cloud belongs is predicted from the internal position tag, the surface position tag, and the external position tag.
[0128] The point cloud position can be decoded according to the encoding features of each point cloud through a position tag decoding network. In some embodiments, the decoder in a variational autoencoder can be used to decode the point cloud position according to the encoding features of each point cloud in the target point cloud data. In some embodiments, when the position tags include the first position tag and the second position tag, decoding the point cloud position can also be understood as predicting whether the position represented by the point cloud is occupied by the reference 3D model (Occupancy). If it is predicted that the position represented by a point cloud is occupied by the reference 3D model, the predicted position tag of the point cloud is correspondingly determined to be the first position tag; conversely, if it is predicted that the position represented by a point cloud is not occupied by the reference 3D model, the predicted position tag of the point cloud is correspondingly determined to be the second position tag.
[0129] Figure 8 An exemplary schematic diagram of processing using a model encoding network and a position tag decoding network is shown, as Figure 8 shown, the target point cloud data of the reference 3D model (including the coordinate information of each point cloud) is input into the model encoding network, and the model encoding network encodes each point cloud therein to obtain the encoding features of each point cloud in the target point cloud data; then, the encoding features of each point cloud in the target point cloud data are input into the position tag decoding network, and the predicted position tags of each point cloud in the target point cloud data are output.
[0130] Step 740, calculate the prediction loss according to the position tags and the corresponding predicted position tags of each point cloud in the target point cloud data.
[0131] The prediction loss reflects the difference between the predicted position labels for the point cloud and the corresponding position labels. The position label difference between the position label of each point cloud and the corresponding predicted position label can be calculated through a loss function. Then, the position label differences corresponding to all the point clouds in the target point cloud data are added together as the prediction loss. The loss function can be a mean squared error loss function, an absolute error loss function, a cross-entropy loss function, etc., which are not specifically limited herein.
[0132] Step 750: Adjust the parameters of the model encoding network according to the prediction loss until the second training end condition is reached.
[0133] Similarly, according to the prediction loss, the parameters of the model encoding network can be adjusted in the manner of gradient descent until the second training end condition is reached. The second training end condition can be that the number of iterations reaches the second number threshold, or the prediction loss converges, which is not specifically limited herein. In some embodiments, according to the prediction loss, the parameters of the model encoding network and the position label decoding network can also be adjusted in the manner of gradient descent until the second training end condition is reached.
[0134] Through the above training process, it can be ensured that the trained model encoding network can accurately perform feature encoding on the point cloud data of the 3D model, ensuring the accuracy of feature encoding. Then, the trained model encoding network can be used to assist in training the 3D model generation network, and subsequently, in combination with the trained 3D model generation network, to perform refined processing on the user-input 3D model to be optimized, generating a 3D model with richer appearance details. Moreover, in the above training process, not only the surface point cloud and the internal point cloud are used to train the model encoding network, but also the external point cloud is used to train the model encoding network, which enables the trained model encoding network to accurately distinguish the point cloud belonging to the 3D model from the point cloud located outside the 3D model during the encoding process, thereby distinguishing the encoding features of the two and improving the encoding accuracy of the model encoding features.
[0135] In some embodiments, before step 710, as Figure 9 shown, the method further includes:
[0136] Step 910: Collect the internal point cloud located inside the reference 3D model to obtain the first point cloud set.
[0137] In some embodiments, the internal point cloud located inside the reference 3D model can be collected, and the sampled internal point cloud is added to the first point cloud set. In this way, the number of point clouds in different three-dimensional space regions with the same volume in the first point cloud set is basically balanced.
[0138] In some embodiments, it is also possible to set the internal point cloud ratio, and then, based on the internal point cloud ratio and the total number of point clouds to be sampled in total, determine the number of internal point clouds to be sampled. After that, according to the number of internal point clouds to be sampled (assumed to be the first number), collect the internal point clouds located inside the reference 3D model (such as uniform collection or random collection), so that the total number of internal point clouds in the obtained first point cloud set is the first number.
[0139] Step 920: Collect the external point clouds located outside the reference 3D model to obtain a second point cloud set.
[0140] In some embodiments, uniformly collect the external point clouds located outside the reference 3D model, and add the sampled external point clouds to the second point cloud set. In this way, the number of point clouds in different three-dimensional space regions with the same volume in the second point cloud set is basically balanced.
[0141] In some embodiments, it is also possible to set the external point cloud ratio, and then, based on the external point cloud ratio and the total number of point clouds to be sampled in total, determine the number of external point clouds to be sampled, so that the total number of external point clouds in the obtained second point cloud set is the second number.
[0142] Step 930: Collect the surface point clouds located on the surface of the reference 3D model to obtain a third point cloud set.
[0143] In some embodiments, it is possible to uniformly collect the surface point clouds located on the surface of the reference 3D model, and add the sampled surface point clouds to the second point cloud set. In this way, the number of point clouds in different surface regions with the same area in the second point cloud set is basically balanced.
[0144] In some embodiments, before step 930, the method further includes: obtaining the curvature information of each surface of the reference 3D model; according to the curvature information, determining the point cloud collection density corresponding to each surface, where the point cloud collection density refers to the number of point clouds collected per unit area; the point cloud collection density is positively correlated with the curvature; correspondingly, step 930 includes: collecting the surface point clouds located on the surface of the reference 3D model according to the point cloud collection density corresponding to each surface to obtain a third point cloud set.
[0145] In the above - mentioned manner, for a surface with a larger curvature, the point - cloud acquisition density is greater. Thus, when the surface areas are the same, a surface with a larger curvature can sample more surface point - clouds. Conversely, for a surface with a smaller curvature, the number of sampled surface point - clouds is less. For a 3D model, a surface with a larger curvature has a stronger constraint on the shape of the 3D model and a stronger expression of the shape of the 3D model. Therefore, in this embodiment, the point - cloud acquisition density of the surface is determined according to the curvature of the surface, such that the larger the curvature of the surface, the greater the point - cloud acquisition density, so as to sample more surface point - clouds from the surface with a larger curvature, while for the surface with a smaller curvature, the point - cloud acquisition density is smaller and fewer surface point - clouds are sampled. In this way, it can be ensured that the third point - cloud set sampled can effectively represent the shape of the reference 3D model.
[0146] In some embodiments, a surface point - cloud ratio is set. Among them, the sum of the surface point - cloud ratio, the internal point - cloud ratio, and the external point - cloud ratio is 1. Then, according to the surface point - cloud ratio and the total number of point - clouds that need to be sampled in total, the number of surface point - clouds that need to be sampled is determined. After that, according to the number of surface point - clouds that need to be sampled (assumed to be the third number), and in combination with the point - cloud acquisition density determined for each surface of the reference 3D model, surface point - clouds are acquired, so that the total number of surface point - clouds in the obtained third point - cloud set is the third number.
[0147] Step 940, combine the first point - cloud set, the second point - cloud set, and the third point - cloud set to obtain the target point - cloud data.
[0148] Considering that the position information differences of point - clouds that are close to each other are small, and the feature similarity after encoding is also high, multiple point - clouds that are close to each other contribute little to improving the encoding accuracy of the model encoding network when used to train the model encoding network. Therefore, based on this consideration, representative partial point - clouds can be acquired to train the model encoding network. In this way, the training efficiency of the model encoding network can be improved, and the training duration of the model encoding network can be shortened.
[0149] Next, a specific embodiment is combined to illustrate the solution of this application. Figure 10 is a flowchart of training a model encoding network and a 3D model generation network shown according to an embodiment of this application. As Figure 10 shown, the method includes: Step 1010, acquire the first training data for training the model encoding network.
[0150] The first training data includes the target point cloud data of multiple reference 3D models and the position labels of each point cloud in the target point cloud data. Based on the 3D model data of the reference 3D models, point cloud acquisition can be performed for the reference 3D models to obtain the acquired point cloud data. Then, internal point clouds, external point clouds, and surface point clouds are respectively sampled from the acquired point cloud data. For example, 500,000 internal point clouds, 500,000 external point clouds, and 500,000 surface point clouds are sampled. The set of the sampled point clouds (internal point clouds, external point clouds, and surface point clouds) is used as the target point cloud data. Among them, during the process of sampling the surface point clouds, more sampling can be performed at positions with larger curvatures and less sampling at positions with smaller curvatures. For specific implementation details, reference can be made to the description above. Average sampling can be performed for the internal point clouds and the external point clouds.
[0151] Step 1020: Use the first training data to train the model encoding network.
[0152] In some embodiments, average random sampling can be performed in the target point cloud data of the reference 3D models. For example, for each reference 3D model, the point clouds in the target point cloud data of the reference 3D model are divided into multiple point cloud groups, and then the point clouds in each point cloud group are input into the model encoding network for encoding. The specific training process can be referred to the corresponding embodiments described above. In this embodiment, the model encoding network is the encoder in the variational autoencoder, and the 3D model generation network includes a diffusion network and an object decoding network. The encoder in the variational autoencoder can be used to perform compression encoding on the 3D models. Step 1030: The trained model encoding network performs compression encoding on the 3D surface point cloud data of multiple first 3D models to obtain the sample encoding features of each first 3D model. Figure 7 After the 3D surface point cloud data of each first 3D model is collected, assuming that there are a total of 500,000 surface point clouds in the 3D surface point cloud data, the 500,000 surface point clouds can be divided into multiple point cloud groups, and then each point cloud group is respectively input into the model encoding network for encoding.
[0153]
[0154] Step 1040: Image acquisition is performed for the first 3D model from multiple different perspectives to obtain two-dimensional model images from multiple perspectives.
[0155] Image acquisition can be performed for the first 3D model from 6 different perspectives to obtain two-dimensional model images of the first 3D model from 6 perspectives. The 6 perspectives are: front view, left view, right view, rear view, top view, and bottom view. In this way, the details of the first 3D model are presented from multiple perspectives, and a comprehensive understanding of the first 3D model can be obtained.
[0156] Step 1050: Generate descriptive text based on the 2D model images of the first 3D model from multiple perspectives to obtain the sample detailed descriptive text of the first 3D model. The image descriptive text of each 2D model image can be generated through an image-to-text model, and then the image descriptive texts of the 2D model images from multiple perspectives are combined as the sample detailed descriptive text of the first 3D model.
[0157] Step 1060: Use the sample encoding features of each first 3D model and the sample detailed descriptive text of the first 3D model to train the 3D model generation network. The process of training the 3D model generation network can refer to the description in the corresponding embodiment above, which will not be elaborated here. Figure 4 Corresponding to the description in the embodiment, it will not be elaborated here.
[0158] In some embodiments, after training the model encoding network and the 3D model generation network, the model encoding network and the 3D model generation network can be tested using test data. As Figure 11 shown, the testing process includes:
[0159] Step 1110: Obtain a test 3D model and test detailed descriptive text.
[0160] The test 3D model refers to a 3D model used to test the model encoding network and the 3D model generation network. It can be a 3D model with a relatively simple structure. For example, it can be a 3D rough model constructed from basic geometric bodies such as spheres, cuboids, cylinders, and cones provided in the basic solid model library. The test detailed descriptive text is used to indicate the detailed content that the 3D model needs to present.
[0161] Step 1120: Perform encoding processing on the test 3D model through the trained model encoding network to obtain the model encoding features of the test 3D model.
[0162] Step 1130: Input the model encoding features of the test 3D model and the text semantic features of the test detailed descriptive text into the 3D model generation network, and the 3D model generation network outputs the generated 3D model.
[0163] If it is evaluated that the 3D model output by the 3D model generation network has a basic shape that conforms to the test 3D model and the details conform to the test detailed descriptive text, it can be determined that the trained model generation network and the 3D model generation network meet the requirements, that is, they can be used for 3D model generation, realizing the refinement of a relatively rough 3D model to obtain a more refined 3D model with richer external details.
[0164] Through the solution of the present application, the detailed description text can be used to refine a three-dimensional model with a rough appearance, obtaining a three-dimensional model with richer details and a more refined appearance. For users, they do not need high three-dimensional model modeling capabilities. They can construct a three-dimensional model of the required shape through simple geometric bodies, then describe the required details through text, and generate a refined three-dimensional model with rich details according to the method of the present application, reducing the design threshold for three-dimensional modeling and improving the generation efficiency of three-dimensional models.
[0165] For example, if a user needs to have strong control over the head size and body proportion of a three-dimensional model, the user can simply piece together a relatively simple three-dimensional model using a cuboid, a cylinder, or a sphere. After that, use the detailed description text. Then, utilize the detailed description text to refine the constructed simple three-dimensional model, and a refined three-dimensional model can be generated with the same basic shape as the input three-dimensional model and details that conform to the detailed description text. In this way, even non-professional modelers can automatically generate the three-dimensional model desired by the user according to the method of the present application.
[0166] In addition, in the solution of the present application, the point cloud data of the three-dimensional model can be encoded by a model encoding network to obtain model encoding features. The point cloud data of the three-dimensional model includes the position information of multiple point clouds, and it is not necessary to use a three-dimensional convolutional network to process the three-dimensional model in three-dimensional space. In this way, the computing resources used by the model encoding network for processing are also less, reducing the cost of three-dimensional model generation.
[0167] The following introduces the device embodiments of the present application, which can be used to execute the methods in the above embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the above method embodiments of the present application.
[0168] Figure 12 is a block diagram of a three-dimensional model processing device shown according to an embodiment of the present application, as Figure 12 shown. The three-dimensional model processing device includes: an acquisition module 1210, configured to acquire a three-dimensional model to be optimized and a detailed description text; an encoding module 1220, configured to perform encoding processing on the three-dimensional model to be optimized to obtain model encoding features; a semantic feature extraction module 1230, configured to extract semantic features from the detailed description text to obtain text semantic features; and a three-dimensional model generation module 1240, configured to generate a model by a three-dimensional model generation network according to the text semantic features and the model encoding features to obtain a target three-dimensional model whose basic shape conforms to the three-dimensional model to be optimized and whose details conform to the detailed description text.
[0169] In some embodiments, the three-dimensional model generation network includes a diffusion network and a model decoding network; the three-dimensional model generation module 1240 includes: a diffusion denoising processing unit, configured to use the diffusion network to perform T rounds of diffusion denoising processing on the model encoding features with the text semantic features as the diffusion condition, to obtain the target model features output by the T-th round of diffusion processing; wherein, in the T rounds of diffusion processing, the model features output by the i-th round of diffusion processing are used as the input of the (i + 1)-th round of diffusion processing, i and T are positive integers, and i + 1 ≤ T; a model decoding unit, configured to use the model decoding network to perform model decoding processing on the target model features, to obtain a target three-dimensional model whose basic shape conforms to the to-be-optimized three-dimensional model and whose details conform to the detail description text.
[0170] In some embodiments, the processing device for the three-dimensional model further includes: a first acquisition module, configured to acquire the sample encoding features of the first three-dimensional model; a second acquisition module, configured to acquire the sample text semantic features of the sample detail description text of the first three-dimensional model; a first diffusion noise addition module, configured to use the diffusion network to perform N rounds of diffusion noise addition processing on the sample encoding features, to obtain a diffusion noise addition result; the diffusion noise addition result includes the target noise-added model features obtained by the N-th round of diffusion noise addition processing and the noises added in each round of diffusion noise addition processing; N is a positive integer not exceeding T; a first diffusion denoising module, configured to use the diffusion network to perform N rounds of diffusion denoising processing on the target noise-added model features with the sample text semantic features as the diffusion condition, to obtain a diffusion denoising result; the diffusion denoising result includes the noises removed in each round of diffusion denoising processing in the N rounds of diffusion denoising processing; a diffusion loss determination module, configured to determine the diffusion loss according to the noises added in each round of diffusion noise addition processing and the noises removed in each round of diffusion denoising processing; a first adjustment module, configured to at least adjust the parameters of the diffusion network according to the diffusion loss until a first training end condition is reached.
[0171] In some embodiments, the diffusion denoising result further includes the target denoised model features output by the N-th round of diffusion denoising processing; the processing device for the three-dimensional model further includes: a first decoding module, configured to use the model decoding network to perform model decoding processing on the target denoised model features, to obtain a second three-dimensional model; a model reconstruction loss determination module, configured to calculate the model reconstruction loss according to the second three-dimensional model and the first three-dimensional model; correspondingly, the first adjustment module is configured to: perform weighted processing on the diffusion loss and the model reconstruction loss to obtain a target loss; at least adjust the parameters of the diffusion network and the model decoding network according to the target loss until a first training end condition is reached.
[0172] In some embodiments, the processing device for the three-dimensional model further includes: an image acquisition module, configured to acquire two-dimensional model images of the first three-dimensional model from at least one perspective; a text generation module, configured to generate text descriptions for the two-dimensional model images from each perspective to obtain image description texts corresponding to the two-dimensional model images; a first combination module, configured to combine the image description texts corresponding to the two-dimensional model images from at least one perspective to obtain a sample detail description text of the first three-dimensional model.
[0173] In some embodiments, the processing device for the three-dimensional model further includes: a collection module, configured to collect three-dimensional surface point cloud data of the first three-dimensional model; a first encoding module, configured to perform encoding processing on the three-dimensional surface point cloud data by a model encoding network to obtain sample encoding features of the first three-dimensional model.
[0174] In some embodiments, the processing device for the three-dimensional model further includes: a sampling module, configured to sample the three-dimensional surface point cloud data to obtain M groups of sub-surface point cloud data; M is a positive integer greater than 1; a second encoding module, configured to perform encoding processing on each group of the M groups of sub-surface point cloud data by the model encoding network to obtain encoding features of each group of sub-surface point cloud data; a second combination module, configured to combine the encoding features of the M groups of sub-surface point cloud data to obtain sample encoding features of the first three-dimensional model.
[0175] In some embodiments, the processing device for the three-dimensional model further includes: a sample acquisition module, configured to acquire a plurality of target samples; the target samples include target point cloud data collected for a reference three-dimensional model and position labels of each point cloud in the target point cloud data; the target point cloud data includes internal point clouds, surface point clouds of the reference three-dimensional model, and external point clouds located outside the reference three-dimensional model; the position labels are used to indicate the relative position relationship of the corresponding point clouds with respect to the reference three-dimensional model; a third encoding module, configured to perform feature encoding on each point cloud in the target point cloud data by the model encoding network to obtain encoding features of each point cloud in the target point cloud data; a position decoding module, configured to perform point cloud position decoding according to the encoding features of each point cloud in the target point cloud data to obtain predicted position labels of each point cloud in the target point cloud data; a prediction loss calculation module, configured to calculate a prediction loss according to the position labels and the corresponding predicted position labels of each point cloud in the target point cloud data; a second adjustment module, configured to adjust the parameters of the model encoding network according to the prediction loss until a second training end condition is reached.
[0176] In some embodiments, the processing device for the three-dimensional model further includes: a first acquisition module, configured to acquire internal point clouds located inside the reference three-dimensional model to obtain a first point cloud set; a second acquisition module, configured to acquire external point clouds located outside the reference three-dimensional model to obtain a second point cloud set; a third acquisition module, configured to acquire surface point clouds located on the surface of the reference three-dimensional model to obtain a third point cloud set; and a third combination module, configured to combine the first point cloud set, the second point cloud set, and the third point cloud set to obtain target point cloud data.
[0177] In some embodiments, the processing device for the three-dimensional model further includes: a curvature information acquisition module, configured to acquire the curvature information of each surface of the reference three-dimensional model; a sampling density determination module, configured to determine the point cloud acquisition density corresponding to each surface according to the curvature information, where the point cloud acquisition density refers to the number of point clouds acquired per unit area; the point cloud acquisition density is positively correlated with the curvature; and a third acquisition module, configured to: acquire surface point clouds located on the surface of the reference three-dimensional model according to the point cloud acquisition density corresponding to each surface to obtain a third point cloud set.
[0178] In some embodiments, the encoding module includes: a first acquisition unit, configured to acquire three-dimensional surface point cloud data of the three-dimensional model to be optimized; and an encoding processing unit, configured to perform encoding processing on the three-dimensional surface point cloud data of the three-dimensional model to be optimized by a model encoding network to obtain model encoding features.
[0179] Figure 13 The structural schematic diagram of the computer system of the electronic device suitable for implementing the embodiments of the present application is shown. It should be noted that Figure 13 The computer system 1300 of the shown electronic device is only an example, and should not bring any limitation to the functions and usage scopes of the embodiments of the present application. The electronic device can be used to execute the processing method for the three-dimensional model provided by the present application. The electronic device can be a terminal, a server, etc., which is not specifically limited herein.
[0180] Such as Figure 13As shown, computer system 1300 includes a Central Processing Unit (CPU) 1301, which can perform various appropriate actions and processes according to a program stored in a Read-Only Memory (ROM) 1302 or a program loaded from a storage section 1308 into a Random Access Memory (RAM) 1303, such as executing the methods in the above embodiments. In the RAM 1303, various programs and data required for system operations are also stored. The CPU 1301, ROM 1302, and RAM 1303 are connected to each other via a bus 1304. An Input / Output (I / O) interface 1305 is also connected to the bus 1304.
[0181] The following components are connected to the I / O interface 1305: an input section 1306 including a keyboard, a mouse, etc.; an output section 1307 including, for example, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc. and a speaker, etc.; a storage section 1308 including a hard disk, etc.; and a communication section 1309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1309 performs communication processing via a network such as the Internet. A drive 1310 is also connected to the I / O interface 1305 as needed. A removable medium 1311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1310 as needed so that a computer program read from it can be installed into the storage section 1308 as needed.
[0182] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1309, and / or installed from the removable medium 1311. When the computer program is executed by a Central Processing Unit (CPU) 1301, various functions defined in the system of the present application are executed.
[0183] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0184] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0185] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. In some cases, the names of these units do not constitute a limitation on the units themselves.
[0186] As another aspect, this application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist alone without being assembled into the electronic device. The above computer-readable storage medium carries computer-readable instructions, and when the computer-readable storage instructions are executed by a processor, the method in any of the above embodiments is implemented.
[0187] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be an overall module or a part of the unit of the function of the module or unit.
[0188] According to one aspect of the embodiments of this application, a computer program product is provided. The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method in any of the above embodiments.
[0189] It should be noted that although several modules or units of the devices for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of this application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0190] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0191] After considering the specification and practicing the embodiments disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application.
[0192] It should be understood that the present application is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A method for processing a three-dimensional model, characterized in that: include: Obtain the 3D model to be optimized and detailed description text; Performing encoding processing on the three-dimensional model to be optimized to obtain model encoding features; Extracting semantic features from the detailed description text to obtain text semantic features; A three-dimensional model generation network is used to generate a model according to the text semantic features and the model encoding features to obtain a target three-dimensional model whose basic shape is consistent with the three-dimensional model to be optimized and whose details conform to the detail description text.
2. The method according to claim 1, characterized in that The three-dimensional model generation network includes a diffusion network and a model decoding network; The three-dimensional model generation network generates a model according to the text semantic features and the model encoding features to obtain a target three-dimensional model whose basic shape is consistent with the three-dimensional model to be optimized and whose details conform to the detail description text, including: The diffusion network uses the text semantic features as diffusion conditions, performs T rounds of diffusion denoising on the model encoding features, and obtains the target model features output by the T rounds of diffusion processing; wherein, in the T rounds of diffusion processing, the model features output by the i-th round of diffusion processing are used as the input of the i+1-th round of diffusion processing, i and T are positive integers, and i+1≤T; The model decoding network performs model decoding processing on the target model features to obtain a target three-dimensional model whose basic shape is consistent with the three-dimensional model to be optimized and whose details conform to the detail description text.
3. The method according to claim 2, characterized in that Before generating a network from a three-dimensional model and generating a model according to the text semantic features and the model encoding features to obtain a target three-dimensional model whose basic shape is consistent with the three-dimensional model to be optimized and whose details conform to the detail description text, the method further comprises: Obtaining sample encoding features of the first three-dimensional model; Acquire sample text semantic features of sample detailed description text of the first three-dimensional model; The diffusion network performs N rounds of diffusion noise processing on the sample coding features to obtain a diffusion noise result; the diffusion noise result includes the target noise model feature obtained by the Nth round of diffusion noise processing and the noise added in each round of the diffusion noise processing; N is a positive integer not exceeding T; The diffusion network uses the semantic features of the sample text as diffusion conditions to perform N rounds of diffusion denoising on the features of the target noisy model to obtain a diffusion denoising result; the diffusion denoising result includes the noise removed in each round of the diffusion denoising in the N rounds of diffusion denoising; Determine diffusion loss according to the noise added in each round of the diffusion noise adding process and the noise removed in each round of the diffusion noise removing process; According to the diffusion loss, at least a parameter of the diffusion network is adjusted until a first training end condition is reached.
4. The method according to claim 3, characterized in that: The diffusion denoising result also includes the target denoising model features output in the Nth round of diffusion denoising processing; According to the diffusion loss, at least adjusting the parameters of the diffusion network until the first training end condition is reached, the method further includes: The model decoding network performs model decoding processing on the target denoising model features to obtain a second three-dimensional model; Calculating a model reconstruction loss according to the second three-dimensional model and the first three-dimensional model; The step of adjusting at least a parameter of the diffusion network according to the diffusion loss until a first training end condition is reached includes: Performing weighted processing on the diffusion loss and the model reconstruction loss to obtain a target loss; According to the target loss, at least parameters of the diffusion network and the model decoding network are adjusted until a first training end condition is reached.
5. The method according to any one of claims 2 to 4, characterized in that Before obtaining a plurality of first training samples, the method further includes: Acquire a two-dimensional model image of the first three-dimensional model at at least one viewing angle; Generating text descriptions for the two-dimensional model images at each of the viewing angles to obtain image description texts corresponding to each of the two-dimensional model images; The image description texts corresponding to the two-dimensional model image under the at least one viewing angle are combined to obtain the sample detail description text of the first three-dimensional model.
6. The method according to any one of claims 2 to 4, characterized in that Before obtaining a plurality of first training samples, the method further includes: Collecting three-dimensional surface point cloud data of the first three-dimensional model; The model encoding network performs encoding processing according to the three-dimensional surface point cloud data to obtain sample encoding features of the first three-dimensional model.
7. The method according to claim 6, characterized in that Before the model encoding network performs encoding processing according to the three-dimensional surface point cloud data to obtain the sample encoding features of the first three-dimensional model, the method further includes: Sampling the three-dimensional surface point cloud data to obtain M groups of sub-surface point cloud data; M is a positive integer greater than 1; The model encoding network performs encoding processing on each group of sub-surface point cloud data in the M groups of sub-surface point cloud data respectively to obtain encoding features of each group of sub-surface point cloud data; The coding features of the M groups of sub-surface point cloud data are combined to obtain sample coding features of the first three-dimensional model.
8. The method according to claim 6, characterized in that Before the model encoding network performs encoding processing according to the three-dimensional surface point cloud data to obtain the sample encoding features of the first three-dimensional model, the method further includes: Acquire multiple target samples; the target samples include target point cloud data collected for a reference three-dimensional model and position tags of each point cloud in the target point cloud data; the target point cloud data include an internal point cloud, a surface point cloud and an external point cloud located outside the reference three-dimensional model; the position tags are used to indicate the relative position relationship of the corresponding point cloud with respect to the reference three-dimensional model; The model encoding network performs feature encoding on each point cloud in the target point cloud data to obtain encoding features of each point cloud in the target point cloud data; Decoding the point cloud position according to the coding features of each point cloud in the target point cloud data to obtain a predicted position label of each point cloud in the target point cloud data; Calculating prediction loss based on the position label of each point cloud in the target point cloud data and the corresponding predicted position label; According to the prediction loss, the parameters of the model encoding network are adjusted until a second training end condition is reached.
9. The method according to claim 8, characterized in that Before acquiring a plurality of target samples, the method further includes: Collecting an internal point cloud located inside the reference three-dimensional model to obtain a first point cloud set; Collecting an external point cloud located outside the reference three-dimensional model to obtain a second point cloud set; Collecting a surface point cloud located on the surface of the reference three-dimensional model to obtain a third point cloud set; The first point cloud set, the second point cloud set and the third point cloud set are combined to obtain the target point cloud data.
10. The method according to claim 9, characterized in that Before collecting the surface point cloud located on the surface of the reference three-dimensional model to obtain the third point cloud set, the method further includes: Acquiring curvature information of each surface of the reference three-dimensional model; Determine the point cloud acquisition density corresponding to each of the surfaces according to the curvature information, wherein the point cloud acquisition density refers to the number of point clouds collected per unit area; the point cloud acquisition density is positively correlated with the curvature; The collecting of the surface point cloud located on the surface of the reference three-dimensional model to obtain a third point cloud set includes: According to the point cloud collection density corresponding to each of the surfaces, the surface point cloud located on the surface of the reference three-dimensional model is collected to obtain a third point cloud set.
11. The method according to any one of claims 1 to 4, characterized in that: The encoding process of the three-dimensional model to be optimized to obtain the model encoding features includes: Collecting three-dimensional surface point cloud data of the three-dimensional model to be optimized; The model encoding network performs encoding processing according to the three-dimensional surface point cloud data of the three-dimensional model to be optimized to obtain the model encoding features.
12. A three-dimensional model processing device, characterized in that: include: An acquisition module is used to acquire the three-dimensional model to be optimized and the detailed description text; A coding module, used for coding the three-dimensional model to be optimized to obtain a model coding feature; A semantic feature extraction module is used to extract semantic features from the detailed description text to obtain text semantic features; The three-dimensional model generation module is used to generate a network from a three-dimensional model, generate a model according to the text semantic features and the model encoding features, and obtain a target three-dimensional model whose basic shape is consistent with the three-dimensional model to be optimized and whose details conform to the detail description text.
13. An electronic device, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 11 is implemented.
14. A computer-readable storage medium having computer-readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, the method according to any one of claims 1 to 11 is implemented.
15. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the method according to any one of claims 1 to 11 is implemented.