Three-dimensional grid generation method and device, computer program product and electronic equipment

Through the method of feature extraction and attention fusion, combined with global information and generated grid features, the problem of low quality of the existing three-dimensional grid generation method is solved and high-quality three-dimensional grid generation is achieved.

CN120163940APending Publication Date: 2025-06-17NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510219013.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing three-dimensional grid generation methods have problems such as fuzzy grids and poor edges, which make it difficult to generate high-quality topological structures and details.

Method used

By obtaining the grid cell sequence of the grid to be generated, feature extraction and attention fusion are performed, combining global information features and generated grid cell features for autoregression generation, and finally feature mapping is performed to generate the target three-dimensional grid.

Benefits of technology

The refined control of the three-dimensional grid is realized, and the generation of coherent, rich in details and context-dependent grid structures are generated, which improves the generation quality of the three-dimensional grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163940A_ABST
    Figure CN120163940A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, in particular to a three-dimensional grid generation method and device, a computer program product and electronic equipment. The three-dimensional grid generation method comprises the steps of obtaining a grid unit sequence of a grid to be generated, and performing feature extraction on the grid unit sequence through a grid generation model to obtain grid abstract features; performing attention fusion processing on the grid abstract features to obtain global information features; performing autoregression generation according to the global information features and the generated grid unit features to obtain current grid unit features; and performing feature mapping on the current grid unit features to obtain grid unit attribute information, and performing rendering based on the grid unit attribute information to generate a target three-dimensional grid. According to the invention, the quality of the generated three-dimensional grid can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and more particularly, to a three-dimensional mesh generation method, a three-dimensional mesh generation device, a computer program product, and an electronic device. Background Art

[0002] A three-dimensional mesh model is an important way to represent a three-dimensional object, and its quality directly affects the effects of making three-dimensional objects in various scenarios, such as film and television animation, game development, virtual reality, and high-precision digital twins. Although there are currently various three-dimensional mesh generation methods, such as manual modeling, automatic generation schemes based on traditional algorithms, and generation schemes based on deep learning, there are still problems with low quality, such as blurred generated meshes and unrefined edges.

[0003] It should be noted that the information disclosed in the above background art is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0004] The purpose of the present disclosure is to provide a three-dimensional mesh generation method and device, a computer program product, and an electronic device, thereby at least to some extent improving the quality of the generated three-dimensional mesh.

[0005] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be partially learned through the practice of the present disclosure.

[0006] According to one aspect of the present disclosure, there is provided a three-dimensional mesh generation method, including: obtaining a sequence of mesh units to be generated, and extracting features of the sequence of mesh units through a mesh generation model to obtain mesh abstract features; performing attention fusion processing on the mesh abstract features to obtain global information features; performing autoregressive generation based on the global information features and the features of the already generated mesh units to obtain the features of the current mesh unit; and performing feature mapping on the features of the current mesh unit to obtain mesh unit attribute information, so as to render based on the mesh unit attribute information to generate a target three-dimensional mesh.

[0007] In an exemplary embodiment of the present disclosure, the method further includes: obtaining mesh generation guidance information, where the mesh generation guidance information includes at least one of three-dimensional point cloud information, mesh unit quantity information, and mesh unit topology information; performing encoding fusion on the mesh generation guidance information to obtain mesh generation guidance features; and performing attention fusion processing on the mesh abstract features to obtain global information features, including: performing attention fusion processing on the mesh abstract features and the mesh generation guidance features to obtain global information features.

[0008] In an exemplary embodiment of the present disclosure, an attention fusion process is performed on the grid abstract feature and the grid generation guidance feature to obtain a global information feature, including: mapping the grid abstract feature to a query space to obtain a query vector; mapping the grid generation guidance feature to a key-value space to obtain a key vector, and mapping the grid generation guidance feature to a key-value space to obtain a value vector, where different mapping matrices are used to obtain the key vector and the value vector; performing attention calculation based on the query vector, the key vector, and the value vector to obtain the global information feature.

[0009] In an exemplary embodiment of the present disclosure, feature extraction is performed on the grid cell sequence to obtain a grid abstract feature, including: performing depthwise separable convolution processing on the grid cell sequence to obtain a convolution feature; performing a non-linear transformation process on the convolution feature to obtain a processed convolution feature, and performing layer normalization processing on the processed convolution feature to obtain the grid abstract feature.

[0010] In an exemplary embodiment of the present disclosure, autoregressive generation is performed based on the global information feature and the generated grid cell feature to obtain the current grid cell feature, including: for each time step of autoregressive generation, concatenating the global information feature and the generated grid cell feature corresponding to the time step to obtain a concatenated feature, and performing deconvolution processing on the concatenated feature to obtain a deconvolution feature; performing a non-linear conversion on the deconvolution feature to obtain a conversion feature, and performing layer normalization processing on the conversion feature to obtain the current grid cell feature.

[0011] In an exemplary embodiment of the present disclosure, obtaining a grid cell sequence of the grid to be generated includes: obtaining the grid cells of the grid to be generated, and arranging the grid cells according to a predefined sorting rule to obtain a grid cell sequence; performing feature mapping on the current grid cell feature to obtain grid cell attribute information, further including: verifying the grid cell attribute information based on the predefined sorting rule, and adjusting the grid cell attribute information of the target grid cell that does not meet the predefined sorting rule; where the predefined sorting rule at least includes a vertex generation order and an in-plane vertex order, and the in-plane vertex order is used to indicate the consistent order of vertices within different patches.

[0012] In an exemplary embodiment of the present disclosure, the grid cell attribute information is verified based on a predefined sorting rule, and the grid cell attribute information of a target grid cell that does not meet the predefined sorting rule is adjusted, including: performing coordinate resampling processing based on the vertex coordinates corresponding to the target grid cell to obtain vertex updated coordinates that conform to the predefined sorting rule; if vertex updated coordinates that conform to the predefined sorting rule are not obtained through the coordinate resampling processing, determining a sorting constraint loss based on the vertex coordinates, and determining a target loss according to the generation loss of the grid generation model and the sorting constraint loss, so that the grid generation model outputs vertex updated coordinates that conform to the predefined sorting rule based on the target loss; wherein, the generation loss is used to measure the difference between the grid cell attribute information output by the grid generation model and the true grid attributes.

[0013] In an exemplary embodiment of the present disclosure, the method further includes: if vertex updated coordinates that conform to the predefined sorting rule cannot be obtained based on the target loss, backtracking to the target network layer that predicts the target grid cell; adjusting the output features of the target network layer to obtain adjusted output features, so as to predict vertex updated coordinates that conform to the predefined sorting rule according to the adjusted output features.

[0014] According to one aspect of the present disclosure, there is provided a three-dimensional grid generation device, including: an information acquisition module, configured to acquire a grid cell sequence of a grid to be generated, and perform feature extraction on the grid cell sequence through a grid generation model to obtain grid abstract features; a feature extraction module, configured to perform attention fusion processing on the grid abstract features to obtain global information features; an information reconstruction module, configured to perform autoregressive generation according to the global information features and the generated grid cell features to obtain current grid cell features; a rendering processing module, configured to perform feature mapping on the current grid cell features to obtain grid cell attribute information, so as to perform rendering based on the grid cell attribute information to generate a target three-dimensional grid.

[0015] According to one aspect of the present disclosure, there is provided a computer program product, including a computer program, where when the computer program is executed by a processor, the method of any one of the above is implemented.

[0016] According to one aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory, configured to store executable instructions of the processor; wherein, the processor is configured to execute the method of any one of the above by executing the executable instructions.

[0017] The 3D mesh generation method in the exemplary embodiments of the present disclosure extracts mesh abstract features from a sequence of mesh cells, performs attention fusion on the mesh abstract features to obtain global information features, and performs autoregressive generation based on the global information features and the generated mesh cell features to obtain the current mesh cell features, and maps the current mesh cell features to mesh cell attribute information, so that rendering can be performed based on the mesh cell attribute information to generate a target 3D mesh. On the one hand, through hierarchical feature processing, mesh abstract features, global information features, and current mesh cell features are sequentially extracted, realizing adaptive feature extraction and information transfer of the mesh cell sequence, and realizing fine control of the generated 3D mesh. At the same time, in the autoregressive generation process, the global information features and the generated mesh cell features are combined, that is, when generating each mesh cell, it is carried out under the dual guidance of the global information features and the previously generated mesh cell features. The global information features provide the overall information of the 3D mesh, and the generated mesh cell features provide local detail information and context information, so as to be able to generate a coherent, detailed, and context-dependent mesh structure, improving the quality of the generated 3D mesh. On the other hand, the multi-granularity processing mechanism through hierarchical processing can efficiently process long mesh sequences. In addition, through the autoregressive generation process, it can be ensured that only past information (i.e., the generated mesh cell features) is relied on when generating the mesh, ensuring the causality of the output sequence while avoiding information leakage.

[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings

[0019] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner.

[0020] Figure 1 An application environment according to an exemplary embodiment of the present disclosure is shown.

[0021] Figure 2 A network structure diagram of a mesh generation model according to an exemplary embodiment of the present disclosure is shown.

[0022] Figure 3 A flowchart of a 3D mesh generation method according to an exemplary embodiment of the present disclosure is shown.

[0023] Figure 4 A schematic diagram of a multi-head attention module according to an exemplary embodiment of the present disclosure is shown.

[0024] Figure 5Shows a flowchart of a guiding control method for grid generation according to an exemplary embodiment of the present disclosure.

[0025] Figure 6 Shows a schematic structural diagram of another grid generation model according to an exemplary embodiment of the present disclosure.

[0026] Figure 7 Shows a schematic diagram of a separable convolution module according to an exemplary embodiment of the present disclosure.

[0027] Figure 8 Shows a schematic diagram of a causal convolution module according to an exemplary embodiment of the present disclosure.

[0028] Figure 9 Shows a schematic diagram of the composition of a three-dimensional grid generation device according to an exemplary embodiment of the present disclosure.

[0029] Figure 10 Shows a block diagram of an electronic device according to an exemplary embodiment of the present disclosure.

[0030] In the drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed implementation manners

[0031] Now, exemplary embodiments will be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The same reference numerals in the drawings denote the same or similar structures, and thus their detailed descriptions will be omitted.

[0032] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known structures, methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.

[0033] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more software-hardened modules, or in different networks and / or processor devices and / or microcontroller devices.

[0034] Currently, methods for generating 3D mesh models include manual modeling, automatic generation methods based on traditional algorithms such as the Marching Cubes algorithm, and automatic generation methods based on deep learning such as Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), Transformer models, etc. However, manual modeling requires high skills of modelers and is difficult to meet the needs of large-scale 3D content production. Automatic generation methods based on traditional algorithms are difficult to generate high-quality topological structures when dealing with complex geometric structures, and cannot achieve fine control of 3D meshes. While automatic generation methods based on deep learning are difficult to generate high-quality mesh details when dealing with meshes with high resolution or complex topologies, and there are problems such as mesh blurring and rough edges. Moreover, when dealing with long sequences of mesh data, the computational complexity is high and the memory consumption is large. For example, in the generation method based on the Transformer model, when the number of mesh faces exceeds a certain threshold, the computational amount will increase exponentially, and this global self-attention mechanism cannot well handle 3D mesh structures with local characteristics.

[0035] Based on this, the exemplary embodiments of the present disclosure provide a 3D mesh generation method. Through hierarchical processing, mesh abstract features, global information features, and current mesh cell features are sequentially extracted, realizing adaptive feature extraction and information transfer for the mesh cell sequence, and achieving fine control of the generated 3D mesh. When generating each mesh cell, it is carried out under the dual guidance of the global information features and the features of the previously generated mesh cells, capable of generating a coherent, detail-rich, and context-dependent mesh structure, improving the quality of the generated 3D mesh, and efficiently processing long mesh sequences through the multi-granularity processing mechanism of hierarchical processing.

[0036] It should be noted that the 3D mesh generation method of the exemplary embodiments of the present disclosure can be applied to fields such as film and television animation, game development, virtual reality, and high-precision digital twins, without limitation thereto.

[0037] The 3D mesh generation method provided by the exemplary embodiments of the present disclosure can be applied to an application environment as Figure 1 shown. Among them, the terminal 101 communicates with the server 102 through the network. The data storage system can store the data that the server 102 needs to process. The data storage system can be integrated on the server 102, or placed in the cloud or other network servers.

[0038] In an exemplary embodiment, the three-dimensional mesh generation method provided by the exemplary embodiment of the present disclosure can be executed by the server 102, and the corresponding three-dimensional mesh generation device is disposed in the server 102. Correspondingly, in this manner executed by the server 102, the server 102 can start executing the technical solution in the exemplary embodiment of the present disclosure in response to a trigger command, where the trigger command can be sent by the terminal used by the user, or can be locally triggered by the server in response to some automation events.

[0039] Among them, the server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, and can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server 102 can execute background tasks.

[0040] Furthermore, in another exemplary embodiment, the terminal 101 can also have a similar function to the server 102, so as to execute the three-dimensional mesh generation method provided by the exemplary embodiment of the present disclosure.

[0041] Among them, the terminal 101 can be a smart phone, a tablet computer, a notebook computer, or a desktop computer. The terminal 101 can also be referred to as a mobile terminal, a terminal device, a mobile device, etc. The exemplary embodiment of the present disclosure does not limit the type of the terminal 101.

[0042] In addition, the technical solution of the exemplary embodiment of the present disclosure can also be executed collaboratively by the terminal 101 and the server 102. In this manner executed collaboratively by the terminal 101 and the server 102, some steps in the technical solution provided by the exemplary embodiment of the present disclosure are executed by the terminal 101, and some other steps are executed by the server 102. It should be noted that in this manner executed collaboratively by the terminal 101 and the server 102, the steps respectively executed by the terminal 101 and the server 102 can be dynamically adjusted according to the actual situation, and no special limitation is imposed thereon.

[0043] Among them, the terminal 101 and the server 102 can be directly or indirectly connected through a wireless communication method, and the exemplary embodiment of the present disclosure does not impose special limitations thereon.

[0044] First, the concepts or terms involved in the exemplary embodiment of the present disclosure will be introduced below.

[0045] 3D mesh (3D mesh model), a 3D geometric shape composed of vertices, edges, and faces, can be used to describe the shape of 3D objects. A 3D mesh can be composed of multiple triangles, quadrilaterals, or other simple polygons, which are connected to each other to form a complex three-dimensional structure. For example, a triangular mesh consists of three vertices forming a face, and these vertices have x, y, z coordinates, and the vertices are connected by edges.

[0046] Mesh element, which is the basic component of a 3D mesh, can refer to a vertex, a patch, or a combination of a vertex and a patch. Among them, a vertex is represented by its three-dimensional coordinates (x, y, z), and each patch can be represented by the vertex indices that make up the patch. The information of the vertex and the patch can also be encoded as a mesh element.

[0047] Multi-head attention mechanism, a method used to capture long-range dependencies in sequential data, can enable the model to focus on information at different positions in the input sequence and use this information for better prediction.

[0048] Transposed convolution, similar to the convolution operation, is used to map a low-resolution feature map to a high-resolution feature map.

[0049] Causal convolution, a special convolution operation, can be used to process time series data. When processing sequential data, the output at the current moment only depends on the input data at the current moment and its previous data, and does not depend on future data.

[0050] As Figure 2 shows a network structure diagram of a mesh generation model, as Figure 2 shown, including a separable convolution module, a multi-head attention module, a causal convolution module, and a decoding module cascaded in sequence. At the same time, it should be understood that according to the embodiments of the present disclosure, the features and functions of two or more modules can be embodied in one module. Conversely, the features and functions of one module can be further divided and embodied by multiple modules. The exemplary embodiments of the present disclosure include, but are not limited to, the network structure of the mesh generation model described above, taking Figure 2 the network architecture shown as an example for illustration.

[0051] Next, in combination with Figure 2 the network architecture, the 3D mesh generation method of the exemplary embodiments of the present disclosure will be specifically described.

[0052] As Figure 3 shown is a flowchart of a 3D mesh generation method of an exemplary embodiment of the present disclosure. The 3D mesh generation method includes steps S310 to S340, specifically as follows:

[0053] In step S310, a grid cell sequence of the grid to be generated is obtained, and the grid generation model is used to extract features of the grid cell sequence to obtain grid abstract features.

[0054] In an exemplary embodiment of the present disclosure, the grid to be generated consists of a grid cell sequence of vertices, edges, and faces. To facilitate neural network processing, the structured data of the three-dimensional grid is converted into a serialized manner, that is, the grid cells of the grid to be generated are arranged according to a predefined sorting rule to obtain a grid cell sequence. For example, they are arranged in the yzx scan order. For example, the grid cells are first sorted according to the y-axis coordinate, then sorted according to the z-axis coordinate, and finally sorted according to the x-axis coordinate to obtain a grid cell sequence.

[0055] Among them, the separable convolution module in the grid generation model can be used to extract the grid abstract features of the grid cell sequence. The separable convolution module can be used to process the grid cell sequence, extract abstract features based on the relationship between the grid cells in the sequence, and the use of separable convolution can reduce the number of model parameters and the amount of computation, thereby reducing the amount of computation of the subsequent processing module.

[0056] The input of the separable convolution module is the grid cell sequence Among them, N is the length of the grid cell sequence, and din is the dimension of each grid cell, which is related to the type of grid cell. For example, if the grid cell is a vertex, the dimension is 3. If the grid cell is a face, the dimension is the number of vertex indices of the face (for example, a triangular face includes three vertices, so the dimension of the triangular face is 3), and the index of each vertex needs to be encoded.

[0057] Exemplarily, if a three-dimensional model consists of the following vertices and faces: vertices (v1, v2, v3, v4) and faces (f1(v1, v2, v3), f2(v1, v3, v4)), each face is a grid cell. Furthermore, the input of the separable convolution module is the grid cell sequence generated by these grid cells, such as the sequence [v1, v2, v3, v4, f1, f2], or a sequence containing only vertices [v1, v2, v3, v4], or a sequence containing only faces [f1, f2]. Among them, if the input of the separable convolution module is the sequence of vertices [v1, v2, v3, v4], the dimension of each grid cell is 3. If the input of the separable convolution module is the sequence of faces [f1, f2], the dimension of each grid cell is the number of vertices in the face multiplied by the index dimension of the vertex. In actual implementation, the input of the separable convolution module can be a grid cell sequence containing thousands or even millions of grid cells.

[0058] In step S320, attention fusion processing is performed on the grid abstract features to obtain global information features.

[0059] In an exemplary embodiment of the present disclosure, attention fusion processing is performed on the grid abstraction features, that is, multi-dimensional information aggregation is performed to aggregate the global information of the grid cell sequence and establish long-range dependencies, that is, global information is aggregated from the grid abstraction features to obtain global information features. For example, the attention fusion processing can be performed on the grid abstraction features through the multi-head attention module of the grid generation model to obtain global information features.

[0060] Among them, Figure 4 A schematic diagram of a multi-head attention module is shown, as Figure 4 , including a cascaded multi-layer perceptron layer and a multi-head attention layer. The multi-layer perceptron layer can include multiple linear transformation layers and activation layers (such as ReLU activation function layers) for performing non-linear transformation on the feature vector (grid abstraction feature) and increasing the expression ability of the feature. The multi-head attention layer can enable the grid generation model to focus on the information at different positions in the grid cell sequence and use this information for prediction.

[0061] Exemplarily, the processing process of the multi-head attention module can be represented by Formula 1:

[0062] MHSA(Q, K, V) = Concat(head1,..., head h )W O Formula 1

[0063] Among them, The calculation formula of Attention() is: d k is the dimension of the Key vector. The attention weights are normalized through the softmax function so that their sum is 1. Q = F a W Q is the query vector, K = F a W K is the key vector, V = F a W V is the value vector, F a is the grid abstraction feature, W Q , W K and W V are learnable weight matrices for mapping F a to the query space (Q) and the key-value space (K, V) respectively. W O is the output projection matrix, which is a learnable parameter. By splicing the output results of multiple attention heads and performing a linear transformation on the splicing result using the output projection matrix, the final output of the multi-head attention module is obtained, that is, the global information feature is obtained d m is the output dimension, and N' is the length of the grid abstraction feature.

[0064] This process performs attention fusion on the grid abstract features, captures the long-range dependencies between grid cells, aggregates global information to generate global information features, and guides the subsequent decoding process.

[0065] In step S330, grid information reconstruction is performed based on the global information features and the already generated grid cell features to obtain the current grid cell features.

[0066] In an exemplary embodiment of the present disclosure, the global information features carry the global information of the entire grid cell sequence and the long-range dependencies learned through the multi-head attention module, and can provide the overall shape and structure information of the grid to be generated. The already generated grid cell features refer to the accumulation of the grid cell features generated in the previous steps at each time step in the autoregressive generation operation. Specifically, when refining the attributes of the grid cell step by step through autoregressive generation, in addition to relying on the global information features, the output of the previous step is also used as the input for the current reconstruction time step.

[0067] Specifically, in the autoregressive process of the autoregressive model, the output of the model will be used as the input for the next step. For example, when generating an image, a part of the image is first generated, and then the already generated part is used as a condition to guide the generation of the subsequent image; in text generation, a part of the text is first generated, and then the already generated text is used as a condition to guide the generation of the subsequent text. The exemplary embodiment of the present disclosure can perform autoregressive generation through a causal convolution module. In the autoregressive process, when the causal convolution module is first executed (corresponding to the first grid cell of the grid sequence), the already generated grid cell features are an initialized zero vector. In subsequent time steps, the already generated grid cell features refer to the output of the causal convolution module in the previous step (i.e., the previous grid cell). The already generated grid cell features can provide local details and context information. By combining the global information features and the already generated grid cell features during network information reconstruction, the causal convolution module can generate coherent, detail-rich, and context-dependent current grid cell features.

[0068] Since only past information (already generated grid cell features) is utilized in the autoregressive generation process, the causality and temporal consistency of the results are ensured, and information leakage is avoided.

[0069] In step S340, the current grid cell features are feature-mapped to obtain grid cell attribute information, and based on the grid cell attribute information, rendering is performed to generate the target three-dimensional grid.

[0070] In an exemplary embodiment of the present disclosure, performing feature mapping on the current mesh cell features refers to mapping the implicit representation of the feature vector (the current mesh cell features) to the explicit representation of the mesh cell attributes. For vertices, the mesh cell attribute is the three-dimensional coordinates of the vertex, and for patches, the mesh cell attribute is the indices of the vertices included in the patch. This step of feature mapping processing can be performed by the decoding module of the mesh generation model.

[0071] Among them, the decoding module may include a linear network layer and an activation layer. The current mesh cell features are linearly transformed through the linear network layer to be converted into the vector space where the target attributes are located, and the result after the linear transformation is processed through the activation layer to obtain the mesh cell attribute information. The activation layer can adopt the Sigmoid activation function to limit the output attribute value between 0 and 1. Of course, the activation layer can also be adjusted according to actual needs. For example, if the output is the vertex index value, the softmax function can be used, and there is no limitation on this.

[0072] Exemplarily, if the mesh cell is a vertex, a feature vector with a dimension of 3 is mapped to three-dimensional coordinates (x, y, z). If the mesh cell is a patch, the current mesh cell features are mapped to the vertex indices that make up the patch. The three-dimensional coordinates or vertex indices can be used as attribute values for rendering and constructing a three-dimensional mesh model.

[0073] In the three-dimensional mesh generation method in the exemplary embodiment of the present disclosure, on the one hand, through hierarchical feature processing, the mesh abstract features, global information features, and current mesh cell features are sequentially extracted, realizing the adaptive feature extraction and information transmission of the mesh cell sequence, and realizing the refined control of the generated three-dimensional mesh. At the same time, in the autoregressive generation process, the global information features and the already generated mesh cell features are combined. That is, when generating each mesh cell, it is carried out under the dual guidance of the global information features and the previously generated mesh cell features. The global information features provide the overall information of the three-dimensional mesh, and the already generated mesh cell features provide local detail information and context information, so as to be able to generate a coherent, detail-rich, and context-dependent mesh structure, improving the quality of the generated three-dimensional mesh. On the other hand, the multi-granularity processing mechanism through hierarchical processing can efficiently process long mesh sequences. In addition, through the autoregressive generation process, it can be ensured that only the past information (i.e., the already generated mesh cell features) is relied on when generating the mesh, ensuring the causality of the output sequence while avoiding information leakage.

[0074] In an exemplary embodiment, a guiding control method for mesh generation is also provided, allowing flexible control of the shape, details, topological structure, etc. of the generated three-dimensional mesh by obtaining mesh generation guiding information. As Figure 5 shown, this guiding control method may include:

[0075] Step S510: Obtain mesh generation guidance information, which includes at least one of three-dimensional point cloud information, the number of mesh cells information, and mesh cell topology information.

[0076] Optionally, the mesh generation guidance information can be provided by the user. For example, the user can input the mesh generation guidance information through a graphical interface or an API (Application Programming Interface). For example, the user can obtain three-dimensional cell information from a three-dimensional scanner or a three-dimensional data acquisition tool and use it as the mesh generation guidance information. Optionally, three-dimensional point cloud information can also be sampled from an existing three-dimensional mesh model, such as using the FPS (Farthest Point Sampling) or Poisson (Poisson disk sampling) sampling algorithm to sample the mesh surface. Optionally, other forms of three-dimensional data provided by the user can also be converted into three-dimensional point cloud information, such as depth images, voxels, or implicit functions, etc. The exemplary embodiments of the present disclosure can select the method for obtaining the mesh generation guidance information according to actual needs, and there is no limitation on this.

[0077] Step S520: Encode and fuse the mesh generation guidance information to obtain mesh generation guidance features.

[0078] As Figure 6 shown in the structural schematic diagram of another mesh generation model, the mesh generation model may further include a parameter encoding module for encoding and fusing the mesh generation guidance information. Different types of mesh generation guidance information can be encoded separately, and the encoding results can be fused.

[0079] Exemplarily, the features of each three-dimensional point cloud are extracted through a multi-layer perceptron of the parameter encoding module, and then an activation layer is cascaded after the multi-layer perceptron to increase the non-linearity of the features. Finally, the obtained features are aggregated using a max pooling layer to obtain a global feature vector C S , to extract important features in the three-dimensional point cloud information. Another example is that the number of mesh cells information and the mesh cell topology information are encoded into a shape encoding vector C P through a multi-layer perceptron of the parameter encoding module. Similarly, an activation layer is connected after the multi-layer perceptron layer. Furthermore, the global feature vector C S can be connected to the shape encoding vector C P to perform a connection operation to obtain the mesh generation guidance feature Concat(C S , C P ).

[0080] Based on this, the grid generation guidance feature and the grid abstraction feature can be subjected to attention fusion, that is, the grid abstraction feature is subjected to attention fusion processing to obtain the global information feature, including: the grid abstraction feature and the grid generation guidance feature are subjected to attention fusion processing to obtain the global information feature. That is to say, since the global information feature incorporates the grid generation guidance feature, when processing network features, external grid generation guidance information can be considered simultaneously, thereby achieving flexible control over the shape, details, and topological structure of the generated three-dimensional grid, making the generation of the three-dimensional grid controllable and customizable.

[0081] Specifically, the process of subjecting the grid abstraction feature and the grid generation guidance feature to attention fusion processing to obtain the global information feature may include:

[0082] First, the grid abstraction feature is mapped to the query space to obtain a query vector, then the grid generation guidance feature is mapped to the key-value space to obtain a key vector, and the grid generation guidance feature is mapped to the key-value space to obtain a value vector, where the mapping matrices used to obtain the key vector and the value vector are different; finally, attention calculation is performed based on the query vector, the key vector, and the value vector to obtain the global information feature.

[0083] Among them, the grid abstraction feature and the grid generation guidance feature can be subjected to attention fusion processing through a multi-head attention module to obtain the global information feature. The query vector can be obtained through formula 2 below, the key vector can be obtained through formula 3, and the value vector can be obtained through formula 4:

[0084]

[0085] Among them, and are learnable parameters (mapping matrices), and different attention heads capture different cross-dependency relationships. Similar to the processing process without introducing grid generation guidance information, the output results of multiple attention heads are concatenated, and the concatenated result is linearly transformed using the output projection matrix to obtain the final output of the multi-head attention module, that is, the global information feature. The processing process of the multi-head attention module can be represented by formula 5:

[0086] where

[0087]

[0088] where the is also the output projection matrix, which is used to linearly transform the concatenated result of the output results of multiple attention heads to obtain the global information feature.

[0089] By introducing grid generation guiding features in the generation process of global information features, the generation process of the grid generation model can be controlled, the controllability of grid generation can be increased, and a three-dimensional grid that meets the user's expectations can be obtained.

[0090] In an exemplary embodiment, feature extraction on the grid cell sequence to obtain grid abstract features may include:

[0091] First, perform depthwise separable convolution processing on the grid cell sequence to obtain convolution features, then perform non-linear transformation processing on the convolution features to obtain processed convolution features, and perform layer normalization processing on the processed convolution features to obtain grid abstract features. Among them, this step can be executed through the separable convolution module of the grid generation model. The separable convolution module includes multiple cascaded depthwise separable convolution layers. Feature extraction on the grid cell sequence through the separable convolution module of the grid generation model to obtain grid abstract features may include:

[0092] Input the grid cell sequence into the first depthwise separable convolution layer to sequentially perform convolution operations using each depthwise separable convolution layer to obtain grid abstract features.

[0093] Among them, in each convolution process, a depthwise convolution kernel can be used to perform convolution on each input channel, and then a 1×1 convolution kernel can be used to perform a linear combination on all output channels.

[0094] Specifically, as Figure 7 shown is a schematic diagram of a separable convolution module. As Figure 7 , an activation layer and a layer normalization network are cascaded after each depthwise separable convolution layer. The process of obtaining grid abstract features can be represented by formula 6:

[0095] F a =LN(DSConv n (LeakyReLU(...LN(DSConv1(LeakyReLU(X)))...))) Formula 6

[0096] Among them, DSConv i represents the i-th depthwise separable convolution layer, LN represents the layer normalization operation, the formula of LeakyReLU can be LeakyReLU(x) = max(x, αx), α is the leakage coefficient, and the value range is (0, 1), and can be selected as 0.2. "..." means using a depthwise separable convolution layer (including the cascaded activation layer for non-linear transformation processing and the layer normalization network for layer normalization processing) for convolution each time, and repeating the execution until the output of the layer normalization layer corresponding to the last depthwise separable convolution layer is used as the grid abstract feature.

[0097] Using multiple cascaded separable convolutional layers can reduce the number of model parameters and the amount of computation. After each convolutional operation, the sequence length is reduced through normalization operations and a multi-granularity processing mechanism, which can accelerate the model training and inference processes.

[0098] In one exemplary embodiment, autoregressive generation based on the global information features and the generated grid cell features can obtain the current grid cell features, which may include:

[0099] First, for each time step of autoregressive generation, the global information features and the generated grid cell features corresponding to the time step are concatenated to obtain a concatenated feature. Then, the concatenated feature is deconvolved to obtain a deconvolved feature. Finally, the deconvolved feature is non-linearly transformed to obtain a transformed feature, and the transformed feature is layer-normalized to obtain the current grid cell feature. This process is repeated for each time step of autoregressive generation, that is, grid information reconstruction is performed based on the global information features and the generated grid cell features to obtain the current grid cell feature. Specifically:

[0100] First, for each time step of autoregressive generation, the global information features and the generated grid cell features corresponding to the time step are concatenated, and the obtained concatenated feature is deconvolved to obtain a deconvolved feature. Then, the deconvolved feature is non-linearly transformed, and the obtained transformed feature is layer-normalized to obtain the current grid cell feature.

[0101] Among them, a masking operation is performed on the output feature through deconvolution processing (that is, a masking matrix (also called a convolutional kernel) is used during the deconvolution process to mask the output feature), ensuring that the deconvolution operation only uses the feature information before the current reconstruction time step, thereby ensuring the causality of the generation process, ensuring the temporal consistency of the output sequence, and avoiding information leakage. For example, for one-dimensional data, the output feature at time t is only related to the features at time step t and before, and has nothing to do with future features. As Figure 8 shown in the figure is a schematic diagram of a causal convolution module. The causal convolution module includes at least one set of cascaded deconvolution layers, an activation network, and a layer normalization network.

[0102] The processing process of the causal convolution module can be represented by Equation 7:

[0103] R t =DRU t (F m ,R <t )=LN(CTransConv n (ReLU(...LN(CTransConv1(ReLU(Concat(F m ,R <t )))...))))

[0104] Formula 7

[0105] Wherein, R t is the current grid cell feature, DRU t refers to the causal convolution module, CTransConv i is the i-th deconvolution layer, ReLU is the activation network, LN is the layer normalization network, R <t represents the generated grid cell feature, F m is the global information feature, Concat(F m , R <t ) means concatenating the global information feature and the generated grid cell feature corresponding to the reconstruction time step. Similarly, after each deconvolution operation, a layer normalization operation is applied, which can accelerate the training and inference processes of the model. It should be understood that at each time step, the grid information reconstruction (i.e., autoregressive generation) is performed based on the global information feature and the generated grid cell feature corresponding to the reconstruction time step, and this is not listed one by one here.

[0106] By autoregressive generation based on the global information feature and the previously generated grid cell feature to reconstruct higher-resolution detailed information, a coherent, detailed and context-dependent grid structure can be generated. Since the deconvolution process only uses past information, the causality and temporal consistency of the output results are ensured.

[0107] In an exemplary embodiment, obtaining the grid cell sequence of the grid to be generated may be to obtain the grid cells of the grid to be generated and arrange the grid cells according to a predefined sorting rule to obtain the grid cell sequence. Wherein, the predefined sorting rule includes at least the vertex generation order and the in-plane vertex order, and the in-plane vertex order is used to indicate the consistent order of vertices within different patches. The consistent order means ensuring the vertex order within each patch is consistent, which is beneficial to maintaining the connectivity of the grid model, avoiding situations that do not conform to the topological structure, and ensuring that the grid can be correctly rendered.

[0108] Specifically, the in-plane vertex order refers to how to sort the vertices that make up a patch within a patch of a three-dimensional mesh model (e.g., a triangle). For example, in lexicographical order, the vertices can be sorted from smallest to largest according to their index values. For example, if the vertex indices of a triangle are [3, 1, 2], then the result of sorting in lexicographical order is to arrange the vertices as [1, 2, 3]. If the index values of multiple vertices are the same, then ensure that they are sorted according to their order in the original data, so as to ensure that the index of the first vertex is the smallest. The order of vertex indices determines the drawing direction of the patch. If the order of vertex indices is inconsistent, it may cause confusion between the front and back of the patch, thereby affecting the rendering effect. Therefore, ensuring that the vertex with the smallest index is in the front is usually to ensure the consistency of vertex indices and the correctness of rendering.

[0109] Exemplarily, for a triangular patch, its vertex indices are V1, V2, and V3, and the original indices of these vertices are: V1 = 5, V2 = 2, V3 = 1. If sorted in lexicographical order, the sorted result is V3 (index 1), V2 (index 2), V1 (index 5). The index of the first vertex V3 is the smallest, meeting the requirement that the index of the first vertex is the smallest. Therefore, the vertex order inside the triangular patch is [V3, V2, V1], corresponding to the indices [1, 2, 5]. Of course, the mesh cells can also be sorted according to other predefined sorting rules, and the exemplary embodiments of the present disclosure do not limit this.

[0110] Based on this, in an exemplary embodiment, an implementation manner for verifying timing integrity is further provided, that is, mapping the current mesh cell features to obtain mesh cell attribute information, and further including:

[0111] Verifying the mesh cell attribute information based on a predefined sorting rule, and adjusting the mesh cell attribute information of the target mesh cell that does not meet the predefined sorting rule.

[0112] Specifically, through the verification process, it can be ensured that the output grid cell attribute information follows the predefined sorting rules, avoiding situations that do not conform to physical laws, such as generating incorrect topological structures and other problems, thereby ensuring the validity of the grid model. Among them, the coordinate resampling method can be first adopted to perform coordinate resampling processing based on the vertex coordinates corresponding to the target grid cell to obtain the updated vertex coordinates that conform to the predefined sorting rules. For example, if the predicted vertex coordinates do not conform to the yzx scanning order, or the predicted face indices do not conform to the lexicographical order, resampling is performed based on the vertex coordinates corresponding to the target grid cell to generate new coordinates / indices (i.e., updated vertex coordinates) until the updated vertex coordinates that conform to the predefined sorting rules are obtained or the maximum number of samplings is reached. Optionally, the new coordinates / indices can be randomly selected from the original coordinates / indices. Optionally, the new coordinates / indices can be obtained by inserting new coordinates / indices between the original coordinates / indices, etc., and there is no limitation on this.

[0113] Among them, the lexicographical order means sorting the elements in the grid according to the rules of the lexicographical order. The lexicographical order is a sorting method based on character encoding or numerical size, which stipulates the order between elements. The lexicographical order of the exemplary embodiments of the present disclosure can be pre-constructed according to the relevant attributes of the grid cells of the historical three-dimensional grid (such as the coordinates of the grid nodes and the identifiers of the grid cells), and can be set according to specific task scenarios and requirements, and there is no limitation on this.

[0114] Still taking the above example as an example, during coordinate resampling, it is necessary to satisfy V1.y coordinate <= V2.y coordinate <= V3.y coordinate. If the y coordinates are the same, then sort according to the z coordinate. If they are still the same, then sort according to the x coordinate. For the vertices of the face, it is also necessary to ensure that the vertex with the smallest index is in the front.

[0115] Furthermore, if the resampling strategy is invalid and the updated vertex coordinates that conform to the predefined sorting rules cannot be obtained through coordinate resampling processing, then the sorting constraint loss is determined based on the vertex coordinates, and the target loss is determined according to the generation loss and the sorting constraint loss of the grid generation model, so that the grid generation model outputs the updated vertex coordinates that conform to the predefined sorting rules based on the target loss.

[0116] Specifically, the generation loss and the sorting constraint loss of the grid generation model can be expressed as Formula 8:

[0117] Loss = Lgeneration + λLordering Formula 8

[0118] Loss is the target loss, Lgeneration is the generation loss, which is used to measure the difference between the grid cell attribute information output by the grid generation model and the true grid attributes, such as the cross-entropy loss; Lordering is the sorting constraint loss, which is used to measure whether the model follows the predefined sorting rules, and λ is a hyperparameter used to balance the contributions of the generation loss and the sorting constraint loss. By adjusting λ, the model's preference for grid cell sorting can be adjusted, for example, it can be 0.1.

[0119] Among them, the sorting constraint loss can be represented by Equation 9:

[0120] Lordering = ∑i = 1n max(0, yi + 1 - yi) Equation 9

[0121] Among them, yi represents the coordinate of the vertex on the y-axis. If the y coordinates are the same, the sorting is performed according to the z and x coordinates. For the three-dimensional coordinates of xyz, the loss terms in the three directions can be added up.

[0122] By adding the sorting constraint loss to the loss of the grid generation model, the vertex update coordinates that conform to the predefined sorting rules are guided in the form of gradient penalty, which can further improve the accuracy of the grid cell attribute information and make it conform to the physical laws.

[0123] Furthermore, if the vertex update coordinates that conform to the predefined sorting rules cannot be obtained based on the target loss, at least one target network for predicting the target grid cell in the grid generation model is traced back, and the output features of the target network layer are adjusted to output the vertex update coordinates that conform to the predefined sorting rules according to the adjusted output features. Among them, a backtracking algorithm can be used to return to the previous prediction stage and regenerate the grid cell attribute information so that the re-predicted grid cell attribute information conforms to the predefined sorting rules. The backtracking algorithm can trace back layer by layer from the output layer to the input layer to correct the errors in the intermediate layer (at least one target network) or optimize the model parameters, including but not limited to the variational backtracking algorithm, the constraint-based backtracking algorithm, and the backtracking algorithm of inverse reinforcement learning.

[0124] In actual implementation, the resampling strategy and the gradient penalty strategy can be preferentially considered to balance the computational complexity and the generation quality of the grid.

[0125] The following uses a specific example to specifically illustrate the implementation method of the timing integrity verification. Among them, if the vertices for generating a triangular patch are V1, V2, and V3, the coordinate values of each vertex are output through the decoding module.

[0126] First, the coordinate values are verified based on the vertex generation order (taking the yzx scan order as an example). That is, it is necessary to ensure that the y coordinate of vertex V1 is less than the y coordinate of vertex V2, and the y coordinate of vertex V2 is less than the y coordinate of vertex V3. If the coordinate values of each vertex output by the decoding module are: V1.y = 0.2, V2.y = 0.4, V3.y = 0.1, then V3.y < V1.y, which violates the vertex generation order.

[0127] Secondly, resample the coordinate of V3.y. If the number of resampling times is set to 3, and after three resamplings, the updated coordinates of the vertex that still do not conform to the predefined sorting rule are not obtained, then the next step of gradient penalty is carried out.

[0128] Next, determine the sorting constraint loss based on the vertex coordinates, and determine the target loss according to the generation loss and sorting constraint loss of the mesh generation model, so as to prompt the mesh generation model to output the updated coordinates of the vertex that conform to the predefined sorting rule based on the target loss. The sorting constraint loss can be:

[0129] Lordering = max(0, V1.y - V2.y) + max(0, V2.y - V3.y)Lordering = max(0, V1.y - V2.y) + max(0, V2.y - V3.y)

[0130] Formula 10

[0131] Among them, the larger Lordering is, the more predefined sorting rules the model violates.

[0132] Furthermore, it is also necessary to ensure that the vertices within each patch are arranged in lexicographical order. If the index of V1 is 5, the index of V2 is 2, and the index of V3 is 1, then the order is [1, 2, 5]. If after the above steps, the predicted order is still [5, 2, 1], which does not conform to the lexicographical order, then resampling and gradient penalty are carried out again until the updated coordinates (indexes) of the vertex that conform to the predefined sorting rule are output, ensuring the consistency of the internal structure of the patch.

[0133] Again, if the above strategies all fail and the updated coordinates of the vertex that conform to the predefined sorting rule are not obtained, then backtrack to the target network layer that predicts the target grid cell, and adjust the output features of the target network layer to re-predict the vertices (or indexes) of the target grid cell based on the adjusted output features, so as to obtain the updated coordinates of the vertex.

[0134] Through the above steps, by strictly defining the generation order of the grid cell and using methods such as resampling, gradient penalty, and backtracking, it is ensured that the generated grid follows the predefined sorting rules (vertex generation order and in-patch vertex order), avoiding the generation of grid structures that do not conform to physical laws, thus ensuring the correctness of the topological structure of the grid.

[0135] The training and inference processes of the mesh generation model for the exemplary embodiments of the present disclosure will be described below.

[0136] First, prepare the dataset. Collect various types of 3D models, such as organic models, architectural models, human models, art models, etc. The 3D models in the dataset are in mesh format, such as OBJ (Object File Format), PLY (Polygon File Format), etc.

[0137] Second, mesh sampling. Sample 3D point cloud information from each 3D mesh. For example, use the FPS algorithm to uniformly sample on the surface of the 3D mesh, such as sampling 8192 points.

[0138] Next, mesh serialization. Arrange the mesh cells of each 3D model according to a predefined sorting rule to obtain a mesh cell sequence. Among them, the vertices can be sorted according to the yzx space scanning order, and the vertices inside the face are sorted according to the vertex index lexicographical order. The mesh cells in the mesh cell sequence can be the vertices of the 3D mesh or the faces of the 3D mesh. If the mesh cell is a vertex, record the x, y, z coordinates of the vertex. If the mesh cell is a face, record the vertex indices that make up the face.

[0139] Then, data preprocessing. The coordinates of the vertices can be normalized. For example, scale the vertex coordinates to the range of [-1, 1]. Data augmentation can also be performed, such as transformation operations like rotation, translation, and scaling, which are not limited here. Divide the processed dataset into a training set, a validation set, and a test set. For example, 80% is used for training, 10% is used for validation, and 10% is used for testing, which is also not limited here.

[0140] Furthermore, obtain mesh generation guidance information. The mesh generation guidance information can include at least one of 3D point cloud information, mesh cell quantity information, and mesh cell topology information. For example, receive the 3D point cloud data provided by the user. The mesh cell quantity information is a numerical value specifying the number of faces in the mesh to be generated. The mesh cell topology information can be topological preference information, a numerical value between (0, 1), used to specify the proportion of quadrilateral faces in the generated mesh. For example, "0" means only triangular faces are generated, and "1" means only quadrilateral faces are generated.

[0141] Moreover, the network architecture and model loss function of a mesh generation model will be exemplarily described.

[0142] Among them, the separable convolution module may include 4 depthwise separable convolution layers. The convolution kernel size of each layer is 3×3. The number of output channels of the depthwise convolution is 64, and the number of output channels of the pointwise convolution is 64, 128, 256, and 512 respectively. After each depthwise separable convolution layer, a cascade activation layer (such as the LeakyReLU activation function) and a layer normalization network are cascaded. The dimension of the input grid cell sequence is N×3, and the dimension of the output sequence is N′×512 (grid abstract feature), where N′ < N.

[0143] The multi-head attention module includes a multi-layer perceptron layer (such as one) and a multi-head attention layer (such as two). The hidden layer dimension of the multi-layer perceptron layer is, for example, 1024, and the ReLU is used as the activation function. The number of heads of the multi-head attention layer is, for example, 8, and the dimension of the attention head is, for example, 64. The input is the grid abstract feature with the dimension of N′×512, and the dimension of the output global information is N′×1024.

[0144] The causal convolution module includes at least one set of cascaded transposed convolution layers, an activation network, and a layer normalization network. For example, it includes 4 transposed convolution layers. The convolution kernel size of each layer is 3×3, and the number of output channels of each layer is 512, 256, 128, and 64 respectively. The ReLU activation function is used after each transposed convolution layer, and layer normalization is performed. The input dimension of the causal convolution module is the global information feature with the dimension of N′×1024, and the output dimension is N×dunit, where dunit is the attribute dimension of the grid cell. For example, if the grid cell is a vertex, then dunit = 3.

[0145] The parameter encoding module may include a spatial feature processing module and a shape parameter processing module. The spatial feature processing module may adopt a 4-layer fully connected network (MLP). The hidden layer dimension of each layer is 128, and the ReLU is used as the activation function. The final output dimension is 128. The shape parameter processing module adopts a 4-layer fully connected network (MLP). The hidden layer dimension of each layer is 64, and the ReLU is used as the activation function. The final output dimension is 64. The features output by the spatial feature processing module and the shape parameter processing module can be fused, and then fused with the grid feature to obtain the global information feature.

[0146] The sorting constraint loss of the grid generation model can be represented by the above formula 9, and the generation loss can adopt the cross-entropy loss, which can be represented by formula 11:

[0147] Lgeneration = -1n∑i = 1nyilog(pi)+(1 - yi)log(1 - pi) Formula 11 where, Lgeneration is the generation loss, yi is the true value, and pi is the predicted value of the model.

[0148] Furthermore, the objective loss of the grid generation model can be obtained through formula 8.

[0149] Finally, model training and inference are performed. The model is trained using the dataset in multiple training stages. For example, different learning rates and training sequence lengths are used in each stage. First, it is trained on shorter sequences, and then the sequence length is gradually increased. Finally, a trained mesh generation model is obtained, which can input mesh generation guidance information such as point clouds and the mesh cell sequence of the mesh to be generated, and make predictions through the trained mesh generation model to obtain mesh cell attribute information, so as to perform rendering based on the mesh cell attribute information to generate the target 3D mesh.

[0150] Among them, during the model inference process, to ensure that the generated 3D mesh has a correct structure, the mesh cell attribute information can also be verified based on a predefined sorting rule, and the mesh cell attribute information of the target mesh cells that do not meet the predefined sorting rule can be adjusted, which will not be elaborated here.

[0151] Exemplarily, the point cloud data obtained by sampling 8192 points from the three-dimensional surface using the FPS algorithm, and represented as the coordinates (x, y, z) of each point, with the number of mesh cells being 10000 and the quadrilateral face ratio set to 0.8, is input into the trained mesh generation model, and a mesh model containing 10000 faces is generated, where the ratio of quadrilateral faces is 0.8, that is, the generated structure can truly reflect information such as the shape of the input point cloud, and a 3D mesh with a reasonable structure and fine details is generated.

[0152] It should be noted that the specific details of each step have been described in the above exemplary embodiments and will not be elaborated here.

[0153] The 3D mesh generation method in the exemplary embodiments of the present disclosure, on the one hand, through a separable convolution module, a multi-head attention module, and a causal convolution module for hierarchical processing, realizes adaptive feature extraction and information transfer of the mesh cell sequence, and realizes refined control of the generated 3D mesh. At the same time, in each reconstruction time step, the global information feature and the features of the generated mesh cells are combined. That is, when each mesh cell is generated, it is carried out under the dual guidance of the global information feature and the features of the previously generated mesh cells. The global information feature provides the overall information of the 3D mesh, and the features of the generated mesh cells provide local detail information and context information, so as to be able to generate a coherent, detail-rich, and context-dependent mesh structure, improving the quality of the generated 3D mesh. On the other hand, the multi-granularity processing mechanism for hierarchical processing can efficiently process long mesh sequences, and the separable convolution module reduces the computational amount and complexity, thereby improving the generation efficiency while ensuring the generation quality. In addition, through the causal convolution module, it is possible to rely only on past information (i.e., the features of the generated mesh cells) when generating the mesh, ensuring the causality of the output sequence while avoiding information leakage. On the other hand, by verifying and adaptively adjusting the mesh cell attribute information based on a predefined sorting rule, structural anomalies and topological errors during the generation process can be avoided, and it has high robustness.

[0154] In the exemplary embodiments of the present disclosure, a 3D mesh generation device is also provided. Referring to Figure 9 as shown, the 3D mesh generation device 900 may include an information acquisition module 910, a feature extraction module 920, an information reconstruction module 930, and a rendering processing module 940. Specifically:

[0155] The information acquisition module 910 is configured to acquire the mesh cell sequence of the mesh to be generated, and perform feature extraction on the mesh cell sequence through a mesh generation model to obtain mesh abstract features; the feature extraction module 920 is configured to perform attention fusion processing on the mesh abstract features to obtain global information features; the information reconstruction module 930 is configured to perform autoregressive generation based on the global information features and the features of the generated mesh cells to obtain the current mesh cell features; the rendering processing module 940 is configured to perform feature mapping on the current mesh cell features to obtain mesh cell attribute information, and render based on the mesh cell attribute information to generate a target 3D mesh.

[0156] In an exemplary embodiment of the present disclosure, the information acquisition module 910 is further configured to perform: obtaining grid generation guidance information, where the grid generation guidance information includes at least one of three-dimensional point cloud information, the number of grid cell information, and grid cell topology information; encoding and fusing the grid generation guidance information to obtain grid generation guidance features; the feature extraction module 920 is configured to perform: performing attention fusion processing on the grid abstract features and the grid generation guidance features to obtain global information features.

[0157] In an exemplary embodiment of the present disclosure, the feature extraction module 920 is configured to perform: mapping the grid abstract features to a query space to obtain query vectors; mapping the grid generation guidance features to a key-value space to obtain key vectors, and mapping the grid generation guidance features to a key-value space to obtain value vectors, where the mapping matrices used to obtain the key vectors and the value vectors are different; performing attention calculation based on the query vectors, the key vectors, and the value vectors to obtain global information features.

[0158] In an exemplary embodiment of the present disclosure, the information acquisition module 910 is further configured to perform: performing depthwise separable convolution processing on the grid cell sequence to obtain convolution features; performing non-linear transformation processing on the convolution features to obtain processed convolution features, and performing layer normalization processing on the processed convolution features to obtain grid abstract features.

[0159] In an exemplary embodiment of the present disclosure, the information reconstruction module 930 is configured to perform: for each autoregressive generation time step, concatenating the global information features and the generated grid cell features corresponding to the time step to obtain concatenated features, and performing deconvolution processing on the concatenated features to obtain deconvolution features; performing non-linear conversion on the deconvolution features to obtain conversion features, and performing layer normalization processing on the conversion features to obtain the current grid cell features.

[0160] In an exemplary embodiment of the present disclosure, the information acquisition module 910 is configured to perform: obtaining grid cells of the grid to be generated, and arranging the grid cells according to a predefined sorting rule to obtain a grid cell sequence; the rendering processing module 940 is further configured to perform: verifying the grid cell attribute information based on the predefined sorting rule, and adjusting the grid cell attribute information of the target grid cells that do not meet the predefined sorting rule; where the predefined sorting rule includes at least a vertex generation order and an in-plane vertex order, and the in-plane vertex order is used to indicate the consistent order of vertices within different patches.

[0161] In an exemplary embodiment of the present disclosure, the rendering processing module 940 is further configured to perform: performing coordinate resampling processing based on the vertex coordinates corresponding to the target grid cell to obtain vertex update coordinates that conform to a predefined sorting rule; if vertex update coordinates that conform to the predefined sorting rule are not obtained through the coordinate resampling processing, determining a sorting constraint loss based on the vertex coordinates, and determining a target loss based on the generation loss of the grid generation model and the sorting constraint loss, so that the grid generation model outputs vertex update coordinates that conform to the predefined sorting rule based on the target loss; wherein, the generation loss is used to measure the difference between the grid cell attribute information output by the grid generation model and the true grid attributes.

[0162] In an exemplary embodiment of the present disclosure, the rendering processing module 940 is further configured to perform: if vertex update coordinates that conform to the predefined sorting rule are not obtained based on the target loss, then backtracking to the target network layer that predicts the target grid cell; adjusting the output features of the target network layer to output vertex update coordinates that conform to the predefined sorting rule according to the adjusted output features.

[0163] Since the detailed content of each functional module of the three-dimensional grid generation device in the exemplary embodiment of the present disclosure has been described in the exemplary embodiment of the above three-dimensional grid generation method, it will not be repeated here.

[0164] It should be noted that although several modules or units of the three-dimensional grid generation device are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0165] The exemplary embodiment of the present disclosure also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the above three-dimensional grid generation method.

[0166] In one embodiment, the computer program product may be a tangible product containing a computer program, such as a computer-readable storage medium storing the computer program. The readable storage medium may be a storage medium based on signals such as electricity, magnetism, light, electromagnetic, infrared, etc., including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory (Flash), mechanical hard disk (HDD), solid state drive (SSD), and so on. Exemplarily, the computer program product may be implemented as a non-volatile storage medium storing the computer program, such as read-only memory, Nand Flash, etc.

[0167] In one embodiment, a computer program product can be an intangible product containing a computer program. Exemplarily, the computer program product can be implemented as a virtual digital product, such as an executable file storing the computer program, a digital file like an installation package, etc.

[0168] The code of the computer program can be written in one or more programming languages. Examples of programming languages include C, Java, C++, etc. The program code can be executed entirely on the user's computing device, or partially on the user's computing device, or executed as an independent software package, or partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case involving a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, such as a local area network (LAN), a wide area network (WAN), etc., or can be connected to an external computing device (e.g., through an Internet connection provided by an operator).

[0169] The computer program can be carried or transmitted by signals such as electricity, magnetism, light, electromagnetic, infrared, etc. The electronic device can convert the signal carrying the computer program into a digital signal and then run the computer program. When the computer program runs on the electronic device, its code is used to cause the electronic device to execute (more specifically, can cause the processor of the electronic device to execute) the method steps of various exemplary embodiments of the present disclosure. For example, it can execute the above-mentioned three-dimensional mesh generation method, which includes the following steps:

[0170] Obtain a sequence of mesh cells of the mesh to be generated, and perform feature extraction on the sequence of mesh cells through a mesh generation model to obtain mesh abstract features; perform attention fusion processing on the mesh abstract features to obtain global information features; perform autoregressive generation based on the global information features and the features of the already generated mesh cells to obtain the features of the current mesh cell; perform feature mapping on the features of the current mesh cell to obtain mesh cell attribute information, and render based on the mesh cell attribute information to generate the target three-dimensional mesh.

[0171] In an exemplary embodiment of the present disclosure, the method further includes: obtaining mesh generation guidance information, where the mesh generation guidance information includes at least one of three-dimensional point cloud information, mesh cell quantity information, and mesh cell topology information; performing encoding fusion on the mesh generation guidance information to obtain mesh generation guidance features; performing attention fusion processing on the mesh abstract features to obtain global information features, including: performing attention fusion processing on the mesh abstract features and the mesh generation guidance features to obtain global information features.

[0172] In an exemplary embodiment of the present disclosure, an attention fusion process is performed on the grid abstract feature and the grid generation guidance feature to obtain the global information feature, including: mapping the grid abstract feature to the query space to obtain a query vector; mapping the grid generation guidance feature to the key-value space to obtain a key vector, and mapping the grid generation guidance feature to the key-value space to obtain a value vector, where different mapping matrices are used to obtain the key vector and the value vector; performing attention calculation according to the query vector, the key vector, and the value vector to obtain the global information feature.

[0173] In an exemplary embodiment of the present disclosure, feature extraction is performed on the grid cell sequence to obtain the grid abstract feature, including: performing depthwise separable convolution processing on the grid cell sequence to obtain a convolution feature; performing a non-linear transformation process on the convolution feature to obtain the processed convolution feature, and performing layer normalization processing on the processed convolution feature to obtain the grid abstract feature.

[0174] In an exemplary embodiment of the present disclosure, autoregressive generation is performed according to the global information feature and the generated grid cell feature to obtain the current grid cell feature, including: for each time step of autoregressive generation, concatenating the global information feature and the generated grid cell feature corresponding to the time step to obtain a concatenated feature, and performing deconvolution processing on the concatenated feature to obtain a deconvolution feature; performing non-linear conversion on the deconvolution feature to obtain a conversion feature, and performing layer normalization processing on the conversion feature to obtain the current grid cell feature.

[0175] In an exemplary embodiment of the present disclosure, obtaining the grid cell sequence of the grid to be generated includes: obtaining the grid cells of the grid to be generated, and arranging the grid cells according to a predefined sorting rule to obtain the grid cell sequence; mapping the current grid cell feature to obtain grid cell attribute information, further including: verifying the grid cell attribute information based on the predefined sorting rule, and adjusting the grid cell attribute information of the target grid cell that does not meet the predefined sorting rule; where the predefined sorting rule at least includes the vertex generation order and the in-plane vertex order, and the in-plane vertex order is used to indicate the consistent order of vertices within different patches.

[0176] In an exemplary embodiment of the present disclosure, the grid cell attribute information is verified based on a predefined sorting rule, and the grid cell attribute information of the target grid cell that does not meet the predefined sorting rule is adjusted, including: performing coordinate resampling processing based on the vertex coordinates corresponding to the target grid cell to obtain vertex update coordinates that conform to the predefined sorting rule; if vertex update coordinates that conform to the predefined sorting rule are not obtained through the coordinate resampling processing, determining a sorting constraint loss based on the vertex coordinates, and determining a target loss based on the generation loss of the grid generation model and the sorting constraint loss, so that the grid generation model outputs vertex update coordinates that conform to the predefined sorting rule based on the target loss; wherein, the generation loss is used to measure the difference between the grid cell attribute information output by the grid generation model and the true grid attributes.

[0177] In an exemplary embodiment of the present disclosure, the method further includes: if vertex update coordinates that conform to the predefined sorting rule are not obtained based on the target loss, then backtracking to the target network layer that predicts the target grid cell; adjusting the output features of the target network layer to obtain adjusted output features, so as to predict vertex update coordinates that conform to the predefined sorting rule based on the adjusted output features.

[0178] In addition, in an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided. Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0179] The following refers to Figure 10 to describe the electronic device 1000 according to this embodiment of the present disclosure. Figure 10 The shown electronic device 1000 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0180] As Figure 10 shown, the electronic device 1000 is presented in the form of a general-purpose computing device. The components of the electronic device 1000 may include, but are not limited to: at least one of the above processing units 1010, at least one of the above storage units 1020, a bus 1030 connecting different system components (including the storage unit 1020 and the processing unit 1010), and a display unit 1040.

[0181] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 1010, so that the processing unit 1010 executes the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section of this specification. For example, the processing unit 1010 can execute the following steps:

[0182] Obtain the grid cell sequence of the grid to be generated, and perform feature extraction on the grid cell sequence through a grid generation model to obtain grid abstract features; perform attention fusion processing on the grid abstract features to obtain global information features; perform autoregressive generation based on the global information features and the generated grid cell features to obtain the current grid cell features; perform feature mapping on the current grid cell features to obtain grid cell attribute information, so as to perform rendering based on the grid cell attribute information to generate a target three-dimensional grid.

[0183] In an exemplary embodiment of the present disclosure, the method further includes: obtaining grid generation guidance information, where the grid generation guidance information includes at least one of three-dimensional point cloud information, grid cell quantity information, and grid cell topology information; performing encoding fusion on the grid generation guidance information to obtain grid generation guidance features; performing attention fusion processing on the grid abstract features to obtain global information features, including: performing attention fusion processing on the grid abstract features and the grid generation guidance features to obtain global information features.

[0184] In an exemplary embodiment of the present disclosure, performing attention fusion processing on the grid abstract features and the grid generation guidance features to obtain global information features includes: mapping the grid abstract features to a query space to obtain a query vector; mapping the grid generation guidance features to a key-value space to obtain a key vector, and mapping the grid generation guidance features to a key-value space to obtain a value vector, where the mapping matrices used to obtain the key vector and the value vector are different; performing attention calculation based on the query vector, the key vector, and the value vector to obtain global information features.

[0185] In an exemplary embodiment of the present disclosure, performing feature extraction on the grid cell sequence to obtain grid abstract features includes: performing depthwise separable convolution processing on the grid cell sequence to obtain convolution features; performing non-linear transformation processing on the convolution features to obtain processed convolution features, and performing layer normalization processing on the processed convolution features to obtain grid abstract features.

[0186] In an exemplary embodiment of the present disclosure, autoregressive generation is performed based on the global information feature and the generated grid cell features to obtain the current grid cell feature, including: for each time step of autoregressive generation, concatenating the global information feature and the generated grid cell feature corresponding to the time step to obtain a concatenated feature, and performing deconvolution processing on the concatenated feature to obtain a deconvolution feature; performing a non-linear transformation on the deconvolution feature to obtain a transformed feature, and performing layer normalization processing on the transformed feature to obtain the current grid cell feature.

[0187] In an exemplary embodiment of the present disclosure, obtaining the grid cell sequence of the grid to be generated includes: obtaining the grid cells of the grid to be generated, and arranging the grid cells according to a predefined sorting rule to obtain a grid cell sequence; performing feature mapping on the current grid cell feature to obtain grid cell attribute information, further including: verifying the grid cell attribute information based on the predefined sorting rule, and adjusting the grid cell attribute information of the target grid cell that does not meet the predefined sorting rule; wherein the predefined sorting rule includes at least a vertex generation order and an in-plane vertex order, and the in-plane vertex order is used to indicate the consistent order of vertices within different patches.

[0188] In an exemplary embodiment of the present disclosure, verifying the grid cell attribute information based on the predefined sorting rule and adjusting the grid cell attribute information of the target grid cell that does not meet the predefined sorting rule includes: performing coordinate resampling processing on the vertex coordinates corresponding to the target grid cell to obtain vertex update coordinates that conform to the predefined sorting rule; if vertex update coordinates that conform to the predefined sorting rule are not obtained through the coordinate resampling processing, determining a sorting constraint loss based on the vertex coordinates, and determining a target loss based on the generation loss of the grid generation model and the sorting constraint loss, so that the grid generation model outputs vertex update coordinates that conform to the predefined sorting rule based on the target loss; wherein the generation loss is used to measure the difference between the grid cell attribute information output by the grid generation model and the real grid attributes.

[0189] In an exemplary embodiment of the present disclosure, the method further includes: if vertex update coordinates that conform to the predefined sorting rule are not obtained based on the target loss, then backtracking to the target network layer that predicts the target grid cell; adjusting the output feature of the target network layer to obtain an adjusted output feature, so as to predict vertex update coordinates that conform to the predefined sorting rule based on the adjusted output feature.

[0190] The storage unit 1020 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 1021 and / or a cache storage unit 1022, and may further include a read-only storage unit (ROM) 1023.

[0191] The storage unit 1020 may also include a program / utilities 1024 having a set (at least one) of program modules 1025. Such program modules 1025 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0192] The bus 1030 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.

[0193] The electronic device 1000 may also communicate with one or more external devices 1100 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 1000, and / or may communicate with any device that enables the electronic device 1000 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be through the input / output (I / O) interface 1050. Also, the electronic device 1000 may communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 1060. As shown in the figure, the network adapter 1060 communicates with other modules of the electronic device 1000 through the bus 1030. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1000, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0194] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by the way of software combined with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0195] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed, for example, synchronously or asynchronously in multiple modules.

[0196] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

Claims

1. A three-dimensional grid generation method, characterized in that: include: Obtaining a grid cell sequence of a grid to be generated, and extracting features of the grid cell sequence through a grid generation model to obtain grid abstract features; Performing attention fusion processing on the grid abstract features to obtain global information features; Perform autoregression generation based on the global information features and the generated grid unit features to obtain current grid unit features; The current grid unit features are feature mapped to obtain grid unit attribute information, so as to perform rendering based on the grid unit attribute information to generate a target three-dimensional grid.

2. The method according to claim 1, characterized in that The method further comprises: Acquire mesh generation guidance information, wherein the mesh generation guidance information includes at least one of three-dimensional point cloud information, mesh unit quantity information, and mesh unit topology information; Encoding and fusing the grid generation guidance information to obtain a grid generation guidance feature; The attention fusion processing is performed on the grid abstract features to obtain global information features, including: The grid abstract feature and the grid generation guidance feature are subjected to attention fusion processing to obtain the global information feature.

3. The method according to claim 2, characterized in that The step of performing attention fusion processing on the grid abstract feature and the grid generation guidance feature to obtain the global information feature includes: Mapping the grid abstract feature to a query space to obtain a query vector; Mapping the grid generation guidance feature to a key value space to obtain a key vector, and mapping the grid generation guidance feature to a key value space to obtain a value vector, wherein the mapping matrices used to obtain the key vector and the value vector are different; Attention calculation is performed according to the query vector, the key vector and the value vector to obtain the global information feature.

4. The method according to claim 1, characterized in that: The step of extracting features from the grid unit sequence to obtain grid abstract features includes: Performing depth-wise separable convolution processing on the grid unit sequence to obtain convolution features; The convolution feature is subjected to nonlinear transformation processing to obtain processed convolution feature, and the processed convolution feature is subjected to layer normalization processing to obtain the grid abstract feature.

5. The method according to claim 1, characterized in that The step of performing autoregressive generation according to the global information feature and the generated grid unit feature to obtain the current grid unit feature includes: For each time step generated by the autoregression, the global information feature and the generated grid unit feature corresponding to the time step are spliced ​​to obtain a spliced ​​feature, and the spliced ​​feature is deconvolved to obtain a deconvolution feature; The deconvolution feature is nonlinearly transformed to obtain a transformation feature, and the transformation feature is layer-normalized to obtain the current grid unit feature.

6. The method according to claim 1, characterized in that The step of obtaining a grid unit sequence of a grid to be generated includes: Obtaining grid cells of a grid to be generated, and arranging the grid cells according to a predefined sorting rule to obtain the grid cell sequence; The performing feature mapping on the current grid unit feature to obtain grid unit attribute information further includes: Verifying the grid cell attribute information based on the predefined sorting rule, and adjusting the grid cell attribute information of the target grid cells that do not meet the predefined sorting rule; The predefined sorting rules include at least a vertex generation order and an intra-face vertex order, and the intra-face vertex order is used to indicate a consistent order of vertices in different facets.

7. The method according to claim 6, characterized in that The verifying the grid unit attribute information based on the predefined sorting rule and adjusting the grid unit attribute information of the target grid unit that does not meet the predefined sorting rule includes: Performing coordinate resampling processing based on the vertex coordinates corresponding to the target grid unit to obtain vertex update coordinates that conform to the predefined sorting rule; If the vertex update coordinates that meet the predefined sorting rule are not obtained through the coordinate resampling process, a sorting constraint loss is determined based on the vertex coordinates, and a target loss is determined according to the generation loss of the mesh generation model and the sorting constraint loss, so that the mesh generation model outputs the vertex update coordinates that meet the predefined sorting rule based on the target loss; The generation loss is used to measure the difference between the grid unit attribute information output by the grid generation model and the actual grid attributes.

8. The method according to claim 7, characterized in that The method further comprises: If the vertex update coordinates that meet the predefined sorting rule cannot be obtained based on the target loss, backtracking to the target network layer that predicts the target grid unit; The output features of the target network layer are adjusted to obtain adjusted output features, so as to predict the updated coordinates of the vertices that meet the predefined sorting rule according to the adjusted output features.

9. A three-dimensional grid generation device, characterized in that: The device comprises: An information acquisition module is used to acquire a grid cell sequence of a grid to be generated, and to extract features of the grid cell sequence through a grid generation model to obtain grid abstract features; A feature extraction module is used to perform attention fusion processing on the grid abstract features to obtain global information features; An information reconstruction module, used for performing autoregressive generation according to the global information features and the generated grid unit features to obtain current grid unit features; The rendering processing module is used to perform feature mapping on the current grid unit features to obtain grid unit attribute information, so as to perform rendering based on the grid unit attribute information to generate a target three-dimensional grid.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

11. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 8 by executing the executable instructions.