Model generation, model training methods, equipment, storage media and program products

Through the combination of autoregressive model and encoder, multi-scale marker maps are generated, which solves the problem of fixed resolution of the existing 3D mesh model, and realizes the flexibility of 3D mesh models with different resolutions, reducing costs and improving generation efficiency.

CN119919605BActive Publication Date: 2025-08-22TAOBAO CHINA SOFTWARE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510397757.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-08-22
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

The existing 3D mesh model generation scheme based on artificial intelligence has fixed resolution, which is difficult to meet the display needs of different 3D scenes, and the existing methods have high labor and time costs.

Method used

Using a combination of autoregressive model and encoder, a geometric model is obtained by obtaining feature information of geometric model samples, marking diagrams of multiple scales are generated, and autoregressive generation operations are performed using autoregressive models to decode the geometric model that meets the needs according to the target resolution.

Benefits of technology

It realizes flexible generation of 3D mesh models with different resolutions, meets the needs of multiple scenarios, reduces labor and time costs, and improves generation efficiency and model resolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919605B_ABST
    Figure CN119919605B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a model generation, model training method, device, storage medium and program product. In the model generation method, the autoregressive model can perform an autoregressive generation operation of spatial features based on the feature information of the input geometric model description data to accurately output multiple label maps of different scales. Among them, the scale of the label map is positively correlated with the resolution of the geometric model, so that the label maps of multiple different scales can not only express the global structural features of the geometric model sample, but also express the local detail features of the geometric model. Therefore, it is convenient to select the target scale label map corresponding to the target resolution for decoding according to the needs, and obtain a geometric model that meets the expected resolution, thereby flexibly meeting the needs of different scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of human intelligence technology, and in particular to a model generation, model training method, device, storage medium and program product. Background Art

[0002] A three-dimensional (3D) mesh model is a mesh-like geometric structure composed of multiple vertices, edges, and faces, which is used to define the three-dimensional shape and structural details of an object. The 3D mesh model is one of the important forms of expression of 3D assets and a universal format natively supported by a variety of 3D image software and hardware. It is widely used in the film and television, design, e-commerce, and gaming industries to achieve high-quality visual effects and interactive experiences. For example, in the e-commerce field, by displaying the three-dimensional effects of products, the details and characteristics of the products can be displayed from all directions and angles, thereby providing buyers with more comprehensive product information. This method not only enhances the shopping experience, but also effectively helps buyers make more informed decisions, thereby increasing purchase conversion rates. Currently, the market usually relies on artists to manually create high-quality 3D mesh models. The manpower and time costs required for this manual creation method are very large.

[0003] With the continuous advancement of artificial intelligence technology in the fields of language, images, and video, a solution has emerged that uses AI models to generate 3D mesh models. This solution can automatically generate 3D mesh models based on input images and text, reducing labor and time costs. However, existing AI-based solutions have a relatively fixed resolution for the generated 3D mesh models, which may not meet the display requirements of some 3D scenes. Therefore, a new solution is needed. Summary of the Invention

[0004] Embodiments of the present application provide a model generation and model training method, device, storage medium, and program product for flexibly generating geometric models of different resolutions.

[0005] An embodiment of the present application provides a model generation method, including: obtaining geometric model description data in response to a model generation operation; performing feature extraction on the geometric model description data to obtain first feature information; utilizing an autoregressive model to perform multiple autoregressive generation operations of spatial features based on the first feature information to obtain first labeled maps of multiple scales, wherein spatial features generated by different autoregressive generation operations correspond to first labeled maps of different scales; the scale of the first labeled map is positively correlated with the resolution of the geometric model; and decoding the first labeled map of at least one target scale among the multiple scales based on at least one target resolution to obtain at least one target geometric model.

[0006] An embodiment of the present application also provides a model training method, including: obtaining a geometric model sample and a corresponding model description data sample; using an encoder to encode the spatial data corresponding to the geometric model sample to obtain reference label maps of multiple scales, and the reference label maps of multiple scales are used to express the spatial features of the geometric model sample at multiple resolutions; performing feature extraction on the model description data sample to obtain target feature information, and using an autoregressive model to perform multiple autoregressive generation operations of spatial features according to the target feature information to obtain target label maps of multiple scales, the spatial features output by different autoregressive generation operations correspond to target label maps of different scales, the target label maps of the multiple scales correspond to the multiple resolutions, and the scale of the target label map is positively correlated with the resolution of the geometric model; training the autoregressive model based on the difference between the reference label maps of the multiple scales and the target label maps of the multiple scales.

[0007] An embodiment of the present application also provides an electronic device, comprising: a memory and a processor; the memory is used to store one or more computer instructions; the processor is used to execute the one or more computer instructions to: execute the steps in the method provided in the embodiment of the present application.

[0008] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the steps of the method provided in the embodiment of the present application.

[0009] An embodiment of the present application further provides a computer program product, comprising: a computer program / instructions, which, when executed by a processor, can implement the steps of the method provided in the embodiment of the present application.

[0010] In the embodiments of the present application, the autoregressive model can perform an autoregressive generation operation on spatial features based on the feature information of the input geometric model description data to accurately output multiple labeled maps of different scales. The scale of the labeled map is positively correlated with the resolution of the geometric model, allowing the labeled maps of different scales to express not only the global structural characteristics of the geometric model sample but also the local detail characteristics of the geometric model. This facilitates the selection of a labeled map of a target scale corresponding to the target resolution for decoding, resulting in a geometric model that meets the desired resolution, thereby flexibly meeting the needs of different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0012] Figure 1A flowchart of a model training method provided by an exemplary embodiment of the present application;

[0013] Figure 2 A schematic diagram of a processing flow of an encoder provided as an exemplary embodiment of the present application;

[0014] Figure 3 A schematic diagram of a processing flow of an autoregressive model provided for an exemplary embodiment of the present application;

[0015] Figure 4 A flowchart of a model generation method provided by an exemplary embodiment of the present application;

[0016] Figure 5 A schematic structural diagram of an electronic device provided as an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0017] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0018] The terms used in the examples of this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a," "the," and "the" used in the examples of this application and the appended claims are also intended to include plural forms, and unless the context clearly indicates otherwise, "a plurality" generally includes at least two, but does not exclude the inclusion of at least one.

[0019] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.

[0020] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a product or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such a product or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the product or system comprising the element.

[0021] Some solutions that use AI models to generate 3D mesh models typically sample a fixed number of surface point clouds from the geometric model due to hardware resource constraints, and then perform feature compression during the model training process. This approach results in limited resolution and easily loses structural information. Other solutions rely on diffusion models to generate the geometric model, but these models suffer from slow training and inference stages, hindering efficiency.

[0022] In response to the above technical problems, a solution is provided in some embodiments of the present application. The technical solutions provided in each embodiment of the present application are described in detail below with reference to the accompanying drawings.

[0023] Figure 1 This is a flow chart of a model training method provided by an exemplary embodiment of the present application. The model training method is used to train an autoregressive model. The autoregressive model is used to output a label map of multiple scales based on the target feature information of the input geometric model description data. The label map of multiple scales is used to decode to obtain a multi-resolution geometric model, so as to realize a solution of automatically generating a geometric model based on the input geometric model description data. Figure 1 As shown, the method may mainly include the following steps:

[0024] Step 101: Obtain a geometric model sample and a corresponding model description data sample.

[0025] Step 102: Encode the spatial data corresponding to the geometric model sample using an encoder to obtain reference marker images of multiple scales, where the reference marker images of multiple scales are used to express the spatial features of the geometric model sample at multiple resolutions.

[0026] Step 103: extract features from the model description data samples to obtain target feature information, and use the autoregressive model to perform multiple autoregressive generation operations of spatial features based on the target feature information to obtain target label maps of multiple scales. The spatial features output by different autoregressive generation operations correspond to target label maps of different scales. The target label maps of the multiple scales correspond to the multiple resolutions, and the scale of the target label map is positively correlated with the resolution of the geometric model.

[0027] Step 104: Train the autoregressive model according to the differences between the reference labeled images at the multiple scales and the target labeled images at the multiple scales.

[0028] The model training method provided in this embodiment can be executed by an electronic device, which can be a terminal device or a server device, and this embodiment does not impose any restrictions.

[0029] In step 101, the geometric model sample can be a 3D model of any real object or a 3D model of any virtual object. The geometric model sample and the model description data sample are paired training samples, and the model description data sample is used to describe the geometric model sample. The model description data sample may include data of one or more modalities, for example, including but not limited to at least one of text description data, voice description data, point cloud data and image description data. Among them, the image description data may include pictures of the geometric model sample at one or more different perspectives. For example, when the geometric model sample is a 3D model of any object, its model description data sample may include images of the object at different perspectives, such as a front view, a side view, a top view, etc.

[0030] The geometric model samples used in this embodiment are 3D mesh models extracted from a pre-built 3D asset library, each with a certain texture. In this embodiment, the pre-built 3D asset library can include millions or more assets. The 3D models in the 3D asset library can be high-precision 3D mesh models obtained through manual adjustment or semi-automatic adjustment using an algorithm. This large number of high-precision samples enables large-scale model parameter training, significantly improving the resolution of the generated 3D models.

[0031] In step 102, an encoder is used to encode the input geometric model samples to obtain token maps at multiple scales. In this embodiment, a token map is a feature map obtained by encoding the input data into discrete units (tokens) and mapping the discrete units to a latent feature space. In this embodiment, to distinguish the token maps generated in different steps, the token map of the geometric model sample output by the encoder is described as a reference token map. This reference token map is used as a supervisory signal in the process of training the autoregressive model. The scale of the token map refers to the number of discrete units contained in the token map. The larger the scale of the token map, the more discrete units it contains. The smaller-scale token map is obtained by encoding the spatial data of the geometric model sample with a higher data compression rate, and is used to express the global structural features of the geometric model using fewer discrete units. The larger-scale token map is obtained by encoding the spatial data of the geometric model sample with a lower data compression rate, and is used to express the local detail features of the geometric model using more discrete units. In other words, the scale of a label map is proportional to the amount of information it conveys about the spatial features, resulting in a positive correlation between the scale of the label map and the resolution of the geometric model. Smaller label map scales result in coarser details and lower resolution, while larger label map scales allow for finer details and higher resolution. Reference label maps at multiple scales can be used to represent the spatial features of geometric model samples at various resolutions.

[0032] In some optional embodiments, the encoder can be a variational autoencoder (VAE). A variational autoencoder is a generative model that combines a probabilistic graphical model and a neural network. In this embodiment, the encoder is primarily used to learn a latent representation of the sample data (i.e., a reference labeled map). This latent representation is used as a supervisory signal to guide the training process of the autoregressive model.

[0033] In some optional embodiments, when encoding the spatial data corresponding to the geometric model sample using an encoder, the encoder can be used to organize the spatial data corresponding to the geometric model sample into multiple levels of spatial nodes, where the spatial nodes in different levels represent different spatial ranges, and reference marker maps of multiple scales are obtained based on the spatial nodes of the multiple levels. When the spatial ranges represented by the spatial nodes in different levels are different, it means that the spatial nodes of different levels have different compression rates for the spatial data corresponding to the geometric model sample, thereby enabling the reference marker maps of multiple scales to express the spatial features of the geometric model sample at multiple resolutions. An exemplary explanation will be given below.

[0034] Optionally, the encoder may organize the spatial data corresponding to the geometric model sample into multiple levels of spatial nodes according to the target structure. The target structure used to organize the spatial data is a hierarchical data organization structure, which may include but is not limited to a tree structure or a structured data description structure. The tree structure is a nonlinear data structure consisting of nodes and edges. In the data organization scenario of a 3D model, an octree may be used to organize the spatial data of the 3D model, and each node in the octree has eight child nodes. A structured description structure refers to a data organization structure that uses nested tags or key-value pairs to represent hierarchical relationships, such as Extensible Markup Language (XML). The following will take the implementation of the geometric model sample as a 3D model of any object as an example to illustrate the optional ways in which the encoder organizes the model description data sample.

[0035] After obtaining the 3D model of the object, you can sample from the surface of the 3D model points, forming a point cloud ,in For the The coordinates of the points, For the The normal vector of a point.

[0036] Optionally, the encoder may determine an octree structure with a target depth based on the target resolution, and recursively divide the spatial data corresponding to the geometric model sample into spatial nodes based on the octree structure until the target depth is reached. The spatial data of the geometric model sample may include the coordinates and normal vectors of the points in the point cloud obtained by sampling the geometric model sample. The encoder may recursively divide the three-dimensional space where the geometric model sample is located into eight sub-cubes, thereby facilitating the efficient representation and query of three-dimensional data. During the recursive process, the three-dimensional space where the geometric model sample is located may be used as the root node. With each additional layer of depth, each divided cube may be divided into two parts in terms of length, width, and height, respectively, to evenly divide each cube into eight smaller cubes. With d representing the depth of the octree, at depth d, the three-dimensional space where the geometric model sample is located is divided into 2^3d=8^d cubes, each of which corresponds to a spatial node, which is the smallest resolvable unit. The spatial ranges represented by spatial nodes in different levels are different. As the depth of the level increases, the spatial range represented by the spatial nodes becomes smaller, which can more accurately represent the structural information in the space and facilitate the capture of spatial details at a higher resolution. In this embodiment, the depth of the octree can be selected according to the target resolution, and the target resolution can be the highest resolution of the desired output. For example, when the target resolution is 1024, the depth d of the octree can be determined to be 10. When the target resolution is 4096, the depth d of the octree can be determined to be 12. The examples are not listed here.

[0037] After dividing the spatial data corresponding to the geometric model samples into spatial nodes according to the target depth, the encoder can organize the spatial nodes obtained by dividing the geometric model samples according to the depth level to which the spatial nodes belong, resulting in multiple levels of spatial nodes. Among them, spatial nodes belonging to the same depth level are organized into spatial nodes of the same level. The deeper the level, the smaller the spatial range represented by each spatial node in the level.

[0038] For example, if the depth of the target structure is 10, the spatial nodes are organized to obtain 10 levels of spatial nodes. Optionally, for any level, the empty nodes in the level can be identified and filtered out to compress the number of nodes. For any spatial node, it can be determined whether the cube corresponding to the spatial node contains the points in the geometric model sample. If it does, the spatial node is determined to be a non-empty node; if not, the node is determined to be an empty node. In other words, this hierarchical data organization method does not need to store empty areas in the geometric model samples and can effectively compress spatial data. Among them, the first Hierarchical spatial nodes It can be described as:

[0039]

[0040] in, Indicates the total number of levels, is a The matrix contains the The characteristics of all spatial nodes in the hierarchy, express The feature space of Indicates the the number of spatial nodes in the hierarchy, Indicates the number of feature channels for each node.

[0041] In some optional embodiments, after the spatial data corresponding to the geometric model samples are organized into multiple levels of spatial nodes using an encoder, the multiple levels of spatial nodes can be directly used as reference marker maps at multiple scales.

[0042] In other optional embodiments, after the spatial data corresponding to the geometric model samples are organized into multiple levels of spatial nodes by an encoder, the encoder can be used to compress the multiple levels of spatial nodes to obtain reference label maps of multiple scales to compress the data volume of the reference label maps. Optionally, the encoder can use a cross-scale attention mechanism to compress the multiple levels of spatial nodes to obtain reference label maps of multiple scales. Among them, the cross-scale attention mechanism is a mechanism that enables the encoder to dynamically focus on key parts when processing input data. Its core idea is to imitate human attention and assign different weights to different parts of the input data according to task requirements, so as to more effectively extract useful information and achieve compression of the input data.

[0043] The following will take any target level in multiple levels as an example to illustrate an optional implementation method of compressing spatial nodes of the multiple levels based on the cross-scale attention mechanism.

[0044] In some optional embodiments, the encoder may encode the position data in the spatial nodes of the target level to obtain the position codes corresponding to the spatial nodes of the target level. The position coding operation is used to encode the position data of the spatial nodes of the target level into high-dimensional spatial vectors to enrich the position expression of the spatial nodes of the target level. The position data of the spatial nodes include the coordinates of the geometric center point of the cube corresponding to the spatial nodes. In this embodiment, the scales of spatial nodes at different levels are different, and compressed query vectors of different scales can be defined for spatial nodes at different levels to obtain compressed query vectors of multiple scales. Figure 2 As shown, the compressed query vectors at multiple scales can be expressed as , For the Compressed query vectors corresponding to spatial nodes at each level. Compressed query vectors at different scales contain learnable parameters, which are used to learn to assign different weights to features at that scale based on their importance to capture key features. These learnable parameters are updated during the encoder training process to continuously learn to capture key features from the data until the encoder's training loss converges to a specified range.

[0045] Continuing with the target level as an example, the encoder can obtain the compressed query vector corresponding to the target level. The compressed query vector contains a learnable first parameter, which is used to learn to assign different weights to the features of the target level according to feature importance to capture the knowledge of key features. The encoder can perform cross-scale attention calculations on the position encoding and the compressed query vector to capture key features from the position encoding and obtain the first compressed feature corresponding to the spatial node of the target level. Taking the spatial nodes of the hierarchy as an example, the above attention calculation process can be expressed by the following formula:

[0046]

[0047] in, For the The first compressed feature of the spatial node of the level, is the position encoding function, For the The position encoding corresponding to the spatial nodes of the hierarchy, Calculates the function for cross attention.

[0048] After obtaining the first compressed feature based on the above implementation, the encoder can quantize the first compressed feature to obtain a reference marker map of the target scale corresponding to the target level. Quantization refers to the process of mapping continuous values ​​to discrete values. In this embodiment, when the encoder quantizes the first compressed feature, it can map the first compressed feature to a discrete value in the codebook, and the discrete value is a feature vector in the codebook. Among them, the codebook is a predefined set of feature vectors, which is usually learned through clustering (such as K-means) or other methods. When quantizing the first compressed feature, the encoder can use calculation methods such as Euclidean distance or cosine similarity to find a feature that is more similar to the first compressed feature from the codebook, and map the first compressed feature to the more similar feature. Based on the quantization operation, the compressed features of multiple levels can be quantized into reference marker maps of multiple scales, and the reference marker maps of different scales have different scales. The reference marker maps of different scales of multiple levels can be expressed as:

[0049]

[0050] in, Indicates the Reference marker diagram for scale, , is the quantization function. express The feature space of Indicates the The number of feature points in , Indicates the number of feature channels of each reference label map. The number of feature points in the reference marker map of the scale Much smaller than The number of spatial nodes in the hierarchy , achieving a large amount of compression of spatial nodes. Based on this implementation, features can be effectively compressed to reduce storage and computing costs.

[0051] Optionally, before quantizing the first compressed features, the encoder may further fuse features at different levels to enhance the feature expression capability of the model. Figure 2 As shown, the encoder obtains the The first compressed feature of the spatial node of the level After that, the cross-scale attention mechanism can be used to compress the first feature Perform feature fusion operation to obtain the updated first compression feature After that, the first compressed feature is updated Quantify and obtain .

[0052] Continuing with the target level as an example, optionally, the encoder may obtain the second compressed features corresponding to the spatial nodes of the adjacent levels of the target level, sample the second compressed features, and obtain a third compressed feature having the same scale as the first compressed feature. In some optional embodiments, the adjacent levels of the target level may include the upper level of the target level. Taking the octree structure as an example, in the recursive process from the root node to the leaf node, the greater the depth of the octree, the larger the level. For the target level, the scale of the second compressed features corresponding to the spatial nodes of the upper level is small, and the scale and detail features of the second compressed features can be increased by upsampling. In other optional embodiments, the adjacent levels of the target level may include the lower level of the target level. For the target level, the scale of the second compressed features corresponding to the spatial nodes of the lower level is large, and the scale and detail features of the second compressed features can be reduced by downsampling. Among them, the method of sampling the second compressed features may include: interpolation method or convolution method, which is not limited in this embodiment.

[0053] After obtaining the third compressed feature, the encoder can use a cross-scale attention mechanism to fuse the first compressed feature and the third compressed feature to update the first compressed feature. The above fusion process can be expressed by the following formula:

[0054]

[0055] in, For the The first compressed feature of the spatial node of the level, For the The third compressed feature is obtained by upsampling the second compressed feature of the spatial node of the level.

[0056] In this embodiment, after obtaining the updated first compressed feature, the encoder may quantize the first compressed feature to obtain a reference marker map of the target scale. That is, the quantized first compressed feature contains feature data of adjacent levels. The first compressed feature and the third compressed feature correspond to different levels. The compressed features corresponding to the higher level can better capture local details, while the compressed features corresponding to the lower level can better capture the features of the global context. The first compressed feature and the third compressed feature are fused so that local details and global context information can be fused, thereby improving the model's ability to understand and generate complex data. Among them, the cross-scale attention mechanism is a mechanism that uses a cross-scale attention mechanism to process multiple scale features, which can establish associations between features of different scales, capture local details and global context information, dynamically adjust the weights of features of different scales, and enhance the flexibility of the model.

[0057] After obtaining reference label maps at multiple scales based on the above-described embodiment, in step 103, an autoregressive model can be used to perform multiple spatial feature autoregressive generation operations on the target feature information to obtain target label maps at multiple scales. An autoregressive model is a statistical model that can predict the next possible output data based on previous data. In this embodiment, the autoregressive model performs multiple spatial feature autoregressive generation operations to recursively predict the label map at the next possible output based on the label map at the previous scale until a recursive termination condition is met. Multiple spatial feature autoregressive operations can correspond one-to-one to multiple scales, and the spatial features generated by any autoregressive generation operation are stored in the label map at the corresponding scale. For example, the spatial features output by the first autoregressive generation operation are stored in the target label map at the first scale, the spatial features output by the second autoregressive generation operation are stored in the target label map at the second scale, and so on. The target label maps at multiple scales are used to decode and obtain geometric models corresponding to multiple resolutions. This geometric model refers to the geometric model predicted by the autoregressive model based on the model description data sample.

[0058] In some optional embodiments, an autoregressive model can be constructed based on a pre-trained model. The pre-trained model can be a large language model, which is a natural language processing (NLP) model that has been trained on a large scale. Large language models are typically built based on deep learning techniques and trained on large training datasets, enabling them to demonstrate strong performance in natural language processing tasks. Given a context (i.e., a labeled image at the previous scale and target feature information for the model's description of a data sample), the large language model attempts to predict the most likely labeled image at the next scale. The number of parameters in the large language model is greater than a set threshold, typically in the millions or billions. During training, these parameters are continuously adjusted and optimized based on the discrepancy between the large language model's predictions and actual results to improve the large language model's prediction accuracy. In some embodiments, the pre-trained large language model can be constructed using a decoder-only transformer architecture. This architecture enables the large language model to capture complex patterns in the model's descriptions and handle long-range dependencies. After sufficient training, the large language model has powerful generation capabilities and can output the expected target labeling graph based on the target feature information of the given model description data sample.

[0059] In this embodiment, a large language model pre-trained on a large number of general data sets can be used as the base model of the autoregressive model. In order to adapt to the specific application scenarios of this application, that is, for generating target labeling maps of multiple scales based on the target feature information of the input model description data samples, the pre-trained large language model can be trained on a data set formed by description data samples of a large number of geometric model samples and reference labeling maps of geometric model samples, so that the trained large language model is more suitable for performing feature prediction tasks.

[0060] In this embodiment, the autoregressive model is used to predict the target label map of the current scale based on the target label map of the previous scale that has been generated, which will be exemplified below.

[0061] Optionally, an autoregressive model can be used to perform feature extraction on the model description data samples to obtain target feature information. Optionally, the model description data samples may include model description data of at least one different modality, such as but not limited to text description data, point cloud data, image data, etc. When performing feature extraction on the model description data samples, target feature information of each of the different modalities of model description data can be extracted according to the feature extraction methods corresponding to the model description data of each modality. For example, semantic features of text description data can be extracted, geometric features of point cloud data can be extracted, and image features of image data can be extracted. The features obtained by feature extraction of the model description data samples are expressed as target feature information.

[0062] Optionally, when any autoregressive generation operation is performed in the autoregressive model, if the current scale corresponding to any autoregressive generation operation is the first scale, the autoregressive model can map the target feature information into discrete initial spatial features, and use the feature map formed by the discrete initial spatial features as the target label map of the current scale. If the current scale is not the first scale, the autoregressive model can perform autoregressive generation based on the target feature information and the target label map of the previous scale of the current scale to obtain the target label map of the current scale. In this embodiment, the autoregressive model generates target label maps of different scales by first discretizing the target feature information and then performing autoregressive generation on the discretized target feature information. The continuous target feature information can be converted into discrete spatial features so as to establish an association between the discrete spatial features and the voxels of the geometric model, thereby facilitating the construction of the structure of the geometric model.

[0063] Optionally, the autoregressive model may obtain feature dimensions corresponding to each of the multiple scales based on the reference label maps at multiple scales output by the encoder. If the current scale is the first scale, when mapping the target feature information to discrete initial spatial features, the autoregressive model may linearly map the target feature information based on the feature dimensions corresponding to the first scale to obtain discrete initial spatial features that match the feature dimensions corresponding to the first scale.

[0064] If the current scale is not the first scale, the autoregressive model can perform linear mapping on the target feature information when performing autoregressive generation based on the target feature information and the target label map of the previous scale of the current scale to obtain target feature information that matches the feature dimension of the current scale. The autoregressive model can then concatenate the linearly mapped target feature information with the target label map of the previous scale of the current scale to obtain concatenated features, and perform autoregressive generation based on the concatenated features to obtain the target label map of the current scale. The autoregressive model includes a large number of learnable model parameters that can be continuously updated during the model training process to learn the mapping relationship between the target feature information of the input geometric model description data and the target label maps of multiple scales of the desired output, thereby accumulating the knowledge of the autoregressive model for feature prediction tasks.

[0065] When the model describes that the data sample is multimodal data, feature extraction can be performed on the multimodal data to obtain multimodal target feature information, which is marked as .like Figure 3 As shown, the multimodal target feature information can be Input autoregressive model. Target feature information After linear mapping, multiple target feature information matching multiple levels of feature dimensions is obtained, such as Figure 3The differences shown are Level, Level and Hierarchical feature dimension matching target feature information 、 、 The autoregressive model can be based on the input target feature information Start predicting the target label map of the first scale, and predict the target label map of the next scale based on the target feature information and the target label map of the first scale. Figure 3 As shown, the autoregressive model outputs After that, the target labeling map can be tagged using a cross-scale attention mechanism. Up-sample and get the same Hierarchical scale-matched target labeling graph , and With the Hierarchical target feature information After splicing, input the autoregressive model. The autoregressive model can predict the first . Get the target marker map After that, the target labeling map can be tagged using a cross-scale attention mechanism. Up-sample and get the same Hierarchical scale-matched target labeling graph , and With the Hierarchical target feature information After splicing, input the autoregressive model. The autoregressive model can predict the first , and so on, no more details. Figure 3 In

[15] , the mask is used to ensure that the predicted label map of the forward scale is used instead of the label map of the backward scale when predicting the label map of the next scale, so as to maintain the causality of the autoregressive process.

[0066] In step 104, the reference label maps at multiple scales output by the encoder may be used as supervisory signals, and the autoregressive model may be trained based on the differences between the reference label maps at multiple scales and the target label maps at multiple scales. During the training of the autoregressive model, the learnable parameters in the autoregressive model may be continuously adjusted until the differences between the target label maps at multiple scales output by the autoregressive model and the reference label maps at multiple scales output by the encoder are less than a set threshold. At this point, the training of the autoregressive model is terminated, and the trained autoregressive model is output.

[0067] Optionally, during the training process, the difference between the reference label map and the target label map of the same scale can be obtained based on the scale correspondence, for example, the difference between the reference label map of the first scale and the target label map of the first scale, and the difference between the reference label map of the second scale and the target label map of the second scale can be obtained. Afterwards, the cross entropy loss between the reference label map of the multiple scales and the target label map of the multiple scales can be calculated based on the difference between the reference label map and the target label map of the same scale. In the process of training the autoregressive model, the cross entropy loss can be converged to a specified range as an optimization goal to train the autoregressive model. In each round of optimization process, it can be determined whether the cross entropy loss of the autoregressive model converges to the specified range. If it does not converge, the learnable parameters in the autoregressive model can be adjusted and the next round of optimization can be performed. If the cross entropy loss of the autoregressive model converges to the specified range after a certain round of optimization, the optimization operation of the autoregressive model can be stopped to obtain a trained autoregressive model. In this embodiment, the reference label maps of multiple scales obtained by the encoder encoding the geometric model samples are used as supervisory signals to train the autoregressive model, so that the autoregressive model can learn the ability to generate the desired target label maps of multiple scales based on the target feature information of the input model description data samples, thereby facilitating the generation of high-quality geometric models based on the target label maps of multiple scales.

[0068] In some optional embodiments, the encoder can be further trained to obtain a more accurate supervisory signal, thereby correctly guiding the parameter update of the autoregressive model and reducing the difficulty of parameter debugging of the autoregressive model. The following is an exemplary description of the encoder optimization process.

[0069] Optionally, the distribution characteristics of the compression features corresponding to each of the multiple levels can be obtained, and based on the distribution characteristics of the compression features corresponding to each of the multiple levels, the first loss corresponding to the encoder can be obtained. Optionally, the distribution characteristics may include mean and / or variance. Among them, the compression features corresponding to any level may include: compression features obtained by compressing the spatial nodes of this level, or updated compression features obtained by fusing the compression features of this level with compression features of other levels, which is not limited in this embodiment. For example, the mean and variance of the first compression feature corresponding to the first level can be obtained, the mean and variance of the second compression feature corresponding to the second level can be obtained, and so on. Alternatively, the mean and variance of the updated first compression feature corresponding to the first level can be obtained, the mean and variance of the updated second compression feature corresponding to the second level can be obtained, and so on. As Figure 2 As shown, get the The updated first compressed feature of the level Afterwards, the linear mapping layer can be used to obtain the The mean and variance corresponding to each level. The regularization loss of the encoder can be obtained based on the mean and variance of the compressed features corresponding to multiple levels.

[0070] Optionally, in addition to the regularization term loss, the training process of the encoder can also be constrained according to the prediction loss of the autoregressive model. Optionally, a decoding unit can be used to decode the target label map of multiple scales output by the autoregressive model to obtain decoding results corresponding to multiple scales, and the prediction loss of the autoregressive model can be determined according to the difference between the decoding results and the geometric model samples. The decoding unit can be a decoder independent of the encoder and the autoregressive model; or, as Figure 2 As shown, the decoding unit can be a decoding module built into the aforementioned encoder, and this embodiment does not impose any restrictions on this. If the decoding unit is an independent decoder, the decoder can be trained synchronously with the encoder to improve decoding capabilities. If the decoding unit is a decoding module built into the aforementioned encoder, the parameters of the decoding module can be gradually optimized during the encoder training process.

[0071] The following is an example of how to obtain the predicted loss.

[0072] Optionally, the decoding unit can obtain the signed distance prediction value of any point in the space based on the target marker map of multiple scales output by the autoregressive model. The signed distance prediction value is the value of the signed distance function (SDF) of any point in the space obtained based on the target marker map. SDF is a scalar field used to describe the relationship between a point and an object surface in three-dimensional space. For any point in space, SDF returns the distance value from the point to the nearest object surface. If the point is outside the object, a positive value is returned; if it is inside the object, a negative value is returned; if it is exactly on the surface, the SDF value is 0. In this embodiment, an SDF decoder can be used to decode the target marker map of multiple scales output by the autoregressive model to obtain the SDF prediction values ​​corresponding to each of the multiple levels. For example, mark the first The target marker map of the scale is , for the target labeling graph The decoding process can be described by the following formula:

[0073]

[0074]

[0075] in, is the spatial coordinate of any point in three-dimensional space, is the encoding of the spatial coordinates of any point, For the SDF decoder, is the SDF prediction value of any point. According to the signed distance prediction value and the actual signed distance value of the geometric model sample, the second loss corresponding to the encoder is obtained. The actual signed distance value is the SDF value of any spatial point actually obtained from the geometric model sample, which is used to represent the actual surface shape of the geometric model sample. After obtaining the first loss and the second loss, the encoder is trained with the first loss and the second loss converging to the specified range training target. The first loss and the second loss converging to the specified range may include: both the first loss and the second loss converge to the specified range, or the joint loss formed by the first loss and the second loss converges to the specified range. In this embodiment, the joint loss corresponding to the encoder can be expressed as:

[0076]

[0077] in, is the regularization loss with respect to the mean and variance of the predictions, is the SDF loss.

[0078] After obtaining the joint loss, the encoder can be updated based on the magnitude of the joint loss until the encoder's joint loss converges to a specified range. When the encoder's joint loss converges to the specified range and the autoregressive model's loss converges to the specified range, the encoder and autoregressive model optimization process can be stopped, resulting in a trained autoregressive model. The autoregressive model can be deployed in an online production environment to assist in providing services that generate 3D mesh models based on data samples described by the model.

[0079] In some optional embodiments, such as Figure 2 As shown, before decoding the target label map at different scales, the encoder can use upsampled query vectors at different scales The target label maps of different scales are upsampled to restore the compressed low-dimensional target label map to a high-dimensional target label map, that is, . Among them, the upsampled query vector The decoder includes a learnable second parameter, which is used during encoder training to learn the importance weights of different locations or channels in the target marker image. This allows for dynamic adjustment of the contributions of different features or channels during upsampling, improving detail recovery. After upsampling, the decoder module decodes the upsampled target marker image to obtain a higher-fidelity geometric model.

[0080] Based on the above embodiment, the present application provides a training method including an encoder and an autoregressive model, wherein the encoder encodes the geometric model sample to obtain reference label maps of multiple scales, and the reference label maps at different scales can not only express the global structural features of the geometric model sample, but also capture the local detail features of the geometric model sample. The reference label maps of multiple scales obtained by the encoder encoding the geometric model sample are used as supervisory signals to train the autoregressive model, so that the autoregressive model can effectively learn the relationship and regularity between feature information of different scales in space based on the reference label maps of multiple scales, thereby improving the autoregressive model's ability to understand and generate spatial features. Based on the above capabilities, the autoregressive model can accurately output target label maps of different scales according to the target feature information of the input model description data sample, so as to facilitate the selection of the first label map of the target scale corresponding to the target resolution for decoding according to needs, and obtain a geometric model that meets the expected resolution, thereby flexibly meeting the needs of different scenarios.

[0081] Furthermore, in the process of encoding the geometric model samples, the encoder organizes the spatial data corresponding to the geometric model samples into multiple levels of spatial nodes, so that the disordered spatial data can be efficiently organized. On the one hand, it is convenient to support efficient spatial calculations. On the other hand, this hierarchical structure can express multiple scale information of the geometric model samples and support high-level expression of the structure of the geometric model samples, thereby facilitating the extraction of structural details of the geometric model samples.

[0082] Furthermore, the encoder uses a cross-scale attention mechanism to compress spatial nodes at multiple levels, effectively addressing the lack of intrinsic spatial continuity in spatial data. This allows for efficient compression of spatial data into a low-dimensional representation while preserving the structural details of the geometric model samples. When using the encoder to supervise the training of the autoregressive model, a low-dimensional reference marker map containing structural information at different levels is used to guide the training process. This not only guides the autoregressive model's ability to capture structural information at different scales, but also effectively improves the model's training and inference speeds, facilitating the use of lower computational costs and reducing the resolution constraints on the generated 3D mesh model.

[0083] Furthermore, when the autoregressive model uses a pre-trained large language model as its base model, it can fully utilize the large language model's ability to understand information such as images and text, as well as its powerful generation capabilities, and output target labeling maps of multiple scales and meeting multiple expected resolutions based on the target feature information of the given model description data sample.

[0084] In addition to the model training method described in the above embodiments, the present invention also provides a model generation method. Figure 4 As shown, the method may include:

[0085] Step 401: In response to a model generation operation, obtain geometric model description data.

[0086] Step 402: Extract features from the geometric model description data to obtain first feature information.

[0087] Step 403: Using the autoregressive model, perform multiple autoregressive generation operations of spatial features based on the first feature information to obtain first label maps of multiple scales. The spatial features generated by different autoregressive generation operations correspond to first label maps of different scales, and the scale of the first label map is positively correlated with the resolution of the geometric model.

[0088] Step 404: Decode the first label map of at least one target scale among the multiple scales according to at least one target resolution to obtain at least one target geometric model.

[0089] The method for generating a three-dimensional model provided in this embodiment can be applied to a variety of different application scenarios, such as but not limited to e-commerce scenarios, game and film and television scenarios, industrial design and manufacturing scenarios, etc. For example, in an e-commerce scenario, based on the method provided in this embodiment, a 3D virtual fitting model can be obtained based on a user's image photo or clothing size, so as to facilitate real-time viewing of the effect of clothing on the body. Alternatively, a 3D model of a product can be generated based on the merchant's product description text to fully display the product effect and enhance the user experience. Alternatively, a 3D model of a product can be generated based on the user's customized product description information or hand-drawn picture to provide a convenient product customization service. For another example, in game and film and television scenarios, based on the method provided in this embodiment, 3D characters or props can be generated based on scripts or concept maps to speed up the game and film and television production process, which will not be listed one by one.

[0090] In step 401, the model generation operation may be triggered by the user, or may be triggered by certain applications or systems. In some embodiments, the target service for executing the model generation method may run on a terminal device or partially run on a terminal device. The terminal device may display a human-computer interaction interface and may detect the model generation operation initiated by the user by detecting the operation on the human-computer interaction interface. The geometric model description data may be uploaded when the user initiates the model generation operation, or may be obtained from a specified database. The address of the specified database may be provided by the user or may be a default address, which is not limited in this embodiment. Optionally, the autoregressive model may be deployed on the terminal device, or may be deployed on a server device that communicates with the terminal device, which is not limited in this embodiment. In other embodiments, the target service for executing the model generation method may provide an externally open interface. When a call operation of the interface by an application or system is detected, it may be considered that a model generation operation has been detected. The target service may obtain the geometric model description data based on the call parameters of the interface.

[0091] The geometric model description data is data describing the desired output geometric model. In some optional embodiments, the geometric model description data may include: model description data of at least one modality used to describe the desired output geometric model. For example, it may include, but is not limited to, at least one of text description data, voice description data, point cloud data, and image description data. The image description data may include images of the geometric model sample from one or more different perspectives. For example, if the geometric model sample is a 3D model of any object, its geometric model description data may include images of the object from different perspectives, such as a front view, a side view, and a top view. For example, in e-commerce scenarios, when providing users with virtual fitting services, user-provided image photos or clothing size data may be obtained as geometric model description data; when providing merchants with smart product display services, product description text provided by merchants may be obtained as geometric model description data; when providing users with product customization services, user-provided description information or hand-drawn drawings of customized products may be obtained as geometric model description data. For another example, in gaming and film scenarios, scripts or concept drawings may be obtained as geometric model description data.

[0092] After obtaining the geometric model description data, step 402 may be executed to perform feature extraction on the geometric model description data to obtain first feature information. Optionally, target feature information for each of the different modalities of model description data may be extracted using feature extraction methods corresponding to the model description data of each modality. For example, semantic features of text description data, geometric features of point cloud data, and image features of image data may be extracted. Accordingly, the target feature information may include at least one of semantic features, geometric features, and image features.

[0093] After obtaining the target feature information, in step 402, the autoregressive model can be used to perform multiple autoregressive generation operations of spatial features based on the first feature information to obtain first label maps of multiple scales. The spatial features generated by any autoregressive generation operation are stored in the first label map of the corresponding scale. For example, the spatial features output by the first autoregressive generation operation are stored in the first label map of the first scale, the spatial features output by the second autoregressive generation operation are stored in the first label map of the second scale, and so on. The scale of the first label map is positively correlated with the resolution of the geometric model. The larger the scale of the first label map, the more discrete feature units it contains, and the higher the resolution of the decoded geometric model. The first label maps of multiple scales can be used to decode and obtain geometric models of multiple resolutions.

[0094] In this embodiment, the autoregressive model is trained based on the differences between reference labeled maps at multiple scales corresponding to the geometric model sample and second labeled maps at multiple scales output by the autoregressive model for the model description data of the geometric model sample. During training, the autoregressive model learns the relationships between spatial features at different scales, facilitating the execution of multiple autoregressive spatial feature generation operations based on these relationships.

[0095] In some optional embodiments, when training an autoregressive model, a geometric model sample and a corresponding model description data sample may be obtained. The spatial data corresponding to the geometric model sample may be encoded using an encoder to obtain reference label maps at multiple scales. The reference label maps at multiple scales are used to represent the spatial features of the geometric model sample at multiple resolutions. Feature extraction may be performed on the model description data sample to obtain second feature information. The autoregressive model may then perform multiple autoregressive spatial feature generation operations based on the second feature information to obtain second label maps at multiple scales. The second label maps at multiple scales are used to decode and obtain geometric models at multiple resolutions. Based on the differences between the reference label maps at multiple scales and the second label maps at multiple scales, the autoregressive model may be trained to learn the relationships between spatial features at different scales. In this embodiment, the terms "first" and "second" are used to define feature information to facilitate distinguishing feature information extracted from different stages. The first feature information is obtained by extracting features from the geometric model description data during the model inference phase, while the second feature information is obtained by extracting features from the model description data of the geometric model sample during the model training phase. These are the same target feature information as described in the previous embodiments.

[0096] Optionally, in the process of encoding the spatial data corresponding to the geometric model samples using the encoder, the encoder can be used to organize the spatial data corresponding to the geometric model samples into multiple levels of spatial nodes, and a cross-scale attention mechanism can be used to compress the multiple levels of spatial nodes to obtain reference marker maps of multiple scales. The specific implementation of the above training process can refer to the records of the aforementioned embodiments and will not be described again here. Based on the training process, the autoregressive model can learn the ability to generate the desired first marker maps of different scales based on the target feature information of the input geometric model description data under the supervision of the reference marker maps of multiple scales. The first marker maps of different scales can be decoded to obtain geometric models of different resolutions that meet the user's expectations.

[0097] Optionally, when the autoregressive model performs any autoregressive generation operation, if the current scale corresponding to the autoregressive generation operation is the first scale, the autoregressive model may map the target feature information into discrete initial spatial features, and use the feature map formed by the discrete initial spatial features as the first labeled map at the current scale. Optionally, if the current scale is not the first scale, the autoregressive model may perform autoregressive generation based on the target feature information and the first labeled map of the previous scale of the current scale to obtain the first labeled map at the current scale.

[0098] Optionally, during the process of performing autoregressive generation based on the target feature information and the first labeled map at the previous scale before the current scale, the autoregressive model may perform linear mapping on the target feature information so that the target feature information matches the feature dimension of the current scale. The autoregressive model may then concatenate the linearly mapped target feature information with the reference labeled map at the previous scale before the current scale to obtain a concatenated feature, and perform autoregressive generation based on the concatenated feature to obtain the first labeled map at the current scale.

[0099] After obtaining first labeled maps at multiple scales based on the above embodiment, in step 404, the first labeled map at at least one target scale among the multiple scales may be decoded based on at least one target resolution to obtain at least one target geometric model. The at least one target scale and the at least one target resolution may have a one-to-one correspondence, and the first labeled map at one scale may be decoded to obtain a geometric model at one resolution.

[0100] Optionally, the operation of decoding the first labeled image of the target scale corresponding to any target resolution may be implemented based on an SDF decoder and a marching cubes algorithm, which will be exemplarily described below.

[0101] Optionally, for any target scale among the multiple scales, the signed distance function value for any point in three-dimensional space can be obtained based on the first labeled map for that target scale, and a continuous signed distance field can be reconstructed based on the signed distance function for any point in space. A marching cubes algorithm can then be used to extract a triangulated network from the signed distance field to obtain a geometric model of the target resolution corresponding to that target scale.

[0102] Optionally, the target scale can be any one of multiple scales, some of the scales, or all of the scales, and this embodiment does not impose any limitation thereto. In some optional embodiments, the target scale can be the default full scale. The decoder can decode the first labeled images at multiple scales to obtain geometric models at multiple resolutions corresponding to the multiple scales. The geometric models at multiple resolutions can all be output to the user for flexible selection.

[0103] In other optional embodiments, the target scale may be a scale corresponding to a portion of the target resolution specified by the entity initiating the model generation operation (e.g., a merchant user or consumer user on an e-commerce platform). Optionally, before decoding the first labeled map for at least one target scale from multiple scales based on at least one target resolution, the decoder may obtain at least one expected target resolution from the geometric model description data and determine, from the multiple scales, the scale corresponding to the at least one target resolution as the at least one target scale. The decoder may identify keywords describing resolution, such as high-definition, high-pixel, ultra-high-definition, or numerical values ​​containing resolution units, from the geometric model description data. Based on a preset correspondence between keywords and resolutions, the decoder may convert the identified keywords into a target resolution. The decoder may store correspondences between multiple scales and multiple resolutions to facilitate determination of the target scale based on the identified target resolution. Based on this embodiment, the decoder can flexibly and selectively decode the first labeled map for the target scale based on the user's generation requirements, reducing the computational complexity required for decoding and improving the alignment of the decoded geometric model with the user's requirements.

[0104] Based on this implementation, the autoregressive model is trained under the supervision of reference marker maps of multiple scales. The reference marker maps of multiple scales are obtained by encoding geometric model samples. The reference marker maps at different scales can not only express the global structural features of the geometric model samples, but also capture the local detail features of the geometric model samples. Furthermore, the autoregressive model can effectively learn the relationship and regularity between feature information of different scales in space during the training phase, thereby improving the ability to understand and generate spatial features. Based on the above capabilities, the autoregressive model can perform autoregressive generation operations on spatial features based on the feature information of the input geometric model description data to accurately output multiple marker maps of different scales. Among them, the scale of the marker map is positively correlated with the resolution of the geometric model, so that the marker maps of multiple scales can not only express the global structural features of the geometric model samples, but also express the local detail features of the geometric model. Therefore, it is convenient to select the marker map of the target scale corresponding to the target resolution for decoding according to the needs, and obtain a geometric model that meets the expected resolution, thereby flexibly meeting the needs of different scenarios.

[0105] In some optional embodiments, the multi-scale reference marker maps are obtained by an encoder organizing the spatial data corresponding to the geometric model samples into multiple levels of spatial nodes and compressing the multiple levels of spatial nodes using a cross-scale attention mechanism. The encoder organizes the spatial data corresponding to the geometric model samples into multiple levels of spatial nodes, which can efficiently organize the disordered spatial data. This facilitates efficient spatial computation and, on the other hand, this hierarchical structure can express the multi-scale information of the geometric model samples and support a high-level expression of the structure of the geometric model samples, thereby facilitating the extraction of the structural details of the geometric model samples. The encoder uses a cross-scale attention mechanism to compress the spatial nodes of the multiple levels, effectively resolving the problem of the lack of intrinsic spatial continuity in spatial data. This allows the spatial data to be efficiently compressed into a low-dimensional representation while preserving the structural details of the geometric model samples. Furthermore, using a reference marker map containing structural information at different levels and with a low data volume to guide the training process of the autoregressive model not only guides the autoregressive model in learning to capture structural information at different scales, but also effectively improves the inference speed of the autoregressive model.

[0106] It should be noted that the execution entity of each step of the method provided in the above embodiment can be the same device, or the method can be executed by different devices. For example, the execution entity of steps 101 to 104 can be device A; for another example, the execution entity of steps 101 and 102 can be device A, and the execution entity of step 103 can be device B; and so on.

[0107] In addition, some of the processes described in the above embodiments and the accompanying drawings include multiple operations that appear in a specific order, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0108] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0109] Figure 5 The structure diagram of an electronic device provided by an exemplary embodiment of the present application is shown, and the electronic device is applicable to the model generation method or model training method provided by the above embodiment. Figure 5 As shown, the electronic device includes: a memory 501 and a processor 502.

[0110] The memory 501 is used to store computer programs and can be configured to store various other data to support operations on the electronic device. Examples of such data include instructions for any application or method operating on the electronic device.

[0111] In some exemplary embodiments, Figure 5 The electronic device shown is used to execute a model generation method, wherein the processor 502 is coupled to the memory 501 and is used to execute the computer program in the memory 501, so as to: obtain geometric model description data in response to a model generation operation; perform feature extraction on the geometric model description data to obtain first feature information; use an autoregressive model to perform multiple autoregressive generation operations of spatial features according to the first feature information to obtain first label maps of multiple scales, wherein the spatial features generated by different autoregressive generation operations correspond to first label maps of different scales; the scale of the first label map is positively correlated with the resolution of the geometric model; and decode the first label map of at least one target scale among the multiple scales according to at least one target resolution to obtain at least one target geometric model.

[0112] Optionally, the processor 502 is also used to: obtain the geometric model sample and the corresponding model description data sample; use an encoder to encode the spatial data corresponding to the geometric model sample to obtain reference label maps of multiple scales, and the reference label maps of multiple scales are used to express the spatial features of the geometric model sample at multiple resolutions; perform feature extraction on the model description data sample to obtain second feature information, and use an autoregressive model to perform multiple autoregressive generation operations of spatial features according to the second feature information to obtain second label maps of multiple scales, and the second label maps of multiple scales correspond to the multiple resolutions; and train the autoregressive model based on the difference between the reference label maps of the multiple scales and the second label maps of the multiple scales.

[0113] Optionally, when the processor 502 uses the encoder to encode the spatial data corresponding to the geometric model sample to obtain reference label maps of multiple scales, it is specifically used to: use the encoder to organize the spatial data corresponding to the geometric model sample into multiple levels of spatial nodes, and the spatial nodes in different levels represent different spatial ranges; use a cross-scale attention mechanism to compress the spatial nodes of the multiple levels to obtain reference label maps of multiple scales.

[0114] Optionally, when the processor 502 uses the autoregressive model to perform multiple autoregressive generation operations of spatial features according to the target feature information, it is specifically used to: when performing any autoregressive generation operation in the autoregressive model, if the current scale corresponding to any autoregressive generation operation is the first scale, map the target feature information into discrete initial spatial features, and use the feature map formed by the discrete initial spatial features as the first labeled map of the current scale; if the current scale is not the first scale, perform autoregressive generation according to the target feature information and the first labeled map of the previous scale of the current scale to obtain the first labeled map of the current scale.

[0115] Optionally, when the processor 502 performs autoregressive generation based on the target feature information and the first label map of the previous scale of the current scale to obtain the first label map of the current scale, it is specifically used to: linearly map the target feature information so that the target feature information matches the feature dimension of the current scale; splice the linearly mapped target feature information with the reference label map of the previous scale of the current scale to obtain a spliced ​​feature; and perform autoregressive generation based on the spliced ​​feature to obtain the first label map of the current scale.

[0116] Optionally, when the processor 502 decodes the first labeled map of at least one target scale among the multiple scales according to at least one target resolution to obtain at least one target geometric model, the processor 502 is specifically configured to: obtain, for any target resolution among the at least one target resolution, a signed distance function value of an arbitrary point in a three-dimensional space according to the first labeled map of the target scale corresponding to the target resolution; reconstruct a continuous signed distance field according to the signed distance function of the arbitrary point in the space; and extract a triangulated network from the signed distance field using a marching cubes algorithm to obtain the target geometric model corresponding to the target resolution.

[0117] Optionally, before decoding the first labeled image of at least one target scale among the multiple scales according to at least one target resolution, the processor 502 is further configured to: obtain the expected at least one target resolution according to the geometric model description data; and determine, among the multiple scales, a scale corresponding to the at least one target resolution as the at least one target scale.

[0118] Optionally, the geometric model description data includes: model description data of at least one modality for describing a geometric model of desired output.

[0119] In this embodiment, the autoregressive model performs an autoregressive generation operation on spatial features based on the feature information of the input geometric model description data, accurately outputting multiple labeled maps at different scales. The scale of the labeled map is positively correlated with the resolution of the geometric model, enabling the labeled maps of different scales to express not only the global structural features of the geometric model sample but also the local details of the geometric model. This facilitates the selection of a labeled map of a target scale corresponding to the target resolution for decoding, yielding a geometric model of the desired resolution, thus flexibly meeting the needs of different scenarios.

[0120] In other exemplary embodiments, Figure 5The electronic device shown is used to execute a model training method, wherein the processor 502 is coupled to the memory 501 and is used to execute the computer program in the memory 501, so as to: obtain geometric model samples and corresponding model description data samples; use an encoder to encode the spatial data corresponding to the geometric model samples to obtain reference label maps of multiple scales; perform feature extraction on the model description data samples to obtain target feature information, and use an autoregressive model to perform multiple autoregressive generation operations of spatial features according to the target feature information to obtain second label maps of multiple scales, and the spatial features output by any autoregressive generation operation are stored in the second label map of the corresponding scale, the second label maps of the multiple scales correspond to the multiple resolutions, and the scale of the second label map is positively correlated with the resolution of the geometric model; and train the autoregressive model according to the difference between the reference label maps of the multiple scales and the second label maps of the multiple scales.

[0121] Optionally, when the processor 502 uses the encoder to encode the spatial data corresponding to the geometric model sample to obtain reference label maps of multiple scales, it is specifically used to: use the encoder to organize the spatial data corresponding to the geometric model sample into multiple levels of spatial nodes, and the spatial nodes in different levels represent different spatial ranges; use a cross-scale attention mechanism to compress the spatial nodes of the multiple levels to obtain reference label maps of multiple scales.

[0122] Optionally, when the processor 502 uses the encoder to organize the spatial data corresponding to the geometric model sample into multiple levels of spatial nodes, it is specifically used to: use the encoder to determine an octree structure with a target depth according to the target resolution; recursively divide the spatial data corresponding to the geometric model sample into spatial nodes according to the octree structure until the target depth is reached; organize the spatial nodes obtained by dividing the geometric model sample according to the depth level to which the spatial nodes belong, and obtain the multiple levels of spatial nodes.

[0123] Optionally, when the processor 502 uses a cross-scale attention mechanism to compress the spatial nodes of the multiple levels to obtain reference label maps of multiple scales, it is specifically used to: encode the position data in the spatial nodes of any target level among the multiple levels to obtain the position code corresponding to the spatial nodes of the target level; obtain a compressed query vector corresponding to the target level, the compressed query vector includes a learnable first parameter, and the first parameter is used to learn the knowledge of assigning different weights to the features of the target level according to feature importance to capture key features; perform attention calculation on the position code and the compressed query vector to capture key features from the position code to obtain a first compressed feature corresponding to the spatial node of the target level; quantize the first compressed feature to obtain a reference label map of the target scale corresponding to the target level.

[0124] Optionally, before quantizing the first compressed feature, the processor 502 is further used to: obtain a second compressed feature corresponding to a spatial node of an adjacent level of the target level; sample the second compressed feature to obtain a third compressed feature having the same scale as the first compressed feature; and use a cross-scale attention mechanism to fuse the first compressed feature and the third compressed feature to update the first compressed feature.

[0125] Optionally, when the processor 502 uses the autoregressive model to perform multiple autoregressive generation operations of spatial features according to the target feature information to obtain second label maps of multiple scales, it is specifically used to: when performing any autoregressive generation operation in the autoregressive model, if the current scale corresponding to any autoregressive generation operation is the first scale, map the target feature information into discrete initial spatial features to obtain the second label map of the current scale; if the current scale is not the first scale, perform autoregressive generation according to the target feature information and the second label map of the previous scale of the current scale to obtain the second label map of the current scale.

[0126] Optionally, when the processor 502 performs autoregressive generation based on the target feature information and the second label map of the previous scale of the current scale to obtain the second label map of the current scale, it is specifically used to: linearly map the target feature information so that the target feature information matches the feature dimension of the current scale; splice the linearly mapped target feature information with the reference label map of the previous scale of the current scale to obtain a spliced ​​feature; and predict the second label map of the current scale based on the spliced ​​feature and the learnable model parameters.

[0127] Optionally, the processor 502 is further used to: obtain distribution characteristics of the compression features corresponding to each of the multiple levels, the distribution characteristics including mean and / or variance; obtain a first loss corresponding to the encoder based on the distribution characteristics of the compression features corresponding to each of the multiple levels; obtain a signed distance prediction value of any point in the space based on the second labeling graphs of the multiple scales; obtain a second loss corresponding to the encoder based on the signed distance prediction value and the actual signed distance value of the geometric model sample; and train the encoder with the first loss and the second loss converging to a specified range training target.

[0128] Optionally, when the processor 502 trains the autoregressive model based on the difference between the reference labeled images of the multiple scales and the second labeled images of the multiple scales, the processor 502 is specifically used to: obtain the difference between the reference labeled image and the second labeled image of the same scale based on the scale correspondence; calculate the cross entropy loss between the reference labeled images of the multiple scales and the second labeled images of the multiple scales based on the difference between the reference labeled image and the second labeled image of the same scale; and train the autoregressive model with the cross entropy loss converging to a specified range as an optimization goal.

[0129] In this embodiment, the encoder encodes the geometric model samples to obtain reference marker maps of multiple scales. The reference marker maps at different scales can not only express the global structural features of the geometric model samples, but also capture the local detail features of the geometric model samples. The reference marker maps of multiple scales obtained by the encoder encoding the geometric model samples are used as supervisory signals to train the autoregressive model, so that the autoregressive model can effectively learn the relationship and regularity between feature information of different scales in space based on the reference marker maps of multiple scales, thereby improving the autoregressive model's ability to understand and generate spatial features. Based on the above capabilities, the autoregressive model can accurately output multiple first marker maps of different scales according to the first feature information of the input geometric model description data, and enable the first marker maps of different scales to be decoded to obtain geometric models of different resolutions that meet the expectations, so as to facilitate the selection of the first marker map of the target scale corresponding to the target resolution for decoding according to the needs, and obtain a geometric model that meets the expected resolution, thereby flexibly meeting the needs of different scenarios.

[0130] Further, if Figure 5 As shown, the electronic device also includes other components such as a communication component 503, a power component 504, a display component 505 and an audio component 506. Figure 5 Only some components are shown schematically, which does not mean that the electronic device only includes Figure 5 Components shown. Figure 5Components in the dotted box are optional components, not mandatory components, and may depend on the product form of the electronic device. The electronic device of this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone or an IOT device, or a server device such as a conventional server, a cloud server or a server array. If the electronic device of this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, etc., it may include Figure 5 If the electronic device of this embodiment is implemented as a conventional server, cloud server or server array and other server-side devices, it may not include Figure 5 Components within the dotted box.

[0131] The memory 501 may be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0132] The communication component 503 is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as 2G (such as Global System for Mobile Communications (GSM)), 3G (such as Wideband Code Division Multiple Access (WCDMA), 4G (such as Long Term Evolution (LTE)), 4G+ (such as LTE-Advanced (LTE-A)), or 5G (5th Generation Mobile Communication Technology), or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel.

[0133] The power supply component 504 is used to provide power to various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device where the power supply component is located.

[0134] The display assembly includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, it may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensors can detect not only the boundaries of a touch or slide action, but also the duration and pressure associated with the touch or slide action.

[0135] The audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), which is configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signal can be further stored in a memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0136] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed, can implement the steps that can be performed by the electronic device in the above-mentioned method embodiment. The computer-readable storage medium includes volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, tape, disk storage or other magnetic storage devices, or any other non-transmission medium.

[0137] The present application embodiment also provides a computer program product, including: a computer program / instruction, which, when executed by a processor, can implement the steps in the method provided in the embodiment of the present application. It should be understood that each process or a combination of multiple processes in the above-mentioned method flow can be implemented by a computer program or instruction. In addition, these computer programs or instructions can be applied to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device, so that the processor of the general-purpose computer, the special-purpose computer, the embedded processor or other programmable data processing device can be implemented as a device for implementing the corresponding functions in the above-mentioned method embodiment.

[0138] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, product, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, product, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not preclude the presence of additional identical elements in the process, method, product, or apparatus that includes the element.

[0139] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A model generation method, characterized in that: include: In response to the model generation operation, the geometric model description data is obtained; Performing feature extraction on the geometric model description data to obtain first feature information; Using an autoregressive model, perform multiple autoregressive generation operations of spatial features based on the first feature information to obtain first labeled maps of multiple scales, wherein spatial features generated by different autoregressive generation operations correspond to first labeled maps of different scales; The scale of the first label map is positively correlated with the resolution of the geometric model; decoding, according to at least one target resolution, a first label map of at least one target scale among the plurality of scales to obtain at least one target geometric model; The autoregressive model is trained based on the difference between reference labeled images of multiple scales corresponding to the geometric model sample and second labeled images of multiple scales output by the autoregressive model for the model description data of the geometric model sample.

2. The method according to claim 1, characterized in that Also includes: Obtaining the geometric model sample and the corresponding model description data sample; Encoding the spatial data corresponding to the geometric model sample using an encoder to obtain reference labeled images at multiple scales, wherein the reference labeled images at multiple scales are used to express spatial features of the geometric model sample at multiple resolutions; Performing feature extraction on the model description data sample to obtain second feature information, and performing multiple autoregressive generation operations of spatial features based on the second feature information using an autoregressive model to obtain second labeled maps at multiple scales, wherein the second labeled maps at multiple scales correspond to the multiple resolutions; The autoregressive model is trained based on the differences between the reference label maps at the multiple scales and the second label maps at the multiple scales.

3. The method according to claim 2, characterized in that Encoding the spatial data corresponding to the geometric model sample using an encoder to obtain reference marker images of multiple scales, including: Organizing the spatial data corresponding to the geometric model sample into multiple levels of spatial nodes using an encoder, where the spatial nodes in different levels represent different spatial ranges; A cross-scale attention mechanism is used to compress the spatial nodes of the multiple levels to obtain reference label maps of multiple scales.

4. The method according to claim 1, wherein Using an autoregressive model, performing multiple autoregressive spatial feature generation operations based on the first feature information includes: When any autoregressive generation operation is performed in the autoregressive model, if a current scale corresponding to the any autoregressive generation operation is a first scale, mapping the first feature information into discrete initial spatial features, and using a feature map formed by the discrete initial spatial features as a first labeled map of the current scale; If the current scale is not the first scale, autoregressive generation is performed according to the first feature information and a first labeled map of a previous scale of the current scale to obtain the first labeled map of the current scale.

5. The method according to claim 4, characterized in that Performing autoregressive generation based on the first feature information and a first labeled map of a previous scale of the current scale to obtain the first labeled map of the current scale includes: Performing linear mapping on the first feature information so that the first feature information matches the feature dimension of the current scale; Splicing the first feature information after linear mapping with the reference label image of the previous scale of the current scale to obtain a spliced ​​feature; Autoregressive generation is performed according to the splicing features to obtain a first labeled image of the current scale.

6. The method according to any one of claims 1 to 5, characterized in that Decoding a first label map of at least one target scale among the plurality of scales according to at least one target resolution to obtain at least one target geometric model includes: For any target resolution of the at least one target resolution, obtaining a signed distance function value of any point in the three-dimensional space according to a first labeled map of a target scale corresponding to the target resolution; reconstructing a continuous signed distance field according to the signed distance function of any point in the space; A marching cube algorithm is used to extract a triangulated network from the signed distance field to obtain a target geometric model corresponding to the target resolution.

7. The method according to any one of claims 1 to 5, characterized in that Before decoding the first signature map of at least one target scale among the multiple scales according to at least one target resolution, the method further includes: Obtaining the expected at least one target resolution according to the geometric model description data; Among the multiple scales, a scale corresponding to the at least one target resolution is determined as the at least one target scale.

8. The method according to any one of claims 1 to 5, characterized in that The geometric model description data includes: model description data of at least one mode for describing a geometric model expected to be output.

9. A model training method, characterized in that: include: Obtaining a geometric model sample and a corresponding model description data sample; Encoding the spatial data corresponding to the geometric model sample using an encoder to obtain reference labeled images at multiple scales, wherein the reference labeled images at multiple scales are used to express spatial features of the geometric model sample at multiple resolutions; Performing feature extraction on the model description data sample to obtain target feature information, and using an autoregressive model to perform multiple autoregressive generation operations of spatial features based on the target feature information to obtain target label maps of multiple scales, wherein the spatial features output by different autoregressive generation operations correspond to target label maps of different scales, the target label maps of the multiple scales correspond to the multiple resolutions, and the scale of the target label map is positively correlated with the resolution of the geometric model; The autoregressive model is trained based on the differences between the reference label maps at the multiple scales and the target label maps at the multiple scales.

10. The method according to claim 9, characterized in that Encoding the spatial data corresponding to the geometric model sample using an encoder to obtain reference marker images of multiple scales, including: Organizing the spatial data corresponding to the geometric model sample into multiple levels of spatial nodes using an encoder, where the spatial nodes in different levels represent different spatial ranges; A cross-scale attention mechanism is used to compress the spatial nodes of the multiple levels to obtain reference label maps of multiple scales.

11. The method according to claim 10, characterized in that The spatial data corresponding to the geometric model sample is organized into multiple levels of spatial nodes using an encoder, including: Using the encoder to determine an octree structure with a target depth according to the target resolution; Recursively dividing the spatial data corresponding to the geometric model samples into spatial nodes according to the octree structure until the target depth is reached; The spatial nodes obtained by dividing the geometric model samples are organized according to the depth levels to which the spatial nodes belong, to obtain the spatial nodes of the multiple levels.

12. The method according to claim 10, characterized in that A cross-scale attention mechanism is used to compress the spatial nodes of the multiple levels to obtain reference labeled graphs of multiple scales, including: For any target level among the multiple levels, encoding the position data in the spatial node of the target level to obtain a position code corresponding to the spatial node of the target level; Obtaining a compressed query vector corresponding to the target level, the compressed query vector comprising a learnable first parameter, the first parameter being used to learn knowledge of assigning different weights to features of the target level according to feature importance to capture key features; Performing attention calculation on the position code and the compressed query vector to capture key features from the position code and obtain a first compressed feature corresponding to the spatial node of the target level; The first compressed features are quantized to obtain a reference label map of a target scale corresponding to the target level.

13. The method according to claim 12, characterized in that Before quantizing the first compression feature, the method further includes: Obtaining second compressed features corresponding to spatial nodes of a level adjacent to the target level; Sampling the second compressed feature to obtain a third compressed feature having the same scale as the first compressed feature; A cross-scale attention mechanism is used to fuse the first compressed feature and the third compressed feature to update the first compressed feature.

14. The method according to claim 10, characterized in that The autoregressive model is used to perform multiple autoregressive generation operations of spatial features according to the target feature information to obtain target label maps of multiple scales, including: When any autoregressive generation operation is performed in the autoregressive model, if a current scale corresponding to the any autoregressive generation operation is a first scale, mapping the target feature information into discrete initial spatial features to obtain a target label map of the current scale; If the current scale is not the first scale, autoregressive generation is performed based on the target feature information and a target label map of a previous scale of the current scale to obtain a target label map of the current scale.

15. The method according to claim 14, characterized in that Performing autoregressive generation based on the target feature information and a target labeled map of a previous scale of the current scale to obtain the target labeled map of the current scale includes: Performing linear mapping on the target feature information so that the target feature information matches the feature dimension of the current scale; Splicing the target feature information after linear mapping with the reference marker image of the previous scale of the current scale to obtain a spliced ​​feature; The target label map of the current scale is predicted based on the splicing features and learnable model parameters.

16. The method according to claim 15, characterized in that Also includes: Distribution characteristics of the compression features corresponding to each of the multiple levels may be obtained, where the distribution characteristics include a mean and / or a variance; Obtaining a first loss corresponding to the encoder according to distribution characteristics of the compression features corresponding to each of the multiple levels; Obtaining a predicted signed distance value for any point in space based on the target labeled graphs at the multiple scales; Obtaining a second loss corresponding to the encoder according to the predicted sign distance value and the actual sign distance value of the geometric model sample; The encoder is trained with the first loss and the second loss converging to a specified range training target.

17. The method according to any one of claims 9 to 16, characterized in that: Training the autoregressive model based on differences between the reference labeled maps at the multiple scales and the target labeled maps at the multiple scales includes: According to the scale correspondence, the difference between the reference labeled image and the target labeled image of the same scale is obtained; Calculating a cross entropy loss between the reference labeled images at multiple scales and the target labeled images at the multiple scales based on a difference between the reference labeled image and the target labeled image at the same scale; The autoregressive model is trained with the cross entropy loss converging to a specified range as an optimization goal.

18. An electronic device, characterized in that: include: memory and processor; The memory is used to store one or more computer instructions; The processor is configured to execute the one or more computer instructions to perform the steps of the method according to any one of claims 1 to 17.

19. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 17 can be implemented.

20. A computer program product, characterized in that include: A computer program / instruction, which, when executed by a processor, can implement the steps of the method according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • Image generation method and device, equipment, medium and program product

    CN119313762A

  • Three-dimensional model generation method and apparatus, computer device, and storage medium

    WO2025055514A1