Model generation method, storage medium, and program product
By generating human body shape vectors and clothing vectors, decoding and dividing the mesh model for texturing, the problem of poor fit between clothing and human body is solved, achieving accurate matching of clothing parts and improving the detail expression and generation efficiency of 3D character models.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHASING DREAM TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-21
Smart Images

Figure CN122435103A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of 3D modeling technology, specifically to a model generation method, storage medium, and program product. Background Technology
[0002] In the field of automatic generation of 3D characters, related technologies typically employ generative large models based on diffusion models, directly outputting complete 3D mesh models from input images. However, these technologies usually model and texture the human body and clothing as a whole. Limited by the overall generation resolution, individual components of clothing (such as accessories, folds, ribbons, etc.) suffer from blurred details and unclear textures. Furthermore, because clothing and the human body lack independent geometric constraints within the same generation framework, the resulting fit between clothing and the human body is poor, leading to issues such as clipping and misalignment, which negatively impacts the overall realism and usability of the generated model. Summary of the Invention
[0003] This application provides a model generation method, storage medium, and program product, which can solve at least one of the above-mentioned technical problems.
[0004] On one hand, embodiments of this application provide a model generation method, the method comprising: Based on the shape parameter vector of the target image and the preset shape basis vector, a human body shape vector is generated. The human body shape vector is used to characterize the three-dimensional geometric features of the target character corresponding to the target image. The image feature vector of the target image, the human body feature vector corresponding to the human body shape vector, and the initial clothing vector are decoded to generate at least one clothing vector for a part of the human body. The clothing vectors are decoded to generate clothing mesh models for each part of the garment; The human body mesh model is divided into multiple human body region models, wherein the human body mesh model is obtained by transforming the human body shape vector; Based on the target image, textures are applied to each of the human body region models and the corresponding clothing mesh models to generate the target character model.
[0005] On the other hand, embodiments of this application provide a model generation apparatus, the apparatus comprising: The first generation module is used to generate a human body shape vector based on the shape parameter vector of the target image and a preset shape basis vector. The human body shape vector is used to characterize the three-dimensional geometric features of the target character corresponding to the target image. The second generation module is used to decode the image feature vector of the target image, the human feature vector corresponding to the human body shape vector, and the initial clothing vector to generate at least one clothing vector for a part of the human body. The third generation module is used to decode the clothing vectors of each of the aforementioned parts and generate a clothing mesh model. The partitioning module is used to partition the human body mesh model to obtain multiple human body region models, wherein the human body mesh model is obtained by transforming the human body shape vector; The fourth generation module is used to apply textures to each of the human body region models and the corresponding clothing mesh models based on the target image, so as to generate a target character model.
[0006] On the other hand, embodiments of this application provide a computer-readable storage medium storing a computer program adapted for loading by a processor to execute the model generation method as described in any of the above embodiments.
[0007] On the other hand, embodiments of this application provide a computer device, the computer device including a processor and a memory, the memory storing a computer program, the processor executing the model generation method as described in any of the above embodiments by calling the computer program stored in the memory.
[0008] On the other hand, embodiments of this application provide a computer program product, including computer instructions, which, when executed by a processor, implement the model generation method as described in any of the above embodiments.
[0009] The model generation method provided in this application generates a human body shape vector based on the shape parameter vector of the target image and a preset shape basis vector. This allows the generated human body mesh model to accurately match the character morphological features of the target image, ensuring consistency between the human body model and the target image. It fuses the image feature vector of the target image, the human body feature vector corresponding to the human body shape vector, and a preset initial clothing vector for decoding, generating clothing vectors for different body parts. This achieves the adaptation of clothing features to the image style and the three-dimensional geometric features of the human body, ensuring that the generated clothing matches the clothing style of the target image and the human morphology of the target character, solving the problems of poor fit between clothing and the human body and the disconnect between clothing style and image. The method decodes the clothing vectors for each body part to generate a clothing mesh model, completing the conversion from features to a visualized three-dimensional model. The human body mesh model is then divided into multiple human body region models, and textures are applied to each human body region model and its corresponding clothing mesh model according to the target image to generate the target character model. This achieves accurate matching between clothing parts and human body regions, resulting in higher fit between the model's texture, clothing, and human body regions. This reduces manual design costs, improves the generation efficiency of the target character model, and meets the needs for batch and rapid model generation.
[0010] In other words, by jointly decoding image feature vectors, human body feature vectors, and initial clothing vectors, at least one clothing vector for a human body part is generated. This achieves independent modeling of clothing components, avoiding the resolution limitations caused by generating clothing and the entire human body as a whole, and significantly improving the detail representation of each clothing component (such as fine structures like folds, accessories, and ribbons). By dividing the human body mesh model into multiple human body regions and applying textures to the corresponding clothing vectors, precise matching between clothing components and human body regions is achieved. This avoids texture offset and clipping issues caused by geometric misalignment in the overall texture mapping method, ensuring the fit and structural rationality between clothing and the human body. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram of an example generation system provided in an embodiment of this application.
[0013] Figure 2 This is a flowchart illustrating the model generation method provided in an embodiment of this application.
[0014] Figure 3 This is a schematic diagram of a scenario for the model generation method provided in an embodiment of this application.
[0015] Figure 4 This is a schematic diagram of the structure of the model generation device provided in the embodiments of this application.
[0016] Figure 5 A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0018] The application background of the embodiments of this application will be further explained below.
[0019] Anime-style 3D mesh models have wide applications in games and figurine production. These models typically feature unique styles, rich details, and complex color schemes. Traditional design processes rely on manual operation by modelers, which consumes a significant amount of time sculpting the detailed textures of the model's surface and hand-painting individual parts.
[0020] For example, in a scheme for generating 3D animation models based on a diffusion model, one or more images are input into a generative large model, which outputs a corresponding 3D animation character model, thereby improving the model generation efficiency.
[0021] However, when generating 3D anime models, due to insufficient anime-style data in the training data, the generated 3D anime models usually lean more towards a realistic style, especially in facial features and body proportions. For example, they have rounder facial contours and thicker arms, while anime-style characters often have sharper facial contours and thinner arms, making it difficult to adapt to the anime style. In other words, the output 3D anime models do not conform to the characteristics of anime style. Furthermore, the model is often generated by outputting the body and clothing of the 3D anime model as a whole. Due to the limitation of the overall generation resolution, the individual clothing components lack clear details of anime-style clothing, such as various accessories, folds, and ribbons.
[0022] In view of this, embodiments of this application provide a model generation method, apparatus, storage medium, device, and program product. Specifically, the model generation method of this application embodiment can be executed by a computer device, wherein the computer device can be a terminal or a server, etc. The terminal can be a smartphone, tablet computer, laptop computer, smart TV, wearable smart device, smart vehicle terminal, etc. The terminal can also include a client, which can be a 3D modeling client, browser client, instant messaging client, or applet, etc. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0023] For example, when the model generation method runs on a terminal device, the terminal device may include a display screen and a processor. The display screen is used to present the model generation screen and receive user commands generated by the display screen. The processor is used to store the modeling application, run the program, generate model images, respond to commands, and control the display of the character model on the display screen. When the user operates on the target image or character model through the display screen, the screen can control the local content of the terminal device in response to the received operation commands. The terminal device can provide the graphical user interface to the user in various ways, such as rendering the display on the terminal device's screen or presenting the graphical user interface through holographic projection.
[0024] For example, when this model generation method runs on a server, it can be implemented and executed based on a cloud generation system. A cloud generation system refers to a generation method based on cloud computing. A cloud generation system includes a server and client devices. The main body running the model generation application and the main body displaying the target character model are separate. The storage and execution of the model generation method are completed on the server. The display is completed on the client, which is mainly used for receiving and sending target images and parameter data, as well as displaying the image. For example, the client can be a display device with data transmission capabilities located close to the user, such as a mobile terminal, television, computer, PDA, personal digital assistant, head-mounted display device, etc. However, the terminal device performing the model generation processing is the server in the cloud. The user can operate the client to send instructions to the server. The server controls the generation of the target character model according to the instructions, encodes and compresses the generated target character model data, returns it to the client via the network, and finally, the client decodes and outputs the target character model.
[0025] It should be noted that, in this embodiment, the executing entity of the model generation method can be a terminal device or a server. The terminal device can be a local terminal device or the client device mentioned above in cloud generation. This embodiment does not limit the type of executing entity.
[0026] It is understood that in the specific implementation of this application, user object data, context data and other related data are involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0027] For example, in conjunction with the above description, Figure 1This application illustrates a generation system 1000 for implementing a model generation method, as provided in an embodiment of this application. The generation system 1000 may include at least one terminal 1001, at least one server 1002, at least one database 1003, and a network. The user-held terminal 1001 can connect to different servers via the network. The terminal is any device with computing hardware capable of supporting and executing software application tools corresponding to model generation.
[0028] In possible application scenarios, different terminals 1001 may be served by different servers 1002. Therefore, in order to distinguish the servers 1002 corresponding to different game terminals 1001, the embodiments of this application will use the first and second methods for description. In fact, the servers 1002 corresponding to different game terminals 1001 can be the same server 1002.
[0029] Furthermore, when the generation system 1000 includes multiple terminals, multiple servers, and multiple networks, different terminals can connect to each other through different networks and servers. The network can be a wireless network or a wired network, such as a wireless local area network (WLAN), local area network (LAN), cellular network, 2G network, 3G network, 4G network, 5G network, etc. Additionally, different terminals can also connect to other terminals or servers using their own Bluetooth networks or hotspot networks. Furthermore, the system 100 can include multiple databases coupled to different servers, and can continuously store information related to model generation in the databases as different users create target character models online.
[0030] It should be noted that, Figure 1 The schematic diagram of the generation system shown is merely an example. The generation system 1000 described in this application embodiment is for the purpose of more clearly illustrating the technical solutions of this application embodiment and does not constitute a limitation on the technical solutions provided in this application embodiment. As those skilled in the art will know, with the evolution of generation systems and the emergence of new business scenarios, the technical solutions provided in this application embodiment are also applicable to similar technical problems.
[0031] It should be noted that the triggering operations mentioned in the subsequent detailed description of the model generation method provided in the embodiments of this application can all be regarded as triggering operations performed by the user through a finger or by controlling a medium such as a mouse, keyboard, or stylus. The specific medium used can be determined according to the type of computer device. For example, when the computer device is a touchscreen device such as a mobile phone, tablet computer, or game console, the user can operate on the touchscreen using any suitable object or accessory such as a finger or stylus. When the terminal device is a non-touchscreen terminal device such as a desktop computer or laptop computer, the user can operate using an external device such as a mouse or keyboard.
[0032] The technical solution of this application will be described in detail below through specific embodiments. It should be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0033] In this embodiment, the target character model can be applied in the gaming field as a game character controlled by a player. That is, the player operates the virtual character corresponding to the target character model to perform various game activities in the game scene, such as picking up items, engaging in combat, exploring, or solving puzzles. The target character model can also be applied in the field of figurine production, serving as a figurine model created by a user based on a two-dimensional image. A printer can then print the figurine corresponding to the target character model based on that model. This embodiment does not specifically limit the application of this method.
[0034] Please see Figure 2 , Figure 2 This is a flowchart illustrating the model generation method provided in an embodiment of this application. It should be noted that the steps shown may be executed in a different logical order than those shown in the flowchart. The method may include the following steps: Step 011: Generate a human body shape vector based on the shape parameter vector of the target image and the preset shape basis vector. The human body shape vector is used to represent the three-dimensional geometric features of the target character corresponding to the target image.
[0035] The target image may include a two-dimensional image of the character the user wants to create. For ease of explanation, this application's implementation uses the generation of an anime-style target character model as an example.
[0036] Among them, the shape parameter vector can be a numerical vector obtained from a two-dimensional target image by fitting it through a pre-trained neural network, and can be used as a control coefficient to control the changes in the shape of the anime human body. The shape parameter vector determines the superposition strength (degree of change) of the shape basis vector.
[0037] Among them, the shape basis vector can be the basic vector used to represent the shape features of the anime human body, including the human structural characteristics of the display style (such as anime style) of the target character model that the user wants to create (e.g., the sharp facial contours, slender limbs, exaggerated head-to-body ratio, etc. corresponding to the anime style). Different styles of human body structures correspond to different shape basis vectors.
[0038] Among them, the human body shape vector can be a numerical vector generated by a linear combination of the shape parameter vector and the shape basis vector. It is used to characterize the three-dimensional geometric features of the target character corresponding to the target image and to convert the human body shape parameters into a visualized human body mesh model.
[0039] Specifically, the target image can be fitted to obtain the shape parameter vector of the target image, and the human morphological features in the target image can be transformed into a digital human shape vector to represent the three-dimensional geometric features of the target character corresponding to the target image. The shape basis vector carries the style attributes of the model required by the user. Based on the fusion calculation of the shape parameter vector and the shape basis vector, the generated human shape vector can closely match the target character in the target image while ensuring that the human shape does not deviate from the style required by the user (such as the anime style).
[0040] Step 012: Decode the image feature vector of the target image, the human body feature vector corresponding to the human body shape vector, and the initial clothing vector to generate at least one clothing vector for a part of the human body.
[0041] Among them, the image feature vector can be a feature vector obtained by encoding the target image through a pre-trained image encoder such as the Transformer architecture, and is used to characterize the visual features of the target character's clothing style, design, color, texture, etc.
[0042] Among them, the human body feature vector can be a feature vector obtained by encoding and calculation after converting the human body shape vector into a human body mesh model. It is used to characterize the spatial structural information of the human body, such as its three-dimensional geometric shape, contour, proportion, and limb position.
[0043] The initial clothing vector is a randomly initialized vector that serves as the initial noise and basic seed for clothing generation. This provides diversity for clothing generation and avoids generating completely identical clothing results under the same input conditions.
[0044] Among them, the clothing part vector can include feature vectors corresponding to different clothing parts such as upper garment, lower garment, hair, and accessories, which are used to represent the three-dimensional structure and shape information of each clothing component.
[0045] Specifically, by dimensional alignment and feature fusion of image style features, human body geometric features, and initial clothing vectors, clothing vectors for each body part are obtained. This ensures that clothing generation is simultaneously constrained by image style and human body shape, guaranteeing that the clothing matches both the image style and the human body structure. For example, the feature input obtained by fusing image style features, human body geometric features, and initial clothing vectors can be based on a DiT-based diffusion model. The DiT-based diffusion model replaces the traditional single-head output MLP with a multi-head MLP expert head, corresponding to body parts such as upper garments, lower garments, hair, and accessories. Based on the Transformer attention mechanism, the image vector features of the target image (corresponding to clothing) are associated with the human body feature vector (representing human body geometric features). In each diffusion iteration, the features are denoised and optimized. Finally, each expert head independently calculates and outputs the clothing features corresponding to the body part, thus obtaining at least one clothing vector matching the body part.
[0046] Step 013: Decode the clothing vectors for each part to generate a clothing mesh model.
[0047] Among them, the clothing mesh model can be a three-dimensional mesh model generated by transforming the clothing part vectors, corresponding to clothing parts such as tops, bottoms, and hair.
[0048] Specifically, since the feature vectors of clothing parts cannot be directly used for visualization modeling, the feature vectors can be mapped to geometric modeling through the VAE decoder. The abstract features are decoded into a directed distance field (SDF) that can represent the three-dimensional structure of the clothing. The SDF can accurately describe the three-dimensional surface of the clothing in a mathematical way. Then, through the isosurface extraction algorithm, the continuous three-dimensional surface of the clothing is restored from the SDF and transformed into a structured clothing mesh model.
[0049] For example, the clothing vectors of each part can be input into a pre-trained VAE decoder. The VAE decoder decodes and calculates the clothing vectors of each part, transforming the abstract feature vectors into the SDF of the corresponding clothing part to represent the three-dimensional geometry of the clothing part (such as folds, curvature, contour, etc.). Then, the isosurface extraction (Marching Cube) operation is performed on the SDF of each clothing part to extract the clothing surface with a distance value of 0 in the SDF and transform it into a three-dimensional mesh structure composed of vertices, edges, and faces, thereby generating independent clothing mesh models such as upper garment mesh, lower garment mesh, arm mesh, and accessory mesh, completing the conversion from features to a visualized three-dimensional model.
[0050] Step 014: Divide the human body mesh model into multiple human body region models, where the human body mesh model is obtained by transforming the human body shape vector.
[0051] The human body mesh model can be a three-dimensional human body model generated by analyzing the three-dimensional coordinates of the vertices obtained from the analysis of the human body shape vector and combining them with a preset fixed topological connection relationship. It does not include clothing structure and only represents the geometric structure of the human body itself.
[0052] Among them, the human body region model is a sub-mesh model obtained by dividing the complete human body mesh model according to parts, semantics or functions, such as body region, face region, facial features region, limb region, etc., and each region is an independent mesh unit.
[0053] Specifically, the human shape vector corresponds to a human mesh model with a fixed topology. By dividing the mesh according to semantic parts, the overall model can be split into multiple independent human region models. For example, the human shape vector can be converted into a human mesh model first, and then the human mesh model can be divided into multiple local regions according to the preset semantic parts rules. For example, it can be divided into body region, face region, eye region, mouth region, limb region, etc., to obtain multiple independent human region models.
[0054] Step 015: Based on the target image, apply textures to each human body region model and the corresponding clothing mesh model to generate the target character model.
[0055] The target character model can be a complete 3D character model that combines the style required by the user (such as anime style) with clothing and body fit.
[0056] It is understandable that the target image contains the human body features of the target character (such as the shape and proportion of the hands, face, and limbs), clothing style features (such as clothing color, texture, and style), and the correspondence between the human body and clothing. The target image intuitively presents the visual style of each part of the human body, the appearance details of each component of the clothing, and the spatial correspondence between the human body parts and clothing. This provides clear style, form, and correspondence constraints for texture generation, ensuring that the texture is visually consistent with the target character.
[0057] Specifically, based on the division of human body region model and clothing mesh model into parts, the method of independent texturing of each part is adopted. With the semantics of human body parts and clothing features in the target image as constraints, textures that fit the corresponding mesh models are generated to avoid blurring and misalignment of details caused by overall texturing. After the texturing is completed, the human body region model and clothing mesh model are precisely fitted according to the correspondence between human body and clothing in the target image, and finally a target character model that is highly consistent with the visual style and morphological features of the target image is generated.
[0058] In other words, using the target image as a visual reference standard provides the constraints required for texturing (e.g., details of human body parts, clothing style, correspondence between human body and clothing, etc.). Therefore, texturing can be completed based on the target image first, and then combined with part-specific texturing to ensure that the texturing fits the 3D mesh and the visual effect is natural.
[0059] It should be noted that the texturing of each human body region model and the texturing of each clothing mesh model are two independent processes. The independent texturing method, which separates the parts and processes, can further optimize the texture resolution for human body parts (such as the face and hands) and fine clothing parts (such as accessories and folds), improve the model's detail, and avoid problems such as texture blurring caused by overall texturing.
[0060] For example, we can first use the semantic information and visual features (such as facial texture and hand details) of the corresponding human body parts in the target image as constraints to generate human body part textures that match each human body region model (such as the hand region model corresponding to the hand texture in the target image, and the face region model corresponding to the face texture in the target image). Then, we can attach each generated human body part texture to the surface of the corresponding human body region model to complete the human body mesh texture and obtain each texture human body region sub-model. Next, we can use the style, color, and texture features of the corresponding clothing in the target image, as well as the surface normal vector of the clothing mesh model (representing the concavity and folds of the clothing) as constraints to generate clothing part textures that match each clothing mesh model. Then, we can attach each generated clothing part texture to the surface of the corresponding clothing mesh model to complete the clothing mesh texture and obtain each texture clothing sub-model. Finally, based on the correspondence between the human body and clothing in the target image, we can combine and integrate each texture human body region sub-model with the corresponding texture clothing sub-model according to human body parts to finally generate the target character model.
[0061] Thus, by generating human shape vectors based on the shape parameter vectors of the target image and the preset shape basis vectors, the generated human body mesh model can accurately match the character morphological features of the target image, ensuring the consistency between the human body model and the target image. By fusing the image feature vectors of the target image, the human feature vectors corresponding to the human body shape vectors, and the preset initial clothing vectors for decoding, clothing vectors for different body parts are generated. This achieves the adaptation of clothing features to image style and human 3D geometric features, ensuring that the generated clothing matches the clothing style of the target image and the human morphology of the target character, solving the problems of poor clothing fit and clothing style disconnect from the image. Decoding the clothing vectors for each body part generates a clothing mesh model, completing the conversion from features to a visualized 3D model. The human body mesh model is then divided into multiple human body region models, and textures are applied to each human body region model and its corresponding clothing mesh model according to the target image to generate the target character model. This achieves accurate matching between clothing parts and human body regions, resulting in higher fit between the model's texture, clothing, and human body regions, reducing manual design costs, improving the generation efficiency of the target character model, and meeting the needs for batch and rapid model generation.
[0062] In other words, by jointly decoding image feature vectors, human body feature vectors, and initial clothing vectors, at least one clothing vector for a human body part is generated. This achieves independent modeling of clothing components, avoiding the resolution limitations caused by generating clothing and the entire human body as a whole, and significantly improving the detail representation of each clothing component (such as fine structures like folds, accessories, and ribbons). By dividing the human body mesh model into multiple human body regions and applying textures to the corresponding clothing vectors, precise matching between clothing components and human body regions is achieved. This avoids texture offset and clipping issues caused by geometric misalignment in the overall texture mapping method, ensuring the fit and structural rationality between clothing and the human body.
[0063] In some implementations, the method further includes: Step 015: Fit the target image according to the preset mapping relationship between image features and human body shape to obtain the shape parameter vector.
[0064] The pre-defined mapping relationship between image features and human body shape can be obtained by training and learning a neural network on a large number of anime 2D images and anime 3D human body shape sample pairs. This mapping relationship can be used to characterize the correspondence between the visual features of the target image and the three-dimensional shape parameters of the anime human body.
[0065] Specifically, a neural mesh model can be trained to learn the correspondence between image features in a two-dimensional image and the shape of an anime human body. By inputting the target image into the neural mesh model, the neural mesh model fits the target image to output an n-dimensional shape parameter vector, thereby significantly improving the generation efficiency of the shape parameter vector. Moreover, this fitting method can accurately capture the anime human body morphological features in the target image, ensuring that the output shape parameter vector is highly matched with the target image.
[0066] For example, a neural network model can be built by inputting a massive amount of 2D anime images and corresponding 3D human body shape parameters as training samples. The neural network model can then iteratively learn the mapping relationship between the image features of the 2D images and the human body shape parameters, thus completing the model pre-training. Then, the user-input target image can be fed into the pre-trained neural network, and the model will automatically extract visual features such as the target character's body shape, proportions, and posture from the target image.
[0067] In some implementations, the method further includes: Step 016: Obtain the human body model matrix, which is used to represent the shape information of the three-dimensional human body model; Step 017: Analyze and process the human body model matrix to generate M eigenvalues and M eigenvectors. The eigenvalues and eigenvectors correspond one-to-one. M is a positive integer and greater than 0. M is determined based on the number of dimensions of the vertex coordinates of the human body 3D model. Step 018: Determine the eigenvector corresponding to the target eigenvalue as the shape basis vector. The target eigenvalue includes the first n largest eigenvalues among the M eigenvalues.
[0068] Optionally, M eigenvalues and M eigenvectors are generated using the Skinned Multi-Person Linear model (SMPL, a parametric, vertex-skinned linear 3D human modeling method) based on principal component analysis (PCA).
[0069] The human body model matrix can be an N×M dimensional numerical matrix constructed from the shape data of N anime-style human body 3D models (without clothing) with the same number of vertices and the same topological structure. Each row corresponds to a 1×M row vector of the three-dimensional x / y / z coordinates of the V vertices of an anime human body model, which are expanded in a fixed order. It is a standardized data carrier representing the shape information of a batch of anime human bodies, and is denoted as an N*3V matrix (N is the number of models, and V is the number of vertices in a single model).
[0070] Among them, the feature value can be used to measure the importance of the anime human shape information carried by the corresponding feature vector. The larger the feature value, the better the corresponding feature vector (shape feature) matches the morphological features of the anime human body required by the user.
[0071] The target feature value can be the top n largest feature values selected from M feature values after sorting them from largest to smallest, where n ranges from 10 to 200. The corresponding feature vector is the core feature vector that determines the shape of the anime human body.
[0072] Among them, the feature vector can be used to represent a common shape change pattern in a 3D human body model, such as the height, weight, shoulder width, head-to-body ratio, limb proportions and other model shape differences of the human body model; the feature vector dimension can be 1×M, representing the shape change pattern of the anime human body.
[0073] Where M is a positive integer, the value of which is determined by the number of dimensions of the vertex coordinates of the human body 3D model. For example, if a single human body 3D model includes V vertices, and each vertex contains three-dimensional coordinates of x, y and z, then M = 3V, which is the number of columns in the human body model matrix.
[0074] Specifically, the human body model matrix contains a massive amount of shape information of anime human bodies, which contains a large number of redundant and secondary shape differences. By standardizing the human body model matrix and solving the covariance matrix through PCA, the high-dimensional shape data can be decomposed into eigenvalues and eigenvectors ordered by importance, realizing the extraction of core shape features and data dimensionality reduction, retaining only the feature information that plays a decisive role in the shape of anime human bodies. Based on the SMPL parametric 3D human body modeling method (which has the characteristics of fixed topology and vertex skinning, and can generate human body models with different shapes but consistent topology), the basic data of the real-life style is discarded and replaced with a human body model matrix constructed from a massive amount of anime-style human body 3D models. Then, PCA analysis is used to generate eigenvalues and eigenvectors. The selected shape basis vectors naturally carry the core morphological features of anime human bodies (such as sharp facial contours, slender limbs, exaggerated head-to-body ratio, etc.). The magnitude of the eigenvalue corresponds to the importance of the shape feature. By selecting the eigenvectors corresponding to the n largest eigenvalues as shape basis vectors, the shape variation features of anime human bodies can be fully preserved while reducing computational complexity. The number of n determines the detail expressiveness of the human body template. The larger the number, the richer the detail of the generated human body model.
[0075] Optionally, by relying on the parametric framework of the SMPL modeling method, and combining the PCA principle within the framework, the anime-style human body model matrix is analyzed and processed to generate M eigenvalues and M corresponding eigenvectors. The mature modeling system of SMPL is used to improve the standardization and feasibility of anime human body template construction.
[0076] Thus, by performing PCA analysis on the standardized human body model matrix, the eigenvectors corresponding to the top n largest eigenvalues are selected as shape basis vectors, eliminating minor and accidental shape differences, and ensuring that the shape basis vectors carry the stylistic features of the anime human body (such as sharp faces and slender limbs). The one-to-one correspondence between eigenvalues and eigenvectors and the selection rules for basis vectors make the process of obtaining shape basis vectors standardized and reproducible, avoiding the randomness of manual selection, and providing a unified and reliable shape basis for the parametric generation of anime human body shape vectors.
[0077] In some implementations, step 011: generating a human body shape vector based on the shape parameter vector of the target image and a preset shape basis vector, includes: Step 0111: Linearly combine the shape parameter vector and the shape basis vector to generate the human body shape vector.
[0078] Specifically, the overall three-dimensional shape of the human body model can be seen as a weighted superposition of several basic shape components (i.e., shape basis vectors); and each value in the shape parameter vector can represent the weight of the corresponding basis shape in the final human body model. Through linear combination, the relatively low-dimensional shape basis vector can be used to control the relatively high-dimensional vertex coordinates (i.e., shape parameter vectors), and the human body shape matching the target image can be generated quickly while maintaining the consistency of the animation style.
[0079] In other words, by taking the i-th parameter pi in the shape parameter vector and multiplying it with the corresponding i-th shape basis vector bi, we can obtain the weighted basis components (pi×bi). Then, we can sequentially perform addition and summation on all weighted components from i=1 to n: S=p1×b1+p2×b2+...+pn×bn. The resulting vector S is the final human body shape vector. That is, by multiplying each basis vector by its corresponding shape parameter and then adding all the results together, we finally obtain the digital vector of the new human body shape (the vector corresponds to the coordinates of the 3V vertices of the human body, which can be directly converted into a 3D mesh model). The linear operation between the shape parameter vector and the shape basis vector significantly reduces the computational complexity of human body shape generation and improves the generation efficiency. At the same time, the linear combination method can flexibly adjust the values of each shape parameter, realize precise control of the anime human body shape features (such as adjusting the sharpness of the face, the proportion of the limbs, etc.), and generate diverse human body shapes that conform to the anime style.
[0080] In some implementations, step 012: Decoding the image feature vector of the target image, the human feature vector corresponding to the human body shape vector, and the initial clothing vector to generate at least one clothing vector for a human body part, including: Step 0121: Fuse the image feature vector, human body feature vector, and initial clothing vector to generate the input feature vector, wherein the image feature vector is generated based on the target image encoding, and the human body feature vector is generated based on the human body mesh model transformation; Step 0122: Based on the attention mechanism, perform feature capture on the input feature vector to obtain at least one human body part clothing vector, and match the human body part clothing vector with the preset human body part.
[0081] The attention mechanism is used to extract information related to clothing generation and human body part segmentation from image feature vectors, human body feature vectors and initial feature vectors, so as to realize the correlation modeling between image style, human body structure and clothing parts.
[0082] Optionally, the Transformer self-attention mechanism of the Diffusion Transformer Model (DiT) can be used to capture the correlation between image clothing features, human body geometric features and clothing parts in the input feature vector, thereby achieving accurate feature matching and extraction.
[0083] Optionally, by replacing the single-head MLP output at each step of the DiT model with a multi-head MLP, a hybrid expert architecture (MoE architecture) is obtained to achieve independent generation of vectors for each clothing part. A multi-head MLP consists of multiple independent multilayer perceptrons, each corresponding to a preset human clothing part such as upper garment, lower garment, hair, and accessories. Each head is independently responsible for feature calculation and vector generation for its corresponding clothing part.
[0084] Among them, the clothing vectors of human body parts and the clothing feature vectors corresponding to preset human body parts (such as upper body, lower body, hair, accessories, etc.) are used to represent the clothing structure and texture information of the corresponding parts. The clothing vectors of human body parts can be decoded into directed distance fields (SDF) or undirected distance fields (UDF) by a VAE decoder.
[0085] Optionally, the human body feature vector can be a feature vector generated by sampling the human body mesh model through point cloud and then calculating it through a pre-trained VAE (variational autoencoder) framework clothing 3D structure encoder. It represents the geometric structure information of the human body, such as the three-dimensional contour, proportion, and limb position, and can be regarded as the human body shape adaptation constraint condition generated by clothing.
[0086] The initial clothing vector can be a randomly initialized clothing structure vector, which can be regarded as the initial noise seed for clothing generation, providing diversity for clothing generation.
[0087] Specifically, by fusing image style information, human body geometry information, and initial clothing information, an input feature vector is generated. This input feature vector represents that clothing generation is simultaneously constrained by both image content and human body structure (three-dimensional human morphology), ensuring that the generated clothing matches the target image style and accurately adapts to the human body structure, thus improving the fit between clothing and the human body. Furthermore, an attention mechanism automatically learns and captures the correspondence between image features, human body features, and clothing parts, achieving end-to-end part-specific clothing feature generation. This ensures that the clothing style matches the image and the structure adapts to the human body. For example, by using a DiT diffusion model as the backbone network, [the following is achieved]. The Transformer attention mechanism is used to accurately capture the correlation information in the input feature vector. By leveraging the denoising properties of the diffusion model, iterative optimization is performed from the random initial clothing vector to generate clothing features that meet the dual constraints. Then, the DiT model is modified using the MoE architecture. The single-head MLP output of each step of the DiT model is replaced with a multi-head MLP that corresponds one-to-one with the preset clothing parts. This allows each clothing part to be handled by an independent MLP head for feature calculation and vector generation, breaking the limitation of the overall generation resolution, improving the detail representation of individual clothing part vectors, and solving the problem of missing details in clothing components in existing technologies.
[0088] The image feature vector of the target image, the human body feature vector corresponding to the human body shape, and the randomly initialized initial clothing vector are fused into a unified input feature vector. This allows the generation process of clothing part vectors to be constrained by both the clothing style in the image and the human body shape adaptation constraint. This ensures that the generated clothing part vectors not only highly match the clothing style and pattern in the target image, but also accurately adapt to the three-dimensional geometric features of the human body (such as shoulder width, waist size, head-to-body ratio, etc.), solving the problems of poor clothing fit and clothing style disconnect from the target image. At the same time, the attention mechanism based on the diffusion model independently captures features for each part, allowing the generation process of each clothing part vector to be independent of each other. This breaks the limitation of the overall generation resolution, improves the detail expression of individual clothing part vectors, and can accurately reproduce the features of clothing components such as tops, bottoms, and accessories.
[0089] In some implementations, the method further includes: Step 019: Convert the human body mesh model into human body point cloud data through point cloud sampling; Step 020: Calculate the human body point cloud data using an encoder to generate human body feature vectors.
[0090] Among them, the human body mesh model can be a three-dimensional human body model generated by parsing the three-dimensional coordinates of the vertices based on the human body shape vector and according to a preset fixed topological relationship, used to represent the three-dimensional geometric shape of the target character.
[0091] Human point cloud data can be a set of three-dimensional coordinate points of the human body obtained by point cloud sampling, which is used to characterize the spatial structural features of the human body such as outline, proportion, and limb position.
[0092] The encoder can be a clothing 3D structure encoder under the VAE framework, used to extract and calculate features from human point cloud data, and to transform discrete point cloud data into standardized human feature vectors to represent the three-dimensional geometric features of the human body.
[0093] Specifically, the structured human body mesh model is transformed into lightweight point cloud data through point cloud sampling, stripping away redundant topological information and retaining core geometric features; then, the discrete human body point cloud data is processed by an encoder to extract features and map dimensions, transforming the unstructured point cloud data into standardized human body feature vectors that can be used for subsequent model calculations, thus providing precise human body shape constraints for clothing generation.
[0094] For example, the generated human body mesh model can first undergo point cloud sampling processing to convert the mesh-based 3D human body model into discrete human body point cloud data, preserving the 3D geometric information of the human body such as contours, proportions, and limb positions. The obtained human body point cloud data is then input into an encoder, which extracts and calculates features from the data, outputting standardized human body feature vectors. These feature vectors can be used for subsequent fusion with image feature vectors and initial clothing vectors.
[0095] Thus, by sampling the human body mesh model and converting it into human body point cloud data, redundant topological connections in the human body mesh model can be removed, retaining only the core three-dimensional geometric coordinate information of the human body. At the same time, by sampling the vertex data, the problem of excessive computational load and high computing power consumption caused by directly using the human body mesh model with all vertices for feature calculation is avoided, which greatly improves the computational efficiency of subsequent human body feature extraction. During the point cloud sampling process, the core three-dimensional geometric features of the human body, such as the outline, limb proportions, and key part shapes, are completely preserved, and only minor vertex deviations that are meaningless for clothing adaptation are removed, ensuring that the human body point cloud data can realistically and accurately represent the overall morphological features of the human body mesh model. The encoder calculates and generates human body feature vectors from the human body point cloud data, which can transform the discrete three-dimensional coordinate information of the human body point cloud into a standardized feature vector form that can be recognized and fused by the subsequent clothing generation model.
[0096] In some implementations, the image feature vector includes a clothing image feature vector. Step 015: Based on the target image, textures are applied to each of the human body region models and the corresponding clothing mesh models to generate a target character model, including: Step 0151: Using a multi-view diffusion model, generate corresponding human body part textures based on the semantic information of each human body region model; Step 0152: Apply the textures of each human body part to the corresponding human body region model to generate a textured human body mesh model. Step 0153: Using a multi-view diffusion model, generate clothing part textures based on the surface normal vectors of each clothing mesh model and the feature vectors of the clothing image; Step 0154: Apply the textures of each clothing part to the clothing mesh model corresponding to the human body part to generate a textured clothing mesh model. Step 0155: Apply the textured human body mesh model and textured clothing mesh model to the human body parts to generate the target character model.
[0097] Among them, the clothing image feature vector can be used to characterize the visual features such as style, color, texture, and design of clothing in the target image.
[0098] Among them, the multi-view diffusion model can be a generative model used to generate textures for human body parts and clothing parts. It can generate textures that fit the surface of the 3D model from multiple perspectives, avoiding problems such as stretching, deformation and misalignment of textures.
[0099] The semantic information of the human body region model can be the part attribute information corresponding to the human body region model, that is, the information used to distinguish different human body regions such as head, torso, and limbs, which can provide regional constraint information for the generation of human body part textures.
[0100] Among them, the human body part texture map can be a texture map generated based on a multi-view diffusion model and matched with the corresponding human body region model, which is used to attach to the surface of the human body region model.
[0101] Among them, the textured human body mesh model can be a complete human body mesh model after the textures of each human body part have been attached.
[0102] Among them, the surface normal vector can be used to characterize the geometric features of the clothing mesh model surface, such as bumps, curves, and wrinkles, and is used to improve the fit between the clothing part texture and the clothing mesh surface.
[0103] Among them, the clothing part texture map can be a texture map generated based on the multi-view diffusion model and matched with the corresponding clothing mesh model.
[0104] Among them, the textured clothing mesh model can be a complete clothing mesh model after the textures of each clothing part have been attached.
[0105] The target character model can be a complete 3D character model generated by attaching textured human body mesh models and textured clothing mesh models according to human body parts.
[0106] Specifically, semantic information can be used as a constraint to generate human body part textures that match the corresponding regions through a multi-view diffusion model, ensuring that the human body textures are adapted to the attributes of the parts. At the same time, clothing image feature vectors are used as style constraints and surface normal vectors are used as geometric constraints to generate clothing part textures that fit the clothing structure through a multi-view diffusion model. After completing the human body part textures and clothing part textures respectively, the two types of models are combined according to the spatial correspondence of the human body parts to finally form a complete target character model.
[0107] For example, by calling a multi-view diffusion model, human body part textures matching each human body region model can be generated based on the semantic information corresponding to each human body region model. These textures are then attached to the surfaces of their respective human body region models, and integrated to obtain a textured human body mesh model. Next, the clothing vectors for each part are transformed to generate clothing mesh models corresponding to each clothing part. Due to the complexity of clothing structures, clothing surface normals are added as input to the multi-view diffusion model to improve the fit between the generated textures and the clothing surface. The multi-view diffusion model is then used in conjunction with the surface normals of each clothing mesh model and the clothing image feature vectors to generate clothing part textures matching each clothing mesh model. These clothing part textures are then attached to the surfaces of their respective clothing mesh models, and integrated to obtain a textured clothing mesh model. Finally, according to the correspondence of human body parts, the textured human body mesh models and textured clothing mesh models are precisely fitted together to generate the target character model.
[0108] Thus, independent texturing of the human body mesh model and the clothing mesh model avoids problems such as texture blurring, stretching deformation, and misalignment caused by overall texturing. Generating textures for the human body region model using a multi-view diffusion model, combined with semantic information, ensures a high degree of matching between the human body texture and the features of each region. Furthermore, the multi-view generation method ensures that the human body texture appears natural and undistorted in 3D, improving the fit and clarity of the human body texture. For the clothing mesh model, texturing is generated by combining surface normal vectors and clothing image feature vectors. Surface normal vectors accurately describe the complex geometric features of clothing, such as concavity, folds, and curvature, allowing the texture of the clothing texture to follow the distribution of the clothing structure, solving the problem of poor fit of textures for complex clothing structures. Simultaneously, the clothing image feature vectors ensure a high degree of consistency between the clothing texture and the clothing style of the target image. Finally, the textured human body and clothing mesh models are precisely fitted according to human body parts, achieving seamless integration of clothing and human body. The generated target character model possesses high-definition textures, a distinctive anime style, and a high degree of fit between clothing and human body.
[0109] In some implementations, the method further includes: Step 021: Analyze the human body shape vector to obtain a three-dimensional coordinate set, which includes each vertex corresponding to the human body shape; Step 022: Connect the vertices according to the preset human body topology to generate a human body mesh model.
[0110] Among them, the three-dimensional coordinate set can be the x, y, z three-dimensional coordinate data of all vertices obtained after analysis, which constitutes the basic point information of the three-dimensional shape of the human body.
[0111] The preset human topology can be a pre-defined, fixed connection rule between vertices, including the connection methods of edges and faces between vertices, used to construct a complete mesh from discrete vertices.
[0112] Specifically, the human body shape vector compactly stores the three-dimensional coordinate information of all vertices in one-dimensional numerical form. By parsing according to fixed rules, the spatial position of all vertices can be reconstructed. Then, using unified and fixed human body topological relationships to connect the vertices, the discrete coordinate points are transformed into a continuous and structured three-dimensional human body mesh model, realizing the conversion from digital shape vectors to visualized three-dimensional models. For example, by parsing the human body shape vector segment by segment, according to fixed coordinate arrangement rules, the values in the vector are split and combined to obtain the three-dimensional coordinate set of all vertices corresponding to the human body shape. Then, according to the pre-set human body topological relationships, each vertex in the three-dimensional coordinate set is connected sequentially to construct a complete three-dimensional model composed of vertices, edges, and faces, i.e., the human body mesh model.
[0113] Thus, by standardizing the analysis of human shape vectors, the set of three-dimensional coordinates of all vertices of the animated human body that matches the target image can be accurately obtained, ensuring that the morphological features of the human body mesh model are consistent with the target image. At the same time, by connecting vertices according to fixed human body topological relationships, the topological structure of the generated human body mesh model is consistent with that of the animated human body model that constructs the shape basis vectors. This allows subsequent steps such as point cloud sampling, feature extraction, and clothing fitting to be carried out based on a unified topological foundation, avoiding problems such as poor clothing fitting and feature extraction errors caused by inconsistent topological structures. This generation method realizes the automated and accurate conversion from digital vectors of human body shape to a visualized 3D mesh model without manual modeling, which greatly improves the generation efficiency of human body mesh models. Moreover, the conversion process is entirely based on numerical analysis and topological reuse, ensuring the reproducibility and standardization of model generation.
[0114] Please see Figure 3 To better understand the model generation method of this application, we will take the generation of an anime-style character mesh model in response to a user-input image (target image) as an example for explanation.
[0115] After the user inputs a single target image of an anime character, the process can be divided into two parallel branches, which respectively complete human body morphology modeling and image style feature extraction.
[0116] First, the target image can be input into a pre-trained neural network fitting model. The pre-trained neural network fitting model learns the mapping relationship between image features and the shape of the anime human body, and fits the shape parameter vector. Then, based on the linear combination of the shape parameter vector and the preset shape basis vector, a human body mesh model (containing only the three-dimensional geometric shape of the target character, without clothing) is generated.
[0117] Simultaneously, the target image is input into an image encoder (a pre-trained encoder based on the Transformer architecture) to extract visual features such as style, color, texture, and design of clothing in the image, generating an image feature vector.
[0118] Next, the generated human body mesh model is sampled from point clouds (removing redundant topological information and retaining core geometric features) and then input into a VAE encoder (clothing 3D structure encoder) to extract three-dimensional geometric features such as the human body contour, proportions, and limb positions, generating a human body feature vector. At the same time, randomly initialized clothing vectors are prepared (as noise seeds for clothing generation, providing diversity to the results). Finally, the image feature vector, human body feature vector, and randomly initialized clothing vector are fused to form a unified input feature vector, which serves as a constraint condition for the subsequent clothing generation model.
[0119] Next, the fused input feature vector is fed into the DiT model and VAE decoder module. The DiT model accurately captures the correlation between image style features and human geometric features (such as the matching relationship between shoulder width features and upper garment shoulder lines) through a self-attention mechanism. Combined with random initial clothing vectors, the feature vectors of each clothing part are generated through the gradual denoising iteration of the diffusion model. After being decoded by the VAE decoder, the SDF (directed distance field) of each clothing part (such as upper garment SDF, lower garment SDF, etc.) is output, representing the 3D geometric shape (contour, folds, curvature, etc.) of each clothing part.
[0120] Next, an isosurface extraction (marchingcube algorithm) operation is performed on the SDF of each clothing part, that is, the surface with a distance value of 0 in the SDF is extracted and transformed into a 3D mesh structure composed of vertices, edges and faces. Finally, independent clothing part mesh models such as upper garment mesh and lower garment mesh are generated, and these models are adapted to the morphological characteristics of the human body mesh.
[0121] Finally, based on the semantic information of each human body region (such as face, torso, limbs, etc.), matching human body part textures are generated and attached to obtain textured human body mesh textures; based on image feature vectors (style constraints) and clothing mesh surface normal vectors (geometric constraints, describing the concavity, folds, and other features of clothing), matching clothing part textures are generated and attached to obtain textured clothing mesh textures; the textured human body mesh model and clothing part textures are precisely fitted and integrated according to the spatial relationship of human body parts, and finally a complete anime character mesh model is generated. The anime character mesh model is a 3D character model that combines anime style, high-detail textures, and clothing that fits the human body very closely.
[0122] Compared with traditional 3D anime character development, the implementation method of this application does not require manual design of 3D anime character models, which greatly reduces the time consumption of batch design; at the same time, it can generate highly customized models based on a single image, and can also generate multiple models with detailed differences from a single image at one time, providing users with a wider range of choices.
[0123] Compared with the commonly used general generative 3D large models on the market, the implementation method of this application introduces a human body shape template based on anime data, which can produce character models that are more in line with the human body structure of anime style, such as more exaggerated facial contours and limb proportions. At the same time, the hybrid expert architecture adopted by this solution can independently generate each part of the clothing when generating clothing models, which improves the expressiveness of the model and brings more clothing details.
[0124] To facilitate better implementation of the model generation method of this application, this application also provides a model generation apparatus. Please refer to... Figure 4 , Figure 4 A schematic diagram of the structure of the model generation apparatus provided in an embodiment of this application. The model generation apparatus 200 may include: The first generation module 201 is used to generate a human body shape vector based on the shape parameter vector of the target image and a preset shape basis vector. The human body shape vector is used to characterize the three-dimensional geometric features of the target character corresponding to the target image. The second generation module 202 is used to decode the image feature vector of the target image, the human body feature vector corresponding to the human body shape vector, and the initial clothing vector to generate at least one part of the human body clothing vector. The third generation module 203 is used to decode the clothing vectors of each part and generate a clothing mesh model. The partitioning module 204 is used to partition the human body mesh model to obtain multiple human body region models, wherein the human body mesh model is obtained by transforming the human body shape vector; The fourth generation module 205 is used to apply textures to each human body region model and the corresponding clothing mesh model based on the target image, so as to generate the target character model.
[0125] Each unit in the aforementioned model generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These units can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each unit.
[0126] The model generation device 200 can be integrated into a terminal or server that has storage and a processor and thus computing power, or the model generation device 200 can be the terminal or server.
[0127] Optionally, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0128] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. The computer device may be a terminal or a server. Figure 5 As shown, the computer device 300 includes a processor 301 with one or more processing cores, a memory 302 with one or more computer-readable storage media, and a computer program stored in the memory 302 and executable on the processor. The processor 301 is electrically connected to the memory 302. Those skilled in the art will understand that the computer device structure shown in the figures does not constitute a limitation on the computer device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0129] The processor 301 is the control center of the computer device 300. It connects various parts of the computer device 300 through various interfaces and lines. By running or loading software programs and / or modules stored in the memory 302, and calling data stored in the memory 302, it performs various functions of the computer device 300 and processes data, thereby performing overall processing of the computer device 300.
[0130] In this embodiment, the processor 301 in the computer device 300 loads the instructions corresponding to the processes of one or more computer programs into the memory 302 according to the following steps, and the processor 301 runs the computer programs stored in the memory 302 to realize various functions: Based on the shape parameter vector of the target image and the preset shape basis vector, a human body shape vector is generated. The human body shape vector is used to characterize the three-dimensional geometric features of the target character corresponding to the target image. The image feature vector of the target image, the human body feature vector corresponding to the human body shape vector, and the initial clothing vector are decoded to generate at least one clothing vector for a part of the human body. Decode the clothing vectors of each of the aforementioned parts to generate a clothing mesh model; The human body mesh model is divided into multiple human body region models, wherein the human body mesh model is obtained by transforming the human body shape vector; Based on the target image, textures are applied to each of the human body region models and the corresponding clothing mesh models to generate the target character model.
[0131] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0132] Optional, such as Figure 5 As shown, the computer device 300 also includes: a display screen 303, a radio frequency circuit 304, an audio circuit 305, an input unit 306, and a power supply 307. The processor 301 is electrically connected to the display screen 303, the radio frequency circuit 304, the audio circuit 305, the input unit 306, and the power supply 307. Those skilled in the art will understand that... Figure 5 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0133] The display screen 303 can be used to display a graphical user interface (GUI) and receive operation commands generated by the user interacting with the GUI. The display screen 303 may include a display panel and a touch panel. The display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the computer device. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. The touch panel can be used to collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel), generate corresponding operation commands, and execute the corresponding program. Optionally, the touch panel may include a touch detection device and a touch controller. The touch detection device detects the user's touch location and the signal generated by the touch operation, and transmits the signal to the touch controller. The touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to the processor 301, and can receive and execute commands from the processor 301. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it transmits the information to the processor 301 to determine the type of touch event. Subsequently, the processor 301 provides corresponding visual output on the display panel according to the type of touch event. In this embodiment, the touch panel and the display panel can be integrated into the display screen 303 to achieve input and output functions. However, in some embodiments, the touch panel and the display screen 303 can be implemented as two independent components to achieve input and output functions. That is, the display screen 303 can also be used as part of the input unit 306 to achieve input functions.
[0134] The radio frequency circuit 304 can be used to transmit and receive radio frequency signals to establish wireless communication with network devices or other computer devices, and to transmit and receive signals with network devices or other computer devices.
[0135] Audio circuitry 305 can be used to provide an audio interface between a user and a computer device via a speaker and a microphone. Audio circuitry 305 converts received audio data into electrical signals, transmits them to the speaker, and the speaker converts them into sound signals for output. Conversely, the microphone converts collected sound signals into electrical signals, which are then received by audio circuitry 305, converted back into audio data, and output to processor 301 for processing. The audio data is then transmitted via radio frequency circuitry 304 to, for example, another computer device, or output to memory 302 for further processing. Audio circuitry 305 may also include an earphone jack to facilitate communication between peripheral headphones and the computer device.
[0136] The input unit 306 can be used to receive input numbers, characters, or object feature information (such as fingerprints, irises, facial information, etc.), and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control.
[0137] Power supply 307 is used to supply power to various components of computer device 300. Optionally, power supply 307 can be logically connected to processor 301 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. Power supply 307 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0138] although Figure 5 As not shown in the diagram, computer equipment 300 may also include a camera, sensor, wireless fidelity module, Bluetooth module, etc., which will not be described in detail here.
[0139] This application also provides a computer-readable storage medium for storing a computer program. This computer-readable storage medium can be applied to a computer device, and the computer program causes the computer device to execute the corresponding processes in the model generation method described in the embodiments of this application; for brevity, these will not be elaborated further here.
[0140] This application also provides a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the corresponding process in the model generation method described in the embodiments of this application. For simplicity, further details are omitted here.
[0141] This application also provides a computer program comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the corresponding process in the model generation method of this application. For brevity, further details are omitted here.
[0142] It should be understood that the processor in this application may be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above method embodiments can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor described above can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0143] It is understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0144] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0145] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0146] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0147] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0148] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0149] In addition, the functional units in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0150] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer or a server) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0151] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A model generation method, characterized in that, include: Based on the shape parameter vector of the target image and the preset shape basis vector, a human body shape vector is generated. The human body shape vector is used to characterize the three-dimensional geometric features of the target character corresponding to the target image. The image feature vector of the target image, the human body feature vector corresponding to the human body shape vector, and the initial clothing vector are decoded to generate at least one clothing vector for a part of the human body. Decode the clothing vectors for each of the aforementioned parts to generate a clothing mesh model; The human body mesh model is divided into multiple human body region models, wherein the human body mesh model is obtained by transforming the human body shape vector; Based on the target image, textures are applied to each of the human body region models and the corresponding clothing mesh models to generate the target character model.
2. The model generation method according to claim 1, characterized in that, The method further includes: Based on the preset mapping relationship between image features and human body shape, the target image is fitted to obtain the shape parameter vector.
3. The model generation method according to claim 1 or 2, characterized in that, The method further includes: Obtain the human body model matrix, which is used to represent the shape information of the three-dimensional human body model; The human body model matrix is analyzed and processed to generate M eigenvalues and M eigenvectors. The eigenvalues and eigenvectors correspond one-to-one. M is a positive integer and greater than 0. M is determined based on the number of dimensions of the vertex coordinates of the human body 3D model. The feature vector corresponding to the target feature value is determined as the shape basis vector, and the target feature value includes the first n largest feature values among the M feature values.
4. The model generation method according to any one of claims 1-3, characterized in that, The step of generating a human body shape vector based on the shape parameter vector of the target image and a preset shape basis vector includes: The shape parameter vector and the shape basis vector are linearly combined to generate a human body shape vector.
5. The model generation method according to claim 1, characterized in that, Decoding the image feature vector of the target image, the human body feature vector corresponding to the human body shape vector, and the initial clothing vector to generate at least one clothing vector for a human body part includes: The image feature vector, the human body feature vector, and the initial clothing vector are fused to generate an input feature vector, wherein the image feature vector is generated based on the target image encoding, and the human body feature vector is generated based on the human body mesh model. Based on the attention mechanism, feature capture is performed on the input feature vector to obtain at least one clothing vector of the human body part, and the clothing vector of the human body part is matched with a preset human body part.
6. The model generation method according to claim 5, characterized in that, The method further includes: The human body mesh model is converted into human body point cloud data through point cloud sampling; The human body point cloud data is processed by an encoder to generate the human body feature vector.
7. The model generation method according to claim 1, characterized in that, The image feature vector includes a clothing image feature vector. The step of applying textures to each of the human body region models and the corresponding clothing mesh models based on the target image to generate a target character model includes: Using a multi-view diffusion model, corresponding human body part textures are generated based on the semantic information of each human body region model. Each of the aforementioned human body parts is individually mapped onto the corresponding human body region model to generate a textured human body mesh model. Using a multi-view diffusion model, clothing part textures are generated based on the surface normal vectors of each clothing mesh model and the feature vectors of the clothing image. Each of the clothing part textures is applied to the clothing mesh model corresponding to the human body part to generate a textured clothing mesh model. The textured human body mesh model and the textured clothing mesh model are respectively attached to the human body parts to generate the target character model.
8. The model generation method according to claim 1, characterized in that, The method further includes: The human body shape vector is parsed to obtain a three-dimensional coordinate set, which includes each vertex corresponding to the human body shape; According to the preset human body topology, the vertices are connected to generate the human body mesh model.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted for loading by a processor to perform the model generation method as described in any one of claims 1-8.
10. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the model generation method according to any one of claims 1-8.