Multi-color three-dimensional hair style image generation method based on artificial intelligence
By constructing a hybrid drive architecture of three-dimensional point cloud data and dynamic expressions on the model head, the problem of inaccurate hairstyle design is solved, and the generation of personalized multi-color three-dimensional hairstyles is realized, which improves the accuracy of hairstyle design and user satisfaction.
Patent Information
- Application Number
- CN202510466647.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-29
AI Technical Summary
The existing hairstyle design software lacks data support for individual head shape, scalp and hair, which leads to differences in the hairstyle matching effect from the actual situation. It is difficult to accurately determine whether the head shape meets the hairstyle conditions, and the hair distribution of each body is different, making it difficult to reproduce the same effect on different heads.
By creating three-dimensional point cloud data for model heads, building a head shape library and head contour library, using edge computing devices to generate a variety of hairstyles suitable for users, combining dynamic expression hybrid driver architecture and hair-head model to achieve precise matching and rendering.
It improves the accuracy and user satisfaction of the hairstyle design, achieves sub-mm-level hair/scalp separation accuracy, meets the needs of three-dimensional hairstyle design, and enhances the sense of fit between the hairstyle and the individual head.
Smart Images

Figure CN120388134A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of hairstyle design, and particularly relates to a method for generating multi-color three-dimensional hairstyle images based on artificial intelligence. Background Art
[0002] With the continuous improvement of people's living standards, people's requirements for hairstyles are also getting higher and higher. It is no longer limited to the simple life needs in the past, but has higher requirements for it. Often, it is necessary to reflect different personalities and different aesthetic standards, so as to achieve the purpose of improving personal temperament, image and charm.
[0003] Hairstyle design is a comprehensive art, and its design is often affected by various factors, such as the head shape, face shape, facial features, body shape, age, etc. of the human body, and is also affected by personal occupation, skin color, clothing, personal hobbies, season, hair quality, applicability and the trend of the times. In addition, the hairstyle design of the human body is also affected by the state of the human head. Different head states have great differences in the presentation of the same hairstyle. Therefore, when designing a hairstyle, it is necessary to comprehensively consider various factors.
[0004] In the prior art, the presentation of a hairstyle often needs to be realized by a hairstylist or a barber. However, the presentation of a hairstyle by each hairstylist or barber is affected by their technical proficiency and working state, and there will often be certain differences in the multiple presentations of the same hairstyle. Especially when applying the same hairstyle to different head shapes, the above differences will be particularly obvious, which will inevitably reduce the accuracy of hairstyle reuse.
[0005] Although there are some hairstyle design software in the prior art that can help people realize online hairstyle design, such as uploading an image or taking a photo on the Internet, and performing a series of hairstyle designs for the design requester to print out the hairstyle or write down the serial number, and then go to the barbershop and let the hairstylist refer to the design. Most of the existing hairstyle design software applies the already designed hairstyle models, and often lacks the support of data such as the individual's head shape, scalp and hair, or the individual data is not comprehensive, fine and specific enough. Therefore, such hairstyle design can only be regarded as a kind of hairstyle wearing, and the already designed hairstyle model is like a hat being worn on the head, lacking a sense of reality, and the hairstyle often cannot fully fit the individual's head shape. The head shapes of people are very different, so the effect of the hairstyle matching often has a large difference from the actual situation. Moreover, the data such as the hair distribution, quantity, thickness, density, moisture, softness, elasticity, etc. on each individual's head are all different. It is very difficult to completely reproduce the hairstyle model designed according to a certain individual on another individual's head.
[0006] Therefore, when an individual designs and selects a hairstyle through current hairstyle design software, it is often very difficult to accurately judge whether their own head shape, hair and other data meet the conditions of the hairstyle model, determine whether the hairstyle can be trimmed on their head, and ensure that the effect of the trimmed hairstyle is consistent with the designed hairstyle. Summary of the Invention
[0007] The purpose of the present invention is to provide a method for generating a multi-color three-dimensional hairstyle image based on artificial intelligence. By making three-dimensional point cloud data of the model's head, classifying the head and head shape three-dimensional point clouds respectively to make a head shape library and a head contour library, identifying and matching the corresponding head shape library for the facial image uploaded by the user, and making multiple three-dimensional hairstyle models suitable for the user's face, the problems of inaccurate existing hairstyle design and inability to meet customer needs are solved.
[0008] To solve the above technical problems, the present invention is realized through the following technical solutions:
[0009] The present invention is a method for generating a multi-color three-dimensional hairstyle image based on artificial intelligence, including the following steps:
[0010] Step S1: Use an industrial-grade scanning device to scan the model's head at 360° from multiple angles to generate three-dimensional point cloud data of the model's head;
[0011] Step S2: Construct a dynamic expression mixing drive architecture according to the three-dimensional point cloud data of the head;
[0012] Step S3: Perform micro-detail generation and rendering on the generated three-dimensional point cloud data of the head;
[0013] Step S4: Deploy the rendered model on an edge computing device and use TensorRT to optimize the inference speed;
[0014] Step S5: The user uploads their own facial video or image to the edge computing device, and the edge server generates three-dimensional point cloud data of the user's head;
[0015] Step S6: The edge server matches the three-dimensional point cloud data of the user's head with the models in the hair-head classification model to generate multiple hairstyles suitable for the user's face.
[0016] As a preferred technical solution, in the step S1, the processing flow of the three-dimensional point cloud data of the head is as follows:
[0017] Step S11: Pre-select anatomically stable regions (such as the midline of the frontal bone, the highest point of the zygomatic bone, the pre-tragal point, etc.) on the head skin surface as reference points. Usually, at least 3 non-collinear regions need to be covered to ensure the stability of the coordinate system;
[0018] Step S12: Paste reflective markers at the head key points for cross-modal data alignment;
[0019] Step S13: Identify the coordinates of all reflective marker points and calculate the rigid transformation matrix for data from different perspectives;
[0020] Step S14: Correct the errors of the generated point cloud and remove abnormally matched points;
[0021] Step S15: Separate the point cloud of the hair from the point cloud of the head;
[0022] Step S16: After separation, create a head shape library and a head contour library respectively.
[0023] As a preferred technical solution, in the step S13, through rigid transformation pre-alignment, the search space of iterative algorithms such as ICP is reduced, and the time consumption of multi-view fusion is reduced by 40%-60%. The specific steps for calculating the rigid transformation matrix of data from different perspectives are as follows:
[0024] Step S131: Input the original scanned point cloud into the neural network and output a binary labeled point cloud; the neural network adopts a lightweight network architecture and outputs the binary classification result of each point by including coordinates, reflection intensity, and normal vector in the input channels;
[0025] When the neural network model is trained, data augmentation is required to generate synthetic training data to simulate different lighting conditions and the occlusion scenarios of hair bangs, and the specular highlight effect of the reflective markers is rendered using OpenGL to generalize the data augmentation ability;
[0026] Step S132: Encode the spatial distribution features of the labeled points and output a labeled point map with features to improve the robustness of cross-view correspondence;
[0027] Specifically, for each labeled point p i , calculate its relative vector Δp ij = p j - p i with its k = 5 nearest neighbor points to form a local hyperedge E i = {Δp i1 ,..., Δp ik}; Use GraphSAGE to aggregate neighbor information to generate the feature vector f i ∈ R 128 of each node; perform feature fusion to generate f i = MLP([x i ; max j∈N(i) (f j )]), where x i is the original coordinate; the output labeled point map with features is G = (P, F, E).
[0028] Step S133: Establish the correspondence of marked points from different perspectives based on the attention mechanism;
[0029] Specifically, construct a cross-image similarity matrix and calculate the similarity of node features between two images from perspective A and perspective B Introduce the Sinkhorn algorithm to solve the optimal transport matrix T ∈ R n×m , construct a rigid transformation hypothesis pool: randomly sample 4 groups of corresponding points, calculate the candidate transformation matrices {T k}, calculate the scores of the candidate matrices: Select the top five with the highest scores to enter the next step;
[0030] Step S134: Select the optimal solution from multiple candidate transformations to avoid local optima and obtain the optimal rigid transformation matrix. The specific calculation formula is as follows:
[0031]
[0032] In the formula, represents the Bayesian model averaging result of the optimal rigid transformation matrix, T m represents the m-th candidate rigid transformation matrix, w m represents the normalized weight of the m-th transformation hypothesis, P(T m |D) represents the posterior probability of the m-th transformation hypothesis given the observed data D, and M represents the total number of candidate transformations participating in the fusion; compared with single transformation estimation, weighted averaging can suppress the influence of outlier hypotheses and improve the registration robustness (especially when marked points are missing or there is a lot of noise); use the transformation matrix to force-align the marked points from each perspective (such as scalp fiducial points) to ensure the spatial topological consistency of anatomical key landmarks (such as hairline, auricle), providing a stable reference system for subsequent hair-scalp separation.
[0033] To improve the registration accuracy under complex postures, dynamic weight iteration refinement can be carried out. The specific implementation is as follows:
[0034] Allocate weights sensitive to errors, calculate the local curvature K of each corresponding point i , and assign higher weights to high-curvature regions. The calculation formula is as follows:
[0035]
[0036] Perform weighted SVD solution, and the centered point cloud is Calculate the covariance matrix as: Then the specific formula for SVD decomposition is:
[0037] In the formula, Denote the original three-dimensional coordinates (XYZ) of the $i$-th / $j$-th point in point clouds A and B; Denote the centroid of the point cloud after weighted centralization, eliminating translational bias; Denote the deviation vector of the $i$-th point in point cloud A corresponding to the centroid; Denote the deviation vector of the $j$-th point in point cloud B corresponding to the centroid, H represents the weighted covariance matrix; U, V represent orthogonal matrices, ∑ represents a diagonal matrix, R represents a rotation matrix, and t represents a translation vector.
[0038] As a preferred technical solution, in step S15, after correcting the generated point cloud, divide the point cloud space based on the octree, and divide the hair and scalp into independent super-voxel units through point adjacency relationships; construct a hair-scalp classification model, and the input features include coordinates, normals, and RGB values;
[0039] The specific process of dividing the point cloud space based on the octree is as follows:
[0040] Step B1: Establish a dynamic resolution octree, set the initial voxel size to 0.5mm 3 , automatically encrypt the dense hair area (such as the hair tips) to 0.2mm 3 , relax the flat scalp area to 1.0mm 3 ;
[0041] Step B2: Through the curvature-sensitive encryption strategy, trigger subdivision when the local curvature standard deviation > 0.15 to ensure the accuracy of the hair strand edges is retained;
[0042] Step B3: Adopt the point adjacency mode (non-face / line adjacency), and establish a connection relationship between each voxel and its 26-neighborhood voxels;
[0043] Step B4: For the point cloud clusters across voxels, expand the connectivity through Kd-Tree nearest neighbor search to eliminate topological breaks;
[0044] Step B5: Perform feature fusion on color similarity, normal consistency, and curvature continuity; the feature fusion formula is as follows:
[0045] S = w c ×C + w n ×N + w k ×K; (w c + w n + w k = 1);
[0046] In the formula, C, N, and K are the normalized color, normal, and curvature similarities respectively, and w c , w n , w k represent the weights of the normalized color, normal, and curvature similarities respectively;
[0047] Step B6: Place the initial seeds in the flat area of the scalp (i.e., the normal angle < 5°), and the density of the initial seeds is 1 per 2 cm 3 , and the seed points in the hair area are placed based on the density gradient: When the local point density > 200 points / cm 3 , new seeds are automatically generated;
[0048] Perform improved k-means clustering. In each iteration, calculate the mean of the feature vectors of all points within the neighborhood (radius 3 mm) of the seed points, and dynamically adjust the seed positions to the feature centroids; During the clustering optimization process, introduce the motion consistency constraint. For the dynamically scanned data, force the displacement of the seed points between adjacent frames < 0.1 mm;
[0049] Step B7: For the fragment units with an area < 50 mm 2 , perform hierarchical clustering based on feature similarity. In the hairline transition area, use bidirectional region growing, expand outward from both the hair and scalp seeds simultaneously, and take the median boundary; After the segmentation process, it is also necessary to apply bidirectional statistical filtering for noise filtering, that is, remove the sparse units (stray hairs) with a point density < 30 points / cm 3 and the abnormally dense units with a density > 500 points / cm 3 .
[0050] As a preferred technical solution, in the step S2, the dynamic expression mixing and driving architecture includes a bone driving layer and a deformation driving layer; the bone driving layer (high-density distribution in areas such as eyelids and corners of the mouth) is used to construct a simplified facial skeleton containing 62 bones; the deformation driving layer is used to predefine 52 Blend Shapes; the dynamic expression mixing and driving architecture predicts the mixing coefficient α through LSTM, and the inputs are the expression intensity and the head pose. The specific formula is as follows:
[0051] α t = Sigmoid(W h ×LSTM(e t , p t ));
[0052] In the formula, α t represents the scalar output, which is used to weight and fuse different models, e t represents the expression parameter, and p t represents the head Euler angle.
[0053] As a preferred technical solution, the dynamic expression mixing and driving architecture needs to introduce an elastic deformation energy function to make the dynamic model conform to biomechanical modeling. The specific implementation process is as follows:
[0054] Step S21: Calculate the initial displacement vector Δv of each vertex according to the bone drive or Blend Shape parameter i ;
[0055] Step S22: Discretize the continuum mechanics model into grid vertices using the finite element method, and calculate the local deformation energy through the Green strain tensor;
[0056] Step S23: Apply displacement constraint energy to each vertex; the specific formula is as follows:
[0057]
[0058] When ||Δv i || > d max At this time, the energy term shows quadratic growth, and the forced displacement is retracted within the threshold;
[0059] Step S24: Define the vertex gradient as the weighted average of the displacements of adjacent vertices, and the specific formula is as follows:
[0060]
[0061] Step S25: Convert the smoothing term ||Δv i || 2 into the form of Laplacian matrix multiplication:
[0062] E smooth = μ × V T LV;
[0063] In the formula, L is the Laplacian matrix of the grid, and V is the vertex displacement vector;
[0064] Step S26: Perform nonlinear optimization, and perform iterative projection on the vertices with over-limited displacements ||Δv i || > d max Execute:
[0065]
[0066] In the formula, β is the step size, which is determined by line search to ensure convergence.
[0067] As a preferred technical solution, in step S4, calculate the Gaussian curvature of the input three-dimensional grid to generate a curvature distribution map; the calculation formula is as follows:
[0068]
[0069] In the formula, A is the area of the vertex neighborhood, and θ i is the included angle between adjacent triangular patches; the generated curvature distribution map is 32 * 32, and the quantization is divided into [-1, 1];
[0070] The input to the generator is a low-resolution mesh, and the output is a high-resolution vertex displacement map with details; the format of the generator is as follows:
[0071] Input: Low-resolution 3D mesh vertex coordinates (N×3) + curvature feature vectors (N×128).
[0072] Output: High-resolution vertex displacement map (Δx, Δy, Δz), with the resolution increased to 4 times that of the original mesh;
[0073] Construct a multi-scale discriminator based on the mesh curvature map; the multi-scale discriminator adopts a three-branch PatchGAN architecture, processes different-scale features (16×16→8×8→4×4) respectively, aggregates multi-scale features through an adaptive pooling layer, and outputs true / false probabilities and curvature consistency scores.
[0074] As a preferred technical solution, in step S4, when the model is deployed, first perform pruning and quantization processing on the rendered model, convert the model to the ONNX format; use TensorRT to build an optimization engine, through layer fusion, dynamic Shape adaptation, and INT8 calibration; in the deployment stage, use containerization technology to encapsulate the inference service, combine multi-stream parallelism and a persistent engine to reduce latency; finally, achieve a balance between throughput and accuracy through resource isolation, dynamic batching, and real-time monitoring.
[0075] As a preferred technical solution, in step S6, the edge server uses the methods of steps S1 to S3 to generate 3D point cloud data of the user's head for the pictures or videos uploaded by the user, align the point cloud coordinate system, perform rigid registration based on the tip of the nose and the tragus points, extract the user's facial biometric points, construct a facial topology structure, calculate the curvature distribution map and the normal vector field, analyze the geometric attributes of the cheekbone width / jaw contour, and use a 3D CNN enhanced by an attention mechanism to perform non-uniform sampling on the point cloud and then achieve end-to-end classification; and recommend multiple hairstyles suitable for the user's face according to the classification results.
[0076] As a preferred technical solution, the edge server fits the hairstyle mesh to the user's point cloud according to the recommended hairstyle results, maintains the natural drooping feeling of the hair through Laplacian deformation, and performs rendering using the method of step S3; after rendering, the edge server supports 360° rotation and lighting environment simulation. If the user is not satisfied with the recommended result of the hairstyle, the user can select a suitable hairstyle from the hairstyle library and manually perform color rendering. At the same time, the edge server provides a virtual try-on feedback channel to collect user adjustment data (such as hairstyle length, curl preference) for model iteration.
[0077] The present invention has the following beneficial effects:
[0078] (1) The present invention creates three-dimensional point cloud data of a model's head, classifies the head and head shape three-dimensional point clouds respectively to create a head shape library and a head contour library, identifies and matches the corresponding head shape library for the facial images uploaded by users, creates multiple three-dimensional hairstyle models suitable for the user's face, builds a dynamic expression hybrid driving architecture, enables users to view whether the hairstyles are suitable under different facial expressions, completes the comprehensive display of the three-dimensional models, and improves the user satisfaction with hairstyle design.
[0079] (2) The present invention dynamically divides through an octree and fuses multiple features, uses the point adjacency relationship to divide hair and scalp into independent superbody units, thereby constructs a hair-head classification model, achieves sub-millimeter-level hair / scalp separation accuracy, and meets the three-dimensional hairstyle design modeling requirements;
[0080] (3) The present invention calculates the rigid transformation between the point clouds collected at different positions of the scanner (such as 12 circular viewing angles), unifies the local coordinate system to the global coordinate system, solves the coordinate offset problem caused by device movement, and at the same time, for the hair fluttering scenario, updates through a real-time transformation matrix (at the 30fps level), separates the rigid movement of the head and the non-rigid deformation of the hair strands, and avoids the point cloud breakage caused by motion blur.
[0081] (4) The present invention performs micro-detail generation and rendering on the generated three-dimensional point cloud data of the head, calculates the Gaussian curvature of the input three-dimensional mesh, generates a curvature distribution map, in high-curvature regions such as the nose wing and the corner of the eye, generates a pore density higher than that of traditional GANs, improves the detail fidelity, and after introducing curvature constraints, reduces the incidence of mode collapse.
[0082] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0084] Figure 1 It is a flowchart of a method for generating a multi-color three-dimensional hairstyle image based on artificial intelligence of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0085] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0086] In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0087] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further describes the present application in detail with reference to the Figure 1 accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0088] Please refer to Figure 1 as shown. The present invention is a method for generating a multi-color three-dimensional hairstyle image based on artificial intelligence, including the following steps:
[0089] Step S1: Use an industrial-grade scanning device to perform a 360° multi-angle scan on the model's head to generate three-dimensional point cloud data of the model's head;
[0090] Step S2: Construct a dynamic expression blending drive architecture based on the three-dimensional point cloud data of the head;
[0091] Step S3: Perform micro-detail generation and rendering on the generated three-dimensional point cloud data of the head;
[0092] Step S4: Deploy the rendered model on an edge computing device and use TensorRT to optimize the inference speed;
[0093] Step S5: The user uploads their own facial video or image to the edge computing device, and the edge server generates three-dimensional point cloud data of the user's head;
[0094] Step S6: The edge server matches the three-dimensional point cloud data of the user's head with the model in the hair-head classification model to generate multiple hairstyles suitable for the user's face.
[0095] In step S1, the processing flow of the three-dimensional point cloud data of the head is as follows:
[0096] Step S11: Pre-select anatomically stable regions (such as the midline of the frontal bone, the highest point of the zygomatic bone, the pre-tragal point, etc.) on the head skin surface as reference points. Usually, at least 3 non-collinear regions need to be covered to ensure the stability of the coordinate system;
[0097] Step S12: Paste reflective markers at the head key points for cross-modal data alignment;
[0098] Step S13: Identify the coordinates of all reflective marker points and calculate the rigid transformation matrix for data from different perspectives;
[0099] Step S14: Correct the error of the generated point cloud and eliminate abnormally matched points;
[0100] Step S15: Separate the point cloud of the hair from the point cloud of the head;
[0101] Step S16: After separation, create a head shape library and a head contour library respectively.
[0102] As a preferred technical solution, in the step S13, the specific steps for calculating the rigid transformation matrix for data from different perspectives are as follows:
[0103] Step S131: Input the original scanned point cloud into the neural network and output a binary labeled point cloud; the neural network adopts a lightweight network architecture and outputs the binary classification result of each point by including coordinates, reflection intensity, and normal vector in the input channel;
[0104] When the neural network model is trained, it is necessary to perform data augmentation to generate synthetic training data to simulate different lighting conditions and the occlusion scenario of hair bangs, use OpenGL to render the specular highlight effect of the reflective markers, and generalize the data augmentation ability;
[0105] Step S132: Encode the spatial distribution characteristics of the marker points and output a marker point map with features to improve the robustness of cross-view correspondence;
[0106] Specifically, for each marker point p i , calculate its relative vector Δp ij = p j - p i with its k = 5 nearest neighbor points to form a local hyperedge E i = {Δp i1 ,..., Δp ik}; use GraphSAGE to aggregate neighbor information to generate the feature vector f i ∈ R 128 ; perform feature fusion to generate f i = MLP([x i ; max j∈N(i) (f j )]), where x i is the original coordinate; the output marker point map with features G = (P, F, E).
[0107] Step S133: Establish the correspondence of the marked points from different perspectives based on the attention mechanism;
[0108] When specifically implemented, construct a cross-image similarity matrix and calculate the similarity of the node features of the two images from perspective A and perspective B Introduce the Sinkhorn algorithm to solve the optimal transport matrix T ∈ R n×m , construct a rigid transformation hypothesis pool: randomly sample 4 groups of corresponding points and calculate the candidate transformation matrices {T k}, calculate the scores of the candidate matrices: Select the top five with the highest scores to enter the next step;
[0109] Step S134: Select the optimal solution from multiple candidate transformations to avoid local optima and obtain the optimal rigid transformation matrix. The specific calculation formula is as follows:
[0110]
[0111] In the formula, represents the Bayesian model average result of the optimal rigid transformation matrix, T m represents the m-th candidate rigid transformation matrix, w m represents the normalized weight of the m-th transformation hypothesis, P(T m |D) represents the posterior probability of the m-th transformation hypothesis given the observed data D, and M represents the total number of candidate transformations participating in the fusion; compared with single transformation estimation, weighted averaging can suppress the influence of outlier hypotheses and improve the registration robustness (especially when the marked points are missing or there is a lot of noise); use the transformation matrix to force-align the marked points of each perspective (such as scalp fiducial points) to ensure the spatial topological consistency of the anatomical key landmarks (such as hairline, auricle), and provide a stable reference system for subsequent hair-scalp separation.
[0112] To improve the registration accuracy under complex postures, dynamic weight iteration refinement can be performed, and the specific implementation is as follows:
[0113] Allocate the weights sensitive to errors, calculate the local curvature K of each corresponding point i , and assign higher weights to the high-curvature regions. The calculation formula is as follows:
[0114]
[0115] Perform weighted SVD solution. The centered point cloud is Calculate the covariance matrix as: Then the specific formula for SVD decomposition is:
[0116] In the formula, represents the original three-dimensional coordinates (XYZ) of the i-th / j-th point in point clouds A and B; represents the centroid of the point cloud after weighted centralization, eliminating translational bias; represents the deviation vector of the i-th point in point cloud A corresponding to the centroid; represents the deviation vector of the j-th point in point cloud B corresponding to the centroid, H represents the weighted covariance matrix; U, V represent orthogonal matrices, ∑ represents a diagonal matrix, R represents a rotation matrix, and t represents a translation vector.
[0117] As a preferred technical solution, in step S15, after the generated point cloud is corrected, the point cloud space is divided based on an octree, and the hair and scalp are divided into independent super-voxel units through point adjacency relationships; a hair-scalp classification model is constructed, and the input features include coordinates, normals, and RGB values;
[0118] The specific process of dividing the point cloud space based on the octree is as follows:
[0119] Step B1: Establish a dynamic resolution octree with an initial voxel size set to 0.5 mm 3 , and automatically encrypt it to 0.2 mm in the hair-dense area (such as the hair tips) 3 , and relax it to 1.0 mm in the flat area of the scalp 3 ;
[0120] Step B2: Through a curvature-sensitive encryption strategy, trigger subdivision when the local curvature standard deviation > 0.15 to ensure the accuracy of the hair edge retention;
[0121] Step B3: Adopt a point adjacency mode (non-face / line adjacency), and establish a connection relationship between each voxel and its 26 neighboring voxels;
[0122] Step B4: For the point cloud clusters across voxels, expand the connectivity through Kd-Tree nearest neighbor search to eliminate topological breaks;
[0123] Step B5: Perform feature fusion on color similarity, normal consistency, and curvature continuity; the feature fusion formula is as follows:
[0124] S = w c ×C + w n ×N + w k ×K; (w c + w n + w k = 1);
[0125] In the formula, C, N, and K are the normalized color, normal, and curvature similarities respectively, and w c , w n , w k represent the weights of the normalized color, normal, and curvature similarities respectively;
[0126] Step B6: Place initial seeds on a flat area of the scalp (i.e., normal angle < 5°) at a density of 1 seed per 2 cm 3 , the seed points in the hair area are placed based on the density gradient: when the local point density is > 200 points / cm 3 Automatically generate new seeds when
[0127] Perform improved k-means clustering. At each iteration, the mean of the eigenvectors of all points in the neighborhood of the seed point (radius 3mm) is calculated, and the seed position is dynamically adjusted to the feature centroid. A motion consistency constraint is introduced during the clustering optimization process. For dynamic scanning data, the displacement of seed points between adjacent frames is forced to be less than 0.1mm.
[0128] Step B7: For area <50mm 2 The fragment units are clustered hierarchically based on feature similarity. In the hairline transition area, bidirectional region growing is used, expanding outward from the hair and scalp seeds at the same time, taking the median boundary. After the segmentation process, bidirectional statistical filtering is also required for noise filtering, that is, eliminating points with a density of less than 30 points / cm 3 Sparse units (broken hair) and >500 points / cm 3 Abnormally dense units.
[0129] In step S2, the dynamic expression hybrid drive architecture includes a skeleton drive layer and a deformation drive layer. The skeleton drive layer (with high density distribution in areas such as eyelids and mouth corners) is used to construct a simplified facial skeleton containing 62 bones. The simplified facial skeleton with 62 bones is divided into three levels of control (main bones / secondary bones / fine-tuning bones), and the key areas are distributed as follows:
[0130] Eyelid area: 8 bones (control eyelid opening and closing, outer canthal folds);
[0131] Corner of mouth area: 6 bones (drive the contraction of orbicularis oris muscle and the elevation of zygomatic muscle);
[0132] Nose area: 4 bones (simulating nose wing expansion and nose bridge wrinkles); these bones are bound through a parent-child relationship (for example, the zygomatic bone drives the mouth corner bones), achieving anatomically reasonable linkage. For example, when driving a "smile" expression, the zygomatic bone rotates 15°, driving the secondary bones to move 3mm, and the mouth corner bones to lift.
[0133] The deformation drive layer is used to predefine 52 Blend Shapes. The dynamic expression blend drive architecture predicts the blend coefficient α through LSTM, with the input being the expression intensity and head posture. The specific formula is as follows:
[0134] α t =Sigmoid(W h ×LSTM(e t ,pt ));
[0135] Wherein, α t represents a scalar output for weighted fusion of different models, e t represents expression parameters (including 50 labels), p t represents the head Euler angles, and the head rotation state (3-axis Euler angles) represented by Euler angles (roll, pitch, yaw) or quaternions.
[0136] Use LSTM to implicitly model the non-linear coupling relationship between expression parameters (anatomical features) and head pose (kinematic features), for example:
[0137] LSTM receives the input: e t = [AU12: 0.7, AU15: 0.3], p t = [depression angle 10°, deflection -5°]; the network output: the bone drive weight α rig = 0.6, the deformation drive weight α blend = 0.4, and the final vertex position is: V final = 0.6 × T rig (V) + 0.4 × (V + ΔV blend ), T rig represents the bone transformation matrix.
[0138] The dynamic expression blending drive architecture needs to introduce an elastic deformation energy function to make the dynamic model conform to biomechanical modeling. The specific implementation process is as follows:
[0139] Step S21: Calculate the initial displacement vector Δv of each vertex according to the bone drive or Blend Shape parameters i ;
[0140] Step S22: Use the finite element method to discretize the continuous medium mechanics model into grid vertices, and calculate the local deformation energy through the Green strain tensor;
[0141] Step S23: Apply displacement constraint energy to each vertex; the specific formula is as follows:
[0142]
[0143] When ||Δv i || > d max the energy term grows quadratically, forcing the displacement to retract within the threshold;
[0144] Step S24: Define the vertex gradient as the weighted average of the displacements of adjacent vertices. The specific formula is as follows:
[0145]
[0146] Step S25: Convert the smoothing term ||Δv i ‖ 2 into the form of Laplacian matrix multiplication:
[0147] E smooth = μ × V T LV;
[0148] where L is the Laplacian matrix of the grid and V is the vertex displacement vector;
[0149] Step S26: Perform non - linear optimization. For the vertices with over - limited displacement ||Δv i || > d max perform iterative projection:
[0150]
[0151] where β is the step size, determined by line search to ensure convergence.
[0152] In step S4, calculate the Gaussian curvature of the input 3D mesh to generate a curvature distribution map; the calculation formula is as follows:
[0153]
[0154] where A is the area of the vertex neighborhood and θ i is the included angle between adjacent triangular patches; the generated curvature distribution map is 32 * 32, and the quantization is divided into [-1, 1];
[0155] Input the low - resolution mesh into the generator, and output a high - resolution vertex displacement map with details; the format of the generator is:
[0156] Input: Low - resolution 3D mesh vertex coordinates (N × 3)+curvature feature vector (N × 128).
[0157] Output: High - resolution vertex displacement map (Δx, Δy, Δz), and the resolution is increased to 4 times that of the original mesh;
[0158] Construct a multi - scale discriminator based on the mesh curvature map; the multi - scale discriminator adopts a three - branch PatchGAN architecture, processes different scale features (16 × 16 → 8 × 8 → 4 × 4) respectively, aggregates multi - scale features through an adaptive pooling layer, and outputs the true - false probability and curvature consistency score.
[0159] In step S4, when the model is deployed, the rendered model is first pruned and quantized, and the model is converted into the ONNX format. An optimization engine is built using TensorRT, through layer fusion, dynamic Shape adaptation, and INT8 calibration. In the deployment stage, containerization technology is used to encapsulate the inference service, and multi-stream parallelism and a persistent engine are combined to reduce latency. Finally, throughput and accuracy are balanced through resource isolation, dynamic batching, and real-time monitoring.
[0160] Based on the recommended hairstyle results, the edge server fits the hairstyle mesh to the user's point cloud, maintains the natural draping of the hair strands through Laplacian deformation, and performs rendering using the method of step S3. After rendering, the edge server supports 360° rotation and lighting environment simulation. If the user is not satisfied with the recommended hairstyle results, they can select a suitable hairstyle from the hairstyle library and manually perform color rendering. At the same time, the edge server provides a virtual try-on feedback channel to collect user adjustment data (such as hairstyle length and curl preference) for model iteration.
[0161] It should be noted that in the above system embodiments, the various units included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.
[0162] In addition, those of ordinary skill in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the corresponding program can be stored in a computer-readable storage medium.
[0163] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art in the relevant technical field can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.
Claims
1. A method for generating a multi-color three-dimensional hairstyle image based on artificial intelligence, characterized in that, The steps include: Step S1: Use industrial-grade scanning equipment to perform 360° multi-angle scanning on the model's head to generate three-dimensional point cloud data of the model's head; Step S2: constructing a dynamic expression hybrid drive architecture based on the three-dimensional point cloud data of the head; Step S3: Generate and render micro details of the generated three-dimensional point cloud data of the head; Step S4: Deploy the rendered model on the edge computing device and use TensorRT to optimize the inference speed; Step S5: The user uploads his or her facial video or image to the edge computing device, and the edge server generates three-dimensional point cloud data of the user's head; Step S6: The edge server matches the three-dimensional point cloud data of the user's head with the model in the hair-head classification model to generate multiple hairstyles suitable for the user's face.
2. The generating method of a multi-color three-dimensional hairstyle image based on artificial intelligence according to claim 1, wherein In step S1, the three-dimensional point cloud data processing flow of the head is as follows: Step S11: Preselect anatomically stable areas on the head skin surface as reference points, usually covering at least three non-collinear areas to ensure the stability of the coordinate system; Step S12: attaching reflective markers to key points on the head to perform cross-modal data alignment; Step S13: Identify the coordinates of all reflective marker points and calculate the rigid transformation matrix of data at different viewing angles; Step S14: performing error correction on the generated point cloud and removing abnormal matching points; Step S15: Separate the point cloud of the hair and the point cloud of the head; Step S16: After separation, a head shape library and a head contour library are created respectively.
3. The method for generating a multi-color three-dimensional hairstyle image based on artificial intelligence according to claim 2, characterized in that, In step S13, the specific steps of calculating the rigid transformation matrix of different viewing angle data are as follows: Step S131: inputting the original scanned point cloud into the neural network and outputting the binary labeled point cloud; Step S132: Encode the spatial distribution features of the marker points and output a marker point map with features; Step S133: establishing correspondences between markers at different viewing angles based on the attention mechanism; Step S134: Select the optimal solution from multiple candidate transformations to obtain the optimal rigid transformation matrix.
4. A method for generating a multi-color three-dimensional hairstyle image based on artificial intelligence according to claim 2, characterized in that, In step S15, after the generated point cloud is corrected, the point cloud space is divided based on the octree, and the hair and scalp are divided into independent super units based on the point adjacency relationship; a hair-head classification model is constructed, and the input features include coordinates, normals and RGB values; The specific process of dividing the point cloud space based on the octree is as follows: Step B1: Establish a dynamic resolution octree with an initial voxel size of 0.5 mm 3 , and automatically encrypt it to 0.2 mm in the area with dense hair 3 , and relax it to 1.0 mm in the flat area of the scalp 3 ; Step B2: Using a curvature-sensitive encryption strategy, subdivision is triggered when the local curvature standard deviation is greater than 0.15 to ensure that the hairline edge is retained accurately. Step B3: Establish a connection relationship between each voxel and its 26 neighboring voxels; Step B4: For point cloud clusters across voxels, expand connectivity through Kd-Tree nearest neighbor search; Step B5: feature fusion of color similarity, normal consistency, and curvature continuity; Step B6: Place an initial seed on the flat area of the scalp, calculate the mean of the feature vectors of all points in the neighborhood of the seed point, and dynamically adjust the seed position to the feature centroid; Step B7: For the fragment units with an area <50mm 2 , perform hierarchical clustering based on feature similarity. In the hairline transition area, use bidirectional region growing, expanding outward from both hair and scalp seeds simultaneously, and take the median boundary.
5. A method for generating a multi-color three-dimensional hairstyle image based on artificial intelligence according to claim 1, characterized in that, In the step S2, the dynamic expression blending and driving architecture includes a bone driving layer and a deformation driving layer; the bone driving layer is used to construct a simplified facial skeleton including 62 bones; the deformation driving layer is used to predefine 52 BlendShapes; the dynamic expression blending and driving architecture predicts the blending coefficients through LSTM, and the inputs are expression intensity and head pose.
6. The generation method of a multi-color three-dimensional hairstyle image based on artificial intelligence according to claim 1, wherein The dynamic expression blending and driving architecture needs to introduce an elastic deformation energy function to make the dynamic model conform to biomechanical modeling. The specific implementation process is as follows: Step S21: Calculate the initial displacement vector of each vertex according to the bone driving or Blend Shape parameters. Step S22: Use the finite element method to discretize the continuum mechanics model into grid vertices, and calculate the local deformation energy through the Green strain tensor. Step S23: Apply the displacement constraint energy to each vertex. Step S24: Define the vertex gradient as the weighted average of the displacements of adjacent vertices. Step S25: Convert the smoothing term into the form of Laplacian matrix multiplication. Step S26: Perform non-linear optimization.
7. A method for generating a multi-color three-dimensional hairstyle image based on artificial intelligence according to claim 1, characterized in that, In the step S3, calculate the Gaussian curvature of the input three-dimensional mesh to generate a curvature distribution map. The input to the generator is a low-resolution mesh, and the output is a high-resolution vertex displacement map with details. Construct a multi-scale discriminator based on the mesh curvature map; the multi-scale discriminator adopts a three-branch PatchGAN architecture, processes different scale features respectively, aggregates multi-scale features through an adaptive pooling layer, and outputs the true / false probability and the curvature consistency score.
8. A method for generating a multi - color three - dimensional hairstyle image based on artificial intelligence according to claim 1, characterized in that, In the step S4, when the model is deployed, first perform pruning and quantization processing on the rendered model, and convert the model into the ONNX format; use TensorRT to construct an optimization engine through layer fusion, dynamic Shape adaptation, and INT8 calibration. In the deployment stage, use containerization technology to encapsulate the inference service, and combine multi-stream parallelism and a persistent engine to reduce latency; finally, achieve the balance of throughput and accuracy through resource isolation, dynamic batching, and real-time monitoring.
9. A method for generating a multi-color three-dimensional hairstyle image based on artificial intelligence according to claim 1, characterized in that, In the step S6, the edge server uses the methods of steps S1 to S3 to generate three-dimensional point cloud data of the user's head from the pictures or videos uploaded by the user, align the point cloud coordinate system, achieve rigid registration based on the tip of the nose and the tragus points, extract the user's facial biometric points, construct a facial topology structure, calculate the curvature distribution map and the normal vector field, analyze the geometric attributes of the zygomatic width / mandibular contour, and use a 3D CNN enhanced by an attention mechanism to perform non-uniform sampling on the point cloud and then achieve end-to-end classification; and recommend multiple hairstyles suitable for the user's face according to the classification results.
10. A method for generating a multi-color three-dimensional hairstyle image based on artificial intelligence according to claim 9, characterized in that, The edge server fits the hairstyle mesh to the user's point cloud according to the recommended hairstyle result, maintains the natural drooping feeling of the hair through Laplacian deformation, and uses the method of step S3 for rendering; after the rendering is completed, the edge server supports 360° rotation and light environment simulation.