Multi-modal virtual image generation monitoring and evaluation verification method
By constructing a dataset of original geometric features and comparing it with the target style template, and combining aesthetic knowledge graphs and historical cases, the problem of balancing stylization and realism in virtual character generation was solved, achieving high-quality virtual character generation and improved user satisfaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-09
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies struggle to strike a balance between artistic stylization and realism when generating virtual avatars, resulting in a lack of credibility and artistic appeal, and a poor user experience.
By constructing a geometric dataset of the original appearance and comparing it with the target style template, analyzing the deformation depth value, triggering an aesthetic knowledge graph warning, combining aesthetic standards and historical cases to conduct a multi-dimensional adaptability assessment, outputting an image blueprint document, and realizing the deployment of virtual images through stylized rendering and 3D reconstruction technology.
It has improved the quality of virtual avatar generation and user satisfaction in different scenarios, ensuring a balance between stylistic intensity and recognizability, and a harmonious unity between personalization and aesthetics.
Smart Images

Figure CN122049306A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information technology, and in particular to a method for monitoring, evaluating and verifying the generation of multimodal virtual avatars. Background Technology
[0002] In the fields of digital content creation and virtual avatar design, researching how to generate virtual avatars that possess both artistic style and realistic features is crucial. This field is not only about technological innovation but also a core driving force for the development of the cultural and creative industries. Especially in scenarios such as games, film and television, and virtual reality, the quality of virtual avatars directly affects user experience and emotional resonance, making its importance self-evident. However, many current methods, while pursuing artistic stylization, often neglect the balance with real-life appearance, resulting in visually unconvincing generated results. Many solutions pursue exaggerated stylistic effects, such as oil painting-like brushstrokes or anime-like simplified lines, while ignoring the complete preservation of key facial features. As a result, while the generated virtual avatar may be highly artistic, its similarity to the user's original appearance drops sharply. This limitation is not simply due to insufficient technology but rather a lack of in-depth exploration of the dynamic relationship between stylization and realism. Increasing the degree of stylization usually means greater adjustments to facial geometry, such as exaggerated deformation of contour lines or deliberate alteration of facial proportions. Such adjustments often weaken the recognizability of the original appearance. Conversely, overemphasizing the preservation of original features can lead to a lack of stylistic prominence and an artistically unappealing image. The core factor in this contradiction lies in the depth of facial geometric deformation, as controlling this depth directly determines the balance between stylization and realism. Excessive deformation can overly distort facial features, causing them to lose their original recognizability; insufficient deformation, on the other hand, will fail to achieve the desired stylization effect. For example, in customizing virtual game avatars, users might upload selfies hoping for a cute, cartoonish avatar. If the deformation depth is out of control, the original slender, almond-shaped eyes might be stretched into large, round cartoon eyes, and the chin might change from a soft curve to a sharp V-shape, making it impossible for friends to recognize the virtual avatar in the game lobby, leading to user dissatisfaction and abandonment. Similarly, in film and television virtual stunt doubles, directors might require actors to have their faces stylized in a watercolor style. Excessive deformation can soften and blur the original high cheekbones and hooked nose, making it difficult for viewers to associate them with the original celebrity and compromising the realism of the storyline. Therefore, how to reasonably control the deformation depth of facial geometry during the stylization process, so as to find the best balance between artistic effect and realistic features, has become a key problem that urgently needs to be solved in the field of virtual image generation. Summary of the Invention
[0003] This invention provides a method for monitoring and evaluating the generation of multimodal virtual avatars, mainly including: A geometric dataset based on the user's real image data is generated. The original geometric dataset is compared with the target style template to obtain the deformation depth value of each part. The risk of facial proportion imbalance is analyzed based on the deformation depth value, and the level of reduced recognizability is output. When the level of recognizability decreases beyond the threshold allowed by the stylization intensity, potential aesthetic risk points are identified through aesthetic knowledge graphs. Based on the location of aesthetic risk points, corresponding historical cases are identified, and relevant reference cases are selected based on the similarity between current user characteristics and historical successful cases. By combining aesthetic standards and historically relevant reference cases to drive the aesthetic knowledge graph to conduct multi-dimensional adaptability assessment, an image blueprint document is output. After the original geometry dataset is passed to the stylization rendering stage for stylization rendering, it is deployed to the application scenario, and the rendering stability and attribute fusion effect of each deformation depth range are recorded as online feedback data. Based on offline and online feedback data, the optimal range of stylization intensity within the range of maintaining recognizability is evaluated, and the virtual avatar generation quality evaluation result is output.
[0004] Furthermore, the original geometric dataset is compared with the target style template to obtain the deformation depth value of each part. Based on the deformation depth value, the risk of facial proportion imbalance is analyzed, and the level of reduced recognizability is output, including: Read the standard facial feature parameters in the target style template, calculate the difference between the feature values in the original geometric dataset and the standard values to obtain the deformation amplitude coefficient; calculate the deformation depth values of the eye region, nose region and cheekbone region based on the deformation amplitude coefficient; for the deformation depth values of each region, calculate the expected change degree of the key facial features after deformation, and determine the level of reduction in recognizability.
[0005] Furthermore, the identification of potential aesthetic risk points includes: When the level of recognizability reduction exceeds a preset threshold, an early warning mechanism is triggered. An aesthetic standard node library is extracted from the aesthetic knowledge graph. This library includes nodes related to the golden ratio, facial harmony, and differences between Eastern and Western aesthetics. A query is used to retrieve the set of standard nodes that best match the current facial features. Based on the aesthetic parameter benchmark values in the set of standard nodes, the deviation of the current deformation depth value in the dimensions of the eyes, nose, and cheekbones is calculated to form an aesthetic deviation vector. A threshold judgment method is used to assess the risk level of the aesthetic deviation vector, marking high-risk areas and extracting combinations of deviation features as potential aesthetic risk points. The knowledge graph is traversed to find correction strategy nodes that match the risk point type, and the deformation depth correction coefficient in the suggested adjustment parameters is read to identify the specific risk point location and risk level in the current deformation scheme.
[0006] Furthermore, the process of identifying corresponding historical cases based on aesthetic risk point location, and selecting relevant reference cases based on the similarity between current user characteristics and historical successful cases, includes: Based on the coordinates of the aesthetic risk points, records with the same risk point type are retrieved from the historical case database. Body shape feature vectors, stylization processing parameter sequences, user rating values, and usage scenario tags are obtained to construct a historical case feature matrix. Cosine similarity is used to calculate the similarity between the current user's body shape feature vector and the historical case feature matrix, marking highly similar cases. Corresponding data pairs of deformation depth parameters and satisfaction scores are extracted from highly similar cases. Historical cases are filtered based on the corresponding data pairs, and the aesthetic preference distribution data of the filtered cases under the user's selected scenario are statistically analyzed. Relevant reference cases that meet the difference threshold and satisfaction standard are output.
[0007] Furthermore, the output blueprint document includes: By combining the location of aesthetic risk points and relevant reference cases, the rule matching process in the aesthetic knowledge graph is initiated. Aesthetic rule nodes corresponding to the current user's body shape and contour dimensions are extracted, and the deviation between the current deformation depth value and the standard value is calculated to obtain a body shape fit score. User skin color gamut data is extracted, and style preference records for the corresponding skin color category are queried in the aesthetic knowledge graph. The matching degree between the visual effect parameters generated by the current deformation depth value and skin color features is calculated to obtain a style fit index. User hair texture density parameters are read, and the standard requirements for hairstyle processing under different usage occasions are matched. The conformity of hair texture parameters with scene standards is calculated to obtain an occasion fit value. Combining the body shape fit score, style fit index, and occasion fit value, a multi-dimensional fit assessment result is generated through weighted calculation, constructing an image blueprint document.
[0008] Furthermore, after the original geometry dataset is transferred to the stylized rendering stage for stylized rendering, it is deployed to the application scenario, and the rendering stability and attribute fusion effect of each deformation depth range are recorded as online feedback data, including: The original geometric dataset is input into the VToonify renderer, and a weighted fusion of feature vectors and style templates is performed to obtain a two-dimensional stylized facial image. The two-dimensional stylized facial image is used as a texture map, and three-dimensional point cloud data is constructed by combining deformation depth values. A three-dimensional stereoscopic image is reconstructed using a 3D Gaussian splash distribution Gaussian kernel. Adaptation processing is performed based on the three-dimensional stereoscopic image, and the adapted virtual image data is output. The rendering performance of the virtual image data on the target platform is monitored, and the performance indicators and fusion effect data are stored as online feedback data.
[0009] Furthermore, the output virtual avatar generation quality assessment results include: User satisfaction ratings and return rate statistics are extracted from offline delivery records. Online feedback data is read to construct a comprehensive dataset containing stylistic intensity, recognizability, and satisfaction. The comprehensive dataset is divided into stylistic intensity intervals, and the mean and variance of recognizability values within each interval are calculated to mark candidate optimal intervals. Based on the mean recognizability and mean satisfaction values of the candidate optimal intervals, a quality score is obtained by weighted summation. The interval with the highest quality score is selected as the optimal stylistic intensity interval, and the virtual avatar generation quality assessment result is output.
[0010] Furthermore, the method also includes: feeding back the quality assessment results of virtual image generation to the aesthetic knowledge graph, updating the body shape, style, and occasion matching records in the deformation depth adaptability case library, and realizing continuous optimization of multimodal virtual image generation assessment and verification.
[0011] Furthermore, the quality assessment results generated through virtual avatars are fed back to the aesthetic knowledge graph to update the body shape, style, and occasion matching records in the deformation depth adaptability case library, including: The virtual avatar generation quality assessment results are transmitted to the update interface of the aesthetic knowledge graph. New records containing user body shape characteristics, style type, deformation depth value, and satisfaction score are added to the aesthetic rule nodes. Based on the new records, the deformation depth adaptability case library is accessed to calculate the deviation and update the case set and successful matching statistics of the corresponding body shape category. The distribution of stylization intensity values under each body shape category is statistically analyzed through the updated case library to determine the recommendation intensity and retention threshold range as stable parameters for output.
[0012] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: This invention discloses a multimodal virtual avatar generation monitoring and evaluation verification method. Addressing the challenges of balancing stylistic intensity and recognizability, and balancing aesthetic adaptability and personalization needs in virtual avatar generation, the method integrates real user image data with target style templates to construct a primordial geometric dataset. It then analyzes the risk of facial proportion imbalance using deformation depth values, triggering an aesthetic knowledge graph warning. Furthermore, it performs multi-dimensional adaptability assessments by matching aesthetic standards with historical cases, outputting an avatar blueprint document. Simultaneously, it utilizes stylized rendering and 3D reconstruction technologies to deploy the virtual avatar. This invention leverages online and offline feedback data to construct a closed-loop evaluation mechanism, continuously optimizing stylistic intensity and retention thresholds to ensure the stability and suitability of the virtual avatar in different scenarios. Ultimately, it achieves a harmonious unity of personalization and aesthetics, significantly improving the quality of virtual avatar generation and user satisfaction. Attached Figure Description
[0013] Figure 1 This is a flowchart of a multimodal virtual image generation monitoring and evaluation verification method according to the present invention.
[0014] Figure 2 This is a schematic diagram of a multimodal virtual image generation monitoring and evaluation verification method according to the present invention.
[0015] Figure 3 This is another schematic diagram of a multimodal virtual image generation monitoring and evaluation verification method according to the present invention. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0017] In this embodiment, a 1:1 digital twin model of the user is first constructed using a multimodal high-precision data acquisition system, providing a geometric and appearance benchmark for subsequent stylization generation. Specifically, a structured light or laser scanning device is used to perform high-precision 3D scanning of the user's face and whole body, acquiring raw point cloud data with millimeter-level precision. Simultaneously, a high-definition camera array is deployed to simultaneously capture images of the user from multiple perspectives, including front, side, and 45-degree angles, to extract multi-angle facial / whole-body appearance information and skin color gamut data. Subsequently, the ICP iterative nearest point algorithm is used to perform point cloud registration, statistical outlier removal filtering, and Poisson surface reconstruction algorithm mesh reconstruction to generate a continuous and topologically complete 3D mesh model. Based on this, multi-view stereo vision technology is combined for texture mapping, accurately projecting the color and detail information acquired by the high-definition camera array onto the 3D mesh surface, constructing a 1:1 digital twin model with accurate geometry, realistic texture, and consistent proportions. This digital twin model serves as a high-fidelity source of the original geometric dataset, and its key features can be further extracted using lightweight algorithms to adapt to the virtual image generation needs under different computing resources.
[0018] like Figures 1-3 This embodiment of a multimodal virtual avatar generation monitoring and evaluation verification method may specifically include: S101. Identify the coordinate points of facial contours and the ratio of distances between facial features in the user's real image, and simultaneously extract the body contour size, skin color gamut range, and hair texture density. Combine these with the stylization intensity level selected by the user to integrate them into the original geometric dataset.
[0019] The Mediapipe face detector was used to locate 68 key points in the user-uploaded image, obtaining the coordinate sequence of the forehead arc control points, the coordinates of the extreme points of the cheekbone protrusions, the coordinates of the sampling points of the jaw contour curve, and the coordinates of the chin endpoint. Based on the key point coordinates, the ratio of the distance between the pupils to the facial width, the vertical distance from the lower edge of the eyebrow to the upper edge of the eyelid, the lateral span of the left and right endpoints of the nose, the longitudinal length from the upper to the lower end of the philtrum, and the ratio of the thickness of the upper lip to the thickness of the lower lip were calculated to obtain the facial geometric feature vector. OpenPose human pose estimation was used to identify 18 skeletal nodes in the full-body image, obtaining the distance between the left and right acromion points as the shoulder width value, the lateral span at the widest point of the chest as the chest circumference benchmark, the lateral span at the narrowest point of the waist as the waist circumference benchmark, the lateral span at the widest point of the hip as the hip circumference benchmark, and the vertical height from the top of the head to the bottom as the height value. The facial geometric feature vector and the body contour parameter sequence were numerically normalized and then merged into a morphological data matrix. Based on the facial region coordinate range recorded by the morphological data matrix, the RGB pixel value set corresponding to the image position is extracted. The average pixel brightness is calculated as the skin tone brightness index, the average hue angle is calculated as the hue tendency index, and the saturation value is used to determine the warm or cool attribute label. The average saturation S is greater than 0.5 and is warm color, otherwise it is cool color, where S ranges from 0 to 1. Gabor filtering convolution operation is performed on the hair region to obtain the filter response amplitude distribution as texture direction feature, response spectrum peak as texture density feature, and response variance as texture roughness feature. The morphological data matrix, skin tone index sequence, hair texture feature, and stylization intensity value in the 0 to 1 range selected by the user interface slider jointly construct the original geometric dataset.
[0020] Specifically, when processing user-uploaded frontal photos using the Mediapipe facial detector, the 68 key points are distributed according to facial anatomy. The forehead arc includes 17 control points, extending from the left temple through the midline of the forehead to the right temple. Each control point records two-dimensional coordinates to form the forehead contour curve. The cheekbone position is determined by key points 2 to 16, located at the most prominent positions on both sides of the cheek. The algorithm identifies the extreme points of cheekbone protrusion by calculating curvature changes. The jawline contour is a continuous curve formed by key points 5 to 13, with the chin endpoint corresponding to key point 9, whose coordinates are recorded as a marker of the lowest point on the face.
[0021] Specifically, the calculation of the facial feature distance ratio involves multiple geometric measurements. The interpupillary distance is obtained through the Euclidean distance between key points 37 and 46. The facial width is taken as the horizontal span between key points 1 and 17, and the ratio of the two reflects the relative width characteristics of the interpupillary distance. The eyebrow-eye distance is measured as the vertical distance from the lower edge of the eyebrow (point 20) to the upper eyelid (point 38). The nasal wing width is calculated as the horizontal span between the endpoints of the nasal wings (points 32 and 36). The philtrum length is obtained as the vertical distance from the midpoint of the nasal base (point 34) to the midpoint of the upper lip (point 52). The upper and lower lip thicknesses are measured by the vertical heights of points 51 to 53 and 57 to 59, respectively, and the proportional relationship is calculated. These measurements are combined to form a facial geometric feature vector, with each component representing a specific facial structural feature.
[0022] It should be noted that OpenPose human pose estimation, when processing full-body images, identifies 18 skeletal nodes through a convolutional neural network, including key locations such as the head, neck, shoulders, elbows, wrists, hips, knees, and ankles. Shoulder width is obtained by the horizontal distance between the 2nd and 5th acromion points. Chest circumference is measured at the widest lateral span approximately 15 cm below the 1st neck node. Waist circumference is measured at the narrowest point above the 8th and 11th hip nodes. Hip circumference is measured at the widest span at the horizontal position of the hip nodes. These measurements, together with the vertical height from the 0th head node to the 19th foot node, constitute the body contour parameter sequence.
[0023] In one possible implementation, Gabor filtering uses a multi-scale, multi-directional filter bank to process hair texture, containing 40 filter kernels across 5 scales and 8 directions. Each filter kernel is convolved with the hair region to generate a response image. The spatial distribution of the filter response amplitude reflects the directional characteristics of the hair strands. The response spectrum is obtained through Fast Fourier Transform, with the peak frequency corresponding to the texture density. The variance of the response value measures the roughness of the texture. These features are integrated with the aforementioned morphological data matrix, skin color index sequence, and user-defined stylization intensity values to construct a primordial geometric dataset containing multi-dimensional information on facial geometry, body shape, skin color, and hair texture.
[0024] S102. Compare the original geometric dataset with the target style template to obtain the deformation depth value of each part. Analyze the risk of facial proportion imbalance based on the deformation depth value and output the level of reduced recognizability.
[0025] Read the standard facial feature parameters from the target style template, including the standard ratio of eye distance to face width, standard value of nasal wing height, and standard radius of zygomatic bone curvature for different style types. Calculate the difference between the eye distance width value in the original geometric dataset and the standard value of the style template to obtain the horizontal stretching amplitude coefficient. Simultaneously, extract the ratio of the standard value of nasal wing height in the style template to the nasal wing height in the original dataset as the vertical compression amplitude coefficient. Obtain the radius of curvature by fitting the zygomatic bone contour points using a B-spline curve. Calculate the curvature change amplitude coefficient by comparing the zygomatic bone curvature radius of the style template. Based on the lateral stretching amplitude coefficient, longitudinal compression amplitude coefficient, and curvature change amplitude coefficient, the deformation depth values of the eye region, nose region, and cheekbone region are calculated respectively. The deformation depth value of the eye region is obtained by multiplying the lateral stretching amplitude coefficient by a first preset weight, the deformation depth value of the nose region is obtained by multiplying the longitudinal compression amplitude coefficient by a second preset weight, and the deformation depth value of the cheekbone region is obtained by multiplying the curvature change amplitude coefficient by a third preset weight. When the deformation depth value of any region exceeds a preset deformation threshold, the region is marked as having a risk of facial feature disproportion, and the level of reduced recognizability is calculated.
[0026] Specifically, in one implementation, when reading the target style template, the corresponding parameter configuration file needs to be loaded according to the style type selected by the user. Cartoon style emphasizes enlarged eyes and rounded faces, realistic style maintains near-realistic facial proportions, and anime style pursues the ultimate effect of large eyes and a small mouth. The horizontal stretching coefficient is obtained by calculating the ratio of the original eye distance to the standard eye distance of the style template. When the original eye distance is 0.35 times the face width and the cartoon style requires 0.48, the horizontal stretching coefficient is 1.37, indicating that the eyes need to be expanded horizontally by 37%. The vertical compression coefficient reflects the adjustment requirements for nose height, calculated by dividing the style template's nose wing height by the original nose wing height. If the anime style requires simplifying the nose to a dot shape, the compression coefficient can reach 0.3, meaning the nose height is compressed to 30% of its original size.
[0027] Specifically, in the process of fitting the cheekbone contour using B-spline curves, eight control points in the cheekbone region are first extracted, and a cubic B-spline basis function is constructed. A smooth curve equation is obtained through least squares fitting. 100 points are uniformly sampled on the curve, and the curvature value of each point is calculated. The position corresponding to the maximum curvature is taken as the cheekbone vertex, and the radius of curvature at this point is the cheekbone arc feature value. The curvature change amplitude coefficient is equal to the difference between the style template's radius of curvature and the original radius of curvature divided by the original radius of curvature, reflecting the degree of deformation of the cheekbone contour. The deformation depth value is calculated using a weighted combination method, where the weight allocation is adjusted according to the feature emphasis of different style types. For cartoon styles, the focus is on eye changes, with an eye weight set to 0.5, and the nose and cheekbone each at 0.25. For anime styles, the nose is extremely simplified, with the nose weight reduced to 0.15, the eye weight increased to 0.6, and the cheekbone weight at 0.25. When the deformation depth value of a certain area exceeds a preset threshold, it indicates that the deformation of that area may lead to facial feature distortion, requiring risk labeling and close attention in subsequent processing. Optionally, for each region's deformation depth value, the expected degree of change in key facial features after deformation is calculated. Risk scores are calculated for each region: eye risk score = min(|interocular distance change rate / first preset proportional threshold|, 1.0), nose risk score = min(|nasal wing width change rate / second preset proportional threshold|, 1.0), and cheekbone risk score = min(|cheekbone position offset rate / third preset proportional threshold|, 1.0). Optionally, the first, second, and third preset proportional thresholds are 0.15, 0.2, and 0.12, respectively. A weighted summation method is used to calculate the overall level of reduced recognizability: Reduction in recognizability = α × eye risk score + β × nose risk score + γ × cheekbone risk score, where α + β + γ = 1.0, and the weights are α = 0.4, β = 0.3, and γ = 0.3. The result is a value within the range of 0-1; a larger value indicates a higher risk of reduced recognizability.
[0028] In one possible implementation, the assessment of the reduction in recognizability considers the degree of change across multiple dimensions, quantifying it by calculating the relative positional changes of key feature points before and after deformation. Changes in interocular distance are measured by the ratio of the lateral displacement of the pupil point to the original interocular distance; changes in nasal wing width are assessed by the scaling ratio of the distance between the nasal wing endpoints; and zygomatic bone position shift is determined by the lateral displacement of the zygomatic bone apex relative to the facial centerline. These indicators collectively determine the degree of similarity between the virtual avatar and the real face.
[0029] S103. When the level of recognizability decreases beyond the threshold allowed by the stylization intensity, an early warning is triggered. The aesthetic knowledge graph is then used to query the aesthetic standard node library to identify potential aesthetic risk points.
[0030] When the level of recognizability reduction exceeds a preset threshold allowed by stylization intensity, an early warning mechanism is triggered, and the specific numerical difference exceeding the threshold is recorded. The index address of the aesthetic standard node library is extracted from a pre-constructed aesthetic knowledge graph. This aesthetic standard node library includes golden ratio nodes recording the standard ratio between eye width and face width, facial harmony nodes storing coordination parameters of facial feature spacing, and East-West aesthetic difference nodes marking aesthetic preference data from different cultural backgrounds. A graph database query is used to retrieve the set of standard nodes closest to the current facial features. Based on the aesthetic parameter benchmark values in the set of standard nodes, the deviation of the current deformation depth value in each dimension is calculated. Specifically, the eye deformation depth value is subtracted from the standard eye width value of the golden ratio node to obtain the eye deviation; the nose deformation depth value is divided by the standard nose height value of the harmony node to obtain the nose deviation; and the cheekbone deformation depth value is calculated by dividing the contour curvature standard value of the aesthetic difference node to obtain the cheekbone deviation. These three deviation values constitute the aesthetic deviation vector. A threshold-based method is used to assess the risk level of the aesthetic deviation vector. Multiple risk level threshold ranges are preset. When the deviation of the eyes exceeds the first threshold, the deviation of the nose exceeds the second threshold, or the deviation of the cheekbone exceeds the third threshold, the corresponding area is marked as a high-risk region. When multiple areas are simultaneously marked as high-risk, the combination of deviation features of these areas is extracted as potential aesthetic risk points. For these potential aesthetic risk points, the correction strategy nodes matching the risk point type in the knowledge graph are traversed. Each correction strategy node contains a risk type identifier, suggested adjustment parameters, and expected improvement effect attributes. The corresponding correction strategy is found through attribute matching. The deformation depth correction coefficient in the suggested adjustment parameters is read to identify the specific risk point location and risk level that exceeds the aesthetic standard in the current deformation scheme.
[0031] Specifically, in one implementation, the warning mechanism is triggered based on a tiered threshold determination. When a user selects a chibi style and sets the intensity level to 0.7 in a game character customization scenario, the preset allowable threshold for the reduction in recognizability is 0.65. If the actual calculated reduction in recognizability reaches 0.82, exceeding the threshold of 0.17, an warning is immediately triggered and the difference is recorded. The aesthetic knowledge graph is constructed using the Neo4j graph database. Nodes are connected by relational edges to form a network structure. Each node stores specific aesthetic parameters and applicable conditions. The golden ratio node records the 0.618 ratio between facial width and eye distance, the 0.382 ratio between nose length and face length, and the 1.618 ratio between the upper lip to chin and the base of the nose to the upper lip in Western classical aesthetics. These values are derived from statistical analysis of artworks from the Renaissance period. The facial harmony node includes the three courts and five eyes principle, namely the spatial distribution pattern of dividing the face into three equal parts vertically and five equal parts horizontally, as well as the equidistant relationships from the center of the eyebrows to the pupil, from the pupil to the base of the nose, and from the base of the nose to the chin. The nodes representing differences in aesthetic preferences between East and West record the range of curvature of the oval face preferred by Asian users and the three-dimensional skeletal structure parameters favored by European and American users, distinguishing aesthetic preferences from different cultural backgrounds through regional tags.
[0032] Specifically, the graph database query process achieves precise retrieval through Cypher statements. After inputting the current facial feature parameters, the query first locates a subset of nodes matching the user's selected style type, and then finds the top 10 closest standard nodes by calculating the cosine similarity between feature vectors. This set of standard nodes contains widely accepted aesthetic parameter benchmarks for that style; for example, the proportion of large eyes in a cartoon style face is typically between 0.25 and 0.35, while in a realistic style it remains between 0.15 and 0.20. The query results return nodes containing not only numerical parameters but also descriptions of applicable scenarios for that aesthetic standard, user group characteristics, and links to historical application cases, facilitating subsequent deviation analysis and risk assessment. The deviation calculation employs differentiated measurement methods for different facial regions. Eye deviation is obtained by calculating the absolute difference between the current eye deformation depth value and the standard eye width value in the golden ratio nodes; this difference reflects the distance the lateral stretching of the eye deviates from the ideal aesthetic standard. Nasal deviation is calculated using a ratio, which is the current nasal deformation depth divided by the standard nasal height value at the harmony node. A ratio less than 1 indicates that the nose is over-compressed, while a ratio greater than 1 indicates that the nose is elongated. The greater the deviation of the ratio from 1, the higher the risk. Cheekbone deviation is obtained by calculating the Euclidean distance between the curvature value corresponding to the current cheekbone deformation depth and the standard curvature value at the aesthetic difference node. The larger the distance, the more severe the contour deformation.
[0033] In one possible implementation, the three deviation values are combined into a three-dimensional aesthetic deviation vector. The position of this vector in three-dimensional space visually reflects the overall degree of deviation. The magnitude of the vector represents the overall deviation intensity, and the direction angle indicates the facial area where the deviation is most pronounced.
[0034] For example, when the vector mainly extends along the eye dimension, it indicates that the eye is the main source of aesthetic risk, and the deformation parameters related to eye distance need to be adjusted.
[0035] Preferably, the risk level assessment adopts a multi-level threshold determination mechanism, with five preset threshold ranges corresponding to risk levels: extremely low risk corresponds to a deviation of 0 to 0.15, low risk corresponds to 0.15 to 0.30, medium risk corresponds to 0.30 to 0.50, high risk corresponds to 0.50 to 0.75, and extremely high risk corresponds to above 0.75. When the deviation of the eyes is 0.62, the deviation of the nose is 0.38, and the deviation of the cheekbones is 0.71, the system determines that the eyes and cheekbones are high-risk areas, and the nose is a medium-risk area. If two or more parts reach the high-risk level at the same time, the combined features of these parts are extracted, such as "excessively enlarged eyes + excessively prominent cheekbones" as a descriptive label for potential aesthetic risk points.
[0036] For example, the correction strategy nodes in the knowledge graph are stored according to risk type, with each node containing three core attribute fields. Risk type identifiers are recorded using an coded method, such as "EH-01" representing high-risk type 1 for the eye, facilitating rapid indexing and matching. Suggested adjustment parameters include specific numerical adjustment suggestions, such as "reduce the horizontal stretching coefficient of the eye by 15%" or "increase the radius of curvature of the cheekbone by 20 pixels." These parameters can be directly applied to the correction calculation of deformation depth values. The expected improvement effect is quantified as a percentage to represent the degree of risk reduction, helping the system evaluate the effectiveness of the correction strategy.
[0037] Understandably, the attribute matching process iterates through all correction strategy nodes, calculates the string similarity between the current risk point description and the node risk type identifier, and selects the strategy with the highest similarity as the recommended solution. When multiple strategies with close similarity exist, the system will comprehensively consider the user's style preference strength and historical adjustment records, prioritizing correction strategies with moderate adjustment magnitude and minimal impact on the overall style.
[0038] In one embodiment, the reading and application of the deformation depth correction coefficient follows a progressive adjustment principle. Initially, 50% of the coefficient is applied for fine-tuning, and the adjustment effect is evaluated before deciding whether further correction is needed. This progressive method avoids style distortion caused by over-correction while preserving the artistic effect desired by the user. During the correction process, the aesthetic deviation vector is updated in real time. When the deviation in all dimensions drops below medium risk, the final coordinates of the aesthetic risk point and the corresponding quantified risk level are output. This information constitutes the complete aesthetic risk identification result.
[0039] S104. Identify corresponding historical cases based on aesthetic risk point location, and select relevant reference cases based on the similarity between current user characteristics and historical successful cases.
[0040] Based on the risk point location coordinates in the aesthetic risk identification results, records with the same risk point type are retrieved from the historical case database. Body shape feature vectors containing user height-weight index, shoulder-to-hip ratio, and face length-to-width ratio are obtained. Simultaneously, the stylization processing parameter sequence, user rating value, and usage scenario tag for each record are extracted. The usage scenario tags are divided into four categories: game character, social avatar, virtual meeting, and live stream image. A historical case feature matrix is constructed. Cosine similarity is used to calculate the similarity between the current user's body shape feature vector and each row of vectors in the historical case feature matrix. If the similarity exceeds a preset first threshold, it is marked as a high-similarity case. From the high-similarity cases, corresponding data pairs of deformation depth parameter sequences and user satisfaction scores are extracted. Linear regression fitting is performed on each part, such as the eyes, nose, and cheekbones, with deformation depth parameters as independent variables and satisfaction scores as dependent variables, to obtain the correlation function between deformation depth and satisfaction. The average of the correlation functions for multiple parts is taken to obtain a unified function. Based on the unified function, the deformation depth value range corresponding to a satisfaction score higher than a preset standard is determined. Based on the deformation depth range, highly similar cases are selected. The aesthetic preference distribution data of the selected cases in the user's selected scenario is queried. The degree of eye magnification, nose shaping depth, and cheekbone contour sharpness in the scenario are statistically analyzed. The value with the highest frequency of cheekbone edge clarity is defined as the preference feature value. The difference between the current user's deformation parameters and the preference feature value is calculated. When the difference is less than the preset second threshold and the satisfaction score of the corresponding case exceeds the preset standard, the historical case is output as a relevant reference case.
[0041] Specifically, in one implementation, the historical case database adopts a relational database storage structure, with each record containing four main fields: a unique identifier, a timestamp, basic user information, and processing result. During the extraction of body shape feature vectors, the height-to-weight ratio is calculated by dividing weight by the square of height; the shoulder-to-hip ratio reflects the upper-to-lower body proportion; and the face length-to-width ratio reflects facial features. These values are normalized to form a three-dimensional vector, facilitating subsequent similarity calculations. The stylization parameter sequence records the specific values of deformation depth in various regions such as the eyes, nose, and cheekbones in each case, and the user rating uses a ten-point scale to record the user's satisfaction with the final result.
[0042] Specifically, cosine similarity is calculated by dividing the vector dot product by the product of the vector magnitudes. Its value ranges from -1 to 1, with values closer to 1 indicating greater similarity in body shape features. When the body shape feature vector of game player A is [0.23, 0.67, 0.45] and the feature vector of historical case B is [0.25, 0.65, 0.48], the calculated similarity is 0.996, far exceeding the preset threshold of 0.85. Therefore, case B is marked as a highly similar case. Data pairs of deformation depth parameters and satisfaction scores are extracted from all highly similar cases. For example, an eye deformation depth of 0.3 corresponds to a satisfaction score of 7, 0.4 to 8, and 0.5 to 6. A linear relationship is obtained by fitting the least squares method. The slope and intercept parameters reflect the degree of influence of deformation depth changes on satisfaction.
[0043]
[0044] S represents the satisfaction rating, D represents the deformation depth parameter, a represents the slope coefficient of the linear relationship, and b represents the intercept parameter. This formula describes the linear relationship between deformation depth and satisfaction, with the slope 'a' reflecting the degree of influence of changes in deformation depth on satisfaction. Optional,
[0045] n represents the total number of sample data, Di represents the i-th deformation depth value, Si represents the i-th satisfaction score value, and a represents the slope parameter calculated by the least squares method. This formula is the analytical solution for solving the linear regression slope using the least squares method, used to quantify the strength of the influence of deformation depth on satisfaction. The correlation function is established based on the statistical regularity of a large amount of historical data. When the satisfaction score is set to 8 points or above, the corresponding deformation depth value range is obtained by inversely solving the fitted function, such as the deformation depth of the eye between 0.35 and 0.45, the nose between 0.2 and 0.3, and the cheekbone between 0.15 and 0.25.
[0046] Preferably, the statistical analysis of aesthetic preference distribution data employs a frequency analysis method. In game character scenarios, 1000 historical cases are analyzed. The most frequent value for eye magnification is 1.3 times, the peak value for facial roundness is 0.7, and the contour sharpness is concentrated around 0.4. These high-frequency values serve as the preference feature values for this scenario. The current user's deformation parameters are compared item by item with the preference feature values, and the absolute value of the difference for each item is calculated. When all difference values are less than a preset threshold and the satisfaction score of the corresponding case exceeds 8 points, the case is selected as a reference case, providing users with successful stylization processing experience for reference.
[0047] S105. Combining aesthetic standards and historically relevant reference cases to drive the aesthetic knowledge graph to conduct multi-dimensional adaptability assessment and output an image blueprint document.
[0048] Combining the user-inputted locations of aesthetic risk points with relevant reference cases, the rule matching process in the aesthetic knowledge graph is initiated. Aesthetic rule nodes corresponding to the current user's body shape and dimensions are extracted from the graph. These nodes include the shoulder-to-height ratio range, the chest-to-waist difference range, and the hip-to-shoulder-width coordination coefficient. A body shape fit score is obtained by calculating the deviation between the user-input deformation depth value and the standard values in the aesthetic rule nodes. Brightness, hue angle, and saturation data from the user's skin tone color gamut are extracted. The style preference records for the corresponding skin tone category in the aesthetic knowledge graph are queried. Visual effect parameters generated by the current deformation depth value are calculated, including the matching degree between the brightness adjustment value and saturation offset and skin tone features. The product of brightness and style saturation reflects color harmony, and the correspondence between hue angle and lighting treatment determines rendering suitability. A style fit index is obtained through weighted summation. The system reads the user's hair texture density parameters, including hair thickness, curl curvature, and hair density distribution. Based on the style type determined by the style compatibility index, it matches the user's preset usage occasion with the required hairstyle for that style, calculating the degree of conformity between the hair texture parameters and the scene specifications to obtain the occasion matching value. Combining the body shape fit score, style compatibility index, and occasion matching value, a multi-dimensional adaptability assessment result is generated through weighted calculation using preset weighting coefficients. An image blueprint document is constructed, including final deformation depth parameters for each facial region, body contour adjustment suggestions, skin tone rendering schemes, and hairstyle processing guidelines. The output image blueprint document includes virtual avatar preview data and a parameter configuration list. It should be noted that the image blueprint document output in this application can be divided into two main branches based on actual application: one is offline delivery based on image design, and the other is based on online applications, such as games, VR, and mobile applications. The difference lies in selecting different types of target style templates in step S102. Of course, the target style templates for different branches can also be set to more subdivided types according to the needs of each branch.
[0049] Specifically, in one implementation, the rule matching process of the aesthetic knowledge graph is implemented based on the node traversal mechanism of the graph database. Each aesthetic rule node stores the standard parameter range for a specific body type category. After the user inputs their body contour dimensions, the ratio of shoulder width to height is first calculated. If the ratio is between 0.23 and 0.25, it is classified as a standard body type; 0.25 to 0.28 is a broad-shouldered body type; and 0.20 to 0.23 is a narrow-shouldered body type. The difference between chest circumference and waist circumference reflects the upper body curve characteristics. A difference greater than 15 cm indicates a pronounced curve, while a difference less than 8 cm indicates a relatively straight curve. The coordination coefficient between hip circumference and shoulder width is calculated by dividing hip circumference by shoulder width. A coefficient close to 1.0 indicates a balanced upper and lower body; a coefficient greater than 1.2 leans towards a pear-shaped body; and a coefficient less than 0.8 indicates an inverted triangle body type.
[0050] Specifically, the calculation process for body shape fit score involves a comparative analysis of deformation depth values and aesthetic rules. For users with broad shoulders, if the deformation depth value of the horizontal stretching of the eyes exceeds 0.4, it will further enhance the visual effect of being wider at the top and narrower at the bottom, leading to an imbalance in proportions. In this case, the system calculates the deviation as the difference between the deformation depth value and the recommended upper limit of 0.3; the larger the deviation, the lower the fit score. Conversely, users with narrow shoulders are suitable for moderate horizontal expansion to balance the overall proportions, and a higher fit score is obtained when the deformation depth value is in the range of 0.3 to 0.4. The score uses a 100-point scale, with 100 points awarded for complete compliance with aesthetic rules, and 15 points deducted for each 0.1 unit deviation, ensuring the accuracy of the quantitative assessment.
[0051] It's important to note that matching skin tone characteristics with visual effects involves the application of color theory principles. Brightness values are represented using the L component of the HSL color space, ranging from 0 to 100. Asian skin tones typically range from 55 to 75, while those of Caucasian users range from 65 to 85. Hue angle reflects the warm or cool undertone of skin tone: 0 degrees is pure red, 60 degrees is yellowish, and 180 degrees is bluish. Asian skin tones are mostly concentrated in the warm yellow hue range of 30 to 50 degrees. Saturation data represents the purity of color; low saturation presents a natural skin tone, while high saturation leans towards an artistic effect. The style fit index is calculated by multiplying the brightness value and style saturation by 100 to obtain a color harmony score. This score is then combined with the matching degree of hue angle and preset lighting treatment methods, and the weighted sum is normalized to the 0-1 range.
[0052] Preferably, the hair texture density parameter is obtained through image processing algorithms. The thickness of hair strands is identified by edge detection algorithms, using the pixel width of a single hair strand: 1 to 2 pixels for fine hair and 3 to 5 pixels for thick hair. The curvature value is calculated by fitting the radius of curvature of the hair strand trajectory; straight hair has a curvature close to 0, large waves have a curvature between 0.2 and 0.4, and small curls have a curvature exceeding 0.6. The density distribution of hair is assessed by statistically analyzing the number of hair strands per unit area: dense hair has more than 150 strands per square centimeter, while sparse hair has less than 80 strands per square centimeter.
[0053] In one possible implementation, scene specifications are defined based on the socio-cultural requirements of different usage scenarios. Business meeting scenarios require a professional and rigorous image, with hairstyle deformation depth controlled within 0.2 to avoid overly exaggerated styles; social gathering scenarios allow for personalized expression, and deformation depth can be relaxed to 0.5, supporting fashionable and avant-garde styles; game character scenarios pursue fun and recognizability, with deformation depth reaching 0.8 to achieve cartoonish or artistic effects. The occasion matching score is obtained by calculating the degree of conformity between the actual hair quality processing parameters and the scene specification requirements. Full conformity earns full marks, and points are deducted for each 0.1 unit exceeding the specification range.
[0054] For example, the multi-dimensional suitability assessment uses a weighted average method to combine the scores of the three dimensions. The preset weighting coefficients are dynamically adjusted based on the user's chosen primary application scenario. If the user primarily uses it for social media, the style fit weight increases to 0.5, while body type fit and occasion fit each account for 0.25. If used for virtual meetings, the occasion fit weight increases to 0.4, with the other two each accounting for 0.3. The weighted overall score reflects the overall suitability of the virtual avatar in a specific application scenario. A score above 80 indicates ideal performance, 60-80 requires fine-tuning, and a score below 60 suggests reselecting style parameters.
[0055] Understandably, the image blueprint document uses a JSON format to organize its data structure, with the top level containing three metadata fields: user identifier, creation time, and style type. The facial deformation parameters section records detailed final deformation depth values for each area, including the eyes, nose, mouth, cheekbones, and jawline, with each value accurate to three decimal places. Body contour adjustment suggestions are described in text form, outlining the optimization direction for each part, such as "It is recommended to moderately narrow the shoulder contour to balance the overall proportions." The skin tone rendering scheme includes basic skin tone RGB values, highlight color parameters, and shadow depth settings to ensure consistent rendering results.
[0056] For example, the hairstyle guidance section provides specific styling parameters based on hair type and scene requirements, including detailed settings such as the overall hairstyle outline, bangs length ratio, hair end curl level, and hair volume. The virtual avatar preview data includes links to thumbnails of rendered images from multiple angles—front, side, and 45-degree angles—allowing users to view the effect from all angles. The parameter configuration list presents all adjustable parameters in tabular form, including their names, current values, suggested values, and adjustment ranges. Users can use this list for fine-tuning to achieve personalized customization.
[0057] In one embodiment, the avatar blueprint document also supports version management. Each parameter adjustment is saved as a new version, and users can view historical versions to compare changes and select a satisfactory version to restore or continue optimization. This iterative optimization mechanism helps users gradually approach the ideal virtual avatar effect.
[0058] S106. Transfer the original geometric dataset to the stylization rendering stage. Use VToonify to complete the 2D facial style transfer rendering to transform the real facial features into a stylized planar image. Use 3D Gaussian splashing to reconstruct the 2D stylization rendering result under the deformation depth value constraint into a 3D stereoscopic image. Deploy it to the application scene and record the rendering stability and attribute fusion effect of each deformation depth range as online feedback data.
[0059] The facial contour coordinates, facial feature spacing ratios, body contour dimensions, skin color gamut range, and hair texture density parameters from the original geometric dataset are input into the VToonify renderer. A convolutional neural network extracts facial feature vectors and maps them to a style space. Based on the user-selected cartoon or artistic style, corresponding style template parameters are loaded. A weighted fusion of the feature vectors and style templates is performed to obtain a 2D stylized facial image. This 2D stylized facial image retains the key recognition features of the original face while presenting the visual effect of the target style. Using this 2D stylized facial image as a texture map, 3D point cloud data is constructed by combining deformation depth values. Gaussian kernels are distributed in 3D space using 3D Gaussian splashing. The position coordinates of each Gaussian kernel are determined by the point cloud data. The rotation quaternion and scaling factor are calculated based on the deformation depth values, and the opacity is set according to the texture pixel brightness. The Gaussian kernel parameters are optimized using gradient descent until a 3D stereoscopic image is reconstructed. Based on the data scale of the 3D avatar, adaptation processing is performed for different application scenarios. If deployed to a game scene, the number of polygons is reduced to a preset upper limit using a vertex merging algorithm. If deployed to a VR environment, geometric details are increased using a subdivision surface algorithm. If deployed to a mobile interface, the resolution is reduced using a texture compression algorithm, outputting the adapted virtual avatar data. The rendering performance of the virtual avatar data on the target platform is monitored. When the deformation depth value is less than a first threshold, the average and standard deviation of the rendering frame rate are recorded. When the deformation depth value is greater than a second threshold, the texture sampling error and geometric edge smoothness are statistically analyzed. The attribute fusion coherence is evaluated by calculating the angle between the normal vectors of adjacent vertices. The performance indicators and fusion effect data are stored as online feedback data.
[0060] Specifically, in one implementation, the VToonify renderer uses an encoder-decoder architecture to process the original geometry dataset. The encoder consists of a multi-layer convolutional neural network, with each layer containing convolution operations, batch normalization, and activation functions. Facial contour coordinates are normalized and then input into the first convolutional layer with a 3×3 kernel size and a stride of 1, outputting 64 feature channels. The facial feature spacing ratio and body contour dimensions are used as conditional vectors and mapped to the same dimensional space as the convolutional features through a fully connected layer. The two are then fused through feature concatenation. The skin color gamut is converted to RGB three-channel values, with each channel ranging from 0 to 255, and mapped to the color reference in the style space after color space transformation. Hair texture density parameters are extracted using a texture encoder to extract high-frequency detail features, which are then aligned with facial features in the latent space.
[0061] Specifically, the loading process of style template parameters is dynamic based on user selection. Cartoon style templates include attributes such as line simplification coefficients, color block smoothing parameters, and contour exaggeration, while artistic style templates focus on features such as brushstroke texture, color gradients, and light and shadow contrast. The weighted fusion of feature vectors and style templates is achieved through an attention mechanism, calculating the correlation weights between each dimension of the feature vector and the style template. These weights are normalized using a softmax function to ensure a sum of 1. The fused features are then upsampled layer by layer through a decoder network. Each upsampling doubles the feature map resolution while halving the number of channels, outputting a 512×512 resolution 2D stylized facial image. This image retains key recognition features of the original face, such as eye position, nose shape, and mouth contour, while also presenting the line style, color characteristics, and artistic effects of the target style. The core of the 3D Gaussian splashing technology in reconstructing a 3D model from a 2D image lies in the spatial distribution optimization of the Gaussian kernel. Initial point cloud data is obtained by projecting each pixel of the 2D image into 3D space based on depth estimates. The depth values are predicted by a monocular depth estimation network, ranging from 0.5 meters to 3 meters. The position coordinates of each Gaussian kernel are initialized to the 3D coordinates of the corresponding point cloud. The rotation quaternion is randomly initialized and then normalized to ensure unit length. The scaling factor is set according to the deformation depth value; the larger the deformation depth value, the larger the scaling factor value in the corresponding direction, achieving the stretching or compression effect of local areas. The opacity parameter is related to the alpha channel of the texture pixel; completely opaque is set to 1, completely transparent is set to 0, and semi-transparent areas are linearly mapped according to the actual transparency.
[0062] Preferably, the gradient descent optimization process employs an adaptive learning rate strategy, with an initial learning rate of 0.01. After every 100 iterations, the learning rate decays to 0.9 times its original value. The optimization objective function comprises two parts: rendering loss and a regularization term. The rendering loss calculates the pixel difference between the rendered image and the target image, using the L1 norm as a metric. The regularization term constrains the Gaussian kernel parameters within a reasonable range, preventing excessively large or small outliers. During the iteration process, the gradient of the loss function with respect to each parameter is calculated using differentiable rendering technology, and the parameter values are updated according to the gradient direction. When the change in the loss function over 10 consecutive iterations is less than a preset threshold, the optimization is considered converged, and the final 3D image is output.
[0063] In one possible implementation, a differentiated strategy is adopted for adaptation to different application scenarios. When deployed in game scenarios, the vertex merging algorithm calculates the Euclidean distance between adjacent vertices; when the distance is less than a preset threshold, they are merged into a single vertex, and the topological connectivity of related triangles is updated simultaneously. The maximum number of polygons is set according to the performance of the target platform; mobile games are typically limited to 5000 triangles, while PC games can be relaxed to 20000. For VR environment deployment, a subdivision surface algorithm is used. Using the Catmull-Clark subdivision rule, each quadrilateral is subdivided into four sub-faces, with special rules used at the boundaries to preserve contour features. Texture compression for mobile interfaces uses the ETC2 format, compressing the original RGBA texture to 1 / 4 of its original size while maintaining acceptable visual quality.
[0064] For example, rendering performance monitoring encompasses the collection of metrics across multiple dimensions. Frame rate monitoring is achieved by recording the number of frames rendered per second. In the low deformation range (deformation depth value less than 0.3), the rendering time for 100 frames is statistically analyzed, and the average frame rate and standard deviation are calculated. Texture sampling error is assessed by comparing the difference between the actual sampled value and the ideal sampled value. When the deformation depth value exceeds 0.7 and enters the high deformation range, texture stretching in edge areas can easily lead to increased sampling error. Geometric edge smoothness is evaluated by calculating the angle between the normal vectors of adjacent triangular facets. The smaller the angle, the smoother the transition; an angle exceeding 30 degrees may result in noticeable sharp edges.
[0065] Understandably, the evaluation of attribute fusion coherence focuses on the visual harmony between different attributes. Color transitions between the eye area and other facial areas are evaluated by calculating the color difference of boundary pixels, using Euclidean distance in the CIELab color space. The connection between hairstyle and head contour is determined by detecting the continuity of the hairline; the number of discontinuous points reflects the fusion quality. These performance metrics and fusion effect data are organized according to dimensions such as timestamps, scene types, and deformation parameters, and stored in a structured database as online feedback data.
[0066] S107. Based on offline and online feedback data, evaluate the optimal range of stylization intensity within the range of recognizability maintenance, and output the virtual image generation quality evaluation result.
[0067] User satisfaction ratings and return rate statistics are extracted from offline delivery records. The satisfaction ratings record the user's acceptance of the virtual avatar, and the return rate is calculated by dividing the number of returned items by the total number of deliveries. Simultaneously, rendering stability indicators and attribute fusion effect parameters from online feedback data are read. The delivery time of offline data is matched with the rendering time of online data to construct a comprehensive dataset including stylization intensity, recognizability, satisfaction, and performance indicators. For this comprehensive dataset, stylization intensity values are sorted from low to high and divided into a preset number of intervals. The mean and variance of recognizability values within each interval are calculated. When the mean recognizability is higher than a first preset threshold and the variance is less than a second preset threshold, the stylization intensity interval is marked as a candidate optimal interval. The correlation coefficient is obtained by dividing the covariance of the satisfaction rating sequence and the recognizability sequence by the product of the standard deviations of the two sequences to determine the strength of their association. The quality score of each candidate interval is obtained by multiplying the mean recognizability, mean satisfaction, and correlation coefficient of the candidate optimal interval by the corresponding preset weight coefficients and summing them. The weight coefficients are preset according to the application scenario. The interval with the highest quality score is selected as the optimal stylization intensity interval, and the virtual image generation quality evaluation result containing the intensity range of the interval, the expected recognizability level, and the comprehensive quality score is output.
[0068] Specifically, in one implementation, offline delivery record data is extracted by accessing the order management database. Each delivery record contains four core fields: order number, delivery time, user rating, and return flag. User satisfaction ratings are recorded using a numerical scale, with consecutive integers from 1 to 10 representing rating levels from extremely dissatisfied to very satisfied. The return rate is calculated statistically according to a time window, dividing the number of returns within a month by the total number of deliveries during the same period to obtain the return rate percentage for that time period. Online feedback data is read from rendering log files, including performance metrics such as timestamps for each rendering task, frame rate values, memory usage, and texture loading time.
[0069] Specifically, the time matching mechanism achieves data alignment by establishing a mapping relationship. Offline delivery times are typically recorded to the day, while online rendering times are accurate to the millisecond. By rounding the rendering time up to the day, the corresponding delivery batch is found. When multiple rendering tasks exist on the same day, the average performance metric of all rendering tasks on that day is calculated as a representative value. The comprehensive dataset is organized in a tabular structure, with each row representing a user case and columns containing data from multiple dimensions such as stylization intensity, identifiability value, satisfaction rating, rendering frame rate, and memory usage, forming a structured dataset for subsequent analysis.
[0070] It should be noted that the stylized intensity range is divided using an equal-width binning method, uniformly dividing the intensity value range from 0 to 1 into 10 intervals, each with a width of 0.1. Within each interval, the mean discernibility is obtained by summing the discernibility values of all samples and dividing by the number of samples, while the variance reflects the dispersion of discernibility within that interval. The first preset threshold is typically set to 0.7, indicating that a discernibility of 70% or higher is considered acceptable; the second preset threshold is set to 0.05 to ensure that the discernibility within the interval is relatively stable and does not fluctuate significantly.
[0071] Preferably, the calculation of the correlation coefficient involves solving for the covariance and standard deviation. The covariance is obtained by averaging the products of the deviations of the satisfaction scores and the identifiability; a positive value indicates a positive correlation, and a negative value indicates a negative correlation. The standard deviations of the two series are calculated by taking the square root of the sum of the squares of the differences between their respective values and the means, and are used to normalize the covariance. The correlation coefficient ranges from -1 to 1; the closer the absolute value is to 1, the stronger the linear correlation, while a value close to 0 indicates that there is no obvious linear relationship between the two.
[0072] For example, the weighting coefficients of the quality score are adjusted according to different application scenarios. Game scenarios prioritize the artistry of visual effects, with stylistic intensity weighted at 0.4, recognizability weighted at 0.3, and satisfaction weighted at 0.3. Business scenarios require maintaining a professional image, so recognizability weight is increased to 0.5, stylistic intensity weight is decreased to 0.2, and satisfaction weight remains at 0.3. The comprehensive quality score for each candidate interval is calculated by weighted summation, and the interval with the highest score is selected as the recommended stylistic intensity range. The output quality assessment result includes the upper and lower limits of the interval, the corresponding average recognizability level, and the comprehensive quality score.
[0073] S108. Feedback the quality assessment results of virtual image generation to the aesthetic knowledge graph, update the body shape, style, and occasion matching records in the deformation depth adaptability case library, and realize continuous optimization of multimodal virtual image generation assessment and verification.
[0074] The optimal stylization intensity range, recognizability level, and overall quality score from the virtual avatar generation quality assessment results are transmitted to the update interface of the aesthetic knowledge graph. A new record containing the user's body shape characteristics, selected style type, final deformation depth value, and user satisfaction score is added to the corresponding aesthetic rule node. This new record is stored as a successful matching case through the node attribute update operation of the graph database. Based on the successful matching case, historical records similar to the current body shape characteristics are accessed in the deformation depth adaptability case library. The deviation between the deformation depth value of the new case and the average of the historical records is calculated. If the deviation is within a preset range, the new case is added to the case set of the corresponding body shape category. If the deviation exceeds the range, it is marked as an abnormal case and stored separately. Simultaneously, the statistical value of the number of successful matches for each style type and usage scenario combination under that body shape category is updated. The updated case library is used to statistically analyze the distribution of stylization intensity values under each body type category. The intensity values are sorted in ascending order, and the value in the middle position is taken as the recommended intensity. The retention threshold sequence is calculated, and the first quartile and the third quartile are determined as the boundary of the stable interval. When the number of cases in the same body type category exceeds the preset number threshold, the recommended stylization intensity value and retention threshold range of that category are locked as stable parameters for output, so as to achieve continuous optimization of virtual image generation evaluation and verification.
[0075] Specifically, in one implementation, the feedback mechanism for the virtual avatar generation quality assessment results is asynchronously transmitted via a message queue. The assessment results, including the optimal stylization intensity range, recognizability level, and overall quality score, are first encapsulated into a JSON format data packet. The aesthetic knowledge graph is stored using the Neo4j graph database. Each aesthetic rule node has a unique identifier and attribute list. New records are added to the corresponding node using the MERGE operation of the Cypher statement to avoid duplicate records. User body features are stored as numerical vectors, including dimensions such as height-to-weight ratio, shoulder-to-hip ratio, and face length-to-width ratio. Style types are represented by enumerated values, such as cartoon style as 1, realistic style as 2, and artistic style as 3. Deformation depth values and satisfaction scores are stored directly as floating-point numbers.
[0076] Specifically, access to the deformation depth adaptation case library is achieved through index lookup for rapid location. Body shape feature similarity is determined using Euclidean distance, calculating the distance between the body shape vector of a new case and historical body shape vectors. Records with a distance less than a preset threshold are identified as similar cases. In the deviation calculation process, the arithmetic mean of the deformation depth values in the similar case set is first calculated, and then the absolute value of the difference between the new case's deformation depth value and this mean is calculated. The preset range is typically set to ±20% of the average value. New cases within this range are considered to conform to the standard processing pattern for that body shape category and are directly added to the existing case set; cases outside this range indicate the use of unconventional deformation depth settings and need to be stored separately for subsequent analysis of their specificity.
[0077] It's important to note that the success rate statistics employ a three-dimensional counting matrix structure, with the three dimensions corresponding to body type, style type, and usage scenario, respectively. Each time a new case is successfully stored, the count value at the corresponding position is incremented by 1. This statistical method facilitates quick lookup of the historical success rate for specific combinations, providing a reference for new users.
[0078] Preferably, the statistical analysis of stylization intensity distribution begins when the number of cases reaches 30 or more. All cases are sorted by stylization intensity values from smallest to largest, and the value in the middle of the sequence is the median, serving as a reference benchmark for recommendation intensity. The quartiles are calculated by dividing the sorted data into four equal parts. The first quartile, Q1, represents 25% of the data below this value, the third quartile, Q3, represents 75% of the data below this value, and the interval between Q1 and Q3 is defined as the stable interval, indicating that most users' choices are concentrated within this range.
[0079] For example, the parameter locking mechanism is triggered after a certain number of cases have been accumulated, typically with a threshold of 100 cases. The locked parameters serve as the standard configuration for that body type category, allowing new users to directly adopt these validated parameter values and reducing trial-and-error costs. Continuous optimization is achieved through periodic recalculation, with parameter updates triggered every 50 new cases to ensure that the recommended values always reflect the latest user preference trends, thus achieving dynamic optimization of the virtual avatar generation quality.
[0080] In summary, the overall architecture of this application is as follows: First, a 1:1 digital twin model of the user is constructed through a multimodal high-precision data acquisition system. Then, in step S101, the original geometric dataset is obtained by integrating the digital twin model. In steps S102 to S105, corresponding image blueprint documents are generated based on the original geometric dataset and target style templates of different branches. Then, in step S106, the desired virtual image is obtained by stylized rendering. Finally, in steps S107 and S108, the evaluation is performed and fed back to the aesthetic knowledge graph to optimize the quality of virtual image generation.
[0081] Obviously, those skilled in the art can make various modifications and variations to the embodiments of this application without departing from the spirit and scope of the embodiments of this application. Therefore, if these modifications and variations to the embodiments of this application fall within the scope of the claims of this application and their equivalents, this application also intends to include these modifications and variations.
Claims
1. A multimodal avatar generation monitoring and evaluation validation method, characterized in that, The method comprises: forming a real appearance geometry dataset according to user real appearance data; comparing the real appearance geometry dataset with a target style template to obtain a deformation depth value of each part, analyzing a facial feature proportion imbalance risk according to the deformation depth value, and outputting a recognizable degree reduction level; when the recognizable degree reduction level exceeds a threshold value allowed by the stylization intensity, identifying a potential aesthetic risk point through an aesthetic knowledge graph; locating and identifying corresponding historical cases based on the aesthetic risk point, and filtering out relevant reference cases according to the similarity between the current user features and the historical successful cases; performing multidimensional adaptability evaluation on the aesthetic knowledge graph by combining aesthetic standard constraints and historical relevant reference cases to drive the aesthetic knowledge graph, and outputting an image blueprint document; after the real appearance geometry dataset is transmitted to a stylization rendering link for stylization rendering and then deployed to an application scenario, recording the rendering stability and attribute fusion effect of each deformation depth interval as online feedback data; evaluating the optimal interval of the stylization intensity within the recognizable degree maintenance range according to the offline feedback data and the online feedback data, and outputting a virtual image generation quality evaluation result.
2. The multi-modal avatar generation monitoring and evaluation validation method of claim 1, wherein, The comparison of the real appearance geometry dataset with the target style template to obtain the deformation depth value of each part, the analysis of the facial feature proportion imbalance risk according to the deformation depth value, and the output of the recognizable degree reduction level comprise: reading standard facial feature parameters in the target style template, calculating a deformation amplitude coefficient by taking the difference between the feature values in the real appearance geometry dataset and the standard values, calculating the deformation depth values of the eye region, the nose region and the cheekbone region according to the deformation amplitude coefficient, and calculating the expected change degree of the key features of the deformed face according to the deformation depth values of the regions to determine the recognizable degree reduction level.
3. The method of claim 1, wherein, The identification of the potential aesthetic risk point comprises: triggering an early warning mechanism when the recognizable degree reduction level value exceeds a preset threshold value, extracting an aesthetic standard node library from the aesthetic knowledge graph, the aesthetic standard node library including a golden ratio node, a facial harmony node and an east-west aesthetic difference node, retrieving a standard node set closest to the current facial features through a query statement, calculating an aesthetic deviation vector composed of the deviation degrees of the current deformation depth values in the eye, nose and cheekbone dimensions according to the aesthetic parameter benchmark values in the standard node set, performing risk level evaluation on the aesthetic deviation vector by threshold determination method, marking a high-risk area, extracting a deviation feature combination as a potential aesthetic risk point, and traversing the correction strategy nodes in the knowledge graph that match the risk point types, reading deformation depth correction coefficients in the recommended adjustment parameters, and identifying the specific risk point position and risk degree in the current deformation scheme.
4. The multi-modal avatar generation monitoring and evaluation validation method of claim 1, wherein, The locating and identifying of the corresponding historical cases based on the aesthetic risk point and the filtering out of the relevant reference cases according to the similarity between the current user features and the historical successful cases comprise: Based on the coordinates of the aesthetic risk points, records with the same risk point type are retrieved from the historical case database. Body shape feature vectors, stylization processing parameter sequences, user rating values, and usage scenario tags are obtained to construct a historical case feature matrix. Cosine similarity is used to calculate the similarity between the current user's body shape feature vector and the historical case feature matrix, marking highly similar cases. Corresponding data pairs of deformation depth parameters and satisfaction scores are extracted from highly similar cases. Historical cases are filtered based on the corresponding data pairs, and the aesthetic preference distribution data of the filtered cases under the user's selected scenario are statistically analyzed. Relevant reference cases that meet the difference threshold and satisfaction standard are output.
5. The multi-modal avatar generation monitoring and evaluation validation method of claim 1, wherein, The output blueprint document includes: By combining the location of aesthetic risk points and relevant reference cases, the rule matching process in the aesthetic knowledge graph is initiated. Aesthetic rule nodes corresponding to the current user's body shape and contour dimensions are extracted, and the deviation between the current deformation depth value and the standard value is calculated to obtain a body shape fit score. User skin color gamut data is extracted, and style preference records for the corresponding skin color category are queried in the aesthetic knowledge graph. The matching degree between the visual effect parameters generated by the current deformation depth value and skin color features is calculated to obtain a style fit index. User hair texture density parameters are read, and the standard requirements for hairstyle processing under different usage occasions are matched. The conformity of hair texture parameters with scene standards is calculated to obtain an occasion fit value. Combining the body shape fit score, style fit index, and occasion fit value, a multi-dimensional fit assessment result is generated through weighted calculation, constructing an image blueprint document.
6. The multi-modal avatar generation monitoring and evaluation validation method of claim 1, wherein, The process involves transferring the original geometric dataset to the stylized rendering stage for stylized rendering, then deploying it to the application scenario. The rendering stability and attribute fusion effects for each deformation depth range are recorded as online feedback data, including: The original geometric dataset is input into the VToonify renderer, and a weighted fusion of feature vectors and style templates is performed to obtain a two-dimensional stylized facial image. The two-dimensional stylized facial image is used as a texture map, and three-dimensional point cloud data is constructed by combining deformation depth values. A three-dimensional stereoscopic image is reconstructed using a 3D Gaussian splash distribution Gaussian kernel. Adaptation processing is performed based on the three-dimensional stereoscopic image, and the adapted virtual image data is output. The rendering performance of the virtual image data on the target platform is monitored, and the performance indicators and fusion effect data are stored as online feedback data.
7. The multi-modal avatar generation monitoring and evaluation validation method of claim 1, wherein, The output virtual avatar generation quality assessment results include: User satisfaction ratings and return rate statistics are extracted from offline delivery records. Online feedback data is read to construct a comprehensive dataset containing stylistic intensity, recognizability, and satisfaction. The comprehensive dataset is divided into stylistic intensity intervals, and the mean and variance of recognizability values within each interval are calculated to mark candidate optimal intervals. Based on the mean recognizability and mean satisfaction values of the candidate optimal intervals, a quality score is obtained by weighted summation. The interval with the highest quality score is selected as the optimal stylistic intensity interval, and the virtual avatar generation quality assessment result is output.
8. The multi-modal avatar generation monitoring and evaluation validation method of claim 1, wherein, The method further includes: feeding back the quality assessment results of virtual image generation to the aesthetic knowledge graph, updating the body shape, style, and occasion matching records in the deformation depth adaptability case library, and realizing continuous optimization of multimodal virtual image generation assessment and verification.
9. The multi-modal avatar generation monitoring and evaluation validation method of claim 8, wherein, The process of generating quality assessment results through virtual avatars and feeding them back to the aesthetic knowledge graph to update the body shape, style, and occasion matching records in the deformation depth adaptability case library includes: The virtual avatar generation quality assessment results are transmitted to the update interface of the aesthetic knowledge graph. New records containing user body shape characteristics, style type, deformation depth value, and satisfaction score are added to the aesthetic rule nodes. Based on the new records, the deformation depth adaptability case library is accessed to calculate the deviation and update the case set and successful matching statistics of the corresponding body shape category. The distribution of stylization intensity values under each body shape category is statistically analyzed through the updated case library to determine the recommendation intensity and retention threshold range as stable parameters for output.