A recognition analysis method and system based on facial cosmetic results
By separating static and dynamic features and combining neural networks and 3D models to perform multi-dimensional recognition and analysis of facial cosmetic surgery results, the rigidity and lack of adaptability of existing methods are solved, and a more accurate and scientific evaluation of cosmetic surgery effects is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2026-03-24
AI Technical Summary
Existing methods for identifying and analyzing facial cosmetic surgery results are too rigid and cannot adapt to the diverse changes in facial expressions, resulting in insufficient accuracy and adaptability.
By collecting facial images before and after cosmetic surgery, static and dynamic features are separated, and a pre-trained neural network is used for feature analysis. Static and dynamic image sets are constructed, and multi-dimensional recognition analysis is performed by combining a three-dimensional structural model. The parameters of motion continuity and motion amplitude are introduced for feedback correction, and a comprehensive analysis result is formed.
It achieves multi-dimensional recognition and analysis of facial cosmetic surgery effects, improving the accuracy and scientific nature of the recognition, and possesses robustness and high computational efficiency, enabling accurate evaluation of cosmetic surgery effects under dynamic conditions.
Smart Images

Figure CN120375443B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of facial recognition analysis, in particular to a recognition analysis method and system based on facial cosmetic results. BACKGROUND
[0002] In the fields of face recognition, image processing and biometric analysis, automatic analysis technology of facial images has been widely applied in recent years.
[0003] With the development of computer vision and deep learning algorithms, identity recognition, emotion recognition, expression recognition and other tasks based on facial images have shown significant value in security, medical, beauty and social industries. In particular, in the field of face recognition, facial feature extraction, key point positioning and expression change analysis technologies have matured, laying a solid foundation for image intelligent analysis.
[0004] However, when facial images are subject to external interference (such as expression movements, angle changes, light changes) or internal changes (such as natural aging, surgical adjustments, etc.), traditional recognition algorithms still face challenges in stability and adaptability.
[0005] In addition, how to consider individual differences and dynamic expression factors in image analysis process to ensure recognition accuracy and evaluation credibility has also become an important issue in the industry. Therefore, the research on comprehensive recognition analysis method of facial images under multi-dimensional features is particularly crucial. SUMMARY
[0006] In view of the above problems, the present application is proposed.
[0007] Therefore, the technical problem solved by the present application is that the existing recognition analysis method of facial cosmetic results has the problems of too stereotyped recognition results and inability to adapt to diversified facial expressions.
[0008] To solve the above technical problems, the present application provides the following technical scheme: a recognition analysis method based on facial cosmetic results, comprising:
[0009] Collecting facial images of a target object before and after cosmetic surgery; the facial images include, under different actions, obtaining the feature markers of each image through feature analysis;
[0010] Dividing the feature markers into static features and dynamic features;
[0011] Respectively constructing an image set A under the static features and an image set B under the dynamic features;
[0012] According to the image set A, analyzing the completion degree of the expected effect and the preservation degree of the individual characteristics to obtain the recognition results of the completion degree and the preservation degree;
[0013] By analyzing the features of image set B under each dynamic feature, the evaluation parameters of each image are obtained. The evaluation parameters are then used to perform multi-dimensional action evaluation to obtain multi-dimensional action recognition results.
[0014] The multi-dimensional action recognition results are fed back into the completion recognition results to obtain the analysis results after the completion adjustment under the adaptive dynamic features.
[0015] As a preferred embodiment of the facial plastic surgery result recognition and analysis method described in this invention, the feature analysis includes: pre-setting L action categories, and using a pre-trained neural network to take the collected continuous frame images as the input of the neural network;
[0016] The output contains information pairs including the action category and confidence level.
[0017] As a preferred embodiment of the facial plastic surgery result recognition and analysis method of the present invention, wherein: the feature label is: for each image, the information pair output by the neural network;
[0018] The static features include information pairs whose action category is no facial expression action; the dynamic features include information pairs whose action category is facial expression action.
[0019] Let u be the index of the action category, representing a positive integer with a maximum value of L; in image set A, u = L; in image set B, u ∈ (0, L-1]; the image set A contains subsets: the image subset A+ before plastic surgery and the image subset A- after plastic surgery;
[0020] The image set B contains subsets: image subsets under each action category, B = {B1, B2, ..., B}. L-1}; where B u This represents a subset of the u-th action category;
[0021] In each subset, the image with the highest confidence level is used as the seed image; let f u The seed image for the u-th action category;
[0022] In the sequence of acquired consecutive frames, each seed image f is labeled. u ;
[0023] The growth process of the seed image is as follows:
[0024] From f u Initially, the frame sequence is grown frame by frame to both sides. During the growth process, each frame image is judged. If the action category of the image matches f... uthe action category of the image is same as the u-th action category, the growth is confirmed; if the action category of the image is different from the u-th action category, the growth is terminated; u the action category of the image is same as the u-th action category, the growth is confirmed; if the action category of the image is different from the u-th action category, the growth is terminated;
[0025] the seed image f u After the growth is completed, the length of the seed image growth is evaluated; if the total length of the growth to both sides is less than the preset frame number, the seed image of the u-th action category is reselected, the original seed image is marked as unselectable, the image with the highest confidence is taken as the seed image, and the growth is reperformed until the preset frame number is met;
[0026] After the growth is completed, each seed image and the continuous frame image grown to is obtained as a to-be-analyzed image sequence set of each action category; the to-be-analyzed image sequence set of the u-th action category is represented as: AB u ={A0,B u0};
[0027] Wherein, if the u-th action category is: no expression action occurs, B u0 is an empty set and A0=A0+,A0-}; if the u-th action category is: expression action occurs, A0 is an empty set; A0 represents the to-be-analyzed image sequence set of the static feature, including the to-be-analyzed image sequence set A0+ of the static feature before the cosmetic and the to-be-analyzed image sequence set A0- of the static feature after the cosmetic; B u0 represents the to-be-analyzed image sequence set of the u-th action category under the dynamic feature;
[0028] When the u-th action category has no seed image with a total length meeting the preset frame number, the seed image with the longest total length is selected to construct the to-be-analyzed image sequence set under the u-th action category.
[0029] As a preferred scheme of the recognition analysis method based on the facial cosmetic result, the completion degree includes: obtaining a three-dimensional structure model of the face of the target object, marking a cosmetic region, performing two-dimensional plane projection conversion on the three-dimensional structure model to obtain a two-dimensional image containing the cosmetic region mark;
[0030] The completion degree is recognized for each frame image in A0-.
[0031] Wherein, n represents the number of images in A0-; k represents the image index in A0-; C total represents the completion degree of the cosmetic region after the cosmetic; max(C k ) represents the maximum value of the similarity between the cosmetic region of the k-th image in A0- and the marked region in the two-dimensional image under all projection angles;
[0032] The preservation degree includes, for each frame image in A0+, the calculation of the completion degree, to obtain C total+ represents the completion degree of the cosmetic region before cosmetic completion; C total is the completion degree of the cosmetic region after cosmetic completion; and C total+ is the completion degree of the cosmetic region after cosmetic completion. total .
[0033] As a preferred scheme of the recognition analysis method based on the facial cosmetic result, wherein: the multi-dimensional action evaluation includes but is not limited to, in B u0 , the evaluation of the continuity and action amplitude of the image sequence to be analyzed, to obtain a set of evaluation parameters (G u , Q u ); after weighted average of the evaluation parameters (G u , Q u ), the comprehensive evaluation parameters (G z , Q Z ) are obtained.
[0034] Wherein, G u represents the continuity parameter of B u0 , Q u represents the action amplitude parameter of B u0 , R represents the number of images in B u0 ; r represents the image index in B u0 ; S r,r+1 represents the similarity between the rth image and the r+1th image in the sequence B u0 ; S r,h represents the similarity between the rth image and the hth image in the sequence B u0 ; min represents the minimum value; Su0 is the minimum similarity threshold of the uth action category, which represents the maximum action amplitude.
[0035] As a preferred scheme of the recognition analysis method based on the facial cosmetic result, wherein: the analysis result after the completion degree adjustment is: C ’ total =G z ×C total ÷Q Z .
[0036] As a preferred scheme of the recognition analysis method based on the facial cosmetic result, wherein: after the completion recognition, all results are divided into three parts:
[0037] Part one: the recognition results of the completion degree and the preservation degree; part two: the multi-dimensional action recognition result; part three: the analysis result after the completion degree adjustment.
[0038] A facial plastic surgery result recognition and analysis system employing any of the methods described in this invention, characterized in that:
[0039] The acquisition unit acquires facial images of the target subject before and after cosmetic surgery.
[0040] The recognition unit analyzes the completion rate of the expected effect and the preservation rate of its own characteristics based on the image set under static features before and after the cosmetic surgery, and obtains the recognition results of the completion rate and the preservation rate.
[0041] The analysis unit obtains evaluation parameters for each image by performing feature analysis on each dynamic feature, and uses the evaluation parameters to perform multi-dimensional action evaluation to obtain multi-dimensional action recognition results.
[0042] The adjustment unit feeds back the multi-dimensional action recognition results to the completion recognition results, thereby obtaining the analysis results of the completion adjusted under the dynamic features.
[0043] A computer device includes: a memory and a processor; the memory stores a computer program, wherein: when the processor executes the computer program, it implements the steps of the method described in any one of the present invention.
[0044] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of the present invention.
[0045] The beneficial effects of this invention: The facial cosmetic surgery result recognition and analysis method provided by this invention integrates static and dynamic features to perform multi-dimensional recognition and analysis of cosmetic surgery results, achieving coordinated evaluation at the structural and functional levels. It employs neural networks for action recognition, comparing the similarity of each frame of image with the expected 3D model from different angles to obtain completion and preservation results. Simultaneously, it introduces action coherence and action amplitude parameters to construct a dynamic evaluation mechanism and provides feedback correction to the static completion results. Ultimately, it forms a comprehensive analysis result including cosmetic surgery completion, individual feature preservation, and dynamic performance adaptability, significantly improving the accuracy, scientific rigor, and clinical applicability of cosmetic surgery result recognition. The system has the advantages of strong robustness, high computational efficiency, and comprehensive evaluation dimensions. Attached Figure Description
[0046] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0047] Figure 1 The first embodiment of the present invention provides an overall flowchart of a method for identifying and analyzing facial cosmetic surgery results. Detailed Implementation
[0048] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0049] Example 1, referring to Figure 1 As an embodiment of the present invention, a method for identification and analysis based on facial cosmetic surgery results is provided, comprising:
[0050] S1: Collect facial images of the target object before and after cosmetic surgery; the facial images include feature markers for each image obtained through feature analysis under different actions.
[0051] Furthermore, image acquisition devices with high-definition imaging capabilities, such as digital cameras, industrial cameras, structured light scanners, or depth cameras (such as Kinect, RealSense, etc.), are employed to acquire high-quality facial images. In this embodiment, images are acquired from seven standard angles: frontal, 45-degree side views, 90-degree side views, upward views, and downward views, to comprehensively reconstruct the three-dimensional structural features of the face as comprehensively as possible. In other feasible solutions, acquisition from other angles can be used.
[0052] Basic images are captured in a static expression state, while the subject is guided to sequentially perform preset facial expressions (such as smiling, closing eyes, opening mouth, raising eyebrows, etc.), and consecutive frames are captured in each expression state to construct a dynamic feature image sequence. Image acquisition should be carried out under conditions of uniform lighting and a simple background, avoiding interference factors such as shadows and reflections, to ensure image clarity and complete features. Pre-operative image acquisition should be completed within one week before the surgery, and post-operative image acquisition should be carried out after the surgical recovery period (such as more than one month) to ensure facial structure stability.
[0053] The feature analysis includes pre-setting L action categories, using a pre-trained neural network, and taking the acquired consecutive frame images as input to the neural network; the output contains information pairs containing action categories and confidence scores. The neural network classifies the "action category" based on the confidence score for each action. Therefore, the output will inevitably generate both confidence scores and classification results; this solution outputs both as a single information combination.
[0054] In this embodiment, the neural network structure used is either I3D, TCN, or a hybrid CNN-LSTM structure. This is because the solution requires not just a standard image classification network, but a neural network structure with temporal modeling capabilities. The classification of consecutive frames needs to be considered. I3D is suitable for action recognition tasks, with input being a sequence of consecutive frame images. It extends a 2D convolutional network (such as Inception-v1) to a 3D convolution, allowing the convolutional kernels to extract features in the temporal dimension. It effectively captures the differences between consecutive image frames. The output is the classification result of the action category plus the confidence score for each category (softmax output). The CNN-LSTM network structure involves using a CNN (such as ResNet) to extract spatial features from each frame. The extracted frame feature sequence is then input into an LSTM (Long Short-Term Memory) network for temporal modeling. LSTM can model the temporal context, learning the relationship between the action in the current frame and the preceding and following frames. The output is also a category plus the corresponding confidence score. It is suitable for video or frame sequence inputs, has a simple and flexible structure, and is suitable for a small number of categories of cosmetic surgery actions.
[0055] TCN, on the other hand, uses recurrent neural networks and one-dimensional temporal convolution to extract temporal dependencies in frame sequences; it supports faster parallel training and has stable performance; it is often used for action recognition and facial expression recognition; and it can also output category + confidence.
[0056] The feature label is defined as: for each image, the information pair output by the neural network.
[0057] The features are categorized into static features and dynamic features. Static features include information pairs whose action category is no facial expression; dynamic features include information pairs whose action category is facial expression.
[0058] For example, when the neural network identifies "no movement" or "static" (e.g., a calm face, no obvious muscle changes), the action category in the corresponding output information pair is "no facial expression." When the neural network identifies actions such as laughing, frowning, blinking, or opening the mouth, the action category in the corresponding output information pair is the corresponding action, categorized as "facial expression present."
[0059] It's important to understand that in static images, facial expressions remain unchanged, muscle states are stable, and contours are clear, making it easy to extract structural features such as key facial points and regional outlines. Therefore, static features are primarily used to analyze whether the face achieves the expected morphological effect after cosmetic surgery in a static state, such as nose height, chin contour, and facial symmetry, thereby assessing the completion and retention of the cosmetic surgery. Additionally, facial expressions (such as smiling, frowning, and blinking) cause muscle traction, skin stretching, and deformation of facial tissues, serving as important indicators for evaluating the naturalness and harmony of the cosmetic surgery effect. Dynamic feature images can reflect the performance of the surgical area during actual use, such as whether it is stiff, asymmetrical, or has facial expression breaks. Therefore, dynamic features are used for multi-dimensional action evaluation, such as analysis of continuity, movement amplitude, and stability. Without classification, mixing static and dynamic images will cause the analysis results to be interfered with by facial expression factors, leading to feature extraction confusion. For example, movement amplitude may be misjudged as structural deformation, and static contours may be distorted due to movement, resulting in inaccurate analysis. Therefore, it is necessary to separate the two types of feature images and model and process them separately to achieve stable and accurate recognition and analysis.
[0060] S2: Construct image set A under the static features and image set B under the dynamic features respectively.
[0061] Let u be the index of the action category, representing a positive integer with a maximum value of L; in image set A, u = L; in image set B, u ∈ (0, L-1).
[0062] The image set A contains a subset: a subset A+ of images before cosmetic surgery and a subset A- of images after cosmetic surgery.
[0063] The image set B contains subsets: image subsets B = {B1, B2, ..., B} under each action category. L-1}
[0064] Among them, B u Let B represent the subset of the u-th action category, where u ∈ (0, L-1) in the image set B (where one of the L categories is "static features where no action occurs"); if the subset of action categories B u empty set This indicates that the action category cannot be identified and is marked as... This indicates an empty set. The presence of an empty set signifies a lack of expression, prompting alert and a separate warning. At that time, the “multi-dimensional action evaluation” in the following text will not measure the empty set part.
[0065] Furthermore, within each subset, the image with the highest confidence level is selected as the seed image; let f u The seed image for the u-th action category.
[0066] In the sequence of acquired consecutive frames, each seed image f is labeled. u .
[0067] The growth process of the seed image is as follows:
[0068] From f u Initially, the frame sequence is grown frame by frame to both sides. During the growth process, each frame image is judged. If the action category of the image matches f... u If the action category is the same as f, then growth is confirmed; if the action category of the image is the same as f u If the action categories are different, growth will stop.
[0069] Seed image f u After growth is complete, the length of the seed image growth is evaluated. If the total length grown to both sides is less than the preset number of frames, the seed image for the u-th action category is reselected. The original seed image is marked as unselectable, and the image with the highest confidence is used as the seed image. Growth is repeated until the preset number of frames is met. If the number of frames grown to both sides of the seed image is too small, it may only cover the beginning or end of the action, failing to fully reflect the complete change process of the action, leading to distorted analysis results. Therefore, by setting a minimum growth length, it is ensured that the extracted image sequence can cover the main stages of the entire action process. By dynamically reselecting the new image with the highest confidence as the seed and restricting the original seed from being selected again, automatic iteration and optimization of seed image selection are achieved, ultimately obtaining the most representative and continuous image sequence under the action category.
[0070] After growth is complete, each seed image and its corresponding consecutive frame images are acquired, serving as the set of image sequences to be analyzed for each action category; the set of image sequences to be analyzed for the u-th action category is denoted as: AB u ={A0, B u0}
[0071] If the u-th action category is: no facial expression action occurred, then B u0 A0 is an empty set and A0 = A0+, A0-; if the u-th action category is: facial expression action, then A0 is an empty set; A0 represents the set of image sequences to be analyzed for the static features, including the set of image sequences to be analyzed for the static features before plastic surgery A0+ and the set of image sequences to be analyzed for the static features after plastic surgery A0-; B u0 This represents the set of image sequences to be analyzed for the u-th action category under the dynamic features described (the u-th action category includes one static category, so this index can only be selected up to L-1 positions, and the L-th action category is "no facial expression action").
[0072] It's important to note that there's only one A0, so in the design process, A0 is generally not set to a specific action category, but rather directly set to the Lth action. This won't affect the sequence order of u in B (B). u0 There are multiple sets, depending on the different u.
[0073] If there is no seed image with a total length that meets the preset number of frames for the u-th action category, then the seed image with the longest total length is selected to construct the set of image sequences to be analyzed under the u-th action category.
[0074] It's important to note that when performing recognition and analysis for each action category, it's necessary to extract image segments representing that action category from a continuous image sequence. Since the entire video or image sequence contains a large number of redundant frames, processing them without filtering will increase the system's computational burden, and some frames may contain noise or atypical actions, affecting recognition accuracy. This solution, by selecting a seed, can find the most representative continuous frame segment for each expression, thus reducing computational load while maintaining accuracy.
[0075] Seed images are often located in the "high-expression region" of a particular action category, meaning the frames with the clearest and most typical expressions. Growing from these seeds accurately captures the most expressive time segment of the action, avoiding misjudgments caused by using the entire video or low-quality frames. Instead of analyzing the entire image sequence, it focuses only on the small segment of high-quality continuous frames grown from the seed image, significantly reducing the computational burden of feature extraction, similarity calculation, and coherence evaluation, thus improving overall processing efficiency. Since the image with the highest confidence level is more clearly categorized and has clear boundaries, the image sequence constructed from it is more stable and representative in terms of feature coherence and action amplitude, contributing to improved accuracy and reliability of subsequent analyses (such as completion adjustment).
[0076] S3: Based on the image set A, analyze the completion degree of the expected effect and the preservation degree of its own characteristics respectively to obtain the recognition results of the completion degree and the preservation degree.
[0077] The completion process includes acquiring a three-dimensional structural model of the target object's face, marking the cosmetic surgery area, and performing a two-dimensional plane projection transformation on the three-dimensional structural model to obtain a two-dimensional image containing the cosmetic surgery area markings.
[0078] It's worth noting that during the actual cosmetic surgery plan design phase, doctors or designers typically use professional 3D modeling software (such as FaceGen, 3D-Me, ZBrush, or medical-grade simulation systems) to pre-determine the client's post-surgery facial appearance. This 3D modeling process is based on the target subject's facial scan data, combined with the cosmetic surgery intention, and generates a 3D model of the changed face through parameter adjustments. This model is not only used for pre-operative simulation and client preview but also serves as the source of the "target expected morphology" in this invention.
[0079] Therefore, a three-dimensional structural model is established before the cosmetic surgery, possessing complete facial structural feature data, serving as the target model for subsequent completion comparison analysis. During the modeling process, a professional doctor or designer selects the areas involved in the cosmetic surgery on the surface of the 3D model, labeling them as a set of region tags on the facial mesh (e.g., nose region, jaw region, cheek regions, etc.). Alternatively, automatic labeling can be used: facial analysis tools (such as dlib, FaceMesh, OpenFace) are used to automatically detect and divide key facial regions, and then the system settings automatically select the areas involved in the cosmetic surgery as "marked regions."
[0080] The purpose of 2D projection conversion is to project the 3D structure onto a standard 2D image for matching and comparison with the actual captured image. A standard viewpoint (e.g., frontal, 45° left, 45° right, etc.) is selected, and the virtual camera's pose matrix (rotation + translation) is set. Using a perspective projection model or orthographic projection model, the 3D mesh points (X,Y,Z) are converted into 2D image coordinates (x,y). The marker points of the cosmetic surgery area are simultaneously projected onto the 2D image, forming a 2D image "containing cosmetic surgery area markers," which can be used for subsequent feature comparison with the corresponding areas in the actual image.
[0081] For each frame of image in A0-, perform completion assessment;
[0082] Where n represents the number of images in A0-; k represents the image index in A0-; C total This indicates the degree of completion of the cosmetic surgery area; max(C) k Let ) represent the maximum similarity between the cosmetic region of the k-th image in A0 and the marked region in the two-dimensional image, across all projection angles.
[0083] It's important to note that each frame of the actual image is taken from a different angle. Directly comparing it to a target image at a fixed angle can easily lead to errors. Therefore, the design selects the angle that best matches the current frame from multiple projection angles for comparison, making it more adaptable and objective. Dynamically selecting the projection image with the optimal angle as a reference helps identify the visual state that truly expresses the cosmetic surgery effect, avoiding low scores due to angle errors. Each frame is scored independently, which can be used for subsequent statistical analysis, trend judgment, or average completion assessment, enabling fine-grained dynamic cosmetic surgery analysis.
[0084] The similarity can be calculated using SSIM or cosine similarity based on deep features (such as FaceNet and ArcFace) to extract feature vectors of the cosmetic surgery area.
[0085] The preservation degree includes calculating the completion degree for each frame of image in A0+, to obtain C. total+ This indicates the degree of completion of the cosmetic surgery area before the surgery is completed (calculated using the same method as C). total Similar, only the parameters need to be adaptively replaced; C total With C total+ By subtracting the values, we obtain the retention rate P. total Traditional methods only compare the differences before and after surgery, but this comparison cannot achieve a comprehensive effect. The "completeness calculation" of images acquired before and after surgery actually analyzes the actual similarity between the facial region before surgery and the target object. The calculations described above provide a more comprehensive analysis. Subtracting the two "completeness" values before and after surgery yields the most comprehensive result in "feature preservation." Analyzing only different frames before and after surgery may lead to biases due to angle and other factors. Therefore, using the target object as a transition provides higher robustness. In other words, it's about "either the surgically altered or un-surgically altered"—this method allows for the fastest and most effective measurement of "preservation rate."
[0086] S4: By analyzing the features of image set B under each dynamic feature, the evaluation parameters of each image are obtained. The evaluation parameters are then used to perform multi-dimensional action evaluation to obtain multi-dimensional action recognition results.
[0087] The multi-dimensional action assessment includes, but is not limited to, in B u0 In this process, the coherence and motion amplitude of the image sequence to be analyzed are evaluated, resulting in a set of evaluation parameters (G). u Q u ).
[0088] In other alternative embodiments, the overall adaptability of the face can be analyzed using neural networks to determine whether there is any unnaturalness in the face after cosmetic surgery. Alternatively, the degree of organ distortion can be analyzed to determine whether the face exhibits excessively exaggerated movements during actions.
[0089] Among them, G u Indicates the coherence parameter, Q u This represents the amplitude of movement parameter. The larger the value, the smaller the range of movement that can be performed, indicating a greater constraint on the face after cosmetic surgery. Since the evaluation sequence is obtained through seed node selection and represents the most characteristic movement, the actual amplitude of movement should be the maximum amplitude under that expression (which can also be understood as the degree of freedom of facial movement). R represents B u0 The number of images in B; r represents the number of images in B. u0 Image index in S; r,r+1 B u0 The similarity between the r-th image and the (r+1)-th image in the sequence; S r,h B u0 The similarity between the r-th image and the h-th image in the sequence; min represents the minimum value; Su0 is the minimum similarity threshold for the u-th action category, and represents the maximum action amplitude.
[0090] For the evaluation parameter (G) u Q u After weighted averaging, the comprehensive evaluation parameter (G) is obtained. z Q Z ).
[0091] It is important to know that the weighted average process is achieved by dividing each motion continuity parameter (or motion amplitude parameter) by its standard value.
[0092]
[0093] Among them, G u0 B u0 The standard value of the coherence parameter, Q u0 B u0 The standard value of the motion amplitude parameter. L-1 represents the number of seeds under dynamic characteristics.
[0094] S5: Feed the multi-dimensional action recognition results back to the completion recognition results to obtain the analysis results after the completion adjustment under the adaptive dynamic features.
[0095] The analysis result after completion adjustment is: C ’ total =Gz ×C total ÷Q Z .
[0096] In actual facial cosmetic surgery effect evaluation, relying solely on completion recognition based on static images is insufficient to fully reflect the performance of the surgery in real-world use. Because the face is in frequent motion in daily life (such as smiling, frowning, and speaking), the dynamic performance of the surgical area (whether it appears natural, symmetrical, and stable) directly impacts the overall evaluation of the surgery's effect. By using dynamic features such as continuity parameters and movement amplitude parameters to correct static completion, the final analysis results retain both structural evaluation and functional performance, possessing higher reliability and practical value. The adjusted completion results can help doctors or designers identify potential problem areas under dynamic conditions.
[0097] After the recognition is completed, all results are displayed in three parts:
[0098] Part 1: The recognition results of the completion rate and the preservation rate; Part 2: The multi-dimensional action recognition results; Part 3: The analysis results after adjusting the completion rate.
[0099] On the other hand, this embodiment also provides a recognition and analysis system based on facial cosmetic surgery results, which includes:
[0100] The acquisition unit acquires facial images of the target subject before and after cosmetic surgery.
[0101] The recognition unit analyzes the degree of completion of the expected effect and the degree of preservation of its own characteristics based on the image set under static features before and after the cosmetic surgery, and obtains the recognition results of the degree of completion and the degree of preservation.
[0102] The analysis unit obtains evaluation parameters for each image by performing feature analysis on each dynamic feature, and uses the evaluation parameters to perform multi-dimensional action evaluation to obtain multi-dimensional action recognition results.
[0103] The adjustment unit feeds back the multi-dimensional action recognition results to the completion recognition results, thereby obtaining the analysis results of the completion adjusted under the dynamic features.
[0104] If the above functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0105] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0106] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0107] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0108] Example 2 is an embodiment of the present invention, which provides a method for identification and analysis based on facial plastic surgery results. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.
[0109] Six volunteers (numbered P1 to P6) who had undergone facial cosmetic surgery were invited to participate in the test. The procedures covered common areas, including rhinoplasty, jaw reduction, double eyelid surgery, and facial fillers. All surgeries were completed and the recovery period exceeded one month, with stable facial conditions, meeting the requirements for objective identification and subjective evaluation.
[0110] Personalized 3D expected models were generated using the FaceGen modeling tool, and the cosmetic surgery target was confirmed using a standard pre-operative styling process. Image data before and after the procedure was collected using an Intel RealSense D435i depth camera, capturing images from seven perspectives: front, 45-degree left / right, 90-degree side view, upward view, and downward view. Simultaneously, continuous frame sequences were recorded for five types of dynamic expressions: smiling, open mouth, raised eyebrows, frowning, and closed eyes. Each expression lasted 1.5 seconds at a frame rate of 30fps, for a total of 1050 frames per object.
[0111] Facial feature vectors are extracted from each frame of the image using the ArcFace neural network. An I3D model is then used for action classification, outputting the expression category and corresponding confidence score. Images are divided into static and dynamic feature sets according to preset rules, creating image set A (static) and image set B (dynamic). For each action category, seed images are selected based on confidence scores, and an image sequence is automatically generated using action coherence and amplitude determination mechanisms to construct the subset of images to be analyzed.
[0112] The completion score is calculated by extracting features from the FaceNet embedding and comparing them using cosine similarity based on the two-dimensional image projected onto the target model. The dynamic feature index evaluates the coherence and naturalness of the action based on the mean and range of cosine similarity between adjacent frames. In the final stage, the dynamic results are fed back into the completion score to form a "completion-adjusted index" to measure the adaptability and expressiveness of the actual cosmetic surgery effect in dynamic real-world scenarios.
[0113] During the experiment's completion phase, cosmetic surgeons were organized to fill out a "Recognition Result Satisfaction Scale" to subjectively evaluate whether the provided recognition analysis results truly reflected the cosmetic surgery effect and dynamic performance. A 5-point scale was used (1 point for extreme dissatisfaction and 5 points for extreme satisfaction). Some data were recorded as shown in Table 1.
[0114] Table 1 Data Record Table
[0115] Parameter name / number P1 P2 P3 P4 P5 P6 Static completion (0-1) 0.86 0.83 0.91 0.84 0.88 0.92 Action continuity parameter (0-1) 0.91 0.89 0.95 0.88 0.90 0.96 Action amplitude parameter (the larger the value, the more limited) 1.30 2.00 0.50 1.70 1.10 0.55 Score after completion adjustment (automatically calculated) 0.60 0.37 1.73 0.44 0.72 1.61 Score of saving degree (0-1) 0.81 0.77 0.85 0.78 0.82 0.88 Doctor score - structure restoration (1-5) 5 4.5 5 4.5 5 5 Doctor score - dynamic naturalness (1-5) 5 4.5 5 4.5 5 5 Doctor score - overall evaluation of credibility (1-5) 5 5 5 5 5 5
[0116] As can be seen, the average static completion score is 0.89, indicating that the system has good accuracy in multi-view structural reconstruction. The scores of P3 and P6 are as high as 0.91 and 0.92, respectively, showing that the system is very accurate in recognizing geometric changes in delicate cosmetic surgery areas such as the nasal alae and jaw.
[0117] The average dynamic coherence is 0.915, reflecting the system's high stability in seed image growth and continuous frame action sequence extraction. Combined with the multi-dimensional comparison structure of this invention, this parameter successfully captures the temporal continuity in natural actions, ensuring that subsequent analysis is based on natural processes.
[0118] The amplitude of movement parameter reflects the "degree of freedom" of facial movements. The system automatically calculates the amplitude of movement based on the range of changes in key points between frames. A large amplitude value (e.g., 2.00 for P2, 1.70 for P4) indicates that the range of movement is somewhat restricted after cosmetic surgery; while a small amplitude value (e.g., 0.50 for P3, 0.55 for P6) indicates a high degree of freedom of movement in that area, and cosmetic surgery has not imposed any restrictions. The system will therefore appropriately lower or increase the completion score to better reflect actual usage.
[0119] For example, the adjusted completion score for P3 was 1.73, which is much higher than the static score of 0.91, indicating that the system recognized that the expression was natural and the movement was complete, and gave positive feedback. However, due to the limited range of motion (2.00), even though the static completion score was high (0.83), the adjusted score for P2 dropped to 0.37, which reminded doctors to pay attention to the risk of "good structure but limited expression".
[0120] This dynamic feedback mechanism is a dimensional fusion capability that is difficult to achieve with traditional methods, demonstrating the technological novelty of this approach.
[0121] According to the doctors' scores, all subjects scored above 4.5 in terms of structural fidelity, dynamic naturalness, and overall credibility. The average dynamic naturalness score was 4.83, and the credibility score was 5.00. This indicates that the system's recognition results not only match the objective parameters but also closely match the doctors' clinical observations.
[0122] Of particular note is that in samples with limited amplitude, such as P2 and P4, doctors confirmed slight limitations in facial expression, and the system score reflected this difference. In samples with sufficient amplitude, such as P3 and P6, the system gave extremely high scores, and doctors also stated that the expressions were "natural and not stiff." This dual consistency between system recognition and clinical validation is the most valuable part of this invention, demonstrating that this method not only has data-driven accuracy but also interpretability for real-world medical problems.
[0123] In summary, this invention, through a dynamic adjustment mechanism, identifies the degree of limitation of different cosmetic surgery areas on the dynamic expression ability of the face. It does not rely on a single judgment of static structure and constructs an evaluation system that is highly adapted to the actual cosmetic surgery results, demonstrating significant creativity and practicality.
[0124] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for identifying and analyzing facial cosmetic surgery results, characterized in that, include: Facial images of the target subject were collected before and after cosmetic surgery. The facial images include feature markers for each image obtained through feature analysis under different actions; The feature labels are divided into static features and dynamic features; Construct image set A under the static features and image set B under the dynamic features respectively; Based on the image set A, the completion rate of the expected effect and the preservation rate of its own characteristics are analyzed respectively to obtain the recognition results of the completion rate and the preservation rate; By analyzing the features of image set B under each dynamic feature, the evaluation parameters of each image are obtained. The evaluation parameters are then used to perform multi-dimensional action evaluation to obtain multi-dimensional action recognition results. The multi-dimensional action recognition results are fed back into the completion recognition results to obtain the analysis results after completion adjustment under the adaptive dynamic features; The completion includes obtaining a three-dimensional structural model of the target object's face, marking the cosmetic surgery area, and performing a two-dimensional plane projection transformation on the three-dimensional structural model to obtain a two-dimensional image containing the cosmetic surgery area markings. Let the set of image sequences to be analyzed, representing the static features described before cosmetic surgery, be denoted as . The set of image sequences to be analyzed for the static features after plastic surgery is as follows: ; against Completeness assessment is performed on each frame of the image. ; Where n represents The number of images in; k represents Image index in; This indicates the degree of completion of the surgically treated area. express The maximum similarity between the cosmetic surgery region in the k-th image and the marked region in the two-dimensional image across all projection angles; The preservation level includes, for For each frame of the image, the completion rate is calculated to obtain... This indicates the degree of completion of the surgical area before the procedure is finished; and Differences are calculated to obtain the retention rate. .
2. The identification and analysis method based on facial cosmetic surgery results as described in claim 1, characterized in that: The feature analysis includes: pre-setting L action categories, and using a pre-trained neural network to take the collected consecutive frame images as the input of the neural network; The output contains information pairs including the action category and confidence level.
3. The identification and analysis method based on facial cosmetic surgery results as described in claim 2, characterized in that: The feature label is defined as: for each image, the information pair output by the neural network; The static features include information pairs whose action category is no facial expression action; the dynamic features include information pairs whose action category is facial expression action. Let u be the index of the action category, representing a positive integer with a maximum value of L; in image set A, u = L; in image set B, u ∈ (0, L-1]; the image set A contains subsets: the image subset A+ before plastic surgery and the image subset A- after plastic surgery; The image set B contains subsets: image subsets under each action category. ;in, Represents a subset of the u-th action category; In each subset, the image with the highest confidence level is used as the seed image; let... The seed image for the u-th action category; In the sequence of acquired consecutive frames, each seed image is labeled. ; The growth process of the seed image is as follows: from Initially, the frame sequence is grown frame by frame to both sides. During the growth process, each frame image is judged. If the image's action category is similar to... If the action category is the same, then growth is confirmed; if the action category of the image is the same as... If the action categories are different, growth will stop; Seed image After growth is complete, the length of the seed image growth is evaluated. If the total length of growth to both sides is less than the preset number of frames, the seed image of the uth action category is reselected. The original seed image is marked as unselectable, and the image with the highest confidence is used as the seed image. Growth is repeated until the preset number of frames is met. After growth is complete, each seed image and its corresponding consecutive frame images are acquired, serving as the set of image sequences to be analyzed for each action category; the set of image sequences to be analyzed for the u-th action category is represented as: ; Where the u-th action category is: no facial expression action occurred, then It is an empty set and If the u-th action category is: facial expression action, then It is an empty set; The set of image sequences to be analyzed represents the static features; This represents the set of image sequences to be analyzed for the u-th action category under the aforementioned dynamic features; If there is no seed image with a total length that meets the preset number of frames for the u-th action category, then the seed image with the longest total length is selected to construct the set of image sequences to be analyzed under the u-th action category.
4. The identification and analysis method based on facial cosmetic surgery results as described in claim 3, characterized in that: The multi-dimensional action assessment includes, but is not limited to, in In this process, the coherence and motion amplitude of the image sequence to be analyzed are evaluated, resulting in a set of evaluation parameters. ; Evaluation parameters After weighted averaging, the comprehensive evaluation parameters are obtained. ; in, express The coherence parameter, ; express The amplitude parameters of the movement, R represents The number of images in; r represents Image index in; express The similarity between the r-th image and the (r+1)-th image in the sequence; express The similarity between the r-th image and the h-th image in the sequence; min represents the minimum value; Su0 is the minimum similarity threshold for the u-th action category, and represents the maximum action amplitude.
5. The identification and analysis method based on facial cosmetic surgery results as described in claim 4, characterized in that: The analysis results after completion adjustment are as follows: .
6. The identification and analysis method based on facial plastic surgery results as described in claim 5, characterized in that: After the recognition is completed, all results are displayed in three parts: Part 1: The recognition results of the completion rate and the preservation rate; Part 2: The multi-dimensional action recognition results; Part 3: The analysis results after adjusting the completion rate.
7. A facial plastic surgery result recognition and analysis system employing the method described in any one of claims 1-6, characterized in that: The acquisition unit acquires facial images of the target subject before and after cosmetic surgery. The recognition unit analyzes the completion rate of the expected effect and the preservation rate of its own characteristics based on the image set under static features before and after the cosmetic surgery, and obtains the recognition results of the completion rate and the preservation rate. The analysis unit obtains evaluation parameters for each image by performing feature analysis on each dynamic feature, and uses the evaluation parameters to perform multi-dimensional action evaluation to obtain multi-dimensional action recognition results. The adjustment unit feeds back the multi-dimensional action recognition results to the completion recognition results, thereby obtaining the analysis results after completion adjustment under the adaptive dynamic features.
8. A computer device, comprising: Memory and processor; The memory stores a computer program, characterized in that: when the processor executes the computer program, it implements the steps of the method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Parkinson's disease auxiliary diagnosis method and system based on facial expression static and dynamic characteristics
CN118280554A
Multi-dimensional intelligent analysis method for motion process, computer equipment and storage medium
CN118887738A