Figure scene generation system based on AI interaction
By introducing skin microstructure analysis and facial dynamic expression processing technology into the character scene generation system, combining three-dimensional image model and motion capture technology, the accuracy of user age evaluation and hairstyle and clothing matching in the existing system is solved, and more natural and realistic image generation and higher user experience are achieved.
Patent Information
- Application Number
- CN202510295615.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
AI Technical Summary
The existing character scene generation system has many challenges in user age assessment, hairstyle and clothing matching and dynamic effect simulation, especially the accuracy of age assessment is affected by lighting, shooting angles and expression changes.
By introducing skin microstructure analysis, color space conversion and histogram analysis methods, combined with the analysis and processing of dynamic facial expressions, the optical flow method is used to calculate the displacement of facial key points, remove temporary features related to expressions, and judge the matching degree of clothing and hairstyles in real time through three-dimensional image model and motion capture technology.
It improves the accuracy and generalization ability of age assessment, enhances the fit between hairstyle and clothing and users, makes the generated images more natural and realistic, and improves the user experience.
Smart Images

Figure CN120147460A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image generation, and particularly to a human scene generation system based on AI interaction. Background Art
[0002] In today's digital age, with the rapid development of artificial intelligence (AI) technology, human scene generation technology has been widely applied in multiple fields such as virtual fitting, virtual image design, and game character creation. However, there are still many challenges in the existing human scene generation systems in aspects such as user age assessment, hairstyle and clothing matching, and dynamic effect simulation.
[0003] For example, Chinese Patent Application No. 202311011914.3 discloses a human image scene generation system based on AI interaction. This human image scene generation system includes a user interaction design module, a human image shooting module, a human image analysis and processing module, a human image processing module, an image integration and correction module, an image generation module, and a management database. This system matches corresponding scene themes for users according to the keywords selected by users through the voice recognition interface, and at the same time conducts age assessment on the captured static images of users and, based on this, screens layer by layer to match the most suitable hairstyles and clothing for the user's human images, making up for the defect of low attention to age in the prior art, providing a choice that better conforms to age characteristics and aesthetic preferences, meeting the standards of personalized customization according to user needs and preferences, improving the visual effect and realism, enhancing the vividness and quality of the generated images, and making them more in line with user expectations.
[0004] In the existing patent technology solutions, age assessment methods often rely on a single photo or video, making it difficult to comprehensively capture the facial and body features of users, resulting in deviations in age assessment results. Especially when users make extreme expressions, changes in facial expressions may cause temporary changes in biological markers such as skin texture gloss and wrinkles, further reducing the accuracy of age prediction. In addition, factors such as different lighting conditions, shooting angles, and user postures will also affect the age assessment results. Summary of the Invention
[0005] This application provides a human scene generation system based on AI interaction. By improving the age prediction method, more accurate age assessment results are provided, and personalized clothing and hairstyle suggestions are provided for users according to the assessment results. By introducing methods such as skin microstructure analysis, color space conversion, and histogram analysis, features such as skin texture and pigment deposition can be more accurately quantified, thereby improving the accuracy of age assessment.
[0006] This application provides a person-scene generation system based on AI interaction, including: a user interaction design module, a person image shooting module, a person image analysis and processing module, a person image processing module, an image integration and correction module, an image generation module, and a management database; the person image shooting module is used to obtain a user image and a microscopic structure image of the skin surface by using a high-definition camera, and the microscopic structure includes a skin glossiness reference value, a wrinkle level reference value, pore characteristics, and pigment deposition. The user image contains a hair coverage rate; an algorithm is used to calculate the pore characteristics, and color conversion is performed on the microscopic structure image to obtain the pigment deposition. The user age evaluation correlation coefficient is calculated through the skin glossiness, wrinkle level, pore characteristics, pigment deposition, and hair coverage rate; the person image analysis and processing module is used to obtain user age evaluation correlation data based on the user image and the microscopic structure of the skin surface, predict the age of the user, and then obtain the user age evaluation correlation coefficient; the image integration and correction module is used to integrate the user interaction scene and the user decoration image to obtain a user scene decoration image.
[0007] Preferably, the specific steps for calculating the pore characteristics using an algorithm are as follows: Use the algorithm to detect the edges with drastic gray-scale changes in the image. The edges correspond to the contours of the pores. Segment the identified pores, extract the size characteristics of the segmented pore regions. By measuring the diameter of the pore regions and counting the number of pixels within the pore regions, the area of the pores can be obtained. Extract the shape characteristics of the segmented pore regions. Evaluate the circularity by calculating the ratio of the perimeter to the area of the pore regions. For non-circular pores, calculate the ratio of the longest diameter to the shortest diameter, that is, the aspect ratio.
[0008] Preferably, the formula for the user age evaluation correlation coefficient ρ is: , where ψ is the skin glossiness reference value, ξ is the wrinkle level reference value, is the average pore diameter, is the quantization index of the pigment deposition, λ is the hair coverage rate, is the weight coefficient of each feature for age evaluation. The coefficient is set and adjusted according to actual data or experience. e is the natural constant, is the adjustment parameter affected by the glossiness.
[0009] Preferably, the method for predicting the age of the user further includes: S101, collect a facial dataset and perform facial key point annotation based on the collected facial dataset; S102, calculate the displacement of the facial key points using the optical flow method according to the annotated facial key points; S103, construct a feature vector according to the calculated displacement of the key points. The feature vector includes a first feature vector, a second feature vector, and a third feature vector; S104. According to the characteristics of the data in the facial dataset, train a regression model. Input the obtained third feature vector into the trained regression model through the person image analysis and processing module, and the regression model outputs the correlation coefficient for optimized user age assessment.
[0010] Preferably, the displacement vectors of the facial key points calculated by the optical flow method are arranged in sequence to form a second feature vector. The dimension of the second feature vector is equal to the number of facial key points, and each dimension corresponds to the displacement vector of a key point. Extract global features from the facial image dataset, arrange the extracted global features in sequence to form a first feature vector, compare the first feature vector and the second feature vector, calculate the difference value, i.e., the expression restoration value, identify the inherent features and temporary features, calculate the variance of the temporary features according to the identified temporary features, set a feature threshold according to the calculated variance, remove the vectors in the first feature vector that exceed the feature threshold, organize the remaining features after removing the temporary features, arrange them in sequence, and combine the organized features into a third feature vector.
[0011] Preferably, the method for optimizing the correlation coefficient of user age assessment further includes: S201. Obtain the user's facial and body images through a camera, and use image processing algorithms to extract the user's facial features and body features. S202. Input the extracted facial features and body features into the age assessment algorithm to calculate the age assessment value. S203. Collect images, select candidate hairstyles and clothing from the collected hairstyle images according to the user age assessment correlation coefficient, match the candidate hairstyles and clothing with the user's facial and body images, and adjust the matched hairstyles and clothing using the age assessment value.
[0012] Preferably, use a camera to obtain the user's full-body image, use an algorithm to extract the body lines and contours when standing, and locate the key joint points of the body. Use the distance formula to calculate the distance between two joint points. For the measurement of angles, select the corresponding joint point triple. To measure the elbow bending angle, select the shoulder, elbow, and wrist joint points, calculate the vectors between adjacent joint points. For the shoulder, elbow, and wrist, calculate the vector from the shoulder to the elbow and the vector from the elbow to the wrist, and use the dot product and modulus of the vectors to calculate the angle between the two vectors. The formula is: , where is the vector between adjacent joint points.
[0013] Preferably, the steps for adjusting the age assessment value are: S301. Build a three-dimensional image model according to the collected user's facial and body images. In S302, motion capture technology is used to obtain real motion data, and the changes of the 3D image model are simulated according to the real motion data. When the 3D model performs actions, it is judged in real time whether the clothing and hairstyle match. If they do not match, the age evaluation value is adjusted. If they match, the age evaluation value remains unchanged.
[0014] Preferably, when the 3D image model performs actions, it is judged in real time whether the clothing and hairstyle match, and the judgment is made according to the swing amplitude of the clothing and the change rate of the hairstyle shape. The consistency of the clothing swing amplitude refers to the matching degree between the swing amplitude of the clothing in the 3D model and the swing amplitude of the same type of clothing in the real situation. By measuring the maximum distance of the clothing swing in the 3D model during walking and comparing it with the swing amplitude of the same type of clothing in the real video under the same action, if the swing amplitude does not exceed 5% of the set swing amplitude difference, it is considered that the clothing swing amplitude is consistent with the real situation, otherwise, it is inconsistent; the rationality of the hairstyle shape change rate refers to whether the shape change rate of the hairstyle in the 3D model during the action conforms to the expected physical behavior. By calculating the change rate of the hairstyle shape in the 3D model during walking and comparing it with the change rate of the same type of hairstyle in the real video under the same action, if the difference between the hairstyle shape change rate and the real situation does not exceed 15%, the hairstyle shape change rate is considered reasonable. If it exceeds 15%, the age evaluation value is adjusted.
[0015] Preferably, the user interaction design module includes a voice recognition unit and a scene matching unit. The voice recognition unit is used to collect the voice information of the user, convert the voice information of the user into text, and extract keywords. The scene matching unit is used to match the extracted keywords with the scene keywords in the database to screen and obtain the user interaction scene.
[0016] One or more technical solutions provided in this application have at least the following technical effects or advantages: By improving the age prediction method and introducing methods such as skin microstructure analysis, color space conversion, and histogram analysis, this solution can more accurately quantify features such as skin texture and pigment deposition, improving the accuracy of age assessment. At the same time, combined with the analysis and processing of facial dynamic expressions, the optical flow method is used to calculate the displacement of facial key points, and temporary features related to expressions are eliminated, further improving the accuracy and generalization ability of age prediction. In addition, by adjusting factors affecting age assessment, such as the position of the hairline, the position of the eye corners, and the hair color, and eliminating the influence of photographing shadows, the fit of the hairstyle and clothing to the user is improved. Finally, this solution not only judges the matching degree from the front photo, but constructs a 3D image model to comprehensively simulate the actual image of the user, combines motion capture technology to judge the matching degree of clothing and hairstyle in real time, and ensures the coordination and natural smoothness of the dynamic change effect, providing users with a more real and natural virtual image experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic structural diagram of a character scene generation system based on AI interaction according to the present invention; Figure 2 It is a schematic flowchart for predicting the age of a user in an embodiment of the present invention; Figure 3 It is a schematic flowchart for optimizing the correlation coefficient of user age evaluation in an embodiment of the present invention; Figure 4 It is a schematic flowchart for adjusting the age evaluation value in an embodiment of the present invention. Detailed implementation manners
[0018] To facilitate the understanding of the present invention, the present application will be described more comprehensively below with reference to the relevant drawings; the preferred embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein; on the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0019] It should be noted that the terms "vertical", "horizontal", "upper", "lower", "left", "right" and similar expressions used herein are for illustrative purposes only and do not represent the only embodiments.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs; the terms used in the description of the present invention in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention; the term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0021] Embodiment 1: Figure 1 It is a schematic flowchart of a character scene generation system based on AI interaction in an embodiment of the present invention, including: a user interaction design module, a character image shooting module, a character image analysis and processing module, a character image processing module, an image integration and correction module, an image generation module, and a management database.
[0022] The user interaction design module includes a speech recognition unit and a scene matching unit. The speech recognition unit is used to collect the user's voice information, convert the user's voice information into text, and extract corresponding keywords. Through the speech recognition unit, the user can interact with the system without manually inputting text, which improves the convenience and efficiency of interaction, enables the system to more accurately understand the user's voice input, and thus provides more accurate services. The scene matching unit is used to match the extracted keywords with the scene keywords in the database, and screen to obtain the user interaction scene. The scene matching unit can accurately judge the interaction scene required by the user, enabling the system to provide personalized image generation for different scenes. Through scene matching, the system can better meet the specific needs of users, provide more personalized and considerate services, and enhance user satisfaction.
[0023] The human image shooting module is used to obtain the user's image and the microscopic structure of the skin surface by using a high-definition camera. The microscopic structure includes the skin glossiness reference value, the wrinkle level reference value, the shape and size of pores, and the distribution or color of pigment spots. The user image contains the hair coverage rate.
[0024] The human image analysis and processing module is used to obtain user age assessment correlation data based on the user image and the microscopic structure of the skin surface, predict the user's age, and then obtain the user age assessment correlation coefficient. Furthermore, the captured skin image is preprocessed by the image processing technology Adobe Photoshop. The preprocessing includes filtering and enhancing contrast. Filtering can remove the noise and interference in the image and make the image clearer. Enhancing contrast can highlight details such as pores and pigment spots, making them easier to observe and analyze. The preprocessed image is converted into a visual graph using the visualization software MATLAB to make the skin features more intuitive and easy to understand. Different skin features are assigned different colors to distinguish them more clearly. For example, pores can be represented by one color, and pigment spots can be represented by another color. Using 3D reconstruction technology, the two-dimensional skin image is converted into a three-dimensional solid graph, and the observer can observe and analyze the skin features from different angles to obtain more comprehensive information. Different skin features are assigned different colors to distinguish them more clearly. After obtaining the micro-structured image, the Canny edge detection algorithm is used to detect the edges with drastic gray-scale changes in the image. These edges correspond to the contours of the pores. Then, morphological processing is used to segment the identified pores, and size features of the segmented pore regions are extracted. By measuring the diameter of the pore regions and counting the number of pixels within the pore regions, the area of the pores can be obtained. Shape features of the segmented pore regions are extracted by calculating the ratio of the perimeter to the area of the pore regions to evaluate the roundness. The higher the roundness, the closer the pore shape is to a circle. For non-circular pores, the ratio of the longest diameter to the shortest diameter, i.e., the aspect ratio, is calculated; After converting the obtained micro-structured image from the RGB color space to the Lab color space, the image is decomposed into three independent channels: The L channel represents the luminance information, which reflects the light and dark degree of the image; the a channel and the b channel represent different aspects of chromaticity. Among them, the a channel is related to the opposition of red and green colors, and the b channel is related to the opposition of yellow and blue colors, which helps to reduce the influence of light changes on pigment analysis. The chromaticity (the a and b channels in the ab space) and luminance (the L channel in the Lab space) are used to generate histograms. In the histogram, the horizontal axis represents the pixel value range, and the vertical axis represents the frequency of each pixel value. By counting the number of pixels with different chromaticity or luminance values, information about the concentration and dispersion of pigments in the image can be obtained. Analyze the shape, peak position, peak width and other characteristics of the histogram to quantify the pigment deposition situation. For example, the sharpness of the peak may reflect the uniformity of pigment deposition, and the position of the peak may reveal the type or concentration of the main pigment; The specific analysis method of the skin gloss reference value includes the following steps: By segmenting and detecting the skin color of the user's skin area in the micro-structured image, the skin area and other areas are distinguished, and the skin area is divided into several skin sub-areas for RGB color detection. The red, green, and blue component values of the skin in each skin sub-area are respectively denoted as Ri, Gi, Bi, where i represents the number of the i-th skin sub-area divided, i = 1, 2,..., k. Substitute the red, green, and blue component values of the skin in each skin sub-area into the formula: to analyze and obtain the skin luminance Y of the user. k represents the number of skin sub-areas. Compare the skin luminance of the user with the preset skin luminance ranges corresponding to each glossiness in the management database, and screen out the glossiness corresponding to the skin luminance of the user, denoted as the skin gloss reference value ψ; The specific analysis method of the wrinkle level reference value is: and extract the corresponding wrinkle feature data from the face part of the micro-structured image; the wrinkle feature data includes the number of wrinkles, the wrinkle depth of each wrinkle, and the wrinkle length of each wrinkle. Denote the number of wrinkles as a, and the wrinkle depth corresponding to each wrinkle as , where j represents the number of each wrinkle, j = 1, 2,... a, and the wrinkle length corresponding to each wrinkle is denoted as , and substitute it into the formula: , to obtain the characteristic data matching coefficient between the user's skin area and wrinkles of each level , represents the number of reference wrinkles of the q-th level of wrinkles, q represents the number of each level of wrinkles, q = 1, 2,..., p, reference represents the reference wrinkle depth of the q-th level of wrinkles, represents the reference wrinkle length of the q-th level of wrinkles, η1, η2, and η3 respectively represent the set correction coefficients for the number of wrinkles, wrinkle depth, and wrinkle length, and e represents the natural constant; select the wrinkles of the level corresponding to the maximum characteristic data matching coefficient from the characteristic data matching coefficients between the user's skin area and wrinkles of each level as the reference value of the wrinkle level, denoted as ξ; The specific method for the hair coverage rate includes the following steps: separately divide the hair area of the person in the user image to obtain the hair area image of the person, read the width w_hair and height h_hair of the hair area image of the person, and convert the hair area image of the person into a grayscale image. Detect the grayscale value of each pixel point in the converted grayscale image, compare it with the grayscale value range corresponding to the set standard hair density threshold of the user, and obtain the number of pixel points that meet the range, denoted as σ. Compare the number of pixel points that meet the range with the total number of pixel points in the hair area image of the person, and substitute it into the formula: , dpi represents the pixel density of the image stored in the management database, and then analyze to obtain the hair coverage rate λ. By comparing the ratio of the number of pixel points that meet the range to the total number of pixel points, an accurate hair coverage rate is provided. The sizes of different hair area images may vary, and comparing the number of pixel points that meet the range with the total number of pixel points in the hair area can eliminate this difference, making the evaluation result more accurate and comparable; The user age evaluation correlation coefficient comprehensively analyzes the user age evaluation correlation coefficient through weight assignment of skin glossiness, wrinkle level, pore characteristics, pigment deposition, and hair coverage rate. The formula for the user age evaluation correlation coefficient ρ is: , where ψ is the reference value of skin glossiness, which reflects the brightness and firmness of the skin. Usually, the skin glossiness of young people is higher. ξ is the reference value of the wrinkle level, which represents the number and depth of wrinkles, and usually increases with age. is the average pore diameter, and a larger pore diameter may be related to age growth. Quantification indicators for pigment deposition, such as the area or concentration of pigmented spots, may increase with age. λ is the hair coverage rate. Although it has little direct relationship with age, it can be used as an auxiliary indicator because hair thinning or hair loss is sometimes related to age. are the weight coefficients of each feature for age assessment. These coefficients can be set and adjusted according to actual data or experience. e is the natural constant used to construct the sigmoid function to smooth the impact of gloss on age assessment. is the adjustment parameter for the impact of gloss, used to control the steepness of the sigmoid function.
[0025] The described human image processing module is used to perform corresponding clothing and hairstyle matching according to the user age assessment correlation coefficient of the user, obtain the preselected clothing and preselected hairstyles, compare and process the original hairstyle and original clothing of the user in the user image with the preselected hairstyles and preselected clothing respectively, screen out the set of hairstyles to be determined by the user and the set of clothing to be determined by the user, thereby further screening to obtain the image hairstyle and image clothing of the user, and import the image hairstyle and image clothing of the user into the user's image, and then obtain the user decorated image.
[0026] The described image integration and correction module is used to integrate the user interaction scenario and the user decorated image to obtain the user scenario decorated image, and perform brightness correction on the user scenario decorated image; the image generation module is used to read the brightness-corrected user scenario decorated image, record it as the final user image, and generate and display it.
[0027] The described management database is used to store the pixel density of the image, each scene theme word, the scene corresponding to each scene theme word, each clothing corresponding to the scene, each hairstyle corresponding to the scene, the age evaluation index threshold, the clothing corresponding to each age group, the hairstyle corresponding to each age group, the age evaluation index range corresponding to each age interval, the skin brightness corresponding to each gloss, the chrominance component correction coefficient, the standard hair density threshold of the user, the wrinkle level data corresponding to the wrinkle feature data matching coefficient, and the wrinkle feature data correction coefficient.
[0028] The technical solutions in the embodiments of the present application at least have the following technical effects or advantages: By improving the age prediction method, more accurate age assessment results are provided, and personalized clothing and hairstyle suggestions are provided for users according to the assessment results. By introducing methods such as skin microstructure analysis, gray-level co-occurrence matrix algorithm, color space conversion, and histogram analysis, features such as skin texture and pigment deposition can be more accurately quantified, thereby improving the accuracy of age assessment. According to the more accurate age assessment results, the human image processing module can provide clothing and hairstyle suggestions that are more in line with the actual age of the user. Through the user interaction design module to obtain the user interaction scenario, and the image integration and correction module to perform processing such as brightness correction, the finally generated user scenario decoration image is more natural and realistic, enhancing the user experience.
[0029] Embodiment 2: Based on the optimization of the age prediction result in Embodiment 1, but the image is collected when the user makes extreme expressions (such as laughing, crying, frowning, etc.). The facial expression changes may cause temporary changes in biological markers such as skin texture gloss and wrinkles, interfering with the age prediction result and reducing the accuracy. This embodiment introduces facial dynamic expressions and improves the accuracy of age prediction by removing temporary features related to expressions.
[0030] As Figure 2 shown, the method for predicting the age of a user further includes: S101, collect a facial data set and perform facial key point annotation according to the collected facial data set; Furthermore, download images from publicly available facial image data sets on the Internet, such as LFW (Labeled Faces in the Wild), CelebA, etc. The data sets usually contain rich facial images covering different ages, genders, races, and expressions, or take facial images of different people through a camera or mobile phone to ensure the diversity and authenticity of the data. Screen the collected images to remove blurred, occluded, or low-quality images to ensure the quality of the data set. Use MTCNN for automatic annotation. MTCNN is a deep learning-based facial detection algorithm that can not only detect the facial position but also perform facial key point annotation simultaneously. Determine the key points that need to be annotated, such as the corners of the eyes, the corners of the mouth, the endpoints of the eyebrows, the bridge of the nose, and the cheeks. Input the image into MTCNN, and MTCNN automatically outputs the key point positions. After the annotation is completed, we need to save the annotation result as a file.
[0031] S102, according to the annotated facial key points, use the optical flow method to calculate the displacement of the facial key points; Further, the Lucas-Kanade optical flow method is used. The Lucas-Kanade optical flow method is used to calculate the displacement of facial key points, output the displacement vector of facial key points, capture the changes in facial dynamic expressions, extract facial key points from the image sequence. Facial key points usually include feature points such as the corners of the eyes, the corners of the mouth, and the bridge of the nose, so that there is a one-to-one correspondence between key points in consecutive frames. The extracted facial key points are input into the Lucas-Kanade optical flow method. A small window is selected around each key point, and the formula for calculating the sum of the squares of the brightness differences of the pixels within the window is: , where S represents the sum of the squares of the brightness differences, W represents the set of pixel points within the window, I(x, y, t) represents the brightness value of the t-th frame image at the position (x, y), dx and dy are the assumed motion vector components (i.e., the x and y components of the displacement vector). Through Taylor expansion, the expression of the sum of the squares of the brightness differences is linearized to obtain a system of linear equations about the motion vector components. The least squares method is used to solve the linearized system of equations. The calculated displacement vector is plotted as a vector on the image to visualize the motion trajectory. The direction of the displacement vector represents the direction of the key point movement (right, down, etc.), and the magnitude represents the distance of the key point movement (i.e., the modulus length of the vector).
[0032] S103. Construct a feature vector according to the calculated displacement of the key points. The feature vector includes a first feature vector, a second feature vector, and a third feature vector. Specifically, for the displacement vector of the facial key points calculated by the Lucas-Kanade optical flow method, each key point will have a corresponding displacement vector. This vector contains two components, direction and magnitude, representing the movement of the key point between consecutive frames. Arrange the displacement vectors of all facial key points in a certain order to form a vector. The dimension of the vector is equal to the number of facial key points, and each dimension corresponds to the displacement vector of a key point. This vector is the second feature vector, which reflects the motion characteristics of the facial key points. Extract global features from the facial image dataset. For skin texture, use a texture analysis algorithm. For wrinkles, the depressions or folds on the skin surface can be detected through image processing techniques. For facial shape, use an edge detection algorithm to extract the facial contour. According to the selected feature extraction method, process the facial image, and quantify the extracted features. These features can reflect the overall attributes of the face, such as skin texture, wrinkles, facial shape, etc. Arrange the extracted global features in a certain order to form a vector. The dimension of this vector depends on the number and type of the extracted global features. This vector forms the first feature vector, which reflects the overall features of the face. Compare the first feature vector and the second feature vector, analyze the similarities and differences between them, and use the Euclidean distance to calculate the difference value. The formula is: , where represents the Euclidean distance, which measures the straight-line distance between two n-dimensional feature vectors A and B, where n represents the dimension of the feature vector. represents the i-th element in the feature vector A. The i-th element in the feature vector B. The difference value reflects the degree of change or restoration of the facial expression, that is, the expression restoration value, which is used to measure the restoration effect or degree of change of the facial expression. The larger the expression restoration value, the greater the change or the lower the restoration degree of the facial expression; the smaller the expression restoration value, the smaller the change or the higher the restoration degree of the facial expression. Combine the expression restoration value with the change of the lighting conditions, analyze the correlation between each feature in the first feature vector and the facial expression, identify the inherent features and temporary features. The inherent features are the facial contour and eye shape, and the temporary features are the wrinkles caused by facial expression changes and the shadows caused by lighting changes. According to the identified temporary features, calculate the variance of the temporary features, set the feature threshold according to the calculated variance, remove the vectors in the first feature vector that exceed the feature threshold, sort the remaining features after removing the temporary features in a certain order, and combine the sorted features into a new vector, that is, the third feature vector. The third feature vector reflects the inherent features of the face and removes the interference caused by factors such as facial expression changes, lighting changes, and occlusion.
[0033] S104. According to the characteristics of the data in the facial dataset, train a regression model. Input the obtained third feature vector into the trained regression model through the person image analysis and processing module, and the regression model outputs the correlation coefficient for optimizing the user age assessment. Further, use random forest regression as the regression model. Take the third feature vector as the input feature of the training data, and the output target is the correlation coefficient of the optimized user age assessment. Combine the input feature and the output target to form a training data set. Divide the training data set into a training set and a validation set. The training set is used to train the model, and the validation set is used to adjust the model parameters and evaluate the model performance. Use the training set to train the regression model. During the training process, the model will learn the mapping relationship between the input feature and the output target. By continuously iterating and adjusting the parameters, the performance of the model on the training set will become better and better. Use the validation set to evaluate the performance of the model. Input the third feature vector into the trained regression model, and the regression model outputs the correlation coefficient of the optimized user age assessment. The person image processing module is used to perform corresponding clothing and hairstyle matching according to the correlation coefficient of the optimized user age assessment to obtain the preselected clothing and preselected hairstyle. Compare and process the original hairstyle and original clothing of the user in the user image with the preselected hairstyle and preselected clothing respectively, screen out the set of hairstyles to be determined and the set of clothing to be determined for the user, and further screen out the image hairstyle and image clothing of the user, and import the image hairstyle and image clothing of the user into the user image, and then obtain the user decorated image.
[0034] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages: By introducing the analysis and processing of facial dynamic expressions, the accuracy of age prediction is improved. By calculating the displacement of facial key points using the optical flow method and fusing global features and dynamic expression features, the permanent features of the face can be described more accurately. After removing the temporary features related to expressions, the third feature vector can more accurately reflect the actual age characteristics of the user. By calculating the expression reduction value and correcting the face, the interference of facial expression changes on the age prediction result is effectively reduced. The regression model is trained based on the third feature vector, which can better adapt to the facial expression changes of different users and improve the generalization ability of age prediction.
[0035] Embodiment 3: When matching the hairstyle based on Embodiment 1 and Embodiment 2, the position of the user's hairline, the position of the eye corners, and the hair color will all obscure the improvement result of age in the above conversation. At the same time, when the user uses a high-definition camera for shooting, the shadow of the photo will affect the user's body posture, and the user's body posture will also obscure the improvement result of age in the above conversation. In this embodiment, by analyzing the specific effects of the position of the hairline, the position of the eye corners, and the hair color on age assessment, image processing algorithms (such as shadow removal and light balance) are used to reduce the impact of shadows on body posture recognition, improve the fit of the hairstyle and clothing to the user, make age assessment more accurate, and at the same time enhance the overall image of the user.
[0036] As Figure 3 shown, the method for optimizing the correlation coefficient of user age assessment further includes: S201. Obtain the facial and body images of the user through a camera, and use image processing algorithms to extract the facial features and body features of the user. Specifically, use a high-definition camera to obtain the facial image of the user, and preprocess the image, including adjusting the light, contrast, and removing possible noise and shadows to ensure clear image quality and rich details. Use the Canny edge detector, an edge detection algorithm, to process the preprocessed facial image, identify and extract the boundary line between the hair and the forehead, that is, the hairline, and smooth the extracted hairline to remove irregular edges and noise. Use Haar features to extract the key features of the human face. Use a large number of labeled face and non-face images to train a classifier, the support vector machine. Apply the trained classifier to the image to be detected, and search for the face area through a sliding window. When the classifier determines that a certain area is a face, the detection and positioning of the face are completed. Use shape features (such as circles, ellipses, etc.) to locate the eye area, apply the Canny edge detector, an edge detection algorithm, to extract the contour of the eye area, and further confirm the existence and position of the eyes. Use the Harris corner detector, a corner detection algorithm, to detect the corners on the eye contour to identify a specific eye corner shape, and analyze the degree of its drooping or upward tilt by calculating the angle between the outer eye corner and the horizontal line. Use the HSV (hue, saturation, value) color space for analysis, convert the image from the RGB color space to the HSV color space to more accurately extract the hair color. Use threshold segmentation technology to locate the hair area in the image, extract the color of the located hair area, and determine the hair color by calculating the average color value of all pixels in the hair area and comparing the extracted color information with predefined color categories.
[0037] Use a camera to obtain the full-body image of the user. Use the OpenPose body posture analysis algorithm to extract the body lines and contours when standing or sitting, and locate the key joint points of the body, such as the shoulders, elbows, knees, etc. Use the Euclidean distance formula to calculate the distance between joint point pairs. For angle measurement, select the corresponding joint point triples. To measure the elbow bending angle, select the shoulder, elbow, and wrist joint points, calculate the vectors between adjacent joint points. For the shoulder, elbow, and wrist, calculate the vector from the shoulder to the elbow and the vector from the elbow to the wrist, and use the dot product and modulus of the vectors to calculate the angle between the two vectors. The formula is: , where is the vector between adjacent joint points.
[0038] S202. Input the extracted facial features and body features into the age assessment algorithm to calculate the age assessment value. Specifically, the age assessment algorithm is a machine learning-based model that includes neural networks in machine learning. Facial features and body features are input into the machine learning model, and the machine learning model calculates the age assessment value through a built-in calculation formula. The formula is: , where A is the age assessment value, f is a non-linear function used to combine and transform feature values, are the weights of each feature. These weights are obtained through model training and reflect the importance of each feature for age assessment. are the input facial and body feature values, which may be preprocessed or normalized. b is the bias term used to adjust the output of the model to make it more consistent with the actual age distribution. The age assessment value is calculated according to the formula, and the model outputs the age assessment value.
[0039] S203: Collect images, select candidate hairstyles and clothing from the collected hairstyle images according to the user age assessment correlation coefficient, match the candidate hairstyles and clothing with the user's facial and body images, and use the age assessment value to adjust the matched hairstyles and clothing. Furthermore, collect a variety of hairstyle images from fashion magazines, online image libraries, and professional hairstyle design websites, classify and file the collected hairstyle images, select candidate hairstyles and clothing from the collected hairstyle images according to the user age assessment correlation coefficient, open the user's face image and hairstyle image using image processing software Photoshop, add the hairstyle image as a new layer to the user's face image, and make it roughly align with the position of the user's head by moving the hairstyle layer. Calculate the coincidence degree of the candidate hairstyle image and the user's face image at the hairline and face contour. Calculate the pixel distance between the hairstyle edge and the user's face contour to calculate the coincidence degree, and select the one with the highest coincidence degree as the matching hairstyle; collect a variety of clothing styles and sizes from fashion brand official websites, e-commerce platforms, and clothing designer works, extract the body lines and contours when standing or sitting, compare the user's body posture characteristics with the size information of the clothing one by one, and calculate the fitting degree of the clothing and the user's body posture according to the comparison results. Select the one with the highest fitting degree as the matching clothing.
[0040] The age assessment value adjusts the matched hairstyle and clothing, and divides the age assessment value into different age groups, youth (18-30 years old), middle-aged (31-50 years old) and elderly (over 50 years old), and sets the hairstyle adjustment rules: the youth group tends to be fashionable and lively, and chooses hairstyles with popular elements. The color selection is bolder, and try bright or unique colors. The style selection focuses on individual expression, and chooses asymmetrical and layered hairstyles; the middle-aged group's hairstyle style tends to be stable and mature, and chooses classic and simple hairstyles. The color selection is mainly natural and low-key. The style selection focuses on modifying the face shape and chooses hairstyles suitable for the workplace or daily occasions; the elderly group's hairstyle style tends to be comfortable and easy to care for, and chooses short or medium-long hair, avoiding overly complicated styles. The color selection is mainly natural and soft, and dyeing is considered to cover gray hair, but the color should not be too bright. The style selection focuses on practicality and comfort, and chooses hairstyles that are easy to comb and maintain. According to the user's age assessment value, determine the age group to which he belongs, and adjust the hairstyle according to the corresponding age group. Choose suitable hairstyle styles, colors and designs, and make fine adjustments to the hairstyle to ensure that it meets their personalized needs; set clothing adjustment rules: the clothing style of the young segment tends to be fashionable and trendy, so they choose slim and tight clothing styles, and the color choices are bolder and brighter, and they try contrasting colors or splicing designs. The tailoring choices focus on showing the body lines, and they choose tight pants and short skirts. The clothing style of the middle-aged segment tends to be steady and generous, so they choose classic and simple clothing styles, and the color choices are mainly natural and low-key. The tailoring choices focus on modifying the body, and they choose well-fitting suits and long skirts. The clothing style of the elderly segment tends to be comfortable and loose, so they choose loose tops and trousers, and the color choices are mainly natural and soft, and they choose dark or light-colored clothes. The tailoring choices focus on the comfort and convenience of wearing, and they can choose clothing styles that are easy to put on and take off and move around. According to the user's age assessment value, determine the age group to which they belong, and according to the clothing adjustment rules corresponding to the age group, choose suitable clothing styles, colors and tailoring, and make fine adjustments to the clothing based on the user's personal preferences, body characteristics and needs.
[0041] The technical solutions in the above-mentioned embodiments of the present application have at least the following technical effects or advantages: by adjusting the factors affecting age assessment, the fit between the hairstyle and clothing and the user is improved, the age assessment is made more accurate, and at the same time the overall image of the user is improved; by considering factors such as the hairline position, the position of the corners of the eyes, the hair color, and eliminating the influence of shadows when taking photos, the accuracy of age assessment is significantly improved; the hairstyle and clothing are adjusted according to the age assessment value to make them more suitable for the user's age and image, thereby improving the user's satisfaction and self-confidence; and by providing personalized hairstyle and clothing matching services, the user's pursuit of beauty and youthfulness is met, thereby enhancing market competitiveness.
[0042] Embodiment 4: Based on Embodiments 1 to 3, it is not accurate enough to judge whether the user's age conforms to the user's hairstyle and clothing only from the front photo (a single direction). In this solution, the overall contour and center point of the human face are integrated from the acquired facial images, and the overall contour and center point are matched with the acquired human body image. Through the obtained matching data, a 3D image model is constructed to achieve all-round coordination and natural smoothness of the hairstyle, clothing, and dynamic change effects, as Figure 4 shown.
[0043] S301. Construct a three-dimensional image model according to the acquired user's facial and body images; Specifically, obtain age estimation data according to the acquired user's facial and body images, use the 3D modeling software Blender to create a head model, combine the human body image, and use the 3D modeling software to create a body model. Combine the created head model and body model to form a complete 3D image model, ensuring that the proportion and connection between the head and the body are natural and coordinated. Adjust the parameters of the created 3D model so that the hairstyle and clothing on the user's 3D image model maintain a harmonious and unified effect from all directions (front, side, back, etc.), including facial details, body proportion, skin color, etc., so that the model is highly similar to the user's actual image. Use the age assessment value to adjust the created 3D model, adjust skin texture, facial features, etc., to more accurately reflect the user's age.
[0044] S302. Use motion capture technology to obtain real motion data, and simulate the changes of the three-dimensional image model according to the real motion data. When the 3D model is performing actions, it is judged in real time whether the clothing and hairstyle match. If they do not match, adjust the age assessment value. If they match, the age assessment value remains unchanged.
[0045] Specifically, a mechanical motion capture device is selected, and marker points are attached to key parts (such as joints). When the user walks and turns, data is captured and stored. A skeletal system and an animation controller are set in the three-dimensional image model. The captured data is mapped to the skeletal system of the 3D model. According to the mapped motion data, the 3D model is driven to perform corresponding actions. The consistency of the clothing swing amplitude refers to the matching degree between the swing amplitude of the clothing in the 3D model and the swing amplitude of the same type of clothing in the real situation. By measuring the maximum distance (or angle) of the clothing swing during walking or turning of the 3D model and comparing it with the swing amplitude of the same type of clothing in the real video under the same action, if the swing amplitude does not exceed 5% of the set swing amplitude difference, it is considered that the clothing swing amplitude is consistent with the real situation; otherwise, it is inconsistent. The rationality of the hairstyle shape change rate refers to whether the shape change rate of the hairstyle in the 3D model during the action conforms to the expected physical behavior. By calculating the change rate of the hairstyle shape during walking or turning of the 3D model (i.e., the amount of change in the hairstyle shape per unit time) and comparing it with the shape change rate of the same type of hairstyle in the real video under the same action, if the hairstyle shape change rate differs from the real situation by no more than 15%, it is considered that the hairstyle shape change rate is reasonable; if it exceeds 15%, the age assessment value is adjusted.
[0046] The technical solutions in the above embodiments of the present application at least have the following technical effects or advantages: This solution not only judges the matching degree of the user's age with the hairstyle and clothing from the front photo, but also comprehensively simulates the actual image of the user from all aspects (front, side, back, etc.) by constructing a three-dimensional image model. This ensures that the hairstyle and clothing maintain a harmonious and unified effect on the user's 3D image model, improving the accuracy and authenticity of the simulation. Combining with the motion capture technology, real-time user real motion data is obtained and the changes of the three-dimensional image model are simulated. When the 3D model performs actions, it is judged in real time whether the clothing and hairstyle match. If they do not match, the age assessment value is adjusted to ensure the coordination and natural smoothness of the dynamic change effect. Through the all-round and high-precision image simulation and the real-time judgment and adjustment of the dynamic change effect, this solution can provide users with a more real and natural virtual image experience.
[0047] The above is only the preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A character scene generation system based on AI interaction, characterized in that: include: A user interaction design module, a character image shooting module, a character image analysis and processing module, a character image processing module, an image integration and correction module, an image generation module, and a management database; the character image shooting module is used to obtain a user image and a microstructure image of the skin surface by using a high-definition camera, wherein the microstructure includes a skin gloss reference value, a wrinkle level reference value, pore features, and pigmentation, and the user image contains hair coverage; An algorithm is used to calculate pore features, and color conversion is performed on the microstructure image to obtain pigmentation. The user age assessment correlation coefficient is calculated through skin gloss, wrinkle level, pore features, pigmentation and hair coverage. The character image analysis and processing module is used to obtain user age assessment correlation data based on the user image and the microstructure of the skin surface, predict the user's age, and then obtain the user age assessment correlation coefficient. The image integration and correction module is used to integrate the user interaction scene and the user decoration image to obtain the user scene decoration image.
2. A character scene generation system based on AI interaction as claimed in claim 1, characterized in that: The specific steps of using the algorithm to calculate the pore features are as follows: use the algorithm to detect the edges where the grayscale changes dramatically in the image, the edges correspond to the contours of the pores, segment the identified pores, extract the size features of the segmented pore areas, measure the diameter of the pore areas, and count the number of pixels in the pore areas to obtain the area of the pores, extract the shape features of the segmented pore areas, and evaluate the circularity by calculating the ratio of the circumference of the pore area to the area. For non-circular pores, calculate the ratio of the longest diameter to the shortest diameter, that is, the aspect ratio.
3. The character scene generation system based on AI interaction as claimed in claim 1, characterized in that: The formula for the user age evaluation correlation coefficient ρ is: , where ψ is the reference value of skin glossiness, ξ is the reference value of wrinkle level, is the average pore diameter, is a quantitative indicator of pigment deposition, λ is the hair coverage, is the weight coefficient of each feature for age assessment, which is set and adjusted according to actual data or experience, e is a natural constant, It is the adjustment parameter affecting glossiness.
4. The character scene generation system based on AI interaction according to claim 1, characterized in that: Methods for predicting the user's age also include: S101, collecting a facial data set, and annotating facial key points according to the collected facial data set; S102, calculating the displacement of the facial key points using an optical flow method according to the marked facial key points; S103, constructing a feature vector according to the calculated displacement of the key point, where the feature vector includes a first feature vector, a second feature vector and a third feature vector; S104, training a regression model according to the characteristics of the data in the facial data set, inputting the obtained third feature vector into the trained regression model through the character image analysis and processing module, and the regression model outputs the optimized correlation coefficient of the user age assessment.
5. A character scene generation system based on AI interaction as claimed in claim 4, characterized in that: The displacement vectors of the facial key points calculated by the optical flow method are arranged in order to form a second eigenvector. The dimension of the second eigenvector is equal to the number of facial key points, and each dimension corresponds to the displacement vector of a key point. Extract global features from the facial image data set, arrange the extracted global features in order to form a first feature vector, compare the first feature vector with the second feature vector, calculate the difference value, i.e., the expression restoration value, identify the inherent features and temporary features, calculate the variance of the temporary features based on the identified temporary features, set the feature threshold based on the calculated variance, remove the vectors in the first feature vector that exceed the feature threshold, sort the remaining features after removing the temporary features, arrange them in order, and combine the sorted features into a third feature vector.
6. A character scene generation system based on AI interaction as claimed in claim 4, characterized in that: Methods for optimizing the correlation coefficient of user age assessment also include: S201, acquiring a user's facial and body images through a camera, and extracting the user's facial features and body features using an image processing algorithm; S202, inputting the extracted facial features and body features into an age assessment algorithm to calculate an age assessment value; S203, collect images, select candidate hairstyles and clothing from the collected hairstyle images according to the user age assessment correlation coefficient, match the candidate hairstyles and clothing with the user's face and body images, and use the age assessment value to adjust the matched hairstyles and clothing.
7. A character scene generation system based on AI interaction as claimed in claim 6, characterized in that: Use the camera to obtain the user's full-body image, use the algorithm to extract the body lines and contours when standing, and locate the key joints of the body. Use the distance formula to calculate the distance between two joints. For angle measurement, select the corresponding joint point triplet. To measure the elbow bending angle, select the shoulder, elbow and wrist joints, calculate the vectors between adjacent joints, and for the shoulder, elbow and wrist, calculate the vector from the shoulder to the elbow and the vector from the elbow to the wrist. Use the dot product and modulus of the vector to calculate the angle between the two vectors. The formula is: ,in, is the vector between adjacent joint points.
8. The character scene generation system based on AI interaction as claimed in claim 6, characterized in that: The steps to adjust the age estimate are: S301, constructing a three-dimensional image model based on the collected user's facial and body images; S302, using motion capture technology to obtain real motion data, simulating changes in the three-dimensional image model based on the real motion data, and judging in real time whether the clothing and hairstyle match when the 3D model is in motion. If not, adjusting the age assessment value; if matching, the age assessment value remains unchanged.
9. The character scene generation system based on AI interaction as claimed in claim 8, characterized in that: When the three-dimensional image model is performing an action, it judges in real time whether the clothing and hairstyle match, and makes judgments based on the clothing swing amplitude and the rate of change of the hairstyle shape. The clothing swing amplitude consistency refers to the degree of match between the clothing swing amplitude in the 3D model and the swing amplitude of similar clothing in real situations. By measuring the maximum swing distance of the clothing of the 3D model during walking and comparing it with the swing amplitude of similar clothing in the real video under the same action, if the swing amplitude does not exceed 5% of the set swing amplitude difference, it is considered that the clothing swing amplitude is consistent with the actual situation, otherwise it is inconsistent; the rationality of the hairstyle shape change rate refers to whether the shape change rate of the hairstyle in the 3D model during the action process conforms to the expected physical behavior. By calculating the change rate of the hairstyle shape of the 3D model during walking and comparing it with the shape change rate of similar hairstyles in the real video under the same action, the hairstyle shape change rate is considered reasonable if it does not differ from the actual situation by more than 15%. If it exceeds 15%, the age assessment value is adjusted.
10. The character scene generation system based on AI interaction according to claim 1, characterized in that: The user interaction design module includes a speech recognition unit and a scene matching unit. The speech recognition unit is used to collect the user's voice information, convert the user's voice information into text, and extract keywords. The scene matching unit is used to match the extracted keywords with the scene keywords in the database to screen and obtain the user interaction scene.
Citation Information
Patent Citations
Facial pore detection system and method
CN107679507A
Pore detection method, system, apparatus and storage medium
CN109300105A
Human face image quantitative analysis system and method
CN109730637A
Age prediction method, system and device based on facial features and textural features
CN112329607A
Image processing method and device, electronic equipment and storage medium
CN114445302A
Cited By
Decoration effect display method for interior design and naked eye VR system
CN120543452A
Mandarin fish image grading processing method based on appearance feature extraction
CN120580520A