Personalized digital human image creation and generation system based on deep learning
Through a personalized digital person image creation and generation system based on deep learning, the problem of insufficient digital person image generation efficiency, fidelity and interactive capabilities in the existing technology is solved, and efficient and personalized digital person image generation and automation processing is achieved.
Patent Information
- Application Number
- CN202510143225.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-02-10
AI Technical Summary
The prior art is difficult to accurately capture and reproduce the user's appearance and voice characteristics when creating personalized digital human images, and has limitations in terms of generation efficiency, fidelity and interaction capabilities.
A personalized digital person image creation and generation system based on deep learning is adopted, including a multivariate data acquisition module, an image creation module, a model storage module, a matching module and a rendering interaction module. By constructing multiple neural network models, facial, whole-body, and speech feature vectors are extracted, and a personalized digital human image is generated using Generative Adversarial Network (GAN) technology, combining discriminators to evaluate the matching degree with user needs, and data backtracking is performed.
It realizes the full process automation from data collection to digital person image generation, improves the efficiency and quality of digital person image generation, ensures the high personalization and accuracy of the generation of digital person image, and improves the operating efficiency of the system and the usability of the model.
Smart Images

Figure CN119600162B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital human generation, and in particular to a personalized digital human image creation and generation system based on deep learning. Background Art
[0002] With the development of digital technology, personalized digital human images have shown broad application prospects in many fields such as entertainment, education, and customer service. Users have an increasing demand for personalized digital human images, and expect to create digital humans that are both highly personalized and can interact naturally in different scenarios. However, existing technologies face many challenges in creating personalized digital human images, including how to accurately capture and reproduce the user's appearance and voice features, and how to achieve a high degree of match between the digital human image and user needs. In addition, existing technologies are also limited in terms of the generation efficiency, fidelity, and interaction capabilities of digital human images. These problems limit the application and development of personalized digital human image technology.
[0003] In order to address the above challenges, deep learning technology has been introduced into the process of creating digital human images. Deep learning technology, especially generative adversarial networks (GANs), has shown powerful capabilities in processing image and voice data, making it possible to extract feature vectors from multivariate data and generate realistic digital human images. By constructing a variety of neural network models, feature vectors of the face, whole body, and voice can be extracted separately, and adaptive strategies can be used to optimize the feature extraction process. However, although deep learning technology provides new solutions, in practical applications, it is still necessary to solve problems such as how to efficiently integrate multimodal data, how to accurately evaluate the matching degree between the generated image and user needs, and how to achieve realistic rendering and interaction of digital humans in different scenarios.
[0004] In order to solve the above defects, a technical solution is now provided. Summary of the invention
[0005] The purpose of the present invention is to solve the existing problems and propose a personalized digital human image creation and generation system based on deep learning.
[0006] The purpose of the present invention can be achieved through the following technical solutions:
[0007] A personalized digital human image creation and generation system based on deep learning, including:
[0008] The multivariate data acquisition module collects facial images and full-body photos through cameras, records voice samples using microphones, verifies environmental parameters during collection, and pre-processes the collected data;
[0009] The image creation module builds multiple neural network models to extract facial, whole body, and voice feature vectors respectively, uses adaptive strategies to optimize feature extraction, and uses GAN technology to generate personalized digital human images after fusing feature vectors. The discriminator evaluates the matching degree with user needs in various ways, and performs data backtracking analysis and optimization when it does not meet the standards. At the same time, multiple optimization strategies are used in GAN training to improve model performance;
[0010] Model storage module: establish a distributed storage database to store digital human models and related metadata, classify and manage models and establish an indexing mechanism, regularly monitor models and optimize database performance;
[0011] The matching module collects facial and full-body images to obtain new appearance feature vectors when receiving new requirements, and calculates the similarity between the new appearance feature vectors and the stored model feature vectors. If the similarity is higher than the threshold, the model is fine-tuned, and if the similarity is lower than the threshold, the model is rebuilt.
[0012] The rendering interaction module obtains the model according to the user's needs, analyzes the scene lighting environment and model material properties, uses PBR and ray tracing algorithms for rendering, dynamically adjusts the level of detail and rendering parameters according to the digital human's position and device performance, and realizes the presentation and interaction of digital humans in different scenes.
[0013] Furthermore, the specific steps of the multivariate data acquisition module are as follows:
[0014] The camera is used to collect multi-angle facial images and full-body photos of the user to ensure comprehensive information on appearance features. At the same time, the microphone is used to record multiple types of user voice samples, including daily conversations and keynote speeches, to capture voice features.
[0015] Verification is performed when collecting appearance feature information and voice feature information, and judgment is made by obtaining environmental parameters when collecting appearance features and voice features;
[0016] The collected facial and full-body images are cropped to remove irrelevant background interference; normalization processing is performed to unify the image size and color space parameters; image enhancement algorithms are used, including histogram equalization to enhance contrast and Gaussian filtering to remove noise, to improve image quality; voice samples are denoised, and a voice activity detection algorithm is used to remove silent segments, and the audio format is optimized to a unified standard format.
[0017] Furthermore, the specific operation steps of the multivariate data acquisition module for verification when collecting appearance feature information and voice feature information are as follows:
[0018] Environmental parameters for collecting appearance features include:
[0019] Light intensity: Use lux to measure, set the standard value of light intensity, and calculate the difference between the light intensity and the standard value, which is recorded as the intensity difference;
[0020] Light uniformity: quantified by calculating the variance of light intensity at different locations, using the formula: Light uniformity = ,in, is the light intensity at each measurement point, is the average light intensity, is the number of measurement points;
[0021] Background contrast: quantified by calculating the average brightness and color difference between the background and the user's appearance features;
[0022] Shooting angle deviation: The shooting angle is measured by the angle sensor installed on the camera, and the angle difference between all acquired image angles and the preset standard angle is calculated and summed up, which is recorded as the angle deviation value;
[0023] The obtained intensity difference, illumination uniformity, background contrast and angular deviation are calibrated as qi, gf, hg and jy respectively, and after normalization, they are entered into the following formula: To obtain the external value ZKL, where The preset weight coefficients are respectively the intensity difference, the illumination uniformity, the background contrast and the angle deviation; and the obtained external judgment value ZKL is compared with the preset external judgment threshold. When the external judgment value ZKL is greater than the preset external judgment threshold, it is judged that the environmental parameters when the appearance feature is collected meet the standard. Otherwise, the environmental parameters when the appearance feature is collected are adjusted and collected again;
[0024] The environmental parameters for collecting sound features include:
[0025] Background noise intensity: The background noise is quantified using the sound pressure level and recorded as the background noise value;
[0026] Reverberation time: The time required for the reverberation time sound to decay by 60 dB in the room is measured by acoustic measuring instruments, including sound level meters and impulse response meters, and recorded as the silencing value;
[0027] Sound frequency response: Use an audio analyzer to measure the sound attenuation or amplification of the environment at different frequencies to obtain a frequency response curve. This parameter is quantified by calculating the flatness of the curve and recorded as the audio frequency value.
[0028] The obtained back noise value, sound attenuation value and sound frequency value are quantified to scores between 1 and 10, respectively, and the back noise value score, sound attenuation value score and sound frequency value score are obtained respectively. After normalization, a bottom circle is established with the back noise value score as the radius, and then a cone model is established with the sound attenuation value score as the height. The vertex of the cone model is selected as the center of the sphere, and a spherical model is established with the sound frequency value as the diameter. The volume of the heteromorphic model formed by the cone and the sphere is calculated and recorded as the sound judgment value;
[0029] The obtained sound judgment value is compared with the preset sound judgment threshold. When the sound judgment value is greater than the preset sound judgment threshold, it is judged that the environmental parameters when the sound feature is collected meet the standard. Otherwise, the environmental parameters when the sound feature is collected are optimized and the sound features are collected again.
[0030] Furthermore, the specific operation steps of the image creation module are as follows:
[0031] Construct a convolutional neural network model, train on facial images, learn the shape, proportion, position relationship of facial features and skin texture features, and output facial feature vectors; use a human posture recognition model based on the convolutional neural network model to process full-body photos, extract body proportions, limb morphology, and body posture features, and form a full-body feature vector;
[0032] Use recurrent neural network or long short-term memory network to analyze speech samples, extract fundamental frequency, formant, pitch change, speech rate change, prosodic features and emotional color, and generate speech feature vectors;
[0033] Adopt adaptive feature extraction strategies based on the characteristics of different modal data; for facial images with expression changes greater than the preset standard, increase the depth of the convolutional neural network model to enhance feature learning capabilities; for dialects or accents in speech samples, optimize the model's parameter initialization method to improve the accuracy of feature extraction;
[0034] Build a basic digital human model based on the statistical average human model or the 3D scanned standard human model, and initialize the model parameters according to the basic characteristics of the user, including bone structure, muscle distribution, and basic skin texture;
[0035] The extracted multimodal feature vectors are fused by weighted summation or vector concatenation to form a comprehensive feature vector. Generative adversarial network technology is used to build a model architecture including a generator and a discriminator. The generator uses the comprehensive feature vector as input and generates a personalized digital human image through a multi-layer neural network, including facial features and body shape that matches the body shape.
[0036] The discriminator receives the generated digital human image and real human image data and determines the degree of match with user needs;
[0037] Through adversarial training between the generator and the discriminator, the parameters of the generator are optimized to improve the quality and personalization of the generated digital human image;
[0038] In the adversarial network training process, an optimization strategy is adopted; an attention mechanism is introduced to make the generator pay attention to the user's personalized characteristics when generating digital human images, including highlighting the user's facial identification or clothing details;
[0039] By using multi-scale training methods, the model can learn the characteristics of digital human images at different resolutions, thereby improving the detail richness and overall coordination of the generated image. At the same time, by using transfer learning technology, the model parameters pre-trained on large-scale general data sets are transferred to the model, which speeds up the training convergence and improves the model performance.
[0040] Furthermore, the specific operation steps of the image creation module to determine the degree of matching with user needs are as follows:
[0041] Conduct appearance feature matching evaluation and style feature quantitative comparison;
[0042] Feature vector distance measurement: The appearance features in the user's requirements and the corresponding features of the generated digital human image are converted into feature vectors respectively; for the shape of the facial features, vectors are constructed by the coordinates of the key points; the body proportions are constructed by the proportional values; and then the cosine similarity between the two is calculated: ,in is the user demand vector, Generate a digital human vector; and use cosine similarity as the standard to measure the distance between the digital human image and the user demand feature vector, recorded as the special evaluation value;
[0043] Attribute matching establishes matching rules and scoring systems for hairstyle, skin color, and clothing style attributes. Specifically, hairstyle is judged by hairstyle type, length range, and bangs style attributes. A complete match is scored as 1 point, a partial match is scored as 0.5 points, and a mismatch is scored as 0 points, which are recorded as the hairstyle score. Skin color is calculated in the Lab color space based on the distance from the user's desired skin color. , where and They represent the skin color value expected by the user and the skin color value of the generated digital person respectively; and They represent the green to red color component attribute values expected by the user and the green to red color component attribute values of the generated digital person respectively; and Represent the blue to yellow color component attribute values expected by the user and the blue to yellow color component attribute values of the generated digital person respectively; preset several distance intervals, calculate the distance interval corresponding to the distance d of the user's required skin color in the Lab color space according to the calculated skin color, determine the score, and record it as the skin color score; the clothing style is judged by the predefined style classification model, and the matching probability is higher than 0.8 and gets 1 point, 0.5-0.8 gets 0.5 points, and lower than 0.5 gets 0 points, which is recorded as the clothing score; after normalizing the hairstyle score, skin color score and clothing score, they are respectively recorded as clothing scores; establish a triangle with the hairstyle score, skin color score and clothing score as the three sides of the triangle, calculate the area of the triangle, and record it as the attribute score, and use this attribute score as the result of measuring attribute matching;
[0044] Quantification of style features: Quantify different styles, calculate the Mahalanobis distance between the clothing color of the digital person and the corresponding quantified style clothing color, and record it as the style evaluation value;
[0045] Emotional feature matching evaluation: Based on the theory of emotional psychology, emotional features such as liveliness and calmness are associated with facial expressions and gestures and quantified; the cosine similarity between the generated digital human emotional feature vector and the user demand emotional feature vector is calculated and recorded as the emotional evaluation value;
[0046] The obtained special evaluation value, attribute score, reputation value and sentiment evaluation value are marked as TR, SY, FO and QU respectively, and after normalization, they are put into the following formula: To get the comprehensive judgment score DFQ, The preset weight coefficients are respectively the special evaluation value, attribute score, reputation value and emotional evaluation value. The obtained comprehensive judgment score DFQ is compared with the preset comprehensive judgment score threshold. When the comprehensive judgment score DFQ is greater than the preset comprehensive judgment score threshold, it is judged that the matching degree between the generated digital human image and the user demand meets the standard. Otherwise, data backtracking is performed to determine the non-compliant part.
[0047] Furthermore, the image creation module performs data backtracking and determines the specific operation steps of the non-compliant parts as follows:
[0048] The special evaluation value, attribute score, moral evaluation value and sentiment evaluation value obtained from the analysis are determined separately to determine which parameters lead to the overall matching degree not meeting the standard;
[0049] When the special evaluation value does not meet the standard, check whether the collected user appearance image is clear, whether the multiple angles are complete, and whether there are errors in the feature extraction process; check whether the image acquisition device parameter settings are correct, including camera resolution and focus; check whether the feature extraction algorithm does not meet the standard for processing certain special appearance features;
[0050] When the attribute score does not meet the standard, the corresponding generation module is adjusted for the attribute that does not meet the standard; including: analyzing the parameters and algorithms for skin color generation, re-adjusting the probability distribution or color space conversion method for skin color generation; enriching the training data of clothing styles, and adding parameters for clothing style control in the generator to ensure that clothing styles that meet user needs can be generated;
[0051] When the style rating does not meet the standard, rebuild or train the style quantification model; collect more sample data of different styles, re-extract and quantify the features of these styles; use unsupervised learning methods, including clustering algorithms, to assist in discovering the intrinsic features of different styles, so that the style quantification model can more accurately reflect various styles; adjust the style generation mechanism; add a style control vector to the input of the generator, and control the style of the generated digital human by adjusting the value of the vector; at the same time, during the training process, add a style loss function so that the generator pays more attention to the generation of style features when generating the digital human image, reducing the deviation from the style required by the user;
[0052] When the emotion rating value does not meet the standard, the emotion generation module is improved; the parameters of the emotion-related neural network layer in the generator are adjusted, including adjusting the parameters related to facial muscle movements according to the emotional characteristics of user needs when generating facial expressions; at the same time, during the training process, the supervisory signal of the emotional characteristics is added so that the generator can better learn and generate digital human images that meet the emotional characteristics of user needs.
[0053] Furthermore, the specific operation steps of the model storage module are as follows:
[0054] Establish a model storage database, adopt a distributed storage architecture, store the generated digital human model in a binary format or a custom structured format in the database, and associate the user's identity information, collected data features, and model training parameter metadata for subsequent query, call, and management;
[0055] According to different application scenarios and user needs, the stored models are classified and managed, and a corresponding indexing mechanism is established to improve the efficiency of model retrieval;
[0056] Regularly monitor the stored digital human models and check the integrity and accuracy of the models; at the same time, optimize the performance of the model storage database, such as data compression, index reconstruction and other operations, to ensure the efficiency of system operation.
[0057] Furthermore, the specific operation steps of the matching module are as follows:
[0058] When receiving a new user image creation requirement, the facial and full-body images collected by the multivariate data acquisition module are processed to obtain a new appearance feature vector; for each model in the storage database, its similarity with the new feature vector is calculated. The specific process is as follows:
[0059] Feature vector normalization: normalize the new feature vector and the model feature vector so that each feature dimension has the same weight and scale;
[0060] Distance metric selection: Cosine similarity is used to calculate the distance between two feature vectors, using the formula: ,in and are the new feature vector and the model feature vector respectively, and The magnitude of a vector;
[0061] Similarity evaluation: According to the calculated distance value, it is converted into a similarity score. The smaller the distance, the higher the similarity. The similarity threshold is set. When the similarity score is higher than the threshold, it is considered that the matching degree between the model and the user needs meets the standard. Then the model is retrieved for fine-tuning. The specific process is as follows:
[0062] According to the difference analysis of the feature vectors in the similarity calculation process, the difference between the feature parameters of the current model and the new requirements is quantified to a value between 0 and 1 and recorded as the difference judgment value. The difference judgment value is compared with the preset difference judgment threshold. When the difference judgment value is greater than the preset difference judgment threshold, the feature is judged to be a key parameter;
[0063] Adjust the weights of key parameters; the weight adjustment is based on preset adjustment rules or adjustment strategies learned through machine learning algorithms;
[0064] Use the gradient-based method to perform incremental updates. For the determined key parameters, calculate the gradient of the loss function relative to the new requirement. Let the loss function L represent the difference between the digital human image generated by the model and the appearance features of the new requirement. For the key parameters ,calculate , which represents the loss function for the L model parameters The partial derivative of; then, according to the principle of the gradient descent algorithm, with the learning rate Update the parameters: ;
[0065] In the incremental update process, an adaptive learning rate strategy is adopted; including the use of Adagrad, Adadelta or Adam adaptive learning rate optimization algorithms; the learning rate is automatically adjusted according to the historical gradient information of the parameters, so that when updating key parameters, the update step size can be dynamically adjusted according to the importance and update history of the parameters;
[0066] When the similarity score is lower than the threshold, it is determined that the model needs to be rebuilt.
[0067] Furthermore, the specific operation steps of the rendering interaction module are as follows:
[0068] When the user issues a request for retrieval, the rendering interaction module first receives the request; the request contains the user's relevant information on the digital human application scenario; according to the user's retrieval request, the model is obtained from the model storage module or the matching module;
[0069] When the user has no new requirements, the pre-stored model that meets the user's basic requirements is directly selected from the model storage module; if the user has new requirements and the matching module has fine-tuned the model, the fine-tuned model in the matching module is obtained;
[0070] After obtaining the model, analyze the lighting environment of the virtual scene where the digital human is located; including determining the position, intensity, color and type of the light source; setting it according to the material properties of the digital human model itself; material properties include the roughness, glossiness, transparency of the skin, and the texture and feel of the clothes; for the skin, set different roughness values to simulate the texture of real skin; for clothes, load pre-designed texture patterns to express the material of the clothes;
[0071] Using physically based rendering technology combined with ray tracing algorithms, the physical phenomena of light propagation, reflection, refraction and scattering are calculated according to scene lighting and model material properties, giving digital humans realistic light and shadow performance;
[0072] Monitor the position of the digital human in the virtual scene in real time and calculate the distance between the digital human and the viewing angle; dynamically adjust the detail level of the model according to the distance between the digital human and the scene and the user's focus factor; when the digital human is in the distance, use the preset low-detail model to reduce the amount of rendering calculation and improve rendering efficiency; when the digital human is in the focus position of the scene, switch to the preset high-detail model to show facial expressions and clothing texture details to ensure high quality of visual effects;
[0073] Detect the graphics processing capabilities of the user's device, including GPU model, video memory size and hardware parameters; automatically adjust rendering parameters based on hardware performance.
[0074] Compared with the prior art, the present invention has the following beneficial effects:
[0075] (1) The present invention realizes the full process automation from data collection to digital human image generation by constructing a multivariate data acquisition module, an image creation module, a model storage module, a matching module and a rendering interaction module. The multivariate data acquisition module can fully capture the user's appearance and voice features, while the image creation module uses deep learning technology to extract feature vectors and adopts generative adversarial network technology to generate personalized digital human images. This end-to-end system design greatly improves the efficiency and quality of digital human image generation. At the same time, through the introduction of the discriminator, the system can automatically evaluate the matching degree between the generated image and the user's needs, and perform data backtracking analysis and optimization when it does not meet the standards, ensuring the high personalization and accuracy of the generated digital human image.
[0076] (2) In the present invention, the model storage module adopts a distributed storage architecture to effectively store and manage digital human models and related metadata. By establishing a classification management and indexing mechanism, the model retrieval efficiency is improved, making the call and management of digital human models more convenient. The system regularly monitors the integrity and accuracy of the model and performs performance optimization, such as data compression and index reconstruction, to ensure the high efficiency of system operation and the high availability of the model. This optimized storage and management mechanism not only improves the storage efficiency of digital human images, but also provides a solid foundation for rapid response to new needs.
[0077] (3) In the present invention, the rendering interaction module dynamically adjusts the detail level and rendering parameters of the digital human image according to user needs, thereby achieving realistic presentation and smooth interaction of the digital human in different scenarios. By analyzing the scene lighting environment and model material properties, and using physically based rendering technology combined with ray tracing algorithms, it can accurately calculate physical phenomena such as light propagation, reflection, refraction and scattering, giving the digital human a realistic light and shadow performance. It can also adjust the rendering parameters in real time according to the position of the digital human in the virtual scene and the graphics processing capabilities of the user's device, ensuring that while maintaining high-quality visual effects, the rendering efficiency is improved to meet the needs of different application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] In order to facilitate understanding by those skilled in the art, the present invention is further described below in conjunction with the accompanying drawings;
[0079] Figure 1 This is the overall system block diagram of the present invention. DETAILED DESCRIPTION
[0080] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0081] It should be understood that the terms “include” and “comprising” used in the specification and claims of the present disclosure indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0082] It should also be understood that the terms used in this disclosure are only for the purpose of describing specific embodiments and are not intended to limit the disclosure. As used in this disclosure and claims, the singular forms of "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should also be further understood that the term "and / or" used in this disclosure and claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations.
[0083] like Figure 1 As shown, a personalized digital human image creation and generation system based on deep learning includes a multivariate data acquisition module, an image creation module, a model storage module, a matching module, and a rendering interaction module;
[0084] The multivariate data acquisition module collects facial images and full-body photos through the camera, and records voice samples through the microphone. It verifies environmental parameters during the acquisition (appearance feature acquisition involves light intensity, uniformity, background contrast, and shooting angle deviation, and voice feature acquisition involves background noise intensity, reverberation time, and sound frequency response), and performs pre-processing such as cropping, normalization, and enhancement on the collected data to ensure data quality in order to fully obtain appearance and voice feature information; uses the camera to collect multi-angle facial images (including expressions) and full-body photos (covering a variety of postures and different lighting conditions) of the user to ensure that the appearance feature information is fully obtained; at the same time, uses the microphone to record multiple types of user voice samples, including daily conversations and keynote speeches, to capture voice features;
[0085] Verification is performed when collecting appearance feature information and voice features, and judgment is made by obtaining environmental parameters when collecting appearance features and voice features, where the environmental parameters when collecting appearance features include:
[0086] Light intensity: Lux is used to measure, a standard value for light intensity is set, and the difference between the light intensity and the standard value is calculated, which is recorded as the intensity difference; Light uniformity: quantified by calculating the variance of light intensity at different locations, using the formula: Light uniformity = ,in, is the light intensity at each measurement point, is the average light intensity, is the number of measurement points; Background contrast: quantified by calculating the average brightness difference and color difference between the background and the user's appearance features (such as face or body); Shooting angle deviation: The shooting angle is measured by the angle sensor installed on the camera, and the angle difference between all acquired image angles and the preset standard angle is calculated and summed up, which is recorded as the angle deviation value;
[0087] The obtained intensity difference, illumination uniformity, background contrast and angular deviation are calibrated as qi, gf, hg and jy respectively, and after normalization, they are entered into the following formula: To obtain the external value ZKL, where The preset weight coefficients are respectively the intensity difference, the illumination uniformity, the background contrast and the angle deviation; and the obtained external judgment value ZKL is compared with the preset external judgment threshold. When the external judgment value ZKL is greater than the preset external judgment threshold, it is judged that the environmental parameters when the appearance feature is collected meet the standard. Otherwise, the environmental parameters when the appearance feature is collected are adjusted and collected again;
[0088] The environmental parameters for collecting sound features include:
[0089] Background noise intensity: Use sound pressure level (in decibels, dB) to quantify background noise and record it as background noise value; Reverberation time: Use acoustic measuring instruments, including sound level meters and impulse response meters, to measure the time required for reverberation time sound to decay 60 dB in the room and record it as the attenuation value; Sound frequency response: Use an audio analyzer to measure the sound attenuation or amplification of the environment at different frequencies, obtain the frequency response curve, and quantify this parameter by calculating the flatness of the curve and record it as the audio frequency value;
[0090] The obtained back noise value, sound attenuation value and audio frequency value are quantified to scores between 1 and 10 respectively to obtain the back noise value score, sound attenuation value score and audio frequency value score respectively. After normalization, a bottom circle is established with the back noise value score as the radius, and a cone model is established with the sound attenuation value score as the height. The vertex of the cone model is selected as the center of the sphere, and a spherical model is established with the audio frequency value as the diameter. The volume of the heteromorphic model formed by the cone and the sphere is calculated and recorded as the sound judgment value. The obtained sound judgment value is compared with the preset sound judgment threshold. When the sound judgment value is greater than the preset sound judgment threshold, it is judged that the environmental parameters when the sound feature is collected meet the standard. Otherwise, the environmental parameters when the sound feature is collected are optimized and the sound feature is collected again.
[0091] The collected facial and full-body images are cropped to remove irrelevant background interference; normalization processing is performed to unify the image size and color space parameters; image enhancement algorithms are used, including histogram equalization to enhance contrast and Gaussian filtering to remove noise, to improve image quality; voice samples are denoised, and a voice activity detection algorithm is used to remove silent segments, and the audio format is optimized to a unified standard format.
[0092] The image creation module extracts facial, whole body, and voice feature vectors by constructing multiple neural network models, and uses adaptive strategies to optimize feature extraction. After fusing feature vectors, it uses GAN technology to generate personalized digital human images. The discriminator evaluates the matching degree with user needs in multiple ways, and performs data backtracking analysis and optimization when it does not meet the standards. At the same time, multiple optimization strategies are used in GAN training to improve model performance.
[0093] Construct a convolutional neural network (CNN) model, train on facial images, learn the shape, proportion, positional relationship of facial features and skin texture features, and output facial feature vectors; use a CNN-based human posture recognition model to process full-body photos, extract body proportions (such as shoulder width to waist width ratio, height to leg length ratio, etc.), limb morphology (such as arm curvature, leg lines, etc.), body characteristics (such as center of gravity offset when standing, walking posture characteristics, etc.), and form a full-body feature vector; use a recurrent neural network (RNN) or a long short-term memory network (LSTM) to analyze speech samples, extract fundamental frequency, resonance peaks, pitch changes, speech speed changes, rhythmic features and emotional colors (such as joy, sadness, anger, etc.), and generate speech feature vectors;
[0094] Adopt adaptive feature extraction strategies based on the characteristics of different modal data; for data with large expression changes in facial images, increase the depth of the convolution layer of the CNN model to enhance feature learning capabilities; for dialects or special accents in voice samples, optimize the parameter initialization method of the RNN / LSTM model to improve the accuracy of feature extraction; build a basic digital human model based on the statistical average human model or the 3D scanned standard human model, and initialize the model parameters according to the basic characteristics of the user (such as gender, age range, approximate height and weight, etc.), including bone structure, muscle distribution, and basic skin texture;
[0095] The extracted multimodal feature vectors are fused, and a weighted summation or vector concatenation fusion method is used to form a comprehensive feature vector. The generative adversarial network (GAN) technology is used to build a model architecture including a generator and a discriminator. The generator uses the comprehensive feature vector as input and generates a personalized digital human image through a multi-layer neural network, including facial features (such as unique eye shape, personalized lip contour, etc.) and a body shape that conforms to the body shape (generating a suitable clothing effect according to the user's body shape). The discriminator receives the generated digital human image and real human image data to determine the degree of match with user needs. The specific process is as follows:
[0096] Perform appearance feature matching evaluation and style feature quantitative comparison respectively;
[0097] Feature vector distance measurement: The appearance features in user requirements (such as facial features, body proportions, etc.) and the corresponding features of the generated digital human image are converted into feature vectors respectively; for facial features, vectors are constructed by key point coordinates; for body proportions, vectors are directly constructed by proportional values; then the cosine similarity between the two is calculated: ,in is the user demand vector, Generate a digital human vector; and use cosine similarity as the standard to measure the distance between the digital human image and the user's required feature vector, recorded as the special evaluation value; attribute matching is based on the attributes of hairstyle, skin color, and clothing style, and establishes matching rules and scoring systems; specifically: the hairstyle is judged by the hairstyle type (straight hair, curly hair), length range, and bangs style attributes, with 1 point for a complete match, 0.5 points for a partial match, and 0 points for a mismatch, which are recorded as the hairstyle score; the skin color is calculated in the Lab color space to be the distance from the user's required skin color , where and They represent the skin color value expected by the user and the skin color value of the generated digital person respectively; and They represent the green to red color component attribute values expected by the user and the green to red color component attribute values of the generated digital person respectively; and Represent the blue to yellow color component attribute values expected by the user and the blue to yellow color component attribute values of the generated digital person respectively; preset several distance intervals, calculate the distance interval corresponding to the distance d of the user's required skin color in the Lab color space according to the calculated skin color, determine the score, and record it as the skin color score; the clothing style is judged by the predefined style classification model, and the matching probability is higher than 0.8 and gets 1 point, 0.5-0.8 gets 0.5 points, and lower than 0.5 gets 0 points, which is recorded as the clothing score; after normalizing the hairstyle score, skin color score and clothing score, they are respectively recorded as clothing scores; establish a triangle with the hairstyle score, skin color score and clothing score as the three sides of the triangle, calculate the area of the triangle, and record it as the attribute score, and use this attribute score as the result of measuring attribute matching;
[0098] Style feature quantification: quantify different styles, calculate the Mahalanobis distance between the clothing color of the digital human and the corresponding quantified style clothing color, and record it as the style evaluation value; emotional feature matching evaluation: based on the theory of emotional psychology, associate and quantify emotional features such as liveliness and calmness with facial expressions (such as the degree of eye opening and closing, the angle of the corners of the mouth, etc.) and posture movements (such as gesture amplitude, body swing frequency, etc.); calculate the cosine similarity between the emotional feature vector of the digital human and the emotional feature vector of the user's needs, and record it as the emotional evaluation value;
[0099] The obtained special evaluation value, attribute score, reputation value and sentiment evaluation value are marked as TR, SY, FO and QU respectively, and after normalization, they are put into the following formula: To get the comprehensive judgment score DFQ, The preset weight coefficients are the special evaluation value, attribute score, reputation value and sentiment evaluation value respectively. The obtained comprehensive judgment score DFQ is compared with the preset comprehensive judgment score threshold. When the comprehensive judgment score DFQ is greater than the preset comprehensive judgment score threshold, it is judged that the matching degree between the generated digital human image and the user demand meets the standard. Otherwise, the data is backtracked to determine the non-compliant part. The specific process is as follows:
[0100] The special evaluation value, attribute score, moral evaluation value and sentiment evaluation value obtained from the analysis are determined separately to determine which parameters lead to the overall matching degree not meeting the standard;
[0101] When the special evaluation value does not meet the standard, check whether the collected user appearance image is clear, whether the multiple angles are complete, and whether there are errors in the feature extraction process; check whether the image acquisition device parameter settings are correct, including camera resolution and focus; check whether the feature extraction algorithm does not meet the standard for processing certain special appearance features; when the attribute score does not meet the standard, adjust the corresponding generation module for the attributes that do not meet the standard; including: analyzing the parameters and algorithms for skin color generation, re-adjusting the probability distribution or color space conversion method for skin color generation; enriching the training data of clothing styles, and adding parameters for clothing style control in the generator to ensure that clothing styles that meet user needs can be generated; when the style evaluation value does not meet the standard, rebuild or train the style quantification model. Collect more sample data of different styles, such as image data of fashion, classical and other styles, and re-extract and quantify these styles. Unsupervised learning methods, such as clustering algorithms, can be used to assist in discovering the intrinsic characteristics of different styles, so that the style quantification model can more accurately reflect various styles; adjust the style generation mechanism; add a style control vector to the input of the generator, and accurately control the style of the generated digital person by adjusting the value of the vector. At the same time, during the training process, the style loss function is added so that the generator pays more attention to the generation of style features when generating digital human images, reducing the deviation from the style required by users; when the emotional evaluation value does not meet the standard, the emotion generation module is improved. The parameters of the neural network layer related to emotions in the generator are adjusted. For example, when generating facial expressions, the parameters related to facial muscle movements are adjusted according to the emotional characteristics required by users. At the same time, during the training process, the supervisory signal of emotional features is added so that the generator can better learn and generate digital human images that meet the emotional characteristics required by users.
[0102] Through adversarial training of the generator and the discriminator, the parameters of the generator are continuously optimized to improve the quality and personalization of the generated digital human image; in the GAN training process, an optimization strategy is adopted; an attention mechanism is introduced to enable the generator to pay more attention to the user's specific personalized features when generating digital human images, including highlighting the user's unique facial features or emphasizing clothing details; a multi-scale training method is used to allow the model to learn the characteristics of digital human images at different resolutions, thereby improving the richness of details and overall coordination of the generated image; at the same time, the transfer learning technology is used to migrate the model parameters pre-trained on a large-scale general data set to the model of this system, thereby accelerating the training convergence speed and improving the model performance.
[0103] The model storage module is used to establish a distributed storage database, store digital human models and related metadata, classify and manage models and establish an indexing mechanism, regularly monitor models and optimize database performance to ensure storage efficiency, availability and convenience of model management;
[0104] Establish a model storage database and adopt a distributed storage architecture to ensure high availability and scalability of storage; store the generated digital human model in a binary format or a custom structured format in the database, and associate the user's identity information, collected data features, and model training parameter metadata for subsequent query, call, and management; classify and manage the stored models according to different application scenarios and user needs. For example, store digital human models used for virtual live broadcasts and models for education and training scenarios in different storage areas, and establish corresponding indexing mechanisms to improve model retrieval efficiency; regularly monitor the stored digital human models to check the integrity and accuracy of the models; at the same time, optimize the performance of the model storage database, such as data compression, index reconstruction, and other operations to ensure the efficiency of system operation.
[0105] The matching module is used to collect facial and full-body images to obtain new appearance feature vectors when receiving new requirements, and calculate their similarity with the stored model feature vectors. If the similarity is higher than the threshold, the model is fine-tuned (including determining key parameters, adjusting weights, incremental updates, etc.), and if the similarity is lower than the threshold, the model is rebuilt to quickly adjust the model to adapt to new requirements;
[0106] When receiving a new user image creation requirement, the facial and full-body images collected by the multivariate data acquisition module are processed to obtain a new appearance feature vector; for each model in the storage database, its similarity with the new feature vector is calculated. The specific process is as follows:
[0107] Feature vector normalization: normalize the new feature vector and the model feature vector so that each feature dimension has the same weight and scale for easy comparison; distance metric selection: use cosine similarity to calculate the distance between two feature vectors, through the formula: ,in and are the new feature vector and the model feature vector respectively, and The modulus of the vector; Similarity evaluation: According to the calculated distance value, it is converted into a similarity score. The smaller the distance, the higher the similarity; a similarity threshold is set. When the similarity score is higher than the threshold, it is considered that the matching degree between the model and the user's needs meets the standard; then the model is retrieved for fine-tuning. The specific process is as follows:
[0108] According to the difference analysis of the feature vectors in the similarity calculation process, the difference between the feature parameters of the current model and the new requirements is quantified to a value between 0 and 1 and recorded as the difference judgment value. The difference judgment value is compared with the preset difference judgment threshold. When the difference judgment value is greater than the preset difference judgment threshold, the feature is judged to be a key parameter;
[0109] Adjust the weights of key parameters. If the new requirement emphasizes a certain feature, such as the user wants the digital human to have brighter eyes, then the weights of parameters related to eye appearance can be increased. This weight adjustment is based on preset adjustment rules or adjustment strategies learned through machine learning algorithms; incremental updates are performed using gradient-based methods. For example, for the determined key parameters, calculate the gradient of the loss function relative to the new requirement; assuming that the loss function L represents the difference between the digital human image generated by the model and the appearance characteristics of the new requirement, for the key parameters ,calculate , which represents the loss function for the L model parameters The partial derivative of; Then, according to the principle of the gradient descent algorithm, with a smaller learning rate Update the parameters: This small update can avoid having too big an impact on the overall structure of the model;
[0110] In the incremental update process, an adaptive learning rate strategy is adopted, including the use of Adagrad, Adadelta or Adam adaptive learning rate optimization algorithms. These algorithms automatically adjust the learning rate according to the historical gradient information of the parameters, so that when updating key parameters, the update step size can be dynamically adjusted according to the importance and update history of the parameters, so as to more efficiently converge to the parameter value that meets the new requirements.
[0111] When the similarity score is lower than the threshold, it is determined that the model needs to be rebuilt.
[0112] The rendering interaction module is used to obtain the model according to the user's needs, analyze the scene lighting environment and model material properties, use PBR and ray tracing algorithms for rendering, and dynamically adjust the level of detail and rendering parameters according to the digital human's position and device performance to achieve realistic presentation and smooth interaction of the digital human in different scenes;
[0113] When the user issues a request for retrieval, the rendering interaction module first receives the request. The request may contain relevant information about the user's application scenario for the digital human; according to the user's retrieval requirements, the model is obtained from the model storage module or the matching module; if the user has no new requirements, the pre-stored model that meets the application scenario and the user's basic requirements is directly selected from the model storage module; if the user has new requirements and the matching module has fine-tuned the model, then the fine-tuned model in the matching module is obtained;
[0114] After obtaining the model, analyze the lighting environment of the virtual scene where the digital human is located. This includes determining the position, intensity, color and type of the light source (such as point light source, parallel light source, etc.); setting it according to the material properties of the digital human model itself. Material properties include the roughness, glossiness, transparency of the skin, and the texture and feel of the clothes. For the skin, set different roughness values to simulate the texture of real skin; for clothes, load pre-designed texture patterns to express the material of the clothes; use physically based rendering (PBR) technology combined with ray tracing algorithms to accurately calculate physical phenomena such as light propagation, reflection, refraction and scattering according to the scene lighting and model material properties, giving the digital human a realistic light and shadow performance;
[0115] Monitor the position of the digital human in the virtual scene in real time and calculate the distance between the digital human and the viewing angle; dynamically adjust the level of detail of the model based on factors such as the distance between the digital human and the scene and the user's focus; when the digital human is in the distance, use a low-detail model to reduce the amount of rendering calculations and improve rendering efficiency. Low-detail models may simplify details such as facial expressions and clothing textures; when the digital human is close to or in the focus of the scene, switch to a high-detail model to show rich facial expressions, clothing textures and other details to ensure high-quality visual effects; detect the graphics processing capabilities of the user's device, including hardware parameters such as GPU model and video memory size; automatically adjust rendering parameters based on hardware performance.
[0116] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only specific implementation methods. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A personalized digital human image creation and generation system based on deep learning, characterized by: include: The multivariate data acquisition module collects facial images and full-body photos through cameras, records voice samples using microphones, verifies environmental parameters during collection, and pre-processes the collected data; The image creation module builds multiple neural network models to extract facial, whole body, and voice feature vectors respectively, uses adaptive strategies to optimize feature extraction, and uses GAN technology to generate personalized digital human images after fusing feature vectors. The discriminator evaluates the matching degree with user needs in various ways, and performs data backtracking analysis and optimization when it does not meet the standards. At the same time, multiple optimization strategies are used in GAN training to improve model performance; Model storage module: establish a distributed storage database to store digital human models and related metadata, classify and manage models and establish an indexing mechanism, regularly monitor models and optimize database performance; The matching module collects facial and full-body images to obtain new appearance feature vectors when receiving new requirements, and calculates the similarity between the new appearance feature vectors and the stored model feature vectors. If the similarity is higher than the threshold, the model is fine-tuned, and if the similarity is lower than the threshold, the model is rebuilt. The rendering interaction module obtains the model according to the user's needs, analyzes the scene lighting environment and model material properties, uses PBR and ray tracing algorithms for rendering, and dynamically adjusts the level of detail and rendering parameters according to the digital human's position and device performance to achieve the presentation and interaction of digital humans in different scenes; The specific steps of the multivariate data acquisition module are as follows: The camera is used to collect multi-angle facial images and full-body photos of the user to ensure comprehensive information on appearance features. At the same time, the microphone is used to record multiple types of user voice samples, including daily conversations and keynote speeches, to capture voice features. Verification is performed when collecting appearance feature information and voice feature information, and judgment is made by obtaining environmental parameters when collecting appearance features and voice features; The collected facial and full-body images are cropped to remove irrelevant background interference; normalization is performed to unify image size and color space parameters; image enhancement algorithms are used, including histogram equalization to enhance contrast and Gaussian filtering to remove noise, to improve image quality; voice samples are denoised, voice activity detection algorithms are used to remove silent segments, and audio formats are optimized to a unified standard format; The specific operation steps of the multivariate data acquisition module for verification when collecting appearance feature information and voice feature information are as follows: Environmental parameters for collecting appearance features include: Light intensity: Use lux to measure, set the standard value of light intensity, and calculate the difference between the light intensity and the standard value, which is recorded as the intensity difference; Light uniformity: quantified by calculating the variance of light intensity at different locations, using the formula: Light uniformity = ,in, is the light intensity at each measurement point, is the average light intensity, is the number of measurement points; Background contrast: quantified by calculating the average brightness and color difference between the background and the user's appearance features; Shooting angle deviation: The shooting angle is measured by the angle sensor installed on the camera, and the angle difference between all acquired image angles and the preset standard angle is calculated and summed up, which is recorded as the angle deviation value; The obtained intensity difference, illumination uniformity, background contrast and angular deviation are calibrated as qi, gf, hg and jy respectively, and after normalization, they are entered into the following formula: To obtain the external value ZKL, where are respectively the preset weight coefficients of intensity difference, illumination uniformity, background contrast and angular deviation; and the obtained external judgment value ZKL is compared with the preset external judgment threshold. When the external judgment value ZKL is greater than the preset external judgment threshold, it is judged that the environmental parameters when the appearance feature is collected meet the standard. Otherwise, the environmental parameters when the appearance feature is collected are adjusted and collected again; The environmental parameters for collecting sound features include: Background noise intensity: The background noise is quantified using the sound pressure level and recorded as the background noise value; Reverberation time: The time required for the reverberation time sound to decay by 60 dB in the room is measured by acoustic measuring instruments, including sound level meters and impulse response meters, and recorded as the silencing value; Sound frequency response: Use an audio analyzer to measure the sound attenuation or amplification of the environment at different frequencies to obtain a frequency response curve. This parameter is quantified by calculating the flatness of the curve and recorded as the audio frequency value. The obtained back noise value, sound attenuation value and sound frequency value are quantified to scores between 1 and 10, respectively, and the back noise value score, sound attenuation value score and sound frequency value score are obtained respectively. After normalization, a bottom circle is established with the back noise value score as the radius, and then a cone model is established with the sound attenuation value score as the height. The vertex of the cone model is selected as the center of the sphere, and a spherical model is established with the sound frequency value as the diameter. The volume of the heteromorphic model formed by the cone and the sphere is calculated and recorded as the sound judgment value; The obtained sound judgment value is compared with the preset sound judgment threshold. When the sound judgment value is greater than the preset sound judgment threshold, it is judged that the environmental parameters when the sound feature is collected meet the standard. Otherwise, the environmental parameters when the sound feature is collected are optimized and the sound feature is collected again. The specific operation steps of the image creation module are as follows: Construct a convolutional neural network model, train on facial images, learn the shape, proportion, position relationship of facial features and skin texture features, and output facial feature vectors; use a human posture recognition model based on the convolutional neural network model to process full-body photos, extract body proportions, limb morphology, and body posture features, and form a full-body feature vector; Use recurrent neural network or long short-term memory network to analyze speech samples, extract fundamental frequency, formant, pitch change, speech rate change, prosodic features and emotional color, and generate speech feature vectors; Adopt adaptive feature extraction strategies based on the characteristics of different modal data; for facial images with expression changes greater than the preset standard, increase the depth of the convolutional neural network model to enhance feature learning capabilities; for dialects or accents in speech samples, optimize the model's parameter initialization method to improve the accuracy of feature extraction; Build a basic digital human model based on the statistical average human model or the 3D scanned standard human model, and initialize the model parameters according to the basic characteristics of the user, including bone structure, muscle distribution, and basic skin texture; The extracted multimodal feature vectors are fused by weighted summation or vector concatenation to form a comprehensive feature vector. Generative adversarial network technology is used to build a model architecture including a generator and a discriminator. The generator uses the comprehensive feature vector as input and generates a personalized digital human image through a multi-layer neural network, including facial features and body shape that matches the body shape. The discriminator receives the generated digital human image and real human image data and determines the degree of match with user needs; Through adversarial training between the generator and the discriminator, the parameters of the generator are optimized to improve the quality and personalization of the generated digital human image; In the adversarial network training process, an optimization strategy is adopted; an attention mechanism is introduced to make the generator pay attention to the user's personalized characteristics when generating digital human images, including highlighting the user's facial identification or clothing details; Using multi-scale training methods, the model can learn the characteristics of digital human images at different resolutions, improving the detail richness and overall coordination of the generated image. At the same time, the transfer learning technology is used to transfer the model parameters pre-trained on large-scale general data sets to the model, speeding up the training convergence and improving the model performance. The specific operation steps of the image creation module to determine the matching degree with the user's needs are as follows: Conduct appearance feature matching evaluation and style feature quantitative comparison; Feature vector distance measurement: The appearance features in the user's requirements and the corresponding features of the generated digital human image are converted into feature vectors respectively; for the shape of the facial features, vectors are constructed by the coordinates of the key points; the body proportions are constructed by the proportional values; and then the cosine similarity between the two is calculated: ,in is the user demand vector, Generate a digital human vector; and use cosine similarity as the standard to measure the distance between the digital human image and the user demand feature vector, recorded as the special evaluation value; Attribute matching establishes matching rules and scoring systems for hairstyle, skin color, and clothing style attributes. Specifically, hairstyle is judged by hairstyle type, length range, and bangs style attributes. A complete match is scored as 1 point, a partial match is scored as 0.5 points, and a mismatch is scored as 0 points, which are recorded as the hairstyle score. Skin color is calculated in the Lab color space based on the distance from the user's desired skin color. , where and They represent the skin color value expected by the user and the skin color value of the generated digital person respectively; and They represent the green to red color component attribute values expected by the user and the green to red color component attribute values of the generated digital person respectively; and Represent the blue to yellow color component attribute values expected by the user and the blue to yellow color component attribute values of the generated digital person respectively; preset several distance intervals, calculate the distance interval corresponding to the distance d of the user's required skin color in the Lab color space according to the calculated skin color, determine the score, and record it as the skin color score; the clothing style is judged by the predefined style classification model, and the matching probability is higher than 0.8 and gets 1 point, 0.5-0.8 gets 0.5 points, and lower than 0.5 gets 0 points, which is recorded as the clothing score; after normalizing the hairstyle score, skin color score and clothing score, they are respectively recorded as clothing scores; establish a triangle with the hairstyle score, skin color score and clothing score as the three sides of the triangle, calculate the area of the triangle, and record it as the attribute score, and use this attribute score as the result of measuring attribute matching; Quantification of style features: Quantify different styles, calculate the Mahalanobis distance between the clothing color of the generated digital person and the color of the corresponding quantified style clothing, and record it as the style evaluation value; Emotional feature matching evaluation: Based on the theory of emotional psychology, the lively and calm emotional features are associated with facial expressions and gestures and quantified; the cosine similarity between the generated digital human emotional feature vector and the user demand emotional feature vector is calculated and recorded as the emotional evaluation value; The obtained special evaluation value, attribute score, reputation value and sentiment evaluation value are marked as TR, SY, FO and QU respectively, and after normalization, they are put into the following formula: To get the comprehensive judgment score DFQ, The preset weight coefficients are respectively the special evaluation value, attribute score, reputation value and emotional evaluation value. The obtained comprehensive judgment score DFQ is compared with the preset comprehensive judgment score threshold. When the comprehensive judgment score DFQ is greater than the preset comprehensive judgment score threshold, it is judged that the matching degree between the generated digital human image and the user demand meets the standard. Otherwise, data backtracking is performed to determine the non-compliant part.
2. The personalized digital human image creation and generation system based on deep learning according to claim 1 is characterized in that: The image creation module performs data backtracking and determines the specific operation steps of the non-compliant parts as follows: The special evaluation value, attribute score, moral evaluation value and sentiment evaluation value obtained from the analysis are determined separately to determine which parameters lead to the overall matching degree not meeting the standard; When the special evaluation value does not meet the standard, check whether the collected user appearance image is clear, whether the multiple angles are complete, and whether there are errors in the feature extraction process; check whether the image acquisition device parameter settings are correct, including camera resolution and focus; check whether the feature extraction algorithm does not meet the standard for processing certain special appearance features; When the attribute score does not meet the standard, the corresponding generation module is adjusted for the attribute that does not meet the standard; including: analyzing the parameters and algorithms for skin color generation, re-adjusting the probability distribution or color space conversion method for skin color generation; enriching the training data of clothing styles, and adding parameters for clothing style control in the generator to ensure that clothing styles that meet user needs can be generated; When the style rating does not meet the standard, rebuild or train the style quantification model; collect more sample data of different styles, re-extract and quantify the features of these styles; use unsupervised learning methods, including clustering algorithms, to assist in discovering the intrinsic features of different styles, so that the style quantification model can more accurately reflect various styles; adjust the style generation mechanism; add a style control vector to the input of the generator, and control the style of the generated digital human by adjusting the value of the vector; at the same time, during the training process, add a style loss function so that the generator pays more attention to the generation of style features when generating the digital human image, reducing the deviation from the style required by the user; When the emotion rating value does not meet the standard, the emotion generation module is improved; the parameters of the emotion-related neural network layer in the generator are adjusted, including adjusting the parameters related to facial muscle movements according to the emotional characteristics of user needs when generating facial expressions; at the same time, during the training process, the supervisory signal of the emotional characteristics is added so that the generator can better learn and generate digital human images that meet the emotional characteristics of user needs.
3. The personalized digital human image creation and generation system based on deep learning according to claim 1 is characterized in that: The specific operation steps of the model storage module are as follows: Establish a model storage database, adopt a distributed storage architecture, store the generated digital human model in a binary format or a custom structured format in the database, and associate the user's identity information, collected data features, and model training parameter metadata for subsequent query, call, and management; According to different application scenarios and user needs, the stored models are classified and managed, and a corresponding indexing mechanism is established to improve the efficiency of model retrieval; Regularly monitor the stored digital human models to check the integrity and accuracy of the models; at the same time, optimize the performance of the model storage database, including data compression and index reconstruction operations, to ensure the efficiency of system operation.
4. The personalized digital human image creation and generation system based on deep learning according to claim 1 is characterized in that: The specific operation steps of the matching module are as follows: When receiving a new user image creation requirement, the facial and full-body images collected by the multivariate data acquisition module are processed to obtain a new appearance feature vector; for each model in the storage database, its similarity with the new feature vector is calculated. The specific process is as follows: Feature vector normalization: normalize the new feature vector and the model feature vector so that each feature dimension has the same weight and scale; Distance metric selection: Cosine similarity is used to calculate the distance between two feature vectors, using the formula: ,in and are the new feature vector and the model feature vector respectively, and The magnitude of a vector; Similarity evaluation: According to the calculated distance value, it is converted into a similarity score. The smaller the distance, the higher the similarity. The similarity threshold is set. When the similarity score is higher than the threshold, it is considered that the matching degree between the model and the user needs meets the standard. Then the model is retrieved for fine-tuning. The specific process is as follows: According to the difference analysis of the feature vectors in the similarity calculation process, the difference between the feature parameters of the current model and the new requirements is quantified to a value between 0 and 1 and recorded as the difference judgment value. The difference judgment value is compared with the preset difference judgment threshold. When the difference judgment value is greater than the preset difference judgment threshold, the feature is judged to be a key parameter; Adjust the weights of key parameters; Weight adjustment is performed based on preset adjustment rules or adjustment strategies learned through machine learning algorithms; Use the gradient-based method to perform incremental updates. For the determined key parameters, calculate the gradient of the loss function relative to the new requirement. Let the loss function L represent the difference between the digital human image generated by the model and the appearance features of the new requirement. For the key parameters ,calculate , which represents the loss function for the L model parameters The partial derivative of; then, according to the principle of the gradient descent algorithm, with the learning rate Update the parameters: ; In the incremental update process, an adaptive learning rate strategy is adopted; including the use of Adagrad, Adadelta or Adam adaptive learning rate optimization algorithms; the learning rate is automatically adjusted according to the historical gradient information of the parameters, so that when updating key parameters, the update step size can be dynamically adjusted according to the importance and update history of the parameters; When the similarity score is lower than the threshold, it is determined that the model needs to be rebuilt.
5. The personalized digital human image creation and generation system based on deep learning according to claim 1 is characterized in that: The specific operation steps of the rendering interaction module are as follows: When the user issues a request for retrieval, the rendering interaction module first receives the request for retrieval; the request contains the user's relevant information on the application scenario of the digital human; according to the user's retrieval demand, the model is obtained from the model storage module or the matching module; When the user has no new requirements, the pre-stored model that meets the user's basic requirements is directly selected from the model storage module; if the user has new requirements and the matching module has fine-tuned the model, the fine-tuned model in the matching module is obtained; After obtaining the model, analyze the lighting environment of the virtual scene where the digital human is located; including determining the position, intensity, color and type of the light source; setting it according to the material properties of the digital human model itself; material properties include the roughness, glossiness, transparency of the skin, and the texture and feel of the clothes; for the skin, set different roughness values to simulate the texture of real skin; for clothes, load pre-designed texture patterns to express the material of the clothes; Using physically based rendering technology combined with ray tracing algorithms, the physical phenomena of light propagation, reflection, refraction and scattering are calculated according to scene lighting and model material properties, giving digital humans realistic light and shadow performance; Monitor the position of the digital human in the virtual scene in real time and calculate the distance between the digital human and the viewing angle; dynamically adjust the model's detail level based on the distance between the digital human and the scene and the user's focus; when the digital human is in the distance, use the preset low-detail model to reduce the amount of rendering calculations and improve rendering efficiency; When the digital human is in the focus of the scene, it switches to the preset high-detail model to show the facial expressions and clothing texture details to ensure high-quality visual effects; Detect the graphics processing capabilities of the user's device, including GPU model, video memory size and hardware parameters; Automatically adjust rendering parameters according to hardware performance.
Citation Information
Patent Citations
Interactive digital human generation method and system based on artificial intelligence
CN119206001A