Desktop digital human interactive studio method based on JavaFX
Through depth camera scanning and JavaFX technology, an accurate digital human facial model is established and the synchronization of phonemes and lip animations is achieved, which solves the problem of insufficient authenticity and nature of digital human lip animations, and realizes efficient coordinated control of lip and tongue.
Patent Information
- Application Number
- CN202411679293.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-11-22
AI Technical Summary
The prior art is difficult to accurately capture and reproduce the subtle changes in the human face during pronunciation, especially the coordinated movement of lips, tongue and teeth, resulting in insufficient authenticity and naturalness of digital human lip animations.
Three-dimensional shape data is obtained by scanning the real person's face through a depth camera, and combining medical anatomy knowledge to establish an accurate digital human facial model. JavaFX is used to synchronize phonemes and lip animations, optimize lip control parameters through comparative learning, build a tongue control model and tooth visibility control rules, and realize coordinated control of lip and tongue.
It significantly improves the authenticity and nature of digital human lip animations, making it closer to the characteristics of real-person pronunciation, and provides key technical support for the interaction between virtual anchors and digital humans.
Smart Images

Figure CN119937773A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to a desktop digital human interactive broadcasting method based on JavaFX. Background Art
[0002] In the study of interactive broadcasting methods for desktop digital humans based on JavaFX, building a high-precision digital human face model and realizing realistic broadcast lip animation is a complex technical challenge. The core issue is how to accurately capture and reproduce the subtle changes in the human face during pronunciation, especially the coordinated movement of the lips, tongue and teeth. This involves accurate modeling of multiple dimensions such as geometric shapes, texture features, and anatomical structures. At the same time, it is also necessary to establish an accurate correspondence between phonemes and lip shape changes to achieve natural synchronization of lip animation and speech. However, the movement of facial muscles when humans pronounce is extremely complex, and the transition between different phonemes is also very subtle. How to ensure smooth and natural animation while reflecting the subtle differences in the pronunciation of different phonemes is a huge challenge. Solving these technical problems is crucial to improving the interactive experience and application value of digital humans. Summary of the invention
[0003] The present invention provides a desktop digital human interactive presentation method based on JavaFX, which mainly includes:
[0004] Scan the real face with a depth camera to obtain the three-dimensional shape data of facial tissues, segment and annotate the data with medical anatomical knowledge, and build a digital human face model, including the geometric shape and texture features of the facial tissue structure of the lips, tongue, and teeth;
[0005] Obtain the lip shape changes of different phonemes in the digital human broadcast, establish amplitude, speed and trajectory templates, use JavaFX animation and timeline functions to obtain the lip shape parameter sequence corresponding to the sound sequence, and realize the synthesis of the broadcast lip shape animation synchronized with the phonemes by controlling the lip shape changes of the digital human face model to obtain the broadcast lip shape animation;
[0006] According to the obtained digital human lip animation, a phoneme-to-lip parameter mapping model is trained to obtain a digital human lip animation with high lip control accuracy and natural and coordinated changes. The similarity between the digital human lip animation and the real person's lip shape is judged through comparative learning. If the similarity is lower than the preset threshold, the lip control parameters are fine-tuned to make the lip changes during the broadcast closer to the real person, improve the visual realism of the lip animation, and obtain the optimized broadcast lip animation;
[0007] According to the optimized digital human lip animation, the lip shape feature vector of the key frame is extracted to obtain the lip shape change template representing different phonemes in the state of pronunciation. The template is used to guide the generation of lip shape animation to ensure the coherence of the lip shape change of the digital human when performing different phonemes.
[0008] According to the lip shape change template, a tongue position control model is constructed to obtain the tongue position parameters of different phonemes when the digital human is performing pronunciation. The shape and position of the digital human tongue are controlled by the parameters to achieve the synchronization of the tongue position animation and the phonemes during the performance, and obtain the mapping relationship between the tongue position and the phonemes during the performance.
[0009] According to the mapping relationship between tongue position and phonemes in broadcasting, the teeth visibility control rules are established. According to the teeth visibility parameters of the sequence, the visibility of the teeth of the digital human in broadcasting is dynamically adjusted to make the changes in teeth visibility match the characteristics of the pronunciation of the real person in broadcasting.
[0010] The shape and texture features of the digital human's lips, tongue and teeth are tracked in real time during the broadcast, compared with the target broadcast lip shape, tongue position and teeth visibility parameters, and the digital human facial model parameters are dynamically adjusted to perform coordinated control of the lips, tongue and teeth during the broadcast.
[0011] The technical solution provided by the embodiment of the present invention may have the following beneficial effects:
[0012] The present invention discloses a desktop digital human interactive broadcasting method based on JavaFX. A depth camera is used to scan a real person's face to obtain three-dimensional shape data, and an accurate digital human facial model is established in combination with medical anatomical knowledge. JavaFX is used to achieve synchronization between phonemes and lip shape animation, and lip shape control parameters are optimized through comparative learning. Key frame feature vectors are extracted to construct a lip shape change template to guide the generation of broadcast lip shape animation. At the same time, a refined tongue position control model and tooth visibility control rules are established to achieve precise synchronization of tongue position and tooth animation with phonemes. The facial model parameters are dynamically adjusted to achieve refined coordinated control of lips, tongue and teeth. The present invention significantly improves the authenticity and naturalness of digital human broadcast lip animation, and provides key technical support for applications such as virtual anchors and digital human interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 The present invention is a flowchart of a JavaFX-based desktop digital human interactive presentation method.
[0014] Figure 2 The figure is a schematic diagram of a JavaFX-based desktop digital human interactive presentation method of the present invention.
[0015] Figure 3 It is another schematic diagram of a JavaFX-based desktop digital human interactive presentation method of the present invention. DETAILED DESCRIPTION
[0016] In order to further understand the content of the present invention, the present invention is described in detail in conjunction with the accompanying drawings and embodiments. The present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It is understood that the specific embodiments described herein are only used to explain the relevant inventions, rather than to limit the invention. It is also necessary to explain that, for ease of description, only the parts related to the invention are shown in the accompanying drawings.
[0017] like Figure 1-3 In this embodiment, a JavaFX-based desktop digital human interactive presentation method may specifically include:
[0018] Step S101, scan the real person's face with a depth camera to obtain the three-dimensional shape data of the facial tissue, subdivide and annotate the data in combination with medical anatomical knowledge, and establish a digital human face model, including the geometric shape and texture features of the facial tissue structure of the lips, tongue, and teeth.
[0019] The method comprises the following steps: obtaining point cloud information carrying multi-angle three-dimensional scanning data of a real face, wherein the point cloud information is obtained by collecting facial tissues using a non-contact 3D laser scanner; determining the boundary areas of the lip, tongue, and tooth facial tissues according to the point cloud information and a pre-established medical anatomical knowledge base; converting the point cloud information into a voxel grid structure using an octree voxelization method to obtain voxelized data of the facial tissues; using a region growing algorithm to segment the lip, tongue, and tooth regions according to the voxelized data, and applying a Loop subdivision algorithm to the lip region if the lip region is segmented; applying a deformation field simulation to the tongue region if the tongue region is segmented; and applying a Machi cube algorithm to the tooth region if the tooth region is segmented; receiving facial tissue data including texture image information and feature descriptors, wherein the feature descriptors are obtained by calculating the curvature and normal vector geometric features of the facial tissues and constructing them using a point feature histogram; and performing fine segmentation and annotation of the facial tissues using a PointNet++ network according to the facial tissue data to obtain a digital human facial model including the lip, tongue, and tooth facial tissues.
[0020] Exemplarily, a non-contact 3D laser scanner is used to perform a multi-angle 3D scan of a real face to obtain point cloud data of facial tissues including lips, tongue, and teeth, and the complete facial 3D shape information is obtained by merging through a multi-view registration algorithm. The acquired facial 3D point cloud data is subdivided according to medical anatomical knowledge to divide the boundary areas of facial tissues such as lips, tongue, and teeth, and the octree voxelization method is used to convert the point cloud data into a voxel grid structure. Based on the voxelized data, the lip, tongue, and tooth regions are segmented using a regional growing algorithm, the Loop subdivision algorithm is used to increase the geometric details of the lips, the deformation field is used to simulate the soft characteristics of the tongue, and the March cube algorithm is used to reconstruct the precise geometric shape of the teeth. The feature extraction of the voxelized facial tissue structure is performed, and the geometric features such as the curvature and normal vector of each facial tissue are calculated, and the feature descriptor of the facial tissue is constructed using a point feature histogram (PFH). Combining texture image information and feature descriptors, the PointNet++ network is used to perform fine segmentation and annotation of facial tissues, and a high-precision three-dimensional geometric model of facial tissues such as lips, tongue, and teeth is constructed. The segmentation results are combined with the original texture image using the pix2pix generative adversarial network to generate a high-precision texture map and map it to the model surface.
[0021] When constructing a high-precision digital human face model, a non-contact 3D laser scanner with a resolution of 0.1mm was first used to perform a 360-degree full-scale scan of the real face. The scanning angle interval for each scan was 30 degrees, and a total of 12 sets of point cloud data were obtained. The iterative closest point algorithm (ICP) was used for point cloud registration, and the registration error threshold was set to 0.05mm to achieve accurate stitching of multi-view point clouds. The octree voxelization algorithm was applied to the merged point cloud data, and the voxel size was set to 0.5mm to convert the continuous point cloud into a discrete voxel structure. The density-based region growing algorithm was then used to segment the facial tissue, and the growth threshold was set to 0.8mm, and the lip, tongue and tooth areas were successfully extracted. The Loop subdivision algorithm was applied to the lips for 3 iterations to increase the number of triangular facets by 4 times, significantly improving the geometric details. The tongue was simulated by a deformation field based on a mass-spring system, and the elastic coefficient was set to 0.5N / m to achieve soft characteristics. The tooth area was reconstructed using the Machi cube algorithm with a grid resolution of 0.1mm to accurately restore the sharp edges of the teeth. In the feature extraction stage, the average curvature and Gaussian curvature of each vertex are calculated, and the normal vector is estimated using the K nearest neighbor method (K=10). A 153-dimensional feature descriptor is constructed using a point feature histogram (PFH) with a radius of 3 mm. These features are input into the PointNet++ network, which contains 4 layers of Set Abstraction layers and 2 layers of Feature Propagation layers to achieve semantic segmentation of the point cloud. Finally, the pix2pix generative adversarial network is used. Both the discriminator and the generator use the U-Net structure, with the original texture image and segmentation label map of 256x256 resolution as input, and a fine texture map of the same resolution is output. The generated texture is accurately mapped to the surface of the three-dimensional model through parametric mapping, completing the construction of a high-precision digital human face model.
[0022] Step S102, obtain the lip shape changes of different phonemes in the digital human performance, establish amplitude, speed and trajectory templates, use the animation and timeline functions of JavaFX to obtain the lip shape parameter sequence corresponding to the sound sequence, and realize the synthesis of the performance lip shape animation synchronized with the phonemes by controlling the lip shape changes of the digital human facial model to obtain the performance lip shape animation.
[0023] A high-speed camera is used to obtain a lip shape change sequence corresponding to different phonemes during the pronunciation of a real person, and the lip shape change sequence contains image data of 240 frames per second. Edge detection is performed according to the lip shape change sequence to obtain the lip contour, and feature points are extracted from the lip contour to obtain key point coordinates. The key point coordinates are fitted using the least squares method to obtain a lip shape change curve, and the amplitude, speed and acceleration parameters of the lip shape change are calculated according to the lip shape change curve. If the lip shape change parameters are obtained, a phoneme lip shape change template library is constructed; wherein the phoneme lip shape change template library includes the amplitude, speed and trajectory template corresponding to each phoneme. A lip shape parameter sequence is read from the phoneme lip shape change template library, and the lip shape parameter sequence includes key point coordinates, change amplitude and speed information; the radial basis function interpolation method is used to map the lip shape parameter sequence to the vertex coordinates of the digital human face model, and the change of the vertex coordinates is controlled by the animation function to obtain a lip shape animation synchronized with the phoneme.
[0024] For example, a high-speed camera is used to shoot the pronunciation process of a real person at a rate of 240 frames per second to obtain the lip shape change sequence corresponding to different phonemes, the lip contour is extracted by the Canny edge detection algorithm, and then the SURF feature point detection algorithm is applied to obtain the key point coordinates. The lip shape change curve is fitted by the least squares method to calculate the amplitude, speed and acceleration parameters of the lip shape change. According to the extracted lip shape change parameters, a phoneme lip shape change template library is constructed, and a refined amplitude, speed and trajectory template is established for each phoneme, which is stored as a JSON format file to improve reading efficiency. Using the timeline function of JavaFX, the sequence data is read and the hidden Markov model is used for phoneme recognition. The dynamic time warping algorithm is used to achieve the time alignment of the phoneme and the lip shape template, and the corresponding lip shape parameter sequence is generated, including the key point coordinates, change amplitude and speed information. Through the animation function of JavaFX, the radial basis function interpolation method is used to map the generated lip shape parameter sequence to the vertex coordinates of the digital human face model, the JavaFX KeyFrame and KeyValue classes are used to set the key frames, and the animation playback is controlled by the Timeline class to achieve the lip shape animation synthesis synchronized with the phonemes.
[0025] When realizing the lip shape animation synthesis of digital human, a high-speed camera with a frame rate of 240fps is first used to shoot the pronunciation process of real people, and a pronunciation video sequence containing 20 basic phonemes is collected. The Canny edge detection algorithm is applied to each frame of the image, and the double thresholds are set to 50 and 150 to extract the lip contour. Then, the SURF feature point detection algorithm is used, and the Hessian matrix threshold is set to 300 to extract 64 key points in the lip area. These key points are fitted by the least squares method to obtain the lip shape change curve, and the amplitude, velocity and acceleration parameters of the curve are calculated. Lip shape change templates are established for the 20 basic phonemes respectively. Each template contains 10 key frames, and each frame stores the coordinates, amplitude and velocity information of 64 key points and saves them in JSON format. When processing the input voice data, a 5-state hidden Markov model is used for phoneme recognition, and the recognition accuracy rate reaches 95%. The dynamic time warping algorithm is used, and the window size is set to 5 to achieve time alignment of phonemes and lip shape templates and generate a complete lip shape parameter sequence. In JavaFX, the radial basis function interpolation method is used to select 50 control points and map the lip shape parameters to a digital human face model containing 5,000 vertices. The keyframes are set through the JavaFX KeyFrame class, with 60 frames per second, the KeyValue class defines the interpolation of vertex coordinates, and the Timeline class controls the 30-second animation clip to achieve smooth lip shape animation synthesis. The entire process from video acquisition to animation generation takes about 2 minutes on a processor with a main frequency of 3.5GHz, and the size of the generated animation file is about 20MB.
[0026] Step S103, based on the obtained digital human broadcast lip shape animation, train the phoneme to lip shape parameter mapping model to obtain the digital human broadcast lip shape animation with high lip shape control accuracy and natural and coordinated changes, and judge the similarity between the digital human broadcast lip shape animation and the real person's pronunciation lip shape through comparative learning. If the similarity is lower than the preset threshold, fine-tune the lip shape control parameters to make the lip shape changes during broadcast closer to the real person, improve the visual realism of the lip shape animation, and obtain the optimized broadcast lip shape animation.
[0027] Acquire the phoneme sequence, use the time convolution network to build the phoneme to lip shape parameter mapping model, and obtain the optimized model parameters through the back propagation algorithm. Receive the real-person pronunciation video, use the 3D deformation model to reconstruct the 3D facial model in the video, extract the key features of lip shape contour, opening and closing degree, and lip corner position, and establish the real-person pronunciation lip shape feature database. According to the digital human broadcast lip shape sequence and the real-person pronunciation lip shape sequence, determine the time alignment relationship through the dynamic time warping algorithm, and judge the visual similarity of each frame in combination with the structural similarity index. Calculate the similarity score between the digital human broadcast lip shape animation and the real-person pronunciation lip shape feature. If the similarity score is lower than the preset threshold, enter the parameter fine-tuning process. Use the genetic algorithm to fine-tune the lip shape control parameters. The genetic algorithm includes chromosome encoding of the lip shape control parameters; generating new parameter combinations through roulette selection, single-point crossover and Gaussian mutation operations; iterative optimization with the similarity score as the fitness function; when the similarity score exceeds the preset threshold or reaches the maximum number of iterations, output the optimized broadcast lip shape animation.
[0028] Exemplarily, according to the digital human lip animation and the corresponding phoneme sequence, a mapping model from phonemes to lip parameters is constructed using a temporal convolutional network, and the model parameters are optimized by a back-propagation algorithm to obtain a digital human lip animation with high lip control accuracy and natural and coordinated changes. The 3D facial model in the real-person pronunciation video is reconstructed using a 3D deformation model, and key features such as lip contour, opening and closing, and lip corner position are extracted to establish a real-person pronunciation lip feature database. The time alignment of the digital human and real-person lip sequences is calculated using a dynamic time warping algorithm, and the visual similarity of each frame is evaluated in combination with the structural similarity index. The digital human lip animation and the real-person pronunciation lip features are compared to calculate the similarity score between the two. If the similarity is lower than the preset threshold of 0.85, the parameter fine-tuning process is entered. Genetic algorithm is used to fine-tune the lip shape control parameters. The chromosome encoding represents the numerical range of the lip shape control parameters. New parameter combinations are generated through roulette wheel selection, single-point crossover and Gaussian mutation operations. The similarity score is used as the fitness function. Through iterative optimization, the lip shape changes during broadcasting are made closer to real people. When the similarity exceeds the preset threshold or reaches the maximum number of iterations, the optimized broadcast lip animation is output.
[0029] In the process of optimizing the lip animation of digital human, a phoneme-to-lip shape parameter mapping model containing an 8-layer temporal convolutional network was first constructed, with 40-dimensional phoneme features as input and 68 lip key point coordinates as output. The Adam optimizer was used with a learning rate of 0.001, and 50 epochs were trained on 100,000 phoneme-lip shape sample pairs to obtain a high-precision lip shape control model with an average error of less than 0.5 mm. Subsequently, a 3D deformable model was used to reconstruct the 3D face of the real person pronunciation video, using 199 shape parameters and 40 expression parameters to accurately capture the lip shape changes. 68 lip key points were extracted, and 10 feature parameters such as opening and closing degree and lip corner position were calculated to construct a real person lip shape feature database containing 50,000 frames. The dynamic time warping algorithm was used with a window size of 15 frames to calculate the optimal alignment path of the digital human and real person lip shape sequences. Combined with the structural similarity index, the weights were set to 0.3 for brightness, 0.3 for contrast, and 0.4 for structure to evaluate the visual similarity of each frame. If the average similarity is lower than 0.85, the genetic algorithm is started to fine-tune the parameters. In the genetic algorithm, the chromosome encoding uses 10 floating-point numbers to represent the lip shape control parameters, the population size is set to 100, and roulette wheel selection, single-point crossover with a crossover probability of 0.8 and Gaussian mutation with a mutation probability of 0.1 are used. Iteration optimization is performed for 50 generations or until the similarity exceeds 0.9, and the optimized lip animation is finally output to make the lip shape changes of the digital human highly consistent with the pronunciation of the real person.
[0030] Step S104, based on the optimized digital human lip animation, extract the lip shape feature vector of the key frame, obtain the lip shape change template representing different phonemes in the broadcast pronunciation state, and use the template to guide the generation of the broadcast lip shape animation to ensure the continuity of the lip shape change of the digital human when broadcasting different phonemes.
[0031] The video sequence of lip animation is obtained, and the key frames are extracted by using the motion amplitude calculation method based on the optical flow method. The extraction threshold is set according to 1.5 times the average motion amplitude. The lip shape feature vector is processed by the hierarchical clustering algorithm, and Ward's minimum variance method is used as the link criterion. The lip shape change template representing different phonemes is obtained by calculating the class center. The mapping relationship between phonemes and lip shape change templates is established, and the optimal path search is performed based on the continuous phoneme sequence using the dynamic programming algorithm. The state is defined as the lip shape template corresponding to the current phoneme. For the selected lip shape change template sequence, a multidimensional cubic spline interpolation algorithm is used to generate a smooth lip shape transition curve, and each dimension of the lip shape feature is interpolated, and the interpolation result is mapped to the digital human face model. The generated lip shape sequence is processed by the Kalman filter to eliminate the mutation and jitter in the lip shape sequence, and a coherent and natural live lip shape animation is obtained.
[0032] Exemplarily, according to the optimized digital human lip animation, the motion amplitude calculation based on the optical flow method is used to extract key frames, the threshold is set to 1.5 times the average motion amplitude, and the low-dimensional lip shape feature vector of each key frame is obtained by dimensionality reduction processing using the principal component analysis method. The extracted lip shape feature vectors are grouped using a hierarchical clustering algorithm, and the Ward's minimum variance method is used as the link criterion. The number of clusters is automatically determined according to the number of phonemes, and the lip shape change template representing different phonemes in the state of broadcast pronunciation is obtained by calculating the class center. A mapping relationship between phonemes and lip shape change templates is established, and a dynamic programming algorithm is used to search for the optimal path for continuous phoneme sequences. The state is defined as the lip shape template corresponding to the current phoneme. The transfer equation considers the context information of the previous and next phonemes and selects the most appropriate lip shape change template sequence. According to the selected lip shape change template sequence, a multidimensional cubic spline interpolation algorithm is used to generate a smooth lip shape transition curve, and each dimension of the lip shape feature is interpolated respectively, and a time weight is added to adjust the transition speed, and the interpolation result is mapped to the digital human face model. The generated lip shape sequence is filtered by Kalman filter to eliminate mutations and jitters, and generate coherent and natural lip shape animation. In the process of optimizing the lip shape animation of digital human, the Horn-Schunck optical flow algorithm is first used to calculate the motion amplitude between adjacent frames, and the threshold is set to 1.5 times the average motion amplitude to extract 200 key frames. The principal component analysis is applied to the coordinates of the 68 lip key points of each key frame, and the first 20 principal components that explain 95% of the variance are retained to obtain a 20-dimensional lip shape feature vector. Subsequently, Ward's minimum variance method is used for hierarchical clustering, and the number of clusters is set to 40. Assuming that there are 40 phonemes, 40 lip shape change templates are obtained. After establishing the mapping relationship between phonemes and lip shape templates, the Viterbi algorithm is used to implement dynamic programming. The state space is 40 lip shape templates, the transition probability is based on the phoneme co-occurrence frequency, and the observation probability is determined by the phoneme recognition confidence. For a 10-second speech clip containing about 300 phonemes, the algorithm completes the optimal path search within 0.1 seconds. After selecting the lip template sequence, the cubic spline interpolation algorithm is used to interpolate the 20-dimensional feature vector, and the number of interpolation points is three times that of the original sequence. The time weight function w(t) = sin(πt / 2) is introduced to adjust the transition speed. Finally, the Kalman filter is applied, and the measurement noise covariance R is set to 0.01 and the process noise covariance Q is set to 0.001. The generated lip sequence is smoothed to eliminate high-frequency jitter and obtain the final coherent and natural live lip animation.
[0033] Step S105, construct a tongue position control model based on the lip shape change template, obtain the tongue position parameters of different phonemes when the digital human is performing pronunciation, control the changes in the shape and position of the digital human's tongue through the parameters, realize the synchronization of the tongue position animation and the phonemes during the performance, and obtain the mapping relationship between the tongue position and the phonemes during the performance.
[0034] The digital human tongue position control model is constructed by NURBS surface modeling technology, and the control point grid is set according to the lip shape change template; the bidirectional LSTM network is used to perform phoneme recognition on the broadcast voice to obtain the phoneme sequence and timestamp information; the tongue movement is simulated by a physical simulation method based on a mass spring system, and if the lip shape changes, the tongue position parameters are optimized according to the phoneme sequence; the deep deterministic policy gradient algorithm is used to construct the Actor-Critic network structure, and the tongue position control strategy is trained according to the naturalness of the tongue position animation and the phoneme synchronization; the tongue position parameters are interpolated using the cubic Hermite interpolation algorithm, and the tongue position parameter sequence is obtained by adding interpolation points.
[0035] For example, according to the lip shape change template, the NURBS surface modeling technology is used to build a digital human tongue model, a 30x20 control point grid is set, the tongue contour is described by a spline curve, and the tongue position change is parameterized by the control points to obtain a refined tongue position control model. The bidirectional LSTM network is used to perform phoneme recognition on the broadcast voice, extract the phoneme sequence and its timestamp information, and combine the tongue position description in the phonetics knowledge base to establish the initial mapping relationship between phonemes and tongue position parameters. Through the physical simulation method based on the mass spring system, the tongue movement is simulated according to the lip shape change and the phoneme sequence, the tongue position parameters are optimized, and the precise synchronization of the tongue position animation and the phoneme pronunciation is achieved. The deep deterministic policy gradient algorithm is used to construct the Actor-Critic network structure, and the naturalness of the tongue position animation and the phoneme synchronization are used as the reward function to train the tongue position control strategy and optimize the mapping relationship between the tongue position and the phoneme in the broadcast. The tongue position parameters are interpolated using the cubic Hermite interpolation algorithm, and 100 interpolation points are added to improve the continuity and naturalness of the tongue movement, and the final refined tongue position parameter sequence is obtained.
[0036] In the process of building the digital human tongue position control model, the NURBS surface modeling technology is first used to create a 3D model of the tongue, and a 30x20 control point grid is set. The surface order is 3, and the tongue shape is accurately described by adjusting the weights and positions of 600 control points. Subsequently, a 4-layer bidirectional LSTM network is used for phoneme recognition. Each layer contains 256 hidden units, the input is a 40-dimensional Mel spectrum feature, and the output is a phoneme probability distribution. The recognition accuracy rate reaches 95%. Based on the recognition results, a tongue position parameter mapping table containing 44 phonemes is established, and each phoneme corresponds to 20 tongue position control parameters. Next, a mass spring system is used to simulate the movement of the tongue. The tongue is divided into 500 mass points, 1500 spring connections are set, the elastic coefficient is 100N / m, the damping coefficient is 0.5Ns / m, and the time step is 0.001s. The tongue deformation is calculated by numerical integration. In the optimization stage, the DDPG algorithm is used to train the tongue position control strategy. The Actor network adopts a 3-layer fully connected structure (256-128-64), the Critic network is a 4-layer (256-256-128-1), the learning rate is set to 0.0001, the discount factor is 0.99, and a stable control strategy is obtained after 100,000 iterations. Finally, the optimized tongue position parameter sequence is interpolated to 300 frames / second based on the original 30 frames / second to ensure C1 continuity and generate smooth and natural tongue position animation. The whole process is run on a workstation configured with an RTX 3080 GPU and 32GB RAM. The optimization of the tongue position parameters of a single phoneme takes about 0.5 seconds, achieving real-time performance.
[0037] Step S106, based on the mapping relationship between tongue position and phonemes in broadcasting, establish teeth visibility control rules, and dynamically adjust the visibility of the digital human's teeth during broadcasting according to the teeth visibility parameters of the sequence, so that the changes in teeth visibility match the characteristics of real people's pronunciation during broadcasting.
[0038] The tongue position parameters and phoneme types are obtained, and the random forest algorithm is used to construct the tooth visibility control rules according to the tongue position parameters and the phoneme types to obtain the tooth visibility prediction model; the studio voice signal is received, and the phoneme recognition is performed on the studio voice signal through the convolutional neural network-long short-term memory network hybrid model to extract the phoneme sequence and its duration; according to the tooth visibility prediction model and the phoneme sequence, a tooth visibility parameter sequence corresponding to each phoneme is generated; the three-dimensional tooth model of the digital human is constructed by using the subdivision surface technology, and if the tooth visibility parameter sequence is received, the transparency of the three-dimensional tooth model is controlled by the vertex shader and the pixel shader combined with the masking technology; the tooth visibility parameter sequence is processed by cubic spline interpolation to obtain a smoothed tooth visibility parameter sequence; if the smoothed tooth visibility parameter sequence does not match the preset real person pronunciation characteristics, the dynamic time warping algorithm is used to fine-tune the smoothed tooth visibility parameter sequence to obtain a tooth visibility parameter sequence that matches the preset real person pronunciation characteristics.
[0039] Exemplarily, based on the mapping relationship between tongue position and phonemes in broadcasting, the random forest algorithm is used to construct the tooth visibility control rules, and the tongue position parameters and phoneme types are used as input features to train the tooth visibility prediction model. The convolutional neural network-long short-term memory network hybrid model is used to perform phoneme recognition on the broadcast voice, extract the phoneme sequence and its duration, and combine it with the tooth visibility prediction model to generate the tooth visibility parameter sequence corresponding to each phoneme. The subdivision surface technology is used to build a refined three-dimensional tooth model of the digital human, and the transparency of the tooth model is controlled by the vertex shader and pixel shader combined with the masking technology to achieve dynamic and fine adjustment of tooth visibility. The tooth visibility parameter sequence is smoothed by the cubic spline interpolation algorithm to eliminate the mutation phenomenon, and the transparency of the tooth model is updated in real time through the OpenGL graphics API. The generated tooth visibility sequence is fine-tuned using the dynamic time warping algorithm to match the tooth visibility changes with the characteristics of the real person's pronunciation during the broadcast. In the process of realizing the tooth visibility control of digital human, the prediction model was first constructed using the random forest algorithm. The input features included 20-dimensional tongue position parameters and 44 phoneme types. The forest was composed of 500 decision trees. The maximum depth of each tree was set to 10, the minimum number of leaf node samples was 5, and the training sample size was 100,000. The tooth visibility prediction model with an accuracy of 95% was obtained. Subsequently, a 1D-CNN and bidirectional LSTM hybrid network was used for phoneme recognition. The CNN part contained 3 convolutional layers with convolution kernel sizes of 3, 5, and 7 respectively. The LSTM layer contained 256 hidden units. The input was 40-dimensional Mel spectrum features and the output was phoneme probability distribution. The phoneme recognition accuracy was 97% after training on the LibriSpeech dataset. A high-precision tooth model was constructed using NURBS surface technology. The number of control points was 20x30x10, and the subdivision level was set to 4, generating a fine tooth model with about 1 million vertices. In the OpenGL environment, tooth transparency control is achieved through vertex shaders and pixel shaders. Alpha blending technology is used, and the blending function is set to (GL_SRC_ALPHA, GL_ONE_MINUS_SRC_ALPHA). Cubic spline interpolation is applied to the generated tooth visibility parameter sequence, and the original 30 frames / second is interpolated to 300 frames / second to ensure C2 continuity. Finally, the FastDTW algorithm is used to fine-tune the generated tooth visibility sequence, with the window size set to 20 and the tolerance threshold set to 0.1. It is compared with the reference sequence extracted from the live studio video and iterated and optimized 5 times, so that the average Euclidean distance between the generated sequence and the real-person feature is less than 0.05. The entire process is run on a workstation configured with an RTX3090 GPU and 64GB RAM, achieving real-time rendering performance of 60 frames per second.
[0040] Step S107, real-time tracking of the shape and texture features of the digital human's lips, tongue and teeth during the broadcast, comparing them with the target broadcast lip shape, tongue position and teeth visibility parameters, dynamically adjusting the digital human facial model parameters, and performing coordinated control of the lips, tongue and teeth during the broadcast.
[0041] The key point detection algorithm of the human face is used to obtain the key point information of the digital human face image, and the key point information includes the position data of the lip, tongue and teeth areas. According to the key point information obtained, the mapping relationship between the two-dimensional image features and the three-dimensional model parameters is established. If it is detected that there is a deviation between the target broadcast lip shape, tongue position and tooth visibility parameters and the actual feature parameters, the adjustment amount of the facial model control point is calculated. The adjustment amount is minimized by the iterative optimization algorithm to obtain the adjustment result of the digital human face model. According to the adjustment result, a mass spring system is introduced to connect the key control points of the lips, tongue and teeth, and the mass spring system is used to realize the coordinated control of the lips, tongue and teeth. The spring force action data is obtained from the mass spring system, and the final adjustment parameters of the digital human face model are determined according to the spring force action data.
[0042] Exemplarily, the 68-point facial key point detection algorithm of the Dlib library is used to extract the key points in the digital human face image, and the two-dimensional key points are mapped to the three-dimensional space in combination with the three-dimensional deformation model to locate the lip, tongue and tooth areas, and the shape and texture feature parameters of these areas are extracted by principal component analysis. According to the preset target broadcast lip shape, tongue position and tooth visibility parameters, the active appearance model is used to establish the mapping relationship between the two-dimensional image features and the three-dimensional model parameters, and the extracted actual feature parameters are calculated with the target parameters by Euclidean distance calculation to obtain the feature deviation value. The Levenberg-Marquardt algorithm is used to calculate the adjustment amount of the facial model control point based on the feature deviation value, and the adjustment amount is minimized by iterative optimization to achieve the fine adjustment of the digital human face model. The extended Kalman filter is used to smooth the adjustment process of the facial model, and the appropriate state transfer matrix and observation matrix are set to eliminate the jitter caused by feature extraction and parameter comparison. The mass spring system is introduced as a physical constraint model to connect the key control points of the lips, tongue and teeth, and the coordinated control of the lips, tongue and teeth is achieved through the action of the spring force to ensure the naturalness and consistency of the action. In the process of realizing the refined control of digital human face, the 68-point face key point detection algorithm of Dlib library is first used to locate the key points on the image with a resolution of 1920x1080, and the average detection time is 15ms. Subsequently, BaselFaceModel is used as a 3D deformable model, which contains 199 shape parameters and 40 expression parameters, to map the 2D key points to 3D space, and the error is controlled within 0.5mm. The shape and texture features of the lip, tongue and teeth area are extracted by principal component analysis, and the first 20 principal components that explain 95% of the variance are retained. The mapping relationship between 2D features and 3D parameters is established using the active appearance model, with 100,000 training samples and a convergence threshold of 0.001. The Euclidean distance between the actual features and the target parameters is calculated to obtain a 68-dimensional feature deviation vector. The Levenberg-Marquardt algorithm is used for optimization, with the initial damping factor λ set to 0.01, the maximum number of iterations to 50, and the convergence condition that the relative error is less than 1e-6. The extended Kalman filter is used for smoothing. The state vector dimension is 239 (199 + 40), the observation vector dimension is 68, and the process noise covariance Q and observation noise covariance R are set to 0.01 times the unit matrix through debugging. A mass spring system is introduced, with 20 mass points set on the lips, 10 on the tongue, and 8 on the teeth. The spring stiffness coefficient k = 100 N / m and the damping coefficient c = 0.5 Ns / m. The entire process runs on the RTX3090 GPU, achieving 60fps real-time performance, the facial model control accuracy reaches 0.1mm, and the similarity score with real-life performance reaches 0.95, with a full score of 1.
[0043] The above is only a preferred embodiment of the present invention. It should be pointed out that ordinary technicians in this technical field can make several improvements and supplements without departing from the principle of the present invention. These improvements and supplements should also be regarded as the scope of protection of the present invention.
Claims
1. A desktop digital human interactive presentation method based on JavaFX, characterized in that: The method comprises: Scan the real face with a depth camera to obtain the three-dimensional shape data of facial tissues, segment and annotate the data with medical anatomical knowledge, and build a digital human face model, including the geometric shape and texture features of the facial tissue structure of the lips, tongue, and teeth; Obtain the lip shape changes of different phonemes in the digital human broadcast, establish amplitude, speed and trajectory templates, use JavaFX animation and timeline functions to obtain the lip shape parameter sequence corresponding to the sound sequence, and realize the synthesis of the broadcast lip shape animation synchronized with the phonemes by controlling the lip shape changes of the digital human face model to obtain the broadcast lip shape animation; According to the obtained digital human lip animation, a phoneme-to-lip parameter mapping model is trained to obtain a digital human lip animation with high lip control accuracy and natural and coordinated changes. The similarity between the digital human lip animation and the real person's lip shape is judged through comparative learning. If the similarity is lower than the preset threshold, the lip control parameters are fine-tuned to make the lip changes during the broadcast closer to the real person, improve the visual realism of the lip animation, and obtain the optimized broadcast lip animation; According to the optimized digital human lip animation, the lip shape feature vector of the key frame is extracted to obtain the lip shape change template representing different phonemes in the state of pronunciation. The template is used to guide the generation of lip shape animation to ensure the coherence of the lip shape change of the digital human when performing different phonemes. According to the lip shape change template, a tongue position control model is constructed to obtain the tongue position parameters of different phonemes when the digital human is performing pronunciation. The shape and position of the digital human tongue are controlled by the parameters to achieve the synchronization of the tongue position animation and the phonemes during the performance, and obtain the mapping relationship between the tongue position and the phonemes during the performance. According to the mapping relationship between tongue position and phonemes in broadcasting, the teeth visibility control rules are established. According to the teeth visibility parameters of the sequence, the visibility of the teeth of the digital human in broadcasting is dynamically adjusted to make the changes in teeth visibility match the characteristics of the pronunciation of the real person in broadcasting. The shape and texture features of the digital human's lips, tongue and teeth are tracked in real time during the broadcast, compared with the target broadcast lip shape, tongue position and teeth visibility parameters, and the digital human facial model parameters are dynamically adjusted to perform coordinated control of the lips, tongue and teeth during the broadcast.
2. The JavaFX-based desktop digital human interactive presentation method according to claim 1, characterized in that: The method involves scanning a real person's face with a depth camera to obtain the three-dimensional shape data of facial tissues, segmenting and annotating the data in combination with medical anatomical knowledge, and establishing a digital human face model, including the geometric shape and texture features of the facial tissue structures of the lips, tongue, and teeth, including: Acquire point cloud information carrying multi-angle three-dimensional scanning data of a real person's face, wherein the point cloud information is obtained by collecting facial tissue using a non-contact 3D laser scanner; Determine the boundary area of the lip, tongue and dental facial tissues according to the point cloud information and a pre-established medical anatomical knowledge base; An octree voxelization method is used to convert the point cloud information into a voxel grid structure to obtain voxelized data of facial tissue; For the voxelized data, a region growing algorithm is used to segment the lip, tongue and teeth regions. If the lip region is segmented, a Loop subdivision algorithm is applied to the lip region. If the tongue region is obtained by segmentation, deformation field simulation is used for the tongue region; If the tooth region is obtained by segmentation, the March cube algorithm is applied to the tooth region; Receiving facial tissue data including texture image information and feature descriptors, wherein the feature descriptors are constructed by calculating curvature and normal vector geometric features of the facial tissue and using a point feature histogram; Based on the facial tissue data, the PointNet++ network is used to perform fine segmentation and annotation of facial tissue, and a digital human facial model including lips, tongue and teeth facial tissue is obtained.
3. The JavaFX-based desktop digital human interactive presentation method according to claim 1, characterized in that: The method of obtaining the lip shape changes of different phonemes in the digital human performance, establishing the amplitude, speed and trajectory template, using the animation and timeline functions of JavaFX, obtaining the lip shape parameter sequence corresponding to the sound sequence, and realizing the synthesis of the performance lip shape animation synchronized with the phonemes by controlling the lip shape changes of the digital human face model to obtain the performance lip shape animation includes: A high-speed camera is used to obtain a lip shape change sequence corresponding to different phonemes during the pronunciation of a real person, wherein the lip shape change sequence includes image data at 240 frames per second; Performing edge detection according to the lip shape change sequence to obtain a lip contour, and extracting feature points from the lip contour to obtain key point coordinates; The coordinates of the key points are fitted using the least square method to obtain a lip shape change curve, and the amplitude, speed and acceleration parameters of the lip shape change are calculated according to the lip shape change curve; If the lip shape change parameters are obtained, a phoneme lip shape change template library is constructed; The phoneme lip shape change template library includes the amplitude, speed and trajectory templates corresponding to each phoneme; Reading a lip shape parameter sequence from the phoneme lip shape change template library, wherein the lip shape parameter sequence includes key point coordinates, change amplitude and speed information; The lip shape parameter sequence is mapped to the vertex coordinates of the digital human face model by using a radial basis function interpolation method, and the change of the vertex coordinates is controlled by an animation function to obtain a lip shape animation synchronized with the phonemes.
4. The JavaFX-based desktop digital human interactive presentation method according to claim 1, characterized in that: The method includes: training a phoneme-to-lip shape parameter mapping model based on the obtained digital human broadcast lip shape animation to obtain a digital human broadcast lip shape animation with high lip shape control accuracy and natural and coordinated changes; judging the similarity between the digital human broadcast lip shape animation and the real person's lip shape through comparative learning; if the similarity is lower than a preset threshold, fine-tuning the lip shape control parameters to make the lip shape changes during broadcast closer to the real person, improving the visual authenticity of the lip shape animation, and obtaining an optimized broadcast lip shape animation, including: Obtain the phoneme sequence, use the time convolutional network to build a phoneme to lip shape parameter mapping model, and obtain the optimized model parameters through the back propagation algorithm; Receive a real-person pronunciation video, use a 3D deformable model to reconstruct a 3D facial model in the video, extract key features of lip contour, opening and closing degree, and lip corner position, and establish a real-person pronunciation lip feature database; Based on the lip shape sequence of the digital human and the lip shape sequence of the real person, the time alignment relationship is determined by the dynamic time warping algorithm, and the visual similarity of each frame is determined by combining the structural similarity index; Calculate the similarity score between the lip shape features of the digital human and the real person's lip shape features. If the similarity score is lower than a preset threshold, enter the parameter fine-tuning process; A genetic algorithm is used to fine-tune the lip shape control parameters, and the genetic algorithm includes chromosome encoding of the lip shape control parameters; Generate new parameter combinations through roulette wheel selection, single-point crossover and Gaussian mutation operations; Iterative optimization is performed using the similarity score as the fitness function; When the similarity score exceeds a preset threshold or reaches the maximum number of iterations, the optimized lip-sync animation is output.
5. The JavaFX-based desktop digital human interactive presentation method according to claim 1, characterized in that: The method extracts the lip shape feature vector of the key frame according to the optimized digital human lip animation, obtains the lip shape change template representing different phonemes in the state of pronunciation, and uses the template to guide the generation of the lip shape animation to ensure the continuity of the lip shape change of the digital human when performing different phonemes, including: Obtain the video sequence of lip-sync animation, extract key frames using the motion amplitude calculation method based on the optical flow method, and set the extraction threshold according to 1.5 times the average motion amplitude; Hierarchical clustering algorithm is used to process lip shape feature vectors, Ward's minimum variance method is used as the linking criterion, and lip shape change templates representing different phonemes are obtained by calculating the cluster center. Establish the mapping relationship between phonemes and lip shape change templates, and use the dynamic programming algorithm to search for the optimal path based on the continuous phoneme sequence. The state is defined as the lip shape template corresponding to the current phoneme. For the selected lip shape change template sequence, a multi-dimensional cubic spline interpolation algorithm is used to generate a smooth lip shape transition curve, interpolate each dimension of the lip shape feature, and map the interpolation result to the digital human face model; The generated lip shape sequence is processed by Kalman filter to eliminate the mutation and jitter in the lip shape sequence and obtain a coherent and natural lip shape animation.
6. The JavaFX-based desktop digital human interactive presentation method according to claim 1, characterized in that: The tongue position control model is constructed according to the lip shape change template, the tongue position parameters of different phonemes when the digital human is performing pronunciation are obtained, the changes of the shape and position of the digital human tongue are controlled by the parameters, the synchronization of the tongue position animation and the phonemes during the performance is achieved, and the mapping relationship between the tongue position and the phonemes during the performance is obtained, including: The digital human tongue position control model is constructed using NURBS surface modeling technology, and the control point grid is set according to the lip shape change template; Use a bidirectional LSTM network to perform phoneme recognition on the broadcast voice and obtain the phoneme sequence and timestamp information; The tongue movement is simulated by a physical simulation method based on a mass-spring system. If the lip shape changes, the tongue position parameters are optimized according to the phoneme sequence. A deep deterministic policy gradient algorithm is used to construct an Actor-Critic network structure, and the tongue position control strategy is trained according to the naturalness of the tongue position animation and the phoneme synchronization; The tongue position parameters are interpolated using the cubic Hermite interpolation algorithm, and the tongue position parameter sequence is obtained by adding interpolation points.
7. The JavaFX-based desktop digital human interactive presentation method according to claim 1, characterized in that: The method of establishing a teeth visibility control rule according to the mapping relationship between tongue position and phonemes in the broadcast, dynamically adjusting the visibility of the teeth of the digital person in the broadcast according to the teeth visibility parameters of the phonetic sequence, and making the teeth visibility change match the characteristics of the real person in the broadcast pronunciation, includes: Obtaining tongue position parameters and phoneme types, and using a random forest algorithm to construct tooth visibility control rules based on the tongue position parameters and phoneme types to obtain a tooth visibility prediction model; Receiving a broadcast voice signal, performing phoneme recognition on the broadcast voice signal through a convolutional neural network-long short-term memory network hybrid model, and extracting a phoneme sequence and its duration; Generate a tooth visibility parameter sequence corresponding to each phoneme according to the tooth visibility prediction model and the phoneme sequence; The three-dimensional tooth model of the digital human is constructed by using subdivision surface technology. If a tooth visibility parameter sequence is received, the transparency of the three-dimensional tooth model is controlled by using vertex shader and pixel shader combined with masking technology; Performing cubic spline interpolation processing on the tooth visibility parameter sequence to obtain a smoothed tooth visibility parameter sequence; If the smoothed tooth visibility parameter sequence does not match the preset real person pronunciation characteristics, the dynamic time warping algorithm is used to fine-tune the smoothed tooth visibility parameter sequence to obtain a tooth visibility parameter sequence that matches the preset real person pronunciation characteristics.
8. The JavaFX-based desktop digital human interactive presentation method according to claim 1, characterized in that: The real-time tracking of the shape and texture features of the lips, tongue and teeth of the digital human during the broadcasting process, comparing them with the target broadcasting lip shape, tongue position and teeth visibility parameters, dynamically adjusting the digital human facial model parameters, and performing coordinated control of the lips, tongue and teeth during the broadcasting, includes: Using a facial key point detection algorithm to obtain key point information of a digital human face image, the key point information includes position data of lips, tongue and teeth areas; According to the key point information obtained, a mapping relationship between two-dimensional image features and three-dimensional model parameters is established; If it is detected that the target broadcast lip shape, tongue position and teeth visibility parameters deviate from the actual feature parameters, the adjustment amount of the facial model control point is calculated; The adjustment amount is minimized through an iterative optimization algorithm to obtain the adjustment result of the digital human face model; According to the adjustment results, a mass spring system is introduced to connect the key control points of lips, tongue and teeth, and the mass spring system is used to achieve coordinated control of lips, tongue and teeth; The spring force action data is obtained from the mass-spring system, and the final adjustment parameters of the digital human face model are determined according to the spring force action data.
Citation Information
Patent Citations
Voice-driven face animation generation method and system
CN115457169A
Facial motion capture system and method based on data analysis
CN118588079A
A system and method of mouth mapping for efficiently teaching language pronunciation
GB2628162A
Oral device for safe & efficient dental treatment
IN202141012903A
KR20240083590A
Cited By
Phoneme time axis driven high-definition video mouth shape automatic synthesis method
CN120640052A
Virtual digital population broadcast optimization method and device
CN121564153A
Automatic test method and system for numerical control system, electronic equipment and medium
CN121657633A