A digital human mouth shape prediction system and method based on a variational autoencoder in smart agriculture
The digital human mouth shape prediction system using variational autoencoders solves the problem of unclear mouth shape prediction under noise interference in smart agriculture, and achieves accurate mouth shape prediction matching of professional terms in complex environments, improving the synchronization and naturalness of mouth shape animation.
Patent Information
- Application Number
- CN202511811338.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2045-12-04
AI Technical Summary
Existing technologies for digital human mouth shape prediction in smart agriculture suffer from poor noise suppression, leading to distortion of speech information and mouth shape matching deviations. Furthermore, state transition models struggle to capture the complex relationship between speech and mouth shape, resulting in unclear mouth shapes and speech asynchrony.
A digital human mouth shape prediction system based on variational autoencoder is adopted. Multi-dimensional speech features are extracted through the speech feature extraction module, and latent representations are generated and decoded into mouth shape key point sequences using the variational autoencoder mapping module. The semantic enhancement module is combined to perform semantic enhancement and temporal filtering of the mouth shape key point sequences. Finally, the mouth shape rendering module maps the key point sequences to a three-dimensional digital human face model to generate mouth shape animation synchronized with the speech signal.
It achieves accurate lip-reading prediction for professional vocabulary in complex and noisy environments, suppresses lip tremors and abnormal frames, ensures synchronization between lip movements and speech, and improves the clarity and naturalness of digital human lip-reading animation.
Smart Images

Figure CN121236246B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a digital human mouth shape prediction system and method based on a variational autoencoder in smart agriculture. BACKGROUND
[0002] In the actual scene of smart agriculture, there are often noises of agricultural machinery working and sounds of wind blowing in the farmland, and the agricultural technology explanation and policy explanation contain many professional vocabulary in the agricultural field. Users need to understand these professional contents through the changes of the digital human mouth shape, which requires the digital human mouth shape prediction method to not only accurately match the pronunciation characteristics of professional vocabulary and ensure the clarity of mouth shape actions, but also to resist the interference of environmental noise on the voice signal, avoid the deformation of the mouth shape or the asynchronization of the mouth shape and the voice, and ensure the accurate and efficient transmission of information.
[0003] At present, the commonly used technical solution for this scenario is a mouth shape prediction method based on a traditional state transition model. The solution first processes the voice signal in a noisy environment through a basic voice noise reduction method, then extracts surface information representing the characteristics of the sound from the voice, such as the changes in pitch and strength, and finally uses a state transition model to establish the correspondence between the voice information and the preset mouth shape key actions, and outputs the corresponding digital human mouth shape change sequence according to the input voice.
[0004] However, the existing solution has obvious deficiencies. The traditional basic noise reduction method does not have good suppression effect on the complex noise in the farmland, and the residual noise can easily cause distortion of the extracted voice information, which in turn causes deviation in the matching of the mouth shape corresponding to the professional vocabulary. At the same time, the state transition model is difficult to capture the complex relationship between the voice and the mouth shape, and has weak correlation processing capability for the sequence of mouth shape changes. It not only cannot accurately restore the subtle mouth shape actions when pronouncing professional vocabulary, but also easily causes problems such as mouth shape shaking and sudden jumping of actions, resulting in unclear digital human mouth shape and asynchronization with the voice, which ultimately affects the user's understanding of agricultural professional information. SUMMARY
[0005] The present application aims to provide a digital human mouth shape prediction system and method based on a variational autoencoder in smart agriculture to solve the problem of insufficient clarity and synchronization of digital human mouth shape in the prior art.
[0006] To solve the above technical problems, in a first aspect, the present application provides a digital human mouth shape prediction system based on a variational autoencoder in smart agriculture, comprising a voice feature extraction module, a variational autoencoder mapping module, a semantic enhancement module, and a mouth shape rendering module:
[0007] The voice feature extraction module is configured to receive an input voice signal generated in an intelligent agricultural scene and extract multi-dimensional voice features from the input voice signal, and generate a voice representation by fusing the multi-dimensional voice features.
[0008] The variational autoencoder mapping module is configured to input the voice representation into a variational autoencoder, obtain a latent representation by using the variational autoencoder, and decode the latent representation into a sequence of mouth key points.
[0009] The semantic enhancement module is configured to perform semantic enhancement on the sequence of mouth key points based on an agricultural specific term in the input voice signal and a preset standard mouth template, and perform time series filtering on the enhanced sequence of mouth key points to suppress mouth shape jitter and abnormal frames.
[0010] The mouth shape rendering module is configured to map the semantically enhanced sequence of mouth key points to a three-dimensional digital human face model, convert the sequence of mouth key points into deformation data of facial soft tissues based on facial parameters corresponding to different target users, calculate position and rotation parameters of corresponding skeletal nodes according to a preset correspondence between mouth shapes and bones, and perform combined rendering based on the deformation data and the position and rotation parameters to generate a digital human mouth shape animation synchronized with the input voice signal.
[0011] Optionally, in the process of mapping the semantically enhanced sequence of mouth key points to a three-dimensional digital human face model and converting the sequence of mouth key points into deformation data of facial soft tissues based on facial parameters corresponding to different target users, the mouth shape rendering module is specifically configured to perform the following process:
[0012] mapping the semantically enhanced sequence of mouth key points to a three-dimensional digital human face model;
[0013] loading corresponding facial parameters from a parameter database according to a target user identifier, establishing associated motion characteristics of facial soft tissues based on the facial parameters, and establishing a biomechanical model of facial soft tissues conforming to individual characteristics of the target user;
[0014] inputting the sequence of mouth key points into the biomechanical model, calculating deformation responses of facial soft tissues under key point driving according to the associated motion characteristics, and outputting a sequence of deformation data corresponding to the sequence of mouth key points.
[0015] Optionally, in the process of establishing associated motion characteristics of facial soft tissues based on the facial parameters and establishing a biomechanical model of facial soft tissues conforming to individual characteristics of the target user, the mouth shape rendering module is specifically configured to perform the following process:
[0016] establish a deformation retention characteristic describing deformation resistance of the facial soft tissue in each direction based on a muscle tension coefficient in the facial parameter;
[0017] establish a motion stability characteristic describing motion stability in a motion process of the facial soft tissue based on a soft tissue control parameter in the facial parameter;
[0018] combine the deformation retention characteristic and the motion stability characteristic to form a correlation motion characteristic;
[0019] determine a thickness distribution characteristic of each part of the facial soft tissue based on a sparse distribution of the facial soft tissue in the facial parameter;
[0020] divide the facial soft tissue into a plurality of region units connected with each other, and assign a corresponding correlation motion characteristic and thickness distribution characteristic to each region unit according to an anatomical position;
[0021] establish unit shape data according to the correlation motion characteristic of the region unit, and establish unit distribution data according to the thickness distribution characteristic;
[0022] establish a biomechanical model of the facial soft tissue by assembling the unit shape data and the unit distribution data of all region units.
[0023] Optionally, in the process of rendering the mouth shape according to the preset correspondence relationship between the mouth shape and the skeleton, calculating the position and rotation parameter of the corresponding skeleton node, combining the deformation data and the position and rotation parameter to render, and generating the digital human mouth shape animation synchronized with the input speech signal, the mouth shape rendering module is specifically configured to perform the following process:
[0024] determine a skeleton node affected by each key point and a corresponding influence degree according to a preset correlation rule between the mouth shape key point sequence and the skeleton node;
[0025] calculate a position offset and a rotation change of the skeleton node based on the influence degree;
[0026] dynamically fuse facial soft tissue deformation and facial skeleton motion represented by the skeleton node based on the deformation data and the offset and rotation change of the skeleton node, and perform rendering processing based on a fusion result to generate a preliminary mouth shape animation sequence;
[0027] perform time synchronization control on the preliminary mouth shape animation sequence and the input speech signal to ensure that a time synchronization error between the mouth shape animation and the speech signal is within a preset threshold range, and output a digital human mouth shape animation sequence synchronized with the input speech signal.
[0028] Optionally, in the process of performing dynamic fusion of face soft tissue deformation and face skeleton motion represented by the bone nodes based on the morphing data and the offset and rotation changes of the bone nodes, the mouth shape rendering module is specifically configured to perform the following process:
[0029] Layered fusion of the soft tissue vertex displacement represented by the morphing data and the rigid motion generated by the bone node transformation is performed to establish a layered fusion framework;
[0030] In the layered fusion framework, a morphing data weight value and a bone transformation weight value are calculated according to a preset spatial distance of a vertex from a mouth core region, wherein the morphing data weight value is negatively correlated with the spatial distance, and the bone transformation weight value is positively correlated with the spatial distance;
[0031] Based on the offset and rotation changes of the bone nodes, a rigid transformation matrix corresponding to the bone nodes is calculated, and a rigid transformation position of a vertex is calculated based on the rigid transformation matrix;
[0032] A local morphing position of a vertex is determined based on the morphing data;
[0033] The local morphing position and the rigid transformation position are weighted and superimposed according to the morphing data weight value and the bone transformation weight value;
[0034] For a vertex affected by multiple bone nodes, a transformation contribution value of each bone node is calculated according to the binding weight of the vertex from each bone node, and the transformation contribution values are weighted and averaged.
[0035] Optionally, in the process of performing semantic enhancement of the mouth shape key point sequence based on the agricultural specific terms in the input speech signal, combining a preset standard mouth shape template, and performing time series filtering on the enhanced mouth shape key point sequence to suppress mouth shape jitter and abnormal frames, the semantic enhancement module is specifically configured to perform the following process:
[0036] Text conversion is performed on the input speech signal, and agricultural specific terms in the converted text are identified, and a corresponding standard mouth shape template is extracted from a preset template library according to the agricultural specific terms;
[0037] Similarity calculation is performed on the mouth shape key point sequence and the standard mouth shape template, and adaptive weighted adjustment is performed on the key point positions related to the agricultural specific terms in the mouth shape key point sequence based on the calculation result to generate a semantically enhanced mouth shape key point sequence;
[0038] Consecutive consistency analysis is performed on the key points in the enhanced mouth shape key point sequence, and based on the analysis result, the key point trajectory is smoothed to suppress mouth shape jitter and eliminate abnormal frames that do not conform to the human face movement rules.
[0039] Optionally, in the process of inputting the speech representation into the variational autoencoder, obtaining the latent representation through the variational autoencoder, and decoding the latent representation into the mouth shape key point sequence, the variational autoencoder mapping module is specifically configured to perform the following process:
[0040] The speech representation is input into a first network path of the variational autoencoder, and a feature distribution of the latent representation is obtained by performing feature transformation on the speech representation;
[0041] Based on the feature distribution, the latent representation is obtained through a random sampling operation;
[0042] The latent representation is input into a second network path of the variational autoencoder, and a feature sequence matching the dimension of the mouth shape key point is obtained by performing layer-by-layer feature expansion on the latent representation;
[0043] The feature sequence is subjected to coordinate mapping to output a mouth shape key point sequence.
[0044] In a second aspect, the present application provides a digital human mouth shape prediction method based on a variational autoencoder in smart agriculture, comprising:
[0045] An input speech signal generated in a smart agriculture scene is received, and multi-dimensional speech features are extracted therefrom, and the multi-dimensional speech features are fused to generate a speech representation;
[0046] The speech representation is input into a variational autoencoder, and a latent representation is obtained through the variational autoencoder, and the latent representation is decoded into a mouth shape key point sequence;
[0047] Based on the agricultural proper nouns in the input speech signal, the mouth shape key point sequence is semantically enhanced in combination with a preset standard mouth shape template, and the enhanced mouth shape key point sequence is subjected to time series filtering to suppress mouth shape jitter and abnormal frames;
[0048] The mouth shape key point sequence after semantic enhancement is mapped to a three-dimensional digital human face model, the mouth shape key point sequence is converted into deformation data of facial soft tissue in combination with facial parameters corresponding to different target users, and the position and rotation parameters of the corresponding skeletal nodes are calculated according to the preset correspondence between the mouth shape and the skeleton, and the deformation data and the position and rotation parameters are combined and rendered to generate a digital human mouth shape animation synchronized with the input speech signal.
[0049] In a third aspect, the present application provides an electronic device, comprising:
[0050] a memory for storing a computer program;
[0051] a processor for implementing the steps of the method for predicting digital human mouth shape in smart agriculture based on variational autoencoder according to the first aspect when executing the computer program.
[0052] In a fourth aspect, the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program can implement the steps of the method for predicting digital human mouth shape in smart agriculture based on variational autoencoder according to the first aspect when executed by a processor.
[0053] The method for predicting digital human mouth shape in smart agriculture based on variational autoencoder provided by the present application can receive input speech signals of a smart agriculture scene through a speech feature extraction module, extract multi-dimensional speech features and generate speech representation by fusion, can provide comprehensive and accurate speech basic data for subsequent mouth shape prediction, and avoid information loss caused by single feature; the speech representation is input into a model through a variational autoencoder mapping module, latent representation is obtained and decoded into a mouth shape key point sequence, which can establish effective association between speech and mouth shape, and realize preliminary conversion from speech to mouth shape key information; the semantic enhancement module combines agricultural proper nouns and a preset standard mouth shape template to enhance the mouth shape key point sequence, and performs time sequence filtering, which can improve the mouth shape accuracy of agricultural professional vocabulary, suppress mouth shape jitter and abnormal frames, and ensure the stability of the mouth shape sequence; the mouth shape rendering module maps the enhanced mouth shape key point sequence to a three-dimensional digital human face model, combines user face parameters to convert into face soft tissue deformation data, calculates skeleton node parameters and combines rendering, which can convert abstract mouth shape key points into three-dimensional digital human mouth shape animation that fits the individual characteristics of the user, and realize synchronization with the input speech.
[0054] Further, the mouth shape rendering module first maps the mouth shape key point sequence after semantic enhancement to a three-dimensional digital human face model, then loads corresponding face parameters from a parameter database according to a target user identifier, establishes a face soft tissue associated motion characteristic and a biomechanical model that meets the individual characteristics of the user based on the parameters, and finally inputs the mouth shape key point sequence into the model to calculate and output the corresponding face soft tissue deformation data sequence according to the associated motion characteristic. This process can establish a biomechanical model that fits the individual face characteristics of the user by loading face parameters exclusive to the target user, make the calculated face soft tissue deformation data more personalized and realistic, avoid the problem of mismatch between the mouth shape and the face characteristics of the user caused by a general model, and further improve the accuracy and naturalness of the digital human mouth shape animation. BRIEF DESCRIPTION OF DRAWINGS
[0055] In order to make the technical scheme of the embodiments of the present application or the prior art clearer, the accompanying drawings needed in the description of the embodiments or the prior art will be briefly introduced. Obviously, the accompanying drawings in the description below only aim to some embodiments of the present application, and for those skilled in the art, other drawings can be obtained based on these accompanying drawings without creative effort.
[0056] Figure 1 A structural schematic diagram of a digital human mouth shape prediction system based on a variational autoencoder in smart agriculture provided by an embodiment of the present application is shown.
[0057] Figure 2 A specific implementation flowchart of a digital human mouth shape prediction system based on a variational autoencoder in smart agriculture provided by an embodiment of the present application is shown.
[0058] Figure 3 A flowchart of a digital human mouth shape prediction method based on a variational autoencoder in smart agriculture provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0059] In order to make those skilled in the art better understand the present application, the present application will be further described in detail below in combination with the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort fall within the scope of protection of the present application.
[0060] Figure 1 An architecture diagram of a digital human mouth shape prediction system based on a variational autoencoder in smart agriculture provided by an embodiment of the present application is shown. As shown in the figure, Figure 1 The system includes a speech feature extraction module 11, a variational autoencoder mapping module 12, a semantic enhancement module 13, and a mouth shape rendering module 14.
[0061] The speech feature extraction module 11 is configured to receive an input speech signal generated in a smart agriculture scenario and extract multi-dimensional speech features therefrom, and to fuse the multi-dimensional speech features to generate a speech representation.
[0062] The input voice signal refers to audio information generated in the smart agriculture scene and used for digital person broadcasting, such as voice content of agricultural technology guidance and policy interpretation; the multi-dimensional voice feature is a set of information extracted from the input voice and reflecting different attributes of the voice, including time domain features, frequency domain features, and cepstrum domain features such as mel-frequency cepstral coefficients; the voice representation is a unified feature vector or matrix formed by fusion processing of the multi-dimensional voice features and used as input of a subsequent mouth shape prediction model, which can comprehensively reflect the key information of the input voice.
[0063] Specifically, first, the voice feature extraction module 11 receives the input voice signal generated in the smart agriculture scene. The process acquires voice stream data such as "greenhouse crop pest control technology guidance" recorded by an agricultural technician through an audio acquisition interface or a data receiving port built in the module, ensuring that the voice signal is not lost or distorted during transmission.
[0064] Secondly, the module calls a multi-feature extraction algorithm to extract multi-dimensional voice features from the received voice signal. For example, the voice signal is decomposed into the frequency domain by short-time Fourier transform, and the spectral peak, frequency bandwidth, and other frequency domain features of different frequency bands are calculated. At the same time, the short-time energy and zero-crossing rate of the voice signal are calculated by time domain analysis algorithm, and the mel-frequency cepstral coefficients and other cepstrum domain features are obtained by mel filtering and discrete cosine transform, forming a multi-dimensional original feature set.
[0065] Finally, the module uses a feature fusion algorithm such as weighted average fusion and principal component analysis fusion to process the extracted multi-dimensional voice features. For example, according to the importance of different features to mouth shape prediction, corresponding weights are assigned, time domain, frequency domain, and cepstrum domain features are converted into feature vectors of uniform dimensions, and voice representation that can be directly input into the subsequent variational autoencoder mapping module is generated, completing the entire voice feature extraction process.
[0066] The variational autoencoder mapping module 12 is configured to input the voice representation into a variational autoencoder, obtain a latent representation through the variational autoencoder, and decode the latent representation into a mouth shape key point sequence.
[0067] The voice representation is a unified feature vector or matrix output by a voice feature extraction module and integrating multi-dimensional voice features such as time domain, frequency domain and cepstrum domain features, serving as input data of the module; the variational autoencoder is a generative model combining deep learning and Bayesian inference, including an encoder and a decoder, and capable of learning the latent distribution of data and generating target data based on the distribution; the latent representation is a low-dimensional vector obtained after processing by the variational autoencoder, containing core information for converting a voice signal into mouth shape action; and the mouth shape key point sequence is data formed by arranging three-dimensional coordinates of multiple mouth shape key positions in time sequence, and capable of dynamically reflecting the changing state of the mouth shape with the voice, serving as the final output result of the module.
[0068] The variational autoencoder mapping module 12 is specifically configured to perform the following processes: inputting the voice representation into a first network path of a variational autoencoder, performing feature transformation on the voice representation to obtain a feature distribution of the latent representation; obtaining the latent representation through a random sampling operation based on the feature distribution; inputting the latent representation into a second network path of the variational autoencoder, performing layer-by-layer feature expansion on the latent representation to obtain a feature sequence matching the dimension of the mouth shape key point; and performing coordinate mapping on the feature sequence to output the mouth shape key point sequence.
[0069] In the above specific processes, the first network path, i.e., the encoder part of the variational autoencoder, is composed of multiple layers of neural networks, and is configured to perform feature compression and transformation on the input voice representation; the feature transformation is a process of gradually extracting key information related to the mouth shape from the voice representation and compressing the dimension through linear operation and nonlinear activation function of the neural network; and the feature distribution of the latent representation is a distribution form output by the encoder and describing the probability characteristics of the latent variable, usually represented by a Gaussian distribution parameterized by mean and variance.
[0070] The random sampling operation is a process of extracting specific values of the latent variable based on the feature distribution of the latent representation, and needs to be realized by the reparameterization technique to be gradient derivable to meet the model training requirement; the second network path, i.e., the decoder part of the variational autoencoder, is symmetrical to the encoder structure, and is configured to reconstruct the low-dimensional latent representation into a high-dimensional feature; the layer-by-layer feature expansion is a processing mode of the decoder for gradually increasing the feature dimension through multiple layers of neural networks, each layer being configured to expand the dimension through linear transformation and enhance the expression ability through the activation function; the feature sequence matching the dimension of the mouth shape key point is feature data with a dimension consistent with the preset number of mouth shape key points, ensuring that the feature dimension can directly correspond to the number of mouth shape key positions; and the coordinate mapping is a process of converting the abstract feature sequence into specific coordinates of the mouth shape key points in the three-dimensional space, and is realized according to the corresponding relationship between the features and the coordinates learned during the model training stage.
[0071] In this embodiment, the speech representation is first input to the encoder corresponding to the first network path of the variational autoencoder through the variational autoencoder mapping module 12. The encoder uses a multi-layer fully connected neural network or a convolutional neural network to perform feature transformation on the speech representation. By alternating operations of linear activation and non-linear activation functions such as ReLU, the feature dimension is gradually compressed, and finally the feature distribution of the latent representation is output, which is usually the mean vector of a Gaussian distribution. Sum of logarithmic variance vector For example, the 128-dimensional speech representation corresponding to the phrase "greenhouse tomato water and fertilizer management" is input into the encoder, processed through three fully connected layers, and outputs a 32-dimensional mean vector. and a 32-dimensional log-variance vector Together they constitute the feature distribution of the latent representation. .
[0072] Secondly, based on the feature distribution obtained in the first step, a random sampling operation is performed using reparameterization techniques to obtain the latent representation. Specifically, the process begins with a standard normal distribution... Mid-sampling yields random noise vector Then, combined with the mean vector output by the encoder and the variance vector obtained by exponential operation Through formula Generate latent representation ,in, This is the element-wise multiplication symbol, which means multiplying the elements at corresponding positions of two vectors. It is the mean vector; This is the variance vector. For example, use this formula for a 32-dimensional mean vector. Sum of logarithmic variance vector Perform calculations: and Multiply each corresponding element together, then multiply by... Adding the corresponding elements generates a 32-dimensional latent representation. This approach retains the probabilistic characteristics of the feature distribution while also meeting the gradient propagation requirements for model training.
[0073] Next, the latent representation obtained in the second step... The input is fed into the second network path decoder of the variational autoencoder. The decoder employs a multi-layer fully connected neural network symmetrical to the encoder, achieving dimensionality enhancement through layer-by-layer feature expansion. Each layer expands the feature dimension through linear transformations and combines activation functions to enhance non-linear expressive power, until the output is a feature sequence matching the dimension of the lip shape key points. For example, a 32-dimensional latent representation... The input decoder is gradually expanded from 32 dimensions to 68 dimensions through 4 fully connected layers, corresponding to 68 key points of the mouth shape, to obtain a 68-dimensional feature sequence.
[0074] Finally, the feature sequence generated in the third step is processed by a coordinate mapping algorithm, and each feature value in the feature sequence is converted into a coordinate of a mouth shape key point in a three-dimensional space according to the correspondence between the features and the coordinates learned in the model training phase, and a mouth shape key point sequence containing time sequence information is output. For example, a 68-dimensional feature sequence is converted into 68 three-dimensional coordinates of key points by coordinate mapping, and a continuous mouth shape key point sequence is formed by arranging the 68 three-dimensional coordinates according to the time sequence of the speech.
[0075] The present application solves the gradient transmission problem of the random process through the reparameterization technique, guarantees effective training of the model while preserving the feature distribution characteristics, realizes accurate conversion of the low-dimensional latent representation to the high-dimensional mouth shape feature, ensures that the feature dimension matches the mouth shape key point, establishes a direct association between the speech feature and the mouth shape space coordinate, effectively captures the complex nonlinear relationship between the speech and the mouth shape, and finally generates an accurate mouth shape key point sequence to provide reliable basic data for the subsequent modules and improve the rationality and accuracy of the mouth shape prediction.
[0076] The semantic enhancement module 13 is configured to perform semantic enhancement on the mouth shape key point sequence based on the agricultural specific terms in the input speech signal and in combination with a preset standard mouth shape template, and perform time sequence filtering on the enhanced mouth shape key point sequence to suppress mouth shape jitter and abnormal frames.
[0077] The input speech signal is original audio information generated in the intelligent agriculture scene, such as speech related to agricultural technology guidance and policy interpretation; the agricultural specific terms are professional terms related to the agricultural field in the input speech, such as "drip irrigation and fertilizer ratio" and "green prevention and control of plant diseases and insect pests"; the preset standard mouth shape template is mouth shape key point reference data that is accurately matched with the pronunciation of the agricultural specific terms and is trained and stored in advance, and is used to standardize the mouth shape expression corresponding to the professional vocabulary; the mouth shape key point sequence is a three-dimensional coordinate data of the mouth shape key position arranged in time sequence output by the variational autoencoder mapping module; the semantic enhancement is a process of optimizing and adjusting the mouth shape key point sequence based on the semantic information of the agricultural specific terms, so that it is more in line with the pronunciation characteristics of the professional vocabulary; the time sequence filtering is a smoothing process of the time sequence dimension of the enhanced mouth shape key point sequence, and the core purpose is to suppress the mouth shape jitter phenomenon and eliminate abnormal frames that do not conform to the motion law, and to guarantee the continuity of the mouth shape sequence.
[0078] The semantic enhancement module 13 is specifically configured to perform the following processes: text conversion on the input speech signal, identification of agricultural specific terms in the converted text, extraction of corresponding standard mouth shape templates from a preset template library according to the agricultural specific terms, similarity calculation of the mouth shape key point sequence and the standard mouth shape template, and adaptive weighted adjustment of the key point positions related to the agricultural specific terms in the mouth shape key point sequence based on the calculation result to generate a semantic enhanced mouth shape key point sequence, continuous consistency analysis of the key points in the enhanced mouth shape key point sequence, and smoothing processing of the key point trajectories based on the analysis result to suppress mouth shape jitter and eliminate abnormal frames that do not conform to the human facial movement law.
[0079] In the above specific process, the text conversion is a processing process of converting the input speech signal into a text form; the agricultural specific term identification is an operation of screening agricultural field professional terms from the converted text through natural language processing technology; the preset template library is a database for storing standard mouth shape templates corresponding to various agricultural specific terms, which is indexed by professional vocabulary for quick extraction; and the standard mouth shape template is a mouth shape key point reference sequence preset for the pronunciation characteristics of agricultural specific terms and conforming to the human facial movement law.
[0080] The similarity calculation is an operation process of measuring the degree of fit of the mouth shape key point sequence and the standard mouth shape template in key point position and movement trend; the adaptive weighted adjustment is a processing method of assigning dynamic weights to the key points related to the agricultural specific terms in the mouth shape key point sequence and correcting the positions according to the similarity calculation result, and the weight size is inversely proportional to the similarity, and the lower the similarity, the greater the adjustment amplitude; and the semantic enhanced mouth shape key point sequence is a mouth shape key point time series data that is more consistent with the pronunciation characteristics of professional vocabulary after specific term matching adjustment.
[0081] The continuous consistency analysis is a process of analyzing the time sequence trajectory of the enhanced mouth shape key point sequence to determine whether the key point movement conforms to the continuity rule of human facial muscle movement; the smoothing processing is an operation of correcting the fluctuation of the key point trajectory through a time sequence filtering algorithm; the mouth shape jitter is an irregular small fluctuation phenomenon of the mouth shape key point in time sequence; and the abnormal frame is an invalid data frame in which the mouth shape key point position deviates from the human facial movement law and the connection between the previous and subsequent frames is broken.
[0082] In this embodiment, the input speech signal is first converted into text by the semantic enhancement module 13. End-to-end speech-to-text technology based on deep learning is used to map the temporal features of the audio signal into a text sequence, resulting in text data consistent with the speech content. For example, the speech "Greenhouse tomato gray mold control requires ventilation and dehumidification combined with biological agent spraying" is converted into text with the same content. Secondly, agricultural terminology is identified in the converted text. An agricultural dictionary containing terms such as "greenhouse tomato," "gray mold control," "ventilation and dehumidification," and "biological agent spraying" is loaded. The TF-IDF keyword extraction algorithm and BiLSTM semantic classification model are combined to filter out agricultural terms in the text. Finally, for each identified agricultural term, a corresponding standard mouth shape template is extracted from a preset template library using a keyword matching index. For example, for "biological agent spraying," the mouth shape key point reference sequence, trained from a large number of pronunciation samples and stored in the template library, is extracted as the standard mouth shape template. ( The number of key points for the mouth shape. For the first (3D coordinates of key points).
[0083] Next, based on the standard mouth shape template extracted in the first step... The Euclidean distance algorithm is used to calculate the mouth shape key point sequence output by the variational autoencoder. For the first in the sequence The similarity between the three-dimensional coordinates of each key point and the template is calculated using the following formula:
[0084]
[0085] in, Represents a sequence With template Similarity, range of values The closer the value is to 1, the more similar the two sides are. This represents the total number of key points in the mouth shape. and The number of key points must correspond one-to-one; For sequence The Middle The three-dimensional coordinates of the key points; template The Middle The three-dimensional coordinates of the key points; for and The Middle The three-dimensional Euclidean distance between each corresponding key point reflects the spatial positional difference of a single key point; The maximum possible distance to the preset key points is used for normalization to eliminate the influence of coordinate magnitude.
[0086] Based on the similarity calculation results, the sequence Perform adaptive weighted adjustment, setting a similarity threshold Th (e.g., 0.8). Then, the coordinates of key points related to agricultural proper nouns are adjusted according to the following formula:
[0087]
[0088] in, For the adjusted number The coordinates of the key points; These are the weighting coefficients. The lower the similarity, The larger the template The stronger the guiding effect; template The Middle Coordinates of key points; Original sequence The Middle The coordinates of key points are used to generate a semantically enhanced sequence of lip shape key points after adjustment. .
[0089] Finally, the enhanced sequence , For frame number, For the first For the set of key points in a frame, perform continuity consistency analysis and calculate the displacement change of corresponding key points in adjacent frames:
[0090]
[0091] in, For the first Frame and the The first frame between two adjacent frames Displacement of key points; For frame indexing; For key point indexing; For the first Frame number The three-dimensional coordinates of the key points; For the first Frame number The three-dimensional coordinates of the key points.
[0092] Displacement thresholds are set based on the movement patterns of the human face. If a frame contains more than Key point displacement If the frame is not found to be an abnormal frame, it will be removed. Subsequently, Kalman filtering is used to smooth the keypoint trajectories, with the first frame being the most accurate. Key points Taking coordinates as an example, the filtering formula is as follows:
[0093] Prediction phase:
[0094] Update phase:
[0095] in, For the first frame Predicted values of coordinates; For the first frame Filtered values of the coordinates; For the first Estimated motion velocity of frame keypoints; This refers to the frame interval time. The Kalman gain is calculated from the process noise covariance and the observation noise covariance. For the first frame The original observation values of the coordinates; For the first frame The filtered values of the coordinates. The same operation is performed on the coordinates, eventually resulting in a smooth sequence of mouth shape key points without abnormal frames.
[0096] This application identifies agricultural terminology and matches it with standard mouth shape templates. By combining similarity calculation and adaptive weighted adjustment, it accurately corrects the key mouth shape positions corresponding to professional terms, significantly improving the matching degree between mouth shape and agricultural speech semantics. This makes the mouth shape of professional content clearer, improves the naturalness of digital human mouth shape animation and information transmission efficiency, and helps users easily understand agricultural professional guidance content.
[0097] The mouth shape rendering module 14 is used to map the semantically enhanced mouth shape key point sequence to a three-dimensional digital human face model, combine the facial parameters corresponding to different target users, convert the mouth shape key point sequence into deformation data of facial soft tissue, and calculate the position and rotation parameters of the corresponding bone nodes according to the preset correspondence between mouth shape and bones. Based on the deformation data and the position and rotation parameters, the module performs combined rendering to generate a digital human mouth shape animation synchronized with the input speech signal.
[0098] The semantic-enhanced mouth shape key point sequence is a mouth shape key position time sequence coordinate data that has accuracy, stability, and semantic matching degree after being optimized by a semantic enhancement module; the three-dimensional digital human face model is a digital model that simulates the structure of a human face and is constructed based on three-dimensional modeling technology, and includes components such as a face skeleton and soft tissue; the target user is a service object of the system, such as an agricultural technology learner or a planting practitioner; the face parameters are a parameter set related to the individual of the target user and describing the mechanical properties of the face soft tissue, such as a muscle tension coefficient and a soft tissue control parameter; the deformation data of the face soft tissue is quantitative data representing the shape change of the face soft tissue under the driving of the mouth shape movement; the preset correspondence between the mouth shape and the skeleton is a predefined association rule between the mouth shape movement and the face skeleton node movement; the position and rotation parameters of the skeleton node are quantitative data describing the spatial coordinates and rotation state of the face skeleton node; the combined rendering is a processing process of fusing the soft tissue deformation data and the skeleton movement parameters to generate a visual animation; the input voice signal is original audio data received by the system; and the digital human mouth shape animation is a three-dimensional visual dynamic effect that simulates the mouth shape movement of a human body and is synchronized with the voice signal.
[0099] As shown in Figure 2 The mouth shape rendering module 14 “maps the semantic-enhanced mouth shape key point sequence to a three-dimensional digital human face model, combines the respective face parameters of different target users, and converts the mouth shape key point sequence into deformation data of face soft tissue” is specifically used to perform the following processes:
[0100] The semantic-enhanced mouth shape key point sequence is mapped to a three-dimensional digital human face model; according to the target user identifier, the corresponding face parameters are loaded from the parameter database, the associated movement characteristics of the face soft tissue are established based on the face parameters, and the biomechanical model of the face soft tissue that conforms to the individual characteristics of the target user is established; the mouth shape key point sequence is input to the biomechanical model, the deformation response of the face soft tissue under the driving of the key point is calculated according to the associated movement characteristics, and a deformation data sequence corresponding to the mouth shape key point sequence is output.
[0101] As a possible implementation scheme, in the process of “establishing the associated movement characteristics of the face soft tissue based on the face parameters, and establishing the biomechanical model of the face soft tissue that conforms to the individual characteristics of the target user”, the following is specifically performed:
[0102] Based on the muscle tension coefficient in the facial parameters, deformation retention characteristics describing the resistance of facial soft tissue to deformation in various directions are established; based on the soft tissue control parameters in the facial parameters, motion stability characteristics describing the movement of facial soft tissue are established; the deformation retention characteristics and motion stability characteristics are combined to form associated motion characteristics; based on the density distribution of facial soft tissue in the facial parameters, the thickness distribution characteristics of each part of the facial soft tissue are determined; the facial soft tissue is divided into multiple interconnected regional units, and each regional unit is assigned corresponding associated motion characteristics and thickness distribution characteristics according to its anatomical location; unit morphology data is established based on the associated motion characteristics of the regional units, and unit distribution data is established based on the thickness distribution characteristics; based on the unit morphology data and unit distribution data of all regional units, a biomechanical model of facial soft tissue is established through assembly.
[0103] Among them, the biomechanical model of facial soft tissue is a mathematical model that integrates individual mechanical properties and anatomical structure to simulate the movement response of facial soft tissue; the muscle tension coefficient is an index among facial parameters that describes the ability of facial soft tissue to resist deformation, and the larger the value, the more difficult the deformation.
[0104] Deformation retention characteristics are a quantitative description of the facial soft tissue's resistance to deformation in all directions, based on the muscle tension coefficient; soft tissue control parameters are parameters that describe the energy dissipation characteristics of facial soft tissue movement and affect the rate of movement decay; movement stability characteristics are a quantitative expression of the energy dissipation law of soft tissue movement, constructed based on the soft tissue control parameters.
[0105] Density distribution is information describing the mass distribution of different parts of facial soft tissue in facial parameters; thickness distribution characteristics are mass distribution features of different areas of facial soft tissue determined based on density distribution.
[0106] Element morphology data is the quantitative data of stiffness characteristics of a single region element; element distribution data is the quantitative data of mass distribution of a single region element.
[0107] In the above specific process, the target user identifier is the identity information used to uniquely distinguish different target users, such as user number, account ID, etc.; the associated motion characteristics are quantitative expressions that describe the mechanical characteristics of facial soft tissue based on facial parameters; the assembly method is a method of integrating the stiffness and mass descriptions of all regional units to construct a complete model; the deformation response is the shape change response of facial soft tissue driven by the mouth shape key points; the deformation data sequence is the continuous soft tissue deformation data that corresponds to the mouth shape key point sequence in time.
[0108] In this embodiment, the first step is the mapping of key mouth features and the construction of a biomechanical model. The module first maps the semantically enhanced sequence of key mouth features onto a 3D digital human face model, then uses a coordinate normalization algorithm to associate the key points with their corresponding positions on the model. The mapping formula employs homogeneous coordinate transformation.
[0109]
[0110] in, For the model coordinate system, the first Homogeneous coordinates of key points, including ; for Homogeneous transformation matrix; For the world coordinate system Three-dimensional coordinates of key points; Adding "1" to the homogeneous column vector of key points enables simultaneous translation, scaling, and rotation transformations.
[0111] After mapping is completed, the module loads facial parameters from the parameter library according to the user's identifier, including muscle tension coefficient, soft tissue control parameters, and density distribution. Based on the muscle tension coefficient, stiffness characteristics are constructed, and based on the soft tissue control parameters, damping characteristics are constructed. The two constitute mechanical properties. Then, the facial soft tissue is divided into regional units such as lips and cheeks, and each unit is assigned mechanical properties and mass distribution, and assembled into a user-personalized facial soft tissue biomechanical model.
[0112] Next, facial soft tissue deformation data is generated. The mouth shape keypoint sequence is input into a biomechanical model, and the soft tissue deformation response driven by the keypoints is calculated using kinetic equilibrium equations. A synchronous deformation data sequence is output, with the equations as follows:
[0113]
[0114] in, This is the global quality matrix, reflecting the distribution of soft tissue quality. The acceleration vector of soft tissue; The global damping matrix characterizes energy dissipation; For soft tissue velocity vectors; This is the global stiffness matrix, reflecting the resistance to deformation. It is the displacement vector of the region element and the deformation core; The vector of external forces applied to the key points.
[0115] Simultaneously, skeletal node parameters are calculated. Skeletal nodes affected by key points are determined according to preset association rules, such as the mandibular node corresponding to the lower lip key point, and the degree of influence is calculated by the distance attenuation function, with closer nodes having a stronger influence. Then, the position offset and rotation changes of the skeletal nodes are obtained through vector operations.
[0116] The mouth shape rendering module 14 "calculates the position and rotation parameters of the corresponding skeleton nodes according to the preset correspondence relationship between the mouth shape and the skeleton, combines and renders the morphing data and the position and rotation parameters to generate a digital human mouth shape animation synchronized with the input speech signal” is specifically used to perform the following process:
[0117] According to the preset association rule of the mouth shape key point sequence and the skeleton node, determine the skeleton node affected by each key point and the corresponding influence degree; based on the influence degree, calculate the position offset and rotation change of the skeleton node; based on the morphing data and the offset and rotation change of the skeleton node, dynamically fuse the face soft tissue morphing and the face skeleton motion represented by the skeleton node, and perform rendering processing based on the fusion result to generate a preliminary mouth shape animation sequence; time synchronization control is performed on the preliminary mouth shape animation sequence and the input speech signal, to ensure that the time synchronization error between the mouth shape animation and the speech signal is within a preset threshold range, and a digital human mouth shape animation sequence synchronized with the input speech signal is output.
[0118] Among them, “based on the morphing data and the offset and rotation change of the skeleton node, dynamically fuse the face soft tissue morphing and the face skeleton motion represented by the skeleton node” is specifically through:
[0119] Layered fusion is performed on the soft tissue vertex displacement represented by the morphing data and the rigid motion generated by the skeleton node transformation to establish a layered fusion framework; in the layered fusion framework, according to the preset spatial distance between the vertex and the mouth core area, the morphing data weight value and the skeleton transformation weight value are calculated, wherein the morphing data weight value is negatively correlated with the spatial distance, and the skeleton transformation weight value is positively correlated with the spatial distance; based on the offset and rotation change of the skeleton node, the rigid transformation matrix corresponding to the skeleton node is calculated, and the rigid transformation position of the vertex is calculated based on the rigid transformation matrix; the local morphing position of the vertex is determined based on the morphing data; according to the morphing data weight value and the skeleton transformation weight value, the local morphing position and the rigid transformation position are weighted and superimposed; for the vertex affected by multiple skeleton nodes, the transformation contribution value of each skeleton node is calculated according to the binding weight of the vertex and each skeleton node, and the transformation contribution values are weighted and averaged.
[0120] In the above specific process, the preset correlation rule is an influence relationship and an action degree standard between the mouth shape key point and the face skeleton node set in advance; the position offset is a spatial displacement of the skeleton node relative to the initial position in the mouth shape movement; the rotation change is a rotation angle change of the skeleton node in the mouth shape movement; the face skeleton movement is a whole face skeleton movement composed of the position offset and the rotation change of the skeleton node; the preliminary mouth shape animation sequence is mouth shape animation data obtained by dynamic fusion and rendering, which is not time-synchronized; and the preset threshold is a maximum time synchronization error range allowed by the system.
[0121] The digital human mouth shape animation sequence is continuous mouth shape animation data finally output and synchronized with the voice; the hierarchical fusion framework is a technical framework for hierarchically fusing the soft tissue vertex displacement and the skeleton rigid movement; the spatial distance is a three-dimensional straight line distance between the face soft tissue vertex and the mouth core region; the deformation data weight value is a coefficient representing the influence degree of the deformation data on the final position of the vertex; the skeleton transformation weight value is a coefficient representing the influence degree of the skeleton movement on the final position of the vertex; the rigid transformation matrix is a mathematical matrix describing the rigid movement (translation + rotation) of the skeleton node; the rigid transformation position is a spatial position of the vertex obtained after the rigid movement of the skeleton; the local deformation position is a spatial position of the vertex only affected by the soft tissue deformation; and the transformation contribution value is an influence amount of a single skeleton node on the position of the vertex.
[0122] In the embodiment of the application, first, the module will establish a hierarchical fusion framework to integrate the soft tissue deformation data and the skeleton node movement. When fusing, the weight will be calculated according to the spatial distance between the face soft tissue vertex and the mouth core region. The closer the distance, the greater the influence weight of the deformation data on the vertex position and the smaller the influence weight of the skeleton movement. The weight relationship can be expressed as wherein is the deformation data weight range 0-1, is the skeleton transformation weight range 0-1.
[0123] Then, the module will calculate the rigid transformation position of the vertex under the skeleton movement through the translation and rotation matrix of the skeleton, calculate the local deformation position only affected by the soft tissue deformation through the displacement vector in the deformation data, and then add the two weighted to obtain the final position of the vertex. For the vertex affected by multiple skeleton nodes, such as the cheek vertex affected by the maxilla and the mandible, the transformation contribution value of each skeleton node will be calculated according to the binding weight of the vertex and each skeleton node and the preset correlation coefficient, and the total is 1, and then weighted average. After the fusion is completed, the module will render the result to generate a preliminary mouth shape animation sequence, adjust the animation playing rhythm through the time stamp alignment algorithm, ensure that the time synchronization error of the animation and the input voice is controlled within the preset threshold (such as below 50ms), and finally output the digital human mouth shape animation synchronized with the voice.
[0124] In practical applications, in the wheat stripe rust prevention and control online training scene of smart agriculture, the system needs to generate a digital human to explain the synchronous mouth shape animation of the related voice. The mouth shape rendering module 14 first receives the 46-dimensional mouth shape key point sequence output by the semantic enhancement module. The sequence contains 120 frames of coordinates, a frame rate of 25 fps, a total duration of 4.8 seconds, and clearly corresponds to the pronunciation frame segment of professional terms such as "triazole fungicide" and "sprayer uniform spraying" in the voice. The module calls the homogeneous transformation matrix to map the key points to the three-dimensional digital human face model, and adjusts the coordinates by a scaling factor of 0.1 and a translation correction offset, so that the lower lip midpoint key point accurately matches the corresponding position of the model. According to the user "F003" identifier, load the biomechanical data such as the lip muscle tension coefficient and the face soft tissue control parameter, divide the face soft tissue into 6 regional units, and assemble them into a personalized biomechanical model. Input the mapped key point sequence into the model to calculate the synchronous deformation data of the lip unit, such as the maximum displacement of 0.3mm when pronouncing the "fungi" word with the lower lip. At the same time, it is determined that the lower lip key point affects the mandibular node, and the position offset is 2.3mm and the rotation angle is 5°. Assign deformation and skeletal motion weights according to the distance from the face vertex, fuse the rendered preliminary animation, and correct the 35ms synchronization error through timestamp alignment. Finally, output a mouth shape animation with a resolution of 1920x1080 and a frame rate of 25fps, which can be directly embedded into the training video to assist in understanding professional terms.
[0125] Through the mouth shape rendering module 14, the digital human mouth shape can accurately fit the facial characteristics of the target user, the lip deformation is natural and has no harsh feeling, and the overall movement completely conforms to the human physiological law, and the mandibular rotation and the lip stretching action are coordinated and smooth. The mouth shape animation and the input voice are precisely synchronized, and there is no obvious misplacement between the sound and the picture; under high resolution, the mouth shape change corresponding to the pronunciation of agricultural professional terms is clear and distinguishable. After embedding it into the training video, the understanding of the voice content can be strengthened through visual aid, the learning threshold of agricultural professional knowledge is significantly reduced, the information transmission efficiency in the online training scene of smart agriculture is effectively improved, and the learning experience of the students is further optimized.
[0126] Figure 3 A flowchart of a specific implementation of a digital human mouth shape prediction system based on a variational autoencoder in smart agriculture provided by an embodiment of the present application is shown in Figure 3 The method can include:
[0127] S301, receiving an input voice signal generated in a smart agriculture scene, and extracting multi-dimensional voice features therefrom, and generating a voice representation by fusing the multi-dimensional voice features;
[0128] S302, inputting the voice representation into a variational autoencoder, obtaining a latent representation through the variational autoencoder, and decoding the latent representation into a mouth shape key point sequence;
[0129] S303, based on the agricultural specific term in the input speech signal, combining a preset standard mouth shape template, performing semantic enhancement on the mouth shape key point sequence, and performing time sequence filtering on the enhanced mouth shape key point sequence to suppress mouth shape jitter and abnormal frames.
[0130] S304, mapping the semantic enhanced mouth shape key point sequence to a three-dimensional digital human face model, combining different target users respectively corresponding face parameters, converting the mouth shape key point sequence into face soft tissue deformation data, and calculating the position and rotation parameters of the corresponding skeleton nodes according to the preset correspondence between the mouth shape and the skeleton, combining and rendering based on the deformation data and the position and rotation parameters, generating a digital human mouth shape animation synchronized with the input speech signal.
[0131] The smart agriculture based on the variational autoencoder digital human mouth shape prediction system of the embodiment of the application is used to realize the smart agriculture based on the variational autoencoder digital human mouth shape prediction method described above, and therefore the specific embodiments in the smart agriculture based on the variational autoencoder digital human mouth shape prediction system can be seen from the embodiment part of the smart agriculture based on the variational autoencoder digital human mouth shape prediction method described above, and the specific embodiments can be referred to the description of the corresponding embodiment part, which will not be repeated here.
[0132] The application also provides an electronic device, comprising: a memory for storing a computer program; a processor for executing the computer program to realize the steps of the smart agriculture based on the variational autoencoder digital human mouth shape prediction method described above.
[0133] The application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the steps of the smart agriculture based on the variational autoencoder digital human mouth shape prediction method described above.
[0134] In an exemplary embodiment, the above computer readable storage medium can include, but is not limited to, a U disk, a read-only memory, a random access memory, a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0135] The embodiment of the application also provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to realize the steps in the smart agriculture based on the variational autoencoder digital human mouth shape prediction method described above.
[0136] Those skilled in the art will further realize that the mere concepts, teachings, and embodiments described herein are merely meant to provide an enabling description of embodiments of the present application and are not intended to limit the scope of the present application. Accordingly, embodiments as described herein contemplate all modifications that come within the scope of the present application.
[0137] The above provides a detailed introduction to the digital human mouth shape prediction system and method based on variational autoencoder in smart agriculture. The principles and implementation methods of the present application are described by applying specific examples. The above description of the embodiments is only used to help understand the method and its core idea. It should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the present application.
Claims
1. A digital human mouth shape prediction system based on a variational autoencoder in smart agriculture, characterized in that, The method comprises the following steps: a speech feature extraction module, a variational autoencoder mapping module, a semantic enhancement module, and a lip rendering module: The speech feature extraction module is configured to receive an input speech signal generated in a smart agriculture scenario and extract multi-dimensional speech features from the input speech signal, and fuse the multi-dimensional speech features to generate a speech representation; The variational autoencoder mapping module is configured to input the speech representation into a variational autoencoder, obtain a latent representation through the variational autoencoder, and decode the latent representation into a sequence of lip key points; The semantic enhancement module is configured to perform semantic enhancement on the sequence of lip key points based on an agricultural specific term in the input speech signal and a preset standard lip template, and perform time series filtering on the enhanced sequence of lip key points to suppress lip shaking and abnormal frames; The lip rendering module is configured to map the enhanced sequence of lip key points to a three-dimensional digital human face model, combine respective face parameters of different target users, the face parameters including related data describing face muscle and soft tissue motion characteristics, convert the sequence of lip key points into deformation data of face soft tissue, calculate position and rotation parameters of corresponding bone nodes according to a preset correspondence between lip and bone, and perform combined rendering based on the deformation data and the position and rotation parameters to generate a digital human lip animation synchronized with the input speech signal; In the process of performing semantic enhancement on the sequence of lip key points based on the agricultural specific term in the input speech signal and the preset standard lip template, and performing time series filtering on the enhanced sequence of lip key points to suppress lip shaking and abnormal frames, the semantic enhancement module is specifically configured to perform the following processes: Text conversion is performed on the input speech signal, and an agricultural specific term in the converted text is identified, and a corresponding standard lip template is extracted from a preset template library according to the agricultural specific term; Similarity calculation is performed on the sequence of lip key points and the standard lip template, and adaptive weighted adjustment is performed on key point positions related to the agricultural specific term in the sequence of lip key points based on a calculation result to generate a semantically enhanced sequence of lip key points; Continuous consistency analysis is performed on the key points in the enhanced sequence of lip key points, and smoothing processing is performed on a key point trajectory based on an analysis result to suppress lip shaking and eliminate abnormal frames that do not conform to human face motion rules.
2. The system of claim 1, wherein, In the process of mapping the enhanced sequence of lip key points to a three-dimensional digital human face model and converting the sequence of lip key points into deformation data of face soft tissue based on respective face parameters of different target users, the lip rendering module is specifically configured to perform the following processes: The enhanced sequence of lip key points is mapped to a three-dimensional digital human face model; According to a target user identifier, corresponding face parameters are loaded from a parameter database, associated motion characteristics of face soft tissue are established based on the face parameters, and a biomechanical model of face soft tissue conforming to individual characteristics of the target user is established; The mouth shape key point sequence is input into the biomechanical model, and deformation response of facial soft tissue under key point driving is calculated according to the associated motion characteristics, so as to output a deformation data sequence corresponding to the mouth shape key point sequence.
3. The system of claim 2, wherein, In the process of establishing the associated motion characteristics of the facial soft tissue based on the facial parameters and simultaneously establishing the biomechanical model of the facial soft tissue conforming to the individual characteristics of the target user, the mouth shape rendering module is specifically configured to perform the following process: Based on the muscle tension coefficient in the facial parameters, the muscle tension parameter representing the contraction strength of the facial muscles during movement, the deformation retention characteristics describing the deformation resistance of the facial soft tissue in each direction are established; Based on the soft tissue control parameter in the facial parameters, the soft tissue control parameter representing the internal resistance of the facial soft tissue during deformation, the motion stability characteristics during the movement of the facial soft tissue are established; The deformation retention characteristics and the motion stability characteristics are combined to form the associated motion characteristics; Based on the density distribution of the facial soft tissue in the facial parameters, the thickness distribution characteristics of each part of the facial soft tissue are determined; The facial soft tissue is divided into a plurality of connected region units, and each region unit is assigned with corresponding associated motion characteristics and thickness distribution characteristics according to the anatomical position; Unit morphology data are established according to the associated motion characteristics of the region units, and unit distribution data are established according to the thickness distribution characteristics; Based on the unit morphology data and the unit distribution data of all region units, the biomechanical model of the facial soft tissue is established by assembling.
4. The system of claim 1, wherein, In the process of performing the combination rendering based on the deformation data and the position and rotation parameters of the skeleton nodes to generate the digital human mouth shape animation synchronized with the input speech signal, the mouth shape rendering module is specifically configured to perform the following process: According to the preset association rule between the mouth shape key point sequence and the skeleton nodes, the skeleton nodes affected by each key point and the corresponding influence degree are determined; Based on the influence degree, the position offset and rotation change of the skeleton nodes are calculated; Based on the deformation data and the offset and rotation change of the skeleton nodes, the facial soft tissue deformation and the facial skeleton movement represented by the skeleton nodes are dynamically fused, and the preliminary mouth shape animation sequence is generated based on the fusion result; The preliminary mouth shape animation sequence and the input speech signal are time-synchronized to ensure that the time synchronization error between the mouth shape animation and the speech signal is within a preset threshold range, and a digital human mouth shape animation sequence synchronized with the input speech signal is output.
5. The system of claim 4, wherein, In the process of performing the dynamic fusion of the facial soft tissue deformation and the facial skeleton movement represented by the skeleton nodes based on the deformation data and the offset and rotation change of the skeleton nodes, the mouth shape rendering module is specifically configured to perform the following process: The soft tissue vertex displacement represented by the deformation data and the rigid motion generated by the skeleton node transformation are hierarchically fused to establish a hierarchical fusion framework; In the layered fusion framework, a deformation data weight value and a bone transformation weight value are calculated according to the spatial distance between the vertex and the mouth core region, wherein the deformation data weight value is negatively correlated with the spatial distance, and the bone transformation weight value is positively correlated with the spatial distance; Based on the offset and rotation change of the bone node, a rigid transformation matrix corresponding to the bone node is calculated, and a rigid transformation position of the vertex is calculated based on the rigid transformation matrix; The local deformation position of the vertex is determined based on the deformation data; The local deformation position and the rigid transformation position are weighted and superimposed according to the deformation data weight value and the bone transformation weight value; For the vertex affected by multiple bone nodes, the transformation contribution value of each bone node is calculated according to the binding weight of the vertex and each bone node, and the transformation contribution values are weighted and averaged to obtain the fusion result.
6. The system of claim 1, wherein, The variational autoencoder mapping module is specifically configured to perform the following processes in the process of inputting the speech representation into the variational autoencoder, obtaining the latent representation through the variational autoencoder, and decoding the latent representation into the sequence of mouth key points: The speech representation is input into the first network path of the variational autoencoder, and the speech representation is feature-transformed to obtain a feature distribution of the latent representation; Based on the feature distribution, the latent representation is obtained through a random sampling operation; The latent representation is input into the second network path of the variational autoencoder, and the latent representation is layer-by-layer feature-expanded to obtain a feature sequence matching the dimension of the mouth key point; The feature sequence is subjected to coordinate mapping to output the sequence of mouth key points.
7. A method for predicting the mouth shape of a digital human in smart agriculture based on a variational autoencoder, applied to the smart agriculture-based digital human mouth shape prediction system of any one of claims 1-6, characterized in that, Comprising: Receiving an input speech signal generated in a smart agriculture scene, and extracting multi-dimensional speech features therefrom, and fusing the multi-dimensional speech features to generate a speech representation; Inputting the speech representation into a variational autoencoder, obtaining a latent representation through the variational autoencoder, and decoding the latent representation into a sequence of mouth key points; Based on the agricultural specific terms in the input speech signal, combining a preset standard mouth template, performing semantic enhancement on the sequence of mouth key points, and performing time series filtering on the enhanced sequence of mouth key points to suppress mouth shaking and abnormal frames; Mapping the semantically enhanced sequence of mouth key points to a three-dimensional digital human face model, combining the respective face parameters of different target users, converting the sequence of mouth key points into deformation data of the soft tissue of the face, and calculating the position and rotation parameters of the corresponding bone nodes according to the preset correspondence between the mouth and the skeleton, and combining the deformation data and the position and rotation parameters for rendering to generate a digital human mouth animation synchronized with the input speech signal.
8. An electronic device, comprising: Comprising: A memory for storing a computer program; A processor for executing the computer program to implement the steps of the smart agriculture-based digital human mouth prediction method based on a variational autoencoder according to claim 7.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the digital human mouth shape prediction method based on a variational autoencoder in smart agriculture according to claim 7.
Citation Information
Patent Citations
Voice-driven face key point sequence generation method and device
CN115187705A
Virtual digital population synchronization method, system, device, medium and program product
CN119517071A