Digital human expression driving method and device and electronic equipment
By combining facial feature representation and temporal prediction models with adaptive selection of target coefficient sets based on system delay, the real-time and naturalness issues in digital human facial expression driving are solved, achieving high-quality facial expression synchronization and fluency.
Patent Information
- Application Number
- CN202511481564.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-16
- Publication Date
- 2026-02-17
AI Technical Summary
Existing digital human facial expression-driven technologies suffer from poor real-time performance and stiff facial expression transitions, making it difficult to achieve high-quality real-time interaction.
The facial expression feature representation model extracts facial expression coefficient features, the time-series facial expression prediction model predicts the future coefficient set, and the current system delay adaptive selection of the target coefficient set drives the digital human model to achieve synchronized and natural transition of facial expressions.
It improves the real-time performance and naturalness of digital human expressions, reduces facial jiggling, adapts to latency variations under different hardware configurations, and maintains good driving performance.
Smart Images

Figure CN121545196A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of digital human, in particular to a digital human expression driving method and device and electronic equipment. BACKGROUND
[0002] Visual expression driving technology is a widely researched and applied technology, which refers to driving a two-dimensional or three-dimensional digital human model to make corresponding expressions based on the facial expressions and actions in the video. The most common one is various virtual headgear applications. The existing visual expression driving scheme has poor real-time performance and certain delay problems due to high computational complexity or the use of single-frame recognition scheme. SUMMARY
[0003] The present application provides a digital human expression driving method, device and electronic equipment.
[0004] In a first aspect, the present application provides a digital human expression driving method, comprising: obtaining a collected face image, performing coefficient feature extraction on the face image through an expression feature representation model to obtain expression coefficient features; based on a plurality of the expression coefficient features, obtaining a feature sequence, performing coefficient prediction through a time series expression prediction model based on the feature sequence to obtain a plurality of future coefficient sets of different time spans; based on the time consumption data of the last digital human expression driving processing period, obtaining a current system delay, and based on the current system delay, determining a target coefficient set from the plurality of future coefficient sets of different time spans; driving a digital human model based on the target coefficient set.
[0005] In a second aspect, the present application provides a digital human expression driving device, comprising: a feature acquisition module configured to obtain a collected face image, perform coefficient feature extraction on the face image through an expression feature representation model to obtain expression coefficient features; a coefficient prediction module configured to obtain a feature sequence based on a plurality of the expression coefficient features, perform coefficient prediction through a time series expression prediction model based on the feature sequence to obtain a plurality of future coefficient sets of different time spans; a coefficient determination module configured to obtain a current system delay based on the time consumption data of the last digital human expression driving processing period, and determine a target coefficient set from the plurality of future coefficient sets of different time spans based on the current system delay; an expression driving module configured to drive a digital human model based on the target coefficient set.
[0006] In a third aspect, the present application provides an electronic device, comprising: One or more processors; The processor is used to invoke instructions to cause the electronic device to perform the method described in the first aspect above.
[0007] Fourthly, embodiments of this application provide a storage medium storing instructions that, when executed on an electronic device, cause the electronic device to perform the method described in the first aspect above.
[0008] Fifthly, embodiments of this application provide a program product, including at least one of a program and instructions, wherein when the program and instructions are executed by an electronic device, they implement the steps of the method described in the first aspect.
[0009] According to the technical solution of this application, multiple facial images are subjected to coefficient feature extraction using an expression feature representation model to obtain expression coefficient features. A feature sequence is constructed based on the most recent multiple expression coefficient features, and a temporal expression prediction model is used to predict coefficients based on the feature sequence, resulting in multiple future coefficient sets with different time spans. Based on dynamically acquired current system latency, a target coefficient set is adaptively selected from multiple future coefficient sets with different time spans. The digital human model is driven by the prediction results of future multiple frames in the target coefficient set, which can effectively compensate for the inherent system latency, making the digital human's expression response almost synchronized with the user's expression, improving the real-time performance of the digital human driver and enhancing the user experience. Furthermore, this solution uses historical multi-frame features to predict future multi-frame prediction results, and drives the digital human model based on these prediction results, effectively smoothing expression transitions, reducing jitter that may be caused by single-frame recognition, improving the naturalness and fluency of expressions, and enhancing the user experience. In addition, this solution dynamically acquires the current system latency, exhibiting strong latency adaptability, dynamically adapting to dynamic latency changes under different hardware configurations and system loads, maintaining good driving performance, and demonstrating strong robustness.
[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0011] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0012] Figure 1 This is a flowchart illustrating a digital human expression-driven method according to an exemplary embodiment.
[0013] Figure 2 This is a flowchart illustrating a digital human expression-driven method according to another exemplary embodiment.
[0014] Figure 3 is a block diagram of a digital human expression driving apparatus according to an exemplary embodiment.
[0015] Figure 4 is a block diagram of a digital human expression driving apparatus 400 according to an exemplary embodiment. DETAILED DESCRIPTION
[0016] Embodiments of the present application are described below in detail, examples of which are shown in the drawings, in which the same or similar notations represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by reference to the drawings are exemplary and are intended to explain the present application, and are not to be understood as limiting the present application.
[0017] The embodiments of the present application are not exhaustive, and are only schematic of some embodiments, and are not specific limitations on the scope of protection of the present application. In the case of no contradiction, each step in an embodiment can be implemented as an independent embodiment, and the steps can be combined arbitrarily, for example, the scheme after removing some steps in an embodiment can be implemented as an independent embodiment, and the order of the steps in an embodiment can be exchanged arbitrarily, in addition, the optional implementation in an embodiment can be combined arbitrarily; in addition, the embodiments can be combined arbitrarily, for example, the steps of different embodiments or part of the steps can be combined arbitrarily, an embodiment can be combined with the optional implementation of other embodiments.
[0018] In each embodiment of the present application, the terms and / or descriptions of the embodiments are consistent and can be referred to each other if there is no special description and logical conflict, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.
[0019] It should be noted that the acquisition, transmission, storage, use, processing, etc. of data in the technical solutions of the present application comply with the relevant provisions of national laws and regulations, and do not violate public order and good customs.
[0020] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present application are authorized by the user or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0021] It is worth noting that in the embodiments of the present application, some software, components, models, etc. may be mentioned in the prior art, which should be considered as exemplary, and the purpose is only to illustrate the feasibility of the implementation of the technical solutions of the present application, but it does not mean that the applicant has or will necessarily use the scheme.
[0022] The current digital human expression driving technology usually adopts a single-frame recognition scheme, that is, the expression of each frame of human face image is recognized, the blendshape coefficient (abbreviated as bs coefficient) of the expression is extracted, and the three-dimensional digital human model is driven accordingly. The bs coefficient is usually a vector with tens to hundreds of dimensions, and through these coefficients, the three-dimensional model can be finely driven to express the same expression as the real person through the camera. However, the prior art has the following defects: first, due to the delay of the overall pipeline of data transmission, calculation and driving, the digital human model expression presents obvious hysteresis, which performs poorly in real-time driving scene, that is, the real-time performance is poor, and there is a certain delay problem; second, based on the single-frame recognition scheme, the context understanding of the expression change between consecutive frames is lacking, which may cause the expression transition to be harsh, affect the naturalness of expression driving, and the recognition result has a certain degree of jitter in time sequence, resulting in harsh expression driving and affecting the naturalness of expression driving.
[0023] In summary, the prior art still has challenges in dealing with dynamic system delay, achieving real-time response, and ensuring the continuous and natural flow of expression driving, and it is difficult to meet the needs of high-quality real-time interaction.
[0024] Based on this, the embodiments of the present application provide a digital human expression driving method, device and electronic equipment, which can improve the real-time performance of digital human driving and improve the naturalness and fluency of expression.
[0025] The digital human expression driving method, device and electronic equipment of the embodiments of the present application are described below with reference to the accompanying drawings.
[0026] It should be noted that the execution subject of the digital human expression driving method of the embodiments of the present application can be a digital human expression driving device, which can be realized by software and / or hardware. The device can be configured in an electronic equipment, which can include but is not limited to a terminal, a server, etc.
[0027] Figure 1 is a flowchart of a digital human expression driving method according to an exemplary embodiment. As Figure 1 shown, the digital human expression driving method can include but is not limited to the following steps 101-104.
[0028] In step 101, the collected face image is acquired, and the coefficient feature of the face image is extracted through the expression feature representation model to obtain the expression coefficient feature.
[0029] In some embodiments, the human face image can be collected in real time through a camera, and the human face image is subjected to face detection and cropping processing to obtain the face image.
[0030] In some embodiments, the expression coefficient feature is a bs coefficient feature, and each frame of the face image is input into the expression feature representation model to obtain the bs coefficient feature.
[0031] In some embodiments, a data set needs to be acquired, and the deep learning regression model is trained through the data set to obtain the expression feature representation model.
[0032] In some embodiments, the training method of the expression feature representation model comprises: acquiring a data set, the data set comprising a plurality of data samples, the data sample comprising a face image and a corresponding real expression coefficient; extracting spatial features of the face image through a convolutional neural network; marking spatial features belonging to a key region of the face in the spatial features through a attention mechanism module; extracting features of the marked spatial features through a feature extraction layer to obtain expression features; mapping the expression features to an expression coefficient space through a fully connected layer to obtain predicted expression coefficient features; and adjusting parameters of the convolutional neural network based on a difference between the predicted expression coefficient features and the real expression coefficient.
[0033] In some embodiments, the method for acquiring the data set comprises: collecting real face video data, acquiring a plurality of face images from the real face video data, and labeling each frame of the face image using an existing expression recognition algorithm to obtain corresponding expression bs coefficients, or directly acquiring high-precision expression bs coefficients of the video frame through a professional motion capture device, and constructing a data set of one-to-one correspondence between the face image and the expression bs coefficient based on the acquired large number of face images and the corresponding expression bs coefficients.
[0034] In some embodiments, the model structure of the deep learning regression model for training the expression feature representation model includes a convolutional neural network (CNN) module, an attention mechanism module, a feature extraction layer, and a fully connected layer. The CNN module is used to extract spatial features of a face image. The attention mechanism module is used to focus on key regions of the face. The feature extraction layer is used for key feature extraction. The fully connected layer is used to map the key features to a bs coefficient space. The training target of the deep learning regression model in the training process is to minimize the mean square error between the predicted expression bs coefficients and the labeled real expression bs coefficients. After training, the deep learning regression model removes the fully connected layer, and the output is the expression bs coefficient feature for describing the facial expression action. That is, for each input face image, the forward part of the deep learning regression model (excluding the fully connected layer) is used to obtain the expression bs coefficient feature.
[0035] In some embodiments, after obtaining the expression coefficient feature, the expression coefficient feature is stored in a feature pool, and the feature pool is used to store expression coefficient features of a plurality of recent face images.
[0036] In step 102, based on the plurality of expression coefficient features, a feature sequence is obtained, and based on the feature sequence, coefficient prediction is performed through a time-series expression prediction model to obtain a plurality of future coefficient sets of different time spans.
[0037] In some embodiments, a plurality of expression coefficient features are obtained from the feature pool, and based on the obtained plurality of expression coefficient features, a feature sequence is constructed.
[0038] In this embodiment, the time-series expression prediction model is used to predict BS coefficients of M future frames of different time spans in one step based on past multiple frames of facial expressions.
[0039] In some embodiments, the model architecture of the time-series expression prediction model is a Transformer encoder-decoder architecture. The encoder is used to receive a feature sequence of expression bs coefficients of historical N frames of face images (for example, N=10). The decoder has M parallel output heads, each output head corresponding to a specific future time span ki (for example, k1=1 frame, k2=2 frames,..., kM=5 frames, corresponding to an advance of about 16 ms, 33 ms,..., 83 ms, respectively, at a frame rate of 60 FPS), that is, a plurality of different time spans are represented by frame numbers.
[0040] In some embodiments, the method for predicting the coefficients based on the feature sequence by the time-series expression prediction model to obtain the future coefficient set of multiple different time spans includes: receiving the feature sequence by the encoder in the time-series expression prediction model and performing encoding processing to obtain an encoded feature sequence; performing decoding processing on the encoded feature sequence by multiple parallel output heads in the time-series expression prediction model to obtain the future coefficient set of multiple different time spans, wherein the multiple parallel output heads correspond to the multiple different time spans.
[0041] In some embodiments, the training method of the time-series expression prediction model includes: taking the sum of the mean square errors between the bs coefficients predicted by all output heads and the real bs coefficients as the target, constructing multiple feature sequences by the multiple expression bs coefficient features obtained by the expression feature representation model, training the neural network of the Transformer encoder-decoder architecture to obtain the time-series expression prediction model, the input of the trained time-series expression prediction model being the feature sequence constructed by the multiple expression bs coefficient features, and the output being the future bs coefficient set of M different time spans, wherein the future bs coefficient set of M different time spans can be represented as {BS_coeffs_t+k1, BS_coeffs_t+k2,..., BS_coeffs_t+kM}.
[0042] In step 103, the current system delay is obtained based on the time consumption data of the last digital human expression driving processing period, and the target coefficient set is determined from the future coefficient set of multiple different time spans based on the current system delay.
[0043] In some embodiments, the time consumption data of the last digital human expression driving processing period can be obtained by monitoring the current expression driving pipeline in real time to obtain the timestamps of each key node, and the current system delay can be determined based on the time consumption data.
[0044] In some embodiments, the target coefficient set is determined from the future coefficient set of multiple different time spans based on the current system delay, including: obtaining the target future coefficient set with the smallest difference from the current system delay from the future coefficient set of multiple different time spans. That is, the prediction lead time of the selected target coefficient set is smaller than the current system delay and closest to the current system delay.
[0045] In step 104, the digital human model is driven based on the target coefficient set.
[0046] In some embodiments, after obtaining the target bs coefficient set, the digital human model is driven by the target bs coefficient set to present the facial expression and complete the digital human expression driving.
[0047] In the above embodiments, the expression coefficient features are extracted from the collected multiple facial images by the expression feature representation model, and the expression coefficient features are obtained; a feature sequence is constructed based on the nearest multiple expression coefficient features, and the coefficient prediction is performed based on the feature sequence by the time-series expression prediction model, and multiple future coefficient sets of different time spans are obtained; based on the dynamically obtained current system delay, the target coefficient set is adaptively selected from the multiple future coefficient sets of different time spans, and the digital human model is driven according to the future multiple frame prediction results in the target coefficient set, which can effectively compensate for the inherent delay of the system, so that the digital human expression response is almost synchronized with the user expression, the real-time performance of the digital human driving is improved, and the user experience is improved; moreover, the future multiple frame prediction results are obtained based on the historical multiple frame features, and the digital human model is driven according to the future multiple frame prediction results, which can effectively smooth the expression transition, reduce the jitter caused by single frame recognition, improve the naturalness and fluency of the expression, and improve the user experience; in addition, the current system delay is dynamically obtained, which has strong delay adaptability and can dynamically adapt to the dynamic delay changes under different hardware configurations and system loads, and always maintains good driving effect and strong robustness. The method used in the present application has wide applicability, and is not only suitable for expression driving, but also can be extended to limb driving and other visual driving scenes that require super real-time response.
[0048] In order to further clearly illustrate how the above step 103 obtains the current system delay and how to adaptively determine the target coefficient set from the multiple future coefficient sets of different time spans, in any one or combination of the above embodiments, as shown in the following step 203, the digital human expression driving method can include but is not limited to the following steps 203-207. Figure 2
[0049] In step 201, the collected facial image is obtained, and the coefficient features are extracted from the facial image by the expression feature representation model to obtain the expression coefficient features.
[0050] It should be noted that the implementation of the present step 201 can refer to the step 101 in the above embodiments for details, and the principle is the same, which will not be repeated here.
[0051] In step 202, based on the multiple expression coefficient features, a feature sequence is obtained, and based on the feature sequence, the coefficient prediction is performed by the time-series expression prediction model to obtain multiple future coefficient sets of different time spans.
[0052] It should be noted that the implementation of the present step 202 can refer to the step 102 in the above embodiments for details, and the principle is the same, which will not be repeated here.
[0053] In step 203, the timestamps of a plurality of key nodes of the last digital human expression driving processing cycle are obtained, the plurality of key nodes at least including a human face image acquisition node and a driving completion node.
[0054] In some embodiments, the timestamps of the human face image acquisition node and the driving completion node of the last digital human expression driving processing cycle can be obtained, so that the difference between the timestamp of the driving completion node and the timestamp of the human face image acquisition node is the total time consumption.
[0055] In some embodiments, the digital human expression driving processing cycle is divided into a plurality of stages, and the timestamps of the starting key nodes of the stages are obtained.
[0056] In step 204, based on the timestamps of the plurality of key nodes, the total time consumption of the last digital human expression driving processing cycle is obtained; and the total time consumption is taken as the current system delay.
[0057] For example, the system delay perception module monitors the end-to-end delay of the current expression driving pipeline in real time to dynamically obtain the current system delay; the current system delay is the total time consumption from the start of image acquisition to the completion of digital human driving, which can include image acquisition delay, expression feature representation model inference delay, timing expression prediction model inference delay, BS coefficient transmission delay, rendering engine driving delay, etc. Time stamps are stamped at key nodes of the data flow (such as image acquisition, expression feature representation model output, timing expression prediction model output, and rendering driving before), and the delay of each stage and the total delay are obtained by calculating the timestamp difference. The total time consumption can be directly obtained according to the timestamps of the two ends of the last cycle of the current expression driving pipeline. The total time consumption can also be obtained according to the delay of each stage between a plurality of key nodes.
[0058] In step 205, the current system delay is converted into an equivalent frame number based on the frame interval.
[0059] For example, the current system delay D_system (e.g., in milliseconds) confirmed by the system dynamic perception module is converted into an equivalent frame number D_frames = D_system / frame_interval, where frame_interval is the time interval between adjacent two frames, such as 16.67 ms.
[0060] In step 206, from a plurality of different time span future coefficient sets, a target future coefficient set is obtained, which has a time span less than or equal to the equivalent frame number and has the smallest difference with the equivalent frame number.
[0061] For example, from the M future bs coefficient sets {BS_coeffs_t+k1, BS_coeffs_t+k2,..., BS_coeffs_t+kM} output by the time sequence expression prediction model, an optimal future bs coefficient set is selected according to D_frames, and the future bs coefficient set that makes the prediction lead time, i.e., the future time span k, closest to and less than D_frames can be selected, such as BS_coeffs_t+k2.
[0062] In step 207, the digital human model is driven based on the target coefficient set.
[0063] It should be noted that the implementation of the present step 207 can refer to the step 104 in the above-mentioned embodiments for details, and the principle is the same, which will not be repeated here.
[0064] In the above-mentioned embodiments, by monitoring the time stamps of the key nodes of the current expression driving pipeline, the current system delay is dynamically obtained, which can improve the accuracy and adaptability of the current system delay; by converting the current system delay into an equivalent frame number, from a plurality of future coefficient sets with different time spans, a target future coefficient set with a time span less than the equivalent frame number and a difference from the equivalent frame number is obtained, which can obtain a more accurate prediction lead time, thereby further improving the real-time performance of digital human driving and improving user experience.
[0065] Figure 3 is a block diagram of a digital human expression driving device according to an exemplary embodiment. As shown in Figure 3 The digital human expression driving device can include a feature acquisition module 310, a coefficient prediction module 320, a coefficient determination module 330, and an expression driving module 340.
[0066] The feature acquisition module 310 is configured to acquire a collected face image, perform coefficient feature extraction on the face image through an expression feature representation model, and obtain expression coefficient features. The coefficient prediction module 320 is configured to obtain a feature sequence based on a plurality of expression coefficient features, perform coefficient prediction through a time sequence expression prediction model based on the feature sequence, and obtain a plurality of future coefficient sets with different time spans. The coefficient determination module 330 is configured to obtain a current system delay based on time consumption data of a last digital human expression driving processing period, and determine a target coefficient set from a plurality of future coefficient sets with different time spans based on the current system delay. The expression driving module 340 is configured to drive a digital human model based on the target coefficient set.
[0067] In some implementations, the coefficient determination module 330, when obtaining the current system latency based on the time consumption data of the last digital human expression driving processing period, is configured to: obtain timestamps of a plurality of key nodes of the last digital human expression driving processing period, the plurality of key nodes at least including a face image acquisition node and a driving completion node; obtain total time consumption of the last digital human expression driving processing period based on the timestamps of the plurality of key nodes; and take the total time consumption as the current system latency.
[0068] In some implementations, the coefficient determination module 330, when determining the target coefficient set from the plurality of future coefficient sets of different time spans based on the current system latency, is configured to: obtain, from the plurality of future coefficient sets of different time spans, a target future coefficient set with a time span smaller than the current system latency and a smallest difference from the current system latency.
[0069] In some implementations, the plurality of different time spans are represented by frame numbers, and the coefficient determination module 330, when determining the target coefficient set from the plurality of future coefficient sets of different time spans based on the current system latency, is configured to: convert the current system latency into an equivalent frame number based on a frame interval; obtain, from the plurality of future coefficient sets of different time spans, a target future coefficient set with a time span smaller than the equivalent frame number and a smallest difference from the equivalent frame number.
[0070] In some implementations, the feature obtaining module 310, after obtaining the expression coefficient feature, is further configured to: store the expression coefficient feature into a feature pool, the feature pool being configured to store expression coefficient features of a plurality of recent face images.
[0071] In some implementations, the apparatus further includes a model training module 350, and the model training module 350, when training the expression feature representation model, is configured to: obtain a data set, the data set including a plurality of data samples, each data sample including a face image and a corresponding real expression coefficient; extract spatial features of the face image through a convolutional neural network; label spatial features belonging to a key region of the face in the spatial features through an attention mechanism module; extract features of the labeled spatial features through a feature extraction layer to obtain expression features; map the expression features to an expression coefficient space through a fully connected layer to obtain predicted expression coefficient features; adjust parameters of the convolutional neural network based on a difference between the predicted expression coefficient features and the real expression coefficient.
[0072] In some implementations, the coefficient prediction module 320, when performing coefficient prediction by the temporal expression prediction model to obtain a plurality of future coefficient sets of different time spans, is configured to: receive the feature sequence by an encoder in the temporal expression prediction model and perform encoding processing to obtain an encoded feature sequence; perform decoding processing on the encoded feature sequence by a plurality of parallel output heads in the temporal expression prediction model to obtain a plurality of future coefficient sets of different time spans, wherein the plurality of parallel output heads correspond to the plurality of different time spans.
[0073] As to the apparatus in the above-described embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and thus will not be described in detail here.
[0074] Figure 4 is a block diagram of an apparatus 400 for implementing digital human expression driving according to an example embodiment. For example, the apparatus 400 can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, and the like.
[0075] Referring to Figure 4 , the apparatus 400 can include one or more of the following components: a processing component 402, a memory 404, a power supply component 406, a multimedia component 408, an audio component 410, an input / output (I / O) interface 412, a sensor component 414, and a communication component 416.
[0076] The processing component 402 usually controls overall operations of the apparatus 400, such as operations associated with displaying, making phone calls, data communications, camera operations, and recording operations. The processing component 402 can include one or more processors 420 to execute instructions to complete all or part of steps of the methods described above. In addition, the processing component 402 can include one or more modules to facilitate interaction between the processing component 402 and other components. For example, the processing component 402 can include a multimedia module to facilitate the interaction between the multimedia component 408 and the processing component 402.
[0077] The memory 404 is configured to store various types of data to support the operation of the device 400. Examples of such data include instructions for any application or methods operating on the device 400, contact data, phonebook data, messages, pictures, videos, and the like. The memory 404 can be implemented by any type of volatile or nonvolatile storage devices or a combination thereof such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read only memory (PROM), read only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0078] The power component 406 provides power to the various components of the device 400. The power component 406 can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 400.
[0079] The multimedia component 408 includes a screen providing an output interface between the device 400 and a user. In some embodiments, the screen includes a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touch, swiping, and gestures on the touch panel. The touch sensor can not only sense a boundary of a touching or swiping action, but also detect duration and pressure related to the touching or swiping action. In some embodiments, the multimedia component 408 includes a front camera and / or a rear camera. The front and / or rear camera can receive external multimedia data when the device 400 is in an operation mode, such as a shooting mode or a video mode. Each of the front and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0080] The audio component 410 is configured to output and / or input audio signals. For example, the audio component 410 includes a microphone (MIC) that is configured to receive external audio signals when the device 400 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 404 or transmitted via the communication component 416. In some embodiments, the audio component 410 also includes a speaker for outputting audio signals.
[0081] The I / O interface 412 provides an interface between the processing component 402 and peripheral interface modules, which can be a keyboard, a click wheel, buttons, and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.
[0082] The sensor component 414 includes one or more sensors to provide status assessments for various aspects of the device 400. For example, the sensor component 414 can detect an open / closed position of the device 400, relative positioning of components, such as a display and keypad of the device 400, a change in position of the device 400 or a component of the device 400, the presence or absence of user contact with the device 400, the orientation or acceleration / deceleration of the device 400, and a temperature change of the device 400. The sensor component 414 can include proximity sensor(s) configured to detect the presence of objects in a proximity without any physical contact. The sensor component 414 can also include a light sensor, such as a CMOS or CCD image sensor, utilized in imaging applications. In some embodiments, the sensor component 414 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0083] The communication component 416 is configured to facilitate wired or wireless communication between the device 400 and another device. The device 400 can access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 416 receives a broadcast signal or broadcast related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 416 also includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra-WideBand (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0084] In an exemplary embodiment, the device 400 can be implemented using one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, or other electronic units to perform the above-described methods.
[0085] In an exemplary embodiment, a non-transitory computer readable storage medium, such as the memory 404 including instructions, is also provided, which can be executed by the processor 420 of the device 400 to perform the above-described methods. For example, the non-transitory computer readable storage medium can be a ROM, a RAM, a CD-ROM, a magnetic tape, a floppy disc, and an optical data storage device, etc.
[0086] In an exemplary embodiment, a program product including at least one of a program and an instruction, which is executed by the processor 420 of the device 400 to implement the steps of the above-described methods, is also provided.
[0087] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the application being indicated by the following claims.
[0088] It is to be understood that the application is not limited to the precise construction herein described and as shown in the attached drawings, and that various modifications and changes can be made by those skilled in the art without departing from the scope of the application. The scope of the application is limited only by the claims that follow.
Claims
1. A method for driving digital human expressions, the method comprising: The method comprises: obtaining a collected face image, performing coefficient feature extraction on the face image through an expression feature representation model to obtain expression coefficient features; based on a plurality of expression coefficient features, obtaining a feature sequence, and performing coefficient prediction through a time-series expression prediction model based on the feature sequence to obtain a plurality of future coefficient sets of different time spans; based on the time consumption data of the last digital human expression driving processing cycle, obtaining a current system delay, and based on the current system delay, determining a target coefficient set from the plurality of future coefficient sets of different time spans; based on the target coefficient set, driving a digital human model.
2. The method of claim 1, wherein, The method comprises: obtaining the time stamps of a plurality of key nodes of the last digital human expression driving processing cycle, the plurality of key nodes at least including a face image collection node and a driving completion node; based on the time stamps of the plurality of key nodes, obtaining the total time consumption of the last digital human expression driving processing cycle; and taking the total time consumption as the current system delay.
3. The method of claim 1, wherein, The method comprises: from the plurality of future coefficient sets of different time spans, obtaining a target future coefficient set with a time span smaller than the current system delay and a minimum difference from the current system delay.
4. The method of claim 1, wherein, The plurality of different time spans are represented by frame numbers, and the method comprises: based on the frame interval, converting the current system delay into an equivalent frame number; from the plurality of future coefficient sets of different time spans, obtaining a target future coefficient set with a time span smaller than the equivalent frame number and a minimum difference from the equivalent frame number.
5. The method of claim 1, wherein, After obtaining the expression coefficient features, the method comprises: storing the expression coefficient features into a feature pool, the feature pool being used to store expression coefficient features of a plurality of recent face images.
6. The method of claim 1, wherein, The training method of the expression feature representation model comprises: obtaining a data set, the data set comprising a plurality of data samples, the data samples comprising face images and corresponding real expression coefficient; extracting spatial features of the face images through a convolutional neural network; annotating spatial features belonging to key regions of the face in the spatial features through a attention mechanism module; performing feature extraction on the annotated spatial features through a feature extraction layer to obtain expression features; mapping the expression features to expression coefficient space through a fully connected layer to obtain predicted expression coefficient features; adjusting parameters of the convolutional neural network based on the difference between the predicted expression coefficient features and the real expression coefficient.
7. The method of claim 1, wherein, The method comprises: receiving the feature sequence and performing encoding processing through an encoder in the time-series expression prediction model to obtain an encoded feature sequence; The coding feature sequence is respectively decoded by a plurality of parallel output heads in the time sequence expression prediction model to obtain a plurality of future coefficient sets of different time spans, wherein the plurality of parallel output heads correspond to a plurality of different time spans.
8. A digital human expression driving apparatus, characterized by, The method comprises the following steps: The feature acquisition module is configured to acquire a collected face image, extract a coefficient feature of the face image by an expression feature representation model, and obtain an expression coefficient feature. The coefficient prediction module is configured to obtain a feature sequence based on a plurality of expression coefficient features, perform coefficient prediction by a time sequence expression prediction model based on the feature sequence, and obtain a plurality of future coefficient sets of different time spans. The coefficient determination module is configured to obtain a current system delay based on time consumption data of a last digital human expression driving processing period, determine a target coefficient set from the plurality of future coefficient sets of different time spans based on the current system delay. The expression driving module is configured to drive a digital human model based on the target coefficient set.
9. An electronic device, comprising: The method comprises the following steps: One or more processors; The processor is configured to call instructions to enable the electronic device to perform the method of any one of claims 1 to 7.
10. A storage medium, the storage medium storing instructions, wherein, When the instructions run on the electronic device, the electronic device performs the method of any one of claims 1 to 7.
11. A program product comprising at least one of a program, instructions, characterized in that The program or instructions are executed by the electronic device to implement the steps of the method of any one of claims 1 to 7.