A virtual avatar cloning data processing method and system based on deep learning

By processing facial expression perception data using deep learning methods, a mapping relationship between image frames and key facial regions is established. Dynamic features are extracted and temporal modeling is performed to generate facial expression control parameters. This solves the problem of unstable micro-expression extraction of virtual avatars in low-light and blurry video frames, and achieves stable and natural expression generation of virtual avatars.

CN121326151BActive Publication Date: 2026-04-10CLOUD ATTACK NETWORK TECH HEBEI CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CLOUD ATTACK NETWORK TECH HEBEI CO LTD
Filing Date
2025-10-27
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies are unstable in extracting micro-expressions in low-light and blurry video frames, causing virtual avatars to fail to accurately reflect the user's emotional state.

Method used

By using deep learning methods, the raw data of facial expression perception is acquired and subjected to temporal calibration, anomaly cleaning and normalization mapping. The mapping relationship between image frames and key facial regions is established, dynamic features are extracted and temporal modeling is performed, and the expressive ability of images is evaluated by combining speech parameters. Facial expression control parameters are generated, temporal correction is performed and stability analysis is conducted, and a sequence of facial expression control instructions suitable for virtual avatar generation is output.

Benefits of technology

It achieves precise regional expression control under multimodal input fluctuation conditions, improves the naturalness and matching degree of the driving effect of virtual characters, ensures the coherence and stability of expression generation, and enhances the fault tolerance under low frame rate, partial occlusion or noisy input conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121326151B_ABST
    Figure CN121326151B_ABST
Patent Text Reader

Abstract

The application discloses a kind of virtual image cloning data processing method and system based on deep learning, it is related to cloning data processing technical field.The kind of virtual image cloning data processing method and system based on deep learning, comprising: S1, obtain expression perception original data, and the expression perception original data is preprocessed;S2, establish the mapping relationship of image frame and face key area, extract dynamic feature and carry out time series modeling, evaluate image expression ability, generate preliminary expression control parameter set;S3, evaluate the response state of each face key area in the current time segment, and execute time domain correction of expression control parameter;S4, according to frame-level expression control vector, execute stability analysis to face key area, and according to credibility screening and missing compensation strategy, reconstruct complete control structure.The problem that virtual image cannot truly reflect user emotional state caused by unstable micro-expression extraction in low light and fuzzy video frames is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of clone data processing, in particular to a virtual image clone data processing method and system based on deep learning. BACKGROUND

[0002] With the continuous development of digital people, virtual anchors and immersive interaction technology, the driving and generation method of virtual images has become one of the core key technologies in current multi-modal human-computer interaction systems. Especially under the conditions of multi-source input such as voice, image, action, etc., how to improve the expression credibility and expression consistency of virtual characters in complex environments has become an important direction of virtual image driving research.

[0003] For example, the invention with publication number CN115374298B provides a virtual image data processing method and device based on index, wherein a virtual image data processing method based on index includes: obtaining the index identifier of the virtual image data in the virtual world; determining the target index entry where the target index identifier matching the index identifier is located in the index list; reading the index value mapped by the target index identifier in the target index entry; the index value includes the category identifier of the virtual image data, the image feature and the image data of the image; based on the target index identifier and the image data, the image reconstruction processing of the virtual image is carried out, and the virtual image is obtained.

[0004] For example, the invention with publication number CN119781606A relates to a digital person real-time dialogue method, system, terminal and medium based on image cloning, the method includes: dividing the original user video into closed-mouth silent video and non-silent video; frame-by-frame intercepting the image of the face area in the non-silent video and synthesizing the training video; training the image cloning model based on the training video; converting the answer text generated by the large language model and the speech synthesis model into audio data, combining the trained image cloning model to generate a face image sequence matching the audio data, synthesizing the face image sequence to the non-silent video, and then pushing it out, if the non-silent video is not long enough, a forward-backward circular playing strategy is adopted, and in the non-dialogue state, the closed-mouth silent video is circularly pushed to the front end.

[0005] However, although the prior art has made significant progress in virtual image reconstruction and cloning driving, under non-ideal conditions such as micro-expression changes, voice cooperative control, image blur interference or insufficient illumination, there is still a lack of a complete technical path with real-time response evaluation, regional stability determination and driving control compensation ability.

[0006] Therefore, in view of the above problems, there is an urgent need for a virtual image clone data processing method and system based on deep learning. SUMMARY

[0007] Technical problems to be solved

[0008] In view of the deficiencies in the prior art, the present application provides a virtual image cloning data processing method and system based on deep learning, which solves the problem that micro-expression extraction is unstable in low-light and blurred video frames, resulting in the virtual image being unable to truly reflect the user's emotional state.

[0009] Technical scheme

[0010] To achieve the above object, the present application is implemented by the following technical scheme: a virtual image cloning data processing method and system based on deep learning, comprising: S1, obtaining expression perception raw data, and performing time sequence calibration, abnormal cleaning and normalized mapping processing on the expression perception raw data to obtain preprocessed expression perception raw data; S2, establishing a mapping relationship between image frames and facial key areas based on the preprocessed expression perception raw data, extracting dynamic features and performing time sequence modeling, combining voice parameter evaluation image expression ability, and generating a preliminary expression control parameter set; S3, based on the image expression ability evaluation result, the image frame index relationship and the voice short-time energy parameter, evaluating the response state of each facial key area in the current time segment, and performing time domain correction of the expression control parameter, and constructing a frame-level expression control vector sequence with consistent structure; S4, performing stability analysis on the frame-level expression control vector according to the facial key area, reconstructing the complete control structure according to the credibility screening and missing compensation strategy, and outputting the expression control instruction sequence suitable for virtual image generation.

[0011] Further, the specific steps of obtaining expression perception raw data and performing time sequence calibration, abnormal cleaning and normalized mapping processing on the expression perception raw data to obtain preprocessed expression perception raw data are as follows: obtaining expression perception raw data, which includes image frames, image resolution, ambient light intensity, voice frames and voice signals; performing time sequence correction on the expression perception raw data by a multi-modal timestamp alignment sliding window dynamic resampling algorithm to eliminate sampling delay; performing abnormal cleaning on the expression perception raw data by a Kalman filter algorithm to remove pseudo-features introduced by sudden jitter; testing the synchronization correlation between the expression signal and the voice signal by a mutual information analysis method to identify invalid micro-expression segments and interference records; performing uniform dimension conversion on the expression perception raw data by a min-max normalization algorithm to realize feature standardization and normalization processing.

[0012] Further, the specific steps of establishing a mapping relationship between the image frame and the facial key region based on the pre-processed facial expression perception raw data are as follows: based on the pre-processed facial expression perception raw data, the image frame sequence and the language frame sequence are extracted; the facial local displacement, the facial displacement speed and the facial displacement acceleration are obtained by performing time domain difference processing on the continuous image frames through a face key point tracking algorithm; the sound source fundamental frequency fluctuation amplitude is obtained by performing frequency domain envelope analysis on the language signal through a short-time Fourier transform algorithm; the speech short-time energy sequence is obtained by performing energy curve calculation on the speech frame sequence through a short-time energy analysis algorithm, and the speech short-time energy peak value is extracted; the background mean energy of the mute segment is obtained by performing energy statistics on the mute segment in the speech signal through a voice activity detection algorithm; the sound energy mutation frequency is obtained by extracting the jump features of the speech short-time energy sequence through an energy mutation point detection algorithm; each type of motion parameter is mapped to the mouth corner, the eye corner and the inter-brow region in the face image according to the time stamp, and the motion parameters include the facial local displacement, the facial displacement speed and the facial displacement acceleration; for each image frame, the facial displacement acceleration mean and the facial displacement speed mean of each region are calculated, and a facial local dynamic feature matrix is constructed.

[0013] Further, the specific steps of extracting dynamic features and performing time sequence modeling are as follows: the image frame sequence and the corresponding facial local dynamic feature matrix are taken as time sequence input, a time slice sequence is generated by adopting a sliding time window strategy, the continuous change relationship of the facial local dynamic feature matrix in each time slice is modeled by calling a structure with time sequence modeling capability, and a time feature sequence describing the expression state transition path is output.

[0014] Further, the specific steps of combining the speech parameter to evaluate the image expression expression ability and generating a preliminary expression control parameter set are as follows: for each image frame in the image frame sequence, the product of the facial displacement acceleration mean and the facial displacement speed mean is calculated to obtain a facial dynamic change intensity term; the image acquisition definition is obtained by multiplying the ambient light intensity and the image resolution; the image expression ability normalization index is obtained by dividing the facial dynamic change intensity term by the image acquisition definition; the physiological response enhancement term is constituted by calculating the speech short-time energy peak value divided by the background mean energy of the mute segment, and multiplying the speech correction coefficient; the expression intention response coefficient is obtained by adding one to the physiological response enhancement term; the virtual expression expression evaluation value is obtained by multiplying the image expression ability normalization index and the expression intention response coefficient; the time feature sequence and the virtual expression expression evaluation value corresponding to the image frame are corresponded to construct a frame-level credibility weight vector; the virtual expression expression evaluation value of the image frame is converted into a frame-level weight value in the fusion process through a normalization mapping function, and the time feature sequence is weighted processed accordingly; the weighted time feature sequence is used to extract the dominant motion direction, the average amplitude and the action duration of the facial key region in the time slice, and an expression control parameter set is generated.

[0015] Further, based on the image expression expression ability evaluation result, the image frame index relationship and the speech short-time energy parameter, the specific steps of evaluating the response state of each facial key region in the current time segment are as follows: based on the virtual expression expression evaluation value, the sound energy mutation frequency, the sound fundamental frequency fluctuation amplitude, the speech short-time energy peak value and the background average energy of the mute segment, the response state of each facial key region in the current time segment is evaluated: the weighted average value of the virtual expression expression evaluation value of the region in the current time segment is calculated to obtain the region average distinguishability score item; the sound energy mutation frequency is divided by the sum of the sound fundamental frequency fluctuation amplitude and the minimum item to obtain the physiological rhythm coordination degree; the mean value of the speech short-time energy peak value is calculated, and the speech short-time energy peak value mean value is divided by the background average energy of the mute segment to obtain the speech response intensity item; the region average distinguishability score item, the physiological rhythm coordination degree and the speech response intensity item are multiplied to obtain the region expression response evaluation value.

[0016] Further, the specific steps of performing time domain correction of the expression control parameter and constructing a frame-level expression control vector sequence with consistent structure are as follows: for each facial key region in the current time segment, the corresponding expression control parameter set is extracted, the expression response evaluation value is taken as a scaling factor, the average motion amplitude and the action duration parameter are proportionally adjusted to form the corrected control parameter set of the region; the corrected control parameter sets of the facial key regions are aggregated and arranged in the image frame time sequence to form the expression control vector sequence with uniform structure.

[0017] Further, the specific steps of performing stability analysis on the frame-level expression control vector according to the facial key region are as follows: the expression control vector sequence is called, the time index relationship of the image frame and the corner of the mouth, the corner of the eye and the glabella region is combined, the control parameter of each facial key region in each image frame is structure-unpacked, the corresponding average motion amplitude is extracted, and the average motion amplitude sequence in the time segment is constructed; the absolute value of the average motion amplitude difference between each image frame in the time segment and the previous frame is calculated, and divided by the sum of the average motion amplitude of the previous frame and the minimum item to obtain the inter-frame relative change rate; one is subtracted from the inter-frame relative change rate to obtain the region stability factor of the current frame; the region stability factor is multiplied by the virtual expression expression evaluation value corresponding to the frame to obtain the frame-level weighted stability score; the weighted stability scores of all image frames in the current time segment are arithmetically averaged to obtain the region expression stability evaluation value.

[0018] Further, the specific steps of reconstructing a complete control structure according to the credibility screening and missing compensation strategy and outputting an expression control instruction sequence suitable for virtual image generation are as follows: comparing the regional expression stability evaluation value with the expression stability threshold in real time, when the expression stability evaluation value of a region is lower than the expression stability threshold, marking the correction control parameter of the corresponding region in the current time segment as an untrusted state; when the regional expression stability evaluation value is greater than or equal to the expression stability threshold, marking the correction control parameter of the corresponding region in the current time segment as a trusted state; structuring and aggregating all the correction control parameters in the current time segment marked as trusted states according to the image frame timestamp order to generate a frame-level expression control vector; using an interpolation extrapolation mechanism to compensate for the missing of the region control parameters marked as untrusted states; and finally outputting the constructed expression control vector sequence as the standard input of the virtual image expression.

[0019] The second aspect of the present application provides a virtual image cloning data processing system based on deep learning, comprising: an expression perception raw data acquisition preprocessing module, a local dynamic modeling and distinguishability calculation module, a regional response evaluation and control parameter correction module, and a stability screening and driving instruction construction module, wherein: the expression perception raw data acquisition preprocessing module is used to acquire expression perception raw data, and perform time sequence calibration, abnormal cleaning and normalized mapping processing on the expression perception raw data to obtain preprocessed expression perception raw data; the local dynamic modeling and distinguishability calculation module is used to establish a mapping relationship between image frames and facial key regions based on the preprocessed expression perception raw data, extract dynamic features and perform time sequence modeling, evaluate image expression expression ability in combination with speech parameters, and generate a preliminary expression control parameter set; the regional response evaluation and control parameter correction module is used to evaluate the response state of each facial key region in the current time segment based on the image expression expression ability evaluation result, the image frame index relationship and the speech short-time energy parameter, and perform time domain correction of the expression control parameter to construct a frame-level expression control vector sequence with consistent structure; the stability screening and driving instruction construction module is used to perform stability analysis on the frame-level expression control vector according to the facial key region, reconstruct a complete control structure according to the credibility screening and missing compensation strategy, and output an expression control instruction sequence suitable for virtual image generation.

[0020] Advantageous effects

[0021] The present application has the following beneficial effects:

[0022] (1) The virtual image cloning data processing method and system based on deep learning, by establishing the mapping relationship between the image frame and the facial key area, combining the facial local displacement, velocity and acceleration and other time sequence parameters to construct the dynamic feature matrix, and using the sliding time window strategy for continuous modeling. The expression state transition path is output through the time sequence modeling structure, and the virtual expression expression evaluation value is used as the frame level weighting factor to realize the credibility difference expression in feature fusion, so that the final expression control parameter is more consistent with the actual intention change.

[0023] (2) The virtual image cloning data processing method and system based on deep learning, by proposing the regional expression response evaluation value index, the response ability of each facial key area to the expression intention in the current time segment is comprehensively described. Among them, the image definition and the facial dynamic intensity jointly reflect the quality of visual information acquisition, the short-time energy of speech and the background energy of silent segment reflect the physiological synchronicity, forming a triple criterion of facial response, which still maintains accurate regional expression control under the fluctuation condition of multi-modal input, and improves the naturalness and matching degree of driving effect.

[0024] (3) The virtual image cloning data processing method and system based on deep learning, by proposing the regional expression stability evaluation value, it can real-time evaluate whether each region action exists mutation or fluctuation, through comparison with expression stability threshold, automatically mark credible or non-credible area, and provide decision basis for driving control parameter rejection and correction, improve the coherence and stability of facial driving action, avoid abnormal situation of expression jitter, jump.

[0025] (4) The virtual image cloning data processing method and system based on deep learning, through the credible state screening mechanism, the expression control parameters of low stability area are filtered and processed, and the interpolation extrapolation algorithm is combined to reasonably compensate the missing control amount, so as to ensure that the generated expression control vector structure is complete and the rhythm is stable. This mechanism not only prevents local abnormal data from interfering with the driving instruction, but also enhances the fault tolerance capability under the condition of low frame rate, local shielding or noise input. The finally output control instruction can stably drive the virtual image to generate continuous and credible expression, which significantly improves the overall robustness and adaptability. BRIEF DESCRIPTION OF DRAWINGS

[0026] Figure 1 It is a virtual image cloning data processing method flow chart based on deep learning;

[0027] Figure 2 It is a virtual image cloning data processing system structure diagram based on deep learning;

[0028] Figure 3 It is a virtual expression expression evaluation value trend chart in image frame sequence;

[0029] Figure 4 a face key region expression stability evaluation map. DETAILED DESCRIPTION

[0030] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work are within the protection scope of the present application.

[0031] Please refer to Figures 1-4 The embodiment of the present application provides a technical solution: a virtual image cloning data processing method and system based on deep learning, comprising: S1, acquiring expression perception original data, and performing time sequence calibration, abnormal cleaning and normalized mapping processing on the expression perception original data to obtain preprocessed expression perception original data; S2, establishing a mapping relationship between image frames and face key regions based on the preprocessed expression perception original data, extracting dynamic features and performing time sequence modeling, combining speech parameter evaluation image expression ability, generating a preliminary expression control parameter set; S3, based on the image expression expression ability evaluation result, the image frame index relationship and the speech short-time energy parameter, evaluating the response state of each face key region in the current time segment, and performing time domain correction of the expression control parameter, constructing a frame-level expression control vector sequence with consistent structure; S4, performing stability analysis on the frame-level expression control vector according to the face key region, reconstructing the complete control structure according to the credibility screening and missing compensation strategy, and outputting the expression control instruction sequence suitable for virtual image generation.

[0032] Specifically, the specific steps of obtaining expression perception raw data and performing time sequence calibration, abnormality cleaning and normalization mapping on the expression perception raw data to obtain preprocessed expression perception raw data are as follows: obtaining expression perception raw data, the expression perception raw data including image frames, image resolution, ambient light intensity, speech frames and speech signals, wherein the image frames are collected in real time by a face image collection device and numbered by frame sequence, the image resolution is directly extracted by an imaging device parameter, the ambient light intensity is obtained by global brightness statistics of the image frames, and the speech frames and the speech signals are extracted by a synchronous speech channel collection device and bound with time stamps respectively; performing time axis alignment and synchronous sampling density reconstruction on the image frames, the image resolution, the ambient light intensity, the speech frames and the speech signals respectively by a multi-modal timestamp alignment sliding window dynamic resampling algorithm, so as to ensure the consistency of the multi-source features in the time dimension and eliminate sampling delay and sensing drift error; detecting abnormal fluctuations of the facial key point trajectory changes in the image frames by a Kalman filtering algorithm, identifying and removing the boundary of the sudden high-frequency interference segment in the speech signal, and completely removing the facial pseudo displacement data and audio pseudo energy features introduced by sudden jitter; testing the synchronization correlation between the expression signal and the speech signal by a mutual information analysis method, and screening out invalid micro-expression segments and high-frequency noise interference records without linkage consistency; performing unified dimension mapping and compressing to a standard scale range on the facial displacement velocity, the facial displacement acceleration and the image resolution in the image frames, the speech short-time energy peak value, the background mean energy of the mute segment and the fluctuation amplitude of the sound base frequency in the speech signal, and the ambient light intensity by a min-max normalization algorithm, so as to realize cross-modal feature standardization and normalization processing, and finally obtain the preprocessed expression perception raw data.

[0033] In the embodiment, by performing multi-modal time alignment, abnormal interference removal, effectiveness determination and cross-modal normalization processing on the expression perception raw data composed of image frames, image resolution, ambient light intensity, speech frames and speech signals, the overall processing quality of the multi-source data in terms of time synchronization, noise robustness, feature credibility and dimensionality consistency is effectively improved, an accurate, stable and standardized data basis is provided for subsequent image expression ability modeling, virtual expression expression evaluation value calculation and regional expression response evaluation, and thus the reliability and practicality of the whole virtual image expression generation are enhanced.

[0034] Specifically, the specific steps of establishing the mapping relationship between the image frame and the facial key region based on the preprocessed facial expression perception raw data are as follows: based on the preprocessed facial expression perception raw data, the image frame sequence and the language frame sequence are extracted, the image frame sequence is used to construct the dynamic change track of the facial key region, and the language frame sequence is used to assist the modeling of the speech expression feature; the continuous image frame is positioned and regionally tracked frame by frame through the face key point tracking algorithm, and time domain difference processing is performed to obtain the facial local displacement, facial local displacement speed and facial local displacement acceleration of the mouth corner, eye corner and inter-brow region, and construct a region dynamic feature sequence corresponding to the timestamp; the frequency domain envelope analysis of the language signal is performed through the short-time Fourier transform algorithm, the frequency energy distribution is calculated frame by frame, and the vocalization fundamental frequency fluctuation amplitude reflecting the speech fluctuation is obtained; the speech frame sequence is calculated frame by frame through the short-time energy analysis algorithm, the speech short-time energy sequence is constructed, and the speech short-time energy peak value is extracted based on the peak value recognition algorithm, which is used to measure the instantaneous change of the speech intensity; the speech signal is analyzed and the energy is counted segment by segment through the speech activity detection algorithm, the background mean energy of the silent segment is calculated, and the speech environment noise level in the silent state is reflected; the speech short-time energy sequence is extracted through the energy mutation point detection algorithm, the energy mutation event in the vocalization behavior is identified, the vocalization energy mutation frequency is obtained, and the time index reflecting the speech rhythm characteristic is constructed; the facial local displacement, facial local displacement speed and facial local displacement acceleration are mapped to the corresponding mouth corner, eye corner and inter-brow region in the image frame according to the acquisition timestamp, so as to ensure that the region dynamic parameters and the image frame are accurately corresponding; for each image frame, the facial displacement acceleration mean and the facial displacement speed mean are calculated in the mapping region respectively, and the multi-region time sequence aligned facial local dynamic feature matrix is constructed, which provides accurate input for subsequent dynamic modeling.

[0035] In the embodiment, by mapping modeling of the image frame and the facial key region on the preprocessed facial expression perception raw data, the face key point tracking algorithm, the short-time Fourier transform algorithm, the short-time energy analysis algorithm, the speech activity detection algorithm and the energy mutation point detection algorithm are combined, the facial local displacement, the facial local displacement speed, the facial local displacement acceleration, the vocalization fundamental frequency fluctuation amplitude, the speech short-time energy peak value, the silent segment background mean energy and the vocalization energy mutation frequency are extracted and accurately synchronized item by item; by constructing the dynamic feature matrix of the mouth corner region, the eye corner region and the inter-brow region, it is ensured that the motion parameters and the image frame are strictly aligned in the time dimension, which provides high-quality feature input for facial local dynamic modeling and virtual expression expression ability evaluation, and enhances the response accuracy and time sequence stability of the expression driving process.

[0036] Specifically, the specific steps of extracting dynamic features and performing time series modeling are as follows: taking the image frame sequence and the corresponding face local dynamic feature matrix as time series input; adopting a sliding time window strategy to generate a time segment sequence, ensuring that each time segment has a stable local dynamic change trend at the frame level; the sliding time window strategy is implemented by setting a fixed length time window and a fixed step size, and performing overlapping sliding segmentation on the image frame time axis; calling a structure with time series modeling capability to model the continuous change relationship of the face local dynamic feature matrix in each time segment, wherein the time series modeling structure adopts a bidirectional gated recurrent unit network, effectively captures the dynamic evolution characteristics of micro-expression on the time axis through the forward and backward information propagation mechanism, constructs the bidirectional dependence relationship of local motion trend, and improves the accuracy and expression ability of time series modeling; finally outputting a time feature sequence describing the expression state transition path, providing time series continuity support for subsequent expression control parameter generation and response evaluation index construction.

[0037] In the present embodiment, by introducing the construction of the face local dynamic feature matrix and the time series mapping of the image frame sequence, combining the sliding time window strategy and the structure with time series modeling capability, the stable capture and state evolution modeling of the face local dynamic feature in the time segment are realized, which effectively enhances the continuity and expression consistency of the expression state transition path, and provides high time series resolution and dynamic precision guarantee for subsequent expression control parameter generation and expression response evaluation value extraction.

[0038] Specifically, the specific steps of generating the preliminary expression control parameter set in combination with the voice parameter evaluation of the image expression are as follows: for each image frame in the image frame sequence, the product of the face displacement acceleration mean value and the face displacement speed mean value is calculated respectively to obtain a face dynamic change intensity term, which is used to depict the dynamic activity degree of the local face region corresponding to the current image frame; the ambient light intensity and the image resolution extracted from the preprocessed expression perception original data are multiplied item by item to obtain an image acquisition definition term, which is used to reflect the recognizable level of the image frame under the current acquisition condition; the face dynamic change intensity term is divided by the image acquisition definition term to obtain an image expression capacity normalization index, which is used as a quantitative reference for measuring the visual expression intensity of the current image frame; the peak value of the short-time energy of the voice is divided by the background mean energy of the mute segment, and combined with the voice correction coefficient obtained by modeling the voice short-time energy change range in the historical multi-round pronunciation process of the same speaker to form a physiological response enhancement term, which enhances the contribution of the voice behavior to the expression expression driving process; the physiological response enhancement term is added by one to obtain an expression intention response coefficient, which further amplifies the adjustment effect of the physiological response on the image expression capacity; the image expression capacity normalization index is multiplied by the expression intention response coefficient to obtain a virtual expression expression evaluation value corresponding to each image frame, which is used as an important criterion for the credibility of the current image frame; the time feature sequence and the virtual expression expression evaluation value of the corresponding image frame are matched frame by frame to construct a frame-level credibility weight vector, ensuring that the information reliability difference in the feature modeling process is effectively included in the evaluation system; the virtual expression expression evaluation value of each frame is converted into a frame-level weight value in the fusion process through a normalization mapping function, realizing dynamic weighted processing of the time feature sequence; finally, the weighted time feature sequence is used to extract the dominant motion direction, average amplitude and action duration of the key face region in the time segment to generate the expression control parameter set for driving the virtual image.

[0039] wherein the specific calculation formula of the virtual expression expression evaluation value is:

[0040] ;

[0041] In the formula, S represents the virtual expression expression evaluation value, represents the face displacement acceleration mean value, represents the face displacement speed mean value, L represents the ambient light intensity, and R represents the image resolution, represents the voice correction coefficient, E represents the peak value of the short-time energy of the voice, represents the background mean energy of the mute segment.

[0042] In the embodiment, Table 1 is a virtual expression evaluation value data table, which lists in detail the key indicators of 5 image frames in the process of evaluating the virtual image expression expression, including the average face displacement acceleration, the average face displacement speed, the ambient light intensity, the image resolution, the short-time energy peak value of the voice, the background average energy of the silent segment, and the finally calculated virtual expression evaluation value. Among them: the average face displacement acceleration of frame 1 is 0.032, the average face displacement speed is 0.045, the ambient light intensity is 0.6, the image resolution is 0.70, the short-time energy peak value of the voice is 0.030, the background average energy of the silent segment is 0.616, and the finally calculated virtual expression evaluation value is 0.636; the average face displacement acceleration of frame 2 is 0.028, the average face displacement speed is 0.042, the ambient light intensity is 0.5, the image resolution is 0.65, the short-time energy peak value of the voice is 0.028, the background average energy of the silent segment is 0.616, and the virtual expression evaluation value is 0.601; the average face displacement acceleration of frame 3 is 0.034, the average face displacement speed is 0.047, the ambient light intensity is 0.6, the image resolution is 0.70, the short-time energy peak value of the voice is 0.032, the background average energy of the silent segment is 0.616, and the virtual expression evaluation value is 0.765; the average face displacement acceleration of frame 4 is 0.038, the average face displacement speed is 0.049, the ambient light intensity is 0.7, the image resolution is 0.75, the short-time energy peak value of the voice is 0.035, the background average energy of the silent segment is 0.616, and the virtual expression evaluation value is 1.002; the average face displacement acceleration of frame 5 is 0.030, the average face displacement speed is 0.044, the ambient light intensity is 0.55, the image resolution is 0.68, the short-time energy peak value of the voice is 0.031, the background average energy of the silent segment is 0.616, and the virtual expression evaluation value is 0.827.

[0043] Table Virtual expression evaluation value data table

[0044]

[0045] As Figure 3As shown, it is a virtual expression evaluation value trend chart in the image frame sequence, which shows the change trend of virtual expression evaluation value in five consecutive image frames. The virtual expression evaluation value is calculated based on the face local motion state, image clarity factor and speech expression dynamic characteristics, which comprehensively reflects the expression ability of each frame at the current time. The virtual expression evaluation value of frame 4 is the highest, which is 1.598, indicating that the expression at this moment is the most clear and reliable; The evaluation value of frame 1 is the lowest, which is 0.735, which may be affected by image blur and speech interference, and the expression reliability is relatively low. The virtual expression evaluation value sequence provides a quantitative basis for subsequent expression control parameter weighted processing and region stability judgment.

[0046] In the embodiment, by combining the image expression ability normalization index with the expression intention response coefficient, the virtual expression evaluation value is calculated, and the frame-level reliability weight vector is constructed accordingly, which significantly improves the fusion precision between the dynamic characteristics of the face key region and the time sequence. In the expression control parameter generation process, the ambient light intensity, image resolution, speech short-time energy peak value, silent segment background mean energy and speech correction coefficient and other key data are introduced, the joint modeling of image frame expression ability and speech driving intensity is realized, the comprehensive utilization ability of facial expression recognition to multi-source heterogeneous information is effectively enhanced, the time sequence consistency and reliable weight distribution rationality of expression control parameter set are guaranteed, and the expression accuracy and stability of subsequent virtual image driving process are ensured.

[0047] Specifically, based on the image expression expression ability evaluation result, the image frame index relationship and the speech short-time energy parameter, the specific steps of evaluating the response state of each facial key region in the current time segment are as follows: first, the virtual expression evaluation value corresponding to each image frame in the current time segment, the sound energy mutation frequency, the sound fundamental frequency fluctuation amplitude, the speech short-time energy peak value and the background average energy of the silent segment are obtained, and according to the time index relationship between the image frame and the facial key region, a multi-frame data set corresponding to each facial key region is established; secondly, the weighted average value of the virtual expression evaluation value in the multi-frame image associated with each facial key region is calculated to obtain the regional average distinguishability score item, which is used to measure the expressible degree of the region in the image sequence; thirdly, the sound energy mutation frequency in the current time segment is divided by the sum of the sound fundamental frequency fluctuation amplitude and the minimum item to construct the physiological rhythm coordination degree reflecting the coordination between the sound stability and the frequency fluctuation; further, the speech short-time energy peak value of each facial key region associated image frame in the current time segment is extracted, the arithmetic average value is calculated and divided by the background average energy of the silent segment to form the speech response intensity item used to measure the speech-driven response ability of the region; finally, the regional average distinguishability score item, the physiological rhythm coordination degree and the speech response intensity item are multiplied item by item to output the regional expression response evaluation value, which is used to represent the comprehensive response level of the region to the image expression ability and the speech-driven signal in the current time segment.

[0048] wherein the specific calculation formula of the regional expression response evaluation value is:

[0049] ;

[0050] In the formula, Q represents the regional expression response evaluation value, S represents the virtual expression evaluation value, n represents the number of image frames in the time segment, represents the sound energy mutation frequency, represents the sound fundamental frequency fluctuation amplitude, represents the minimum item, represents the speech short-time energy peak value average, represents the background average energy of the silent segment.

[0051] In the embodiment, by introducing the virtual expression evaluation value, the sound energy mutation frequency, the sound fundamental frequency fluctuation amplitude, the speech short-time energy peak value and the background average energy of the silent segment, the regional average distinguishability score item, the physiological rhythm coordination degree and the speech response intensity item are constructed, and the regional expression response evaluation value is further calculated, which realizes the fine modeling and accurate evaluation of the response state of each facial key region in the current time segment, improves the perception sensitivity of the weak emotional driving signal and the credibility of the response determination in the expression control parameter generation process, and enhances the driving stability and expression matching effect of the virtual image in the complex expression environment.

[0052] Specifically, the specific steps of performing time domain correction of expression control parameters and constructing a structure-consistent frame-level expression control vector sequence are as follows: for each facial key region in the current time segment, including the corner of the mouth region, the corner of the eye region and the glabella region, the corresponding expression control parameter set is extracted, including the dominant motion direction, the average motion amplitude and the action duration; the region expression response evaluation value is taken as a scaling factor, and the average motion amplitude and the action duration parameters in the expression control parameter set are respectively subjected to linear proportional adjustment to obtain the corrected average motion amplitude and the corrected action duration, which constitute the corrected control parameter set of the region; under the guidance of the timestamp index of all image frames, the corrected control parameter sets of each facial key region are sequentially aggregated in the frame-level order, and are arranged and combined based on the region structure consistency requirement to constitute the expression control vector sequence with unified data format and time sequence structure.

[0053] In the embodiment, by introducing the region expression response evaluation value as a scaling factor, the average motion amplitude and the action duration in the expression control parameter set are dynamically adjusted, which effectively enhances the response accuracy of the facial key region, and further improves the individualized adaptation ability of the frame-level expression control parameter; at the same time, by the construction method of the expression control vector sequence with unified structure, the synchronous coordination of the control parameters of different facial key regions in the time dimension is ensured, which provides a high-consistency and high-robustness control input basis for the virtual image expression, and significantly improves the dynamic stability and expression authenticity in the virtual expression generation process.

[0054] Specifically, the specific steps of performing stability analysis on the frame-level expression control vector according to the facial key region are as follows: calling an expression control vector sequence, combining the image frame with the time index relationship of the mouth corner, eye corner and inter-brow region, and constructing a one-to-one mapping index between the image frame and the facial key region based on the preprocessed expression perception raw data; for each image frame, performing a structure unpacking operation to separate the control parameter set corresponding to each facial key region in the frame, and extracting the average motion amplitude of each region to construct an average motion amplitude sequence within the current time segment in time order; for each image frame, calculating the absolute value of the average motion amplitude difference between the frame and the previous frame, and dividing the sum of the average motion amplitude of the previous frame and a minimum term to generate the inter-frame relative change rate sequence of the frame; wherein the minimum term is a positive real number less than 1, which is used to prevent numerical instability caused by a zero denominator; by taking the difference of one for the inter-frame relative change rate of each frame, the region stability factor of the current frame is obtained, and the frame-level stability sequence of the facial key region is constructed; the region stability factor and the virtual expression evaluation value of the corresponding image frame are multiplied frame by frame to form the weighted stability score value after weighted processing, and a complete frame-level weighted stability score set within the region is constructed; the weighted stability scores of all image frames within the current time segment are subjected to an arithmetic average operation to generate the regional expression stability evaluation value of the facial key region, which is used to measure the continuity and stability of the expression control instruction of the region within the time segment.

[0055] wherein the specific calculation formula of the regional expression stability evaluation value is:

[0056] ;

[0057] In the formula, W represents the regional expression stability evaluation value, S represents the virtual expression evaluation value, n represents the number of image frames within the time segment, represents the minimum term, represents the average motion amplitude in the i-th frame, represents the average motion amplitude in the i-1-th frame.

[0058] In the embodiment, Table 2 is a region expression stability evaluation value data table, which details the key parameters in the process of evaluating the expression stability of the five facial key regions, including the number of image frames in the time segment, the frame-level weighted stability score set, and the finally calculated region expression stability evaluation value. Among them: the number of image frames in the time segment of region 1 is 5, the frame-level weighted stability score set is [0.426, 0.543, 0.537, 0.881], and the finally calculated region expression stability evaluation value is 0.7703; the number of image frames in the time segment of region 2 is 5, the frame-level weighted stability score set is [0.787, 0.556, 0.658, 0.942], and the region expression stability evaluation value is 0.7134; the number of image frames in the time segment of region 3 is 5, the frame-level weighted stability score set is [0.470, 0.687, 0.465, 0.518], and the region expression stability evaluation value is 0.7586; the number of image frames in the time segment of region 4 is 5, the frame-level weighted stability score set is [0.827, 0.938, 0.247, 0.532], and the region expression stability evaluation value is 0.7983; the number of image frames in the time segment of region 5 is 5, the frame-level weighted stability score set is [0.606, 0.502, 0.546, 0.450], and the region expression stability evaluation value is 0.8182.

[0059] Table Region expression stability evaluation value data table

[0060] Region number n Set of frame-level weighting stability scores W Region 1 5 [0.426,0.543,0.537,0.881] 0.7703 Region 2 5 [0.787,0.556,0.658,0.942] 0.7134 Region 3 5 [0.470,0.687,0.465,0.518] 0.7586 Region 4 5 [0.827,0.938,0.247,0.532] 0.7983 Region 5 5 [0.606,0.502,0.546,0.450] 0.8182

[0061] As Figure 4 shown, it is a facial key region expression stability evaluation diagram, which shows the region expression stability evaluation values of the five facial key regions, reflecting the average motion amplitude fluctuation of each region in the current time segment and its influence on the virtual expression evaluation value. The blue dotted line indicates the expression stability threshold, the green color indicates the trusted state, and the red color indicates the untrusted state. As can be seen from the figure, the region expression stability evaluation value of region 2 is lower than the expression stability threshold and is marked as an untrusted state; the region expression stability evaluation values of the remaining regions are all higher than the expression stability threshold and are determined as trusted states, which can be used for subsequent construction and execution of expression control instructions. The diagram clearly expresses the distribution characteristics and screening logic of the region response stability.

[0062] In the embodiment, by constructing the time index relationship between the image frame and the facial key region, the expression control vector sequence is structurally unpacked according to the image frame, the average motion amplitude of the mouth corner region, the eye corner region and the glabella region is extracted, the average motion amplitude sequence of the facial key region is constructed, then the inter-frame relative change rate is calculated and the region stability factor is generated, combined with the corresponding virtual expression evaluation value of each frame, a frame-level weighted stability score sequence is formed, and finally the region expression stability evaluation value is output, which effectively improves the continuity evaluation ability of the frame-level expression control parameter and enhances the stability recognition precision of the facial key region control structure.

[0063] Specifically, according to the credibility screening and missing compensation strategy, the complete control structure is reconstructed, and the specific steps of outputting the expression control instruction sequence suitable for virtual image generation are as follows: comparing the region expression stability evaluation value with the preset expression stability threshold value frame by frame in real time, and recording the stability state of the mouth corner region, the eye corner region and the glabella region in each image frame; when the region expression stability evaluation value of a region is lower than the corresponding expression stability threshold value, the corresponding modified control parameter of the region in the current time segment is marked as untrusted state, and the untrusted region index table is updated synchronously; when the region expression stability evaluation value is greater than or equal to the expression stability threshold value, the corresponding modified control parameter of the region in the current time segment is marked as trusted state, and the time stamp position of the trusted control parameter is recorded; for the region modified control parameters marked as trusted state in the current time segment, the structured aggregation processing is performed according to the image frame time stamp order, the frame-level expression control vector set is constructed, and the time consistency and parameter integrity are verified; for the region modified control parameters marked as untrusted state, a linear interpolation extrapolation mechanism based on the trusted samples of the previous and next frames is called to compensate the control parameters, and the dynamic repair of the missing data segment is completed; the interpolation extrapolation mechanism is a linear interpolation model constructed based on the trusted control parameters in the adjacent frames, the trusted samples before and after the current missing frame are used as the interpolation basis, and when necessary, trend extrapolation is performed to the time segment boundary to complete the dynamic fitting and missing segment repair of the control parameters; finally, the expression control vector sequence with complete structure, time synchronization and data credibility is output as the standard input of virtual image expression motion generation, which drives the virtual image expression to be synthesized in real time.

[0064] In the embodiment, by comparing the region expression stability evaluation value with the expression stability threshold value frame by frame in real time, the credibility state of each facial key region is accurately marked, and the structured aggregation and interpolation extrapolation compensation of the modified control parameter is performed based on the credibility marking result, which can effectively improve the integrity, stability and driving reliability of the frame-level expression control vector, ensure that the finally output expression control vector sequence has time continuity, region consistency and data credibility, and significantly enhance the real restoration ability and expression consistency of micro-expression dynamic changes.

[0065] As Figure 2 shown, the second aspect of the present application provides a deep learning-based virtual image cloning data processing system, comprising: an expression-aware raw data acquisition and preprocessing module, a local dynamic modeling and distinguishability calculation module, a regional response evaluation and control parameter correction module, and a stability screening and driving instruction construction module, wherein: the expression-aware raw data acquisition and preprocessing module is used to acquire expression-aware raw data, and perform time series calibration, abnormal cleaning and normalized mapping processing on the expression-aware raw data to obtain preprocessed expression-aware raw data; the local dynamic modeling and distinguishability calculation module is used to establish a mapping relationship between image frames and facial key regions based on the preprocessed expression-aware raw data, extract dynamic features and perform time series modeling, evaluate image expression expression ability in combination with speech parameters, and generate a preliminary expression control parameter set; the regional response evaluation and control parameter correction module is used to evaluate the response state of each facial key region in the current time segment based on the image expression expression ability evaluation result, the image frame index relationship and the speech short-time energy parameter, and perform time-domain correction of the expression control parameter to construct a frame-level expression control vector sequence with consistent structure; the stability screening and driving instruction construction module is used to perform stability analysis on the frame-level expression control vector according to the facial key region, reconstruct the complete control structure according to the credibility screening and missing compensation strategy, and output an expression control instruction sequence suitable for virtual image generation.

[0066] In this embodiment, by constructing a virtual image cloning data processing system composed of an expression-aware raw data acquisition and preprocessing module, a local dynamic modeling and distinguishability calculation module, a regional response evaluation and control parameter correction module, and a stability screening and driving instruction construction module, the full-process processing of expression-aware raw data can be realized, covering the acquisition and normalized preprocessing of image frames, image resolution, environmental light intensity, speech frames and speech signals, to the dynamic feature extraction and time series modeling of facial key regions, to the expression response state evaluation and control parameter correction based on virtual expression expression evaluation value and speech short-time energy parameter, and finally forming an expression control instruction sequence with consistent structure, stable and reliable, effectively improving the performance of virtual image in driving accuracy, response consistency and emotional expression authenticity.

[0067] It is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting; it is not intended to exclude myriad other embodiments of the present application that other presenters can develop. It is also possible, however, that only a single element can be present. It is further noted that such a term as "comprising" is intended to mean that the embodiments include the recited elements, but not excluding other elements. "Consisting essentially of when used herein in relation to a composition, means that the composition includes the recited elements, and can include additional elements, so long as the additional elements do not materially alter the basic and novel characteristics of the claimed composition. "Consisting of" when used herein in relation to a composition, means that the composition includes the recited elements, and no additional elements.

[0068] The preferred embodiments of the application disclosed above are only to help explain the principles of the present application. The preferred embodiments do not describe all the details of the present application, nor limit the present application to only the specific embodiments described. It is apparent that many modifications and variations can be made to the present application based on the content of the present disclosure. The present disclosure selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present application, so that those skilled in the art can well understand and utilize the present application. The present application is limited only by the claims and their full scope and equivalents.

Claims

1. A deep learning-based virtual avatar cloning data processing method, characterized in that, The method comprises the following steps: S1, obtaining expression perception raw data, and performing time sequence calibration, abnormality cleaning and normalization mapping processing on the expression perception raw data to obtain preprocessed expression perception raw data; S2, establishing a mapping relationship between image frames and facial key regions based on the preprocessed expression perception raw data, extracting dynamic features and performing time sequence modeling, combining speech parameters to evaluate image expression ability, and generating a preliminary expression control parameter set; the speech parameters include sound energy mutation frequency, sound fundamental frequency fluctuation amplitude, speech short-time energy peak value and background mean energy of the mute segment, and the facial key regions include the corners of the mouth, the corners of the eyes and the glabella region; The specific steps of combining speech parameters to evaluate image expression ability and generating a preliminary expression control parameter set are as follows: For each image frame in the image frame sequence, the product of the average facial displacement acceleration and the average facial displacement speed of each region is calculated to obtain a facial dynamic change intensity term; the environmental light intensity is multiplied by the image resolution to obtain an image acquisition definition term; The facial dynamic change intensity term is divided by the image acquisition definition term to obtain an image expression ability normalization index; the speech short-time energy peak value is divided by the background mean energy of the mute segment, and multiplied by a speech correction coefficient to form a physiological response enhancement term; The physiological response enhancement term is added by one to obtain an expression intention response coefficient; The image expression ability normalization index is multiplied by the expression intention response coefficient to obtain an image expression evaluation value of each region; The time feature sequence describing the expression state transition path is corresponded to the image expression evaluation value of the corresponding image frame to construct a frame-level credibility weight vector; through a normalization mapping function, the image expression evaluation value of the image frame is converted into a frame-level weight value in the fusion process, and the time feature sequence is weighted processed accordingly; the weighted time feature sequence is used to extract the dominant motion direction, average motion amplitude and action duration of the facial key region in the time segment to generate a preliminary expression control parameter set; S3, based on the image expression ability evaluation result, the image frame index relationship and the speech parameter, the response state of each facial key region in the current time segment is evaluated, and the time domain correction of the expression control parameter is performed to construct a frame-level expression control vector sequence with consistent structure; S4, performing stability analysis on the frame-level expression control vector sequence according to the facial key region, and reconstructing the complete control structure according to the credibility screening and missing compensation strategy to output the expression control instruction sequence suitable for virtual image generation.

2. The deep learning-based virtual clone data processing method according to claim 1, characterized in that: The specific steps of obtaining expression perception raw data and performing time sequence calibration, abnormality cleaning and normalization mapping processing on the expression perception raw data to obtain preprocessed expression perception raw data are as follows: Obtain expression perception raw data, which includes image frames, image resolution, environmental light intensity, speech frames and speech signals; The multi-modal timestamp alignment sliding window dynamic resampling algorithm is used to correct the time sequence of the expression perception raw data to eliminate sampling delay; the Kalman filtering algorithm is used to clean the expression perception raw data to remove the pseudo features introduced by sudden jitter; The synchronization correlation between the expression signal and the speech signal is tested by a mutual information analysis method, and invalid expression segments and interference records are identified; the original expression perception data is uniformly converted by a min-max normalization algorithm, so as to realize feature standardization and normalization processing.

3. The virtual clone data processing method based on deep learning according to claim 2, characterized in that: The specific steps of establishing the mapping relationship between the image frame and the facial key region based on the preprocessed expression perception original data are as follows: Based on the preprocessed expression perception original data, the image frame sequence and the speech frame sequence are extracted; the facial local displacement, the facial displacement velocity and the facial displacement acceleration are obtained by performing time domain difference processing on the continuous image frames through a face key point tracking algorithm; the vocal fundamental frequency fluctuation amplitude is obtained by performing frequency domain envelope analysis on the speech signal through a short-time Fourier transform algorithm; the speech short-time energy sequence is obtained by performing energy curve calculation on the speech frame sequence through a short-time energy analysis algorithm, and the speech short-time energy peak value is extracted; the background mean energy of the mute segment is obtained by performing energy statistics on the mute segment in the speech signal through a voice activity detection algorithm; the vocal energy mutation frequency is obtained by extracting the jump feature of the speech short-time energy sequence through an energy mutation point detection algorithm; The various motion parameters are mapped to the mouth corner, eye corner and eyebrow region in the image according to the time stamp, and the motion parameters include the facial local displacement, the facial displacement velocity and the facial displacement acceleration; For each image frame, the facial displacement acceleration mean and the facial displacement velocity mean of each region are calculated to construct a facial local dynamic feature matrix.

4. The virtual clone data processing method based on deep learning according to claim 3, characterized in that: The specific steps of extracting the dynamic feature and performing time sequence modeling are as follows: The image frame sequence and the corresponding facial local dynamic feature matrix are taken as time sequence input, a time slice sequence is generated by adopting a sliding time window strategy, the continuous change relationship of the facial local dynamic feature matrix in each time slice is modeled by calling a structure with time sequence modeling capability, and a time feature sequence describing the expression state transition path is output.

5. The virtual clone data processing method based on deep learning according to claim 4, characterized in that: Based on the image expression expression ability evaluation result, the image frame index relationship and the speech short-time energy parameter, the response state of each facial key region in the current time slice is evaluated, including: Based on the image expression expression evaluation value, the vocal energy mutation frequency, the vocal fundamental frequency fluctuation amplitude, the speech short-time energy peak value and the background mean energy of the mute segment, the response state of each facial key region in the current time slice is evaluated: the weighted average value of the image expression expression evaluation value of the region in the current time slice is calculated to obtain a region average distinguishability score item; the vocal energy mutation frequency is divided by the sum of the vocal fundamental frequency fluctuation amplitude and the minimum item to obtain a physiological rhythm coordination degree; the mean value of the speech short-time energy peak value is calculated, and the speech short-time energy peak value mean is divided by the background mean energy of the mute segment to obtain a speech response intensity item; the region average distinguishability score item, the physiological rhythm coordination degree and the speech response intensity item are multiplied to obtain the expression response evaluation value of the region.

6. The virtual clone data processing method based on deep learning according to claim 5, characterized in that: The specific steps of performing time domain correction on the expression control parameter to construct a frame-level expression control vector sequence with consistent structure are as follows: For each facial key region in the current time segment, a corresponding preliminary expression control parameter set is extracted, a corresponding regional expression response evaluation value is taken as a scaling factor, a proportional adjustment is performed on the average motion amplitude and the motion duration parameter, and a modified control parameter set of the region is constructed; The modified control parameter sets of each facial key region are aggregated and arranged in image frame time order to form a consistent expression control vector sequence.

7. The virtual clone data processing method based on deep learning according to claim 6, characterized in that: The specific steps of performing stability analysis on the frame-level expression control vector sequence according to the facial key region are as follows: The expression control vector sequence is called, the control parameters of each facial key region in each image frame are structurally unpacked in combination with the time index relationship of the image frame and the corner of the mouth, the corner of the eye and the glabella region, the corresponding average motion amplitude is extracted, and the average motion amplitude sequence in the time segment is constructed; The absolute value of the average motion amplitude difference between each image frame in the time segment and the previous frame is calculated, and then divided by the sum of the average motion amplitude of the previous frame and the minimum term to obtain the inter-frame relative change rate; The inter-frame relative change rate is subtracted by one to obtain the regional stability factor of the current frame; the regional stability factor is multiplied by the image expression evaluation value corresponding to the frame to obtain the frame-level weighted stability score; the weighted stability scores of all image frames in the current time segment are arithmetically averaged to obtain the regional expression stability evaluation value.

8. The virtual clone data processing method based on deep learning according to claim 7, characterized in that: The specific steps of reconstructing the complete control structure according to the credibility screening and missing compensation strategy and outputting the expression control instruction sequence suitable for virtual image generation are as follows: The regional expression stability evaluation value is compared with the expression stability threshold in real time, when the expression stability evaluation value of a region is lower than the expression stability threshold, the modified control parameters of the corresponding region in the current time segment are marked as untrusted state; when the regional expression stability evaluation value is greater than or equal to the expression stability threshold, the modified control parameters of the corresponding region in the current time segment are marked as trusted state; The modified control parameters in the current time segment which are marked as trusted state are structurally aggregated according to the image frame timestamp order to generate frame-level expression control vector; the region control parameters marked as untrusted state are compensated by interpolation extrapolation mechanism; finally, the constructed expression control instruction sequence is output as the standard input of virtual image expression.

9. A deep learning-based virtual clone data processing system, applying the deep learning-based virtual clone data processing method of any one of claims 1-8. It comprises: an expression perception raw data acquisition and preprocessing module, a local dynamic modeling and distinguishability calculation module, a regional response evaluation and control parameter modification module, and a stability screening and driving instruction construction module, wherein: The expression perception raw data acquisition and preprocessing module is used to acquire expression perception raw data, and to perform time series calibration, abnormal cleaning and normalized mapping processing on the expression perception raw data to obtain preprocessed expression perception raw data; The local dynamic modeling and distinguishability calculation module is used to establish a mapping relationship between the image frame and the facial key region based on the preprocessed expression perception raw data, extract dynamic features and perform time series modeling, evaluate the image expression expression ability in combination with the speech parameter, and generate a preliminary expression control parameter set; The region response evaluation and control parameter modification module is configured to evaluate the response state of each facial key region in the current time segment based on the image expression expression ability evaluation result, the image frame index relationship and the voice parameter, and perform time domain modification on the expression control parameter to construct a frame-level expression control vector sequence with consistent structure. The stability screening and driving instruction construction module is configured to perform stability analysis on the frame-level expression control vector sequence according to facial key regions, reconstruct a complete control structure according to a credibility screening and missing compensation strategy, and output an expression control instruction sequence suitable for virtual image generation.

Citation Information

Patent Citations

  • Index-based virtual image data processing method and device

    CN115374298B

  • Digital human real-time dialogue method and system based on image cloning, terminal and medium

    CN119781606A

  • Real-time interaction 3D digital holographic cabin method based on deep learning and sound cloning

    CN120318437A

  • Lip sync animation creation device, computer program, and face model creation system

    JP2008140364A