Voice-based image driving method and image-driven data processing method

By encoding and transforming reference speech and facial images, high-definition and high-fidelity virtual object images are generated, solving the problem of insufficient image clarity and fidelity in existing technologies and improving the user experience.

CN116363269BActive Publication Date: 2025-12-09ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310252857.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-08
Publication Date
2025-12-09
Estimated Expiration
2043-03-08

AI Technical Summary

Technical Problem

Existing key-point-based virtual object face reconstruction methods generate images with insufficient clarity and fidelity, resulting in a poor user experience.

Method used

By acquiring reference speech and facial images, speech coding and image coding are performed. The image features of the first region are transformed using facial prior features and target speech features to generate a high-definition and high-fidelity target image containing texture features.

Benefits of technology

It improves the clarity and fidelity of generated images, thus enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116363269B_ABST
    Figure CN116363269B_ABST
Patent Text Reader

Abstract

The embodiment of the specification provides a voice-based image driving method and an image-driven data processing method, wherein the voice-based image driving method comprises the following steps: obtaining a reference voice and a reference facial image of a virtual object, performing voice coding on the reference voice to obtain target voice features, performing image coding on the reference facial image to obtain first image features of a first region and second image features of a second region, performing feature transformation on the first image features based on facial prior features and the target voice features to determine first target image features, wherein the facial prior features comprise facial texture features, and generating a target image after driving according to the first target image features and the second image features. The first image features are transformed based on the target voice features and the facial prior features, and a target image with high fidelity and high definition is obtained through decoding, thereby improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present specification relate to the technical field of image processing, in particular to an image driving method based on voice. BACKGROUND

[0002] With the development of computer technology, according to the reference voice input by the user, a target image of a virtual object corresponding to the face change is generated, that is, a virtual object reconstruction technology, which has wide application in virtual live broadcast, text broadcast and content recommendation video and other fields.

[0003] At present, the neural network model based on deep learning is used to generate the target image, which is based on the key points of the face of the virtual object, realizes the face reconstruction of the virtual object corresponding to the reference voice, and obtains the target image.

[0004] However, the face reconstruction of the virtual object based on the key points may only obtain a virtual object face with approximate face structure, and the details of the target image are missing, resulting in insufficient image clarity and fidelity, and insufficient user experience, therefore, there is an urgent need for an image driving method based on voice with high definition and fidelity. SUMMARY

[0005] Therefore, the embodiments of the present specification provide an image driving method based on voice. One or more embodiments of the present specification also relate to an image driving data processing method, an image driving device based on voice, an image driving data processing device, a computing device, a computer readable storage medium and a computer program, to solve the technical defects in the prior art.

[0006] In one embodiment of the present specification, an image driving method based on voice is provided, comprising:

[0007] obtaining a reference voice and a reference face image of a virtual object;

[0008] performing voice coding on the reference voice to obtain target voice features, and performing image coding on the reference face image to obtain first image features of a first region and second image features of a second region, wherein the first region is a region of the reference face image following the voice change, and the second region is a region of the reference face image other than the first region;

[0009] performing feature transformation on the first image features based on face prior features and the target voice features to determine first target image features, wherein the face prior features include face texture features, and the face prior features include face texture features;

[0010] generating a target image after driving according to the first target image features and the second image features.

[0011] In one or more embodiments of the present specification, a reference voice and a reference facial image of a virtual object are obtained, the reference voice is speech coded to obtain target speech features, and the reference facial image is image coded to obtain first image features of a first region and second image features of a second region, wherein the first region is a region of the reference facial image that changes with the voice, and the second region is a region of the reference facial image other than the first region. The first image features are feature transformed based on facial prior features and the target speech features to determine first target image features, wherein the facial prior features include facial texture features. The first target image features and the second image features are used to generate a driven target image. Since the facial prior features include facial texture features, the first image features of the first region that change with the voice are feature transformed based on the facial prior features and the target speech features, so that the obtained first target image features not only correspond to the speech features but also include the texture features. Finally, the second image features and the first target image features are used to generate a complete driven target image that corresponds to the reference voice and includes texture features. The target image has the characteristics of high fidelity and high definition, and the user experience is improved. BRIEF DESCRIPTION OF DRAWINGS

[0012] Figure 1 is a flowchart of a voice-based image driving method provided by one embodiment of the present specification;

[0013] Figure 2 is a flowchart of another voice-based image driving method provided by one embodiment of the present specification;

[0014] Figure 3 is a flowchart of another voice-based image driving method provided by one embodiment of the present specification;

[0015] Figure 4 is a flowchart of an image driving data processing method provided by one embodiment of the present specification;

[0016] Figure 5 is a flowchart of constructing facial prior features in a voice-based image driving method provided by one embodiment of the present specification;

[0017] Figure 6 is a flowchart of a voice-based image driving method provided by one embodiment of the present specification;

[0018] Figure 7 is a flowchart of a voice-based image driving method applied to face animation generation provided by one embodiment of the present specification;

[0019] Figure 8is a structural schematic diagram of a voice-based image driving device provided by an embodiment of the present specification;

[0020] Figure 9 is a structural schematic diagram of another voice-based image driving device provided by an embodiment of the present specification;

[0021] Figure 10 is a structural schematic diagram of another voice-based image driving device provided by an embodiment of the present specification;

[0022] Figure 11 is a structural schematic diagram of a data processing device of image driving provided by an embodiment of the present specification;

[0023] Figure 12 is a structural block diagram of a computing device provided by an embodiment of the present specification. DETAILED DESCRIPTION

[0024] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present specification. However, the present specification can be practiced without the specific details, other than in the examples, set forth in this description. Those skilled in the art, in light of the description, can implement the present specification without limiting the same to the specific details disclosed in this description.

[0025] The terminology used in this description of one or more embodiments of the present specification is for the purpose of describing particular embodiments only and is not intended to be limiting of one or more embodiments of the present specification. As used in this description and the appended claims of one or more embodiments of the present specification, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0026] It will be understood that, although the terms first, second, etc. can be used herein to describe various information, these terms are not intended to denote a temporal or chronological order. Rather, these terms are used solely to distinguish one from another only. For example, without departing from the scope of one or more embodiments of the present specification, first can be termed second, and similarly, second can be termed first. Depending on the context, the word "if' as used herein can be interpreted to mean "when" or "in response to determining" or "in response to a determination".

[0027] First, the noun terms related to one or more embodiments of the present specification are explained.

[0028] MLP (Multilayer Perceptron) model: a neural network model including an input layer, a hidden layer, and an output layer, and each layer is connected by full connection.

[0029] CNN (Convolutional Neural Networks) model: a multi-layer neural network model with forward propagation and back propagation, and has a convolution kernel (filter) for processing feature data.

[0030] RNN (Recurrent Neural Network) model: a recurrent neural network model that recursively processes in the direction of vector representation and each intermediate layer is connected in a chain.

[0031] Pointnet model: a key point space estimation model composed of two parts of a classification network and a segmentation network. The classification network classifies the key points of the input image. The segmentation network is an extension of the classification network, which outputs the spatial features of the key points in different categories through the global features and local features of the key points.

[0032] Pointnet++ model: a key point space estimation model that enhances the ability of local feature extraction based on the Pointnet model.

[0033] GCN (Graph Convolutional Networks) model: a convolutional neural network model applied to processing graph data.

[0034] LSTM (Long Short Term Memory) model: a neural network model with the ability to remember long and short term information, and has a convolution kernel (filter) for processing feature data.

[0035] Transformer model: a neural network model based on attention mechanism, which analyzes the features of the data through attention calculation.

[0036] BERT (Bidirectional Encoder Representation from Transformers) model: a neural network model with bidirectional attention encoding representation function.

[0037] GAN (Generative Adversarial Network) model: a deep learning neural network model, including a generator (Generator) and a discriminator (Discriminator), through the alternating training of the generator and the discriminator, a high-accuracy generator is obtained.

[0038] Diffusion model: a neural network model obtained by adding noise and removing noise for forward and backward derivation training.

[0039] In the present specification, a voice-based image driving method is provided, and the present specification also relates to an image-driven data processing method, a voice-based image driving device, an image-driven data processing device, a computing device, a computer-readable storage medium, and a computer program, which are described in detail one by one in the following embodiments.

[0040] Figure 1 A flowchart of a voice-based image driving method provided by an embodiment of the present specification is shown, which includes the following specific steps:

[0041] Step 102: obtaining a reference face image and a reference voice of a virtual object.

[0042] The embodiment of the present specification is applied to a client or a server with a virtual object face image driving platform, wherein the image driving platform includes but is not limited to: a live application, a webpage or an applet with a virtual anchor, an information broadcasting application, a webpage or an applet with a virtual object, a content recommendation application, a webpage or an applet with a virtual object.

[0043] The virtual object is an object that can be driven, which can be a module object pre-constructed by a module construction tool, for example, a character module or an animal module pre-constructed by a three-dimensional modeling tool, or an image object obtained by image mapping a real object in the physical world, for example, an image object of a real character or a real animal collected by an image collection device.

[0044] The reference face image is a visual image containing a virtual object face. The reference face image contains at least one visual image. The reference face image is in the form of a photo, a picture, or a video frame, etc. The reference face image is a visual image in a specific color space, for example, an RGB (Red-Green-Blue) image, an HIS (Hue-Saturation-Intensity) image, a YUV (luminance chrominance) image, a YCbCr (luminance offset) image, etc. The reference face image can be a visual image collected by an image collection device, for example, a human image, an animal image, etc. collected by an optical shooting device. The reference face image can also be an artificially generated visual image, for example, an animation image, a painting image, etc. The reference face image can also be a visual image generated by an image driving algorithm, which is not limited herein. The reference face image can be a two-dimensional image or a three-dimensional image, which is not limited herein.

[0045] The reference voice is voice data for guiding the face transformation of the virtual object. The reference voice can be natural language voice data, for example, Chinese voice, or non-natural language voice data, for example, animal calls, which is not limited herein.

[0046] The reference face image and the reference voice of the virtual object can be directly obtained from an image database and a voice database. The image database and the voice database can be a local database of the image driving platform, or a remote database, for example, a cloud database or an open source database of the image driving platform, etc. The reference face image and the reference voice of the virtual object can also be directly received from an image collection device and a voice collection device. The reference face image and the reference voice of the virtual object can also be directly received from a user through a front end of the image driving platform.

[0047] For example, a video collection device uploads a human video of a person A (the first N video frames of the video have high-definition and high-fidelity images, and the last M video frames have insufficient definition and fidelity). The Nth video frame in the human video is extracted as the reference face image, and the voice data corresponding to the last M video frames of the video is determined as the reference voice.

[0048] The reference face image and the reference voice of the virtual object lay a data foundation for subsequent encoding of the first image feature, the second image feature, and the target voice feature.

[0049] Step 104: speech coding is performed on the reference speech to obtain target speech features, and image coding is performed on the reference facial image to obtain first image features of a first region and second image features of a second region, wherein the first region is a region of the reference facial image that changes with the speech, and the second region is a region of the reference facial image other than the first region.

[0050] The first region is a region of the virtual object face that changes with the speech, for example, a mouth region of the virtual character, an eye region of the virtual character, a cheek region of the virtual character, etc., and the second region is a region of the virtual object face other than the first region, for example, a forehead region of the virtual character, an ear region of the virtual character, etc. It should be noted that the first region and the second region correspond to different virtual objects when the speaking habits of different virtual objects are different, and cannot be simply considered that the first region is the mouth region and the second region is a region of the virtual object face other than the mouth region. The first region and the second region can come from the same reference facial image or from different reference facial images, which are not limited herein.

[0051] The first image features are encoded feature vectors obtained by image coding on the first region, and the second image coding is encoded feature vectors obtained by image coding on the second region. The first image features and the second image features contain texture features of the image, image spatial features, etc.

[0052] The target speech features are encoded feature vectors obtained by speech coding on the reference speech. The text feature vector obtained by text recognition on the reference speech can also be an audio feature vector obtained by frequency division on the reference speech, which is not limited herein.

[0053] Speech coding is performed on the reference speech to obtain target speech features, and image coding is performed on the reference facial image to obtain first image features of a first region and second image features of a second region, and the specific manner is as follows: a preset coding algorithm is used to perform speech coding on the reference speech to obtain target speech features, and image coding is performed on the reference facial image to obtain first image features of a first region and second image features of a second region, wherein the preset coding algorithm can be a statistical coding algorithm, for example, a hash coding algorithm, a one-hot coding algorithm, etc., or a coding layer of a neural network model can be used for coding.

[0054] For example, a one-hot coding algorithm is used to perform speech coding on the reference speech to obtain target speech features, and a preset image hash coding algorithm is used to perform image coding on the reference facial image to obtain first image features of a mouth region of the A character and second image features of other facial regions of the A character.

[0055] The reference speech is speech coded to obtain target speech features, and the reference face image is image coded to obtain first image features of a first region and second image features of a second region, wherein the first region is a region of the reference face image that changes with the speech, and the second region is a region of the reference face image other than the first region, laying a feature foundation for subsequent feature transformation and providing image features for subsequent decoding to obtain a target image.

[0056] Step 106: using face prior features to perform attention calculation on the first image features to obtain attention image features, and performing feature optimization on the attention image features according to the target speech features to obtain first target image features of the first region, wherein the face prior features include face texture features.

[0057] The face prior features are prior features of virtual object face texture features, and the feature transformation is performed according to the prior features to ensure the accuracy of the target image features obtained after the feature transformation. The face texture features are surface properties of the face region of the virtual object in the image, such as the thickness and density of the image texture.

[0058] The first target image features are feature vectors of the first region corresponding to the reference speech and containing texture features. They correspond to the target speech features of the reference speech and contain texture features. Based on the first target image features, a high-fidelity and high-definition visual image of the first region corresponding to the reference speech can be obtained.

[0059] The first image feature is transformed based on the facial prior feature and the target voice feature to determine the first target image feature. Specifically, the facial prior feature is used to perform attention calculation on the first image feature to obtain an attention image feature. The attention image feature is optimized based on the target voice feature to obtain the first target image feature of the first region. Attention calculation is a feature transformation method of calculating an attention distribution on input information and then calculating a weighted average of the input information based on the attention distribution, so that a feature vector focusing on part of the information can be obtained. The facial prior feature contains facial texture features of virtual objects in a large number of sample images. Through attention calculation, the facial texture features of virtual objects in the global sample images are utilized, and the facial texture features with high correlation to the first image feature are focused on, so that the attention image feature containing the texture features is obtained. For example, the reference facial image is an image of a virtual object with black hair, a wide face, and large features. The facial prior feature contains facial texture features of virtual objects with black hair, a wide face, and large features. The attention image weight obtained by attention calculation focusing on the above texture features focuses on the facial texture features corresponding to the above features. Feature optimization is a feature transformation method of adaptively adjusting the image feature based on other features, so that the obtained first target image feature is adapted to the target voice feature. For example, the first region is the mouth region, and the reference voice is an open sound "a". The corresponding first target image feature represents a large mouth space and a tight mouth texture.

[0060] Exemplarily, the facial prior feature is used to perform attention calculation on the first image feature to obtain an attention image feature. The attention image feature is optimized based on the target voice feature to obtain the first target image feature of the mouth region of the A person.

[0061] The first image feature is transformed based on the facial prior feature and the target voice feature to determine the first target image feature, so that the obtained first target image feature not only corresponds to the voice feature but also contains texture features, thereby laying a foundation for subsequent generation of the driven target image.

[0062] Step 108: generating a driven target image based on the first target image feature and the second image feature.

[0063] The target image is a visual driving image containing a virtual object face corresponding to the reference voice. The target image contains a plurality of visual images. The target image can be a two-dimensional image or a three-dimensional image, which is not limited here.

[0064] According to the first target image feature and the second image feature, a driven target image is generated in the following manner: according to the first target image feature and the second image feature, a preset decoding algorithm is used to decode to obtain the target image of the virtual object. The preset decoding algorithm can be a statistical decoding algorithm, for example, a hash decoding algorithm, a one-hot decoding algorithm, etc., or a decoding layer of a neural network model can be used for decoding.

[0065] Illustratively, according to the first target image feature and the second image feature, an image hash decoding algorithm is used to decode the last M video frames of the A person corresponding to the reference voice, to obtain a complete person video with high definition and high fidelity.

[0066] In the embodiments of the present specification, the reference voice and the reference face image of the virtual object are obtained, the reference voice is speech encoded to obtain target voice features, and the reference face image is image encoded to obtain first image features of a first region and second image features of a second region, wherein the first region is a region of the reference face image that changes with the voice, and the second region is a region of the reference face image other than the first region. Based on the face prior feature and the target voice feature, the first image feature is feature-transformed to determine the first target image feature, wherein the face prior feature includes a face texture feature. According to the first target image feature and the second image feature, a driven target image is generated. The face prior feature including the face texture feature is based on the face prior feature and the target voice feature, and the first image feature of the first region that changes with the voice is feature-transformed, so that the obtained first target image feature not only corresponds to the voice feature but also includes the texture feature. Finally, according to the second image feature and the first target image feature, a complete driven target image corresponding to the reference voice and including the texture feature is generated. The target image has the characteristics of high fidelity and high definition, and the user experience is improved.

[0067] Optionally, before step 106, the following specific steps are further included:

[0068] The calibration face image of the virtual object is obtained;

[0069] The calibration face image is image encoded to obtain calibration image features of a target region, wherein the target region corresponds to the second region;

[0070] Based on the feature deviation between the second image feature and the calibration image feature, the first image feature is calibrated to obtain a calibrated first image feature.

[0071] When the reference face image is multiple different visual images containing the first region and / or the second region of the virtual object face, due to the different poses of the virtual object, i.e., the different spatial features of the face, the texture feature correspondence is transformed, and the matching of the texture features is problematic. For example, when the virtual object is a virtual person, the facial features such as the facial features and the face shape change when the face orientation is inconsistent, and thus need to be calibrated. Considering the structured information of the virtual object face, the geometric consistency of the target image is enhanced, the texture matching problem is not caused by the division of the face region, and the accuracy of the obtained target image features is ensured.

[0072] The calibration face image is a visual image different from the first region in spatial features and containing the virtual object face. The calibration face image is consistent with the reference face image in form and consistent with the reference face image in specific color space. For example, the reference face image is a visual image collected by a 30-degree lateral face-to-image collection device, and the calibration face image is a visual image collected by a forward face-to-image collection device. The two have different spatial features but correspond to the same virtual object. The virtual object face in the calibration face image contains a target region, and the target region corresponds to the second region but is not necessarily completely consistent. For example, the reference face image is an animal image containing part of the virtual animal face features, and the calibration image contains all the virtual animal face features.

[0073] The feature deviation is a spatial feature deviation between the second image feature and the calibration image feature, which is a feature vector with spatial distribution.

[0074] The calibration face image is image encoded to obtain the calibration image feature of the target region. Specifically, the calibration face image is image encoded by using a preset encoding algorithm to obtain the calibration image feature of the target region. The preset encoding algorithm in step 104 is consistent, and details are not described herein.

[0075] The first image feature is calibrated based on the feature deviation between the second image feature and the calibration image feature to obtain the calibrated first image feature. Specifically, the first image feature is calibrated based on the feature deviation between the second image feature and the calibration image feature by using a spatial feature transformation algorithm to obtain the calibrated first image feature. The spatial feature transformation algorithm is an algorithm for realizing spatial feature transformation by interpolating the feature vector, such as a bilinear interpolation algorithm (Bilinear Interpolating), a bilinear sampling algorithm (Bilinear Sampling), etc.

[0076] Exemplarily, the N-10th video frame in the A-person video is extracted as a calibration face image, the image hash coding algorithm is used to code the calibration face image to obtain the calibration image features of other face regions, the second image features are compared with the calibration image features to obtain feature deviations, and the bilinear sampling algorithm is used to calibrate the first image features according to the feature deviations to obtain the calibrated first image features.

[0077] The calibration face image of the virtual object is obtained, the calibration face image is coded to obtain the calibration image features of the target region, the second image features are compared with the calibration image features to obtain feature deviations, and the first image features are calibrated according to the feature deviations to obtain the calibrated first image features. The accuracy of the subsequently obtained first target image features is ensured.

[0078] Optionally, comparing the second image features with the calibration image features to obtain the feature deviations comprises the following specific steps:

[0079] The second image features are key point coded to obtain second key point spatial features, and the calibration image features are key point coded to obtain calibration key point spatial features.

[0080] The feature deviations are determined according to the feature differences between the second key point spatial features and the calibration key point spatial features.

[0081] The first image features are calibrated based on the feature deviations to obtain the calibrated first image features.

[0082] The spatial features can be determined by sampling and coding the key points on the face of the virtual object. For example, the key points such as eyes, nose, and eyebrows on the face of the virtual person are sampled and coded. Thus, the feature deviations can be understood as follows: a spatial coordinate system is constructed on the key points in the second region of the reference face image, each key point has a corresponding coordinate position, the corresponding key points in the second region of the calibration image are mapped to the spatial coordinate system, each corresponding key point has a corresponding coordinate position, and the spatial feature deviations are determined by comparing the differences between the coordinate positions.

[0083] The second key point spatial features are feature vectors representing the key point spatial features of the second region, and the calibration key point spatial features are feature vectors representing the corresponding key point spatial features of the second region in the calibration image.

[0084] The second image feature is key point coded to obtain a second key point space feature, and the calibration image feature is key point coded to obtain a calibration key point space feature. Specifically, the second image feature is key point coded by using an encoding layer of a pre-trained space adaptation model to obtain the second key point space feature, and the calibration image feature is key point coded to obtain the calibration key point space feature. The space adaptation model is a neural network model with a space adaptation adjustment function for image features. The space adaptation model includes an encoding layer, a calculation module, and a decoding layer. The space adaptation model can implement the space feature transformation algorithm described in the above embodiments. The space adaptation model can be an MLP model, a CNN model, an RMM model, a GCN model, a Pointnet model, a Pointnet++ model, or the like.

[0085] The first image feature is calibrated based on the feature deviation between the second key point space feature and the calibration key point space feature to obtain a calibrated first image feature. Specifically, the feature deviation is calculated by using the calculation module of the space adaptation model according to the feature difference between the second key point space feature and the calibration key point space feature. The first image feature is calibrated based on the feature deviation to obtain the calibrated first image feature.

[0086] The second image feature is key point coded to obtain a second key point space feature, and the calibration image feature is key point coded to obtain a calibration key point space feature. Specifically, the second image feature is key point coded by using an encoding layer of a pre-trained space adaptation model to obtain the second key point space feature, and the calibration image feature is key point coded to obtain the calibration key point space feature. The space adaptation model is a neural network model with a space adaptation adjustment function for image features. The space adaptation model includes an encoding layer, a calculation module, and a decoding layer. The space adaptation model can implement the space feature transformation algorithm described in the above embodiments. The space adaptation model can be an MLP model, a CNN model, an RMM model, a GCN model, a Pointnet model, a Pointnet++ model, or the like.

[0087] The first image feature is calibrated based on the feature deviation between the second key point space feature and the calibration key point space feature to obtain a calibrated first image feature. Specifically, the feature deviation is calculated by using the calculation module of the space adaptation model according to the feature difference between the second key point space feature and the calibration key point space feature. The first image feature is calibrated based on the feature deviation to obtain the calibrated first image feature.

[0088] Optionally, step 106 includes the following specific steps:

[0089] The first image feature is attention calculated by using the face prior feature to obtain a first attention image feature.

[0090] The first attention image feature is normalized by using the target voice feature as a constraint condition to obtain a first target image feature.

[0091] The first attention image feature is an image feature containing a corresponding texture feature in the face prior feature. The first attention image feature is more focused on a specific face texture feature under the guidance of the face prior feature.

[0092] The face prior feature is determined as a key vector and a value vector, the first image feature is determined as a query vector, and attention calculation is performed according to the query vector, the key vector and the value vector to obtain the first attention image feature. The specific formula of the attention calculation is shown in formula 1:

[0093]

[0094] Wherein, H is the output feature vector, Q is the input feature vector, that is, the query vector Query, K is the key vector Key, V is the value vector Value, d k is a preset temperature factor.

[0095] The normalization processing is to limit the image feature in a certain range (such as [0, 1] or [-1, 1]) and take the target speech feature as a constraint condition, so as to eliminate the influence of singular values in the first attention image feature. The effect is that the obtained first target image feature is adapted to the target speech feature.

[0096] It should be noted that the above process is implemented in a neural network model with an attention mechanism in the pre-training process, for example, a Transformer model, a BERT model, etc.

[0097] Exemplarily, the face prior feature is determined as a key vector and a value vector, the first image feature is determined as a query vector Query, attention calculation is performed according to the query vector, the key vector and the value vector by using formula 1 to obtain the first attention image feature, and the first attention image feature is normalized with the target speech feature as a constraint condition to obtain the first target image feature of the mouth region of the A person.

[0098] The face prior feature is used to perform attention calculation on the first image feature to obtain the first attention image feature, and the first attention image feature is normalized with the target speech feature as a constraint condition to obtain the first target image feature of the first region. The first target image feature and the target speech feature are improved in pertinence, and the texture feature contained in the first target image feature is improved in pertinence.

[0099] Optionally, the facial prior features include first facial prior features corresponding to the first region and second facial prior features corresponding to the second region.

[0100] Correspondingly, before step 106, the method further includes the following specific steps:

[0101] Discretize the first image features based on the first facial prior features to obtain discretized first image features.

[0102] Discretize the second image features based on the second facial prior features to obtain discretized second image features.

[0103] Since the facial prior features encode the first region and the second region respectively, the facial prior features can be constructed respectively when constructing the facial prior features, thereby improving the pertinence of the facial prior features and the pertinence of the feature transformation.

[0104] The first facial prior features are a set of facial texture features of the first region of the face of the virtual object in the sample image, and the second facial prior features are a set of facial texture features of the second region of the face of the virtual object in the sample image. The first facial prior features and the second facial prior features are a set of discrete feature vectors, which are obtained by discretizing the features of the face of the virtual object in the sample image and include discretized texture features, for example, discretized feature extraction of the face according to discrete regions (such as three-court regions) or discretized feature extraction of the face according to the five organs and the nearby regions.

[0105] The discretization processing is a limited feature vector mapping, which maps a feature vector with more dimensions to a feature vector with fewer dimensions, can reduce the amount of subsequent data processing, improve the processing efficiency, ensure the stability of the generated results, and reduce the training difficulty of the related neural network model.

[0106] Discretize the first image features based on the first facial prior features to obtain discretized first image features, and the specific manner is: discretize the first image features based on the feature similarity of each discrete feature in the first image features and the first facial prior features to obtain discretized first image features.

[0107] Discretize the second image features based on the second facial prior features to obtain discretized second image features, and the specific manner is: discretize the second image features based on the feature similarity of each discrete feature in the second image features and the second facial prior features to obtain discretized second image features.

[0108] Exemplarily, the first image feature is discretized based on the feature similarity of each discrete feature in the first image feature and the first facial prior feature, to obtain a discretized first image feature, and the second image feature is discretized based on the feature similarity of each discrete feature in the second image feature and the second facial prior feature, to obtain a discretized second image feature.

[0109] The first image feature is discretized based on the first facial prior feature, to obtain a discretized first image feature, and the second image feature is discretized based on the second facial prior feature, to obtain a discretized second image feature, which improves the image driving efficiency, ensures the stability of image driving, and reduces the complexity of image driving.

[0110] Optionally, step 104 comprises the following specific steps:

[0111] The pre-trained image driving model is obtained, wherein the image driving model comprises an image encoding layer, a speech encoding layer, a feature transformation layer, and a decoding layer;

[0112] The reference speech is input into the speech encoding layer to obtain target speech features, and the reference facial image is input into the image encoding layer to obtain first image features of a first region and second image features of a second region;

[0113] Correspondingly, step 106 comprises the following specific steps:

[0114] The first image features, the target speech features, and the facial prior features are input into the feature transformation layer, the first image features are transformed based on the facial prior features and the target speech features, and first target image features are determined;

[0115] Correspondingly, step 108 comprises the following specific steps:

[0116] The first target image features and the second image features are merged to obtain merged image features;

[0117] The merged image features are input into the decoding layer to obtain a driven target image.

[0118] The image driving model is a neural network model with the function of driving a virtual object by an image, and is an image processing model, including but not limited to an LSTM model, a Transformer model, a BERT model, a GAN model, a Diffusion model, etc. The image driving model includes an image encoding layer, a speech encoding layer, a feature transformation layer, and a decoding layer. The feature transformation layer includes an attention calculation layer and a conditional normalization layer (CLN, Conditional Layer-Normalization), and the feature transformation layer can be a Transformer model. The image driving model is pre-trained for the task of driving the face of a virtual object with high definition and high fidelity.

[0119] The merged image feature is a feature vector obtained by merging the first target image feature and the second image feature in space. For example, the first region is the mouth region of the virtual character, and the second region is other facial regions of the virtual character. By merging the first target image feature of the mouth region and the second image feature of the other facial regions in space, the merged image feature is an image feature representing the full face of the virtual character.

[0120] The feature merging is a merging method of feature vectors in space, which is specifically implemented by a merging function (Concat function).

[0121] The first image feature, the target speech feature, and the facial prior feature are input into the feature transformation layer. Based on the facial prior feature and the target speech feature, the first image feature is transformed to determine the first target image feature. Specifically, the first image feature and the facial prior feature are input into the attention calculation layer in the feature transformation layer to calculate the first attention image feature. The first attention image feature and the target speech feature are input into the conditional normalization layer in the feature transformation layer to obtain the first target image feature. The specific feature processing method of the attention layer and the conditional normalization layer has been described in the above embodiments, and will not be repeated here.

[0122] Exemplarily, a pre-trained GAN model having an image driving function for a virtual object is acquired, wherein the GAN model comprises an image encoding layer, a speech encoding layer, an attention calculation layer, a conditional normalization layer, and a decoding layer; a reference speech is input into the speech encoding layer to obtain target speech features, and a reference facial image of a person A is input into the image encoding layer to obtain first image features of a mouth region of the person A and second image features of other facial regions of the person A; the first image features and facial prior features are input into the attention calculation layer to calculate first attention image features; the first attention image features and the target speech features are input into the conditional normalization layer to obtain first target image features; the first target image features and the second image features are merged by using a Concat function to obtain merged image features; and the merged image features are input into the decoding layer to obtain the last M video frames of the person A corresponding to the reference speech, i.e., a complete person video with high definition and high fidelity.

[0123] A pre-trained image driving model is acquired, wherein the image driving model comprises an image encoding layer, a speech encoding layer, a feature transformation layer, and a decoding layer; a reference speech is input into the speech encoding layer to obtain target speech features, and a reference facial image is input into the image encoding layer to obtain first image features of a first region and second image features of a second region; the first image features, the target speech features, and facial prior features are input into the feature transformation layer to transform the first image features based on the facial prior features and the target speech features to determine first target image features; the first target image features and the second image features are merged to obtain merged image features; and the merged image features are input into the decoding layer to obtain a target image after driving. The efficiency of image driving is improved, the fidelity and definition of the target image are further improved, and the user experience is further improved.

[0124] Optionally, before acquiring the pre-trained image driving model, the following specific steps are further included:

[0125] A training sample set is acquired, wherein the training sample set comprises a plurality of training sample groups, and any training sample group comprises a sample image of a sample virtual object, a label image of the sample virtual object, and a sample speech corresponding to the label image.

[0126] The facial prior features are taken as prior features for feature transformation, and the image driving model is supervisedly trained according to the sample images, the sample speeches, and the label images of the training sample groups to obtain the trained image driving model.

[0127] The training sample set is a set of sample images used for training of the image-driven model, and includes a plurality of training sample groups. Any training sample group includes a sample image of a sample virtual object, a label image of the sample virtual object, and sample speech corresponding to the label image. The sample image is a visual sample image containing a face of the virtual object. The label image is a target image, which is a visual sample image containing a face of the virtual object corresponding to reference speech. The sample image can be a visual sample image collected by an image collection device, a visual sample image artificially generated, or a visual sample image generated by using an image generation algorithm, without limitation. The label image is generated or collected in the same manner as the sample image. The sample speech is sample speech data for guiding face transformation corresponding to the label image, which can be natural language sample speech data or non-natural language sample speech data, without limitation. To ensure the training effect of the image-driven model, the training sample set has a large size, which is generally obtained from an open source database, for example, a sample video database, a video database of an online video platform, or the like. To ensure the training efficiency of the image-driven model, the training sample set has a small size, which is generally obtained from a local database.

[0128] The supervised training is a manner of training a neural network model by using label data. In the embodiments of the present specification, the face prior feature is taken as the prior feature, the sample image and the sample speech are input into the image-driven model to generate a predicted image, a training loss value is calculated according to the predicted image and the label image, the model parameters of the image-driven model are adjusted according to the training loss value, and the above steps are iteratively repeated until a preset training end condition is reached, to obtain a trained image-driven model. The training loss value includes but is not limited to a cosine loss value, an L1 loss value, an L2 loss value, and a cross-entropy loss value. The training end condition includes but is not limited to a preset loss value threshold, a preset iteration number, and a judgment condition for completing training of each sample group.

[0129] The face prior feature is taken as the prior feature, the sample image and the sample speech are input into the image-driven model to generate a predicted image, and the specific manner is as follows: the face prior feature is taken as the prior feature, the sample image and the sample speech are input into an image encoding layer of the image-driven model to correspondingly obtain a first image feature of a first region, a second image feature of a second region, and a sample speech feature, the face prior feature is taken as the prior feature, the first image feature is feature-transformed by using a feature transformation layer according to the sample speech feature to determine a target image feature of the first region, and the second image feature and the target image feature are decoded by using a decoding layer to obtain the predicted image.

[0130] According to the training loss value, the model parameters of the image driving model are adjusted in the following manner: according to the training loss value, the model parameters of the image encoding layer, the speech encoding layer, the feature transformation layer and the decoding layer are adjusted by using a gradient update method. It should be noted that, in order to ensure the consistency of the image encoding in the training process and the application process, the parameters of the image encoding layer can be fixed.

[0131] Exemplarily, video data in an open-source video database is obtained, a training sample set is constructed according to the video data, a face prior feature is taken as a prior feature, a sample image and a sample speech are input into a GAN model to generate a predicted image, a cross-entropy loss value is calculated according to the predicted image and a label image, and the parameters of a speech encoding layer, a feature transformation layer and a decoding layer in the GAN model are adjusted by using a gradient update method according to the cross-entropy loss value. The above steps are iteratively repeated until a preset loss value threshold is reached, and a trained GAN model is obtained.

[0132] The training sample set is obtained, wherein the training sample set includes a plurality of training sample groups, any training sample group includes a sample image of a sample virtual object, a label image of the sample virtual object and a sample speech corresponding to the label image, a face prior feature is taken as a prior feature for feature transformation, the image driving model is supervised trained according to the sample image, the sample speech and the label image of each training sample group, and a trained image driving model is obtained. The face prior feature is obtained by pre-extracting the face region feature of the virtual object in the sample image set by using the pre-trained face reconstruction model, the face prior feature is taken as a prior feature for feature transformation, the image driving model is supervised trained according to the sample image, the sample speech and the label image, a trained image driving model is obtained, and the feature extraction capability of the image driving model for the speech data and the image data is improved, so that the subsequent generated target image driven has the characteristics of high fidelity and high definition.

[0133] Optionally, the face prior feature is obtained by pre-extracting the face region feature of the sample image by using the face reconstruction model, and correspondingly, the method further includes the following specific steps:

[0134] The sample image set is obtained, wherein the sample image set includes a plurality of sample image pairs, and any sample image pair includes a first region sample image and a second region sample image corresponding to the same virtual object;

[0135] For any sample image pair, the first region sample image and the second region sample image in the sample image pair are respectively subjected to texture feature extraction by using the encoding layer of the pre-trained face reconstruction model, and the texture feature of the first region sample image and the texture feature of the second region sample image are obtained.

[0136] Integrate the texture features of each first region sample image and the texture features of each second region sample image to obtain the face prior feature.

[0137] The face prior feature can be a set of face texture features of the virtual object in the sample image, specifically a set of feature vectors. The face prior feature is constructed by using a pre-trained face reconstruction model to extract features from the face region of a large number of virtual objects in a large number of sample images in the sample image set. It contains a large number of texture features of virtual object faces, such as hair color, and a large number of spatial features of virtual object faces, such as face shape and facial features. The face prior feature realizes feature generalization of the virtual object face in the sample image, and has high migration and universality. The face prior feature is constructed in advance under the supervision of a high-definition and high-fidelity virtual object face reconstruction task.

[0138] The sample image set is a set of sample images constructed in advance for face prior feature construction, including a plurality of sample image pairs, each sample image pair including a first region sample image and a second region sample image corresponding to the same virtual object. The first region sample image and the second region sample image are obtained by region splitting of the same sample image. The sample image corresponding to the first region sample image and the second region sample image can be a visual sample image collected by an image collection device, an artificially generated visual sample image, or a visual sample image generated by an image-driven algorithm, without limitation. To ensure the high migration, universality and accuracy of the constructed face prior feature, the sample image set is large in size, generally obtained from an open source database, such as a sample image database or an image database of an online image platform.

[0139] The face reconstruction model is a neural network model with virtual object face reconstruction function, including an encoding layer with texture feature extraction function and a decoding layer with image-driven function. The face reconstruction model can be an MLP model, a CNN model, an RMM model, a GCN model, a Pointnet model, a Pointnet++ model, etc. It should be noted that the face reconstruction model can be the same as the image-driven model, or different.

[0140] Integrating the texture features of each first region sample image and the texture features of each second region sample image to obtain the face prior feature can be to integrate and construct a first face prior feature and a second face prior feature, or to integrate and construct one face prior feature, without limitation.

[0141] Exemplarily, a sample image set is obtained from an open-source image database, wherein the sample image set includes 100,000 sample image pairs, any sample image pair includes a first region sample image and a second region sample image corresponding to a same virtual character, for any sample image pair, texture feature extraction is performed on the first region sample image and the second region sample image in the sample image pair respectively by using an encoding layer of a pre-trained Diffusion model, to obtain texture features of the first region sample image and texture features of the second region sample image, and the texture features of the 100,000 first region sample images and the texture features of the 100,000 second region sample images are integrated to obtain the face prior feature.

[0142] A sample image set is obtained, wherein the sample image set includes a plurality of sample image pairs, any sample image pair includes a first region sample image and a second region sample image corresponding to a same virtual object, for any sample image pair, texture feature extraction is performed on the first region sample image and the second region sample image in the sample image pair respectively by using an encoding layer of a pre-trained face reconstruction model, to obtain texture features of the first region sample image and texture features of the second region sample image, and the texture features of the first region sample images and the texture features of the second region sample images are integrated to obtain the face prior feature. The face prior feature containing texture features is obtained by pre-training the first region sample images and the second region sample images in the sample image set by using the pre-trained face reconstruction model, so that the subsequently obtained target image feature not only corresponds to the speech feature but also contains the texture feature, and then a complete target image of the virtual object corresponding to the reference speech and containing the texture feature is obtained by subsequent decoding, the target image has the characteristics of high fidelity and high definition, improves the user experience, and improves the transferability and universality of image driving.

[0143] Optionally, before the texture feature extraction of the first region sample image and the second region sample image in the sample image pair by using the encoding layer of the pre-trained face reconstruction model, the following specific steps are further included:

[0144] A pre-training set is obtained, wherein the pre-training set includes a plurality of pre-training pairs, any pre-training pair includes a first training sample image of a first region and a second training sample image of a second region corresponding to a same sample virtual object;

[0145] The face reconstruction model is supervised trained according to the first training sample image and the second training sample image of each pre-training pair, to obtain the trained face reconstruction model.

[0146] The pre-training set is a set of sample images pre-constructed for pre-training of the face reconstruction model, and includes a plurality of pre-training pairs. Any pre-training pair includes a first training sample image corresponding to a first region of a same sample virtual object and a second training sample image corresponding to a second region of the same sample virtual object. The first training sample image and the second training sample image are obtained by regionally splitting a same sample image. The sample image corresponding to the first training sample image and the second training sample image can be a visual sample image collected by an image collection device, an artificially generated visual sample image, or a visual sample image generated by using an image driving algorithm, which is not limited herein. In order to ensure the texture feature extraction capability of the face reconstruction model obtained by pre-training, the pre-training set has a large size, and is generally obtained by accessing an open source database, for example, a sample image database, an image database of an online image platform, and the like.

[0147] The supervised training is a manner of training a neural network model by using labeled data. The first training sample image and the second training sample image in the pre-training pair are labeled data. In the embodiments of the present specification, the face reconstruction model is supervised trained according to the first training sample image and the second training sample image of each pre-training pair, to obtain a trained face reconstruction model. The specific manner is as follows: the first training sample image is input into the face reconstruction model to generate a first prediction image; a training loss value is calculated according to the first prediction image and the second training sample image; the model parameters of the face reconstruction model are adjusted according to the training loss value; the above steps are iteratively repeated until a preset training end condition is reached, to obtain a first-stage trained face reconstruction model; the second training sample image is input into the face reconstruction model to generate a second prediction image; a training loss value is calculated according to the second prediction image and the first training sample image; the model parameters of the face reconstruction model are adjusted according to the training loss value; the above steps are iteratively repeated until a preset training end condition is reached, to obtain a trained face reconstruction model. The training order of the first training sample image and the second training sample image can be changed. The training loss value includes but is not limited to a cosine loss value, an L1 loss value, an L2 loss value, and a cross-entropy loss value. The training end condition includes but is not limited to a preset loss value threshold, a preset iteration number, and a judgment condition for completing training of each sample group.

[0148] The first training sample image is input into the encoding layer of the face reconstruction model to extract a first texture feature, and the first texture feature is input into the decoding layer of the face reconstruction model to generate the first prediction image. The second training sample image is input into the face reconstruction model to generate the second prediction image, and the same applies.

[0149] According to the training loss value, the model parameters of the image driving model are adjusted, specifically: according to the training loss value, the model parameters of the encoding layer and the decoding layer are adjusted using the gradient update method.

[0150] Exemplarily, a pre-training set is obtained from an open-source image database, wherein the sample image set includes 10,000 pre-training pairs, any pre-training pair includes a first training sample image of a first region and a second training sample image of a second region of the same sample virtual object, the first training sample image is input into the Diffusion model to generate a first predicted image, the cross-entropy loss value is calculated according to the first predicted image and the second training sample image, the model parameters of the Diffusion model are adjusted using the gradient update method according to the cross-entropy loss value, the above steps are iteratively repeated until a preset iteration number is reached, and the Diffusion model trained in the first stage is obtained. The second training sample image is input to generate a second predicted image, the cross-entropy loss value is calculated according to the second predicted image and the first training sample image, the model parameters of the Diffusion model are adjusted using the gradient update method according to the cross-entropy loss value, the above steps are iteratively repeated until a preset iteration number is reached, and the trained Diffusion model is obtained.

[0151] A pre-training set is obtained, wherein the pre-training set includes a plurality of pre-training pairs, any pre-training pair includes a first training sample image of a first region and a second training sample image of a second region corresponding to the same sample virtual object, and the face reconstruction model is supervised trained according to the first training sample image and the second training sample image of each pre-training pair to obtain the trained face reconstruction model. The texture feature extraction capability of the encoding layer of the face reconstruction model is improved, and high-accuracy face prior features are ensured to be extracted and constructed subsequently.

[0152] Figure 2 A flowchart of another voice-based image driving method according to an embodiment of the present specification is shown, which is applied to a cloud-side device and includes the following specific steps:

[0153] Step 202: receiving an image driving request for a virtual object sent by an end-side device, wherein the image driving request carries a reference face image and a reference voice of the virtual object;

[0154] Step 204: performing voice coding on the reference voice to obtain target voice features, and performing image coding on the reference face image to obtain first image features of a first region and second image features of a second region, wherein the first region is a region of the reference face image that changes with the voice, and the second region is a region of the reference face image other than the first region;

[0155] Step 206: performing attention calculation on the first image feature by using the face prior feature to obtain an attention image feature, performing feature optimization on the attention image feature according to the target voice feature to obtain a first target image feature of the first region, wherein the face prior feature comprises a face texture feature;

[0156] Step 208: generating a target image after driving according to the first target image feature and the second image feature;

[0157] Step 210: sending the target image to the terminal device for rendering.

[0158] The cloud-side device is a network cloud-side device providing a virtual object face image driving function, and is a virtual device. The terminal device is a terminal device of a client or a server of an application, a webpage or a small program platform providing a virtual object face image driving function, and is an entity device. The cloud-side device and the terminal device are connected through a network transmission channel to perform data transmission. The computing power performance of the cloud-side device is higher than that of the terminal device.

[0159] Steps 204 to 208 have been described in detail in steps 102 to 108 of the embodiment, and will not be described here again. Figure 1 Steps 102 to 108 of the embodiment have been described in detail, and will not be described here again.

[0160] The terminal device realizes the rendering and display of the target image through a renderer.

[0161] In the embodiments of the present specification, the receiving end side device sends an image driving request for the virtual object, wherein the image driving request carries a reference face image and a reference voice of the virtual object, the reference voice is speech coded to obtain target voice features, and the reference face image is image coded to obtain first image features of a first region and second image features of a second region, wherein the first region is a region of the reference face image that changes with the voice, and the second region is a region of the reference face image other than the first region, the first image features are feature transformed based on face prior features and the target voice features to determine first target image features, wherein the face prior features include face texture features, and a target image after driving is generated according to the first target image features and the second image features, and the target image is sent to the end side device for rendering. Since the face prior features include face texture features, the first image features of the first region that change with the voice are feature transformed based on the face prior features and the target voice features, so that the obtained first target image features not only correspond to the voice features but also include the texture features, and finally, the complete target image after driving that corresponds to the reference voice and includes the texture features is generated according to the second image features and the first target image features, the target image has the characteristics of high fidelity and high definition, the user experience is improved, at the same time, the image driving is implemented on the cloud side device with higher computing power, the efficiency of the image driving is improved, and the computing power cost of the end side device is reduced.

[0162] Figure 3 A flowchart of another voice-based image driving method provided according to an embodiment of the present specification is shown, which is applied to an augmented reality (AR) device and includes the following specific steps:

[0163] Step 302: receiving an image driving request for a virtual object, wherein the image driving request carries a reference face image and a reference voice of the virtual object;

[0164] Step 304: speech coding the reference voice to obtain target voice features, and image coding the reference face image to obtain first image features of a first region and second image features of a second region, wherein the first region is a region of the reference face image that changes with the voice, and the second region is a region of the reference face image other than the first region;

[0165] Step 306: using face prior features to perform attention calculation on the first image features to obtain attention image features, and using the target voice features to perform feature optimization on the attention image features to obtain first target image features of the first region, wherein the face prior features include face texture features;

[0166] Step 308: generating a target image after driving according to the first target image features and the second image features;

[0167] Step 310: rendering the target image.

[0168] The embodiment of the present specification is applied to a reality augmented AR (Augmented Reality) device providing a virtual object face image driving function, wherein the platform can be a reality augmented game application, a reality augmented live broadcast platform, a reality augmented information broadcast application, a reality augmented content recommendation application, etc.

[0169] Steps 304 to 308 have been described in detail in the embodiment of the present specification, and will not be repeated here. Figure 1 Steps 104 to 108 of the embodiment are described in detail in the embodiment of the present specification, and will not be repeated here.

[0170] The target image is rendered, specifically: the target image is reality augmented rendering. Wherein, the reality augmented rendering is realized by a reality augmented renderer.

[0171] In the embodiment of the present specification, an image driving request for a virtual object is received, wherein the image driving request carries a reference face image of the virtual object and a reference voice, the reference voice is voice coded to obtain target voice features, and the reference face image is image coded to obtain first image features of a first region and second image features of a second region, wherein the first region is a region of the reference face image that changes with the voice, and the second region is a region of the reference face image other than the first region, the first image features are feature transformed based on face prior features and the target voice features to determine first target image features, wherein the face prior features include face texture features, the driven target image is generated according to the first target image features and the second image features, and the target image is rendered. Since the face prior features include the face texture features, the first image features of the first region that change with the voice are feature transformed based on the face prior features and the target voice features, so that the obtained first target image features not only correspond to the voice features but also contain the texture features, and finally the driven target image that is complete, corresponds to the reference voice and contains the texture features is generated according to the second image features and the first target image features, the target image has the characteristics of high fidelity and high definition, improving the user experience, and the image driving is implemented in the reality augmented AR device with stronger rendering effect, improving the rendering effect of the driven target image and further improving the user experience.

[0172] Figure 4 A flowchart of an image driving data processing method according to an embodiment of the present specification is shown, the method is applied to a cloud side device, including the following specific steps:

[0173] Step 402: obtaining a training sample set, wherein the training sample set includes a plurality of training sample groups, any training sample group includes a sample image of a sample virtual object, a label image of the sample virtual object, and sample speech corresponding to the label image;

[0174] Step 404: using a face prior feature as a feature transformation prior feature, performing supervised training on the image driving model according to the sample image, the sample speech, and the label image of each training sample group, and obtaining a trained image driving model, wherein the face prior feature includes a face region feature extracted from the virtual object in the sample image by using a pre-trained face reconstruction model, and the face region feature includes a texture feature;

[0175] Step 406: sending model parameters of the trained image driving model to an end-side device.

[0176] The cloud-side device is a network cloud-side device providing a model training function, and is a virtual device. The end-side device is a terminal device providing a virtual object face image driving function, and is an entity device. The end-side device and the cloud-side device are connected through a network channel for data transmission. The computing power performance of the cloud-side device is higher than that of the end-side device.

[0177] Steps 402 to 404 have been described in detail in the Figure 1 embodiment and will not be repeated here.

[0178] In the embodiment of the present specification, a training sample set is obtained, wherein the training sample set includes a plurality of training sample groups, any training sample group includes a sample image of a sample virtual object, a label image of the sample virtual object, and sample speech corresponding to the label image, a face prior feature is used as a feature transformation prior feature, the image driving model is supervised trained according to the sample image, the sample speech, and the label image of each training sample group, a trained image driving model is obtained, wherein the face prior feature includes a face region feature extracted from the virtual object in the sample image by using a pre-trained face reconstruction model, and the face region feature includes a texture feature, and model parameters of the trained image driving model are sent to an end-side device. The face region feature extracted from the virtual object in the sample image by using the pre-trained face reconstruction model in advance obtains the face prior feature including the texture feature, the image driving model is supervised trained according to the sample image, the sample speech, and the label image by using the face prior feature as the feature transformation prior feature, the trained image driving model is obtained, the feature extraction capability of the image driving model for the speech data and the image data and the texture feature is improved, the generated target image has the characteristics of high fidelity and high definition, at the same time, the model training is implemented on the cloud-side device with higher computing power, the efficiency of the model training is improved, and the computing power cost of the end-side device is reduced.

[0179] Figure 5 A flowchart of constructing face prior features in a voice-based image driving method provided by one embodiment of the present specification is shown.

[0180] As shown in Figure 5 , a sample image set is obtained, any sample image is extracted, the sample image is split to obtain a first region sample image and a second region sample image, the texture features of the first region sample image and the second region sample image are extracted by using the encoding layer of the face reconstruction model, the texture features of the first region sample image and the texture features of the second region sample image are obtained, and the texture features of the first region sample image and the texture features of the second region sample image are decoded by using the decoding layer of the face reconstruction model to obtain a predicted image. The texture features of each first region sample image are integrated to obtain a first face prior feature, and the texture features of each second region sample image are integrated to obtain a second face prior feature.

[0181] Figure 6 A flowchart of a voice-based image driving method provided by one embodiment of the present specification is shown.

[0182] As shown in Figure 6As shown, the calibration face image of the other face region except the mouth region, the reference face image of the other face region, the reference face image of the mouth region are encoded by using the image encoding layer to obtain the calibration image feature, the second image feature and the first image feature respectively, the target speech feature is obtained by encoding the reference speech by using the speech encoding layer, the calibration image feature and the second image feature are discretized based on the face prior feature of the other face region to obtain the discretized calibration image feature and the discretized second image feature, the first image feature is discretized based on the face prior feature of the mouth region to obtain the discretized first image feature, the discretized calibration image feature and the discretized second image feature are input into the key point encoding layer of the adaptive face alignment model to obtain the calibration key point space feature and the second key point space feature, the feature deviation is obtained by decoding the calibration key point space feature and the second key point space feature by using the key point decoding layer of the adaptive face alignment model, the feature deviation and the discretized first image feature are input into the bilinear sampling module for interpolation processing to obtain the calibrated first image feature, the calibrated first image feature is taken as Q (query vector), the face prior feature of the mouth region is taken as K (key vector) and V (value vector), the attention image feature is obtained by using the attention calculation layer of the Transformer model, the target speech feature is input into the conditional normalization layer as a constraint condition to normalize the attention image feature, and the first target image feature corresponding to the mouth region is obtained through the forward feedback neural network layer (FFN) and the discrete smoothing approximation layer (Gumbel Softmax) processing, the second image feature and the first target image feature are merged to obtain the merged image feature, and the target image after driving is obtained by decoding the merged image feature by using the image decoding layer.

[0183] The following description is made in conjunction with the accompanying drawings Figure 7 The speech-based image driving method provided in the specification is further described taking the application of the speech-based image driving method in face animation generation as an example. Among them, Figure 7 A processing process flow diagram of a speech-based image driving method applied to face animation generation provided by an embodiment of the specification is shown, including the following specific steps:

[0184] Step 702: Obtain a sample image set, wherein the sample image set includes a plurality of sample image pairs, and any sample image pair includes a sample image corresponding to a mouth region of a sample animation character and a sample image of other face regions, wherein the other face regions are other face regions except the mouth region.

[0185] Step 704: For each sample image pair, texture features of the sample image of the mouth region and the sample image of the other facial region are extracted respectively by using the image encoding layer of the pre-trained face reconstruction model, to obtain the texture features of the sample image of the mouth region and the texture features of the sample image of the other facial region.

[0186] In the embodiments of the present specification, the face reconstruction model is a GAN model.

[0187] Step 706: The texture features of each first region sample image are integrated to obtain the mouth region face prior feature, and the texture features of each other facial region sample image are integrated to obtain the other facial region face prior feature.

[0188] Step 708: Obtain the calibration face image of the other facial region of the target animation character, the reference face image of the other facial region, the reference face image of the mouth region, and the reference voice.

[0189] Step 710: The image encoding layer of the face reconstruction model is used to encode the calibration face image of the other facial region to obtain the corresponding calibration image feature of the other facial region, the image encoding layer of the face reconstruction model is used to encode the reference face image of the other facial region to obtain the corresponding second image feature of the other facial region, the image encoding layer of the face reconstruction model is used to encode the reference face image of the mouth region to obtain the corresponding first image feature of the mouth region, and the voice encoding layer of the face reconstruction model is used to encode the reference voice to obtain the target voice feature.

[0190] Step 712: Based on the face prior feature of the other facial region, the calibration image feature and the second image feature are discretized to obtain the discretized calibration image feature and the discretized second image feature, and based on the face prior feature of the mouth region, the first image feature is discretized to obtain the discretized first image feature.

[0191] Step 714: The calibration image feature is key point encoded to obtain the calibration key point space feature, and the second image feature is key point encoded to obtain the second key point space feature.

[0192] Step 716: According to the feature difference between the calibration key point space feature and the second key point space feature, the feature deviation is determined.

[0193] Step 718: According to the feature deviation, the first image feature is calibrated to obtain the calibrated first image feature.

[0194] Step 720: The face prior feature of the mouth region is used to calculate the attention of the calibrated first image feature to obtain the attention image feature.

[0195] Step 722: Using the target speech features as constraints, normalize the attention image features to obtain the first target image features corresponding to the mouth region.

[0196] Step 724: Merge the features of the first target image and the second image to obtain merged image features.

[0197] Step 726: Based on the merged image features, use the decoding layer of the facial reconstruction model to generate video frames driven by the target animated character.

[0198] Step 728: Send the video frames driven by the target animated character to the front-end rendering.

[0199] In the embodiments of this specification, facial region features extracted from sample animated characters in the sample image set are obtained using a pre-trained virtual reconstruction model. This yields facial prior features containing texture features. Based on the target speech features and the aforementioned facial prior features containing texture features, feature transformation is performed on the first image features of the mouth region that follow the speech changes. This results in the first target image features that not only correspond to the speech features but also contain texture features. Finally, based on the first target image features and the second image features, a complete video frame driven by the target animated character is generated, corresponding to the reference speech and containing texture features. The video frame driven by the target animated character has high fidelity and high definition, improving the user experience. Furthermore, the adaptive processing of face alignment ensures texture matching and enhances versatility.

[0200] It should be noted that the embodiments in this specification may involve the use of user data. In practical applications, user-specific personal data may be used in the scheme described herein within the scope permitted by applicable laws and regulations, provided that it complies with the applicable laws and regulations of the country (e.g., with the user's explicit consent, with the user being properly notified, etc.).

[0201] With the above Figure 1 Corresponding to the method embodiments, this specification also provides embodiments of a voice-based image driving device. Figure 8 A schematic diagram of a voice-based image driving device according to one embodiment of this specification is shown. Figure 8 As shown, the device includes:

[0202] The first acquisition module 802 is configured to acquire a reference facial image and reference speech of a virtual object;

[0203] The first encoding module 804 is configured to perform speech coding on the reference speech to obtain target speech features, and perform image coding on the reference face image to obtain first image features of a first region and second image features of a second region, wherein the first region is a region of the reference face image that changes with the speech, and the second region is a region of the reference face image other than the first region.

[0204] The first feature transformation module 806 is configured to perform feature transformation on the first image features based on face prior features and the target speech features to determine first target image features, wherein the face prior features include face texture features.

[0205] The first generation module 808 is configured to generate a driven target image according to the first target image features and the second image features.

[0206] Optionally, the apparatus further comprises:

[0207] The calibration module is configured to obtain a calibration face image of the virtual object, perform image coding on the calibration face image to obtain calibration image features of a target region, wherein the target region corresponds to the second region, and calibrate the first image features based on a feature deviation between the second image features and the calibration image features to obtain calibrated first image features.

[0208] Optionally, the calibration module is further configured to:

[0209] perform key point coding on the second image features to obtain second key point spatial features, perform key point coding on the calibration image features to obtain calibration key point spatial features, and calibrate the first image features based on a feature deviation between the second key point spatial features and the calibration key point spatial features to obtain calibrated first image features.

[0210] Optionally, the first feature transformation module 806 is further configured to:

[0211] perform attention calculation on the first image features using the face prior features to obtain first attention image features, and perform normalization processing on the first attention image features using the target speech features as a constraint condition to obtain the first target image features.

[0212] Optionally, the face prior features include first face prior features corresponding to the first region and second face prior features corresponding to the second region.

[0213] Correspondingly, the apparatus further comprises:

[0214] The discretization module is configured to perform discretization processing on the first image feature based on the first facial prior feature to obtain a discretized first image feature, and perform discretization processing on the second image feature based on the second facial prior feature to obtain a discretized second image feature.

[0215] Optionally, the first encoding module 804 is further configured to:

[0216] Optionally, the first encoding module 804 is further configured to:

[0217] Optionally, the first encoding module 804 is further configured to:

[0218] Optionally, the first encoding module 804 is further configured to:

[0219] Optionally, the first encoding module 804 is further configured to:

[0220] Optionally, the first encoding module 804 is further configured to:

[0221] Optionally, the apparatus further comprises:

[0222] Optionally, the apparatus further comprises:

[0223] Optionally, the apparatus further comprises:

[0224] The facial prior feature construction module is configured to acquire a sample image set, wherein the sample image set includes multiple sample image pairs, and any sample image pair includes a first region sample image and a second region sample image corresponding to the same virtual object; for any sample image pair, the encoding layer of the pre-trained facial reconstruction model is used to extract texture features from the first region sample image and the second region sample image in the sample image pair respectively, to obtain the texture features of the first region sample image and the texture features of the second region sample image; the texture features of each first region sample image and the texture features of each second region sample image are integrated to obtain the facial prior features.

[0225] Optionally, the device further includes:

[0226] The pre-training module is configured to acquire a pre-training set, wherein the pre-training set includes multiple pre-training pairs, and each pre-training pair includes a first training sample image of a first region and a second training sample image of a second region corresponding to the same sample virtual object; the face reconstruction model is supervised and trained based on the first training sample image and the second training sample image of each pre-training pair to obtain a trained face reconstruction model.

[0227] In the embodiments of this specification, since the facial prior features include facial texture features, based on the facial prior features and the target speech features, the first image features of the first region following the speech transformation are transformed so that the obtained first target image features not only correspond to the speech features but also include texture features. Finally, based on the second image features and the first target image features, a complete driven target image that corresponds to the reference speech and includes texture features is generated. The target image has the characteristics of high fidelity and high definition, which improves the user experience.

[0228] The above is an illustrative scheme of a voice-based image driving device according to this embodiment. It should be noted that the technical solution of this voice-based image driving device and the technical solution of the above-described voice-based image driving method belong to the same concept. For details not described in detail in the technical solution of the voice-based image driving device, please refer to the description of the technical solution of the above-described voice-based image driving method.

[0229] With the above Figure 2 Corresponding to the method embodiments, this specification also provides embodiments of a voice-based image driving device. Figure 9 A schematic diagram of another voice-based image driving device provided in one embodiment of this specification is shown. Figure 9 As shown, this device is applied to cloud-side equipment, and the device includes:

[0230] The first receiving module 902 is configured to receive an image driving request for a virtual object sent by an end-side device, wherein the image driving request carries a reference face image of the virtual object and a reference voice;

[0231] The second encoding module 904 is configured to perform voice encoding on the reference voice to obtain target voice features, and perform image encoding on the reference face image to obtain first image features of a first region and second image features of a second region, wherein the first region is a region of the reference face image that changes with the voice, and the second region is a region of the reference face image other than the first region;

[0232] The second feature transformation module 906 is configured to perform feature transformation on the first image features based on face prior features and the target voice features to determine first target image features, wherein the face prior features include face texture features;

[0233] The second generation module 908 is configured to generate a driven target image according to the first target image features and the second image features.

[0234] The first rendering module 910 is configured to send the target image to the end-side device for rendering.

[0235] In the embodiments of the present specification, since the face prior features include face texture features, the first image features of the first region that changes with the voice are transformed based on the face prior features and the target voice features, so that the obtained first target image features not only correspond to the voice features but also contain the texture features. Finally, a complete driven target image that corresponds to the reference voice and contains the texture features is generated according to the second image features and the first target image features. The target image has the characteristics of high fidelity and high definition, which improves the user experience. At the same time, the image driving is implemented on the cloud-side device with higher computing power, which improves the efficiency of image driving and reduces the computing power cost of the end-side device.

[0236] The above is a schematic scheme of another voice-based image driving apparatus according to an embodiment of the present specification. It should be noted that the technical scheme of the voice-based image driving apparatus belongs to the same concept as the technical scheme of the voice-based image driving method described above. The details of the technical scheme of the voice-based image driving apparatus that are not described in detail can be referred to the description of the technical scheme of the voice-based image driving method described above.

[0237] According to the method embodiments described above, Figure 3 the present specification also provides voice-based image driving apparatus embodiments, Figure 10 FIG. 8 shows a structural schematic diagram of another voice-based image driving apparatus according to an embodiment of the present specification. As shown in FIG. 8, Figure 10As shown, the device is applied to an augmented reality (AR) device, and the device comprises:

[0238] The second receiving module 1002 is configured to receive an image driving request for a virtual object, wherein the image driving request carries a reference facial image and a reference voice of the virtual object.

[0239] The third encoding module 1004 is configured to perform voice encoding on the reference voice to obtain target voice features, and perform image encoding on the reference facial image to obtain first image features of a first region and second image features of a second region, wherein the first region is a region of the reference facial image that changes with the voice, and the second region is a region of the reference facial image other than the first region.

[0240] The third feature transformation module 1006 is configured to perform feature transformation on the first image features based on facial prior features and the target voice features to determine first target image features, wherein the facial prior features include facial texture features.

[0241] The third generation module 1008 is configured to generate a driven target image according to the first target image features and the second image features.

[0242] The second rendering module 1010 is configured to render the target image.

[0243] In the embodiments of the present specification, since the facial prior features include facial texture features, the first image features of the first region that changes with the voice are transformed based on the facial prior features and the target voice features, so that the obtained first target image features not only correspond to the voice features but also contain the texture features. Finally, a complete driven target image that corresponds to the reference voice and contains the texture features is generated according to the second image features and the first target image features. The target image has the characteristics of high fidelity and high definition, improves the user experience, and at the same time, the image driving is implemented on the augmented reality (AR) device with a stronger rendering effect, improves the rendering effect of the driven target image, and further improves the user experience.

[0244] The above is a schematic scheme of the image driving device based on the voice according to the present embodiment. It should be noted that the technical scheme of the image driving device based on the voice belongs to the same concept as the technical scheme of the image driving method based on the voice described above. The details of the technical scheme of the image driving device based on the voice that are not described in detail can be referred to the description of the technical scheme of the image driving method based on the voice.

[0245] Corresponding to the method embodiment described above, Figure 4 The present specification also provides an image driving data processing device embodiment, Figure 11A structural diagram of an image-driven data processing apparatus is shown. As shown in Figure 11 The apparatus is applied to a cloud-side device, and the apparatus comprises:

[0246] The second acquisition module 1102 is configured to acquire a training sample set, wherein the training sample set comprises a plurality of training sample groups, any training sample group comprises a sample image of a sample virtual object, a label image of the sample virtual object, and a sample voice corresponding to the label image;

[0247] The model training module 1104 is configured to perform supervised training on the image-driven model according to the sample image, the sample voice, and the label image of each training sample group, by taking the face prior feature as a feature-transformed prior feature, to obtain a trained image-driven model, wherein the face prior feature comprises a face region feature of a virtual object extracted from a sample image set by using a pre-trained face reconstruction model, and the face region feature comprises a texture feature.

[0248] The sending module 1106 is configured to send model parameters of the trained image-driven model to an end-side device.

[0249] In the embodiment of the present specification, the face region feature of a virtual object extracted from a sample image set by using a pre-trained face reconstruction model is used to obtain a face prior feature comprising a texture feature, the face prior feature is taken as a feature-transformed prior feature, and supervised training is performed on the image-driven model according to the sample image, the sample voice, and the label image, to obtain a trained image-driven model, thereby improving the feature extraction capability of the image-driven model for the texture feature and the pertinence between the voice data and the image data, and making the subsequently generated target image have the characteristics of high fidelity and high definition. Meanwhile, the model training is implemented on a cloud-side device with higher computing power, thereby improving the efficiency of model training and reducing the computing power cost of the end-side device.

[0250] The above is a schematic scheme of the image-driven data processing apparatus according to the embodiment. It should be noted that the technical scheme of the image-driven data processing apparatus belongs to the same concept as the technical scheme of the image-driven data processing method described above, and the details of the technical scheme of the image-driven data processing apparatus that are not described in detail can be referred to the description of the technical scheme of the image-driven data processing method.

[0251] Figure 12 A structural block diagram of a computing device is shown according to an embodiment of the present specification. The components of the computing device 1200 include but are not limited to a memory 1210 and a processor 1220. The processor 1220 is connected with the memory 1210 through a bus 1230, and a database 1250 is used to save data.

[0252] The computing device 1200 also includes an access device 1240 that enables the computing device 1200 to communicate via one or more networks 1260. Examples of these networks include PSTN (Public Switched Telephone Network), LAN (Local Area Network), WAN (Wide Area Network), PAN (Personal Area Network), or combinations of these like the Internet. The access device 1240 can include one or more of any type of network interface (for example, NIC (Network Interface Card)) of either wired or wireless, such as an IEEE 802.12 WLAN (Wireless Local Area Networks) wireless interface, a Wi-MAX (World Interoperability for Microwave Access) interface, an Ethernet interface, a USB (Universal Serial Bus) interface, a cellular network interface, a Bluetooth interface, a NFC (Near Field Communication) interface, and so on.

[0253] In one embodiment of the present specification, the above-mentioned components of the computing device 1200 and other components not shown in the Figure 12 may be connected to each other, for example, through a bus. It should be understood that Figure 12 The computing device structure diagram shown is merely for the purpose of example, and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.

[0254] The computing device 1200 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (for example, a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, and so on), a mobile phone (for example, a smartphone), a wearable computing device (for example, a smart watch, smart glasses, and so on), or other types of mobile devices, or a stationary computing device such as a desktop computer or a PC (Personal Computer). The computing device 1200 can also be a mobile or stationary server.

[0255] The processor 1220 is configured to execute computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned voice-based image driving method or the image-driven data processing method.

[0256] The above is a schematic scheme of the computing device of the present embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the voice-based image driving method and the image-driven data processing method described above belong to the same concept, and the details of the technical scheme of the computing device that are not described in detail can be seen from the description of the technical scheme of the voice-based image driving method or the image-driven data processing method described above.

[0257] An embodiment of the present specification also provides a computer readable storage medium storing computer executable instructions, which, when executed by a processor, implement the steps of the voice-based image driving method or the image-driven data processing method described above.

[0258] The above is a schematic scheme of the computer readable storage medium of the present embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the voice-based image driving method and the image-driven data processing method described above belong to the same concept, and the details of the technical scheme of the storage medium that are not described in detail can be seen from the description of the technical scheme of the voice-based image driving method or the image-driven data processing method described above.

[0259] An embodiment of the present specification also provides a computer program, which, when executed in a computer, causes the computer to perform the steps of the voice-based image driving method or the image-driven data processing method described above.

[0260] The above is a schematic scheme of the computer program of the present embodiment. It should be noted that the technical scheme of the computer program and the technical scheme of the voice-based image driving method and the image-driven data processing method described above belong to the same concept, and the details of the technical scheme of the computer program that are not described in detail can be seen from the description of the technical scheme of the voice-based image driving method or the image-driven data processing method described above.

[0261] The above describes specific embodiments of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than the order in which they are recited and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.

[0262] The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, software distribution medium, etc.

[0263] It should be noted that, for the foregoing method embodiments, in order to facilitate description, they are all described as a combination of a series of actions, but those skilled in the art should know that the present application is not limited to the order of the actions described, because according to the present application, certain steps can be performed in other order or at the same time. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0264] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0265] The preferred embodiments of the present application disclosed above are only used to help explain the present application. The alternative embodiments do not describe all the details and limit the present application to the specific embodiments described. Obviously, according to the content of the present application, many modifications and changes can be made. The present application selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present application, so that those skilled in the art can well understand and use the present application. The present application is limited by the claims and their entire scope and equivalents.

Claims

1. A voice-based image driving method, comprising: obtaining a reference voice and a reference facial image of a virtual object; performing voice coding on the reference voice to obtain target voice features, and performing image coding on the reference facial image to obtain first image features of a first region and second image features of a second region, wherein the first region is a region of the reference facial image that changes with the voice, and the second region is a region of the reference facial image other than the first region; performing feature transformation on the first image features based on facial prior features and the target voice features to determine first target image features, wherein the facial prior features include pairs of facial texture features of the first region and the second region of a virtual object face in a sample image, and the performing feature transformation on the first image features based on the facial prior features and the target voice features to determine first target image features comprises: performing attention calculation on the first image features using the facial prior features to obtain attention image features, and performing feature optimization on the attention image features according to the target voice features to obtain the first target image features of the first region; generating a driven target image according to the first target image features and the second image features.

2. The method of claim 1, before the performing feature transformation on the first image features based on the facial prior features and the target voice features to determine first target image features, further comprising: obtaining a calibration facial image of the virtual object; performing image coding on the calibration facial image to obtain calibration image features of a target region, wherein the target region corresponds to the second region; performing calibration on the first image features based on feature deviations between the second image features and the calibration image features to obtain calibrated first image features.

3. The method of claim 2, the performing calibration on the first image features based on feature deviations between the second image features and the calibration image features to obtain calibrated first image features comprises: performing key point coding on the second image features to obtain second key point spatial features, and performing key point coding on the calibration image features to obtain calibration key point spatial features; performing calibration on the first image features based on feature deviations between the second key point spatial features and the calibration key point spatial features to obtain calibrated first image features.

4. The method of claim 1, the performing feature transformation on the first image features based on the facial prior features and the target voice features to determine first target image features comprises: performing attention calculation on the first image features using the facial prior features to obtain first attention image features; performing normalization processing on the first attention image features with the target voice features as a constraint to obtain the first target image features.

5. The method of claim 1, wherein, the facial prior features include first facial prior features corresponding to the first region and second facial prior features corresponding to the second region; Before the feature transformation of the first image feature based on the target voice feature and the face prior feature, the method further comprises: performing discretization processing on the first image feature based on the first face prior feature to obtain a discretized first image feature; performing discretization processing on the second image feature based on the second face prior feature to obtain a discretized second image feature.

6. The method of claim 1, wherein the encoding of the reference voice to obtain a target voice feature and the encoding of the reference face image to obtain a first image feature of a first region and a second image feature of a second region comprises: obtaining a pre-trained image driving model, wherein the image driving model comprises an image encoding layer, a voice encoding layer, a feature transformation layer, and a decoding layer; inputting the reference voice into the voice encoding layer to obtain a target voice feature, and inputting the reference face image into the image encoding layer to obtain a first image feature of a first region and a second image feature of a second region; the feature transformation of the first image feature based on the target voice feature and the face prior feature to determine a first target image feature comprises: inputting the first image feature, the target voice feature, and the face prior feature into the feature transformation layer to perform feature transformation of the first image feature based on the target voice feature and the face prior feature to determine a first target image feature; the generation of a driven target image based on the first target image feature and the second image feature comprises: performing feature merging of the first target image feature and the second image feature to obtain a merged image feature; inputting the merged image feature into the decoding layer to obtain a driven target image.

7. The method of claim 6, wherein before the obtaining of the pre-trained image driving model, the method further comprises: obtaining a training sample set, wherein the training sample set comprises a plurality of training sample groups, and any training sample group comprises a sample image of a sample virtual object, a label image of the sample virtual object, and a sample voice corresponding to the label image; performing supervised training of the image driving model based on the sample image, the sample voice, and the label image of each training sample group, with the face prior feature as a prior feature for feature transformation, to obtain a trained image driving model.

8. The method of any one of claims 1-7, wherein the face prior feature is obtained by pre-extracting a face region feature of a sample image using a face reconstruction model, and accordingly, the method further comprises: obtaining a sample image set, wherein the sample image set comprises a plurality of sample image pairs, and any sample image pair comprises a first region sample image and a second region sample image corresponding to a same virtual object; For any sample image pair, the encoding layer of the pre-trained face reconstruction model is used to extract texture features of the first region sample image and the second region sample image in the sample image pair respectively, to obtain the texture features of the first region sample image and the texture features of the second region sample image; Integrate the texture features of each first region sample image and the texture features of each second region sample image to obtain the face prior feature.

9. The method of claim 8, before the use of the encoding layer of the pre-trained face reconstruction model to extract texture features of the first region sample image and the second region sample image in the sample image pair, further comprising: Obtain a pre-training set, wherein the pre-training set comprises a plurality of pre-training pairs, any pre-training pair comprising a first training sample image of a first region and a second training sample image of a second region corresponding to the same sample virtual object; According to the first training sample image and the second training sample image of each pre-training pair, the face reconstruction model is supervised trained to obtain the trained face reconstruction model.

10. A voice-based image driving method applied to a cloud-side device, comprising: Receiving an image driving request for a virtual object sent by an end-side device, wherein the image driving request carries a reference face image and a reference voice of the virtual object; Performing voice coding on the reference voice to obtain target voice features, and performing image coding on the reference face image to obtain first image features of a first region and second image features of a second region, wherein the first region is a region of the reference face image that changes with the voice, and the second region is a region of the reference face image other than the first region; Based on the face prior feature and the target voice feature, performing feature transformation on the first image feature to determine the first target image feature, wherein the face prior feature comprises a plurality of pairs of face texture features of the first region and the second region of the virtual object face in the sample image, and the feature transformation on the first image feature based on the face prior feature and the target voice feature to determine the first target image feature comprises: performing attention calculation on the first image feature using the face prior feature to obtain an attention image feature, and performing feature optimization on the attention image feature according to the target voice feature to obtain the first target image feature of the first region; Generating a driven target image according to the first target image feature and the second image feature; Sending the target image to the end-side device for rendering.

11. A voice-based image driving method applied to an augmented reality (AR) device, comprising: Receiving an image driving request for a virtual object, wherein the image driving request carries a reference face image and a reference voice of the virtual object; encode the reference speech to obtain target speech features, and encode the reference face image to obtain first image features of a first region and second image features of a second region, wherein the first region is a region of the reference face image that changes with the speech, and the second region is a region of the reference face image other than the first region; perform feature transformation on the first image features based on the face prior features and the target speech features to determine first target image features, wherein the face prior features include pairs of face texture features of the first region and the second region of a virtual object face in a sample image, and the performing of the feature transformation on the first image features based on the face prior features and the target speech features to determine the first target image features includes: performing attention calculation on the first image features based on the face prior features to obtain attention image features, and performing feature optimization on the attention image features based on the target speech features to obtain the first target image features of the first region; generate a driven target image based on the first target image features and the second image features; render the target image.

12. An image driving data processing method applied to a cloud-side device, comprising: obtaining a training sample set, wherein the training sample set includes a plurality of training sample groups, and any training sample group includes a sample image of a sample virtual object, a label image of the sample virtual object, and sample speech corresponding to the label image; supervising training of an image driving model based on sample images, sample speech, and label images of each training sample group, with face prior features as prior features for feature transformation, to obtain a trained image driving model, and sending model parameters of the trained image driving model to an end-side device; wherein the face prior features include pairs of face texture features of the first region and the second region of a virtual object face in the sample image, are used for feature transformation on first image features to determine first target image features, and are used for generation of a driven target image based on the first target image features and second image features, and the determination of the first target image features includes: performing attention calculation on the first image features based on the face prior features to obtain attention image features, and performing feature optimization on the attention image features based on target speech features to obtain the first target image features of the first region; wherein the target speech features are obtained by speech encoding of reference speech, the first image features and the second image features are obtained by image encoding of a reference face image of a virtual object, the first region is a region of the reference face image that changes with the speech, and the second region is a region of the reference face image other than the first region.

13. A computing device, comprising: a memory and a processor; The memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions, and the computer executable instructions, when executed by the processor, implement the steps of the voice-based image driving method according to any one of claims 1 to 11 or the image driving data processing method according to claim 12. 14.A computer readable storage medium storing computer executable instructions, and the computer executable instructions, when executed by a processor, implement the steps of the voice-based image driving method according to any one of claims 1 to 11 or the image driving data processing method according to claim 12.

Citation Information

Patent Citations

  • Image generation and model training method and device, equipment and storage medium

    CN115690238A