Model construction method and apparatus, image generation method and apparatus, device, and medium

By obtaining voice and expression information, using data processing models to predict face features and updating the model, the problem of poor voice image conversion effect in the prior art is solved, and a more realistic and natural image generation is achieved.

WO2025108417A1PCT designated stage expired Publication Date: 2025-05-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/133799
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-22
Filing Date
2024-11-22
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively generate voice-based images, especially in application scenarios such as virtual live broadcasts and movie production, where voice-to-image conversion is poor.

Method used

By obtaining sample speech, expression description information and face supervision information, the data processing model is used to predict facial expressions and texture coefficients, and the model is updated through the discriminator to improve the generation effect.

Benefits of technology

The image generation effect is improved, and the generated face description image is more realistic and natural, which can better meet the rich and full expression needs in the speech image generation scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024133799_30052025_PF_FP_ABST
    Figure CN2024133799_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides a model construction method and apparatus, an image generation method and apparatus, a device, and a medium. The model construction method comprises: first obtaining sample voice, expression description information corresponding to the sample voice and facial supervision information corresponding to the sample voice, wherein the facial supervision information comprises facial expression supervision information and facial texture coefficient supervision information; inputting voice features of the sample voice and expression features of the expression description information into a data processing model to obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model; then using a discriminator corresponding to the data processing model to determine a discrimination result of the predicted facial expression; and finally on the basis of the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the discrimination result, updating the data processing model, such that the updated data processing model has better performance, which is conducive to improving the image generation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Model construction method, image generation method, device, equipment, medium

[0001] This application claims priority to Chinese Patent Application No. 202311568539.2 filed on November 22, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field

[0002] The present disclosure relates to a model construction method, an image generation method, an apparatus, a device, and a medium. Background Art

[0003] For some application scenarios, such as virtual live broadcast and film production, there may be a need to generate corresponding images based on speech. To facilitate understanding, the following examples are used to illustrate.

[0004] As an example, for a longer speech, the speech can be split into several speech segments, and images corresponding to each speech segment can be generated, so that the video corresponding to the speech can be obtained based on these images later, thereby realizing video generation based on speech. Summary of the Invention

[0005] The present disclosure provides a model construction method, image generation method, device, equipment, and medium, which are conducive to improving image generation effects.

[0006] In order to achieve the above objectives, the technical solutions provided by the present disclosure are as follows:

[0007] The present disclosure provides a model building method, the method comprising:

[0008] Acquire a sample voice, expression description information corresponding to the sample voice, and facial supervision information corresponding to the sample voice, wherein the facial supervision information includes facial expression supervision information and facial texture coefficient supervision information;

[0009] Inputting the speech features of the sample speech and the expression features of the expression description information into a data processing model to obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model;

[0010] Determining a discrimination result of the predicted facial expression using a discriminator corresponding to the data processing model;

[0011] The data processing model is updated according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information, and the judgment result of the predicted facial expression.

[0012] In one possible implementation, updating the data processing model based on the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information, and the discrimination result of the predicted facial expression includes:

[0013] determining a facial expression prediction loss based on difference representation data between the predicted facial expression and the facial expression supervision information;

[0014] determining a facial texture coefficient prediction loss based on difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information;

[0015] Determining accuracy representation data of the predicted facial expression based on the discrimination result of the predicted facial expression;

[0016] determining a model loss of the data processing model based on the facial expression prediction loss, the facial texture coefficient prediction loss, and the accuracy characterization data of the predicted facial expression;

[0017] The data processing model is updated according to the model loss.

[0018] In a possible implementation manner, the facial supervision information further includes facial expression coefficient supervision information;

[0019] The step of inputting the speech features of the sample speech and the expression features of the expression description information into a data processing model to obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model includes:

[0020] Inputting the speech features of the sample speech and the expression features of the expression description information into a data processing model, and obtaining the predicted facial expression, the first predicted facial expression coefficient, and the first predicted facial texture coefficient output by the data processing model;

[0021] The updating of the data processing model based on the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information, and the discrimination result of the predicted facial expression comprises:

[0022] The data processing model is updated based on the predicted facial expression, the first predicted facial expression coefficient, the first predicted facial texture coefficient, the facial expression supervision information, the facial expression coefficient supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression.

[0023] In a possible implementation manner, the expression description information corresponding to the sample voice includes the expression category corresponding to the sample voice and the expression intensity of the expression category;

[0024] The process of determining the facial expression features includes:

[0025] Encoding the expression description information to obtain an expression code, wherein the expression code is used to represent the intensity state of at least one preset expression category, wherein the at least one preset expression category includes the expression category corresponding to the sample voice;

[0026] The expression feature is determined according to the expression code.

[0027] In one possible implementation, determining the facial expression feature based on the facial expression code includes:

[0028] Performing feature embedding processing on the expression code to obtain the expression feature, wherein a size value of the expression feature in a first dimension is equal to a size value of the speech feature in the first dimension;

[0029] The step of inputting the speech features of the sample speech and the expression features of the expression description information into a data processing model to obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model includes:

[0030] splicing the speech feature of the sample speech and the expression feature of the expression description information in a second dimension to obtain a first splicing feature, wherein a dimension value of the first splicing feature in the second dimension is equal to a sum of a dimension value of the expression feature in the second dimension and a dimension value of the speech feature in the second dimension, and the second dimension is different from the first dimension;

[0031] The splicing feature is input into the data processing model to obtain the predicted facial expression and the first predicted facial texture coefficient output by the data processing model.

[0032] In one possible implementation, the expression coding includes intensity normalized representation data of at least one preset expression category; the intensity normalized representation data of the expression category corresponding to the sample voice is obtained by normalizing the expression intensity of the expression category corresponding to the sample voice.

[0033] In one possible implementation, the facial supervision information corresponding to the sample voice is determined based on a reference facial image corresponding to the sample voice; the reference facial image is extracted from a reference video corresponding to the sample voice; the reference video includes the sample voice; the number of frames corresponding to the reference facial image in the reference video is the same as the number of frames corresponding to the sample voice in the reference video.

[0034] In one possible implementation, updating the data processing model based on the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information, and the discrimination result of the predicted facial expression includes:

[0035] If a preset model update condition is met, updating the data processing model according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information, and the judgment result of the predicted facial expression;

[0036] The method further comprises:

[0037] If a preset discriminator update condition is met, the discriminator is updated based on the discriminant result of the predicted facial expression and the discriminant result of the facial expression supervisory information; the discriminant result of the facial expression supervisory information is determined using the discriminator.

[0038] The present disclosure provides an image generation method, the method comprising:

[0039] Obtaining the speech to be processed and the expression description information corresponding to the speech to be processed;

[0040] Inputting the speech features of the speech to be processed and the expression features of the expression description information into a pre-constructed data processing model to obtain a second predicted facial expression coefficient and a second predicted facial texture coefficient output by the data processing model; the data processing model is constructed using the model construction method provided by the present disclosure;

[0041] A facial description image corresponding to the speech to be processed is determined based on the second predicted facial expression coefficient and the second predicted facial texture coefficient.

[0042] In one possible implementation manner, determining the facial description image corresponding to the speech to be processed based on the second predicted facial expression coefficient and the second predicted facial texture coefficient includes:

[0043] constructing a full face texture corresponding to the speech to be processed based on the second predicted facial expression coefficient and the second predicted facial texture coefficient;

[0044] A facial description image corresponding to the speech to be processed is generated based on the full face texture.

[0045] In one possible implementation manner, before generating the facial description image corresponding to the speech to be processed based on the full face texture, the method further includes:

[0046] Obtaining a reference facial image corresponding to the speech to be processed;

[0047] removing facial description information from the reference facial image to obtain a removed image;

[0048] Generating a facial description image corresponding to the speech to be processed based on the full face texture includes:

[0049] Full face generation processing is performed based on the full face texture and the removed image to obtain a facial description image corresponding to the speech to be processed.

[0050] In a possible implementation manner, the speech to be processed is any speech segment in a long speech;

[0051] The method further comprises:

[0052] The speech segments in the long speech and the facial description images corresponding to the speech segments are combined to obtain a video corresponding to the long speech; for any speech segment, the number of frames corresponding to the facial description image corresponding to the speech segment in the video is the same as the number of frames corresponding to the speech segment in the video, and the number of frames corresponding to the speech segment in the video is determined based on the arrangement position of the speech segment in the long speech.

[0053] The present disclosure provides a model building device, comprising:

[0054] A first acquisition unit is configured to acquire a sample speech, expression description information corresponding to the sample speech, and facial supervision information corresponding to the sample speech, wherein the facial supervision information includes facial expression supervision information and facial texture coefficient supervision information;

[0055] a first processing unit, configured to input the speech features of the sample speech and the expression features of the expression description information into a data processing model, and obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model;

[0056] a first determining unit, configured to determine a discrimination result of the predicted facial expression using a discriminator corresponding to the data processing model;

[0057] A model updating unit is used to update the data processing model based on the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression.

[0058] The present disclosure provides an image generating device, comprising:

[0059] A second acquiring unit is used to acquire the speech to be processed and the expression description information corresponding to the speech to be processed;

[0060] a second processing unit, configured to input the speech features of the speech to be processed and the expression features of the expression description information into a pre-constructed data processing model, to obtain a second predicted facial expression coefficient and a second predicted facial texture coefficient output by the data processing model; the data processing model is constructed using the model construction method provided by the present disclosure;

[0061] The second determining unit is configured to determine a facial description image corresponding to the speech to be processed based on the second predicted facial expression coefficient and the second predicted facial texture coefficient.

[0062] The present disclosure provides an electronic device, the device comprising: a processor and a memory;

[0063] The memory is used to store instructions or computer programs;

[0064] The processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes the model building method or image generation method provided by the present disclosure.

[0065] The present disclosure provides a computer-readable medium having instructions or computer programs stored therein. When the instructions or computer programs are executed on a device, the device executes the model building method or image generation method provided by the present disclosure.

[0066] The present disclosure provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the model construction method or image generation method provided by the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments recorded in the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0068] FIG1 is a flow chart of a model building method provided by an embodiment of the present disclosure;

[0069] FIG2 is a schematic diagram of a video generation process provided by an embodiment of the present disclosure;

[0070] FIG3 is a flow chart of an image generation method provided by an embodiment of the present disclosure;

[0071] FIG4 is a schematic structural diagram of a model building device provided by an embodiment of the present disclosure;

[0072] FIG5 is a schematic structural diagram of an image generating device provided by an embodiment of the present disclosure; and

[0073] FIG6 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0074] In order to enable those skilled in the art to better understand the solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the embodiments described are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0075] To better understand the technical solutions provided by the present disclosure, the model building method provided by the present disclosure is described below with reference to some figures. As shown in Figure 1, the model building method provided by an embodiment of the present disclosure includes the following steps S101-S104. Figure 1 is a flow chart of a model building method provided by an embodiment of the present disclosure.

[0076] S101: Acquire a sample voice, expression description information corresponding to the sample voice, and facial supervision information corresponding to the sample voice, where the facial supervision information includes facial expression supervision information and facial texture coefficient supervision information.

[0077] Here, the sample speech refers to the speech used in the model training process. Furthermore, this disclosure does not limit the method for obtaining the sample speech. For example, in some application scenarios, for any long speech in a speech training dataset, since such a long speech typically carries a large amount of speech information, the long speech can be segmented into multiple segments, and each segment can be used as the sample speech. It should be noted that this disclosure does not limit the method for segmenting the long speech. For example, a long speech containing multiple frames of speech information can be segmented into multiple segments, such that each segment carries one frame of speech information. The different frames of speech information carried by the long speech refer to those obtained at different speech sampling times. Therefore, in one case, the sample speech can refer to a speech segment within a long speech that carries any frame of speech information. It should also be noted that the speech referred to in this disclosure can refer to speech existing in an existing speech library that a user has authorized for use.

[0078] In addition, for the sample voice above, the expression description information corresponding to the sample voice is used to describe the expression state corresponding to the sample voice; and the present disclosure does not limit the implementation method of the expression description information. For example, it can be implemented using any existing or future data information that can describe an expression, such as some character strings that can describe expressions. It can be seen that under one possible implementation method, the expression description information corresponding to the sample voice may include the expression category corresponding to the sample voice. Among them, the expression category is used to describe what kind of expression the sample voice corresponds to a certain object; and the expression category can be represented by a character string. It should be noted that the present disclosure does not limit the implementation method of the object. For example, the object can be a virtual image, such as a virtual human, virtual animal or other model. For another example, the object can be an object, such as a robot.

[0079] In fact, in order to better improve the effect of facial expression description, the present disclosure also provides a possible implementation method of the facial expression description information corresponding to the above sample voice. Under this implementation method, the facial expression description information corresponding to the sample voice can include the expression category corresponding to the sample voice and the expression intensity of the expression category, so that the facial expression description information can not only indicate the type of expression corresponding to the sample voice, but also indicate the expression intensity corresponding to the sample voice. This allows the facial expression description information to more accurately describe the facial expression state corresponding to the sample voice, which is conducive to improving the accuracy of facial expressions in images generated based on the facial expression description information, thereby improving the image generation effect. Among them, the expression intensity is used to describe the intensity state corresponding to the sample voice under the expression category. It can be seen that in some application scenarios, the facial expression description information corresponding to the sample voice can include M expression categories and the expression intensity of the M expression categories. Among them, the mth expression category is used to describe the mth expression corresponding to the sample voice, and the present disclosure is not limited to the mth expression category. For example, the mth expression category can be implemented using a character string. The expression intensity of the mth expression category is used to describe the intensity state of the sample speech under the mth expression. Furthermore, the present disclosure does not limit the implementation of the expression intensity of the mth expression category. For example, the expression intensity of the mth expression category can be implemented as "10." m is a positive integer, m≤M, and M is a positive integer.

[0080] In addition, the present disclosure does not limit the method of obtaining the expression description information shown in the above paragraph. For example, in some application scenarios, the expression description information corresponding to the above sample voice may refer to the expression label pre-annotated by relevant personnel for the sample voice.

[0081] In addition, in some application scenarios, in order to reduce the difficulty of obtaining training data without affecting the effect of expression description, the present disclosure also provides a possible implementation method of the process of obtaining the expression description information corresponding to the sample speech above. In this implementation method, the process of obtaining the expression description information corresponding to the sample speech can be: performing expression recognition processing on the reference facial image corresponding to the sample speech to obtain the expression description information corresponding to the sample speech, so that the expression description information includes the expression category corresponding to the sample speech and the expression intensity of the expression category. In this way, the expression description information corresponding to the sample speech can be automatically obtained, thereby effectively avoiding the defects caused by manually labeling the expression description information, thereby helping to reduce the difficulty of model training. Wherein, the reference facial image refers to the true value of the facial image set in advance for the sample speech, so that the reference facial image can describe the actual facial state corresponding to the sample speech; and the present disclosure does not limit the process of obtaining the reference facial image. For example, the process of obtaining the reference facial image can be: extracting the reference facial image corresponding to the sample speech from the reference video corresponding to the sample speech, so that the number of frames corresponding to the reference facial image in the reference video is the same as the number of frames corresponding to the sample speech in the reference video. The reference video refers to the video containing the sample speech, such that the sample speech represents the speech information carried by a certain video frame in the reference video. The frame number corresponding to the sample speech in the reference video represents the frame number of the video frame in the reference video, and the "frame number corresponding to the sample speech in the reference video" represents the frame in which the sample speech appears in the reference video. This "frame number corresponding to the sample speech in the reference video" represents the position of the sample speech among all the video speech in the reference video. The reference facial image represents the image information carried by a certain video frame in the reference video, such that the frame number corresponding to the reference facial image in the reference video represents the frame number of the video frame in the reference video, and the "frame number corresponding to the reference facial image in the reference video" represents the position of the reference facial image among all the video images in the reference video. It can be seen that the sample speech and the reference facial image corresponding to the sample speech are both from the same video frame in the reference video.

[0082] It should be noted that the present disclosure does not limit the implementation method of the expression recognition processing in the above paragraph. For example, the method of performing expression recognition processing on an image with the help of a pre-built expression recognition model can be implemented.

[0083] The facial supervision information corresponding to the sample voice above refers to the true value of the facial information set in advance for the sample voice, so that the facial supervision information can represent the actual facial state corresponding to the sample voice, so that the facial supervision information can represent the guidance information required when using the sample voice to train the model, so that the model update process can be better guided based on the facial supervision information; and in order to better guide the face generation performance of the model, the facial supervision information corresponding to the sample voice can at least include facial expression supervision information and facial texture coefficient supervision information, so that when the model is updated based on the facial supervision information, the model can be better guided to learn facial details such as facial expressions and facial texture coefficients, so that the model updated based on the facial supervision information has better facial image generation performance.

[0084] For the facial expression supervision information corresponding to the sample voice above, the facial expression supervision information refers to the true value of the facial expression set in advance for the sample voice, so that the facial expression supervision information can represent the actual expression state corresponding to the sample voice, so that the facial expression supervision information can guide the model to learn facial expressions when the model is trained using the sample voice; and the present disclosure does not limit the implementation method of the facial expression supervision information, for example, the facial expression supervision information may include at least one preset facial expression probability representation data, and the at least one preset facial expression refers to some pre-set facial expressions, and the present disclosure does not limit the at least one preset facial expression. In addition, the present disclosure does not limit the acquisition process of the facial expression supervision information, for example, in some application scenarios, the facial expression supervision information may refer to the facial expression label marked in advance by relevant personnel for the sample voice. For example, in some application scenarios, in order to reduce the difficulty of obtaining training data, the process of obtaining the facial expression supervision information can be: performing expression recognition processing on the reference facial image corresponding to the sample voice to obtain the facial expression supervision information corresponding to the sample voice, so that the facial expression supervision information can include at least one preset facial expression probability representation data, so that the facial expression supervision information can represent the expression state corresponding to the sample voice as accurately as possible.

[0085] It should be noted that the present disclosure does not limit the implementation method of the expression recognition processing in the above paragraph. For example, the expression recognition processing can be implemented using any existing or future method that can perform expression recognition processing on images, such as a method of performing expression recognition processing on images with the help of a pre-built expression recognition model.

[0086] For the facial texture coefficient supervision information corresponding to the sample speech above, the facial texture coefficient supervision information refers to the true value of the facial texture coefficient set in advance for the sample speech, so that the facial texture coefficient supervision information is used to represent the actual facial texture state corresponding to the sample speech, so that the facial texture supervision information can guide the model to learn the facial texture coefficient when using the sample speech to train the model; and the present disclosure does not limit the acquisition process of the facial texture coefficient supervision information. For example, in some application scenarios, the facial texture coefficient supervision information can refer to the facial texture coefficient label pre-annotated by relevant personnel for the sample speech. For another example, in some application scenarios, in order to reduce the difficulty of obtaining training data, the acquisition process of the facial texture coefficient supervision information can be: performing facial texture coefficient extraction processing on the reference facial image corresponding to the sample speech to obtain the facial texture coefficient supervision information corresponding to the sample speech, so that the facial expression supervision information can represent the facial texture state corresponding to the sample speech as accurately as possible.

[0087] It should be noted that the present disclosure is not limited to the implementation method of the facial texture coefficient extraction process in the above paragraph. For example, the method of performing facial texture coefficient extraction processing on an image with the help of a pre-built facial texture coefficient extraction model can be implemented.

[0088] It should also be noted that the present disclosure does not limit the construction process of the facial texture coefficient extraction model in the above paragraph. For example, it can be specifically as follows: for any object, such as a digital image, a facial texture library of the object is first constructed so that the facial texture library includes the facial texture of the object under various expressions and mouth shapes; then, principal component analysis (PCA) is used to extract the principal components of the facial texture library to obtain the principal component extraction results corresponding to the object; then, based on the principal component extraction results, a facial texture coefficient extraction model corresponding to the object is constructed so that the facial texture coefficient extraction model can perform principal component coefficient extraction processing on any image, such as any image used to describe the object, so that the principal component coefficients extracted by the facial texture coefficient extraction model can be used as the facial texture coefficients corresponding to the image. As can be seen, in one possible implementation, the facial texture coefficient supervision information corresponding to the sample speech described above can be implemented using the principal component coefficients of the reference facial image corresponding to the sample speech; and the principal component coefficients of the reference facial image are obtained by performing principal component coefficient extraction processing on the reference facial image using a facial texture coefficient extraction model that matches the reference facial image. The "facial texture coefficient extraction model that matches the reference facial image" refers to a facial texture coefficient extraction model pre-constructed for the object described by the reference facial image.

[0089] Based on the content of the previous paragraph, it can be seen that in one possible implementation method, the process of obtaining the facial texture coefficient supervision information corresponding to the above sample voice can be specifically as follows: after obtaining the reference facial image corresponding to the sample voice, first search for a facial texture coefficient extraction model that matches the reference facial image from some pre-built facial texture coefficient extraction models; then use this matching facial texture coefficient extraction model to perform principal component coefficient extraction processing on the reference facial image to obtain the facial texture coefficient supervision information corresponding to the sample voice, so that the facial texture coefficient supervision information can represent the facial texture state corresponding to the sample voice as accurately as possible.

[0090] Based on the relevant content of the facial supervision information corresponding to the sample voice above, it can be seen that in some application scenarios, such as scenarios where high model construction efficiency is required, the facial supervision information corresponding to the sample voice may include facial expression supervision information and facial texture coefficient supervision information, so that the facial supervision information can represent the two facial state true values ​​pre-set for the sample voice, so as to guide the model update based on these two supervision information when using the sample voice to train the model, thereby helping to improve the model update efficiency.

[0091] In fact, in some application scenarios, in order to better improve model performance, the present disclosure also provides a possible implementation of the facial supervision information corresponding to the above sample voice. In this implementation, the facial supervision information corresponding to the sample voice includes not only facial expression supervision information and facial texture coefficient supervision information, but also facial expression coefficient supervision information, so that the facial supervision information can represent the three facial state true values ​​pre-set for the sample voice, thereby enabling the facial supervision information to more accurately and comprehensively represent the facial details corresponding to the sample voice, thereby enabling the facial supervision information to better guide the model to learn facial details, which is conducive to improving model performance. Among them, the facial expression coefficient supervision information refers to the facial expression coefficient true value pre-set for the sample voice, so that the facial expression coefficient supervision information is used to represent the actual expression details corresponding to the sample voice, so that the facial expression coefficient supervision information can guide the model to learn facial expression coefficients when using the sample voice to train the model; and the present disclosure does not limit the implementation of the facial expression coefficient supervision information. For example, the facial expression coefficient supervision information can be implemented using any expression coefficient, such as the coefficient of ARKit's 52-dimensional expression feature. In addition, the present disclosure does not limit the method for obtaining the facial expression coefficient supervision information. For example, in some application scenarios, the facial expression coefficient supervision information may refer to the facial expression coefficient label pre-labeled by relevant personnel for the sample speech. For another example, in some application scenarios, in order to reduce the difficulty of obtaining training data, the process of obtaining the facial expression coefficient supervision information may be: performing expression coefficient extraction processing on the reference facial image corresponding to the sample speech to obtain the facial expression coefficient supervision information corresponding to the sample speech, so that the facial expression coefficient supervision information can represent the expression details corresponding to the sample speech as accurately as possible.

[0092] Based on the above discussion of facial supervision information corresponding to sample speech, it can be seen that in some application scenarios, the facial supervision information can be determined based on a reference facial image corresponding to the sample speech, so that the facial supervision information can represent the facial state presented by the reference facial image, which helps reduce the difficulty of obtaining training data. Specifically, because the reference facial image can describe the actual facial details corresponding to the sample speech, the facial supervision information determined based on the reference facial image can represent the facial details corresponding to the sample speech as accurately as possible, thereby enabling the facial supervision information to better guide the model to learn facial details, which is conducive to improving model performance.

[0093] Based on the relevant content of S101 above, it can be seen that in some application scenarios, if you want to use sample voice to participate in model training, you can first obtain the sample voice and the reference facial image corresponding to the sample voice; then use the reference facial image to determine the expression description information corresponding to the sample voice and the facial supervision information corresponding to the sample voice, so that you can subsequently complete the training and update process for the model based on the sample voice, the expression description information and the facial supervision information.

[0094] S102: Inputting the speech features of the sample speech and the expression features of the expression description information corresponding to the sample speech into a data processing model to obtain the predicted facial expression and the first predicted facial texture coefficient output by the data processing model.

[0095] The speech features of the sample speech are used to represent the speech information carried by the sample speech. Furthermore, the present disclosure does not limit the method for obtaining the speech features. For example, the method may specifically include: performing speech feature extraction processing on the sample speech using a pre-built speech feature extraction model to obtain the speech features of the sample speech. It should be noted that the present disclosure does not limit the implementation method of the speech feature extraction model. For example, the speech feature extraction model may be implemented using any model having a speech feature extraction function, such as the Automatic Speech Recognition (ASR) model shown in FIG2 .

[0096] In addition, for the facial expression description information corresponding to the sample voice above, the facial expression features of the facial expression description information are used to represent the facial expression state described by the facial expression description information; and the present disclosure does not limit the method of obtaining the facial expression features. For example, it can be implemented using any existing or future method that can extract facial expression features from the facial expression description information.

[0097] In fact, in order to better characterize the facial expression state, the present disclosure also provides a possible implementation method of the above-mentioned facial expression feature acquisition process. Under this implementation method, when the facial expression description information corresponding to the above-mentioned sample voice includes the facial expression category corresponding to the sample voice and the facial expression intensity of the facial expression category, the facial expression feature acquisition process of the facial expression description information may include the following steps 11-12.

[0098] Step 11: Encode the expression description information corresponding to the above sample speech to obtain an expression code corresponding to the sample speech, wherein the expression code is used to represent the intensity state of at least one preset expression category, and the at least one preset expression category includes the expression category corresponding to the sample speech.

[0099] The expression code corresponding to the sample speech is obtained by encoding the expression description information corresponding to the sample speech, so that the expression code can represent the expression state carried by the expression description information.

[0100] For the expression code corresponding to the sample speech above, the expression code can be used to characterize the intensity state of at least one preset expression category, so that the expression code can more accurately represent the expression state carried by the expression description information corresponding to the sample speech. The preset expression category refers to a coding dimension that is pre-set based on the actual application scenario and is required to be referenced when encoding the expression description information; and the at least one preset expression category at least includes the expression category corresponding to the sample speech.

[0101] In addition, in some application scenarios, in order to better improve the effect of expressing the expression state, the present disclosure also provides a possible implementation method of the expression coding corresponding to the above sample voice. Under this implementation method, the expression coding corresponding to the sample voice may include intensity normalized representation data of at least one preset expression category. Among them, the intensity normalized representation data of the kth preset expression category is used to represent the intensity state corresponding to the sample voice under the kth preset expression category; the intensity normalized representation data of the kth preset expression category belongs to the data interval [0, 1]; and the larger the intensity normalized representation data of the kth preset expression category, the higher the expression intensity of the kth preset expression category. k is a positive integer, k≤the number of categories in the at least one preset expression category.

[0102] In addition, the present disclosure does not limit the method of obtaining the expression code corresponding to the above sample voice. For example, when at least one preset expression category above includes the expression category corresponding to the sample voice and the expression intensity of the expression category, the expression code acquisition process can be: normalizing the expression intensity of the expression category corresponding to the sample voice to obtain the intensity normalized representation data of the expression category corresponding to the sample voice, so that the "intensity normalized representation data of the expression category corresponding to the sample voice" is positively correlated with the "expression intensity of the expression category corresponding to the sample voice", so that the "intensity normalized representation data of the expression category corresponding to the sample voice" can better represent the intensity state corresponding to the sample voice under the "expression category corresponding to the sample voice"; then use the intensity normalized representation data of the expression category corresponding to the sample voice to replace the preset data value of the corresponding expression category in the pre-constructed expression coding template to obtain the expression code corresponding to the sample voice, so that the expression code can better represent the expression state corresponding to the sample voice. Among them, the "normalized intensity representation data of the expression category corresponding to the sample voice" is used to represent the intensity state of the sample voice under the "expression category corresponding to the sample voice"; the "normalized intensity representation data of the expression category corresponding to the sample voice" belongs to the data interval [0, 1]; and the larger the "normalized intensity representation data of the expression category corresponding to the sample voice", the higher the expression intensity of the "expression category corresponding to the sample voice". The expression coding template refers to a template that is set in advance according to the actual application scenario and is required to be referenced when performing expression coding processing, so that the expression coding template can represent the preset data values ​​of each preset expression category, and the present disclosure does not limit the expression coding template. For example, when the at least one preset expression category includes K preset expression categories, the expression coding template can be implemented using a template such as [the preset data value of the first preset expression category, the preset data value of the second preset expression category, ..., the preset data value of the Kth preset expression category]. The preset data value of the kth preset expression category refers to the initial value of the intensity normalized representation data set in advance for the kth preset expression category, so that the preset data value of the kth preset expression category can represent the initial intensity state of the kth preset expression category. k is a positive integer, k≤K, and K is a positive integer.

[0103] Based on the content of the above paragraph, it can be seen that in one possible implementation method, when the expression description information corresponding to the above sample voice includes the expression category corresponding to the sample voice and the expression intensity of the expression category, the acquisition process of the expression code corresponding to the sample voice can be specifically: normalizing the expression intensity of the expression category corresponding to the sample voice to obtain the intensity normalized representation data of the expression category corresponding to the sample voice, so that when it is determined that the expression category corresponding to the sample voice matches the target category in at least one preset expression category involved in the above expression coding template, the preset data value of the target category in the expression coding template is directly replaced with the intensity normalized representation data of the expression category corresponding to the sample voice to obtain the expression code corresponding to the sample voice, so that the expression code corresponding to the sample voice includes the intensity normalized representation data of the expression category corresponding to the sample voice, and the preset data values ​​of other categories in the at least one preset expression category except the target category, and the preset data value of the target category does not exist in the expression code corresponding to the sample voice. Among them, the target category refers to a preset expression category that matches the expression category corresponding to the sample voice; and the present disclosure does not limit the method for determining the target category. For example, it can be specifically: first, calculate the similarity between the expression category corresponding to the sample voice and the kth preset expression category, k is a positive integer, k≤K, K is a positive integer; then, the preset expression category with the highest similarity is determined as the target category, so that the target category can represent the preset expression category that matches the expression category corresponding to the sample voice.

[0104] It should be noted that the present disclosure does not limit the implementation method of the step of "normalizing the expression intensity of the expression category corresponding to the sample voice to obtain intensity normalized representation data of the expression category corresponding to the sample voice" in the above two paragraphs. For example, it can be implemented by using any existing or future method that can normalize a piece of data, such as a method of normalizing according to a preset mapping rule. The preset mapping rule refers to a rule set in advance based on the actual application scenario for mapping any data to the data interval [0, 1]; and the present disclosure does not limit the preset mapping rule.

[0105] Based on the relevant content of the expression coding corresponding to the sample voice above, it can be seen that under a possible implementation method, the expression coding corresponding to the sample voice can be implemented using the coding format of [intensity normalized representation data of the first preset expression category, intensity normalized representation data of the second preset expression category,..., intensity normalized representation data of the Kth preset expression category] so that the expression coding can better represent the expression state corresponding to the sample voice.

[0106] Based on the relevant content of step 11 above, it can be known that for some application scenarios, after obtaining the expression description information corresponding to the above sample voice, the expression description information can be encoded to obtain the expression code corresponding to the sample voice, so that the expression code can better express the expression state corresponding to the sample voice.

[0107] Step 12: Based on the above expression coding, determine the expression features of the expression description information corresponding to the above sample speech.

[0108] It should be noted that the present disclosure does not limit the implementation of the above step 12. For example, in some application scenarios, the step 12 may specifically be: determining the above expression code as the expression feature of the expression description information corresponding to the above sample speech.

[0109] In addition, in order to better improve the effect of expressing the expression state, the present disclosure also provides a possible implementation method of the above step 12. Under this implementation method, the step 12 can be specifically: feature embedding the above expression code to obtain the expression features of the expression description information corresponding to the above sample speech, so that the expression features include the embedding feature vectors corresponding to each preset expression category, so that the expression features can better represent the expression state described by the expression code, thereby helping to improve the expression state representation effect. It should be noted that the present disclosure does not limit the implementation method of the feature embedding processing. For example, it can be implemented using any existing or future fully connected layer, such as the deep neural network (DNN) shown in Figure 2.

[0110] In addition, research has found that in some application scenarios, in order to better improve the image generation effect, it is necessary to align the above expression features and the above voice features in a certain feature dimension, such as width, so as to improve the subsequent feature splicing effect. Based on this, the present disclosure also provides a possible implementation method of the above step 12. Under this implementation method, the step 12 can be specifically as follows: feature embedding processing is performed on the above expression coding to obtain the expression features of the expression description information corresponding to the above sample voice, so that the size value of the expression feature in the first dimension is equal to the size value of the voice feature of the sample voice in the first dimension. In this way, the expression feature and the voice feature can be aligned in the first dimension, thereby effectively avoiding the defects caused by the large difference between the size value of the expression feature in the first dimension and the size value of the voice feature in the first dimension (for example, defects such as the inability to perform feature splicing), which is beneficial to improving the image generation effect. Among them, the first dimension refers to the feature dimension that the facial expression feature and the voice feature need to be aligned; and the present disclosure does not limit the implementation method of the first dimension. For example, when the facial expression feature includes at least one feature vector and the voice feature includes at least one feature vector, in order to better improve the feature splicing effect, it is necessary to ensure that the size of each feature vector in the facial expression feature is consistent with the size of each feature vector in the voice feature. In some application scenarios, the width of the facial expression feature can be used to represent the size of each feature vector in the facial expression feature, and the width of the voice feature can be used to represent the size of each feature vector in the voice feature. Based on the foregoing content, it can be seen that in these application scenarios, the first dimension can be implemented using the width dimension, which is conducive to aligning the size of each feature vector in the facial expression feature with the size of each feature vector in the voice feature.

[0111] Based on the content of the previous paragraph, it can be seen that under a possible implementation method, for the speech features of the above sample speech and the expression features of the expression description information corresponding to the sample speech, the two features can satisfy the following constraints: the number of feature vectors in the speech feature is different from the number of feature vectors in the expression feature, but the size of any feature vector in the speech feature is the same as the size of any feature vector in the expression feature. In this way, the speech feature and the expression feature can be aligned in the dimension of the feature vector size, thereby effectively avoiding the defects caused by the large difference between the size of the feature vector in the expression feature and the size of the feature vector in the speech feature, which is beneficial to improving the image generation effect.

[0112] Based on the relevant contents of steps 11 to 12 above, it can be known that in a possible implementation mode, when the expression description information corresponding to the sample voice above includes the expression category corresponding to the sample voice and the expression intensity of the expression category, the process of obtaining the expression features of the expression description information may include: first encoding according to the expression category corresponding to the sample voice and the expression intensity of the expression category to obtain the expression code corresponding to the sample voice, such as [intensity normalized representation data of the first preset expression category, intensity normalized representation data of the second preset expression category, ..., intensity normalized representation data of the Kth preset expression category]; then using DNN to process the expression code to obtain the expression features of the expression description information, so that the size value of the expression feature in the first dimension is equal to the size value of the voice feature of the sample voice in the first dimension, so that the expression feature and the voice feature can be spliced ​​together according to the first dimension, so that the splicing processing result can better represent the facial state corresponding to the sample voice, thereby making the facial related information determined based on the splicing processing result and the data processing model more accurate.

[0113] The data processing model is used to perform some processing on the input data of the data processing model, such as facial expression prediction processing, facial expression coefficient prediction processing, facial texture coefficient prediction processing, etc.

[0114] In addition, the present disclosure does not limit the model structure of the above data processing model. For example, it can be implemented using a recurrent neural network (RNN), a gated recurrent unit (GRU), or a long short-term memory network (LSTM).

[0115] In addition, the present disclosure does not limit the working principle of the above data processing model. For ease of understanding, some examples are used below to illustrate.

[0116] Example 1: In some application scenarios, such as scenarios with high model building efficiency, the working principle of the above data processing model can be specifically as follows: the voice features of the above sample speech and the expression features of the expression description information corresponding to the sample speech are input into the data processing model to obtain the predicted facial expression and the first predicted facial texture coefficient output by the data processing model. The predicted facial expression refers to the facial expression predicted by the data processing model for the sample speech, so that the predicted facial expression can express the predicted facial expression state corresponding to the sample speech, such as the state presented in terms of expression and lip shape; and the implementation method of the predicted facial expression is similar to the implementation method of the facial expression supervision information above. For the sake of brevity, it will not be repeated here. The first predicted facial texture coefficient refers to the facial texture coefficient predicted by the data processing model for the sample speech, so that the first predicted facial texture coefficient can represent the predicted facial texture state corresponding to the sample speech, such as other aspects other than facial contour, distribution of facial features, etc. in addition to expression and mouth shape; and the implementation method of the first predicted facial texture coefficient is similar to the implementation method of the facial texture coefficient supervision information above. For the sake of brevity, it will not be repeated here.

[0117] Example 2: In some application scenarios, such as scenarios with high model performance requirements, the working principle of the above data processing model can be specifically as follows: the speech features of the above sample speech and the expression features of the expression description information corresponding to the sample speech are input into the data processing model to obtain the predicted facial expression, the first predicted facial expression coefficient and the first predicted facial texture coefficient output by the data processing model. Among them, the first predicted facial expression coefficient refers to the facial expression coefficient predicted by the data processing model for the sample speech, so that the first predicted facial expression coefficient can represent the predicted expression details corresponding to the sample speech, such as the details presented in the expression and lip shape; and the implementation method of the first predicted facial expression coefficient is similar to the implementation method of the facial expression coefficient supervision information above. For the sake of brevity, it will not be repeated here.

[0118] It can be seen that, in one possible implementation, for the above data processing model, if the data processing model is used to process the voice features of the above sample speech and the expression features of the expression description information corresponding to the sample speech, then the output result of the data processing model can at least include predicted facial expressions, first predicted facial expression coefficients and first predicted facial texture coefficients, so that the output result can describe the facial state from more angles, so that the output result can better represent the predicted facial state corresponding to the sample speech, so that the performance of the data processing model in facial state prediction can be more accurately determined based on the output result, which is conducive to improving model performance.

[0119] Example 3: In some application scenarios, when the size value of the speech feature of the above sample speech in the first dimension is equal to the size value of the expression feature of the expression description information corresponding to the sample speech in the first dimension, the working principle of the above data processing model can be specifically as follows: after splicing the speech feature and the expression feature in the second dimension to obtain a first spliced ​​feature, the first spliced ​​feature is input into the data processing model to obtain the predicted facial expression and the first predicted facial texture coefficient output by the data processing model. The second dimension refers to the dimension required to refer to when splicing the speech feature and the expression feature, and the second dimension is different from the first dimension. For example, if the first dimension is width, the second dimension is height. The first splicing feature refers to the result obtained by splicing the voice feature and the expression feature, so that the size value of the first splicing feature in the second dimension is equal to the sum of the size value of the expression feature in the second dimension and the size value of the voice feature in the second dimension, so that the first splicing feature includes all feature vectors in the voice feature and all feature vectors in the expression feature, thereby enabling the first splicing feature to better represent the facial state corresponding to the sample voice, such as expression, mouth shape, contour, etc.

[0120] Example 4. In some application scenarios, when the size value of the speech feature of the above sample speech in the first dimension is equal to the size value of the expression feature of the expression description information corresponding to the sample speech in the first dimension, the working principle of the above data processing model can be specifically: after splicing the speech feature and the expression feature in the second dimension to obtain a first spliced ​​feature, the first spliced ​​feature is input into the data processing model to obtain the predicted facial expression, the first predicted facial expression coefficient and the first predicted facial texture coefficient output by the data processing model.

[0121] Based on the above two paragraphs, it can be seen that in one possible implementation, the above S102 can specifically be: splicing the voice features of the above sample speech and the expression features of the expression description information corresponding to the sample speech in the second dimension to obtain a first spliced ​​feature, so that the size value of the first spliced ​​feature in the second dimension is equal to the sum of the size value of the expression feature in the second dimension and the size value of the voice feature in the second dimension; then inputting the first spliced ​​feature into the data processing model to obtain the output result of the data processing model. The output result may include a predicted facial expression and a first predicted facial texture coefficient; or the output result may include a predicted facial expression, a first predicted facial expression coefficient, and a first predicted facial texture coefficient.

[0122] Furthermore, the present disclosure does not limit the implementation of the above data processing model. For example, in some application scenarios, the data processing model can be implemented using the regressor shown in FIG2 . The regressor can be used to map the input data of the regressor to facial expressions, facial expression coefficients, and facial texture coefficients.

[0123] Based on the relevant content of S102 above, it can be known that in one possible implementation mode, after obtaining the voice features of the above sample speech and the expression features of the expression description information corresponding to the sample speech, the voice features and the expression features can be input into the data processing model, so that the data processing model processes the voice features and the expression features to obtain the output result of the data processing model, so that the output result at least includes the predicted facial expression and the first predicted facial texture coefficient, so that the model performance of the data processing model can be evaluated based on the output result subsequently.

[0124] S103: Determine a discrimination result of the predicted facial expression using a discriminator corresponding to the data processing model.

[0125] Among them, the discriminator is used to assist the above data processing model in realizing adversarial learning during the model training process; and the present disclosure does not limit the implementation method of the discriminator. For example, it can be implemented using any existing or future discriminator required for realizing adversarial learning.

[0126] In addition, the present disclosure does not limit the working principle of the above discriminator. For example, the discriminator can be used to perform accuracy discrimination processing on the input data of the discriminator. It can be seen that in one possible implementation, for the discriminator corresponding to the above data processing model, the discriminator can be used to determine the accuracy of the generated expression, that is, the discriminator can be used to determine whether the input data of the discriminator is a predicted facial expression predicted by the data processing model or facial expression supervision information used to describe the real expression, so that the discriminator can be used to distinguish between the predicted facial expression and the facial expression supervision information.

[0127] In addition, for the predicted facial expression obtained by the above data processing model for the sample speech prediction, the discrimination result of the predicted facial expression is obtained by using the discriminator corresponding to the data processing model to perform expression accuracy discrimination processing on the predicted facial expression, so that the discrimination result can indicate whether the predicted facial expression belongs to the true value of the expression, and thus the discrimination result can indicate whether the discriminator can recognize that the predicted facial expression does not belong to the true value of the expression.

[0128] In addition, the present disclosure does not limit the implementation method of S103 above. For example, S103 can be specifically as follows: after obtaining the predicted facial expression determined by the data processing model for the sample speech, the predicted facial expression is input into the discriminator corresponding to the data processing model, so that the discriminator performs expression accuracy discrimination processing on the predicted facial expression, obtains and outputs the discrimination result of the predicted facial expression, so that the discrimination result can indicate whether the predicted facial expression belongs to the true value of the expression, so that the model performance of the data processing model can be evaluated based on the discrimination result subsequently.

[0129] Based on the relevant content of S103 above, it can be seen that in some application scenarios, after obtaining the predicted facial expression output by the data processing model for the sample speech, the discriminator corresponding to the data processing model can be used to determine the discrimination result of the predicted facial expression, so that the model performance of the data processing model can be evaluated based on the discrimination result.

[0130] S104: updating the data processing model according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information, and the judgment result of the predicted facial expression.

[0131] In the present disclosure, for the above sample speech, after obtaining the predicted facial expression corresponding to the sample speech, the first predicted facial texture coefficient corresponding to the sample speech, the facial expression supervision information corresponding to the sample speech, the facial texture coefficient supervision information corresponding to the sample speech, and the judgment result of the predicted facial expression corresponding to the sample speech, the data processing model can be updated based on these data to make the updated data processing model have better model performance, and continue to execute S101 and subsequent steps based on the updated data processing model to achieve the next round of training for the data processing model, and iterate in this way until the preset stop condition is reached. Wherein, the preset stop condition refers to the condition that needs to be achieved when the iterative training process of the data processing model ends.

[0132] It should be noted that the present disclosure does not limit the implementation method of the preset stop condition in the above paragraph. For example, the preset stop condition may specifically include that the model loss of the data processing model is lower than a preset loss threshold. For another example, the preset stop condition may specifically include that the rate of change of the model loss of the data processing model is lower than a preset rate of change threshold. For another example, the preset stop condition may specifically include that the number of updates of the data processing model reaches a preset number threshold. The model loss of the data processing model is used to characterize the model performance of the data processing model, and the present disclosure does not limit the implementation method of the model loss of the data processing model. For example, the model loss of the data processing model may be implemented using the model loss shown in steps 21 to 24 below or the model loss shown in steps 31 to 35 below.

[0133] In fact, in order to better improve the model updating effect, the present disclosure also provides a possible implementation of the above S104. In this implementation, the S104 may include the following steps 21 to 25.

[0134] Step 21: Determine the facial expression prediction loss based on the difference representation data between the predicted facial expression and the facial expression supervision information.

[0135] In the present disclosure, after the predicted facial expression is determined for the sample speech using the above data processing model, the difference representation data between the predicted facial expression and the facial expression supervision information corresponding to the sample speech is first calculated, so that the difference representation data can represent the difference between the predicted facial expression and the facial expression supervision information, thereby making the difference representation data able to represent the gap between the predicted facial expression and the true value of the expression corresponding to the sample speech; then, the facial expression prediction loss is determined based on the difference representation data, so that the facial expression prediction loss can represent the performance of the data processing model in facial expression prediction.

[0136] It should be noted that the present disclosure does not limit the determination process of the above-mentioned "difference representation data between the predicted facial expression and the facial expression supervision information". For example, it can be: according to a preset distance calculation formula, the distance between the predicted facial expression and the facial expression supervision information is calculated as the difference representation data between the predicted facial expression and the facial expression supervision information. Among them, the preset distance calculation formula refers to a formula for calculating the distance between two pieces of information that is set in advance based on the actual application scenario; and the present disclosure does not limit the preset distance calculation formula. For example, it can be implemented using Euclidean distance or cosine distance.

[0137] It should also be noted that the present disclosure does not limit the implementation method of the above step 21. For example, the step 21 can specifically be: determining the difference representation data between the predicted facial expression and the facial expression supervision information as the facial expression prediction loss. For another example, the step 21 can also be: substituting the difference representation data between the predicted facial expression and the facial expression supervision information into a pre-set expression prediction loss function to obtain the facial expression prediction loss. The expression prediction loss function refers to a pre-set function required for calculating the facial expression prediction loss.

[0138] Step 22: Determine a facial texture coefficient prediction loss based on the difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information.

[0139] In the present disclosure, after determining the first predicted facial texture coefficient for the sample speech using the above data processing model, the difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information corresponding to the sample speech is first calculated, so that the difference representation data can represent the difference between the first predicted facial texture coefficient and the facial texture coefficient supervision information, thereby enabling the difference representation data to represent the gap between the first predicted facial texture coefficient and the true value of the texture coefficient corresponding to the sample speech; then, the facial texture coefficient prediction loss is determined based on the difference representation data, so that the facial texture coefficient prediction loss can represent the performance of the data processing model in facial texture coefficient prediction.

[0140] It should be noted that the present disclosure does not limit the determination process of the above "difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information". For example, it can be: according to a preset distance calculation formula, the distance between the first predicted facial texture coefficient and the facial texture coefficient supervision information is calculated as the difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information.

[0141] It should also be noted that the present disclosure does not limit the implementation of the above step 22. For example, the specific implementation of step 22 may be: determining the difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information as the facial texture coefficient prediction loss. For another example, step 22 may also be: substituting the difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information into a pre-set texture coefficient prediction loss function to obtain the facial texture coefficient prediction loss. The texture coefficient prediction loss function refers to a pre-set function required for calculating the facial texture coefficient prediction loss.

[0142] Step 23: Determine the accuracy representation data of the predicted facial expression based on the discrimination result of the predicted facial expression.

[0143] The accuracy characterization data of the predicted facial expression can indicate the accuracy of the predicted facial expression obtained by using the data processing model to predict the sample speech.

[0144] In addition, the present disclosure does not limit the process of determining the accuracy characterization data of the above-mentioned predicted facial expression. For example, it can be implemented by any existing or future method that can determine the degree of accuracy based on the discrimination result. For another example, step 23 can specifically be: directly determining the discrimination result of the above-mentioned predicted facial expression as the accuracy characterization data of the predicted facial expression. For another example, step 23 can specifically be: substituting the discrimination result of the predicted facial expression into a pre-set expression accuracy evaluation function to obtain the accuracy characterization data of the predicted facial expression. Among them, the expression accuracy evaluation function refers to a pre-set function required for use in calculating expression accuracy.

[0145] Step 24: Determine the model loss of the data processing model based on the facial expression prediction loss, the facial texture coefficient prediction loss, and the accuracy representation data of the predicted facial expression.

[0146] It should be noted that the present disclosure does not limit the implementation of step 24 above. For example, it may specifically be: performing a weighted summation of the above-mentioned facial expression prediction loss, the above-mentioned facial texture coefficient prediction loss, and the above-mentioned data representing the accuracy of predicted facial expressions to obtain the model loss of the data processing model. For another example, step 24 may specifically be: combining the facial expression prediction loss, the facial texture coefficient prediction loss, and the above-mentioned data representing the accuracy of predicted facial expressions to obtain the model loss of the data processing model.

[0147] Step 25: Update the data processing model based on the model loss of the above data processing model.

[0148] In the present disclosure, after obtaining the model loss of the above data processing model, the data processing model can be updated based on the model loss so that the updated data processing model has better model performance, which is conducive to improving the model performance of the data processing model.

[0149] Based on the relevant content of steps 21 to 25 above, it can be seen that in some application scenarios, after obtaining the predicted facial expression corresponding to the above sample voice, the first predicted facial texture coefficient corresponding to the sample voice, the facial expression supervision information corresponding to the sample voice, the facial texture coefficient supervision information corresponding to the sample voice, and the judgment result of the predicted facial expression corresponding to the sample voice, the above data processing model can be updated based on the difference representation data between the predicted facial expression and the facial expression supervision information, the difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information, and the judgment result of the predicted facial expression, so that the updated data processing model has better model performance, which is conducive to improving the model performance of the data processing model.

[0150] In fact, in order to better improve the performance of the model, the present disclosure also provides a possible implementation of the above S104. Under this implementation, when the facial supervision information corresponding to the above sample voice includes facial expression supervision information, facial expression coefficient supervision information and facial texture coefficient supervision information, and the above data processing model determines the predicted facial expression, the first predicted facial expression coefficient and the first predicted facial texture coefficient for the sample voice, the S104 can be specifically: based on the predicted facial expression, the first predicted facial expression coefficient, the first predicted facial texture coefficient, the facial expression supervision information, the facial expression coefficient supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression, update the data processing model. Among them, because the information referenced when updating the data processing model is relatively rich, the updated data processing model has better model performance, which is conducive to better improving the model performance of the data processing model.

[0151] In addition, the present disclosure does not limit the implementation method of the step in the previous paragraph "updating the data processing model based on the predicted facial expression, the first predicted facial expression coefficient, the first predicted facial texture coefficient, the facial expression supervision information, the facial expression coefficient supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression". For ease of understanding, a possible implementation method of S104 is used as an example for explanation below.

[0152] As an example, in a possible implementation, the above S104 may specifically include the following steps 31 to 36.

[0153] Step 31: Determine the facial expression prediction loss based on the difference representation data between the predicted facial expression and the facial expression supervision information.

[0154] It should be noted that the relevant content of step 31 can be found in step 21 above, and for the sake of brevity, it will not be repeated here.

[0155] Step 32: Determine the facial expression coefficient prediction loss based on the difference representation data between the first predicted facial expression coefficient and the facial expression coefficient supervision information.

[0156] In the present disclosure, after determining the first predicted facial expression coefficient for the sample speech using the above data processing model, the difference representation data between the first predicted facial expression coefficient and the facial expression coefficient supervision information corresponding to the sample speech is first calculated, so that the difference representation data can represent the difference between the first predicted facial expression coefficient and the facial expression coefficient supervision information, thereby enabling the difference representation data to represent the gap between the first predicted facial expression coefficient and the true value of the expression coefficient corresponding to the sample speech; then, the facial expression coefficient prediction loss is determined based on the difference representation data, so that the facial expression coefficient prediction loss can represent the performance of the data processing model in facial expression coefficient prediction.

[0157] It should be noted that the present disclosure does not limit the determination process of the above-mentioned "difference representation data between the first predicted facial expression coefficient and the facial expression coefficient supervision information". For example, it can be: according to a preset distance calculation formula, calculate the distance between the first predicted facial expression coefficient and the facial expression coefficient supervision information as the difference representation data between the first predicted facial expression coefficient and the facial expression coefficient supervision information.

[0158] It should also be noted that the present disclosure does not limit the implementation method of the above step 32. For example, the step 32 can specifically be: determining the difference representation data between the first predicted facial expression coefficient and the facial expression coefficient supervision information as the facial expression coefficient prediction loss. For another example, the step 32 can also be: substituting the difference representation data between the first predicted facial expression coefficient and the facial expression coefficient supervision information into a pre-set expression coefficient prediction loss function to obtain the facial expression coefficient prediction loss. The expression coefficient prediction loss function refers to a pre-set function required for use in calculating the facial expression coefficient prediction loss.

[0159] Step 33: Determine the facial texture coefficient prediction loss based on the difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information.

[0160] It should be noted that the relevant content of step 33 can be found in step 22 above, and for the sake of brevity, it will not be repeated here.

[0161] Step 34: Determine the accuracy representation data of the predicted facial expression based on the judgment result of the predicted facial expression.

[0162] It should be noted that the relevant content of step 34 can be found in step 23 above, and for the sake of brevity, it will not be repeated here.

[0163] Step 35: Determine the model loss of the data processing model based on the facial expression prediction loss, the facial expression coefficient prediction loss, the facial texture coefficient prediction loss, and the accuracy representation data of the predicted facial expression.

[0164] It should be noted that the implementation of step 35 is similar to the implementation of step 24 above, and for the sake of brevity, it will not be repeated here.

[0165] Step 36: Update the data processing model based on the model loss of the above data processing model.

[0166] It should be noted that the relevant content of step 36 can be found in step 25 above, and for the sake of brevity, it will not be repeated here.

[0167] Based on the relevant contents of steps 31 to 36 above, it can be known that in some application scenarios, after obtaining the predicted facial expression corresponding to the above sample voice, the first predicted facial expression coefficient corresponding to the sample voice, the first predicted facial texture coefficient corresponding to the sample voice, the facial expression supervision information corresponding to the sample voice, the facial expression coefficient supervision information corresponding to the sample voice, the facial texture coefficient supervision information corresponding to the sample voice, and the judgment result of the predicted facial expression corresponding to the sample voice, the above data processing model can be updated based on the difference representation data between the predicted facial expression and the facial expression supervision information, the difference representation data between the first predicted facial expression coefficient and the facial expression coefficient supervision information, the difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information, and the judgment result of the predicted facial expression, so that the updated data processing model has better model performance, which is conducive to improving the model performance of the data processing model.

[0168] Based on the relevant contents of S101 to S104 above, it can be known that for the model construction method provided by the embodiment of the present disclosure, first obtain the sample voice, the expression description information corresponding to the sample voice, and the facial supervision information corresponding to the sample voice, and the facial supervision information includes facial expression supervision information and facial texture coefficient supervision information; then input the voice features of the sample voice and the expression features of the expression description information into the data processing model to obtain the predicted facial expression and the first predicted facial texture coefficient output by the data processing model; then, use the discriminator corresponding to the data processing model to determine the discrimination result of the predicted facial expression; finally, according to the predicted facial expression, the facial expression is judged. The data processing model is updated based on the facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the discrimination result, so that the updated data processing model has better performance, such as better facial expression prediction performance and facial texture coefficient prediction performance, etc., so that when the data processing model is used to perform facial image generation processing on a voice, a better facial description image can be generated, and then the face presented by the facial description image has a more realistic and natural expression, which is conducive to improving the image generation effect, thereby helping to better meet the rich and full expression requirements in the voice image generation scenario.

[0169] In addition, in order to better achieve adversarial learning between the data processing model and the discriminator, the data processing model and the discriminator can be trained iteratively to perform model construction. Based on this, the present disclosure also provides a possible implementation of the above model construction method, under which the model construction method can include the following steps 41-45.

[0170] Step 41: obtaining a sample voice, expression description information corresponding to the sample voice, and facial supervision information corresponding to the sample voice, wherein the facial supervision information includes facial expression supervision information and facial texture coefficient supervision information.

[0171] It should be noted that the relevant content of step 41 can be found in S101 above, and for the sake of brevity, it will not be repeated here.

[0172] Step 42: Input the speech features of the sample speech and the expression features of the expression description information corresponding to the sample speech into a data processing model to obtain the predicted facial expression and the first predicted facial texture coefficient output by the data processing model.

[0173] It should be noted that the relevant content of step 42 can be found in S102 above, and for the sake of brevity, it will not be repeated here.

[0174] Step 43: Using the discriminator corresponding to the data processing model, determine the discrimination result of the predicted facial expression and the discrimination result of the facial expression supervision information.

[0175] Among them, the discrimination result of the facial expression supervision information is obtained by using the discriminator corresponding to the above data processing model to perform expression accuracy discrimination processing on the facial expression supervision information, so that the discrimination result of the facial expression supervision information can indicate whether the facial expression supervision information belongs to the true expression value, thereby enabling the discrimination result of the facial expression supervision information to indicate whether the discriminator can recognize that the facial expression supervision information belongs to the true expression value.

[0176] In addition, the present disclosure does not limit the implementation of the above step 43. For example, the step 43 may specifically include the following steps 431 and 432.

[0177] Step 431: After obtaining the facial expression supervision information corresponding to the above sample speech, the facial expression supervision information is input into the discriminator corresponding to the above data processing model, so that the discriminator performs expression accuracy discrimination processing on the facial expression supervision information, obtains and outputs the discrimination result of the facial expression supervision information, so that the discrimination result can indicate whether the facial expression supervision information belongs to the true value of the expression, so that the performance of the discriminator can be evaluated based on the discrimination result later.

[0178] Step 432: After using the above data processing model to determine the predicted facial expression for the sample speech, the predicted facial expression is input into the discriminator corresponding to the data processing model, so that the discriminator performs expression accuracy discrimination processing on the predicted facial expression, obtains and outputs the discrimination result of the predicted facial expression, so that the discrimination result can indicate whether the predicted facial expression belongs to the true value of the expression, so that the performance of the discriminator can be evaluated based on the discrimination result, or the model performance of the data processing model can be evaluated based on the discrimination result.

[0179] Based on the relevant content of step 43 above, it can be seen that in some application scenarios, after obtaining the predicted facial expression and facial expression supervision information corresponding to the above sample voice, the discriminator corresponding to the data processing model can be used to determine the discrimination result of the predicted facial expression and the discrimination result of the facial expression supervision information, so that the performance of the target that needs to be updated in the current round of training, such as the data processing model or the discriminator, can be evaluated based on part or all of these two discrimination results.

[0180] Step 44: If the preset model update condition is met, the data processing model is updated based on the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information, and the judgment result of the predicted facial expression.

[0181] Among them, the preset model update condition refers to the condition that needs to be met when iteratively updating the data processing model, and the present disclosure does not limit the preset model update condition. For example, the preset model update condition can be: in the current round of training, the parameters in the data processing model are in an updateable state, but the parameters in the discriminator corresponding to the data processing model are in a fixed state.

[0182] In addition, the implementation of step 44 is similar to the implementation of S104 above, and for the sake of brevity, it will not be repeated here.

[0183] Based on the relevant content of step 44 above, it can be known that for the current round of training, if the parameters in the above data processing model are in an updateable state, but the parameters in the discriminator corresponding to the data processing model are in a fixed state, then it can be determined that the target to be updated in the current round of training is the data processing model. Therefore, the data processing model can be updated based on the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the discrimination result of the predicted facial expression, so that the updated data processing model has better model performance, which is conducive to improving the model performance of the data processing model.

[0184] Step 45: If the preset discriminator update condition is met, the discriminator is updated based on the discriminant result of the predicted facial expression and the discriminant result of the facial expression supervision information.

[0185] Among them, the preset discriminator update condition refers to the condition that needs to be met when iteratively updating the above discriminator, and the present disclosure does not limit the preset discriminator update condition. For example, the preset discriminator update condition can be: in the current round of training, the parameters in the data processing model are in a fixed state, but the parameters in the discriminator corresponding to the data processing model are in an updateable state.

[0186] In addition, the present disclosure does not limit the implementation of step 45. For example, it can be implemented using any existing or future discriminator updating method.

[0187] Based on the relevant content of step 45 above, it can be known that for the current round of training, if the parameters in the above data processing model are in a fixed state, but the parameters in the discriminator corresponding to the data processing model are in an updateable state, then it can be determined that the target to be updated in the current round of training is the discriminator. Therefore, the discriminator can be updated based on the discrimination results of the predicted facial expressions and the discrimination results of the facial expression supervision information, so that the updated discriminator has better discrimination performance, so that when the data processing model is subsequently trained with the help of the updated discriminator, the data processing model can better learn facial information, which is beneficial to improve the model performance of the data processing model.

[0188] Based on the relevant content of steps 41 to 45 above, it can be seen that in some application scenarios, adversarial learning between the data processing model and the discriminator can be achieved by iterative training of the data processing model and the discriminator, so that the data processing model finally obtained has better model performance, which is conducive to improving the model construction effect.

[0189] In addition, the present disclosure does not limit the execution subject of the model construction method provided in the embodiments of the present disclosure. For example, the model construction method provided in the embodiments of the present disclosure can be applied to a terminal device or a server. For another example, the model construction method provided in the embodiments of the present disclosure can also be implemented with the help of the data interaction process between the terminal device and the server. The terminal device can be a smart phone, a computer, a personal digital assistant (PDA), a tablet computer, etc. The server can be a stand-alone server, a cluster server, or a cloud server.

[0190] Furthermore, to better improve the image generation effect in speech-image scenarios, the speech-image generation task can be achieved with the help of a constructed data processing model. Based on this, the present disclosure also provides an image generation method, which is described below with reference to the accompanying drawings. As shown in FIG3 , the image generation method provided in an embodiment of the present disclosure may include the following steps S301-S303. FIG3 is a flowchart of an image generation method provided in an embodiment of the present disclosure.

[0191] S301: Acquire a speech to be processed and expression description information corresponding to the speech to be processed.

[0192] Among them, the speech to be processed refers to the speech required to be used in the speech image scenario; and the present disclosure does not limit the method of obtaining the speech to be processed. For example, in some application scenarios, if video generation processing is required for a long speech, the speech to be processed can refer to any speech segment in the long speech, so that the image generation process for the speech to be processed can be used to realize image generation processing for each speech segment in the long speech.

[0193] Furthermore, for the speech to be processed, the facial expression description information corresponding to the speech to be processed is used to describe the facial expression state corresponding to the speech to be processed; and the implementation of the facial expression description information corresponding to the speech to be processed is similar to the implementation of the facial expression description information corresponding to the sample speech above. Therefore, in one possible implementation, the facial expression description information corresponding to the speech to be processed may include the facial expression category corresponding to the speech to be processed and the facial expression intensity of the facial expression category.

[0194] In addition, the present disclosure does not limit the method for obtaining the expression description information corresponding to the speech to be processed above. For example, it can be provided manually by relevant personnel.

[0195] S302: Input the speech features of the speech to be processed and the expression features of the expression description information corresponding to the speech to be processed into a pre-constructed data processing model to obtain a second predicted facial expression coefficient and a second predicted facial texture coefficient output by the data processing model; the data processing model is constructed using any possible implementation method of the model construction method provided by the present disclosure.

[0196] The speech features of the speech to be processed are used to represent the speech information carried by the speech to be processed; and the method for obtaining the speech features of the speech to be processed is similar to the method for obtaining the speech features of the sample speech described above. Therefore, in one possible implementation, the process for obtaining the speech features of the speech to be processed can be: inputting the speech to be processed into an ASR model, and obtaining speech features output by the ASR model, so that the speech features can represent the speech information carried by the speech to be processed.

[0197] For the facial expression description information corresponding to the above-mentioned speech to be processed, the facial expression features of the facial expression description information are used to represent the facial expression state described by the facial expression description information; and the method for obtaining the facial expression features of the facial expression description information corresponding to the above-mentioned sample speech is similar to the method for obtaining the facial expression features of the facial expression description information corresponding to the above-mentioned sample speech. It can be seen that, under a possible implementation method, the process for obtaining the facial expression features of the facial expression description information corresponding to the above-mentioned speech to be processed can be: firstly encode the facial expression description information corresponding to the above-mentioned speech to be processed to obtain the facial expression code corresponding to the above-mentioned speech to be processed; and then use DNN to process the facial expression code to obtain the facial expression features of the facial expression description information corresponding to the above-mentioned speech to be processed, so that the size value of the facial expression feature in the first dimension is equal to the size value of the speech feature of the above-mentioned speech to be processed in the first dimension. Among them, the facial expression code corresponding to the above-mentioned speech to be processed is used to represent the facial expression state corresponding to the above-mentioned speech to be processed; and the implementation method of the facial expression code corresponding to the above-mentioned sample speech is similar to the implementation method of the facial expression code corresponding to the above-mentioned sample speech. For the sake of brevity, it will not be repeated here.

[0198] The data processing model refers to a model constructed using any possible implementation of the model construction method provided in the present disclosure; and the relevant content of the data processing model can be found above.

[0199] The second predicted facial expression coefficient refers to the facial expression coefficient predicted by the data processing model for the speech to be processed, so that the second predicted facial expression coefficient can represent the predicted facial expression state corresponding to the sample speech, such as the state presented in terms of expression and lip shape, etc.; and the implementation method of the second predicted facial expression coefficient is similar to the implementation method of the first predicted facial expression coefficient in S102 above. For the sake of brevity, it will not be repeated here.

[0200] The second predicted facial texture coefficient refers to the facial texture coefficient predicted by the data processing model for the speech to be processed, so that the second predicted facial texture coefficient can represent the predicted facial texture state corresponding to the sample speech, such as the state presented in terms of facial contour, distribution of facial features, etc.; and the implementation method of the second predicted facial texture coefficient is similar to the implementation method of the first predicted facial texture coefficient in S102 above. For the sake of brevity, it is not repeated here.

[0201] In addition, the present disclosure does not limit the implementation of the above S302. For example, the implementation of S302 is similar to the implementation of S102. For ease of understanding, it is not repeated here.

[0202] It can be seen that in one possible implementation, when the size value of the voice feature of the speech to be processed in the first dimension is equal to the size value of the expression feature of the expression description information corresponding to the speech to be processed in the first dimension, the above S302 can be specifically as follows: first, the voice feature of the speech to be processed and the expression feature of the expression description information corresponding to the speech to be processed are spliced ​​in the second dimension to obtain a second spliced ​​feature, so that the size value of the second spliced ​​feature in the second dimension is equal to the sum of the size value of the expression feature in the second dimension and the size value of the speech feature in the second dimension; then, the second spliced ​​feature is input into the data processing model to obtain the second predicted facial expression coefficient and the second predicted facial texture coefficient output by the data processing model. It should be noted that the implementation method of the second spliced ​​feature is similar to the implementation method of the first spliced ​​feature above, and for the sake of brevity, it will not be repeated here.

[0203] S303: Determine a facial description image corresponding to the speech to be processed according to the second predicted facial expression coefficient and the second predicted facial texture coefficient.

[0204] The facial description image corresponding to the speech to be processed is used to describe the facial state corresponding to the speech to be processed.

[0205] In addition, the present disclosure is not limited to the implementation of the above S303. For example, it can be implemented by adopting any existing or future method that can perform facial image generation processing based on facial expression coefficients and facial texture coefficients.

[0206] In addition, in order to better improve the image generation effect, the present disclosure also provides a possible implementation of the above S303. In this implementation, the S303 may include the following steps 51 and 52.

[0207] Step 51: constructing a full face texture corresponding to the speech to be processed based on the second predicted facial expression coefficient and the second predicted facial texture coefficient.

[0208] The full-face texture corresponding to the speech to be processed is used to describe the full-face state corresponding to the speech to be processed.

[0209] In addition, the present disclosure does not limit the implementation of step 51 above. For example, it can be implemented using any existing or future method that can perform full-face texture construction based on the second predicted facial expression coefficient and the second predicted facial texture coefficient, such as a method for performing full-face texture construction using a pre-built full-face texture construction model. The full-face texture construction model is used to perform full-face texture construction based on the second predicted facial expression coefficient and the second predicted facial texture coefficient to obtain and output a full-face texture. Furthermore, the present disclosure does not limit the implementation of the full-face texture construction model.

[0210] Step 52: Generate a facial description image corresponding to the speech to be processed based on the full facial texture.

[0211] It should be noted that the present disclosure does not limit the implementation of step 52. For example, it can be implemented by adopting any existing or future method that can generate and process a facial image based on full facial texture.

[0212] For example, in some application scenarios, step 52 may specifically include inputting the full-face texture into a pre-built full-face generation model, and obtaining a facial description image corresponding to the speech to be processed, output by the full-face generation model. This facial description image can better describe the facial state corresponding to the speech to be processed. The full-face generation model is used to perform full-face generation processing based on its input data; however, this disclosure is not limited to this full-face generation model.

[0213] Based on the relevant content of steps 51 to 52 above, it can be seen that after using the above data processing model to determine the second predicted facial expression coefficient and the second predicted facial texture coefficient for the speech to be processed, the second predicted facial expression coefficient and the second predicted facial texture coefficient can be used to synthesize the full face texture; then the full face texture is input into the full face generation model to obtain the facial description image output by the full face generation model, so that the facial description image can describe the facial state corresponding to the speech to be processed.

[0214] Based on the relevant contents of S301 to S303 above, it can be seen that for the image generation method provided by the present invention, the speech to be processed and the expression description information corresponding to the speech to be processed are first obtained; then the speech features of the speech to be processed and the expression features of the expression description information are input into a pre-constructed data processing model to obtain the second predicted facial expression coefficient and the second predicted facial texture coefficient output by the data processing model; then, based on the second predicted facial expression coefficient and the second predicted facial texture coefficient, the facial description image corresponding to the speech to be processed is determined, so that the facial description image can express the facial texture state described by the second predicted facial texture coefficient and the facial expression state described by the second predicted facial expression coefficient, so that the facial description image can better express the full face state corresponding to the speech to be processed, which is conducive to improving the image generation effect, thereby helping to better meet the rich and full expression requirements in the speech image generation scenario.

[0215] In addition, for some application scenarios, there may be the following requirements: performing speech image generation processing on a specified object so that the final generated image can describe the facial state corresponding to the speech.

[0216] In order to achieve the requirements shown in the previous paragraph, the present disclosure further provides a possible implementation of the image generation method. In this implementation, the image generation method includes the following steps 61 to 65.

[0217] Step 61: Acquire the speech to be processed, the facial expression description information corresponding to the speech to be processed, and the reference facial image corresponding to the speech to be processed.

[0218] The reference facial image corresponding to the speech to be processed is used to describe an object, such as a digital image, that is required as a basis when performing image generation processing on the speech to be processed.

[0219] Furthermore, the present disclosure does not limit the method for obtaining the reference facial image corresponding to the speech to be processed. For example, the reference facial image corresponding to the speech to be processed may be an image provided by a user via an image input device. The image input device is used to input images, and the present disclosure does not limit the implementation of the image input device.

[0220] In addition, for the relevant contents of the speech to be processed in step 61 and the expression description information corresponding to the speech to be processed, please refer to the relevant contents of S301 above.

[0221] Step 62: Input the speech features of the speech to be processed and the expression features of the expression description information corresponding to the speech to be processed into a pre-constructed data processing model to obtain a second predicted facial expression coefficient and a second predicted facial texture coefficient output by the data processing model; the data processing model is constructed using any possible implementation method of the model construction method provided by the present disclosure.

[0222] It should be noted that for the relevant content of step 62, please refer to the relevant content of S302 above.

[0223] Step 63: Constructing a full face texture corresponding to the speech to be processed based on the second predicted facial expression coefficient and the second predicted facial texture coefficient.

[0224] It should be noted that for the relevant content of step 63, please refer to the relevant content of step 51 above.

[0225] Furthermore, in some application scenarios, because different subjects have different facial texture characteristics, to improve the full-face texture construction effect, a corresponding full-face texture construction model can be pre-built for each subject, so that the full-face texture construction models corresponding to different subjects are different. Based on this, the present disclosure also provides a possible implementation of the process for determining the full-face texture corresponding to the speech to be processed described above. In this implementation, the process for determining the full-face texture corresponding to the speech to be processed can include the following steps 631-633.

[0226] Step 631: Perform object recognition processing on the reference facial image corresponding to the processed speech to obtain an object recognition result.

[0227] The object recognition result is used to represent the object described by the reference facial image corresponding to the processed speech. The present disclosure does not limit the implementation method of the object recognition result. For example, the object recognition result can be implemented using an object identifier.

[0228] In addition, the present disclosure does not limit the implementation method of the above object recognition processing. For example, it can be implemented using any existing or future method that can perform object recognition processing on an image, such as a method for performing object recognition processing with the help of a pre-built object recognition model.

[0229] Step 632: Search for a full-face texture construction model that matches the above object recognition result from at least one pre-constructed full-face texture construction model as a target model.

[0230] It should be noted that the present disclosure does not limit the implementation of step 632 above. For example, when the at least one pre-constructed full-face texture construction model has a pre-labeled object label, step 632 may specifically include determining whether the object label of the i-th full-face texture construction model is the same as the object recognition result. If they are the same, it can be determined that the i-th full-face texture construction model matches the object recognition result. Therefore, the i-th full-face texture construction model can be used as the target model to subsequently complete the full-face texture construction process related to the object recognition result using the target model. Wherein, i is a positive integer, i≤the number of models in the at least one full-face texture construction model.

[0231] Step 633: Input the second predicted facial expression coefficient and the second predicted facial texture coefficient into the target model to obtain the full facial texture corresponding to the speech to be processed output by the target model.

[0232] Based on the relevant contents of steps 631 to 633 above, it can be seen that in some application scenarios, in order to better improve the construction effect of the full-face texture, the reference facial image corresponding to the speech to be processed above can be used to search for a full-face texture construction model corresponding to the object described by the reference facial image from at least one pre-constructed full-face texture construction model; and then the full-face texture construction model obtained by the search is used to perform full-face texture construction processing on the second predicted facial expression coefficient and the second predicted facial texture coefficient corresponding to the speech to be processed to obtain the full-face texture corresponding to the speech to be processed, so that the full-face texture can better represent the full-face state corresponding to the object described by the reference facial image under the speech to be processed, which is conducive to improving the full-face texture construction effect.

[0233] Step 64: Remove the facial description information from the reference facial image corresponding to the speech to be processed to obtain a post-removal image.

[0234] The "post-face image" is obtained by removing facial description information from the reference facial image corresponding to the speech to be processed, so that the post-face image only carries image information other than the facial description information from the reference facial image. The facial description information refers to image information present in the reference facial image that describes the subject's face.

[0235] In addition, the present disclosure does not limit the implementation of the above step 64. For example, it can be implemented by using any existing or future method that can remove facial description information from an image.

[0236] Step 65: Perform full face generation processing based on the full face texture and the image after removal to obtain a facial description image corresponding to the speech to be processed.

[0237] It should be noted that the present disclosure does not limit the implementation method of the above step 65. For example, it can be specifically as follows: the above full-face texture and the above removed image are input into a pre-constructed full-face generator to obtain a facial description image corresponding to the speech to be processed output by the full-face generator, so that the facial state described by the facial description image is consistent with the facial state described by the full-face texture, and the other states described by the facial description image except the facial state are consistent with the corresponding states described by the removed image, so that speech image generation processing can be achieved under the premise of specifying the object. Among them, the full-face generator refers to a pre-constructed model for performing facial image generation processing based on the full-face texture and the removed image; and the present disclosure does not limit the implementation method of the full-face generator. For example, the full-face generator can be implemented using the full-face generator shown in Figure 2.

[0238] Based on the relevant content of steps 61 to 65 above, it can be seen that in some application scenarios, voice image generation processing can be performed based on the voice to be processed, the expression description information corresponding to the voice to be processed, and the reference facial image corresponding to the voice to be processed to obtain the facial description image corresponding to the voice to be processed. In this way, voice image generation processing can be achieved under the premise of specifying the object, which is conducive to improving the image generation effect.

[0239] In addition, the image generation method provided by the present disclosure can be applied not only to the field of speech image generation, but also to the field of image video generation. Based on this, the present disclosure also provides a video generation process, which can specifically include the following steps 71 to 75.

[0240] Step 71: After obtaining a long speech with video generation requirements, the long speech is divided into at least one speech segment.

[0241] The long speech refers to the speech that needs to be processed by speech and video generation; and the long speech carries at least one frame of speech information.

[0242] In addition, for the long speech above, in order to ensure the video generation effect, the long speech can be divided into at least one speech segment, so that the video generation process for the long speech can be completed by subsequently performing image processing on each speech segment.

[0243] In addition, the present disclosure does not limit the method for obtaining the at least one speech segment above. For example, if the long speech above carries Q frames of speech information, the process of obtaining the at least one speech segment may include: taking the speech segment carrying the first frame of speech information in the long speech as the first speech segment, so that the first speech segment carries the first frame of speech information; taking the speech segment carrying the second frame of speech information in the long speech as the second speech segment, so that the second speech segment carries the second frame of speech information; ... (and so on); taking the speech segment carrying the Qth frame of speech information in the long speech as the Qth speech segment, so that the Qth speech segment carries the Qth frame of speech information. Wherein, Q is a positive integer.

[0244] Step 72: Determine the speech to be processed from the above at least one speech segment.

[0245] It should be noted that the present disclosure does not limit the implementation of step 72 above. For example, step 72 may specifically include randomly selecting an untraversed speech segment from the at least one speech segment above as the speech to be processed. The untraversed speech segment refers to a speech segment that has not been processed for speech image generation.

[0246] For another example, if the current round of processing is the qth round of speech image generation, then step 72 above may specifically be: determining the qth speech segment as the speech to be processed, where q is a positive integer, q≤Q.

[0247] Step 73: Determine a facial description image corresponding to the speech to be processed based on the speech to be processed and the facial expression description information corresponding to the speech to be processed.

[0248] It should be noted that the present disclosure does not limit the implementation method of step 73 above. For example, step 73 can be implemented by adopting any processing process provided by the present disclosure that can determine the facial description image corresponding to the speech to be processed based on the speech to be processed and the expression description information corresponding to the speech to be processed, such as the processing process shown in S302-S303 above or the processing process shown in steps 62-65 above.

[0249] Step 74: Determine whether there is any speech segment in the at least one speech segment that has not been traversed. If so, return to execute the above step 72 and subsequent steps; if not, execute the following step 75.

[0250] Step 75: Combine the speech segments and the facial description images corresponding to the speech segments in the long speech to obtain a video corresponding to the long speech; for any speech segment, the number of frames corresponding to the facial description image corresponding to the speech segment in the video is the same as the number of frames corresponding to the speech segment in the video, and the number of frames corresponding to the speech segment in the video is determined based on the arrangement position of the speech segment in the long speech.

[0251] In the present disclosure, for the long speech above, after obtaining the facial description images corresponding to each speech segment in the long speech, the video corresponding to the long speech can be determined based on each speech segment and the facial description images corresponding to each speech segment, so that the video includes each speech segment and the facial description images corresponding to each speech segment, and the number of frames corresponding to the facial description image corresponding to the qth speech segment in the video is the same as the number of frames corresponding to the qth speech segment in the video, such as the qth frame video information carried by the video includes the image information carried by the facial description image corresponding to the qth speech segment and the speech information carried by the qth speech segment, etc., q is a positive integer, q≤Q, so that the speech video generation process can be realized. Among them, the number of frames corresponding to the qth speech segment in the video is used to describe the arrangement position of the qth speech segment in all the video speech of the video; and the number of frames corresponding to the qth speech segment in the video is determined by the arrangement position of the qth speech segment in the long speech.

[0252] It should be noted that the present disclosure does not limit the correlation between the number of frames corresponding to the qth voice segment in the video and the arrangement position of the qth voice segment in the long voice. For example, in some application scenarios, when the long voice records multiple frames of voice information in a forward order, the number of frames corresponding to the qth voice segment in the video and the arrangement position of the qth voice segment in the long voice are positively correlated. For another example, in some application scenarios, when the long voice records multiple frames of voice information in a reverse order, the number of frames corresponding to the qth voice segment in the video and the arrangement position of the qth voice segment in the long voice are negatively correlated.

[0253] Based on the relevant contents of steps 71 to 75 above, it can be known that in some application scenarios, for a long speech, facial description images corresponding to each speech segment in the long speech can be generated first; then these speech segments and the facial description images corresponding to these speech segments can be combined into a video corresponding to the long speech, so that the video not only carries the speech information described by the long speech, but also carries the image information described by the facial description images corresponding to these speech segments, thereby realizing speech and video generation processing. Among them, for any speech segment in the long speech, since the facial description image corresponding to the speech segment can represent the facial state corresponding to the speech segment, such as expression, mouth shape, and distribution of facial features, etc., the video including these facial description images can carry rich expressions, thereby effectively avoiding the defects caused by the fact that all frames of video images in the video generated based on speech use the same expression, thereby helping to improve the video generation effect.

[0254] Furthermore, the present disclosure does not limit the execution subject of the image generation method provided in the embodiments of the present disclosure. For example, the image generation method provided in the embodiments of the present disclosure can be applied to a terminal device or a server. For another example, the image generation method provided in the embodiments of the present disclosure can also be implemented through the data exchange process between the terminal device and the server.

[0255] Based on the model building method provided in the embodiments of the present disclosure, the embodiments of the present disclosure also provide a model building device, which will be explained and illustrated below in conjunction with Figure 4. Figure 4 is a schematic diagram of the structure of a model building device provided in the embodiments of the present disclosure. It should be noted that for the technical details of the model building device provided in the embodiments of the present disclosure, please refer to the relevant content of the model building method above.

[0256] As shown in FIG4 , the model building device 400 provided in an embodiment of the present disclosure includes:

[0257] A first acquisition unit 401 is configured to acquire a sample speech, expression description information corresponding to the sample speech, and facial supervision information corresponding to the sample speech, wherein the facial supervision information includes facial expression supervision information and facial texture coefficient supervision information;

[0258] A first processing unit 402 is configured to input the speech features of the sample speech and the expression features of the expression description information into a data processing model to obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model;

[0259] a first determining unit 403, configured to determine a discrimination result of the predicted facial expression using a discriminator corresponding to the data processing model;

[0260] The model updating unit 404 is configured to update the data processing model based on the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information, and the judgment result of the predicted facial expression.

[0261] In one possible implementation, the model updating unit 404 is specifically used to: determine the facial expression prediction loss based on the difference representation data between the predicted facial expression and the facial expression supervision information; determine the facial texture coefficient prediction loss based on the difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information; determine the accuracy representation data of the predicted facial expression based on the discrimination result of the predicted facial expression; determine the model loss of the data processing model based on the facial expression prediction loss, the facial texture coefficient prediction loss and the accuracy representation data of the predicted facial expression; and update the data processing model based on the model loss.

[0262] In a possible implementation manner, the facial supervision information further includes facial expression coefficient supervision information;

[0263] The first processing unit 402 is specifically configured to input the speech features of the sample speech and the expression features of the expression description information into a data processing model to obtain the predicted facial expression, the first predicted facial expression coefficient, and the first predicted facial texture coefficient output by the data processing model;

[0264] The model updating unit 404 is specifically used to update the data processing model based on the predicted facial expression, the first predicted facial expression coefficient, the first predicted facial texture coefficient, the facial expression supervision information, the facial expression coefficient supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression.

[0265] In a possible implementation manner, the expression description information corresponding to the sample voice includes the expression category corresponding to the sample voice and the expression intensity of the expression category;

[0266] The model building device 400 further includes:

[0267] an expression encoding unit, configured to encode the expression description information to obtain an expression code, wherein the expression code is used to represent an intensity state of at least one preset expression category, wherein the at least one preset expression category includes the expression category corresponding to the sample speech;

[0268] A feature determination unit is used to determine the expression feature according to the expression code.

[0269] In a possible implementation manner, the feature determination unit is specifically configured to: perform feature embedding processing on the expression code to obtain the expression feature, wherein a size value of the expression feature in a first dimension is equal to a size value of the speech feature in the first dimension;

[0270] The first processing unit 402 is specifically used to: splice the voice features of the sample voice and the expression features of the expression description information in a second dimension to obtain a first spliced ​​feature, where the size value of the first spliced ​​feature in the second dimension is equal to the sum of the size value of the expression feature in the second dimension and the size value of the voice feature in the second dimension, and the second dimension is different from the first dimension; input the first spliced ​​feature into the data processing model to obtain the predicted facial expression and the first predicted facial texture coefficient output by the data processing model.

[0271] In one possible implementation, the expression coding includes intensity normalized representation data of at least one preset expression category; the intensity normalized representation data of the expression category corresponding to the sample voice is obtained by normalizing the expression intensity of the expression category corresponding to the sample voice.

[0272] In one possible implementation, the facial supervision information corresponding to the sample voice is determined based on a reference facial image corresponding to the sample voice; the reference facial image is extracted from a reference video corresponding to the sample voice; the reference video includes the sample voice; the number of frames corresponding to the reference facial image in the reference video is the same as the number of frames corresponding to the sample voice in the reference video.

[0273] In one possible implementation, the model updating unit 404 is specifically configured to: if a preset model updating condition is met, update the data processing model based on the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information, and the judgment result of the predicted facial expression;

[0274] The model building device 400 further includes:

[0275] A discriminator updating unit is configured to update the discriminator based on a discriminant result of the predicted facial expression and a discriminant result of the facial expression supervisory information if a preset discriminator updating condition is met; the discriminant result of the facial expression supervisory information is determined using the discriminator.

[0276] Based on the relevant contents of the above-mentioned model building device 400, it can be known that for the model building device 400 provided by the embodiment of the present disclosure, the sample speech, the expression description information corresponding to the sample speech, and the facial supervision information corresponding to the sample speech are first obtained, and the facial supervision information includes facial expression supervision information and facial texture coefficient supervision information; then the speech features of the sample speech and the expression features of the expression description information are input into the data processing model to obtain the predicted facial expression and the first predicted facial texture coefficient output by the data processing model; then, the discriminator corresponding to the data processing model is used to determine the discrimination result of the predicted facial expression; finally, according to the predicted facial The facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the discrimination result are used to update the data processing model so that the updated data processing model has better performance, such as better facial expression prediction performance and facial texture coefficient prediction performance, etc., so that when the data processing model is used to perform facial image generation processing on a voice, a better facial description image can be generated, and then the face presented by the facial description image has a more real and natural expression, which is conducive to improving the image generation effect, thereby helping to better meet the rich and full expression requirements in the voice image generation scenario.

[0277] Based on the image generation method provided in the embodiments of the present disclosure, the embodiments of the present disclosure also provide an image generation device, which will be explained and illustrated below in conjunction with Figure 5. Figure 5 is a schematic structural diagram of the image generation device provided in the embodiments of the present disclosure. It should be noted that for the technical details of the image generation device provided in the embodiments of the present disclosure, please refer to the relevant content of the image generation method above.

[0278] As shown in FIG5 , an image generating apparatus 500 provided by an embodiment of the present disclosure includes:

[0279] The second acquisition unit 501 is used to acquire the speech to be processed and the expression description information corresponding to the speech to be processed;

[0280] a second processing unit 502 configured to input the speech features of the speech to be processed and the expression features of the expression description information into a pre-constructed data processing model to obtain a second predicted facial expression coefficient and a second predicted facial texture coefficient output by the data processing model; wherein the data processing model is constructed using the model construction method according to any one of claims 1 to 8;

[0281] The second determining unit 503 is configured to determine a facial description image corresponding to the speech to be processed according to the second predicted facial expression coefficient and the second predicted facial texture coefficient.

[0282] In a possible implementation manner, the second determining unit 503 includes:

[0283] a texture construction subunit, configured to construct a full-face texture corresponding to the speech to be processed based on the second predicted facial expression coefficient and the second predicted facial texture coefficient;

[0284] The image generation subunit is used to generate a facial description image corresponding to the speech to be processed based on the full face texture.

[0285] In a possible implementation manner, the image generating device 500 further includes:

[0286] A third acquiring unit, configured to acquire a reference facial image corresponding to the speech to be processed;

[0287] an information removal unit, configured to remove facial description information from the reference facial image to obtain a post-removal image;

[0288] The image generation subunit is specifically used to perform full face generation processing based on the full face texture and the removed image to obtain a facial description image corresponding to the speech to be processed.

[0289] In a possible implementation manner, the speech to be processed is any speech segment in a long speech;

[0290] The image generating device 500 further includes:

[0291] The video construction unit is used to combine the speech segments in the long speech and the facial description images corresponding to the speech segments to obtain a video corresponding to the long speech; for any speech segment, the number of frames corresponding to the facial description image corresponding to the speech segment in the video is the same as the number of frames corresponding to the speech segment in the video, and the number of frames corresponding to the speech segment in the video is determined based on the arrangement position of the speech segment in the long speech.

[0292] Based on the relevant content of the above-mentioned image generating device 500, it can be known that for the image generating device 500 provided by the embodiment of the present disclosure, the speech to be processed and the expression description information corresponding to the speech to be processed are first obtained; then the speech features of the speech to be processed and the expression features of the expression description information are input into a pre-constructed data processing model to obtain the second predicted facial expression coefficient and the second predicted facial texture coefficient output by the data processing model; then, based on the second predicted facial expression coefficient and the second predicted facial texture coefficient, the facial description image corresponding to the speech to be processed is determined, so that the facial description image can express the facial texture state described by the second predicted facial texture coefficient and the facial expression state described by the second predicted facial expression coefficient, so that the facial description image can better express the full facial state corresponding to the speech to be processed, which is conducive to improving the image generation effect, thereby helping to better meet the rich and full expression requirements in the speech image generation scenario.

[0293] In addition, an embodiment of the present disclosure also provides an electronic device, which includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any implementation of the model building method provided by the embodiment of the present disclosure, or executes any implementation of the image generation method provided by the embodiment of the present disclosure.

[0294] Referring to FIG6 , a schematic diagram of the structure of an electronic device 600 suitable for implementing embodiments of the present disclosure is shown. Terminal devices in embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG6 is merely an example and should not limit the functionality or scope of use of embodiments of the present disclosure.

[0295] As shown in Figure 6, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. Various programs and data required for the operation of electronic device 600 are also stored in RAM 603. Processing device 601, ROM 602, and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to bus 604.

[0296] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 6 shows the electronic device 600 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.

[0297] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0298] The electronic device provided by the embodiment of the present disclosure and the method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0299] The embodiments of the present disclosure also provide a computer-readable medium, in which instructions or computer programs are stored. When the instructions or computer programs are executed on a device, the device executes any implementation of the model building method provided by the embodiments of the present disclosure, or executes any implementation of the image generation method provided by the embodiments of the present disclosure.

[0300] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0301] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0302] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0303] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device can perform the method.

[0304] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0305] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0306] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit / module does not, in some cases, limit the unit itself.

[0307] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0308] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0309] It should be noted that the various embodiments of this disclosure are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the descriptions of the systems or devices disclosed in the embodiments for similarities and differences between them. Since the systems or devices disclosed in the embodiments correspond to the methods disclosed in the embodiments, their descriptions are relatively simple, and reference can be made to the descriptions of the methods for any related details.

[0310] It should be understood that in the present disclosure, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0311] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0312] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0313] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present disclosure. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments shown herein, but is intended to be construed in the widest manner consistent with the principles and novel features disclosed herein.

Claims

1. A model building method, comprising: Acquire a sample voice, expression description information corresponding to the sample voice, and facial supervision information corresponding to the sample voice, wherein the facial supervision information includes facial expression supervision information and facial texture coefficient supervision information; Inputting the speech features of the sample speech and the expression features of the expression description information into a data processing model to obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model; Determining a discrimination result of the predicted facial expression using a discriminator corresponding to the data processing model; The data processing model is updated according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression.

2. The method according to claim 1, wherein: The updating of the data processing model according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression comprises: Determining a facial expression prediction loss based on the difference representation data between the predicted facial expression and the facial expression supervision information; Determining a facial texture coefficient prediction loss according to the difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information; Determining accuracy characterization data of the predicted facial expression based on the discrimination result of the predicted facial expression; Determining a model loss of the data processing model according to the facial expression prediction loss, the facial texture coefficient prediction loss, and the accuracy characterization data of the predicted facial expression; The data processing model is updated according to the model loss.

3. The method according to claim 1, wherein: The facial supervision information also includes facial expression coefficient supervision information; The step of inputting the speech feature of the sample speech and the expression feature of the expression description information into a data processing model to obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model comprises: Inputting the speech features of the sample speech and the expression features of the expression description information into a data processing model to obtain the predicted facial expression, the first predicted facial expression coefficient and the first predicted facial texture coefficient output by the data processing model; The updating of the data processing model according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression comprises: The data processing model is updated according to the predicted facial expression, the first predicted facial expression coefficient, the first predicted facial texture coefficient, the facial expression supervision information, the facial expression coefficient supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression.

4. The method according to claim 1, wherein: The expression description information corresponding to the sample speech includes the expression category corresponding to the sample speech and the expression intensity of the expression category; The process of determining the facial expression features includes: Encoding the expression description information to obtain an expression code, wherein the expression code is used to represent the intensity state of at least one preset expression category, and the at least one preset expression category includes the expression category corresponding to the sample voice; The expression feature is determined according to the expression code.

5. The method according to claim 4, wherein: The step of determining the expression feature according to the expression code comprises: Performing feature embedding processing on the expression code to obtain the expression feature, wherein the size value of the expression feature in the first dimension is equal to the size value of the speech feature in the first dimension; The step of inputting the speech feature of the sample speech and the expression feature of the expression description information into a data processing model to obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model comprises: splicing the speech feature of the sample speech and the expression feature of the expression description information in a second dimension to obtain a first splicing feature, wherein the dimension value of the first splicing feature in the second dimension is equal to the sum of the dimension value of the expression feature in the second dimension and the dimension value of the speech feature in the second dimension, and the second dimension is different from the first dimension; The first splicing feature is input into the data processing model to obtain the predicted facial expression and the first predicted facial texture coefficient output by the data processing model.

6. The method according to claim 4 or 5, wherein: The expression code includes intensity normalized representation data of at least one preset expression category; The intensity normalized representation data of the expression category corresponding to the sample speech is obtained by normalizing the expression intensity of the expression category corresponding to the sample speech.

7. The method according to claim 1, wherein: The facial supervision information corresponding to the sample speech is determined based on a reference facial image corresponding to the sample speech; The reference facial image is extracted from a reference video corresponding to the sample speech; the reference video includes the sample speech; the number of frames corresponding to the reference facial image in the reference video is the same as the number of frames corresponding to the sample speech in the reference video.

8. The method according to claim 1, wherein: The updating of the data processing model according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression comprises: If the preset model update condition is met, the data processing model is updated according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression; The method further comprises: If the preset discriminator update condition is met, the discriminator is updated according to the discriminant result of the predicted facial expression and the discriminant result of the facial expression supervisory information; the discriminant result of the facial expression supervisory information is determined by using the discriminator.

9. A method for generating an image, comprising: Obtaining the speech to be processed and the expression description information corresponding to the speech to be processed; Inputting the speech features of the speech to be processed and the expression features of the expression description information into a pre-constructed data processing model to obtain a second predicted facial expression coefficient and a second predicted facial texture coefficient output by the data processing model; the data processing model is constructed using the model construction method described in any one of claims 1 to 8; A facial description image corresponding to the speech to be processed is determined according to the second predicted facial expression coefficient and the second predicted facial texture coefficient.

10. The method according to claim 9, wherein: The step of determining the facial description image corresponding to the speech to be processed according to the second predicted facial expression coefficient and the second predicted facial texture coefficient comprises: constructing a full face texture corresponding to the speech to be processed according to the second predicted facial expression coefficient and the second predicted facial texture coefficient; A facial description image corresponding to the speech to be processed is generated based on the full face texture.

11. The method according to claim 10, wherein: Before generating the facial description image corresponding to the speech to be processed based on the full face texture, the method further includes: Acquire a reference facial image corresponding to the speech to be processed; Removing facial description information from the reference facial image to obtain a removed image; The step of generating a facial description image corresponding to the speech to be processed based on the full face texture includes: Full face generation processing is performed based on the full face texture and the removed image to obtain a facial description image corresponding to the speech to be processed.

12. The method according to any one of claims 9 to 11, wherein: The speech to be processed is any speech segment in a long speech; The method further comprises: The speech segments in the long speech and the facial description images corresponding to the speech segments are combined to obtain a video corresponding to the long speech; for any speech segment, the number of frames corresponding to the facial description image corresponding to the speech segment in the video is the same as the number of frames corresponding to the speech segment in the video, and the number of frames corresponding to the speech segment in the video is determined based on the arrangement position of the speech segment in the long speech.

13. A model building device, comprising: A first acquisition unit is configured to acquire a sample voice, expression description information corresponding to the sample voice, and facial supervision information corresponding to the sample voice, wherein the facial supervision information includes facial expression supervision information and facial texture coefficient supervision information; A first processing unit is configured to input the speech feature of the sample speech and the expression feature of the expression description information into a data processing model to obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model; a first determining unit configured to determine a discrimination result of the predicted facial expression using a discriminator corresponding to the data processing model; The model updating unit is configured to update the data processing model according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression.

14. An image generating device, comprising: A second acquisition unit is configured to acquire the speech to be processed and the expression description information corresponding to the speech to be processed; A second processing unit is configured to input the speech features of the speech to be processed and the expression features of the expression description information into a pre-constructed data processing model to obtain a second predicted facial expression coefficient and a second predicted facial texture coefficient output by the data processing model; wherein the data processing model is constructed using the model construction method according to any one of claims 1 to 8; The second determining unit is configured to determine a facial description image corresponding to the speech to be processed according to the second predicted facial expression coefficient and the second predicted facial texture coefficient.

15. An electronic device comprising a processor and a memory, wherein: The memory is configured to store instructions or computer programs; The processor is configured to execute the instructions or the computer program in the memory so that the electronic device executes the model building method described in any one of claims 1 to 8 or the image generating method described in any one of claims 9 to 12.

16. A computer readable medium storing instructions or a computer program, wherein: When the instructions or the computer program are executed on a device, the device is enabled to execute the model building method according to any one of claims 1 to 8 or the image generating method according to any one of claims 9 to 12.

Citation Information

Patent Citations

  • Virtual image generation method and system

    CN108171789A

  • Emotion intelligent recognition method and device, electronic equipment and storage medium

    CN111523389A

  • Face driving model training and face mouth shape animation generation method

    CN112396182A

  • Virtual image generation method and related equipment thereof

    CN114332318A

  • Expression recognition model training method and device, electronic equipment and storage medium

    CN115273199A