Model construction method and device, image generation method and device, equipment and medium
By obtaining voice, expression description and face supervision information, using data processing models to predict and update, and generating real and natural facial description images, the problem of poor voice image generation effect in the prior art is solved.
Patent Information
- Application Number
- CN202311568539.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-23
AI Technical Summary
It is difficult for the prior art to effectively generate real and natural images corresponding to voice, especially in application scenarios such as virtual live broadcasts and movie production.
By obtaining sample speech, expression description information and face supervision information, the data processing model is used to predict facial expressions and texture coefficients, and the model is updated based on the discriminant results to generate a more realistic and natural face description image.
It improves the image generation effect and can better meet the rich and full expression needs in the voice image generation scene.
Smart Images

Figure CN120032406A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a model building method, image generation method, device, equipment, and medium. Background Art
[0002] For some application scenarios, such as virtual live broadcast, movie production, etc., these scenarios may have the following requirements: generating corresponding images based on voice. For ease of understanding, the following is an example.
[0003] As an example, for a longer speech, the speech can be split into several speech segments, and images corresponding to each speech segment are generated, so that the video corresponding to the speech can be obtained based on these images later, thereby realizing video generation based on speech. Summary of the invention
[0004] The present application provides a model building method, image generation method, device, equipment, and medium, which are conducive to improving image generation effects.
[0005] In order to achieve the above objectives, the technical solutions provided by this application are as follows:
[0006] The present application provides a model building method, the method comprising:
[0007] Acquire a sample voice, expression description information corresponding to the sample voice, and facial supervision information corresponding to the sample voice, wherein the facial supervision information includes facial expression supervision information and facial texture coefficient supervision information;
[0008] Inputting the speech features of the sample speech and the expression features of the expression description information into a data processing model to obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model;
[0009] Determining a discrimination result of the predicted facial expression using a discriminator corresponding to the data processing model;
[0010] The data processing model is updated according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression.
[0011] In a possible implementation manner, updating the data processing model according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information, and the discrimination result of the predicted facial expression includes:
[0012] Determining a facial expression prediction loss based on the difference representation data between the predicted facial expression and the facial expression supervision information;
[0013] Determining a facial texture coefficient prediction loss according to the difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information;
[0014] Determining accuracy characterization data of the predicted facial expression based on the discrimination result of the predicted facial expression;
[0015] Determining a model loss of the data processing model according to the facial expression prediction loss, the facial texture coefficient prediction loss, and the accuracy characterization data of the predicted facial expression;
[0016] The data processing model is updated according to the model loss.
[0017] In a possible implementation manner, the facial supervision information further includes facial expression coefficient supervision information;
[0018] The step of inputting the speech feature of the sample speech and the expression feature of the expression description information into a data processing model to obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model comprises:
[0019] Inputting the speech features of the sample speech and the expression features of the expression description information into a data processing model to obtain the predicted facial expression, the first predicted facial expression coefficient and the first predicted facial texture coefficient output by the data processing model;
[0020] The updating of the data processing model according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression comprises:
[0021] The data processing model is updated according to the predicted facial expression, the first predicted facial expression coefficient, the first predicted facial texture coefficient, the facial expression supervision information, the facial expression coefficient supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression.
[0022] In a possible implementation manner, the expression description information corresponding to the sample voice includes the expression category corresponding to the sample voice and the expression intensity of the expression category;
[0023] The process of determining the facial expression features includes:
[0024] Encoding the expression description information to obtain an expression code, wherein the expression code is used to represent the intensity state of at least one preset expression category, and the at least one preset expression category includes the expression category corresponding to the sample voice;
[0025] The expression feature is determined according to the expression code.
[0026] In a possible implementation manner, determining the expression feature according to the expression code includes:
[0027] Performing feature embedding processing on the expression code to obtain the expression feature, wherein the size value of the expression feature in the first dimension is equal to the size value of the speech feature in the first dimension;
[0028] The step of inputting the speech feature of the sample speech and the expression feature of the expression description information into a data processing model to obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model comprises:
[0029] splicing the speech feature of the sample speech and the expression feature of the expression description information in a second dimension to obtain a first splicing feature, wherein the dimension value of the first splicing feature in the second dimension is equal to the sum of the dimension value of the expression feature in the second dimension and the dimension value of the speech feature in the second dimension, and the second dimension is different from the first dimension;
[0030] The splicing feature is input into the data processing model to obtain the predicted facial expression and the first predicted facial texture coefficient output by the data processing model.
[0031] In one possible implementation, the expression coding includes intensity normalized representation data of at least one preset expression category; the intensity normalized representation data of the expression category corresponding to the sample voice is obtained by normalizing the expression intensity of the expression category corresponding to the sample voice.
[0032] In one possible implementation, the facial supervision information corresponding to the sample speech is determined based on a reference facial image corresponding to the sample speech; the reference facial image is extracted from a reference video corresponding to the sample speech; the reference video includes the sample speech; the number of frames corresponding to the reference facial image in the reference video is the same as the number of frames corresponding to the sample speech in the reference video.
[0033] In a possible implementation manner, updating the data processing model according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information, and the discrimination result of the predicted facial expression includes:
[0034] If the preset model update condition is met, the data processing model is updated according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression;
[0035] The method further comprises:
[0036] If the preset discriminator update condition is met, the discriminator is updated according to the discriminant result of the predicted facial expression and the discriminant result of the facial expression supervisory information; the discriminant result of the facial expression supervisory information is determined by using the discriminator.
[0037] The present application provides an image generation method, the method comprising:
[0038] Acquire the speech to be processed and the expression description information corresponding to the speech to be processed;
[0039] Inputting the speech features of the speech to be processed and the expression features of the expression description information into a pre-constructed data processing model to obtain a second predicted facial expression coefficient and a second predicted facial texture coefficient output by the data processing model; the data processing model is constructed using the model construction method provided in the present application;
[0040] A facial description image corresponding to the speech to be processed is determined according to the second predicted facial expression coefficient and the second predicted facial texture coefficient.
[0041] In a possible implementation manner, determining the facial description image corresponding to the speech to be processed according to the second predicted facial expression coefficient and the second predicted facial texture coefficient includes:
[0042] constructing a full face texture corresponding to the speech to be processed according to the second predicted facial expression coefficient and the second predicted facial texture coefficient;
[0043] A facial description image corresponding to the speech to be processed is generated based on the full face texture.
[0044] In a possible implementation manner, before generating the facial description image corresponding to the speech to be processed based on the full face texture, the method further includes:
[0045] Acquire a reference facial image corresponding to the speech to be processed;
[0046] Removing facial description information from the reference facial image to obtain a removed image;
[0047] The step of generating a facial description image corresponding to the speech to be processed based on the full face texture includes:
[0048] Full face generation processing is performed based on the full face texture and the removed image to obtain a facial description image corresponding to the speech to be processed.
[0049] In a possible implementation manner, the speech to be processed is any speech segment in a long speech;
[0050] The method further comprises:
[0051] The speech segments in the long speech and the facial description images corresponding to the speech segments are combined to obtain a video corresponding to the long speech; for any speech segment, the number of frames corresponding to the facial description image corresponding to the speech segment in the video is the same as the number of frames corresponding to the speech segment in the video, and the number of frames corresponding to the speech segment in the video is determined based on the arrangement position of the speech segment in the long speech.
[0052] The present application provides a model building device, comprising:
[0053] A first acquisition unit, used to acquire a sample voice, expression description information corresponding to the sample voice, and facial supervision information corresponding to the sample voice, wherein the facial supervision information includes facial expression supervision information and facial texture coefficient supervision information;
[0054] A first processing unit, configured to input the speech feature of the sample speech and the expression feature of the expression description information into a data processing model, and obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model;
[0055] a first determining unit, configured to determine a discrimination result of the predicted facial expression using a discriminator corresponding to the data processing model;
[0056] A model updating unit is used to update the data processing model according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression.
[0057] The present application provides an image generation device, comprising:
[0058] A second acquisition unit, used to acquire the speech to be processed and the expression description information corresponding to the speech to be processed;
[0059] A second processing unit is used to input the speech features of the speech to be processed and the expression features of the expression description information into a pre-constructed data processing model to obtain a second predicted facial expression coefficient and a second predicted facial texture coefficient output by the data processing model; the data processing model is constructed using the model construction method provided in the present application;
[0060] The second determining unit is used to determine the facial description image corresponding to the speech to be processed according to the second predicted facial expression coefficient and the second predicted facial texture coefficient.
[0061] The present application provides an electronic device, the device comprising: a processor and a memory;
[0062] The memory is used to store instructions or computer programs;
[0063] The processor is used to execute the instructions or computer programs in the memory so that the electronic device executes the model building method or image generation method provided in the present application.
[0064] The present application provides a computer-readable medium, in which instructions or computer programs are stored. When the instructions or computer programs are executed on a device, the device executes the model building method or image generation method provided in the present application.
[0065] The present application provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for executing the model building method or image generation method provided by the present application.
[0066] Compared with the related art, this application has at least the following advantages:
[0067] In the technical solution provided by the present application, a sample voice, expression description information corresponding to the sample voice, and facial supervision information corresponding to the sample voice are first obtained, and the facial supervision information includes facial expression supervision information and facial texture coefficient supervision information; then, the voice features of the sample voice and the expression features of the expression description information are input into a data processing model to obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model; then, a discriminator corresponding to the data processing model is used to determine a discrimination result of the predicted facial expression; finally, based on the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information, and the discrimination result, the data processing model is updated so that the updated data processing model has better performance, such as better facial expression prediction performance and facial texture coefficient prediction performance, so that when the data processing model is used to perform facial image generation processing on a voice, a better facial description image can be generated, and then the face presented by the facial description image has a more real and natural expression, which is conducive to improving the image generation effect, thereby helping to better meet the rich and full expression requirements in the voice image generation scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0069] Figure 1 A flowchart of a model building method provided in an embodiment of the present application;
[0070] Figure 2 A schematic diagram of a video generation process provided in an embodiment of the present application;
[0071] Figure 3 A flowchart of an image generation method provided in an embodiment of the present application;
[0072] Figure 4 A schematic diagram of the structure of a model building device provided in an embodiment of the present application;
[0073] Figure 5 A schematic diagram of the structure of an image generating device provided in an embodiment of the present application;
[0074] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0075] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0076] In order to better understand the technical solution provided by this application, the model building method provided by this application is described below with reference to some drawings. Figure 1 As shown, the model building method provided in the embodiment of the present application includes the following S101-S104. Figure 1 A flowchart of a model building method provided in an embodiment of the present application.
[0077] S101: Acquire a sample voice, expression description information corresponding to the sample voice, and facial supervision information corresponding to the sample voice, wherein the facial supervision information includes facial expression supervision information and facial texture coefficient supervision information.
[0078] Among them, the sample speech refers to the speech used in the model training process; and the present application does not limit the method of obtaining the sample speech. For example, in some application scenarios, for any long speech in the speech training data set, because the long speech usually carries more speech information, the long speech can be divided into multiple speech segments, and each speech segment is used as a sample speech. It should be noted that the present application does not limit the segmentation method for the long speech. For example, it can be specifically: for a long speech carrying multiple frames of speech information, the long speech is divided into multiple speech segments so that each speech segment carries one frame of speech information. Among them, the different frames of speech information carried by the long speech refer to those obtained at different speech sampling times. It can be seen that in one case, the sample speech can refer to a speech segment existing in a long speech and carrying any frame of speech information. It should also be noted that the speech involved in the present application can refer to a speech existing in an existing speech library and authorized by the user.
[0079] In addition, for the sample voice above, the expression description information corresponding to the sample voice is used to describe the expression state corresponding to the sample voice; and the present application does not limit the implementation method of the expression description information. For example, it can be implemented using any existing or future data information that can describe an expression, such as some character strings that can describe expressions. It can be seen that in one possible implementation method, the expression description information corresponding to the sample voice may include the expression category corresponding to the sample voice. Among them, the expression category is used to describe what kind of expression the sample voice corresponds to a certain object; and the expression category can be represented by a character string. It should be noted that the present application does not limit the implementation method of the object. For example, the object can be a virtual image, such as a virtual human, virtual animal, or other model. For another example, the object can be an object, such as a robot.
[0080] In fact, in order to better improve the effect of expression description, the present application also provides a possible implementation method of the expression description information corresponding to the above sample voice. Under this implementation method, the expression description information corresponding to the sample voice can include the expression category corresponding to the sample voice and the expression intensity of the expression category, so that the expression description information can not only indicate the type of expression corresponding to the sample voice, but also indicate the expression intensity corresponding to the sample voice, so that the expression description information can more accurately describe the expression state corresponding to the sample voice, which is conducive to improving the accuracy of facial expressions in images generated based on the expression description information, thereby improving the image generation effect. Among them, the expression intensity is used to describe the intensity state corresponding to the sample voice under the expression category. It can be seen that in some application scenarios, the expression description information corresponding to the sample voice can include M expression categories and the expression intensity of the M expression categories. Among them, the mth expression category is used to describe the mth expression corresponding to the sample voice, and the present application does not limit the mth expression category. For example, the mth expression category can be implemented using a character string. The expression intensity of the mth expression category is used to describe the intensity state of the sample speech under the mth expression; and the present application does not limit the implementation method of the expression intensity of the mth expression category. For example, the expression intensity of the mth expression category can be implemented by the number "10". m is a positive integer, m≤M, and M is a positive integer.
[0081] In addition, the present application does not limit the method for obtaining the expression description information shown in the previous paragraph. For example, in some application scenarios, the expression description information corresponding to the above sample speech may refer to the expression label pre-marked by relevant personnel for the sample speech.
[0082] In addition, in some application scenarios, in order to reduce the difficulty of obtaining training data without affecting the effect of expression description, the present application also provides a possible implementation method of the process of obtaining the expression description information corresponding to the above sample voice. In this implementation method, the process of obtaining the expression description information corresponding to the sample voice can be: performing expression recognition processing on the reference facial image corresponding to the sample voice to obtain the expression description information corresponding to the sample voice, so that the expression description information includes the expression category corresponding to the sample voice and the expression intensity of the expression category, so that the expression description information corresponding to the sample voice can be automatically obtained, thereby effectively avoiding the defects caused by manual labeling of the expression description information, and thus helping to reduce the difficulty of model training. Wherein, the reference facial image refers to the true value of the facial image set in advance for the sample voice, so that the reference facial image can describe the actual facial state corresponding to the sample voice; and the present application does not limit the process of obtaining the reference facial image, for example, the process of obtaining the reference facial image can be: extracting the reference facial image corresponding to the sample voice from the reference video corresponding to the sample voice, so that the number of frames corresponding to the reference facial image in the reference video is the same as the number of frames corresponding to the sample voice in the reference video. Among them, the reference video refers to the video carrying the sample voice, so that the sample voice can represent the voice information carried by a certain video frame in the reference video, so that the frame number corresponding to the sample voice in the reference video is used to represent the frame number of this video frame in the reference video, and then the "frame number corresponding to the sample voice in the reference video" can represent which frame the sample voice appears in the reference video, so that the "frame number corresponding to the sample voice in the reference video" can represent the arrangement position of the sample voice in all video voices in the reference video. The reference facial image is used to represent the image information carried by a certain video frame in the reference video, so that the frame number corresponding to the reference facial image in the reference video is used to represent the frame number of the video frame in the reference video, so that the "frame number corresponding to the reference facial image in the reference video" can represent the arrangement position of the reference facial image in all video images in the reference video. It can be seen that the sample voice and the reference facial image corresponding to the sample voice are both from the same video frame in the reference video.
[0083] It should be noted that the present application does not limit the implementation method of the expression recognition processing in the above paragraph. For example, it can be implemented by a method of performing expression recognition processing on an image with the help of a pre-built expression recognition model.
[0084] The facial supervision information corresponding to the above sample voice refers to the true facial information preset for the sample voice in advance, so that the facial supervision information can represent the actual facial state corresponding to the sample voice, so that the facial supervision information can represent the guiding information required when training the model using the sample voice, so as to better guide the model update process based on the facial supervision information in the future; moreover, in order to better guide the facial generation performance of the model, the facial supervision information corresponding to the sample voice can at least include facial expression supervision information and facial texture coefficient supervision information, so that when updating the model based on the facial supervision information, it can better guide the model to learn facial details such as facial expressions and facial texture coefficients, so that the model updated based on the facial supervision information has better facial image generation performance.
[0085] For the facial expression supervision information corresponding to the above sample voice, the facial expression supervision information refers to the true facial expression preset for the sample voice in advance, so that the facial expression supervision information can represent the actual expression state corresponding to the sample voice, so that the facial expression supervision information can guide the model to learn facial expressions when training the model using the sample voice; moreover, the implementation manner of the facial expression supervision information is not limited in this application. For example, the facial expression supervision information can include the occurrence probability characterization data of at least one preset facial expression, and the at least one preset facial expression refers to some facial expressions preset in advance, and the at least one preset facial expression is not limited in this application. In addition, the acquisition process of the facial expression supervision information is not limited in this application. For example, in some application scenarios, the facial expression supervision information can refer to the facial expression label annotated by relevant personnel in advance for the sample voice. Another example is that in some application scenarios, in order to reduce the difficulty of obtaining training data, the acquisition process of the facial expression supervision information can be: performing expression recognition processing on the reference facial image corresponding to the sample voice to obtain the facial expression supervision information corresponding to the sample voice, so that the facial expression supervision information can include the occurrence probability characterization data of at least one preset facial expression, so that the facial expression supervision information can as accurately as possible represent the expression state corresponding to the sample voice.
[0086] It should be noted that the implementation manner of the expression recognition processing in the above paragraph is not limited in this application. For example, the expression recognition processing can be implemented by using any existing or future method capable of performing expression recognition processing on images, such as the method of performing expression recognition processing on images by means of a pre-constructed expression recognition model.
[0087] For the facial texture coefficient supervision information corresponding to the sample speech above, the facial texture coefficient supervision information refers to the true value of the facial texture coefficient set in advance for the sample speech, so that the facial texture coefficient supervision information is used to represent the actual facial texture state corresponding to the sample speech, so that the facial texture supervision information can guide the model to learn the facial texture coefficient when the model is trained using the sample speech; and the present application does not limit the acquisition process of the facial texture coefficient supervision information, for example, in some application scenarios, the facial texture coefficient supervision information can refer to the facial texture coefficient label marked by relevant personnel in advance for the sample speech. For another example, in some application scenarios, in order to reduce the difficulty of obtaining training data, the acquisition process of the facial texture coefficient supervision information can be: extracting the facial texture coefficient of the reference facial image corresponding to the sample speech, and obtaining the facial texture coefficient supervision information corresponding to the sample speech, so that the facial expression supervision information can represent the facial texture state corresponding to the sample speech as accurately as possible.
[0088] It should be noted that the present application does not limit the implementation method of the facial texture coefficient extraction process in the above paragraph. For example, the method of performing facial texture coefficient extraction processing on an image with the help of a pre-built facial texture coefficient extraction model can be implemented.
[0089] It should also be noted that the present application does not limit the construction process of the facial texture coefficient extraction model in the above paragraph. For example, it can be specifically as follows: for any object, such as a digital image, first construct a facial texture library of the object so that the facial texture library includes the facial texture of the object under various expressions and mouth shapes; then use principal component analysis (PCA) to extract the principal components of the facial texture library to obtain the principal component extraction result corresponding to the object; then, based on the principal component extraction result, construct a facial texture coefficient extraction model corresponding to the object, so that the facial texture coefficient extraction model can perform principal component coefficient extraction processing on any image, such as any image used to describe the object, so that the principal component coefficients extracted by the facial texture coefficient extraction model can be used as the facial texture coefficients corresponding to the image in the future. It can be seen that, in a possible implementation, the facial texture coefficient supervision information corresponding to the sample speech above can be implemented using the principal component coefficients of the reference facial image corresponding to the sample speech; and the principal component coefficients of the reference facial image refer to the principal component coefficient extraction processing performed on the reference facial image by the facial texture coefficient extraction model matching the reference facial image. Among them, the "facial texture coefficient extraction model matching the reference facial image" refers to the facial texture coefficient extraction model pre-constructed for the object described by the reference facial image.
[0090] Based on the content of the above paragraph, it can be seen that in a possible implementation mode, the process of obtaining the facial texture coefficient supervision information corresponding to the sample speech above can be specifically as follows: after obtaining the reference facial image corresponding to the sample speech, first search for a facial texture coefficient extraction model that matches the reference facial image from some pre-built facial texture coefficient extraction models; then use this matching facial texture coefficient extraction model to perform principal component coefficient extraction processing on the reference facial image to obtain the facial texture coefficient supervision information corresponding to the sample speech, so that the facial texture coefficient supervision information can represent the facial texture state corresponding to the sample speech as accurately as possible.
[0091] Based on the relevant content of the facial supervision information corresponding to the sample voice above, it can be known that in some application scenarios, such as scenarios with high requirements for model building efficiency, the facial supervision information corresponding to the sample voice may include facial expression supervision information and facial texture coefficient supervision information, so that the facial supervision information can represent the two facial state true values pre-set for the sample voice, so as to guide the model update based on these two types of supervision information when using the sample voice to train the model, thereby helping to improve the model update efficiency.
[0092] In fact, in some application scenarios, in order to better improve the model performance, the present application also provides a possible implementation of the facial supervision information corresponding to the above sample voice, in which the facial supervision information corresponding to the sample voice includes not only facial expression supervision information and facial texture coefficient supervision information, but also facial expression coefficient supervision information, so that the facial supervision information can represent the three facial state true values set in advance for the sample voice, so that the facial supervision information can more accurately and comprehensively represent the facial details corresponding to the sample voice, and then the facial supervision information can better guide the model to learn facial details, which is conducive to improving the model performance. Among them, the facial expression coefficient supervision information refers to the facial expression coefficient true value set in advance for the sample voice, so that the facial expression coefficient supervision information is used to represent the actual expression details corresponding to the sample voice, so that the facial expression coefficient supervision information can guide the model to learn the facial expression coefficient when the sample voice is used to train the model; and the present application does not limit the implementation of the facial expression coefficient supervision information, for example, the facial expression coefficient supervision information can be implemented using any expression coefficient, such as the coefficient of the 52-dimensional expression feature of arkit. In addition, the present application does not limit the method for obtaining the facial expression coefficient supervision information. For example, in some application scenarios, the facial expression coefficient supervision information may refer to the facial expression coefficient label pre-marked by relevant personnel for the sample voice. For another example, in some application scenarios, in order to reduce the difficulty of obtaining training data, the facial expression coefficient supervision information may be obtained by: performing expression coefficient extraction processing on the reference facial image corresponding to the sample voice to obtain the facial expression coefficient supervision information corresponding to the sample voice, so that the facial expression coefficient supervision information can represent the expression details corresponding to the sample voice as accurately as possible.
[0093] Based on the above content about facial supervision information corresponding to sample speech, it can be known that in some application scenarios, the facial supervision information can be determined based on the reference facial image corresponding to the sample speech, so that the facial supervision information can represent the facial state presented by the reference facial image, which is conducive to reducing the difficulty of obtaining training data. Among them, because the reference facial image can describe the actual facial details corresponding to the sample speech, the facial supervision information determined based on the reference facial image can represent the facial details corresponding to the sample speech as accurately as possible, and then the facial supervision information can better guide the model to learn facial details, which is conducive to improving model performance.
[0094] Based on the relevant content of S101 above, in some application scenarios, if one wants to use a sample voice to participate in model training, the sample voice and the corresponding reference face image can be obtained first; then, using the reference face image, the expression description information corresponding to the sample voice and the face supervision information corresponding to the sample voice can be determined, so as to complete the training and updating process of the model based on the sample voice, the expression description information, and the face supervision information subsequently.
[0095] S102: Input the voice features of the sample voice and the expression features of the expression description information corresponding to the sample voice into the data processing model to obtain the predicted facial expression and the first predicted facial texture coefficient output by the data processing model.
[0096] Among them, the voice features of the sample voice are used to represent the voice information carried by the sample voice; moreover, this application does not limit the acquisition method of the voice features. For example, specifically, it can be: using a pre-constructed voice feature extraction model to perform voice feature extraction processing on the sample voice to obtain the voice features of the sample voice. It should be noted that this application does not limit the implementation manner of the voice feature extraction model. For example, the voice feature extraction model can adopt any model with voice feature extraction function, such as Figure 2 the Automatic Speech Recognition (ASR) model shown, for implementation.
[0097] In addition, for the expression description information corresponding to the above sample voice, the expression features of the expression description information are used to represent the expression state described by the expression description information; moreover, this application does not limit the acquisition method of the expression features. For example, it can be implemented using any existing or future method that can extract expression features from the expression description information.
[0098] Actually, in order to better characterize the expression state, this application also provides a possible implementation manner of the acquisition process of the above expression features. In this implementation manner, when the expression description information corresponding to the above sample voice includes the expression category corresponding to the sample voice and the expression intensity of the expression category, the acquisition process of the expression features of the expression description information can include Step 11 - Step 12 below.
[0099] Step 11: Encode the expression description information corresponding to the above sample voice to obtain the expression encoding corresponding to the sample voice. The expression encoding is used to characterize the intensity state of at least one preset expression category, and the at least one preset expression category includes the expression category corresponding to the sample voice.
[0100] The expression code corresponding to the sample speech is obtained by encoding the expression description information corresponding to the sample speech, so that the expression code can represent the expression state carried by the expression description information.
[0101] For the expression coding corresponding to the sample speech above, the expression coding can be used to characterize the intensity state of at least one preset expression category, so that the expression coding can more accurately represent the expression state carried by the expression description information corresponding to the sample speech. Among them, the preset expression category refers to the coding dimension set in advance according to the actual application scenario and required to be referenced when encoding the expression description information; and the at least one preset expression category at least includes the expression category corresponding to the sample speech.
[0102] In addition, in some application scenarios, in order to better improve the expression state representation effect, the present application also provides a possible implementation method of the expression coding corresponding to the above sample voice, in which the expression coding corresponding to the sample voice may include intensity normalized representation data of at least one preset expression category. Among them, the intensity normalized representation data of the kth preset expression category is used to represent the intensity state corresponding to the sample voice under the kth preset expression category; the intensity normalized representation data of the kth preset expression category belongs to the data interval [0, 1]; and the larger the intensity normalized representation data of the kth preset expression category, the higher the expression intensity of the kth preset expression category. k is a positive integer, k≤the number of categories in the at least one preset expression category.
[0103] In addition, the present application does not limit the method for obtaining the expression code corresponding to the sample voice above. For example, when at least one of the preset expression categories above includes the expression category corresponding to the sample voice and the expression intensity of the expression category, the process of obtaining the expression code can be: normalizing the expression intensity of the expression category corresponding to the sample voice to obtain the intensity normalized representation data of the expression category corresponding to the sample voice, so that the "intensity normalized representation data of the expression category corresponding to the sample voice" is positively correlated with the "expression intensity of the expression category corresponding to the sample voice", so that the "intensity normalized representation data of the expression category corresponding to the sample voice" can better represent the intensity state of the sample voice under the "expression category corresponding to the sample voice"; then use the intensity normalized representation data of the expression category corresponding to the sample voice to replace the preset data value of the corresponding expression category in the pre-constructed expression coding template to obtain the expression code corresponding to the sample voice, so that the expression code can better represent the expression state corresponding to the sample voice. Among them, the "normalized intensity representation data of the expression category corresponding to the sample voice" is used to represent the intensity state of the sample voice under the "expression category corresponding to the sample voice"; the "normalized intensity representation data of the expression category corresponding to the sample voice" belongs to the data interval [0, 1]; and the larger the "normalized intensity representation data of the expression category corresponding to the sample voice", the higher the expression intensity of the "expression category corresponding to the sample voice". The expression coding template refers to a template that is set in advance according to the actual application scenario and is required to be referenced when performing expression coding processing, so that the expression coding template can represent the preset data values of each preset expression category, and the present application does not limit the expression coding template. For example, when the at least one preset expression category includes K preset expression categories, the expression coding template can be implemented using a template such as [preset data value of the first preset expression category, preset data value of the second preset expression category, ..., preset data value of the Kth preset expression category]. The preset data value of the kth preset expression category refers to the initial value of the intensity normalization characterization data set in advance for the kth preset expression category, so that the preset data value of the kth preset expression category can represent the initial intensity state of the kth preset expression category. k is a positive integer, k≤K, and K is a positive integer.
[0104] Based on the content of the above paragraph, it can be known that in a possible implementation mode, when the expression description information corresponding to the above sample voice includes the expression category corresponding to the sample voice and the expression intensity of the expression category, the acquisition process of the expression code corresponding to the sample voice can be specifically as follows: normalizing the expression intensity of the expression category corresponding to the sample voice to obtain the intensity normalized representation data of the expression category corresponding to the sample voice, so that when it is determined that the expression category corresponding to the sample voice matches the target category in at least one preset expression category involved in the above expression coding template, the preset data value of the target category in the expression coding template is directly replaced with the intensity normalized representation data of the expression category corresponding to the sample voice to obtain the expression code corresponding to the sample voice, so that the expression code corresponding to the sample voice includes the intensity normalized representation data of the expression category corresponding to the sample voice, and the preset data values of other categories in the at least one preset expression category except the target category, and the preset data value of the target category does not exist in the expression code corresponding to the sample voice. Among them, the target category refers to a preset expression category that matches the expression category corresponding to the sample voice; and the present application does not limit the method for determining the target category. For example, it can be specifically: first, calculate the similarity between the expression category corresponding to the sample voice and the kth preset expression category, k is a positive integer, k≤K, K is a positive integer; then, determine the preset expression category with the highest similarity as the target category, so that the target category can represent the preset expression category that matches the expression category corresponding to the sample voice.
[0105] It should be noted that the present application does not limit the implementation method of the step of "normalizing the expression intensity of the expression category corresponding to the sample voice to obtain the intensity normalized representation data of the expression category corresponding to the sample voice" in the above two paragraphs. For example, it can be implemented by any existing or future method that can normalize a data, such as a method of normalizing according to a preset mapping rule. Among them, the preset mapping rule refers to a rule set in advance according to the actual application scenario for mapping any data to the data interval [0, 1]; and the present application does not limit the preset mapping rule.
[0106] Based on the relevant content of the expression coding corresponding to the sample voice above, it can be known that under a possible implementation mode, the expression coding corresponding to the sample voice can be implemented using the coding format of [intensity normalized representation data of the first preset expression category, intensity normalized representation data of the second preset expression category, ..., intensity normalized representation data of the Kth preset expression category] so that the expression coding can better represent the expression state corresponding to the sample voice.
[0107] Based on the relevant content of step 11 above, it can be known that for some application scenarios, after obtaining the expression description information corresponding to the sample voice above, the expression description information can be encoded to obtain the expression code corresponding to the sample voice, so that the expression code can better express the expression state corresponding to the sample voice.
[0108] Step 12: Based on the above expression coding, determine the expression features of the expression description information corresponding to the above sample speech.
[0109] It should be noted that the present application does not limit the implementation method of the above step 12. For example, in some application scenarios, the step 12 may specifically be: determining the above expression code as the expression feature of the expression description information corresponding to the above sample speech.
[0110] In addition, in order to better improve the expression state representation effect, the present application also provides a possible implementation method of step 12 above. Under this implementation method, step 12 can specifically be: feature embedding the above expression code to obtain the expression feature of the expression description information corresponding to the above sample speech, so that the expression feature includes the Embedding feature vector corresponding to each preset expression category, so that the expression feature can better represent the expression state described by the expression code, thereby helping to improve the expression state representation effect. It should be noted that the present application does not limit the implementation method of the feature embedding processing. For example, it can adopt any existing or future fully connected layer, such as Figure 2 The Deep Neural Networks (DNN) shown in the figure are implemented.
[0111] In addition, it has been found through research that in some application scenarios, in order to better improve the image generation effect, it is necessary to align the above expression features with the above voice features in a certain feature dimension, such as width, so as to improve the subsequent feature splicing effect. Based on this, the present application also provides a possible implementation method of the above step 12. Under this implementation method, the step 12 can be specifically: feature embedding processing is performed on the above expression coding to obtain the expression features of the expression description information corresponding to the above sample voice, so that the size value of the expression feature in the first dimension is equal to the size value of the voice feature of the sample voice in the first dimension. In this way, the expression feature and the voice feature can be aligned in the first dimension, so that the defects caused by the large difference between the size value of the expression feature in the first dimension and the size value of the voice feature in the first dimension can be effectively avoided (for example, defects such as inability to perform feature splicing), which is beneficial to improving the image generation effect. Among them, the first dimension refers to the feature dimension that needs to be aligned between the facial expression feature and the voice feature; and the present application does not limit the implementation method of the first dimension. For example, when the facial expression feature includes at least one feature vector and the voice feature includes at least one feature vector, in order to better improve the feature splicing effect, it is necessary to ensure that the size of each feature vector in the facial expression feature is consistent with the size of each feature vector in the voice feature. In some application scenarios, the width of the facial expression feature can be used to represent the size of each feature vector in the facial expression feature, and the width of the voice feature can be used to represent the size of each feature vector in the voice feature. Based on the foregoing content, it can be seen that in these application scenarios, the first dimension can be implemented using the width dimension, which is conducive to aligning the size of each feature vector in the facial expression feature with the size of each feature vector in the voice feature.
[0112] Based on the content of the previous paragraph, it can be seen that under a possible implementation mode, for the speech features of the above sample speech and the expression features of the expression description information corresponding to the sample speech, the two features can satisfy the following constraints: the number of feature vectors in the speech feature is different from the number of feature vectors in the expression feature, but the size of any feature vector in the speech feature is the same as the size of any feature vector in the expression feature, so that the speech feature and the expression feature can be aligned in the dimension of the size of the feature vector, thereby effectively avoiding the defects caused by the large difference between the size of the feature vector in the expression feature and the size of the feature vector in the speech feature, which is beneficial to improving the image generation effect.
[0113] Based on the relevant contents of steps 11 to 12 above, it can be known that in a possible implementation mode, when the expression description information corresponding to the sample voice above includes the expression category corresponding to the sample voice and the expression intensity of the expression category, the process of obtaining the expression features of the expression description information may include: first encoding according to the expression category corresponding to the sample voice and the expression intensity of the expression category to obtain the expression code corresponding to the sample voice, such as [intensity normalized representation data of the first preset expression category, intensity normalized representation data of the second preset expression category, ..., intensity normalized representation data of the Kth preset expression category]; then using DNN to process the expression code to obtain the expression features of the expression description information, so that the size value of the expression feature in the first dimension is equal to the size value of the voice feature of the sample voice in the first dimension, so that the expression feature and the voice feature can be spliced according to the first dimension, so that the splicing processing result can better represent the facial state corresponding to the sample voice, thereby making the facial related information determined based on the splicing processing result and the data processing model more accurate.
[0114] The data processing model is used to perform some processing on the input data of the data processing model, such as facial expression prediction processing, facial expression coefficient prediction processing, facial texture coefficient prediction processing, etc.
[0115] In addition, the present application does not limit the model structure of the above data processing model. For example, it can be implemented using a recurrent neural network (RNN), a gated recurrent unit (GRU), or a long short-term memory network (LSTM).
[0116] In addition, the present application does not limit the working principle of the above data processing model. For ease of understanding, some examples are provided below for illustration.
[0117] Example 1, in some application scenarios, such as scenarios with high model building efficiency, the working principle of the above data processing model can be specifically as follows: the speech features of the above sample speech and the expression features of the expression description information corresponding to the sample speech are input into the data processing model to obtain the predicted facial expression and the first predicted facial texture coefficient output by the data processing model. Among them, the predicted facial expression refers to the facial expression predicted by the data processing model for the sample speech, so that the predicted facial expression can express the predicted facial expression state corresponding to the sample speech, such as the state presented in terms of expression and lip shape, etc.; and the implementation method of the predicted facial expression is similar to the implementation method of the facial expression supervision information above, and for the sake of brevity, it will not be repeated here. The first predicted facial texture coefficient refers to the facial texture coefficient predicted by the data processing model for the sample speech, so that the first predicted facial texture coefficient can represent the predicted facial texture state corresponding to the sample speech, such as facial contour, distribution of facial features, and other aspects other than expression and lip shape; and the implementation method of the first predicted facial texture coefficient is similar to the implementation method of the facial texture coefficient supervision information above. For the sake of brevity, it will not be repeated here.
[0118] Example 2: In some application scenarios, such as scenarios with high requirements for model performance, the working principle of the above data processing model can be specifically as follows: the speech features of the above sample speech and the expression features of the expression description information corresponding to the sample speech are input into the data processing model to obtain the predicted facial expression, the first predicted facial expression coefficient and the first predicted facial texture coefficient output by the data processing model. Among them, the first predicted facial expression coefficient refers to the facial expression coefficient predicted by the data processing model for the sample speech, so that the first predicted facial expression coefficient can represent the predicted expression details corresponding to the sample speech, such as the details presented in the expression and mouth shape, etc.; and the implementation method of the first predicted facial expression coefficient is similar to the implementation method of the facial expression coefficient supervision information above, and for the sake of brevity, it will not be repeated here.
[0119] It can be seen that, in one possible implementation, for the above data processing model, if the data processing model is used to process the voice features of the above sample speech and the expression features of the expression description information corresponding to the sample speech, then the output result of the data processing model can at least include predicted facial expressions, first predicted facial expression coefficients and first predicted facial texture coefficients, so that the output result can describe the facial state from more angles, so that the output result can better represent the predicted facial state corresponding to the sample speech, so that the performance of the data processing model in facial state prediction can be more accurately determined based on the output result, which is conducive to improving model performance.
[0120] Example 3, in some application scenarios, when the size value of the voice feature of the above sample speech in the first dimension is equal to the size value of the expression feature of the expression description information corresponding to the sample speech in the first dimension, the working principle of the above data processing model can be specifically: after the voice feature and the expression feature are spliced in the second dimension to obtain the first spliced feature, the first spliced feature is input into the data processing model to obtain the predicted facial expression and the first predicted facial texture coefficient output by the data processing model. The second dimension refers to the dimension required to refer to when splicing the voice feature and the expression feature, and the second dimension is different from the first dimension. For example, if the first dimension is width, the second dimension is height. The first spliced feature refers to the result obtained by splicing the speech feature and the expression feature, so that the size value of the first spliced feature in the second dimension is equal to the sum of the size value of the expression feature in the second dimension and the size value of the speech feature in the second dimension, so that the first spliced feature includes all feature vectors in the speech feature and all feature vectors in the expression feature, and thus the first spliced feature can better represent the facial state corresponding to the sample speech, such as expression, mouth shape, contour and other states.
[0121] Example 4. In some application scenarios, when the size value of the speech feature of the above sample speech in the first dimension is equal to the size value of the expression feature of the expression description information corresponding to the sample speech in the first dimension, the working principle of the above data processing model can be specifically: after splicing the speech feature and the expression feature in the second dimension to obtain a first spliced feature, the first spliced feature is input into the data processing model to obtain the predicted facial expression, the first predicted facial expression coefficient and the first predicted facial texture coefficient output by the data processing model.
[0122] Based on the above two paragraphs, it can be known that in a possible implementation, the above S102 can be specifically: splicing the voice features of the above sample voice and the expression features of the expression description information corresponding to the sample voice in the second dimension to obtain a first splicing feature, so that the size value of the first splicing feature in the second dimension is equal to the sum of the size value of the expression feature in the second dimension and the size value of the voice feature in the second dimension; then input the first splicing feature into the data processing model to obtain the output result of the data processing model. The output result may include the predicted facial expression and the first predicted facial texture coefficient; or, the output result may include the predicted facial expression, the first predicted facial expression coefficient and the first predicted facial texture coefficient.
[0123] In addition, the present application does not limit the implementation method of the above data processing model. For example, in some application scenarios, the data processing model can be used. Figure 2The regressor shown in FIG. 1 is implemented by the regressor shown in FIG. 1 , wherein the regressor can be used to map the input data of the regressor to facial expressions, facial expression coefficients, and facial texture coefficients.
[0124] Based on the relevant content of S102 above, it can be known that in one possible implementation mode, after obtaining the voice features of the above sample speech and the expression features of the expression description information corresponding to the sample speech, the voice features and the expression features can be input into a data processing model, so that the data processing model processes the voice features and the expression features to obtain the output result of the data processing model, so that the output result at least includes the predicted facial expression and the first predicted facial texture coefficient, so that the model performance of the data processing model can be evaluated based on the output result later.
[0125] S103: Determine a discriminant result of predicting facial expression using a discriminator corresponding to the data processing model.
[0126] Among them, the discriminator is used to assist the above data processing model to achieve adversarial learning during the model training process; and this application does not limit the implementation method of the discriminator. For example, it can be implemented using any existing or future discriminator required for achieving adversarial learning.
[0127] In addition, the present application does not limit the working principle of the above discriminator. For example, the discriminator can be used to perform accuracy discrimination processing on the input data of the discriminator. It can be seen that in one possible implementation, for the discriminator corresponding to the above data processing model, the discriminator can be used to determine the accuracy of the generated expression, that is, the discriminator can be used to determine whether the input data of the discriminator is the predicted facial expression predicted by the data processing model, or the facial expression supervision information used to describe the real expression, so that the discriminator can be used to distinguish the predicted facial expression from the facial expression supervision information.
[0128] In addition, for the predicted facial expression obtained by the above data processing model for the sample speech prediction, the discrimination result of the predicted facial expression is obtained by using the discriminator corresponding to the data processing model to perform expression accuracy discrimination processing on the predicted facial expression, so that the discrimination result can indicate whether the predicted facial expression belongs to the true value of the expression, and thus the discrimination result can indicate whether the discriminator can recognize that the predicted facial expression does not belong to the true value of the expression.
[0129] In addition, the present application does not limit the implementation method of S103 above. For example, S103 may specifically be: after obtaining the predicted facial expression determined by the data processing model for the sample speech, the predicted facial expression is input into the discriminator corresponding to the data processing model, so that the discriminator performs expression accuracy discrimination processing on the predicted facial expression, obtains and outputs a discrimination result of the predicted facial expression, so that the discrimination result can indicate whether the predicted facial expression belongs to the true value of the expression, so that the model performance of the data processing model can be subsequently evaluated based on the discrimination result.
[0130] Based on the relevant content of S103 above, it can be known that in some application scenarios, after obtaining the predicted facial expression output by the data processing model for the sample speech, the discriminator corresponding to the data processing model can be used to determine the discrimination result of the predicted facial expression, so that the model performance of the data processing model can be evaluated based on the discrimination result.
[0131] S104: updating the data processing model according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression.
[0132] In the present application, for the above sample speech, after obtaining the predicted facial expression corresponding to the sample speech, the first predicted facial texture coefficient corresponding to the sample speech, the facial expression supervision information corresponding to the sample speech, the facial texture coefficient supervision information corresponding to the sample speech, and the discrimination result of the predicted facial expression corresponding to the sample speech, the data processing model can be updated based on these data so that the updated data processing model has better model performance, and continue to execute S101 and its subsequent steps based on the updated data processing model to achieve the next round of training for the data processing model, and repeat the cycle until the preset stop condition is reached. Among them, the preset stop condition refers to the condition that needs to be achieved when the iterative training process of the data processing model ends.
[0133] It should be noted that the present application does not limit the implementation method of the preset stop condition in the above paragraph. For example, the preset stop condition may specifically include that the model loss of the data processing model is lower than a preset loss threshold. For another example, the preset stop condition may specifically include that the rate of change of the model loss of the data processing model is lower than a preset rate of change threshold. For another example, the preset stop condition may specifically include that the number of updates of the data processing model reaches a preset number threshold. Among them, the model loss of the data processing model is used to characterize the model performance of the data processing model, and the present application does not limit the implementation method of the model loss of the data processing model. For example, the model loss of the data processing model can be implemented using the model loss shown in steps 21 to 24 below or the model loss shown in steps 31 to 35 below.
[0134] In fact, in order to better improve the model updating effect, the present application also provides a possible implementation of the above S104. Under this implementation, the S104 may include the following steps 21 to 25.
[0135] Step 21: Determine the facial expression prediction loss based on the difference representation data between the predicted facial expression and the facial expression supervision information.
[0136] In the present application, after the predicted facial expression is determined for the sample speech using the above data processing model, the difference representation data between the predicted facial expression and the facial expression supervision information corresponding to the sample speech is first calculated, so that the difference representation data can represent the difference between the predicted facial expression and the facial expression supervision information, thereby making the difference representation data able to represent the gap between the predicted facial expression and the true value of the expression corresponding to the sample speech; then, the facial expression prediction loss is determined based on the difference representation data, so that the facial expression prediction loss can represent the performance of the data processing model in facial expression prediction.
[0137] It should be noted that the present application does not limit the determination process of the above "difference representation data between the predicted facial expression and the facial expression supervision information". For example, it can be: according to the preset distance calculation formula, the distance between the predicted facial expression and the facial expression supervision information is calculated as the difference representation data between the predicted facial expression and the facial expression supervision information. Among them, the preset distance calculation formula refers to a formula for calculating the distance between two information set in advance according to the actual application scenario; and the present application does not limit the preset distance calculation formula. For example, it can be implemented using Euclidean distance or cosine distance.
[0138] It should also be noted that the present application does not limit the implementation method of step 21 above. For example, step 21 may specifically be: determining the difference representation data between the predicted facial expression and the facial expression supervision information as the facial expression prediction loss. For another example, step 21 may also be: substituting the difference representation data between the predicted facial expression and the facial expression supervision information into a pre-set expression prediction loss function to obtain the facial expression prediction loss. The expression prediction loss function refers to a pre-set function required for calculating the facial expression prediction loss.
[0139] Step 22: Determine the facial texture coefficient prediction loss based on the difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information.
[0140] In the present application, after determining the first predicted facial texture coefficient for the sample speech using the above data processing model, the difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information corresponding to the sample speech is first calculated, so that the difference representation data can represent the difference between the first predicted facial texture coefficient and the facial texture coefficient supervision information, thereby enabling the difference representation data to represent the difference between the first predicted facial texture coefficient and the true value of the texture coefficient corresponding to the sample speech; then, the facial texture coefficient prediction loss is determined based on the difference representation data, so that the facial texture coefficient prediction loss can represent the performance of the data processing model in facial texture coefficient prediction.
[0141] It should be noted that the present application does not limit the determination process of the above "difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information". For example, it can be: according to a preset distance calculation formula, calculate the distance between the first predicted facial texture coefficient and the facial texture coefficient supervision information as the difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information.
[0142] It should also be noted that the present application does not limit the implementation method of the above step 22. For example, the step 22 may specifically be: determining the difference characterization data between the first predicted facial texture coefficient and the facial texture coefficient supervision information as the facial texture coefficient prediction loss. For another example, the step 22 may also be: substituting the difference characterization data between the first predicted facial texture coefficient and the facial texture coefficient supervision information into a pre-set texture coefficient prediction loss function to obtain the facial texture coefficient prediction loss. The texture coefficient prediction loss function refers to a pre-set function required for use in calculating the facial texture coefficient prediction loss.
[0143] Step 23: Determine the accuracy representation data of the predicted facial expression based on the judgment result of the predicted facial expression.
[0144] The accuracy characterization data of the predicted facial expression can indicate the accuracy of the predicted facial expression predicted by the data processing model for the sample speech.
[0145] In addition, the present application does not limit the determination process of the accuracy characterization data of the above-mentioned predicted facial expression. For example, it can be implemented by any existing or future method that can determine the degree of accuracy based on the discrimination result. For another example, step 23 can specifically be: directly determine the discrimination result of the above-mentioned predicted facial expression as the accuracy characterization data of the predicted facial expression. For another example, step 23 can specifically be: substitute the discrimination result of the predicted facial expression into a pre-set expression accuracy evaluation function to obtain the accuracy characterization data of the predicted facial expression. Among them, the expression accuracy evaluation function refers to a pre-set function required for use in calculating expression accuracy.
[0146] Step 24: Determine the model loss of the data processing model based on the facial expression prediction loss, the facial texture coefficient prediction loss, and the accuracy characterization data of the predicted facial expression.
[0147] It should be noted that the present application does not limit the implementation method of the above step 24. For example, it can be specifically: weighted summing the above facial expression prediction loss, the above facial texture coefficient prediction loss, and the above predicted facial expression accuracy representation data to obtain the model loss of the data processing model. For another example, the step 24 can be specifically: combining the facial expression prediction loss, the facial texture coefficient prediction loss, and the predicted facial expression accuracy representation data to obtain the model loss of the data processing model.
[0148] Step 25: Update the data processing model based on the model loss of the above data processing model.
[0149] In the present application, after obtaining the model loss of the above data processing model, the data processing model can be updated based on the model loss so that the updated data processing model has better model performance, which is conducive to improving the model performance of the data processing model.
[0150] Based on the relevant contents of steps 21 to 25 above, it can be known that in some application scenarios, after obtaining the predicted facial expression corresponding to the above sample voice, the first predicted facial texture coefficient corresponding to the sample voice, the facial expression supervision information corresponding to the sample voice, the facial texture coefficient supervision information corresponding to the sample voice, and the judgment result of the predicted facial expression corresponding to the sample voice, the above data processing model can be updated based on the difference representation data between the predicted facial expression and the facial expression supervision information, the difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information, and the judgment result of the predicted facial expression, so that the updated data processing model has better model performance, which is conducive to improving the model performance of the data processing model.
[0151] In fact, in order to better improve the model performance, the present application also provides a possible implementation of S104 above, under which, when the facial supervision information corresponding to the sample speech above includes facial expression supervision information, facial expression coefficient supervision information and facial texture coefficient supervision information, and the data processing model above determines the predicted facial expression, the first predicted facial expression coefficient and the first predicted facial texture coefficient for the sample speech, S104 can specifically be: based on the predicted facial expression, the first predicted facial expression coefficient, the first predicted facial texture coefficient, the facial expression supervision information, the facial expression coefficient supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression, update the data processing model. Among them, because the information referenced when updating the data processing model is relatively rich, the updated data processing model has better model performance, which is conducive to better improving the model performance of the data processing model.
[0152] In addition, the present application does not limit the implementation method of the step "updating the data processing model based on the predicted facial expression, the first predicted facial expression coefficient, the first predicted facial texture coefficient, the facial expression supervision information, the facial expression coefficient supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression" in the previous paragraph. For ease of understanding, a possible implementation method of S104 is used as an example to illustrate below.
[0153] As an example, in a possible implementation manner, the above S104 may specifically include the following steps 31 to 36.
[0154] Step 31: Determine the facial expression prediction loss based on the difference representation data between the predicted facial expression and the facial expression supervision information.
[0155] It should be noted that the relevant content of step 31 can be found in the above step 21, and for the sake of brevity, it will not be repeated here.
[0156] Step 32: Determine the facial expression coefficient prediction loss based on the difference representation data between the first predicted facial expression coefficient and the facial expression coefficient supervision information.
[0157] In the present application, after determining the first predicted facial expression coefficient for the sample speech using the above data processing model, the difference representation data between the first predicted facial expression coefficient and the facial expression coefficient supervision information corresponding to the sample speech is first calculated, so that the difference representation data can represent the difference between the first predicted facial expression coefficient and the facial expression coefficient supervision information, thereby enabling the difference representation data to represent the difference between the first predicted facial expression coefficient and the true value of the expression coefficient corresponding to the sample speech; then, the facial expression coefficient prediction loss is determined based on the difference representation data, so that the facial expression coefficient prediction loss can represent the performance of the data processing model in facial expression coefficient prediction.
[0158] It should be noted that the present application does not limit the determination process of the above-mentioned "difference representation data between the first predicted facial expression coefficient and the facial expression coefficient supervision information". For example, it can be: according to a preset distance calculation formula, calculate the distance between the first predicted facial expression coefficient and the facial expression coefficient supervision information as the difference representation data between the first predicted facial expression coefficient and the facial expression coefficient supervision information.
[0159] It should also be noted that the present application does not limit the implementation method of the above step 32. For example, the step 32 may specifically be: determining the difference representation data between the first predicted facial expression coefficient and the facial expression coefficient supervision information as the facial expression coefficient prediction loss. For another example, the step 32 may also be: substituting the difference representation data between the first predicted facial expression coefficient and the facial expression coefficient supervision information into a pre-set expression coefficient prediction loss function to obtain the facial expression coefficient prediction loss. The expression coefficient prediction loss function refers to a pre-set function required for use in calculating the facial expression coefficient prediction loss.
[0160] Step 33: Determine the facial texture coefficient prediction loss based on the difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information.
[0161] It should be noted that the relevant content of step 33 can be found in step 22 above, and for the sake of brevity, it will not be repeated here.
[0162] Step 34: Determine the accuracy representation data of the predicted facial expression based on the judgment result of the predicted facial expression.
[0163] It should be noted that the relevant content of step 34 can be found in step 23 above, and for the sake of brevity, it will not be repeated here.
[0164] Step 35: Determine the model loss of the data processing model based on the facial expression prediction loss, the facial expression coefficient prediction loss, the facial texture coefficient prediction loss, and the accuracy characterization data of the predicted facial expression.
[0165] It should be noted that the implementation of step 35 is similar to the implementation of step 24 above, and for the sake of brevity, it will not be repeated here.
[0166] Step 36: Update the data processing model based on the model loss of the above data processing model.
[0167] It should be noted that the relevant content of step 36 can be found in step 25 above, and for the sake of brevity, it will not be repeated here.
[0168] Based on the relevant contents of steps 31 to 36 above, it can be known that in some application scenarios, after obtaining the predicted facial expression corresponding to the above sample voice, the first predicted facial expression coefficient corresponding to the sample voice, the first predicted facial texture coefficient corresponding to the sample voice, the facial expression supervision information corresponding to the sample voice, the facial expression coefficient supervision information corresponding to the sample voice, the facial texture coefficient supervision information corresponding to the sample voice, and the discrimination result of the predicted facial expression corresponding to the sample voice, the above data processing model can be updated based on the difference representation data between the predicted facial expression and the facial expression supervision information, the difference representation data between the first predicted facial expression coefficient and the facial expression coefficient supervision information, the difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information, and the discrimination result of the predicted facial expression, so that the updated data processing model has better model performance, which is conducive to improving the model performance of the data processing model.
[0169] Based on the relevant contents of S101 to S104 above, it can be known that for the model building method provided in the embodiment of the present application, a sample voice, expression description information corresponding to the sample voice, and facial supervision information corresponding to the sample voice are first obtained, and the facial supervision information includes facial expression supervision information and facial texture coefficient supervision information; then, the voice features of the sample voice and the expression features of the expression description information are input into the data processing model to obtain the predicted facial expression and the first predicted facial texture coefficient output by the data processing model; then, the discriminator corresponding to the data processing model is used to determine the discrimination result of the predicted facial expression; finally, according to the predicted facial expression, a discriminator is generated to identify the facial expression; and finally, a discriminator is generated to identify the facial expression. The data processing model is updated based on the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the discrimination result, so that the updated data processing model has better performance, such as better facial expression prediction performance and facial texture coefficient prediction performance, etc., so that when the data processing model is used to perform facial image generation processing on a voice, a better facial description image can be generated, and then the face presented by the facial description image has a more realistic and natural expression, which is beneficial to improving the image generation effect, thereby helping to better meet the needs of rich and full expressions in the voice image generation scenario.
[0170] In addition, in order to better realize the adversarial learning between the data processing model and the discriminator, the data processing model and the discriminator can be used for model building processing by iterative training. Based on this, the present application also provides a possible implementation of the above model building method, under which the model building method can include the following steps 41-45.
[0171] Step 41: obtaining a sample voice, expression description information corresponding to the sample voice, and facial supervision information corresponding to the sample voice, wherein the facial supervision information includes facial expression supervision information and facial texture coefficient supervision information.
[0172] It should be noted that the relevant contents of step 41 refer to S101 above, and for the sake of brevity, they will not be repeated here.
[0173] Step 42: Input the speech features of the sample speech and the expression features of the expression description information corresponding to the sample speech into a data processing model to obtain the predicted facial expression and the first predicted facial texture coefficient output by the data processing model.
[0174] It should be noted that the relevant contents of step 42 can be found in S102 above, and for the sake of brevity, they will not be repeated here.
[0175] Step 43: Using the discriminator corresponding to the data processing model, determine the discriminant result of the predicted facial expression and the discriminant result of the facial expression supervision information.
[0176] Among them, the discrimination result of the facial expression supervision information is obtained by using the discriminator corresponding to the above data processing model to perform expression accuracy discrimination processing on the facial expression supervision information, so that the discrimination result of the facial expression supervision information can indicate whether the facial expression supervision information belongs to the true value of expression, thereby making the discrimination result of the facial expression supervision information indicate whether the discriminator can recognize that the facial expression supervision information belongs to the true value of expression.
[0177] In addition, the present application does not limit the implementation method of the above step 43. For example, the step 43 may specifically include the following steps 431 and 432.
[0178] Step 431: After obtaining the facial expression supervision information corresponding to the above sample speech, the facial expression supervision information is input into the discriminator corresponding to the above data processing model, so that the discriminator performs expression accuracy discrimination processing on the facial expression supervision information, obtains and outputs the discrimination result of the facial expression supervision information, so that the discrimination result can indicate whether the facial expression supervision information belongs to the true value of the expression, so that the performance of the discriminator can be evaluated based on the discrimination result later.
[0179] Step 432: After using the above data processing model to determine the predicted facial expression for the sample speech, the predicted facial expression is input into the discriminator corresponding to the data processing model, so that the discriminator performs expression accuracy discrimination processing on the predicted facial expression, obtains and outputs a discrimination result of the predicted facial expression, so that the discrimination result can indicate whether the predicted facial expression belongs to the true value of the expression, so that the performance of the discriminator can be evaluated based on the discrimination result, or the model performance of the data processing model can be evaluated based on the discrimination result.
[0180] Based on the relevant content of step 43 above, it can be known that in some application scenarios, after obtaining the predicted facial expression and facial expression supervision information corresponding to the above sample speech, the discriminator corresponding to the data processing model can be used to determine the discrimination result of the predicted facial expression and the discrimination result of the facial expression supervision information, so that the performance of the target that needs to be updated in the current round of training, such as the data processing model or the discriminator, can be evaluated based on part or all of these two discrimination results.
[0181] Step 44: If the preset model update condition is met, the data processing model is updated according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression.
[0182] Among them, the preset model update condition refers to the condition that needs to be met when iteratively updating the data processing model, and the present application does not limit the preset model update condition. For example, the preset model update condition may be: in the current round of training, the parameters in the data processing model are in an updateable state, but the parameters in the discriminator corresponding to the data processing model are in a fixed state.
[0183] In addition, the implementation of step 44 is similar to the implementation of S104 above, and for the sake of brevity, it will not be repeated here.
[0184] Based on the relevant content of step 44 above, it can be known that for the current round of training, if the parameters in the above data processing model are in an updateable state, but the parameters in the discriminator corresponding to the data processing model are in a fixed state, it can be determined that the target to be updated in the current round of training is the data processing model. Therefore, the data processing model can be updated based on the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the discrimination result of the predicted facial expression, so that the updated data processing model has better model performance, which is beneficial to improve the model performance of the data processing model.
[0185] Step 45: If the preset discriminator update condition is met, the discriminator is updated according to the discriminant result of the predicted facial expression and the discriminant result of the facial expression supervision information.
[0186] Among them, the preset discriminator update condition refers to the condition that needs to be met when iteratively updating the above discriminator, and the present application does not limit the preset discriminator update condition. For example, the preset discriminator update condition can be: in the current round of training, the parameters in the data processing model are in a fixed state, but the parameters in the discriminator corresponding to the data processing model are in an updateable state.
[0187] In addition, the present application does not limit the implementation method of step 45. For example, it can be implemented by using any existing or future discriminator updating method.
[0188] Based on the relevant content of step 45 above, it can be known that for the current round of training, if the parameters in the above data processing model are in a fixed state, but the parameters in the discriminator corresponding to the data processing model are in an updateable state, it can be determined that the target to be updated in the current round of training is the discriminator. Therefore, the discriminator can be updated based on the discrimination results of the predicted facial expressions and the discrimination results of the facial expression supervision information, so that the updated discriminator has better discrimination performance, so that when the data processing model is subsequently trained with the help of the updated discriminator, the data processing model can better learn facial information, which is beneficial to improve the model performance of the data processing model.
[0189] Based on the relevant contents of steps 41 to 45 above, it can be known that in some application scenarios, adversarial learning between the data processing model and the discriminator can be achieved through iterative training of the data processing model and the discriminator, so that the data processing model finally obtained has better model performance, which is beneficial to improve the model building effect.
[0190] In addition, the present application does not limit the execution subject of the model construction method provided in the embodiment of the present application. For example, the model construction method provided in the embodiment of the present application can be applied to a terminal device or a server. For another example, the model construction method provided in the embodiment of the present application can also be implemented with the help of the data interaction process between the terminal device and the server. Among them, the terminal device can be a smart phone, a computer, a personal digital assistant (PDA), a tablet computer, etc. The server can be a stand-alone server, a cluster server, or a cloud server.
[0191] In addition, in order to better improve the image generation effect in the voice image scene, the voice image generation task can be achieved with the help of the constructed data processing model. Based on this, the present application also provides an image generation method, which is described below in conjunction with the accompanying drawings. Figure 3 As shown, the image generation method provided in the embodiment of the present application may include the following S301-S303. Figure 3 A flowchart of an image generation method provided in an embodiment of the present application.
[0192] S301: Acquire a speech to be processed and expression description information corresponding to the speech to be processed.
[0193] Among them, the speech to be processed refers to the speech required to be used in the speech image scenario; and the present application does not limit the method of obtaining the speech to be processed. For example, in some application scenarios, if video generation processing is required for a long speech, the speech to be processed may refer to any speech segment in the long speech, so that the image generation process for the speech to be processed can be used to realize image generation processing for each speech segment in the long speech.
[0194] In addition, for the above-mentioned speech to be processed, the expression description information corresponding to the speech to be processed is used to describe the expression state corresponding to the speech to be processed; and the implementation method of the expression description information corresponding to the speech to be processed is similar to the implementation method of the expression description information corresponding to the above-mentioned sample speech. It can be seen that in a possible implementation method, the expression description information corresponding to the speech to be processed can include the expression category corresponding to the speech to be processed and the expression intensity of the expression category.
[0195] In addition, the present application does not limit the method for obtaining the expression description information corresponding to the above-mentioned voice to be processed. For example, it can be provided manually by relevant personnel.
[0196] S302: Input the speech features of the speech to be processed and the expression features of the expression description information corresponding to the speech to be processed into a pre-constructed data processing model to obtain a second predicted facial expression coefficient and a second predicted facial texture coefficient output by the data processing model; the data processing model is constructed using any possible implementation of the model construction method provided in the present application.
[0197] Among them, the speech features of the speech to be processed are used to represent the speech information carried by the speech to be processed; and the method of obtaining the speech features of the speech to be processed is similar to the method of obtaining the speech features of the sample speech above. It can be seen that under a possible implementation, the process of obtaining the speech features of the speech to be processed can be: input the speech to be processed into the ASR model, and obtain the speech features output by the ASR model, so that the speech features can represent the speech information carried by the speech to be processed.
[0198] For the facial expression description information corresponding to the above-mentioned speech to be processed, the facial expression features of the facial expression description information are used to represent the facial expression state described by the facial expression description information; and the method for obtaining the facial expression features of the facial expression description information corresponding to the above-mentioned speech to be processed is similar to the method for obtaining the facial expression features of the above-mentioned sample speech. It can be seen that, under a possible implementation method, the process of obtaining the facial expression features of the facial expression description information corresponding to the above-mentioned speech to be processed can be: firstly encode the facial expression description information corresponding to the above-mentioned speech to be processed to obtain the facial expression code corresponding to the above-mentioned speech to be processed; and then use DNN to process the facial expression code to obtain the facial expression features of the facial expression description information corresponding to the above-mentioned speech to be processed, so that the size value of the facial expression features in the first dimension is equal to the size value of the speech features of the above-mentioned speech to be processed. Among them, the facial expression code corresponding to the speech to be processed is used to represent the facial expression state corresponding to the above-mentioned speech to be processed; and the implementation method of the facial expression code corresponding to the above-mentioned speech to be processed is similar to the implementation method of the facial expression code corresponding to the above-mentioned sample speech, and for the sake of brevity, it will not be repeated here.
[0199] The data processing model refers to a model constructed using any possible implementation of the model construction method provided in this application; and the relevant content of the data processing model can be found above.
[0200] The second predicted facial expression coefficient refers to the facial expression coefficient predicted by the data processing model for the speech to be processed, so that the second predicted facial expression coefficient can represent the predicted facial expression state corresponding to the sample speech, such as the state presented in terms of expression and lip shape, etc.; and the implementation method of the second predicted facial expression coefficient is similar to the implementation method of the first predicted facial expression coefficient in S102 above. For the sake of brevity, it will not be repeated here.
[0201] The second predicted facial texture coefficient refers to the facial texture coefficient predicted by the data processing model for the speech to be processed, so that the second predicted facial texture coefficient can represent the predicted facial texture state corresponding to the sample speech, such as the state presented in terms of facial contour, distribution of facial features, etc.; and the implementation method of the second predicted facial texture coefficient is similar to the implementation method of the first predicted facial texture coefficient in S102 above, and for the sake of brevity, it will not be repeated here.
[0202] In addition, the present application does not limit the implementation of S302 above. For example, the implementation of S302 is similar to the implementation of S102 above. For ease of understanding, it will not be repeated here.
[0203] It can be seen that in a possible implementation, when the size value of the voice feature of the speech to be processed in the first dimension is equal to the size value of the expression feature of the expression description information corresponding to the speech to be processed in the first dimension, the above S302 can be specifically as follows: firstly, the voice feature of the speech to be processed and the expression feature of the expression description information corresponding to the speech to be processed are spliced in the second dimension to obtain a second spliced feature, so that the size value of the second spliced feature in the second dimension is equal to the sum of the size value of the expression feature in the second dimension and the size value of the speech feature in the second dimension; then, the second spliced feature is input into the data processing model to obtain the second predicted facial expression coefficient and the second predicted facial texture coefficient output by the data processing model. It should be noted that the implementation method of the second spliced feature is similar to the implementation method of the first spliced feature in the above text, and for the sake of brevity, it will not be repeated here.
[0204] S303: Determine a facial description image corresponding to the speech to be processed according to the second predicted facial expression coefficient and the second predicted facial texture coefficient.
[0205] The facial description image corresponding to the speech to be processed is used to describe the facial state corresponding to the speech to be processed.
[0206] In addition, the present application does not limit the implementation method of the above S303. For example, it can be implemented by adopting any existing or future method that can perform facial image generation processing based on facial expression coefficients and facial texture coefficients.
[0207] In addition, in order to better improve the image generation effect, the present application also provides a possible implementation of the above S303. In this implementation, the S303 may include the following steps 51 and 52.
[0208] Step 51: constructing a full face texture corresponding to the speech to be processed according to the second predicted facial expression coefficient and the second predicted facial texture coefficient.
[0209] The full-face texture corresponding to the speech to be processed is used to describe the full-face state corresponding to the speech to be processed.
[0210] In addition, the present application does not limit the implementation of the above step 51. For example, it can be implemented by any existing or future method that can perform full-face texture construction processing based on the second predicted facial expression coefficient and the second predicted facial texture coefficient, such as a method for performing full-face texture construction processing with the help of a pre-built full-face texture construction model. The full-face texture construction model is used to perform full-face texture construction processing based on the second predicted facial expression coefficient and the second predicted facial texture coefficient to obtain and output full-face texture; and the present application does not limit the implementation of the full-face texture construction model.
[0211] Step 52: Generate a facial description image corresponding to the speech to be processed based on the full face texture.
[0212] It should be noted that the present application does not limit the implementation method of step 52. For example, it can be implemented by any existing or future method that can generate and process a facial image based on the full facial texture.
[0213] For example, in some application scenarios, the above step 52 can be specifically: input the above full face texture into a pre-built full face generation model, and obtain the face description image corresponding to the above speech to be processed output by the full face generation model, so that the face description image can better describe the face state corresponding to the speech to be processed. The full face generation model is used to perform full face generation processing on the input data of the full face generation model; and the present application does not limit the full face generation model.
[0214] Based on the relevant contents of steps 51 to 52 above, it can be known that after determining the second predicted facial expression coefficient and the second predicted facial texture coefficient for the speech to be processed using the above data processing model, the second predicted facial expression coefficient and the second predicted facial texture coefficient can be used to synthesize the full face texture; then the full face texture is input into the full face generation model to obtain the facial description image output by the full face generation model, so that the facial description image can describe the facial state corresponding to the speech to be processed.
[0215] Based on the relevant contents of S301 to S303 above, it can be known that for the image generation method provided in the present application, the speech to be processed and the expression description information corresponding to the speech to be processed are first obtained; then the speech features of the speech to be processed and the expression features of the expression description information are input into a pre-constructed data processing model to obtain the second predicted facial expression coefficient and the second predicted facial texture coefficient output by the data processing model; then, based on the second predicted facial expression coefficient and the second predicted facial texture coefficient, the facial description image corresponding to the speech to be processed is determined, so that the facial description image can represent the facial texture state described by the second predicted facial texture coefficient and the facial expression state described by the second predicted facial expression coefficient, so that the facial description image can better represent the full facial state corresponding to the speech to be processed, which is conducive to improving the image generation effect, thereby helping to better meet the needs of rich and full expressions in the speech image generation scenario.
[0216] In addition, for some application scenarios, these scenarios may have the following requirements: performing speech image generation processing on a specified object so that the finally generated image can describe the facial state corresponding to the speech.
[0217] In order to achieve the requirements shown in the previous paragraph, the present application also provides a possible implementation of the image generation method. In this implementation, the image generation method includes the following steps 61 to 65.
[0218] Step 61: Acquire the speech to be processed, the facial expression description information corresponding to the speech to be processed, and the reference facial image corresponding to the speech to be processed.
[0219] The reference facial image corresponding to the speech to be processed is used to describe the object required to be based on when performing image generation processing on the speech to be processed, such as a digital image.
[0220] In addition, the present application does not limit the method of obtaining the reference facial image corresponding to the speech to be processed, for example, the reference facial image corresponding to the speech to be processed may be an image provided by a user via an image input device. The image input device is used to input an image; and the present application does not limit the implementation method of the image input device.
[0221] In addition, for the relevant contents of the speech to be processed in the above step 61 and the expression description information corresponding to the speech to be processed, please refer to the relevant contents of S301 above.
[0222] Step 62: Input the speech features of the speech to be processed and the expression features of the expression description information corresponding to the speech to be processed into a pre-constructed data processing model to obtain a second predicted facial expression coefficient and a second predicted facial texture coefficient output by the data processing model; the data processing model is constructed using any possible implementation of the model construction method provided in the present application.
[0223] It should be noted that for the relevant content of step 62, please refer to the relevant content of S302 above.
[0224] Step 63: constructing a full face texture corresponding to the speech to be processed according to the second predicted facial expression coefficient and the second predicted facial texture coefficient.
[0225] It should be noted that for the relevant content of step 63, please refer to the relevant content of step 51 above.
[0226] In addition, in some application scenarios, because different objects have different facial texture characteristics, in order to improve the construction effect of the full-face texture, a full-face texture construction model corresponding to each object can be constructed in advance, so that the full-face texture construction models corresponding to different objects are different. Based on this, the present application also provides a possible implementation method of the process of determining the full-face texture corresponding to the speech to be processed above, under which the process of determining the full-face texture corresponding to the speech to be processed can include the following steps 631-633.
[0227] Step 631: Perform object recognition processing on the reference facial image corresponding to the above processed speech to obtain an object recognition result.
[0228] The object recognition result is used to represent the object described by the reference facial image corresponding to the above processed speech; and the present application does not limit the implementation method of the object recognition result. For example, the object recognition result can be implemented using an object identifier.
[0229] In addition, the present application does not limit the implementation method of the above object recognition processing. For example, it can be implemented using any existing or future method that can perform object recognition processing on an image, such as a method for performing object recognition processing with the help of a pre-built object recognition model.
[0230] Step 632: searching for a full face texture construction model that matches the above object recognition result from at least one pre-constructed full face texture construction model as a target model.
[0231] It should be noted that the present application does not limit the implementation method of the above step 632. For example, when the at least one full-face texture construction model pre-constructed above has a pre-labeled object label, the step 632 can be specifically as follows: determine whether the object label of the i-th full-face texture construction model is the same as the above object recognition result. If they are the same, it can be determined that the i-th full-face texture construction model matches the object recognition result, so the i-th full-face texture construction model can be used as the target model, so that the target model can be used to complete the full-face texture construction process related to the object recognition result. Wherein, i is a positive integer, i≤the number of models in the at least one full-face texture construction model.
[0232] Step 633: input the second predicted facial expression coefficient and the second predicted facial texture coefficient into the target model to obtain the full facial texture corresponding to the speech to be processed output by the target model.
[0233] Based on the relevant contents of steps 631 to 633 above, it can be known that in some application scenarios, in order to better improve the construction effect of the full-face texture, the full-face texture construction model corresponding to the object described by the reference facial image can be searched from at least one pre-constructed full-face texture construction model based on the reference facial image corresponding to the speech to be processed above; and then the full-face texture construction model obtained by the search is used to perform full-face texture construction processing on the second predicted facial expression coefficient and the second predicted facial texture coefficient corresponding to the speech to be processed, so as to obtain the full-face texture corresponding to the speech to be processed, so that the full-face texture can better represent the full-face state corresponding to the object described by the reference facial image under the speech to be processed, which is conducive to improving the full-face texture construction effect.
[0234] Step 64: remove the facial description information from the reference facial image corresponding to the speech to be processed to obtain a post-removal image.
[0235] The removed image refers to the image obtained by removing the facial description information from the reference facial image corresponding to the speech to be processed above, so that the removed image only carries other image information from the reference facial image except the facial description information. The facial description information refers to the image information existing in the reference facial image for describing the face of the object.
[0236] In addition, the present application does not limit the implementation method of the above step 64. For example, it can be implemented by using any existing or future method that can remove facial description information from an image.
[0237] Step 65: Perform full face generation processing based on the full face texture and the image after the removal of the text, and obtain a facial description image corresponding to the speech to be processed.
[0238] It should be noted that the present application does not limit the implementation method of the above step 65. For example, it can be specifically: input the above full face texture and the above removed image into a pre-constructed full face generator to obtain a facial description image corresponding to the speech to be processed output by the full face generator, so that the facial state described by the facial description image is consistent with the facial state described by the full face texture, and the other states described by the facial description image except the facial state are consistent with the corresponding states described by the removed image, so as to realize the speech image generation processing under the premise of specifying the object. Among them, the full face generator refers to a pre-constructed model for performing facial image generation processing based on the full face texture and the removed image; and the present application does not limit the implementation method of the full face generator. For example, the full face generator can adopt Figure 2 The full face generator shown is implemented.
[0239] Based on the relevant contents of steps 61 to 65 above, it can be known that in some application scenarios, voice image generation processing can be performed based on the voice to be processed, the expression description information corresponding to the voice to be processed, and the reference facial image corresponding to the voice to be processed to obtain the facial description image corresponding to the voice to be processed. In this way, voice image generation processing can be achieved under the premise of specifying the object, which is beneficial to improving the image generation effect.
[0240] In addition, the image generation method provided by the present application can be applied not only to the field of speech image generation, but also to the field of image video generation. Based on this, the present application also provides a video generation process, which can specifically include the following steps 71 to 75.
[0241] Step 71: After obtaining a long speech with video generation requirements, the long speech is divided into at least one speech segment.
[0242] The long speech refers to the speech that needs to be processed by speech and video generation; and the long speech carries at least one frame of speech information.
[0243] In addition, for the long speech above, in order to ensure the video generation effect, the long speech can be divided into at least one speech segment, so that the video generation process for the long speech can be completed by performing image processing on each speech segment.
[0244] In addition, the present application does not limit the acquisition method of the above at least one speech segment. For example, if the above long speech carries Q frames of speech information, the acquisition process of the at least one speech segment may include: taking the speech segment carrying the first frame of speech information in the long speech as the first speech segment, so that the first speech segment carries the first frame of speech information; taking the speech segment carrying the second frame of speech information in the long speech as the second speech segment, so that the second speech segment carries the second frame of speech information; ... (and so on); taking the speech segment carrying the Qth frame of speech information in the long speech as the Qth speech segment, so that the Qth speech segment carries the Qth frame of speech information. Wherein, Q is a positive integer.
[0245] Step 72: Determine the speech to be processed from the above at least one speech segment.
[0246] It should be noted that the present application does not limit the implementation method of the above step 72. For example, the step 72 may specifically be: randomly selecting a speech segment that has not been traversed from at least one speech segment as the speech to be processed. The speech segment that has not been traversed refers to a speech segment that has not been processed for speech image generation.
[0247] For another example, if the current round of processing is the qth round of speech image generation process, the above step 72 may specifically be: determining the qth speech segment as the speech to be processed, where q is a positive integer, q≤Q.
[0248] Step 73: Determine a facial description image corresponding to the speech to be processed based on the speech to be processed and the facial expression description information corresponding to the speech to be processed.
[0249] It should be noted that the present application does not limit the implementation method of step 73 above. For example, step 73 can be implemented by adopting any processing process provided by the present application that can determine the facial description image corresponding to the speech to be processed based on the speech to be processed and the expression description information corresponding to the speech to be processed, such as the processing process shown in S302-S303 above or the processing process shown in steps 62-65 above.
[0250] Step 74: Determine whether there is any speech segment that has not been traversed in at least one of the above speech segments. If so, return to execute the above step 72 and its subsequent steps; if not, execute the following step 75.
[0251] Step 75: Combine the speech segments in the above long speech and the facial description images corresponding to the speech segments to obtain a video corresponding to the long speech; for any speech segment, the number of frames corresponding to the facial description image corresponding to the speech segment in the video is the same as the number of frames corresponding to the speech segment in the video, and the number of frames corresponding to the speech segment in the video is determined based on the arrangement position of the speech segment in the long speech.
[0252] In the present application, for the above long speech, after obtaining the facial description images corresponding to each speech segment in the long speech, the video corresponding to the long speech can be determined based on each speech segment and the facial description images corresponding to each speech segment, so that the video includes each speech segment and the facial description images corresponding to each speech segment, and the number of frames corresponding to the facial description image corresponding to the qth speech segment in the video is the same as the number of frames corresponding to the qth speech segment in the video, such as the qth frame video information carried by the video includes the image information carried by the facial description image corresponding to the qth speech segment and the voice information carried by the qth speech segment, etc., q is a positive integer, q≤Q, so that the speech video generation process can be realized. Among them, the number of frames corresponding to the qth speech segment in the video is used to describe the arrangement position of the qth speech segment in all the video speech of the video; and the number of frames corresponding to the qth speech segment in the video is determined by the arrangement position of the qth speech segment in the long speech.
[0253] It should be noted that the present application does not limit the correlation between the number of frames corresponding to the qth voice segment in the video and the arrangement position of the qth voice segment in the long voice. For example, in some application scenarios, when the long voice records multiple frames of voice information in a forward order, the number of frames corresponding to the qth voice segment in the video and the arrangement position of the qth voice segment in the long voice are positively correlated. For another example, in some application scenarios, when the long voice records multiple frames of voice information in a reverse order, the number of frames corresponding to the qth voice segment in the video and the arrangement position of the qth voice segment in the long voice are negatively correlated.
[0254] Based on the relevant contents of steps 71 to 75 above, it can be known that in some application scenarios, for a long speech, the facial description images corresponding to each speech segment in the long speech can be generated first; then these speech segments and the facial description images corresponding to these speech segments are combined into a video corresponding to the long speech, so that the video not only carries the speech information described by the long speech, but also carries the image information described by the facial description images corresponding to these speech segments, so that the speech and video generation processing can be realized. Among them, for any speech segment in the long speech, because the facial description image corresponding to the speech segment can represent the facial state corresponding to the speech segment, such as expression, mouth shape, and distribution of facial features, so that the video including these facial description images can carry rich expressions, thereby effectively avoiding the defects caused by all frames of video images in the video generated based on speech using the same expression, which is conducive to improving the video generation effect.
[0255] In addition, the present application does not limit the execution subject of the image generation method provided in the embodiment of the present application. For example, the image generation method provided in the embodiment of the present application can be applied to a terminal device or a server. For another example, the image generation method provided in the embodiment of the present application can also be implemented by means of a data interaction process between a terminal device and a server.
[0256] Based on the model building method provided in the embodiment of the present application, the embodiment of the present application also provides a model building device. Figure 4 Explain and illustrate. Figure 4 This is a schematic diagram of the structure of a model building device provided in an embodiment of the present application. It should be noted that for the technical details of the model building device provided in an embodiment of the present application, please refer to the relevant content of the model building method above.
[0257] like Figure 4 As shown, the model building device 400 provided in the embodiment of the present application includes:
[0258] A first acquisition unit 401 is used to acquire a sample speech, expression description information corresponding to the sample speech, and facial supervision information corresponding to the sample speech, wherein the facial supervision information includes facial expression supervision information and facial texture coefficient supervision information;
[0259] A first processing unit 402 is used to input the speech feature of the sample speech and the expression feature of the expression description information into a data processing model to obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model;
[0260] A first determining unit 403, configured to determine a discrimination result of the predicted facial expression using a discriminator corresponding to the data processing model;
[0261] The model updating unit 404 is used to update the data processing model according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression.
[0262] In one possible implementation, the model updating unit 404 is specifically used to: determine the facial expression prediction loss based on the difference representation data between the predicted facial expression and the facial expression supervision information; determine the facial texture coefficient prediction loss based on the difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information; determine the accuracy representation data of the predicted facial expression based on the discrimination result of the predicted facial expression; determine the model loss of the data processing model based on the facial expression prediction loss, the facial texture coefficient prediction loss and the accuracy representation data of the predicted facial expression; and update the data processing model based on the model loss.
[0263] In a possible implementation manner, the facial supervision information further includes facial expression coefficient supervision information;
[0264] The first processing unit 402 is specifically configured to: input the speech feature of the sample speech and the expression feature of the expression description information into a data processing model to obtain the predicted facial expression, the first predicted facial expression coefficient and the first predicted facial texture coefficient output by the data processing model;
[0265] The model updating unit 404 is specifically used to update the data processing model according to the predicted facial expression, the first predicted facial expression coefficient, the first predicted facial texture coefficient, the facial expression supervision information, the facial expression coefficient supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression.
[0266] In a possible implementation manner, the expression description information corresponding to the sample voice includes the expression category corresponding to the sample voice and the expression intensity of the expression category;
[0267] The model building device 400 further includes:
[0268] An expression encoding unit, used for encoding the expression description information to obtain an expression code, wherein the expression code is used to characterize the intensity state of at least one preset expression category, wherein the at least one preset expression category includes the expression category corresponding to the sample voice;
[0269] A feature determination unit is used to determine the expression feature according to the expression code.
[0270] In a possible implementation manner, the feature determination unit is specifically used to: perform feature embedding processing on the expression code to obtain the expression feature, and the size value of the expression feature in the first dimension is equal to the size value of the voice feature in the first dimension;
[0271] The first processing unit 402 is specifically used to: splice the speech features of the sample speech and the expression features of the expression description information in a second dimension to obtain a first spliced feature, wherein the size value of the first spliced feature in the second dimension is equal to the sum of the size value of the expression feature in the second dimension and the size value of the speech feature in the second dimension, and the second dimension is different from the first dimension; input the first spliced feature into the data processing model to obtain the predicted facial expression and the first predicted facial texture coefficient output by the data processing model.
[0272] In one possible implementation, the expression coding includes intensity normalized representation data of at least one preset expression category; the intensity normalized representation data of the expression category corresponding to the sample voice is obtained by normalizing the expression intensity of the expression category corresponding to the sample voice.
[0273] In one possible implementation, the facial supervision information corresponding to the sample speech is determined based on a reference facial image corresponding to the sample speech; the reference facial image is extracted from a reference video corresponding to the sample speech; the reference video includes the sample speech; the number of frames corresponding to the reference facial image in the reference video is the same as the number of frames corresponding to the sample speech in the reference video.
[0274] In a possible implementation manner, the model updating unit 404 is specifically used to: if a preset model updating condition is met, update the data processing model according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression;
[0275] The model building device 400 further includes:
[0276] A discriminator updating unit is used to update the discriminator according to the discriminant result of the predicted facial expression and the discriminant result of the facial expression supervisory information if a preset discriminator updating condition is met; the discriminant result of the facial expression supervisory information is determined using the discriminator.
[0277] Based on the relevant contents of the above-mentioned model building device 400, it can be known that for the model building device 400 provided in the embodiment of the present application, the sample speech, the expression description information corresponding to the sample speech, and the facial supervision information corresponding to the sample speech are first obtained, and the facial supervision information includes facial expression supervision information and facial texture coefficient supervision information; then, the speech features of the sample speech and the expression features of the expression description information are input into the data processing model to obtain the predicted facial expression and the first predicted facial texture coefficient output by the data processing model; then, the discriminator corresponding to the data processing model is used to determine the discrimination result of the predicted facial expression; finally, according to the predicted facial The facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the discrimination result are used to update the data processing model so that the updated data processing model has better performance, such as better facial expression prediction performance and facial texture coefficient prediction performance, etc., so that when the data processing model is used to perform facial image generation processing on a voice, a better facial description image can be generated, and then the face presented by the facial description image has a more realistic and natural expression, which is beneficial to improving the image generation effect, thereby helping to better meet the needs of rich and full expressions in the voice image generation scenario.
[0278] Based on the image generation method provided in the embodiment of the present application, the embodiment of the present application also provides an image generation device. Figure 5 Explain and illustrate. Figure 5 This is a schematic diagram of the structure of an image generating device provided in an embodiment of the present application. It should be noted that for the technical details of the image generating device provided in an embodiment of the present application, please refer to the relevant contents of the image generating method above.
[0279] like Figure 5 As shown, the image generating device 500 provided in the embodiment of the present application includes:
[0280] The second acquisition unit 501 is used to acquire the speech to be processed and the expression description information corresponding to the speech to be processed;
[0281] The second processing unit 502 is used to input the speech features of the speech to be processed and the expression features of the expression description information into a pre-constructed data processing model to obtain a second predicted facial expression coefficient and a second predicted facial texture coefficient output by the data processing model; the data processing model is constructed by using the model construction method described in any one of claims 1 to 8;
[0282] The second determining unit 503 is used to determine the facial description image corresponding to the speech to be processed according to the second predicted facial expression coefficient and the second predicted facial texture coefficient.
[0283] In a possible implementation manner, the second determining unit 503 includes:
[0284] A texture construction subunit, configured to construct a full face texture corresponding to the speech to be processed according to the second predicted facial expression coefficient and the second predicted facial texture coefficient;
[0285] The image generation subunit is used to generate a facial description image corresponding to the speech to be processed based on the full face texture.
[0286] In a possible implementation manner, the image generating device 500 further includes:
[0287] A third acquisition unit, used to acquire a reference facial image corresponding to the speech to be processed;
[0288] An information removal unit, used for removing facial description information from the reference facial image to obtain a post-removal image;
[0289] The image generation subunit is specifically used to perform full face generation processing based on the full face texture and the removed image to obtain a facial description image corresponding to the speech to be processed.
[0290] In a possible implementation manner, the speech to be processed is any speech segment in a long speech;
[0291] The image generating device 500 further includes:
[0292] The video construction unit is used to combine the speech segments in the long speech and the facial description images corresponding to the speech segments to obtain the video corresponding to the long speech; for any speech segment, the number of frames corresponding to the facial description image corresponding to the speech segment in the video is the same as the number of frames corresponding to the speech segment in the video, and the number of frames corresponding to the speech segment in the video is determined according to the arrangement position of the speech segment in the long speech.
[0293] Based on the relevant contents of the above-mentioned image generating device 500, it can be known that for the image generating device 500 provided in the embodiment of the present application, the speech to be processed and the expression description information corresponding to the speech to be processed are first obtained; then the speech features of the speech to be processed and the expression features of the expression description information are input into a pre-constructed data processing model to obtain the second predicted facial expression coefficient and the second predicted facial texture coefficient output by the data processing model; then, based on the second predicted facial expression coefficient and the second predicted facial texture coefficient, the facial description image corresponding to the speech to be processed is determined, so that the facial description image can represent the facial texture state described by the second predicted facial texture coefficient and the facial expression state described by the second predicted facial expression coefficient, so that the facial description image can better represent the full facial state corresponding to the speech to be processed, which is conducive to improving the image generation effect, thereby helping to better meet the needs of rich and full expressions in the speech image generation scenario.
[0294] In addition, an embodiment of the present application also provides an electronic device, which includes a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device executes any implementation of the model building method provided in the embodiment of the present application, or executes any implementation of the image generation method provided in the embodiment of the present application.
[0295] See also Figure 6 , which shows a schematic diagram of the structure of an electronic device 600 suitable for implementing the embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0296] like Figure 6 As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0297] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0298] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0299] The electronic device provided by the embodiment of the present disclosure and the method provided by the above embodiment belong to the same inventive concept, and the technical details not fully described in this embodiment can be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0300] An embodiment of the present application also provides a computer-readable medium, in which instructions or computer programs are stored. When the instructions or computer programs are executed on a device, the device executes any implementation of the model building method provided in the embodiment of the present application, or executes any implementation of the image generation method provided in the embodiment of the present application.
[0301] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0302] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (Hyper Text Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0303] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0304] The computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device can execute the method.
[0305] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0306] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0307] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit / module does not, in some cases, constitute a limitation on the unit itself.
[0308] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0309] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0310] It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other. For the system or device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.
[0311] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0312] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0313] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0314] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A model building method, It is characterized in that The method comprises: Acquire a sample voice, expression description information corresponding to the sample voice, and facial supervision information corresponding to the sample voice, wherein the facial supervision information includes facial expression supervision information and facial texture coefficient supervision information; Inputting the speech features of the sample speech and the expression features of the expression description information into a data processing model to obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model; Determining a discrimination result of the predicted facial expression using a discriminator corresponding to the data processing model; The data processing model is updated according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression.
2. The method according to claim 1, It is characterized in that The updating of the data processing model according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression comprises: Determining a facial expression prediction loss based on the difference representation data between the predicted facial expression and the facial expression supervision information; Determining a facial texture coefficient prediction loss according to the difference representation data between the first predicted facial texture coefficient and the facial texture coefficient supervision information; Determining accuracy characterization data of the predicted facial expression based on the discrimination result of the predicted facial expression; Determining a model loss of the data processing model according to the facial expression prediction loss, the facial texture coefficient prediction loss, and the accuracy characterization data of the predicted facial expression; The data processing model is updated according to the model loss.
3. The method according to claim 1, It is characterized in that The facial supervision information also includes facial expression coefficient supervision information; The step of inputting the speech feature of the sample speech and the expression feature of the expression description information into a data processing model to obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model comprises: Inputting the speech features of the sample speech and the expression features of the expression description information into a data processing model to obtain the predicted facial expression, the first predicted facial expression coefficient and the first predicted facial texture coefficient output by the data processing model; The updating of the data processing model according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression comprises: The data processing model is updated according to the predicted facial expression, the first predicted facial expression coefficient, the first predicted facial texture coefficient, the facial expression supervision information, the facial expression coefficient supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression.
4. The method according to claim 1, It is characterized in that The expression description information corresponding to the sample speech includes the expression category corresponding to the sample speech and the expression intensity of the expression category; The process of determining the facial expression features includes: Encoding the expression description information to obtain an expression code, wherein the expression code is used to represent the intensity state of at least one preset expression category, and the at least one preset expression category includes the expression category corresponding to the sample voice; The expression feature is determined according to the expression code.
5. The method according to claim 4, It is characterized in that Determining the expression feature according to the expression code includes: Performing feature embedding processing on the expression code to obtain the expression feature, wherein the size value of the expression feature in the first dimension is equal to the size value of the speech feature in the first dimension; The step of inputting the speech feature of the sample speech and the expression feature of the expression description information into a data processing model to obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model comprises: splicing the speech feature of the sample speech and the expression feature of the expression description information in a second dimension to obtain a first splicing feature, wherein the dimension value of the first splicing feature in the second dimension is equal to the sum of the dimension value of the expression feature in the second dimension and the dimension value of the speech feature in the second dimension, and the second dimension is different from the first dimension; The first splicing feature is input into the data processing model to obtain the predicted facial expression and the first predicted facial texture coefficient output by the data processing model.
6. The method according to claim 4, It is characterized in that The expression code includes intensity normalized representation data of at least one preset expression category; The intensity normalized representation data of the expression category corresponding to the sample speech is obtained by normalizing the expression intensity of the expression category corresponding to the sample speech.
7. The method according to claim 1, It is characterized in that The facial supervision information corresponding to the sample speech is determined based on the reference facial image corresponding to the sample speech; The reference facial image is extracted from a reference video corresponding to the sample speech; the reference video includes the sample speech; the number of frames corresponding to the reference facial image in the reference video is the same as the number of frames corresponding to the sample speech in the reference video.
8. The method according to claim 1, It is characterized in that The updating of the data processing model according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression comprises: If the preset model update condition is met, the data processing model is updated according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression; The method further comprises: If the preset discriminator update condition is met, the discriminator is updated according to the discriminant result of the predicted facial expression and the discriminant result of the facial expression supervisory information; the discriminant result of the facial expression supervisory information is determined by using the discriminator.
9. A method for generating an image, It is characterized in that The method comprises: Acquire the speech to be processed and the expression description information corresponding to the speech to be processed; Inputting the speech features of the speech to be processed and the expression features of the expression description information into a pre-constructed data processing model to obtain a second predicted facial expression coefficient and a second predicted facial texture coefficient output by the data processing model; the data processing model is constructed using the model construction method described in any one of claims 1 to 8; A facial description image corresponding to the speech to be processed is determined according to the second predicted facial expression coefficient and the second predicted facial texture coefficient.
10. The method according to claim 9, It is characterized in that The step of determining the facial description image corresponding to the speech to be processed according to the second predicted facial expression coefficient and the second predicted facial texture coefficient comprises: constructing a full face texture corresponding to the speech to be processed according to the second predicted facial expression coefficient and the second predicted facial texture coefficient; A facial description image corresponding to the speech to be processed is generated based on the full face texture.
11. The method according to claim 10, It is characterized in that Before generating the facial description image corresponding to the speech to be processed based on the full face texture, the method further includes: Acquire a reference facial image corresponding to the speech to be processed; Removing facial description information from the reference facial image to obtain a removed image; The step of generating a facial description image corresponding to the speech to be processed based on the full face texture includes: Full face generation processing is performed based on the full face texture and the removed image to obtain a facial description image corresponding to the speech to be processed.
12. The method according to claim 9, It is characterized in that The speech to be processed is any speech segment in a long speech; The method further comprises: The speech segments in the long speech and the facial description images corresponding to the speech segments are combined to obtain a video corresponding to the long speech; for any speech segment, the number of frames corresponding to the facial description image corresponding to the speech segment in the video is the same as the number of frames corresponding to the speech segment in the video, and the number of frames corresponding to the speech segment in the video is determined based on the arrangement position of the speech segment in the long speech.
13. A model building device, It is characterized in that include: A first acquisition unit, used to acquire a sample voice, expression description information corresponding to the sample voice, and facial supervision information corresponding to the sample voice, wherein the facial supervision information includes facial expression supervision information and facial texture coefficient supervision information; A first processing unit, configured to input the speech feature of the sample speech and the expression feature of the expression description information into a data processing model, and obtain a predicted facial expression and a first predicted facial texture coefficient output by the data processing model; a first determining unit, configured to determine a discrimination result of the predicted facial expression using a discriminator corresponding to the data processing model; A model updating unit is used to update the data processing model according to the predicted facial expression, the first predicted facial texture coefficient, the facial expression supervision information, the facial texture coefficient supervision information and the judgment result of the predicted facial expression.
14. An image generating device, It is characterized in that include: A second acquisition unit, used to acquire the speech to be processed and the expression description information corresponding to the speech to be processed; a second processing unit, configured to input the speech features of the speech to be processed and the expression features of the expression description information into a pre-constructed data processing model, and obtain a second predicted facial expression coefficient and a second predicted facial texture coefficient output by the data processing model; the data processing model is constructed by the model construction method according to any one of claims 1 to 8; The second determining unit is used to determine the facial description image corresponding to the speech to be processed according to the second predicted facial expression coefficient and the second predicted facial texture coefficient.
15. An electronic device, It is characterized in that The device comprises: a processor and a memory; The memory is used to store instructions or computer programs; The processor is used to execute the instructions or computer programs in the memory so that the electronic device executes the method according to any one of claims 1 to 12.
16. A computer readable medium, It is characterized in that The computer-readable medium stores instructions or computer programs, and when the instructions or computer programs are executed on a device, the device executes the method according to any one of claims 1 to 12.