Personalized Emoji Generation Method, System, Electronic Device and Readable Medium
By extracting and fusion of expression feature libraries and combining VIT visual processing technology, personalized emoticon packages are generated, which solves the problems of hard expression synthesis and cumbersome user selection in the existing technology, and realizes automated expression action matching and display.
Patent Information
- Application Number
- CN202210018166.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-07
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-01-07
AI Technical Summary
In the prior art, the expression synthesis is stiff and requires the user to choose the expression movement displayed, and lacks the function of automatic recognition and matching.
By receiving the source image with the target face, the preset feature marking model is used to extract the features related to expression actions, generate a semantic marking expression feature library, and generate a personalized emoticon package through the feature fusion model and VIT visual processing technology.
It realizes the automatic display of matching expression action images based on the semantic results in the input content, which improves the personalization and automatic generation capabilities of emoticon packages, and avoids the cumbersome process of manual selection by users.
Smart Images

Figure CN114419177B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and deep learning, and particularly to a personalized meme generation method, system, electronic device and readable medium. Background Art
[0002] With the continuous development of social networks, people's communication has gradually evolved from text-based communication to the use of symbols, images, memes, etc. Since memes can enrich the content of user chats, make up for the dullness of text-based communication and the inaccuracy of expressing meanings, they can better improve communication efficiency. For example, users can express emotions that are difficult to describe in words through the facial expression information in memes. There are already many memes with facial images as materials, but most of these memes are existing meme materials saved by users and cannot be recognized and defined based on facial images, and users need to define the facial images themselves. Summary of the Invention
[0003] Embodiments of the present application provide a personalized meme generation method, system, electronic device and readable medium, which solve the defects in the prior art that the expression synthesis is rigid and the user needs to select the displayed expression actions, and achieve the beneficial effect of being able to display the expression action images matching the semantic results by using the semantic results contained in the input content.
[0004] In a first aspect, an embodiment of the present application provides a personalized meme generation method, and the method includes:
[0005] S100, when a meme tool triggers meme production configuration, receive a plurality of source images with target faces, where each of the target faces has different expression action information;
[0006] S200, use a preset feature marking model to extract features related to expression actions of the target faces in each of the source images, generate an expression feature library related to each expression action, and perform semantic marking on each of the expression feature libraries to obtain a plurality of expression feature libraries with semantic results;
[0007] S300, use a feature fusion model to receive the expression feature information in each of the expression feature libraries, and after VIT visual processing, fuse the plurality of expression feature libraries with semantic results into a personalized meme of the target face, so as to display the corresponding target face expression action image according to the semantic result.
[0008] Further, after step S300, it further includes storing the personalized expression pack of the fused target face, so that when the expression pack tool triggers the expression pack application configuration, after parsing the semantic result related to the personalized expression pack from the input content on the screen, the personalized expression pack of the target face is called from the expression library, and the target face expression action image matching the semantic result is displayed.
[0009] Further, after the step S300, it further includes, when the expression pack tool triggers the expression pack application configuration, receiving the text content input on the screen, parsing one of the semantic results with the personalized expression pack in the text content, triggering the call of the personalized expression pack in the expression library, and displaying the expression action image matching the semantic result.
[0010] Further, in the step S200, the feature marking model uses deep learning technology to extract features related to expression actions, generates an expression feature library related to each expression action, and performs semantic marking related to expression actions on each of the expression feature libraries.
[0011] Further, the feature marking model uses a convolutional neural network for feature extraction, and it includes:
[0012] Performing feature extraction after multiple convolutional downsampling processes on each of the source images;
[0013] Performing key point feature extraction after multiple convolutional downsampling processes on the target face in each of the source images;
[0014] After multiple convolutional downsampling processes on each of the target faces, performing feature contour extraction related to expression actions.
[0015] Further, in step S300, the method for fusing multiple expression feature libraries with semantic results into the personalized expression pack of the target face includes:
[0016] S310, obtaining the expression feature library after multiple convolutional downsampling;
[0017] S320, merging the feature information after each convolutional downsampling process in different expression feature libraries, inputting it into a pre-constructed VIT network architecture, and then performing upsampling to obtain the first fusion feature;
[0018] S330, concatenating the first fusion features of different times by channels, and then performing convolutional upsampling to obtain the second fusion feature;
[0019] S340, concatenating the first fusion feature and the second fusion feature by channels again to obtain the third fusion feature;
[0020] S350, perform multiple convolutional upsamplings on the third fusion feature to obtain a fused personalized emoji package.
[0021] Further, after step S340, it further includes obtaining a target face image of the required facial expression action through convolutional downsampling.
[0022] Further, in the VIT network architecture, receive the merged feature information, perform block processing on the feature information with semantic tags to obtain several feature image blocks;
[0023] Perform linear transformation on each of the feature image blocks, achieve feature image block embedding through dimensionality reduction processing, and after performing spatial position encoding on the feature image block embedding, input them into the Transformer encoder in sequence to perform the fusion of various action expression features.
[0024] In a second aspect, an embodiment of the present application provides a personalized emoji package generation system, which adopts the method described in any one of the first aspects. The system includes:
[0025] An image receiving module, configured to receive a plurality of source images with target faces when the emoji tool triggers the emoji production configuration. Among them, each of the target faces has different facial expression action information;
[0026] A semantic tagging module, configured to use a preset feature tagging model to extract features related to facial expression actions for the target faces in each of the source images, generate an emoji feature library related to each facial expression action, and perform semantic tagging on each of the emoji feature libraries to obtain multiple emoji feature libraries with semantic results;
[0027] An emoji generation module, configured to use a feature fusion model to receive the emoji feature information in each of the emoji feature libraries. After VIT visual processing, fuse the multiple emoji feature libraries with semantic results into a personalized emoji package of the target face, so as to display the corresponding target face facial expression action image according to the semantic result.
[0028] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0029] One or more processors;
[0030] A memory for storing one or more programs;
[0031] When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any one of the first aspects.
[0032] Fourthly, an embodiment of the present application provides a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any one of the first aspects is implemented.
[0033] The technical solutions provided in the embodiments of the present application have at least the following technical effects:
[0034] 1. Compared with the traditional expression image synthesis processing technology, the present application replaces the "rigid" nesting of image copying through a fusion technology. In the traditional synthesis technology, by directly copying a face image on the required expression image or selecting existing expression images with different styles. In this embodiment, after extracting the expression feature image, it is fragmented and then deep learning Vision Transformer (ViT) technology is used to form an integrated personalized expression pack with multiple expression actions, and the expressions required by the display user are automatically generated according to the display requirements.
[0035] 2. The present application combines natural language processing technology. By performing semantic annotation on the expression feature image, the expression feature image is provided with corresponding semantic information and enters the ViT network architecture for learning and training as a one-dimensional input sequence, thus forming an expression feature image with semantic results. Therefore, when the user needs a corresponding expression action image, there is no need to select from numerous expression packs, and the expression action image corresponding to the semantic result can be directly displayed according to the content input on the screen. That is to say, for the made personalized expression pack, only by inputting relevant expression keywords, the corresponding expression action image can be presented. For example, when inputting "pouting", the system returns and generates the "pouting" expression of the user-entered face. Of course, the expression action images in this embodiment can be static images or animated images. The generation configuration of the personalized expression pack in this embodiment can customize the action expression images in the expression pack required for generating the face image, such as making expression images of oneself or collecting some public figures into an expression pack to increase affinity and interest. Description of the Drawings
[0036] Figure 1 It is a flowchart of a method for generating a personalized expression pack in the first embodiment of the present application;
[0037] Figure 2 It is a schematic diagram of the fusion algorithm structure in the first embodiment of the present application;
[0038] Figure 3 It is a module diagram of a personalized expression pack generation system in the second embodiment of the present application. Detailed Embodiments
[0039] In order to better understand the above technical solutions, the above technical solutions will be described in detail below in combination with the accompanying drawings of the specification and specific embodiments.
[0040] Example 1
[0041] Reference appendix Figure 1-2 As shown, the embodiment of the present application provides a personalized emoji generation method, and the method includes:
[0042] Step S100, when the emoji tool triggers the emoji production configuration, receive a plurality of source images with target faces, wherein each of the target faces has different expression action information.
[0043] In this embodiment, the emoji tool in the personalized emoji generation method is loaded on the user terminal, and the user terminal can be a smart phone, a tablet computer, etc. The loading method can be an independent APP application, a WeChat mini-program or a plug-in embedded in an application such as an input method. When the emoji tool triggers the emoji production configuration function, it triggers the opening of the source image reception configuration. The source image can be the picture information saved by the user terminal or the picture information collected by the camera on the current user terminal. For example, when triggering the opening of the emoji production configuration, call the camera on the user terminal to collect images with different expression actions according to the required number of source images, such as expression actions of laughing, smiling, tilting the head, pouting, crying, etc.
[0044] Step S200, use a preset feature marking model to extract features related to the expression actions of the target faces in each of the source images, generate an expression feature library related to each expression action, and perform semantic marking on each of the expression feature libraries to obtain a plurality of expression feature libraries with semantic results.
[0045] In the step S200, the feature marking model uses deep learning technology to extract features related to the expression actions, generate an expression feature library related to each expression action, and perform semantic marking related to the expression actions on each of the expression feature libraries.
[0046] The feature marking model uses a convolutional neural network for feature extraction, and it includes: completing feature extraction after performing multiple convolutional downsampling processes on each of the source images; completing key point feature extraction after performing multiple convolutional downsampling processes on the target faces in each of the source images; after performing multiple convolutional downsampling processes on each of the target faces, completing the extraction of the feature contour related to the expression action.
[0047] Specifically, feature extraction is performed on each of the source images. After multiple convolutional downsampling processes, the processing results of each convolutional downsampling are marked as Si; key-point feature extraction is performed on the target faces in each of the source images. After multiple convolutional downsampling processes, the processing results of each convolutional downsampling are marked as Li; the feature contours related to expression actions in each of the target faces are extracted. After multiple convolutional downsampling processes, the processing results of each convolutional downsampling are marked as ei; an expression feature library including Si, Li, and ei is obtained.
[0048] In this embodiment, the number of convolutional downsampling processes is defined as 5 times, Si = {S0, S1, S2, S3, S4}, Li = {L0, L1, L2, L3, L4}, and ei = {e0, e1, e2, e3, e4}.
[0049] That is to say, five convolutional downsamplings are performed on the source images, and the results of each convolutional downsampling are sequentially marked as S0, S1, S2, S3, S4. 68 key points on the target face are extracted; the eye and mouth regions of the target face are extracted. Taking the expression actions of "pouting, smiling, tilting the head" as an example, five convolutions and downsamplings are performed, and the results of each convolution and downsampling are sequentially marked as L0, L1, L2, L3, L4. The eye and mouth contours of the target face are extracted, and it can be seen the eye and mouth contours of pouting, smiling, and tilting the head. And five convolutions and downsamplings are performed on them, and the results of each convolutional downsampling are sequentially marked as e0, e1, e2, e3, e4.
[0050] Step S300: The feature fusion model receives the expression feature information in each of the expression feature libraries. After VIT visual processing, multiple expression feature libraries with semantic results are fused into a personalized expression pack of the target face, so as to display the corresponding target face expression action image according to the semantic result.
[0051] After step S300, it further includes storing the personalized expression pack of the target face generated by fusion, so that when the expression pack tool triggers the expression pack application configuration, after parsing the semantic result related to the personalized expression pack from the input content on the screen, the personalized expression pack of the target face is called from the expression library, and the target face expression action image matching the semantic result is displayed.
[0052] Further, after step S300, it further includes, when the expression pack tool triggers the expression pack application configuration, receiving the text content input on the screen, parsing out one of the semantic results with the personalized expression pack in the text content, triggering the call of the personalized expression pack in the expression library, and displaying the expression action image matching the semantic result.
[0053] In step S300, the method of fusing multiple expression feature libraries with semantic results into a personalized expression pack for the target face includes:
[0054] Step S310, obtaining the expression feature library after multiple convolutional downsamplings.
[0055] Step S320, merging the feature information after each convolutional downsampling in different expression feature libraries, inputting it into a pre-constructed VIT network architecture, and then performing upsampling to obtain the first fusion feature.
[0056] Step S330, concatenating the first fusion features of different times along the channels, and then performing convolution and upsampling to obtain the second fusion feature.
[0057] S340, concatenating the first fusion feature and the second fusion feature again along the channels to obtain the third fusion feature.
[0058] S350, performing multiple convolutional upsamplings on the third fusion feature to obtain the fused personalized expression pack.
[0059] Further explanation, the method of fusing the features extracted after the above 5 - time convolutional downsampling processing includes: merging {S4, L4, e4}, then inputting it into the VIT network architecture, and then performing upsampling to obtain the first fusion feature E1; merging {S3, L3, e3}, inputting it into the ViT network architecture to obtain the first fusion feature E2; concatenating the first fusion feature E1 and the first fusion feature E2 along the channels, and then performing convolutional upsampling to obtain the second fusion feature E3; merging {S2, L2, e2}, inputting it into the ViT network architecture to obtain the first fusion feature E4, and concatenating the second fusion feature E3 and the first fusion feature E4 along the channels to obtain the third fusion feature; performing multiple convolutional upsamplings on the third fusion feature to obtain the fused personalized expression pack. In this embodiment, three convolutional upsamplings are performed on the third fusion feature. The obtained personalized expression pack can display expression actions such as pouting, smiling, and tilting the head, and only one expression action can be displayed each time.
[0060] After step S340, it further includes obtaining the target face image of the required expression action through convolutional downsampling. That is, in order to apply the personalized expression pack, it is necessary to obtain the target face image of the target expression action through the reverse operation of the fusion operation.
[0061] Further, in the VIT network architecture, the merged feature information is received, the feature information with semantic tags is segmented to obtain a number of feature image patches; each of the feature image patches is linearly transformed, and the feature image patch embedding is achieved through dimensionality reduction processing. After spatial position encoding of the feature image patch embedding, it is input into the Transformer encoder in sequence to fuse various action and expression features.
[0062] In the VIT network architecture, based on the Transformer technology, the image information in the expression feature library is segmented into image patches. While performing linear embedding, position embeddings of the expression feature information are carried out and input into the standard Transformer encoder in sequence. Further explanation, the standard Transformer encoder in this embodiment receives the semantic result token representation of a one-dimensional sequence as input. To process the two-dimensional expression feature image, the expression feature image is deformed into a series of flattened two-dimensional feature image patches, denoted as where (H, W) represents the resolution of the original image, and p 2 represents the resolution of each feature image patch. N = HW / p 2 is the effective sequence length of the VIT network structure. The VIT network architecture uses the same width in all layers, that is, a trainable linear projection maps each vectorized feature image patch to the model dimension, and the corresponding output feature image patch embedding is obtained. It is represented by the following formula:
[0063]
[0064] Before the feature image patch sequence embedding , a deep learning embeddable is added, and its state in the output of the Transformer encoder can be used as the image representation y, as shown in the formula: In the pre-training and fine-tuning stages, the classification head is attached to The Transformer encoder consists of multiple interactive layers of multi-head attention (MSA) and MLP. Layer normalization (LN) is adopted before each feature image patch, and residual connections are applied after each block. The MLP contains two layers presenting GELU non-linearity, as shown in the formula: z` l = MSA(LN(z l-1 )) + z l-1 , l = 1...L; z l = MLP(LN(z` l )) + z` l , l = 1...L represents.
[0065] It can be seen that compared with the traditional synthesis processing technology of expression images, in this embodiment, the fusion technology is used to replace the "rigid" nesting of image copying. In the traditional synthesis technology, the face image is directly copied on the required expression image or existing expression images with different styles are selected. In this embodiment, after extracting the expression feature image, it is fragmented and then deep learning VIT (vision Transformer) technology is used to form a personalized expression pack with multiple expression actions as a whole, and the expressions required by the user are automatically generated according to the display requirements.
[0066] Combined with natural language processing technology, by performing semantic annotation on the expression feature image, the expression feature image is provided with corresponding semantic information and enters the VIT network architecture as a one-dimensional input sequence for learning and training, thus forming an expression feature image with semantic results. Therefore, when the user needs an image of a corresponding expression action, there is no need to select from numerous expression packs, and the expression action image corresponding to the semantic result can be directly displayed according to the content input on the screen. That is to say, for the made personalized expression pack, only by inputting relevant expression keywords, the corresponding expression action image can be presented. For example, when inputting "pouting", the system returns and generates the "pouting" expression of the user's input face. Of course, the expression action images in this embodiment can be static images or animated images. In this embodiment, the generation configuration of the personalized expression pack can customize the action expression images in the required expression pack for the generated face image, such as making expression images of oneself or collecting some public figures to make expression packs, increasing affinity and interest.
[0067] Embodiment 2
[0068] Refer to the appendix Figure 3 As shown, this embodiment provides a personalized expression pack generation system, which adopts the method described in any one of Embodiment 1. The system includes:
[0069] An image receiving module 100, configured to receive a plurality of source images with target faces when the expression pack tool triggers the expression pack production configuration, where each of the target faces has different expression action information.
[0070] A semantic marking module 200, configured to use a preset feature marking model to extract features related to expression actions for the target faces in each of the source images, generate an expression feature library related to each expression action, and perform semantic marking on each of the expression feature libraries to obtain a plurality of expression feature libraries with semantic results.
[0071] The expression generation module 300 is configured to receive the expression feature information in each of the expression feature libraries by using the feature fusion model. After VIT vision processing, it fuses multiple expression feature libraries with semantic results into a personalized expression pack of the target face, so as to display the corresponding target face expression action image according to the semantic results.
[0072] Embodiment III
[0073] This embodiment provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any one of Embodiment I.
[0074] This embodiment provides a computer-readable medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method described in any one of Embodiment I.
[0075] Those skilled in the art should understand that the embodiments of the present invention may be provided as a method, a system, or a computer program product. Therefore, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0076] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0077] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0078] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for realizing the processing in the process Figure 1 one process or multiple processes and / or blocks Figure 1 steps for realizing the functions specified in one block or multiple blocks.
[0079] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0080] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A personalized emoji generation method, characterized in that, The method includes: S100, when the emoji tool triggers the emoji production configuration, receiving a plurality of source images with target faces, wherein each of the target faces has different expression action information; S200, using a preset feature marking model to extract features related to expression actions from the target faces in each of the source images, generating an expression feature library related to each expression action, and performing semantic marking on each of the expression feature libraries to obtain a plurality of expression feature libraries with semantic results; wherein, the feature marking model uses deep learning technology to extract features related to expression actions, generates an expression feature library related to each expression action, and performs semantic marking related to expression actions on each of the expression feature libraries; S300, using a feature fusion model to receive the expression feature information in each of the expression feature libraries, and after VIT visual processing, fusing a plurality of expression feature libraries with semantic results into a personalized emoji of the target face, so as to display the corresponding target face expression action image according to the semantic result; In step S300, the method of fusing a plurality of expression feature libraries with semantic results into a personalized emoji of the target face includes: S310, obtaining the expression feature library after multiple convolutional downsamplings; S320, combining the feature information after each convolutional downsampling processing in different expression feature libraries, inputting it into a pre-constructed VIT network architecture, and then performing upsampling to obtain a first fusion feature; in the VIT network architecture, receiving the combined feature information, performing block processing on the feature information with semantic marking to obtain a plurality of feature image blocks; performing linear transformation on each of the feature image blocks, realizing feature image block embedding through dimensionality reduction processing, and after performing spatial position encoding on the feature image block embedding, inputting it into the Transformer encoder in sequence to perform fusion of various action expression features; S330, concatenating the first fusion features of different times by channels, and then performing convolutional upsampling to obtain a second fusion feature; S340, concatenating the first fusion feature and the second fusion feature again by channels to obtain a third fusion feature; S350, performing multiple convolutional upsamplings on the third fusion feature to obtain a fused personalized emoji.
2. The personalized emoji generation method according to claim 1, characterized in that, After step S300, it further includes storing the fused personalized emoji of the target face, so that when the emoji tool triggers the emoji application configuration, after parsing the semantic result related to the personalized emoji from the on-screen input content, calling the personalized emoji of the target face from the emoji library, and displaying the target face expression action image matching the semantic result.
3. The personalized emoji generation method according to claim 1, characterized in that, After the step S300, it further includes, when the emoji tool triggers the emoji application configuration, receiving the text content input on the screen, parsing one of the semantic results with the personalized emoji in the text content, triggering the call of the personalized emoji in the emoji library, and displaying the expression action image matching the semantic result.
4. The personalized emoji generation method according to claim 1, characterized in that, The feature marking model uses a convolutional neural network for feature extraction, and it includes: Feature extraction is completed after multiple convolutional downsampling processes are performed on each of the source images; Key-point feature extraction is completed after multiple convolutional downsampling processes are performed on the target faces in each of the source images; After multiple convolutional downsampling processes are performed on each of the target faces, feature contour extraction related to expression actions is completed.
5. The personalized emoji generation method according to claim 4, characterized in that, After step S340, it further includes obtaining a target face image of the required expression action through convolutional downsampling.
6. A personalized emoji generation system, adopting the method according to any one of claims 1 - 5, characterized in that, The system includes: An image receiving module configured to receive a plurality of source images with target faces when a meme tool triggers meme production configuration, wherein each of the target faces carries different expression action information; A semantic marking module configured to use a preset feature marking model to perform feature extraction related to expression actions on the target faces in each of the source images, generate an expression feature library related to each expression action, and perform semantic marking on each of the expression feature libraries to obtain multiple expression feature libraries with semantic results; An expression generation module configured to use a feature fusion model to receive the expression feature information in each of the expression feature libraries, and after VIT vision processing, fuse the multiple expression feature libraries with semantic results into a personalized meme of the target face, so as to display the corresponding target face expression action image according to the semantic result.
7. An electronic device, characterized in that, It includes: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method according to any one of claims 1-5.
8. A computer-readable medium, on which a computer program is stored, characterized in that, The computer program, when executed by a processor, implements the method according to any one of claims 1-5.
Citation Information
Patent Citations
Expression package automatic generation method and device, computer equipment and storage medium
CN110458916A
Image processing method and device, electronic equipment and storage medium
CN111597926A
Facial expression determination method, expression parameter determination model, medium and equipment
CN112614213A