Model training method and device, control method and device, electronic equipment and computer readable storage medium
By combining deformation prediction of facial images with training of rendering image loss values and optimizing prediction model parameters, the problem of low accuracy of facial deformation prediction is solved, and high-precision facial deformation prediction is achieved.
Patent Information
- Application Number
- CN202510699427.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-05-27
AI Technical Summary
In the existing technology, the accuracy of facial deformation prediction is not high. Traditional methods are difficult to adapt to complex scenes and cannot accurately predict facial deformation of face images.
By performing deformation prediction on multiple facial sample image frames, determining the loss value of the predicted rendered image and the target rendered image, combining geometric deformation and visual rendering dimensions for model training, updating the parameters of the prediction model, and optimizing the prediction ability of continuous facial movements or expression changes.
The accuracy of facial deformation prediction of face images is improved, and high-precision prediction of continuous facial movements or expression changes is achieved.
Smart Images

Figure CN120673222A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a model training method, a control method, an apparatus, an electronic device, and a computer-readable storage medium. Background Art
[0002] With the rapid development of computer vision, graphics, and deep learning, the demand for accurate facial modeling and dynamic deformation prediction has increased dramatically. Traditional methods rely on manual modeling or physical simulation, which are inefficient and difficult to adapt to complex scenarios. Modern technologies (such as deep learning and 3D reconstruction) provide new possibilities for automated, high-precision facial deformation prediction. Driven by both technological innovation and industrial demand, research on facial deformation prediction is becoming a core area of computer vision and graphics. Its significance lies not only in improving the realism and interactive efficiency of virtual content, but also in creating far-reaching value in fields such as medicine, security, and culture.
[0003] In related technologies, the three-dimensional expression of human faces is usually based on vertex or three-dimensional deformable model methods, which cannot accurately predict facial deformation of facial images, resulting in low accuracy of facial deformation prediction. Summary of the Invention
[0004] The embodiments of the present application provide a model training method, a control method, an apparatus, an electronic device, and a computer-readable storage medium that can accurately predict facial deformation of facial images.
[0005] The technical solution of the embodiment of the present application is implemented as follows:
[0006] The present invention provides a model training method, which includes:
[0007] Using a prediction model, deformation prediction is performed on a plurality of first facial sample image frames to obtain predicted deformation information for each of the first facial sample image frames, wherein the plurality of first facial sample image frames are collected when the face of a target object is in motion, and sample labels of the first facial sample image frames include target deformation information and a target rendered image;
[0008] Performing rendering prediction based on the first facial sample image frame to obtain a predicted rendered image;
[0009] determining a first loss value based on the predicted deformation information and the target deformation information, and determining a second loss value based on the predicted rendered image and the target rendered image;
[0010] determining a third loss value based on the predicted deformation information of the plurality of first facial sample image frames and the target deformation information of the plurality of first facial sample image frames;
[0011] Based on the first loss value, the second loss value and the third loss value, the model parameters of the prediction model are updated to obtain a trained prediction model.
[0012] An embodiment of the present application provides a control method, the method comprising:
[0013] receiving a plurality of facial image frames to be processed, wherein the plurality of facial image frames to be processed are collected when the face of the target object is in motion;
[0014] Calling a trained prediction model to perform deformation prediction on each of the facial image frames to be processed to obtain predicted deformation information for each of the facial image frames to be processed, wherein the trained prediction model is trained using the model training method provided in an embodiment of the present application, and the predicted deformation information includes displacement data of multiple facial vertices;
[0015] For each of the facial image frames to be processed, the bionic device is controlled to drive the plurality of facial vertices to move based on the plurality of displacement data.
[0016] The present invention provides a model training device, comprising:
[0017] a deformation prediction module, configured to perform deformation prediction on a plurality of first facial sample image frames using a prediction model, and obtain corresponding predicted deformation information for each of the first facial sample image frames, wherein the plurality of first facial sample image frames are collected when the face of a target object is in motion, and the sample labels of the first facial sample image frames include target deformation information and a target rendered image;
[0018] A rendering prediction module, configured to perform rendering prediction based on the first facial sample image frame to obtain a predicted rendered image;
[0019] A first determination module is configured to determine a first loss value based on the predicted deformation information and the target deformation information, and to determine a second loss value based on the predicted rendered image and the target rendered image;
[0020] a second determining module, configured to determine a third loss value based on the predicted deformation information of the plurality of first facial sample image frames and the target deformation information of the plurality of first facial sample image frames;
[0021] A model updating module is used to update the model parameters of the prediction model based on the first loss value, the second loss value and the third loss value to obtain a trained prediction model.
[0022] An embodiment of the present application provides a control device, including:
[0023] An image acquisition module is configured to receive a plurality of facial image frames to be processed, wherein the plurality of facial image frames to be processed are acquired when the face of the target object is in motion;
[0024] a deformation prediction module, configured to call a trained prediction model to perform deformation prediction on each of the facial image frames to be processed, thereby obtaining predicted deformation information for each of the facial image frames to be processed, wherein the trained prediction model is obtained by training using the model training method according to any one of claims 1 to 9, and the predicted deformation information includes displacement data of a plurality of facial vertices;
[0025] The facial movement module is used to control the bionic device to drive the plurality of facial vertices to move based on the plurality of displacement data for each facial image frame to be processed.
[0026] An embodiment of the present application provides an electronic device, comprising:
[0027] a memory for storing computer-executable instructions or computer programs;
[0028] The processor is used to implement the model training method or control method provided in the embodiment of the present application when executing the computer-executable instructions or computer programs stored in the memory.
[0029] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions or a computer program for implementing the model training method or control method provided in an embodiment of the present application when executed by a processor.
[0030] An embodiment of the present application provides a computer program product, including computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the model training method or the control method provided in the embodiment of the present application is implemented.
[0031] The embodiments of the present application have the following beneficial effects:
[0032] By applying an embodiment of the present application, a prediction model is used to perform deformation prediction on multiple first facial sample image frames, and corresponding predicted deformation information of each first facial sample image frame is obtained. The multiple first facial sample image frames are collected when the face of the target object moves. The sample labels of the first facial sample image frames include target deformation information and target rendering images. Then, rendering prediction is performed based on the first facial sample image frames to obtain a predicted rendering image. Then, based on the predicted deformation information and the target deformation information, a first loss value is determined, and based on the predicted deformation information and the target deformation information, a second loss value is determined. The prediction model can be trained in combination with the geometric deformation prediction dimension and the visual rendering dimension. Based on the predicted deformation information of the multiple first facial sample image frames and the target deformation information of the multiple first facial sample image frames, a third loss value is determined to achieve joint supervision of temporal dynamic deformation, optimize the prediction model's prediction ability for continuous facial movements or expression changes, and then based on the first loss value, the second loss value and the third loss value, update the model parameters of the prediction model to obtain a trained prediction model. In this way, the loss value between the predicted deformation information and the target deformation information, the loss value between the predicted rendered image and the target rendered image, and the loss value between the predicted deformation information and the target deformation information of multiple first facial sample image frames are comprehensively utilized to update the model parameters of the prediction model, thereby improving the accuracy of facial deformation prediction of face images using the prediction model. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 This is a schematic diagram of an application mode of the model training method provided in an embodiment of the present application;
[0034] Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application;
[0035] Figure 3A This is a first flow chart of the model training method provided in an embodiment of the present application;
[0036] Figure 3B This is a second flow chart of the model training method provided in an embodiment of the present application;
[0037] Figure 3C This is a third flow chart of the model training method provided in an embodiment of the present application;
[0038] Figure 3D Schematic diagram of the training process of the prediction model provided in the embodiment of the present application;
[0039] Figure 3E It is a flowchart of the control method provided in the embodiment of the present application;
[0040] Figure 4 This is a schematic diagram of the prediction model structure provided by the embodiment of the present application;
[0041] Figure 5 This is a structural diagram of the prediction model training process guided by the rendering provided in an embodiment of the present application.
[0042] It should be pointed out that the above-mentioned "first" and "second" are only used to distinguish different solutions, and do not represent the degree of distinction between the advantages and disadvantages of the solutions or the priority in the implementation process. DETAILED DESCRIPTION
[0043] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0044] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0045] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0046] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations when applied in examples, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.
[0047] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0048] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0049] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0050] 1) BlendShape: Also known as a deformation target or target shape, it includes the displacement data of multiple facial vertices.
[0051] 2) Rendering image: refers to the conversion of an abstract three-dimensional face model into a visual two-dimensional image through computer algorithms or physical simulation technology.
[0052] The embodiments of the present application provide a model training method, a control method, an apparatus, an electronic device, and a computer-readable storage medium that can accurately predict facial deformation of facial images.
[0053] The following describes exemplary applications of the electronic devices provided in the embodiments of the present application. The electronic devices provided in the embodiments of the present application can be implemented as various types of terminals, such as laptops, tablet computers, desktop computers, set-top boxes, smartphones, smart speakers, smart watches, smart TVs, and in-vehicle terminals. They can also be implemented as servers. The following describes exemplary applications when the device is implemented as a server.
[0054] See also Figure 1 , Figure 1 This is a schematic diagram of the application mode of the model training method provided in the embodiment of the present application, for example, Figure 1 The server 200, the network 300 and the terminal 400 are involved. The terminal 400 is connected to the server 200 via the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.
[0055] During the model training process, the user triggers a model training instruction for the prediction model on the terminal 400. In response to the user's model training instruction, the terminal 400 sends a model training request to the server 200. In response to the model training request, the server 200 obtains multiple first facial sample image frames collected by the terminal 400, and uses the prediction model to perform deformation prediction on the multiple first facial sample image frames, and obtains predicted deformation information of each first facial sample image frame. The multiple first facial sample image frames are collected when the face of the target object moves. The sample labels of the first facial sample image frames include target deformation information and target rendered images; rendering prediction is performed based on the first facial sample image frames to obtain a predicted rendered image; a first loss value is determined based on the predicted deformation information and the target deformation information, and a second loss value is determined based on the predicted rendered image and the target rendered image; a third loss value is determined based on the predicted deformation information of the multiple first facial sample image frames and the target deformation information of the multiple first facial sample image frames; based on the first loss value, the second loss value and the third loss value, the model parameters of the prediction model are updated to obtain a trained prediction model.
[0056] When performing deformation prediction, the server 200 receives multiple facial image frames to be processed, which are collected when the target object's face moves; calls the trained prediction model to perform deformation prediction on each facial image frame to be processed, and obtains predicted deformation information of each facial image frame to be processed, and the predicted deformation information includes displacement data of multiple facial vertices; and sends the predicted deformation information of each facial image frame to be processed to the terminal 400, so that the terminal 400 generates a control instruction based on the predicted deformation information, and the terminal 400 executes the control instruction, thereby driving the multiple facial vertices to move based on the multiple displacement data. The terminal 400 can be a bionic device.
[0057] In some embodiments, the server (e.g., server 200) may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal 400 may be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, car terminal, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.
[0058] See also Figure 2 , Figure 2is a structural diagram of an electronic device provided in an embodiment of the present application, which may be a terminal or a server. Figure 2 The electronic device shown includes: at least one processor 410, a memory 450, and at least one network interface 420. The various components in the electronic device are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 440 is not described in detail. Figure 2 Various buses are labeled as bus system 440 .
[0059] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0060] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.
[0061] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0062] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0063] The operating system 451 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., which are used to implement various basic businesses and process hardware-based tasks.
[0064] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420 . Exemplary network interfaces 420 include Bluetooth, Wireless LAN (WiFi), and Universal Serial Bus (USB).
[0065] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2 A model training device 455 stored in the memory 450 is shown, which can be software in the form of programs and plug-ins, including the following software modules: a deformation prediction module 4551, a rendering prediction module 4552, a first determination module 4553, a second determination module 4554, and a model update module 4555. These modules are logical, and therefore can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be explained below.
[0066] The model training method provided in the embodiment of the present application will be explained in combination with the exemplary application and implementation of the server device provided in the embodiment of the present application.
[0067] The following describes the model training method provided in the embodiments of the present application. For ease of understanding, the model training method provided in the embodiments of the present application is described using the example of facial deformation prediction for human face images. In actual applications, the model training method and control method provided in the embodiments of the present application can also be applied to facial deformation prediction scenarios such as animal face images and cartoon face images.
[0068] As mentioned above, the electronic device that implements the model training method of the embodiment of the present application can be a terminal, a server, or a combination of the two. Next, the model training method provided by the embodiment of the present application is described by taking the electronic device as an example of a server. Figure 3A , Figure 3A This is a first flow chart of the model training method provided in the embodiment of the present application, which will be combined with Figure 3A The steps shown are explained.
[0069] In step 301, deformation prediction is performed on a plurality of first facial sample image frames using a prediction model to obtain corresponding predicted deformation information of each first facial sample image frame.
[0070] Here, the prediction model can be built based on a lightweight convolutional neural network (ShuffleNetV2). Multiple first facial sample image frames are collected during facial motion of the target subject. The bionic device's eye camera can be used to capture multiple facial motion images of the target subject, and each facial motion image is cropped to obtain multiple first facial sample image frames. Deformation prediction involves performing facial deformation (BlendShape) prediction on the multiple first facial sample image frames. The predicted deformation information is 52-dimensional deformation information, i.e., 52 independent facial deformation information. Each facial deformation information corresponds to a specific facial muscle movement or expression unit, and each facial deformation information includes displacement data for multiple facial vertices. The sample labels of the first facial sample image frames include target deformation information and target rendered images. The target deformation information is the actual 52-dimensional deformation information corresponding to the first facial sample image frames, and the target rendered image is obtained by rendering a three-dimensional face model using the target deformation information.
[0071] In some embodiments, a prediction model is used to perform deformation prediction on multiple first facial sample image frames, and the corresponding predicted deformation information of each first facial sample image frame is obtained. This can be achieved by the following steps: for each first facial sample image frame, convolution processing is performed on the first facial sample image frame to obtain a second facial sample image frame; convolution processing and feature enhancement processing are performed on the second facial sample image frame to obtain a third facial sample image frame; convolution processing and feature enhancement processing are performed on the third facial sample image frame to obtain a fourth facial sample image frame; convolution processing is performed on the fourth facial sample image frame to obtain a fifth facial sample image frame; pooling processing is performed on the fifth facial sample image frame to obtain predicted features of the first facial sample image frame; linear processing is performed on the predicted features of the first facial sample image frame to obtain predicted deformation information of the first facial sample image frame.
[0072] Here, for each first facial sample image frame, the convolution module (conv3*3) in the prediction model is used to perform a convolution process on the first facial sample image frame once to obtain a second facial sample image frame. The convolution module (conv3*3) and feature enhancement module (bottleneck) in the prediction model are used to perform a convolution process on the second facial sample image frame once and a feature enhancement process five times, respectively, to obtain a third facial sample image frame. The convolution module (conv3*3) and feature enhancement module (bottleneck) in the prediction model are used to perform a convolution process on the third facial sample image frame once and a feature enhancement process seven times, respectively, to obtain a fourth facial sample image frame. The convolution module (conv3*3) in the prediction model is used to perform a convolution process on the fourth facial sample image frame once to obtain a fifth facial sample image frame. The global mean pooling module in the prediction model is used to perform a pooling process on the fifth facial sample image frame once to obtain the predicted features of the first facial sample image frame. The fully connected module in the prediction model is used to perform quadratic linear processing on the predicted features of the first facial sample image frame to obtain predicted deformation information of the first facial sample image frame, that is, to obtain predicted 52-dimensional deformation information of the first facial sample image frame.
[0073] For example, taking a first facial sample image frame of 112*112*3 (width 112, height 112, number of channels 3), convolution processing is performed on the first facial sample image frame to obtain a second facial sample image frame, which is 56*56*64. Convolution processing and feature enhancement processing are performed on the second facial sample image frame to obtain a third facial sample image frame, which is 28*28*64. Convolution processing and feature enhancement processing are performed on the third facial sample image frame to obtain a fourth facial sample image frame, which is 14*14*16. Convolution processing is performed on the fourth facial sample image frame to obtain a fifth facial sample image frame, which is 7*7*32. Pooling processing is performed on the fifth facial sample image frame to obtain predicted features of the first facial sample image frame, which are represented as features of 32 channels. Linear processing is performed on the predicted features of the first facial sample image frame to obtain predicted deformation information of the first facial sample image frame, where the predicted deformation information of the first facial sample image frame is represented as features of 52 channels.
[0074] In an embodiment of the present application, a series of convolution processing, feature enhancement processing, pooling processing and linear processing are performed on the first facial sample image frame to obtain the predicted features of the first facial sample image frame, and the predicted features of the first facial sample image frame are linearly processed to obtain the predicted deformation information of the first facial sample image frame. Key feature information can be gradually extracted from the original facial sample image frame, thereby achieving high-precision facial deformation prediction of the prediction model.
[0075] Continue to refer Figure 3A In step 302, rendering prediction is performed based on the first facial sample image frame to obtain a predicted rendered image.
[0076] Here, the first facial sample image frame is subjected to the aforementioned convolution, feature enhancement, pooling, and linear processing to sequentially obtain a second facial sample image frame, a third facial sample image frame, a fourth facial sample image frame, and a fifth facial sample image frame. Based on the first facial sample image frame, the second facial sample image frame, the third facial sample image frame, the fourth facial sample image frame, and the fifth facial sample image frame, a rendered image of the first facial sample image frame is predicted to obtain a predicted rendered image.
[0077] In some embodiments, see Figure 3B , Figure 3B This is a second flow chart of the model training method provided in the embodiment of the present application. Figure 3A Step 302 shown can be performed by Figure 3B Steps 3021 to 3026 are implemented as described below.
[0078] In step 3021, feature extraction is performed on the fifth facial sample image frame to obtain a first sample feature, and feature extraction is performed on the fourth facial sample image frame to obtain a second sample feature.
[0079] Here, the convolution module (conv3*3) in the prediction model is used to extract features from the fifth facial sample image frame to obtain a first sample feature. The convolution module (conv3*3) in the prediction model is used to extract features from the fourth facial sample image frame to obtain a second sample feature. For example, the fifth facial sample image frame is 7*7*32, and feature extraction is performed on the fifth facial sample image frame to obtain a first sample feature, which is 7*7*16. The fourth facial sample image frame is 14*14*16, and feature extraction is performed on the fourth facial sample image frame to obtain a second sample feature, which is 14*14*8.
[0080] In step 3022, the first sample feature and the second sample feature are spliced to obtain a first spliced feature, and the first spliced feature is convolved to obtain a first convolution result.
[0081] Here, the first sample feature and the second sample feature are subjected to feature splicing processing to obtain a first splicing feature. The convolution module (conv3*3) in the prediction model is used to perform convolution processing on the first splicing feature to obtain a first convolution result, so that the dimension of the first convolution result is the same as the dimension of the fourth facial sample image frame. Continuing with the above embodiment, the first sample feature is 7*7*16, the second sample feature is 14*14*8, the first sample feature and the second sample feature are spliced to obtain a first splicing feature, the first splicing feature is 14*14*32, the first splicing feature is subjected to convolution processing to obtain a first convolution result, the first convolution result is 14*14*16, and the dimension of the first convolution result is the same as the dimension of the fourth facial sample image frame (14*14*16).
[0082] In step 3023, feature extraction is performed on the third facial sample image frame to obtain a third sample feature, and the third sample feature and the first convolution result are spliced to obtain a second spliced feature, and the second spliced feature is convolved to obtain a second convolution result.
[0083] Here, the convolution module (conv3*3) in the prediction model is used to extract features from the third facial sample image frame to obtain a third sample feature. The third sample feature and the first convolution result are subjected to feature splicing processing to obtain a second splicing feature, and then the convolution module (conv3*3) in the prediction model is used to convolve the second splicing feature to obtain a second convolution result, so that the dimension of the second convolution result is the same as the dimension of the third facial sample image frame. Continuing with the above embodiment, the third facial sample image frame is 28*28*64, and feature extraction is performed on the third facial sample image frame to obtain a third sample feature of 28*28*32. The first convolution result is 14*14*16, and the third sample feature and the first convolution result are spliced to obtain a second splicing feature of 28*28*128, and the second splicing feature is convolved to obtain a second convolution result of 28*28*64.
[0084] In step 3024, feature extraction is performed on the second facial sample image frame to obtain a fourth sample feature, and the fourth sample feature and the second convolution result are spliced to obtain a third spliced feature, and the third spliced feature is convolved to obtain a third convolution result.
[0085] Here, the convolution module (conv3*3) in the prediction model is used to extract features from the second facial sample image frame to obtain a fourth sample feature. The fourth sample feature and the second convolution result are subjected to feature splicing processing to obtain a third splicing feature, and then the convolution module (conv3*3) in the prediction model is used to convolve the third splicing feature to obtain a third convolution result, so that the dimension of the third convolution result is the same as the dimension of the second facial sample image frame. Continuing from the above embodiment, the second facial sample image frame is 56*56*64, and feature extraction is performed on the second facial sample image frame to obtain a fourth sample feature of 56*56*32. The second convolution result is 28*28*64, and the fourth sample feature and the second convolution result are spliced to obtain a third splicing feature of 56*56*128, and the third splicing feature is convolved to obtain a third convolution result of 56*56*64.
[0086] In step 3025, feature extraction is performed on the first facial sample image frame to obtain a fifth sample feature, and the fifth sample feature and the third convolution result are spliced to obtain a fourth spliced feature.
[0087] Here, the convolution module (conv3*3) in the prediction model is used to extract features from the first facial sample image frame to obtain the fifth sample feature. The fifth sample feature is then concatenated with the third convolution result to obtain the fourth concatenated feature. Continuing with the above embodiment, the first facial sample image frame is 112*112*3. Feature extraction is performed on the first facial sample image frame to obtain the fifth sample feature of 112*112*64. The third convolution result is 56*56*64. The fifth sample feature is concatenated with the third convolution result to obtain the fourth concatenated feature of 112*112*128.
[0088] In step 3026, rendering prediction is performed based on the fourth stitching feature to obtain a predicted rendered image.
[0089] Here, the fourth spliced feature is linearly transformed, and the convolution module (conv1*1) in the prediction model is used to adjust the channel dimension to obtain the processed feature, which is then mapped to the target image space (such as RGB channel) to obtain the predicted rendered image.
[0090] In an embodiment of the present application, the sample features of the first facial sample image frame, the second facial sample image frame, the third facial sample image frame, the fourth facial sample image frame and the fifth facial sample image frame are spliced to obtain a fourth spliced feature, and rendering prediction is performed based on the fourth spliced feature to obtain a predicted rendered image, which can improve the detail richness and noise robustness of the predicted rendered image.
[0091] Continue to refer Figure 3AIn step 303, a first loss value is determined based on the predicted deformation information and the target deformation information, and a second loss value is determined based on the predicted rendered image and the target rendered image.
[0092] Here, the loss value is calculated for the predicted deformation information and the target deformation information of the first sample image frame to obtain a first loss value. The loss value between the predicted rendered image and the real rendered image is calculated using the smooth absolute value loss function (L1_Smooth) to obtain a second loss value.
[0093] In some embodiments, determining the first loss value based on the predicted deformation information and the target deformation information can be achieved by the following steps: determining the first error value between the predicted deformation information and the target deformation information, when the absolute value of the first error value is less than the first preset value, performing a logarithmic conversion on the first error value to obtain a first logarithmic value, and determining the product of the first preset value and the first logarithmic value as the first loss value; when the absolute value of the first error value is greater than or equal to the first preset value, determining the difference between the absolute value of the first error value and the second preset value as the first loss value.
[0094] Here, the positional error between corresponding facial vertices in the predicted deformation information and the target deformation information is determined as a first error value between the predicted deformation information and the target deformation information. When the absolute value of the first error value is less than a first preset value, the first error value is divided by a preset parameter, the resulting quotient is logarithmically transformed to obtain a first logarithmic value, and the product of the first preset value and the first logarithmic value is determined as a first loss value. In this case, the first loss value can be expressed as wln(1+|x| / ∈), where x represents the first error value, w represents the first preset value, and ∈ represents the preset parameter.
[0095] When the absolute value of the first error value is greater than or equal to the first preset value, the absolute value of the first error value is subtracted from the second preset value, and the obtained difference is determined as the first loss value. In this case, the first loss value can be expressed as |x|-C, where x represents the first error value and C represents the second preset value. For example, the first loss value can be determined by the following formula (1):
[0096]
[0097] Wherein, WingLoss represents the first loss value, x represents the first error value, w represents the first preset value, ∈ represents the preset parameter, and C represents the second preset value.
[0098] In an embodiment of the present application, a first error value between the predicted deformation information and the target deformation information is determined, and a first loss value of the prediction model is determined based on the absolute value of the first error value and the size of the first preset value. By combining the segmented loss design of the error absolute value and the preset threshold, the prediction model can ensure efficient optimization of small errors and suppress the negative impact of large errors.
[0099] In some embodiments, determining the second preset value can be achieved by the following steps: performing logarithmic conversion on the first preset value to obtain a second logarithmic value, and determining the product of the second logarithmic value and the first preset value as the first parameter value; determining the difference between the first parameter value and the first preset value as the second preset value.
[0100] Here, the first preset value is divided by the preset parameter, the resulting quotient is logarithmically converted to obtain a second logarithmic value, the second logarithmic value is multiplied by the first preset value, and the resulting product is determined as the first parameter value. The first parameter value is subtracted from the first preset value, and the resulting difference is determined as the second preset value. For example, the second preset value can be expressed as w-wln(1+w / ∈), where w represents the first preset value and ∈ represents the preset parameter.
[0101] In some embodiments, determining the second loss value based on the predicted rendered image and the target rendered image can be achieved by the following steps: determining a second error value between the predicted rendered image and the target rendered image; determining the second error value as the input value of a preset smooth absolute value loss function, determining the output value of the smooth absolute value loss function; and determining the output value as the second loss value.
[0102] Here, the difference between the pixel values in the predicted rendered image and the target rendered image is determined as the second error value, and the second error value is input into the preset smooth absolute value loss function (L1_Smooth), and the output value of the smooth absolute value loss function is determined as the second loss value.
[0103] In an embodiment of the present application, a second error value between the predicted rendered image and the target rendered image is determined, and the second error value is input into a preset smoothed absolute value loss function to obtain a second loss value of the prediction model, which can take into account both accuracy and noise resistance, thereby improving the stability and quality of the prediction model for rendering image prediction.
[0104] Continue to refer Figure 3A In step 304, a third loss value is determined based on the predicted deformation information of the plurality of first facial sample image frames and the target deformation information of the plurality of first facial sample image frames.
[0105] In some embodiments, see Figure 3C , Figure 3CThis is a third flow chart of the model training method provided in the embodiment of the present application. Figure 3A Step 304 shown can be performed by Figure 3C Steps 3041 to 3044 are implemented as described below.
[0106] In step 3041, difference constraint is performed based on the predicted deformation information of the i-1th first facial sample image frame, the predicted deformation information of the i-th first facial sample image frame, the target deformation information of the i-1th first facial sample image frame, and the target deformation information of the i-th first facial sample image frame to obtain the i-th first constraint value.
[0107] Here, i is an integer from 1 to N-1, where N is the total number of first facial sample image frames. Two adjacent first facial sample image frames, namely, the i-1th first facial sample image frame and the i-th first facial sample image frame, are determined from the plurality of first facial sample image frames. A difference constraint is applied to the predicted deformation information and the target deformation information corresponding to the i-1th first facial sample image frame and the i-th first facial sample image frame, respectively, to obtain the i-th first constraint value.
[0108] In some embodiments, difference constraints are performed based on the predicted deformation information of the i-1th first facial sample image frame, the predicted deformation information of the i-1th first facial sample image frame, the target deformation information of the i-1th first facial sample image frame, and the target deformation information of the i-th first facial sample image frame to obtain the i-th first constraint value. This can be achieved by the following steps: performing difference processing on the predicted deformation information of the i-1th first facial sample image frame and the predicted deformation information of the i-th first facial sample image frame to obtain a first difference; performing difference processing on the target deformation information of the i-1th first facial sample image frame and the target deformation information of the i-th first facial sample image frame to obtain a second difference; performing difference processing on the first difference and the second difference to obtain the i-th first constraint value.
[0109] Here, the predicted deformation information of the i-1th first facial sample image frame is subtracted from the predicted deformation information of the i-th first facial sample image frame, and the obtained difference is determined as the first difference. The target deformation information of the i-1th first facial sample image frame is subtracted from the target deformation information of the i-th first facial sample image frame, and the obtained difference is determined as the second difference. The first difference and the second difference are subtracted, and the obtained difference is determined as the i-th first constraint value. For example, the i-th first constraint value can be expressed as in, represents the predicted deformation information of the i-1th first facial sample image frame, Represents the predicted deformation information of the i-th first facial sample image frame, y ti-1 Represents the target deformation information of the i-1th first facial sample image frame, y t i Represents the target deformation information of the i-th first facial sample image frame.
[0110] Continue to refer Figure 3C In step 3042, difference constraint is performed based on the predicted deformation information of the j-2th first facial sample image frame, the predicted deformation information of the j-th first facial sample image frame, the target deformation information of the j-2th first facial sample image frame, and the target deformation information of the j-th first facial sample image frame to obtain the j-th second constraint value.
[0111] Here, j is an integer from 2 to N-1. Subtract the predicted deformation information of the j-2th first facial sample image frame from the predicted deformation information of the jth first facial sample image frame, and determine the obtained difference as the third difference. Subtract the target deformation information of the j-2th first facial sample image frame from the target deformation information of the jth first facial sample image frame, and determine the obtained difference as the fourth difference. Subtract the third difference from the fourth difference, and determine the obtained difference as the i-th second constraint value. For example, the i-th second constraint value can be expressed as in, represents the predicted deformation information of the jth first facial sample image frame, Represents the predicted deformation information of the j-2th first facial sample image frame, y t j Represents the target deformation information of the jth first facial sample image frame, y t j-2 Represents the target deformation information of the j-2th first facial sample image frame.
[0112] Continue to refer Figure 3C In step 3043, a first constrained mean of the N-1 first constrained values is determined, and a second constrained mean of the N-2 second constrained values is determined.
[0113] Here, the N-1 first constraint values are summed and averaged to obtain the first constraint mean. The N-2 second constraint values are summed and averaged to obtain the second constraint mean. For example, taking N as 128, the first constraint mean can be expressed as The second constrained mean can be expressed as
[0114] Continue to refer Figure 3C In step 3044, the sum of the first constrained mean and the second constrained mean is determined as the third loss value.
[0115] Here, the first constraint mean and the second constraint mean are added together, and the resulting sum is determined as the third loss value. For example, the third loss value can be determined by the following formula (2):
[0116]
[0117] Among them, Loss t represents the third loss value, represents the predicted deformation information of the i-1th first facial sample image frame, Represents the predicted deformation information of the i-th first facial sample image frame, y t i-1 Represents the target deformation information of the i-1th first facial sample image frame, y t i represents the target deformation information of the i-th first facial sample image frame, represents the predicted deformation information of the jth first facial sample image frame, Represents the predicted deformation information of the j-2th first facial sample image frame, y t j Represents the target deformation information of the jth first facial sample image frame, y t j-2 Represents the target deformation information of the j-2th first facial sample image frame.
[0118] In an embodiment of the present application, a third loss value is determined based on the predicted deformation information of multiple first facial sample image frames and the target deformation information of multiple first facial sample image frames to achieve joint supervision of temporal dynamic deformation and optimize the prediction model's ability to predict continuous facial movements or expression changes.
[0119] Continue to refer Figure 3A In step 305, based on the first loss value, the second loss value and the third loss value, the model parameters of the prediction model are updated to obtain the trained prediction model.
[0120] Here, based on the first loss value, the second loss value and the third loss value, the model parameters of the prediction model are updated to train the prediction model and obtain a trained prediction model.
[0121] In some embodiments, updating the model parameters of the prediction model based on the first loss value, the second loss value, and the third loss value can be achieved by the following steps: determining the first weight value, the second weight value, and the third weight value corresponding to the first loss value, the second loss value, and the third loss value respectively; using the first weight value, the second weight value, and the third weight value to perform weighted fusion on the first loss value, the second loss value, and the third loss value to obtain the total loss value of the prediction model; based on the total loss value, updating the model parameters of the prediction model.
[0122] Here, weight values are set for the first loss value, the second loss value, and the third loss value to obtain their respective first weight values, second weight values, and third weight values. The first loss value, the second loss value, and the third loss value are weighted and summed using the first weight value, the second weight value, and the third weight value to obtain a total loss value of the prediction model, and the model parameters of the prediction model are updated based on the total loss value.
[0123] In an embodiment of the present application, by utilizing the first weight value, the second weight value and the third weight value, the first loss value, the second loss value and the third loss value are fused into the total loss value of the prediction model. Based on the total loss value, the model parameters of the prediction model are updated, thereby improving the accuracy of facial deformation prediction of face images using the prediction model.
[0124] Example, reference Figure 3D , Figure 3D 3 is a structural diagram of the training process of the prediction model provided in an embodiment of the present application. Using the prediction model 310, deformation prediction is performed on multiple first facial sample image frames 311, and corresponding predicted deformation information of each first facial sample image frame is obtained. The sample label of the first facial sample image frame includes target deformation information 312 and target rendered image 313. Rendering prediction is performed based on the multiple first facial sample image frames 311 to obtain a predicted rendered image. Based on the predicted deformation information and the target deformation information 312, a first loss value is determined, and based on the predicted rendered image and the target rendered image 313, a second loss value is determined. Based on the predicted deformation information of the multiple first facial sample image frames 311 and the target deformation information of the multiple first facial sample image frames 311, a third loss value is determined. Based on the first loss value, the second loss value and the third loss value, the model parameters of the prediction model are updated to obtain a trained prediction model.
[0125] The control method provided by the embodiment of the present application is described below. Figure 3E , Figure 3E This is a flow chart of the control method provided in the embodiment of the present application, which will be combined with Figure 3E The steps shown are explained.
[0126] In step 320 , a plurality of facial image frames to be processed are received.
[0127] Here, the multiple facial image frames to be processed are collected when the target subject's face is in motion. The eye camera of the bionic device can be used to collect multiple facial motion images of the target subject, and each facial motion image is cropped to obtain multiple facial image frames to be processed.
[0128] In step 321 , the trained prediction model is called to perform deformation prediction on each facial image frame to be processed, and predicted deformation information of each facial image frame to be processed is obtained.
[0129] Here, the trained prediction model is obtained by training using the model training method provided in an embodiment of the present application, and the predicted deformation information includes displacement data of multiple facial vertices. For each facial image frame to be processed, convolution processing is performed on the facial image frame to obtain a first facial image frame. Convolution processing and feature enhancement processing are performed on the first facial image frame to obtain a second facial image frame. Convolution processing and feature enhancement processing are performed on the second facial image frame to obtain a third facial image frame. Convolution processing is performed on the third facial image frame to obtain a fourth facial image frame. Pooling processing is performed on the fourth facial image frame to obtain predicted features of the facial image frame to be processed. Linear processing is performed on the predicted features of the facial image frame to be processed to obtain predicted deformation information of the facial image frame to be processed.
[0130] In step 322 , for each facial image frame to be processed, the bionic device is controlled to drive the plurality of facial vertices to move based on the plurality of displacement data.
[0131] Here, the bionic device can be a bionic robot, which obtains displacement data of multiple facial vertices from the predicted deformation information of each facial image frame to be processed, and uses the multiple displacement data to control the bionic device to move the multiple facial vertices to drive the bionic device to perform real-time facial imitation movements.
[0132] Based on this step 320 to step 322, by calling the trained prediction model, deformation prediction is performed on multiple facial image frames to be processed, and predicted deformation information of each facial image frame to be processed is obtained. Based on the predicted deformation information of multiple processed facial image frames, the bionic device is controlled to move multiple facial vertices, which can facilitate facial imitation by controlling the bionic device through a motor, thereby improving the efficiency of real-time facial imitation by the bionic device.
[0133] In some embodiments, the model training method provided by the embodiments of the present application can be applied in the field of cloud technology. A prediction model is used on a cloud platform to perform deformation prediction on multiple first facial sample image frames, and corresponding predicted deformation information is obtained for each first facial sample image frame. The multiple first facial sample image frames are collected when the face of the target object is in motion. The sample labels of the first facial sample image frames include target deformation information and target rendered images. Rendering prediction is then performed based on the first facial sample image frames to obtain a predicted rendered image. A first loss value is determined based on the predicted deformation information and the target deformation information, and a second loss value is determined based on the predicted rendered image and the target rendered image. The prediction model can be trained in combination with the geometric deformation prediction dimension and the visual rendering dimension. A third loss value is determined based on the predicted deformation information of the multiple first facial sample image frames and the target deformation information of the multiple first facial sample image frames, thereby achieving joint supervision of temporal dynamic deformation and optimizing the prediction model's ability to predict continuous facial movements or expression changes. Furthermore, model parameters of the prediction model are updated based on the first loss value, the second loss value, and the third loss value to obtain a trained prediction model. In this way, the loss value between the predicted deformation information and the target deformation information, the loss value between the predicted rendered image and the target rendered image, and the loss value between the predicted deformation information and the target deformation information of multiple first facial sample image frames are comprehensively utilized to update the model parameters of the prediction model, thereby improving the accuracy of facial deformation prediction of face images using the prediction model.
[0134] Below, we will explain the exemplary application of the model training method provided in the embodiment of the present application in the scenario of facial deformation prediction of face images.
[0135] In related technologies, the three-dimensional representation of human faces is usually based on vertex or three-dimensional deformable model methods, which cannot be used to imitate faces through motor-controlled biomimetic robots. In this embodiment, a model training method is proposed to address the problems existing in related technologies. Compared with related technologies, it includes the following improvements:
[0136] This embodiment of the present application proposes a facial blend shape prediction model. This model predicts facial blend shape based on motion images of human faces, generates predicted blend shape information, and uses this information to control motors to drive biomimetic facial imitation. During the prediction model training process, a rendering image is used as an auxiliary constraint (the second loss value in the above embodiment) and a multi-frame motion constraint (the third loss value in the above embodiment) is introduced to smooth the biomimetic facial imitation process.
[0137] In the embodiment of the present application, the predicted deformation information is 52-dimensional deformation (BlendShape) information predicted based on the facial image of the human face, that is, 52 independent facial deformation information, each facial deformation information corresponds to a specific facial muscle movement or expression unit, and each facial deformation information contains displacement data of multiple vertices. For example, refer to Figure 4 , Figure 4 It is a schematic diagram of the prediction model structure provided in the embodiment of the present application.
[0138] The facial image of the human face and its corresponding real deformation information (the target deformation information in the above embodiment) are obtained through the motion capture device, and the facial image of the human face is preprocessed to obtain a first facial sample image frame 401. For example, the first facial sample image frame is 112*112*3. The first facial sample image frame 401 is input into the prediction model. The prediction model is built based on a lightweight convolutional neural network (shufflenetv2). The prediction model is used to perform a series of convolution processing, feature enhancement processing, pooling processing and linear processing on the first facial sample image frame 401 to obtain a second facial sample image frame 402, a third facial sample image frame 403, a fourth facial sample image frame 404 and a fifth facial sample image frame 405 in turn. The output of the prediction model is the predicted 52-dimensional deformation information. The processing process of the first facial sample image frame is shown in Table 1:
[0139] Table 1 Processing process of the first facial sample image frame
[0140]
[0141]
[0142] In order to reduce the error of the predicted deformation information, such as optimizing the lip shape change in facial deformation, a loss value is calculated for the predicted deformation information 406 and the real deformation information 407 to obtain a first loss value 408. For example, the first loss value can be determined by the following formula (1):
[0143]
[0144] Where WingLoss represents the first loss value, ∈ represents the set hyperparameter, w represents the control magnification, x represents the position error of the corresponding facial vertex in the predicted deformation information and the actual deformation information, and C is expressed as w-wln(1+w / ∈). Since the response of the ln(1+x) function to subtle differences is much higher than that of the quadratic function Therefore, the above formula (1) can be used to achieve refined prediction of deformation information.
[0145] The predicted deformation information is given its corresponding physical meaning, that is, an auxiliary branch is introduced into the prediction model to predict the rendering of the facial sample image frame, guiding the prediction model to better learn the refined prediction of deformation information. For example, refer to Figure 5 , Figure 5 This is a structural diagram of the prediction model training process guided by the rendering provided in an embodiment of the present application.
[0146] The training of the auxiliary branch refers to the design of the U-type network. First, a 3x3 convolution module is used to extract features from the facial sample image frame of the last level, and the extracted features are spliced with the features of the facial sample image frame of the previous level to obtain spliced features. The 3x3 convolution module is then used to adjust the dimension of the spliced features, and the above steps are repeated to obtain the final predicted rendering image 501. The real deformation information is used for face rendering to obtain a real rendering image 502, and the loss value between the predicted rendering image and the real rendering image (the target rendering image in the above embodiment) is calculated through the smooth absolute value loss function (L1_Smooth) to obtain a second loss value 503, thereby realizing the reconstruction of the rendering image, thereby helping the prediction model to better understand and learn facial deformation information. The auxiliary branch is only used in the training process of the prediction model, and is used to Figure 1 The prediction model in
[15] performs facial deformation reasoning.
[0147] When training the prediction model, the number of facial sample image frames is set to 128, and multiple facial sample image frames are input into the prediction model in chronological order. The loss value between two adjacent facial sample image frames and two facial sample image frames across frames is calculated to obtain a third loss value. For example, the third loss value can be determined by the following formula (2):
[0148]
[0149] Among them, Loss t represents the third loss value, Represents the predicted deformation information, y t Representing the real deformation information, the multi-frame motion loss is established through the above formula (2), which realizes the difference constraint between the predicted deformation information difference of two adjacent frames and the real deformation information difference of two adjacent frames, as well as the difference constraint between the predicted deformation information difference across frames and the real deformation information difference across frames.
[0150] Finally, based on the first loss value, the second loss value, and the third loss value obtained in the above process, the prediction model is trained to obtain a trained prediction model.
[0151] In the above-mentioned scenario of facial deformation prediction for facial images, the eye camera of the bionic robot can be used to capture facial motion images of the human face, and facial deformation prediction can be performed on the facial motion images of the human face based on the prediction model to obtain predicted deformation information. The predicted deformation information can then be used to control the motor to further drive the bionic robot's face to imitate in real time.
[0152] The following continues to describe the exemplary structure of the model training device 455 provided in the embodiment of the present application as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the model training device 455 of the memory 450 may include: a deformation prediction module 4551, which is used to use the prediction model to perform deformation prediction on multiple first facial sample image frames, and obtain predicted deformation information of each first facial sample image frame, wherein the multiple first facial sample image frames are collected when the face of the target object moves, and the sample labels of the first facial sample image frames include target deformation information and target rendered images; a rendering prediction module 4552, which is used to perform rendering prediction based on the first facial sample image frames to obtain a predicted rendered image; a first determination module 4553, which is used to determine a first loss value based on the predicted deformation information and the target deformation information, and to determine a second loss value based on the predicted rendered image and the target rendered image; a second determination module 4554, which is used to determine a third loss value based on the predicted deformation information of the multiple first facial sample image frames and the target deformation information of the multiple first facial sample image frames; and a model updating module 4555, which is used to update the model parameters of the prediction model based on the first loss value, the second loss value and the third loss value to obtain a trained prediction model.
[0153] In some embodiments, the deformation prediction module 4551 is also used to perform convolution processing on the first facial sample image frame for each first facial sample image frame to obtain a second facial sample image frame; perform convolution processing and feature enhancement processing on the second facial sample image frame to obtain a third facial sample image frame; perform convolution processing and feature enhancement processing on the third facial sample image frame to obtain a fourth facial sample image frame; perform convolution processing on the fourth facial sample image frame to obtain a fifth facial sample image frame; perform pooling processing on the fifth facial sample image frame to obtain predicted features of the first facial sample image frame; and perform linear processing on the predicted features of the first facial sample image frame to obtain predicted deformation information of the first facial sample image frame.
[0154] In some embodiments, the rendering prediction module 4552 is further used to perform feature extraction on the fifth facial sample image frame to obtain a first sample feature, and perform feature extraction on the fourth facial sample image frame to obtain a second sample feature; perform splicing processing on the first sample feature and the second sample feature to obtain a first splicing feature, and perform convolution processing on the first splicing feature to obtain a first convolution result; perform feature extraction on the third facial sample image frame to obtain a third sample feature, and perform splicing processing on the third sample feature and the first convolution result to obtain a second splicing feature, and perform convolution processing on the second splicing feature to obtain a second convolution result; perform feature extraction on the second facial sample image frame to obtain a fourth sample feature, and perform splicing processing on the fourth sample feature and the second convolution result to obtain a third splicing feature, and perform convolution processing on the third splicing feature to obtain a third convolution result; perform feature extraction on the first facial sample image frame to obtain a fifth sample feature, and perform splicing processing on the fifth sample feature and the third convolution result to obtain a fourth splicing feature; perform rendering prediction based on the fourth splicing feature to obtain a predicted rendered image.
[0155] In some embodiments, the first determination module 4553 is also used to determine a first error value between the predicted deformation information and the target deformation information. When the absolute value of the first error value is less than a first preset value, the first error value is logarithmically transformed to obtain a first logarithmic value, and the product of the first preset value and the first logarithmic value is determined as the first loss value; when the absolute value of the first error value is greater than or equal to the first preset value, the difference between the absolute value of the first error value and the second preset value is determined as the first loss value.
[0156] In some embodiments, the first determination module 4553 is also used to perform logarithmic conversion on the first preset value to obtain a second logarithmic value, and determine the product of the second logarithmic value and the first preset value as the first parameter value; and determine the difference between the first parameter value and the first preset value as the second preset value.
[0157] In some embodiments, the first determination module 4553 is further used to determine a second error value between the predicted rendered image and the target rendered image; determine the second error value as the input value of a preset smooth absolute value loss function, determine the output value of the smooth absolute value loss function; and determine the output value as the second loss value.
[0158] In some embodiments, the second determination module 4554 is further used to perform difference constraint based on the predicted deformation information of the i-1th first facial sample image frame, the predicted deformation information of the i-1th first facial sample image frame, the target deformation information of the i-1th first facial sample image frame, and the target deformation information of the i-1th first facial sample image frame to obtain the i-th first constraint value, where i is an integer from 1 to N-1, and N is the total number of first facial sample image frames; perform difference constraint based on the predicted deformation information of the j-2th first facial sample image frame, the predicted deformation information of the j-2th first facial sample image frame, the target deformation information of the j-2th first facial sample image frame, and the target deformation information of the j-1th first facial sample image frame to obtain the j-th second constraint value, where j is an integer from 2 to N-1; determine a first constraint mean of the N-1 first constraint values, and determine a second constraint mean of the N-2 second constraint values; and determine the sum of the first constraint mean and the second constraint mean as the third loss value.
[0159] In some embodiments, the model update module 4555 is also used to determine the first weight value, the second weight value and the third weight value corresponding to the first loss value, the second loss value and the third loss value respectively; use the first weight value, the second weight value and the third weight value to weightedly fuse the first loss value, the second loss value and the third loss value to obtain the total loss value of the prediction model; based on the total loss value, update the model parameters of the prediction model.
[0160] The following continues to describe an exemplary structure of a software module implemented as the control device provided in an embodiment of the present application. In some embodiments, the software modules in the control device may include: an image acquisition module for receiving multiple facial image frames to be processed, where the multiple facial image frames to be processed are collected when the face of the target object moves; a deformation prediction module for calling a trained prediction model to perform deformation prediction on each facial image frame to be processed to obtain predicted deformation information for each facial image frame to be processed, where the trained prediction model is trained using the model training method provided in an embodiment of the present application, and the predicted deformation information includes displacement data of multiple facial vertices; a facial movement module for controlling the bionic device to drive multiple facial vertices to move based on the multiple displacement data for each facial image frame to be processed.
[0161] The embodiments of the present application provide a computer program product, which includes computer-executable instructions or a computer program stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions or computer program from the computer-readable storage medium and executes the computer-executable instructions or computer program, causing the electronic device to perform the model training method or control method provided in the embodiments of the present application.
[0162] The embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will be caused to execute the model training method or control method provided by the embodiment of the present application, for example, Figure 3A The model training method shown.
[0163] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various devices including one or any combination of the above memories.
[0164] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0165] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0166] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0167] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A model training method, characterized in that: The method comprises: Using a prediction model, deformation prediction is performed on a plurality of first facial sample image frames to obtain predicted deformation information for each of the first facial sample image frames, wherein the plurality of first facial sample image frames are collected when the face of a target object is in motion, and sample labels of the first facial sample image frames include target deformation information and a target rendered image; Performing rendering prediction based on the first facial sample image frame to obtain a predicted rendered image; determining a first loss value based on the predicted deformation information and the target deformation information, and determining a second loss value based on the predicted rendered image and the target rendered image; determining a third loss value based on the predicted deformation information of the plurality of first facial sample image frames and the target deformation information of the plurality of first facial sample image frames; Based on the first loss value, the second loss value and the third loss value, the model parameters of the prediction model are updated to obtain a trained prediction model.
2. The method according to claim 1, characterized in that The method of using the prediction model to perform deformation prediction on the plurality of first facial sample image frames to obtain corresponding predicted deformation information of each of the first facial sample image frames includes: For each of the first facial sample image frames, performing convolution processing on the first facial sample image frame to obtain a second facial sample image frame; performing convolution processing and feature enhancement processing on the second facial sample image frame to obtain a third facial sample image frame; performing convolution processing and feature enhancement processing on the third facial sample image frame to obtain a fourth facial sample image frame; performing convolution processing on the fourth facial sample image frame to obtain a fifth facial sample image frame; performing pooling processing on the fifth facial sample image frame to obtain predicted features of the first facial sample image frame; Linear processing is performed on the predicted features of the first facial sample image frame to obtain predicted deformation information of the first facial sample image frame.
3. The method according to claim 2, characterized in that The performing rendering prediction based on the first facial sample image frame to obtain a predicted rendered image includes: Performing feature extraction on the fifth facial sample image frame to obtain a first sample feature, and performing feature extraction on the fourth facial sample image frame to obtain a second sample feature; Performing a splicing process on the first sample feature and the second sample feature to obtain a first splicing feature, and performing a convolution process on the first splicing feature to obtain a first convolution result; Performing feature extraction on the third facial sample image frame to obtain a third sample feature, performing splicing processing on the third sample feature and the first convolution result to obtain a second splicing feature, and performing convolution processing on the second splicing feature to obtain a second convolution result; Performing feature extraction on the second facial sample image frame to obtain a fourth sample feature, performing splicing processing on the fourth sample feature and the second convolution result to obtain a third splicing feature, and performing convolution processing on the third splicing feature to obtain a third convolution result; Performing feature extraction on the first facial sample image frame to obtain a fifth sample feature, and performing splicing processing on the fifth sample feature and the third convolution result to obtain a fourth splicing feature; Rendering prediction is performed based on the fourth splicing feature to obtain a predicted rendered image.
4. The method according to claim 1, wherein The determining a first loss value based on the predicted deformation information and the target deformation information includes: determining a first error value between the predicted deformation information and the target deformation information, performing a logarithmic transformation on the first error value to obtain a first logarithmic value when an absolute value of the first error value is less than a first preset value, and multiplying the first preset value by the first logarithmic value to determine the first loss value; When the absolute value of the first error value is greater than or equal to the first preset value, the difference between the absolute value of the first error value and the second preset value is determined as the first loss value.
5. The method according to claim 1, wherein The determining a second loss value based on the predicted rendered image and the target rendered image includes: determining a second error value between the predicted rendered image and the target rendered image; Determine the second error value as an input value of a preset smooth absolute value loss function, and determine an output value of the smooth absolute value loss function; The output value is determined as the second loss value.
6. The method according to claim 4, characterized in that The method further comprises: Performing logarithmic conversion on the first preset value to obtain a second logarithmic value, and determining the product of the second logarithmic value and the first preset value as a first parameter value; The difference between the first parameter value and the first preset value is determined as the second preset value.
7. The method according to any one of claims 1 to 6, characterized in that The determining a third loss value based on the predicted deformation information of the plurality of first facial sample image frames and the target deformation information of the plurality of first facial sample image frames includes: performing difference constraint based on the predicted deformation information of the (i-1)th first facial sample image frame, the predicted deformation information of the (i)th first facial sample image frame, the target deformation information of the (i-1)th first facial sample image frame, and the target deformation information of the (i)th first facial sample image frame to obtain an (i)th first constraint value, where the value of (i) is an integer from 1 to N-1, and N is the total number of first facial sample image frames; performing difference constraint based on the predicted deformation information of the j-2th first facial sample image frame, the predicted deformation information of the j-th first facial sample image frame, the target deformation information of the j-2th first facial sample image frame, and the target deformation information of the j-th first facial sample image frame to obtain a j-th second constraint value, where the value of j is an integer from 2 to N-1; Determine a first constrained mean of the N-1 first constrained values, and determine a second constrained mean of the N-2 second constrained values; A sum of the first constrained mean and the second constrained mean is determined as the third loss value.
8. The method according to claim 7, characterized in that The step of performing difference constraint based on the predicted deformation information of the (i-1)th first facial sample image frame, the predicted deformation information of the (i)th first facial sample image frame, the target deformation information of the (i-1)th first facial sample image frame, and the target deformation information of the (i)th first facial sample image frame to obtain an (i)th first constraint value includes: performing difference processing on the predicted deformation information of the (i-1)th first facial sample image frame and the predicted deformation information of the (i)th first facial sample image frame to obtain a first difference; performing difference processing on the target deformation information of the (i-1)th first facial sample image frame and the target deformation information of the (i)th first facial sample image frame to obtain a second difference; Perform difference processing on the first difference and the second difference to obtain an i-th first constraint value.
9. The method according to any one of claims 1 to 5, characterized in that The updating of the model parameters of the prediction model based on the first loss value, the second loss value, and the third loss value includes: Determine a first weight value, a second weight value, and a third weight value corresponding to the first loss value, the second loss value, and the third loss value, respectively; Performing weighted fusion on the first loss value, the second loss value, and the third loss value using the first weight value, the second weight value, and the third weight value to obtain a total loss value of the prediction model; Based on the total loss value, model parameters of the prediction model are updated.
10. A control method, characterized in that: The method comprises: receiving a plurality of facial image frames to be processed, wherein the plurality of facial image frames to be processed are collected when the face of the target object is in motion; Calling a trained prediction model to perform deformation prediction on each of the facial image frames to be processed to obtain predicted deformation information for each of the facial image frames to be processed, wherein the trained prediction model is trained using the model training method according to any one of claims 1 to 9, and the predicted deformation information includes displacement data of a plurality of facial vertices; For each of the facial image frames to be processed, the bionic device is controlled to drive the plurality of facial vertices to move based on the plurality of displacement data.
11. A model training device, characterized in that: The device comprises: a deformation prediction module, configured to perform deformation prediction on a plurality of first facial sample image frames using a prediction model, and obtain corresponding predicted deformation information for each of the first facial sample image frames, wherein the plurality of first facial sample image frames are collected when the face of a target object is in motion, and the sample labels of the first facial sample image frames include target deformation information and a target rendered image; A rendering prediction module, configured to perform rendering prediction based on the first facial sample image frame to obtain a predicted rendered image; a first determining module, configured to determine a first loss value based on the predicted deformation information and the target deformation information, and to determine a second loss value based on the predicted rendered image and the target rendered image; a second determining module, configured to determine a third loss value based on the predicted deformation information of the plurality of first facial sample image frames and the target deformation information of the plurality of first facial sample image frames; A model updating module is used to update the model parameters of the prediction model based on the first loss value, the second loss value and the third loss value to obtain a trained prediction model.
12. A control device, characterized in that: The device comprises: An image acquisition module is configured to receive a plurality of facial image frames to be processed, wherein the plurality of facial image frames to be processed are acquired when the face of the target object is in motion; a deformation prediction module, configured to call a trained prediction model to perform deformation prediction on each of the facial image frames to be processed, thereby obtaining predicted deformation information for each of the facial image frames to be processed, wherein the trained prediction model is obtained by training using the model training method according to any one of claims 1 to 9, and the predicted deformation information includes displacement data of a plurality of facial vertices; The facial movement module is used to control the bionic device to drive the plurality of facial vertices to move based on the plurality of displacement data for each facial image frame to be processed.
13. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; The processor is used to implement the model training method described in any one of claims 1 to 9 or the control method described in claim 10 when executing the computer-executable instructions or computer programs stored in the memory.
14. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the model training method according to any one of claims 1 to 9 or the control method according to claim 10 is implemented.
Citation Information
Patent Citations
Face display method, device and equipment for virtual character and storage medium
CN110141857A
Facial expression image generation method and device, equipment, medium and product
CN115731327A
Key point prediction model training method and device, electronic equipment and computer readable storage medium
CN119445277A