Data processing method and device, electronic equipment, storage medium and program product

By constructing and adjusting the deformation parameters of the three-dimensional model, the problem of low accuracy of expression parameter migration is solved, and the accurate and natural migration of expression parameters is achieved.

CN120689214APending Publication Date: 2025-09-23MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510517670.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In the prior art, during the expression parameter migration process, the appearance extraction is performed from a macroscopic perspective of the image, resulting in low accuracy in the expression parameter migration.

Method used

By constructing a three-dimensional model of the first object, determining deformation parameters, and adjusting the three-dimensional model based on the deformation parameters, an image of the second object is generated to ensure accurate migration of the expression parameters.

Benefits of technology

The accuracy and naturalness of expression parameter migration are improved, ensuring that the migrated expression parameters conform to the characteristics and expression parameter features of the second object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689214A_ABST
    Figure CN120689214A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device, electronic equipment, a storage medium and a program product. The method comprises the following steps: determining a first three-dimensional model of a first object, wherein expression parameters of the first three-dimensional model of the first object and a third three-dimensional model of a second object are the same; constructing a second three-dimensional model of the first object based on the first image of the first object; determining deformation parameters based on the second three-dimensional model and the first three-dimensional model; based on the deformation parameters, the third three-dimensional model is adjusted, a fourth three-dimensional model of the second object is obtained, and expression parameters of the fourth three-dimensional model and the second three-dimensional model are the same; and generating a second image of the second object based on the fourth three-dimensional model, the expression parameters of the second object in the second image being the same as the expression parameters of the first object in the first image. According to the invention, the accuracy of expression parameter migration can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data processing method, device, electronic device, storage medium, and program product. Background Art

[0002] Expression transfer is a computer graphics and artificial intelligence technique used to transfer the facial expression parameters of one person to the face of another person or virtual character. This technique typically involves capturing the facial expression parameters of the source individual and then applying these expression parameters to the facial model of the target individual or character to create realistic animations or videos.

[0003] In related technologies, for expression parameter migration, appearance extraction is usually performed on the image to be migrated and the target image respectively, and the target image with the expression parameters of the image to be migrated is obtained through splicing and redirection, thereby realizing expression parameter migration. In this way, since the appearance extraction is performed from the macro perspective of the image, the image after expression parameter migration cannot achieve accurate migration of expression parameters, resulting in low accuracy of expression parameter migration. Summary of the Invention

[0004] The embodiments of the present application provide a data processing method, device, electronic device, computer-readable storage medium, and computer program product, which can effectively improve the accuracy of expression parameter migration.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] The present invention provides a data processing method, including:

[0007] constructing a second three-dimensional model of the first object based on a first image of the first object having first expression parameters, wherein the second three-dimensional model has the first expression parameters;

[0008] Determining deformation parameters of the second three-dimensional model with reference to the first three-dimensional model of the first object, wherein the first three-dimensional model and the third three-dimensional model of the second object have second expression parameters;

[0009] adjusting the third three-dimensional model based on the deformation parameter to obtain a fourth three-dimensional model of the second object, wherein the fourth three-dimensional model has the first expression parameter;

[0010] Based on the fourth three-dimensional model, a second image of the second object is generated.

[0011] An embodiment of the present application provides a data processing device, including:

[0012] a construction module, configured to construct a second three-dimensional model of the first object based on a first image of the first object having first expression parameters, wherein the second three-dimensional model has the first expression parameters;

[0013] a determining module, configured to determine a deformation parameter of the second three-dimensional model with reference to a first three-dimensional model of the first object, the first three-dimensional model and a third three-dimensional model of the second object having second expression parameters;

[0014] an adjustment module, configured to adjust the third three-dimensional model based on the deformation parameter to obtain a fourth three-dimensional model of the second object, wherein the fourth three-dimensional model has the first expression parameter;

[0015] A generating module is configured to generate a second image of the second object based on the fourth three-dimensional model.

[0016] An embodiment of the present application provides an electronic device, including:

[0017] a memory for storing computer-executable instructions or computer programs;

[0018] The processor is used to implement the data processing method provided in the embodiment of the present application when executing the computer-executable instructions or computer programs stored in the memory.

[0019] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for causing a processor to execute the instructions to implement the data processing method provided in the embodiment of the present application.

[0020] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the data processing method described in the embodiment of the present application.

[0021] The embodiments of the present application have the following beneficial effects:

[0022] The method involves determining a first 3D model of a first object, constructing a second 3D model of the first object based on a first image of the first object, determining deformation parameters based on the first and second 3D models, and adjusting the third 3D model based on the deformation parameters to obtain a fourth 3D model of the second object. A second image is generated based on the fourth 3D model, and the deformation parameters are determined using the first 3D model of the first object as a reference. The first 3D model and the third 3D model of the second object both have identical expression parameters. By comparing the second 3D model with the first 3D model, deformation patterns based on the expression parameters can be identified. Based on the determined deformation parameters, the third 3D model is adjusted. The third 3D model of the second object is then converted to the fourth 3D model. This adjustment is based on the deformation relationship between the second and first 3D models, ensuring the accuracy and naturalness of the expression parameter transfer. A second image of the second object is then generated based on the fourth 3D model. The adjusted fourth 3D model is then converted into a visual image, completing the transfer of expression parameters from the first object to the second object. The entire process is based on the deformation relationship between models, ensuring the natural transition and accurate reproduction of expression parameters. It takes into account the differences in expression parameters between different objects, making the migrated expression parameters more consistent with the characteristics and expression parameter features of the second object, thereby effectively improving the accuracy of expression parameter migration. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 Schematic diagram of the data processing system provided in the embodiment of the present application;

[0024] Figure 2 is a structural diagram of an electronic device for data processing provided in an embodiment of the present application;

[0025] Figure 3 Schematic diagram of the data processing method provided in the embodiment of the present application;

[0026] Figure 4 This is a schematic diagram of the principle of the data processing method provided in the embodiment of the present application Figure 1 ;

[0027] Figure 5 This is a schematic diagram of the principle of the data processing method provided in the embodiment of the present application Figure 2 ;

[0028] Figure 6 This is a schematic diagram of the principle of the data processing method provided in the embodiment of the present application Figure 3 ;

[0029] Figure 7 This is a schematic diagram of the principle of the data processing method provided in the embodiment of the present application Figure 4 ;

[0030] Figure 8This is a schematic diagram of the principle of the data processing method provided in the embodiment of the present application Figure 5 ;

[0031] Figure 9 This is a schematic diagram of the principle of the data processing method provided in the embodiment of the present application Figure 6 ;

[0032] Figure 10 It is a schematic diagram of the effect of the data processing method provided in the embodiment of the present application. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0034] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0035] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0037] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0038] 1) Expression Transfer: This is a computer graphics and artificial intelligence technique used to transfer one person's facial expression parameters to the face of another person or virtual character. This technique typically involves capturing the facial expression parameters of the source individual and then applying these expression parameters to a facial model of the target individual or character to create realistic animations or videos.

[0039] 2) Artificial Intelligence (AI): This is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technologies, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0040] 3) Machine Learning (ML): This is a multidisciplinary interdisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. Pretrained models are the latest development in deep learning, integrating these techniques.

[0041] 4) In response to: used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.

[0042] 5) 3D Model: A 3D model is a three-dimensional geometric structure represented mathematically or numerically, typically consisting of vertices, edges, and faces. It can be used in fields such as computer graphics, virtual reality, medical imaging, and game development. Depending on the modeling approach, 3D models can be categorized as parametric statistical models, deep learning generative models, and implicit representation models. A 3D model can be a 3D Morphable Model (3DMM), a Skinned Multi-Animal Linear Model (SMPL), or a Faces Learned with an Articulated Model and Expressions (FLAME). The FLAME model is a parametric 3D animal model based on a skinned mesh. It generates animatable animal models by statistically modeling 3D scan data of different animals (such as cats, dogs, and horses). Its core concept is similar to the SMPL model used in human body modeling, but is optimized for quadrupeds. 5) 3D Deformable Model: This is a statistical model used to represent and model 3D facial shape and texture. It collects a large amount of 3D facial scan data and uses statistical methods to learn the variation patterns of facial shape and texture, thereby constructing a model capable of generating a variety of facial shapes and textures.

[0043] 6) Deformation Transfer for Triangle Meshes: This is a computer graphics technique used to apply geometric deformations of one shape to another. It is used to transfer complex, computationally expensive deformations from one 3D mesh to another 3D mesh of similar topology but different geometry.

[0044] 7) UV: In 3D models, UV refers to the texture coordinate system used to map a 2D texture image onto the 3D model surface. Through UV coordinates, each vertex of the 3D model can be mapped to a pixel (texture) of the 2D image, thus displaying texture details on the model.

[0045] 8) Faces: In 3D computer graphics, faces are one of the basic elements that make up a 3D model. A face is typically composed of three or more vertices, which define the shape of a polygon in 3D space. The most commonly used facets are triangles due to their simplicity and ability to stably represent any polygon. Faces are the basic building blocks of a 3D model's surface, collectively forming the model's geometry. Each facet has its own properties, such as color, texture coordinates, and normal vectors, which are used to add detail and realism to the model's surface during rendering. In 3D modeling, the number and quality of faces directly impact the model's appearance and performance. A high facet density provides finer surface detail, but also increases model complexity and computational cost. Therefore, modelers often need to strike a balance between model detail and performance.

[0046] 9) Expression parameters: These are a set of parameters used to describe and control expression changes in a 3D model. These parameters can include facial muscle movement, expression intensity, and shape changes, collectively determining the appearance and dynamic behavior of the 3D model in different expression states. By adjusting these expression parameters, fine-grained control and customization of the 3D model's expression can be achieved. In the technical parameterization of expression parameters, these parameters are quantified and parameterized to facilitate the transfer and adjustment of expressions between different objects, thereby enabling expression synthesis and animation. Quantification of expression parameters typically involves converting various features of an expression into numerical values ​​that can be understood and processed by computer programs. Facial muscle activity quantifies the degree of contraction and relaxation of facial muscles as a value between 0 and 1, with 0 representing complete relaxation and 1 representing maximum contraction. These values ​​can be used to control the movement of facial muscles in the 3D model. Expression intensity quantifies the intensity of an expression as a percentage. For example, the intensity of a smile can range from 10% (a slight smile) to 100% (a broad smile). These values ​​can be used to adjust the exaggeration of an expression. Shape change quantifies changes in facial shape as changes in coordinate values. For example, an upward turn of the mouth corners can be represented by an increase in the y-coordinate of the mouth corner point, while a downward turn of the eye corners can be represented by a decrease in the y-coordinate of the eye corner point. Expression duration can quantify the duration of an expression into seconds or milliseconds. This helps control the dynamic changes and transitions of expressions. Expression transition can quantify the transition between expressions into a transition function, such as a linear transition, an exponential transition, or a Bezier curve transition. These functions determine the speed and smoothness of expression changes. Expression parameters can be precisely controlled and adjusted to achieve a detailed representation of the expressions of 3D models. For example, a 3D character's smile can be achieved by adjusting parameters such as the amplitude of the upward turn of the mouth corners, the shape change of the eye corners, and the duration of the smile. The quantification of these parameters makes the synthesis and animation of expressions more controllable and efficient.

[0047] During the implementation of the embodiments of this application, the applicant discovered that the related technology has the following problems:

[0048] In related technologies, for expression parameter migration, appearance extraction is usually performed on the image to be migrated and the target image respectively, and the target image with the expression parameters of the image to be migrated is obtained through splicing and redirection, thereby realizing expression parameter migration. In this way, since the appearance extraction is performed from the macro perspective of the image, the image after expression parameter migration cannot achieve accurate migration of expression parameters, resulting in low accuracy of expression parameter migration.

[0049] The embodiments of the present application provide a data processing method, device, electronic device, computer-readable storage medium and computer program product, which can effectively improve the accuracy of expression parameter migration. The following describes an exemplary application of the expression parameter migration system provided by the embodiments of the present application.

[0050] See also Figure 1 , Figure 1 It is a schematic diagram of the architecture of the data processing system 100 provided in an embodiment of the present application. The terminal (terminal 400 is shown as an example) is connected to the server 200 via the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0051] The terminal 400 is used for the user to use the client 410 and display the second image on the graphical interface 410-1 (graphic interface 410-1 is shown as an example). The terminal 400 and the server 200 are connected to each other via a wired or wireless network.

[0052] In some embodiments, the server 200 can be an independent physical server, or a server cluster or business system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal 400 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart TV, a smart watch, a car terminal, etc., but is not limited to this. The electronic device provided in the embodiment of the present application can be implemented as a terminal or as a server. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiment of the present application.

[0053] In some embodiments, the server 200 determines a first three-dimensional model of the first object, and constructs a second three-dimensional model based on the first image of the first object, determines deformation parameters based on the first three-dimensional model and the second three-dimensional model, and sends the deformation parameters to the terminal 400. The terminal 400 adjusts the third three-dimensional model based on the deformation parameters to obtain a fourth three-dimensional model of the second object, and generates a second image of the second object based on the fourth three-dimensional model.

[0054] In other embodiments, the terminal 400 determines a first three-dimensional model of the first object, and constructs a second three-dimensional model based on the first image of the first object, determines deformation parameters based on the first three-dimensional model and the second three-dimensional model, and sends the deformation parameters to the server 200. The server 200 adjusts the third three-dimensional model based on the deformation parameters to obtain a fourth three-dimensional model of the second object, and generates a second image of the second object based on the fourth three-dimensional model.

[0055] See also Figure 2 , Figure 2 is a structural diagram of an electronic device 500 for data processing provided in an embodiment of the present application, wherein: Figure 2 The electronic device 500 shown may be Figure 1 The server 200 or the terminal 400 in Figure 2 The electronic device 500 shown includes: at least one processor 430, a memory 450, and at least one network interface 420. The various components in the electronic device 500 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 440 is not described in detail. Figure 2 Various buses are labeled as bus system 440 .

[0056] The processor 430 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0057] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 430.

[0058] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0059] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0060] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;

[0061] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB).

[0062] In some embodiments, the data processing device provided in the embodiments of the present application can be implemented in software. Figure 2 Data processing device 455 stored in memory 450 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: construction module 4551, determination module 4552, adjustment module 4553, and generation module 4554. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.

[0063] In other embodiments, the data processing device provided in the embodiments of the present application can be implemented in hardware. As an example, the data processing device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the data processing method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0064] In some embodiments, the terminal or server can implement the data processing method provided in the embodiment of the present application by running a computer program or computer executable instructions. For example, the computer program can be a native program (for example, a dedicated migration program) or a software module in the operating system, for example, a deblurring module that can be embedded in any program (such as an instant messaging client, a photo album program, an electronic map client, a navigation client); for example, it can be a local (Native) application (APP, Application), that is, a program that needs to be installed in the operating system to run. In short, the above-mentioned computer program can be any form of application, module or plug-in.

[0065] The data processing method provided in the embodiments of the present application will be explained in combination with the exemplary application and implementation of the server or terminal provided in the embodiments of the present application.

[0066] See also Figure 3 , Figure 3 This is a flow chart of the data processing method provided in the embodiment of the present application, which will be combined with Figure 3 Steps 101 to 105 are shown for illustration. The data processing method provided in the embodiment of the present application can be implemented by the server or the terminal alone, or by the server and the terminal in collaboration. The following description will be made using the server alone as an example.

[0067] In step 101 , a first three-dimensional model of a first object is determined, and expression parameters of the first three-dimensional model of the first object and a third three-dimensional model of a second object are the same.

[0068] In some embodiments, a 3D model refers to a three-dimensional geometric structure represented by mathematics or data, typically consisting of vertices, edges, and faces, and can be used in fields such as computer graphics, virtual reality, medical imaging, and game development. Depending on the modeling method, 3D models can be divided into parameterized statistical models, deep learning generative models, implicit representation models, etc. The 3D model can be a 3D Morphable Model (3DMM), a Skinned Multi-Animal Linear Model (SMPL), or a Faces Learned with an Articulated Model and Expressions (FLAME model). The Skinned Multi-Animal Linear Model is a parametric animal 3D model based on a skinned mesh. It generates animatable animal models by performing statistical modeling on 3D scan data of different animals (such as cats, dogs, and horses). Its core concept is similar to the SMPL model used in human body modeling, but is optimized for quadrupeds. 5) 3D Deformable Model: This is a statistical model used to represent and model 3D facial shape and texture. It collects a large amount of 3D facial scan data and uses statistical methods to learn the variation patterns of facial shape and texture, thereby constructing a model capable of generating a variety of facial shapes and textures.

[0069] In some embodiments, determining a first 3D model of a first object refers to the process of creating or selecting a 3D model to represent a specific object in computer graphics or related fields. This 3D model is a geometric structure composed of vertices, edges, and faces that accurately reflects the object's shape and appearance. The first 3D model of the first object and the third 3D model of the second object have the same expression parameters, meaning that the two 3D models share the same parameterization mechanism for expressing expressions. Expression parameters are a set of numerical values ​​that control the expression changes of a 3D model, determining the model's appearance and dynamic behavior in different expression states. Suppose we have two 3D models: one is a human head model and the other is an animal head model. If the expression parameters of these two models are the same, then we can use the same set of parameters to control their expressions, such as smiling, frowning, and surprise. This means that adjusting the smile parameters of one model will cause the other model to smile in the same manner. The same expression parameters are often achieved by using the same parameterization model or technology. For example, if both models are created based on a 3D deformable model (3DMM), their expression parameters may include the same muscle movement parameters, shape parameters, and texture parameters. These parameters can be adjusted and combined to produce a variety of different expressions.

[0070] For example, in a game development scenario, the first object is a virtual human character, and the second object is a virtual animal character, such as a werewolf or an orc. In a role-playing game (RPG), players can choose from different types of virtual characters. To ensure game immersion, the development team needs to ensure that all characters can display the same range of expressions, such as anger, happiness, sadness, etc. By ensuring that the expression parameters of the 3D models of the first and second objects are the same, the development team can use the same expression control system to drive all characters, saving development time and resources.

[0071] For example, in a virtual reality application scenario, the first object is a real-world person. The second object is a character in the virtual world. During the VR experience, users may want to interact with the virtual character, and these characters need to be able to display expressions similar to real-world people. By ensuring that the expression parameters of the 3D models of the first and second objects are identical, developers can create a more realistic and immersive virtual environment, enhancing the user experience.

[0072] For example, the first object is a 3D head model of a healthy person, and the second object is a 3D head model of a patient. In medical research, researchers may need to compare the facial expressions of healthy and patient individuals to assess the impact of a disease on the patients' facial muscles. By ensuring that the expression parameters of the first and second 3D models are identical, researchers can use the same criteria to analyze and compare the facial expressions of both.

[0073] For example, in an animation production scenario, the first object is a 3D animated character, and the second object is another 3D animated character, perhaps of a different race or species. In an animated film, the director may want all characters to express a rich range of emotions. By ensuring that the expression parameters of the 3D models of the first and second objects are identical, the animator can use the same expression animation to drive all characters, maintaining consistency and fluidity in the animation.

[0074] In step 102 , a second three-dimensional model of a first object is constructed based on a first image of the first object.

[0075] In some embodiments, one or more first images of a first object are required. These first images can be photos, video frames, or other forms of two-dimensional image data. Computer vision technology is used to detect key feature points in the image, such as the position and shape of eyes, nose, mouth, etc. The expression parameters of the first object in the image are analyzed to determine its specific expression parameter characteristics. Start with a general three-dimensional face model, or use a predefined model that is close to the first object as a starting point. Adjust the geometry of the model based on the detected feature points and expression parameter characteristics. Map the color and texture information in the image to the three-dimensional model to enhance the realism of the model.

[0076] In some embodiments, the above-mentioned construction of the second three-dimensional model of the first object based on the first image of the first object can be achieved by: performing a first feature extraction on the first image to obtain the facial shape features of the first object in the first image; performing a second feature extraction on the first image to obtain the expression parameter features of the first object in the first image; and constructing the second three-dimensional model of the first object based on the facial shape features and the expression parameter features.

[0077] In some embodiments, the first image is preprocessed, including noise reduction and contrast enhancement, to improve feature extraction accuracy. Computer vision algorithms (e.g., SIFT, SURF, ORB, etc.) are used to detect facial key points of the first subject in the image, such as the corners of the eyes, the tip of the nose, and the corners of the mouth. Based on the detected feature points, a shape description method (e.g., contour, geometric shape, etc.) is used to describe the facial shape of the first subject. An expression parameter recognition algorithm (e.g., a convolutional neural network (CNN) based on deep learning) is used to identify the expression parameter category (e.g., smile, surprise, anger, etc.) of the first subject in the first image. The identified expression parameters are converted into a set of parameters that describe the intensity and specific characteristics of the expression parameters, such as the degree of upturned corners of the mouth or wrinkles at the corners of the eyes. Starting from a general 3D face model or a predefined model that approximates the first subject, the extracted facial shape features are applied to the initialized 3D model, and the model's geometry is adjusted to match the facial shape of the first subject. The extracted expression parameter features are then applied to the adjusted 3D model, driving the model to exhibit the same first expression parameters as in the first image. Color and texture information in the image is mapped onto the 3D model to enhance the model's realism. Optimize the model's details, such as adding wrinkles and highlights, to further improve the model's realism.

[0078] In some embodiments, facial shape features and expression parameter features are two important aspects of describing the facial appearance and expression of a character in a three-dimensional model or image. They are widely used in fields such as computer graphics, computer vision, virtual reality, and game development to build and manipulate three-dimensional character models. Facial shape features refer to the characteristic points, contour lines, and surface shapes that constitute the geometric structure of a character's face. Facial shape features include: facial contour: such as the shape and position of the forehead, cheekbones, and chin. The position, shape, and size of facial features such as the eyes, nose, mouth, and ears. Facial symmetry: such as the degree of symmetry between the left and right cheeks. Facial concavity and convexity: such as the height of the nose bridge and the prominence of the cheekbones.

[0079] In some embodiments, expression parameter features refer to a set of parameters that control changes in a character's facial expressions, including muscle movement parameters: such as the degree of contraction and relaxation of facial muscles; expression intensity parameters: such as the degree of exaggeration of expressions such as smiling, frowning, and surprise; and expression duration and transition parameters: such as the speed and smoothness of expression changes.

[0080] In some embodiments, a first feature extraction is performed on the first image, and the facial shape features of the first subject are extracted using computer vision techniques (such as feature detection and key point recognition). A second feature extraction is performed on the first image, and the expression parameter features of the first subject are extracted. This may involve analyzing facial muscle movement, evaluating the intensity of expression, etc. The extracted facial shape features and expression parameter features are combined, and a second three-dimensional model of the first subject is constructed using three-dimensional modeling software or an algorithm. This model not only has the same facial shape as the first subject, but can also simulate and express the expression of the first subject by adjusting the expression parameter features.

[0081] In some embodiments, the first image refers to a two-dimensional image containing the first object and its first expression parameter, which can be a photo, video frame, etc. Information about the facial geometry of the first object extracted from the first image using image processing technology includes facial contours, the location and shape of feature points, etc. Expression parameter features refer to information about the first expression parameters of the first object extracted from the first image using expression parameter recognition technology, including the type, intensity, and specific characteristics of the expression parameters. The second three-dimensional model refers to a three-dimensional model constructed based on the facial shape features and expression parameter features in the first image, which can reproduce the facial shape and expression parameters of the first object. Feature extraction is the process of identifying and extracting meaningful features from an image. These features can be shapes, colors, textures, expression parameters, etc.

[0082] As an example, consider a photo of a person smiling. The goal is to construct a 3D model based on this photo that not only captures the facial shape of the person in the photo but also captures the facial expression parameters of the smile. First, preprocess the photo, adjusting brightness and contrast to ensure that facial features are clearly visible. Computer vision algorithms, such as facial landmark detection, are used to detect facial landmarks, such as the corners of the eyes, the tip of the nose, and the corners of the mouth. Based on the detected landmarks, shape description methods, such as contours or geometric shapes, are used to describe the facial shape of the person in the photo. An expression parameter recognition algorithm, such as a deep learning-based CNN, is used to identify the facial expression parameters of the person in the photo. In this example, the algorithm identifies that the person is smiling. The smile expression parameters are converted into a set of parameters that describe the intensity and specific features of the smile, such as the upward tilt of the mouth corners and the wrinkles at the corners of the eyes. Starting with a general 3D face model that has basic facial shape and expression parameters, the extracted facial shape features are applied to the initialized 3D model, adjusting the model's geometry to match the facial shape of the person in the photo. The extracted smile expression parameter features are applied to the adjusted 3D model, driving the model to display the smile expression parameters. The color and texture information in the photo are mapped onto the 3D model, ensuring that the model's skin color, lighting, and other effects are consistent with the person in the photo. Detailed optimization of the 3D model, such as adding wrinkles and highlights, is performed to further enhance the model's realism.

[0083] By accurately extracting facial shape and expression parameters, the constructed 3D model can highly reproduce the facial features and expression parameters of the original subject, making the virtual character appear more realistic and natural. In virtual reality and game development, this can enable the virtual character to respond to the user's expression parameters and movements in real time, enhancing the user's immersion and interactive experience. By extracting the facial features and expression parameters of different individuals, personalized 3D models can be customized for each user to meet their needs.

[0084] In some embodiments, constructing a second three-dimensional model of the first object based on the facial shape features and the expression parameter features can be achieved in the following manner: based on the facial shape features and the expression parameter features, driving parameters are predicted by a parameter prediction model to obtain driving parameters; based on the driving parameters, the first three-dimensional model is adjusted to obtain a second three-dimensional model of the first object.

[0085] In some embodiments, the above-mentioned driving parameters include driving parameters of key points on the first exclusive three-dimensional model. The above-mentioned adjustment of the first three-dimensional model based on the driving parameters to obtain the second three-dimensional model of the first object can be achieved as follows: controlling the key points in the first three-dimensional model, adjusting the position according to the driving parameters, and obtaining the second three-dimensional model of the first object.

[0086] In some embodiments, driving parameters refer to a set of parameters used to drive the movement and deformation of a 3D model in animation and motion control. These driving parameters can control the model's joint angles, muscle movements, and facial expressions, enabling the model to display various movements and expressions. Driving parameters are typically associated with the model's skeletal structure, muscular system, or facial expression control system. Adjusting these parameters enables precise control of the model. A parameter prediction model uses facial shape features and facial expression parameter characteristics to predict driving parameters. This process may involve machine learning algorithms, such as regression models and neural networks, which can learn and understand the relationship between facial features and driving parameters to predict appropriate driving parameters. Driving parameters include driving parameters for key points on the first 3D model. These key points are typically important locations on the model, such as joints, muscle attachment points, and facial expression control points. By adjusting the driving parameters of these key points, the model's movement and deformation can be controlled. Based on the predicted driving parameters, the positions of the key points in the first 3D model are adjusted. This can be achieved through animation software or a programming interface. The adjusted model can then display movements and expressions that match the facial shape features and facial expression parameter characteristics. By adjusting the positions of the key points, a second 3D model of the first object is generated. The prediction and adjustment process of driving parameters can realize automatic control of the three-dimensional model, enabling the model to automatically generate corresponding animations and expressions based on the input facial shape features and expression parameter features. It has broad application prospects in virtual reality, game development, film and television animation and other fields.

[0087] In some embodiments, key points on the first three-dimensional model need to be identified. These key points are the most important parts of the model, and their movement and deformation have a decisive impact on the overall appearance and movement of the model. The predicted driving parameters are applied to these key points. The driving parameters may include position offset, rotation angle, scaling ratio, etc., which determine the movement and deformation of the key points in space. Through programming or animation software, the key points in the first three-dimensional model are controlled to adjust their positions according to the driving parameters by adjusting the coordinates of the key points, rotating the key points in the direction, or changing the scaling ratio of the key points. Once the positions of the key points are adjusted, the geometry of the entire model is also updated. The movement and deformation of the key points are transmitted to other parts of the model through the model's skeletal structure or musculature. Through these adjustments, the resulting second three-dimensional model of the first object will reflect the new facial shape features and expression parameter characteristics. This new model not only retains the basic shape of the first object but can also express different movements and expressions through the driving parameters.

[0088] In some embodiments, the parameter prediction model is a pre-trained model, typically based on a machine learning algorithm such as a neural network. The model's learning goal is to understand how facial shape features and expression parameter features affect the positions of key points on the three-dimensional model. Using a large amount of training data, the model learns to predict corresponding driving parameters from the facial shape features and expression parameter features. Facial shape features and expression parameter features are extracted from the first image and include specific features of the facial geometry and expression parameters. Facial shape features may include facial contours, the position and shape of facial features, while expression parameter features may include the lift of eyebrows, the upturn of mouth corners, the degree of eye openness, etc. The extracted facial shape features and expression parameter features are input into the parameter prediction model, which outputs a set of driving parameters. These parameters are instructions for controlling the positions of key points in the first three-dimensional model, determining how the key points should move to reflect the facial shape and expression parameters of the first subject. The first three-dimensional model is an existing three-dimensional model that contains a series of key points whose positions determine the model's shape and expression parameters. By adjusting the positions of these key points according to the driving parameters, the model's shape and expression parameters can be altered to more closely resemble the facial features and expression parameters of the first subject. After adjusting the positions of the key points, the first 3D model is transformed into a second 3D model of the first object. This model now has the facial shape and expression parameter characteristics of the first object and can be used in applications such as virtual reality, game development, and film production.

[0089] As an example, let's assume we're developing a virtual reality (VR) game where players can create their own avatars and interact with them in-game. To provide a personalized experience, we want players to be able to upload their own photos. The system then automatically generates a 3D avatar that looks similar to the player and can express various facial expressions. The player uploads a photo of themselves. The system processes the photo and extracts the player's facial shape and facial expression parameters. This may include using computer vision techniques to detect facial key points, such as the corners of the eyes, the tip of the nose, and the corners of the mouth, as well as identifying facial expressions such as smile and surprise in the photo. Using a pre-trained parameter prediction model, the extracted facial shape and facial expression parameters are input into the model. The model analyzes these features and predicts corresponding driving parameters. These parameters are instructions for controlling the positions of key points on the avatar model. The avatar model (the first 3D model) already contains a series of key points, whose positions determine the avatar's shape and facial expressions. Based on the predicted driving parameters, the system adjusts the positions of these key points. For example, if the corners of the mouth in the player's photo are raised, the model adjusts the positions of the key points so that the avatar displays the facial expression parameters of a smile. After adjusting the positions of the key points, the first 3D model is converted into a second 3D model of the player. This model now has the player's facial shape and expression parameter characteristics and can be used as a virtual character in the game.

[0090] In this way, the parameter prediction model is used to predict the driving parameters of key points on the first three-dimensional model based on facial shape features and expression parameter features, and the key points are controlled to adjust their positions according to the driving parameters, thereby constructing a second three-dimensional model of the first object. This method has significant beneficial effects. This method first uses machine learning technology to learn the mapping relationship between facial shape and expression parameter features and the positions of key points in the three-dimensional model, so that the predicted driving parameters can accurately reflect the facial features and expression parameters of the first object. Secondly, by adjusting the positions of key points in a quantitative manner, the efficiency and consistency of model construction are greatly improved, reducing the need for manual adjustment and possible errors. It can also process large amounts of data and quickly generate multiple personalized three-dimensional models to meet the needs of different application scenarios.

[0091] In some embodiments, before obtaining the driving parameters through the parameter prediction model based on the facial shape features and the expression parameter features, the following processing can also be performed: based on the image samples, the driving parameters are predicted through the initial parameter prediction model to obtain the driving parameters of the third object in the image samples; based on the driving parameters of the third object and the driving parameter label of the third object, the initial parameter prediction model is trained to obtain the parameter prediction model.

[0092] In some embodiments, driving parameter labels refer to manually annotated or automatically calculated driving parameters for a third object in an image sample. These labels provide a correspondence between the facial shape and expression parameter features of the object in the image and the driving parameters. Driving parameter labels typically include information such as the location of key points, the intensity of the expression, and the state of muscle movement. They serve as the basis for training models to learn how to predict driving parameters from image features.

[0093] In some embodiments, image samples refer to a collection of images used to train and validate the parameter prediction model. These images can be still photos or frames from a video. Image samples should encompass a variety of scenes, lighting conditions, expressions, and poses to ensure good model generalization. Each image sample should include one or more third objects, annotated with their driving parameter labels.

[0094] In some embodiments, a set of image samples is prepared. These samples contain facial images of a third subject, along with labels for the driving parameters of the third subject in these images. These labels are pre-labeled and represent the keypoint driving parameters corresponding to the facial shape and expression parameters of the third subject in the images. An initial parameter prediction model is created. This model can be a neural network or other machine learning model suitable for regression tasks. The model's structure and parameters are randomly initialized before training. The initial parameter prediction model is trained using the prepared image samples and the corresponding driving parameter labels. During training, the model adjusts its parameters using an optimization algorithm (such as gradient descent) to minimize the difference between the predicted driving parameters and the actual labels. Various optimization techniques, such as regularization, dropout, and learning rate scheduling, may be used during training to prevent model overfitting and improve model generalization. After training is complete, the model needs to be evaluated using a validation set that was not used in training to test its performance. If the model performance is unsatisfactory, it may be necessary to adjust the model structure, optimize the algorithm parameters, or add more training data and then retrain the model. After training and optimization, the obtained model can predict driving parameters of each key point on the first three-dimensional model of the first object based on the input facial shape features and expression parameter features.

[0095] As an example, suppose you are developing an intelligent assistant application that generates a virtual assistant that looks like the user and can interact with them using facial expressions based on their photos. To achieve this functionality, you need to train a parameter prediction model that predicts the driving parameters of key points in the virtual assistant model based on the user's facial shape and expression parameters. Collect a set of image samples containing different faces with varying shapes and expression parameters. Annotate each image sample with corresponding key point driving parameter labels. These labels indicate the key point locations corresponding to the facial shape and expression parameters of the face in the image. Create an initial parameter prediction model, such as a convolutional neural network (CNN), to learn the mapping between facial shape and expression parameters and key point driving parameters. Train the initial parameter prediction model using the prepared image samples and the corresponding driving parameter labels. During training, the model adjusts its parameters using an optimization algorithm to minimize the difference between the predicted driving parameters and the actual labels. Various optimization techniques, such as data augmentation, regularization, and dropout, are used during training to improve the model's generalization and prevent overfitting. After training is complete, the model's performance is evaluated using a validation set that was not used in the training. If the model's performance is unsatisfactory, it may be necessary to adjust the model structure, optimize the algorithm parameters, or add more training data and retrain the model. After training and optimization, the resulting parameter prediction model can predict the user's key point driving parameters based on the input facial shape features and expression parameter features.

[0096] In this way, an initial parameter prediction model was constructed. This model predicts the key point driving parameters of the third object based on the facial shape and expression parameter features of the third object in the image sample. These predicted driving parameters are then compared with the actual driving parameter labels. The initial model is trained and optimized using supervised learning methods, resulting in a parameter prediction model capable of accurately predicting the key point driving parameters of the first object. This process is technically rigorous, and through extensive data training and model optimization, the accuracy and generalization ability of the parameter prediction model are ensured. Ultimately, this parameter prediction model can, in practical applications, predict the driving parameters of each key point on the first 3D model based on the facial shape and expression parameter features of the first object, thereby constructing a highly realistic and personalized 3D model, improving the efficiency and accuracy of model construction.

[0097] In step 103 , deformation parameters are determined based on the second three-dimensional model and the first three-dimensional model.

[0098] In some embodiments, the first three-dimensional model and the third three-dimensional model of the second object have second expression parameters, the second three-dimensional model includes multiple first surface elements, the first three-dimensional model includes second surface elements corresponding to each of the first surface elements one by one, and the deformation parameters include sub-deformation parameters corresponding to each of the first surface elements, and the sub-deformation parameters are used to indicate the degree of deformation of the first surface element compared to the corresponding second surface element.

[0099] In some embodiments, both the first three-dimensional model and the third three-dimensional model of the second object are capable of expressing a second expression. The "second expression parameter" here may refer to a specific facial expression state, such as smile, surprise, anger, etc., or another expression parameter relative to the first expression parameter. The second-dimensional model is composed of multiple facets, which are the basic units that make up the model surface. Each facet is typically composed of three or more vertices, forming a planar polygon, such as a triangle or quadrilateral. The first three-dimensional model also contains facets, which correspond one-to-one with the first facets in the second three-dimensional model. This means that each first facet has a corresponding second facet in the first three-dimensional model, and they correspond in structure and position to each other. Deformation parameters are a set of parameters that describe the degree of deformation of a model. They can control how various parts of the model deform from one state to another. Sub-deformation parameters are part of the deformation parameters and are specifically used to describe the degree of deformation of each first facet.

[0100] In some embodiments, the first three-dimensional model of the first object refers to a pre-existing three-dimensional model that represents the shape of the first object in a certain specific state. This model will be used as a reference to determine the deformation parameters of another model (the second three-dimensional model). The first three-dimensional model and the third three-dimensional model of the second object have second expression parameters, and both the first three-dimensional model and the third three-dimensional model exhibit the second expression parameters. The second expression parameters may be different from the first expression parameters, or they may be different intensities or variations of the same expression parameters. The second three-dimensional model is composed of multiple small three-dimensional units (which may be polygons, grid units, etc.). These units together constitute the shape of the entire model. The first three-dimensional model has a similar structure, and each of its units corresponds to a unit in the second three-dimensional model. This correspondence is the basis for calculating the deformation parameters. The deformation parameters are a set of numerical values ​​used to describe how to deform the second three-dimensional model into the first three-dimensional model. Each first face element has a corresponding sub-deformation parameter, which determines how and to what extent the unit changes during the deformation process.

[0101] In some embodiments, deformation parameters, also known as deformation gradients, are a set of numerical values ​​or matrices that describe how to deform the geometric shape of a 3D model (source model) into the corresponding shape of another target 3D model. Each model unit (such as a vertex, patch, or local area) has a corresponding sub-deformation parameter that controls the unit's displacement, rotation, scaling, and other transformations during the deformation process. Deformation parameters are the bridge connecting the geometric shapes of two 3D models. Their core is to calculate the transformation rules through local feature comparison or physical simulation.

[0102] In some embodiments, the above step 102 can also be implemented in the following manner: perform the following processing on each first surface element: perform feature extraction on the first surface element to obtain the first shape feature of the first surface element, perform feature extraction on the second surface element corresponding to the first surface element to obtain the second shape feature of the second surface element; based on the first shape feature and the second shape feature, determine the sub-deformation parameter corresponding to the first surface element.

[0103] In some embodiments, the first shape feature includes multiple first sub-shape features, where the first sub-shape features are used to indicate the coordinate differences between different vertices on the first facet, and the second shape feature includes multiple second sub-shape features corresponding one-to-one to the first sub-shape features.

[0104] In some embodiments, the above-mentioned determination of the sub-deformation parameter corresponding to the first surface element based on the first shape feature and the second shape feature can be achieved by multiplying the second shape feature and the inverse feature of the first shape feature to obtain the sub-deformation parameter corresponding to the first surface element.

[0105] In some embodiments, the first sub-shape feature refers to a parameter used to describe the coordinate difference between different vertices on the first face element in the first three-dimensional model. These features are usually expressed as vectors, reflecting the relative positional relationship between the vertices. For example, if a face element has three vertices, the first sub-shape feature may include a vector from vertex 1 to vertex 2, a vector from vertex 2 to vertex 3, and a vector from vertex 3 to vertex 1. The second sub-shape feature refers to a parameter that corresponds one-to-one to the first sub-shape feature, and is used to describe the coordinate difference between different vertices on the second face element in the second three-dimensional model. These features are also expressed as vectors, reflecting the relative positional relationship between the vertices in the second model. By comparing the first sub-shape feature and the second sub-shape feature, the similarities and differences in shape between the two models can be analyzed. These features are very important in fields such as shape matching, model registration, and deformation animation. They can help the algorithm understand the geometric structure of the model and perform corresponding operations.

[0106] In some embodiments, the second shape feature refers to a parameter that describes the coordinate difference between different vertices on the second surface element in the second three-dimensional model. These features are usually expressed as vectors and reflect the relative positional relationship between the vertices in the second model. The inverse feature of the first shape feature refers to the reciprocal or inverse operation result of the coordinate difference between different vertices on the first surface element in the first three-dimensional model. These inverse features are used to describe the opposite change in the relative positional relationship between the vertices in the first model. By multiplying the second shape feature and the inverse feature of the first shape feature, the sub-deformation parameter corresponding to the first surface element can be obtained. This process can be understood as applying the shape feature of the second model to the inverse feature of the first model to calculate how the first surface element is deformed to match the shape of the second surface element. The sub-deformation parameter reflects the degree of deformation of the first surface element relative to the second surface element, including changes such as displacement, scaling, and rotation. These parameters can be used to drive the deformation of the three-dimensional model so that the first model can exhibit a shape and expression similar to the second model.

[0107] In some embodiments, a first element is a three-dimensional unit in the target model, which may be a polygon, mesh, or other unit, and constitutes part of the entire model. A second element is a three-dimensional unit in the reference model, corresponding one-to-one with the first element and providing reference information for deformation. The first shape feature analyzes the shape of the first element to extract features describing its geometric properties. These features may include unit size, angles, curvature, etc. The second shape feature performs a similar analysis on the shape of the second element to extract features describing its geometric properties. Based on the extracted first and second shape features, the parameters required to transform the first element into the second element are calculated. These parameters may include the magnitude and direction of transformations such as translation, rotation, and scaling. Sub-deformation parameters are local deformation parameters for each model unit, which act together on the entire model to achieve global deformation from the second three-dimensional model to the first three-dimensional model. Using geometric analysis or computer vision techniques, shape features are extracted from each model unit. The features of the first element are compared with those of the second element to identify differences. Based on these differences, the deformation parameters required to be applied to the first element are calculated to bring its shape closer to that of the second element. The calculated sub-deformation parameters are applied to the first surface element to achieve local shape deformation.

[0108] In some embodiments, the first shape feature includes multiple first sub-shape features, where the first sub-shape features are used to indicate the coordinate differences between different vertices on the first facet, and the second shape feature includes multiple second sub-shape features corresponding one-to-one to the first sub-shape features.

[0109] As an example, when the shape of the first surface element is a triangle, the expression of the first shape feature of the first surface element may be:

[0110] V=(V2-V1, V3-V2)(1)

[0111] Wherein, V is used to indicate the first shape feature, and V1, V2 and V3 are used to indicate the vertex coordinates of the first surface element.

[0112] As an example, when the shape of the second surface element is a triangle, the expression of the second morphological feature of the second surface element may be:

[0113] V'=(V2'-V1', V3'-V2') (2)

[0114] Wherein, V' is used to indicate the second shape feature, and V1', V2' and V3' are used to indicate the vertex coordinates of the second surface element.

[0115] In some embodiments, the above-mentioned determination of the sub-shape parameter corresponding to the first surface element based on the first shape feature and the second morphological feature can be achieved by multiplying the second shape feature and the inverse feature of the first shape feature to obtain the sub-deformation parameter corresponding to the first surface element.

[0116] As an example, the expression of the sub-deformation parameter of the first surface element can be:

[0117] A=V'V- 1 (3)

[0118] A is used to indicate the sub-deformation parameter of the first surface element, V' is used to indicate the second shape feature, and V-1 is used to indicate the inverse feature of the first shape feature.

[0119] As an example, suppose you're producing an animated film in which a character needs to display a variety of complex facial expressions. To achieve this effect, you create a 3D model of the character using 3D modeling software, and then use deformation techniques to generate different facial expressions for the character. The second 3D model is the character's basic facial expression parameter model, typically a neutral expression. The first 3D model is a reference model for the character's specific facial expressions, such as a smile or surprise. Ensure that the mesh structure of the second and first 3D models is consistent, with each first element corresponding to a second element. Feature extraction is performed on the shape of each first element to obtain first shape features. These features may include element size, angle, curvature, etc. Feature extraction is performed on the shape of each corresponding second element to obtain second shape features. The first shape features of each first element are compared with the second shape features of each second element to identify differences. Based on these differences, the sub-deformation parameters that need to be applied to the first element are calculated to make its shape resemble that of the second element. These parameters may include the magnitude and direction of transformations such as translation, rotation, and scaling. The calculated sub-deformation parameters are applied to the first element to achieve localized shape deformation. This process is repeated to deform all model units, and finally a specific expression parameter model of the character is generated.

[0120] In this way, by extracting the shape features of each first surface element and the shape features of the corresponding second surface element, and determining the sub-deformation parameters based on these features, efficient and accurate three-dimensional model deformation can be achieved. This method first captures subtle shape changes by quantitatively analyzing the geometric characteristics of the model unit, providing precise guidance for deformation. Then, by comparing the shape features of the first surface element and the second surface element, the required deformation parameters are calculated. These parameters can accurately reflect the transition process from the second three-dimensional model to the first three-dimensional model. This feature-based deformation method not only improves the accuracy and naturalness of the deformation, but also greatly reduces the workload of manual adjustment and improves work efficiency.

[0121] In step 104, the third three-dimensional model is adjusted based on the deformation parameter to obtain a fourth three-dimensional model of the second object.

[0122] In some embodiments, the fourth three-dimensional model and the second three-dimensional model have the same expression parameters, the third three-dimensional model includes a plurality of third elements, and the deformation parameters include sub-deformation parameters corresponding to each of the third elements. The third three-dimensional model includes a plurality of third elements, the deformation parameters include sub-deformation parameters corresponding to each of the third elements, and the fourth three-dimensional model and the second three-dimensional model have the same expression parameters.

[0123] In some embodiments, an existing three-dimensional model (third three-dimensional model) is adjusted using deformation parameters to generate a new model (fourth three-dimensional model) having specific expression parameters. The third three-dimensional model contains multiple third face elements. These units may be polygons, mesh units, etc., which constitute the structure of the entire model. The deformation parameters are calculated through the previous process, and they describe how to deform the third three-dimensional model into a model with corresponding expression parameters. Each third face element has a corresponding sub-deformation parameter, which determines how and to what extent the unit changes during the deformation process. The deformation parameters are applied to the third three-dimensional model, and each third face element is deformed accordingly according to its corresponding sub-deformation parameter. Through these deformations, the third three-dimensional model is updated to a fourth three-dimensional model, and the adjusted model inherits the structure of the third three-dimensional model.

[0124] In some embodiments, the above-mentioned adjustment of the third three-dimensional model based on the deformation parameters to obtain the fourth three-dimensional model of the second object can be achieved as follows: for each of the third surface elements in the third three-dimensional model, based on the sub-deformation parameters corresponding to the third surface elements, the shape of the third surface element is adjusted to obtain the fourth surface element corresponding to the third surface element; each of the third surface elements in the third three-dimensional model is replaced by the corresponding fourth surface element to obtain the fourth three-dimensional model.

[0125] In some embodiments, fine control of the entire model is achieved by independently adjusting the shape of each unit in the model. Each third-surface element has a corresponding sub-deformation parameter, which accurately describes the manner and extent of the unit's change during the deformation process. By pre-calculating the deformation parameters, the model can be deformed automatically, reducing the workload of manual adjustment. This improves work efficiency, allowing artists and developers to generate models of various expression parameters and forms more quickly. Since the deformation of each model unit is based on its corresponding sub-deformation parameter, this method can maintain the overall consistency and integrity of the model. The deformed model is not only realistic in local details, but also maintains the original structure and proportions as a whole. It has good scalability and flexibility and can be applied to various complex model deformation tasks. Whether it is a simple change in expression parameters or a complex morphological deformation, it can be achieved by adjusting the corresponding deformation parameters.

[0126] In some embodiments, an existing three-dimensional model (a third three-dimensional model) is precisely adjusted using deformation parameters to generate a new model (a fourth three-dimensional model) with specific expression parameters or morphology. This method has broad application value in fields such as computer graphics, animation, and game development. The deformation parameters are derived through analysis and calculation and describe how to deform the third three-dimensional model into a fourth three-dimensional model with specific expression parameters or morphology. Each third-surface element has a corresponding sub-deformation parameter that precisely describes the manner and extent of the element's change during the deformation process. For each third-surface element in the third three-dimensional model, the shape of the element is adjusted based on its corresponding sub-deformation parameter. This may include transformation operations such as translation, rotation, and scaling. Through these adjustments, the fourth-surface element corresponding to each third-surface element is obtained. Each third-surface element in the third three-dimensional model is replaced with the corresponding fourth-surface element. In this way, the entire third three-dimensional model is updated to a fourth three-dimensional model, which now has specific expression parameters or morphology. A fourth three-dimensional model with specific expression parameters or morphology is generated.

[0127] As an example, suppose you're developing a role-playing game in which a character needs to display a variety of complex facial expressions. To achieve this, 3D modeling software is used to create a basic facial expression parameter model (a third 3D model). The goal is to use deformation technology to generate different facial expression parameters for the character (a fourth 3D model). The third 3D model is the character's basic facial expression parameter model, typically with neutral facial expression parameters. Deformation parameters: These parameters are calculated using the previous process and describe how to deform the third 3D model into a model with the first facial expression parameters. Each third element has a corresponding sub-deformation parameter. For each third element in the third 3D model, the unit's shape is adjusted based on its corresponding sub-deformation parameter. This may include transformations such as translation, rotation, and scaling. These adjustments yield the corresponding fourth element for each third element. Each third element in the third 3D model is then replaced with the corresponding fourth element. The entire third 3D model is updated to the fourth 3D model, now with the first facial expression parameters. Game developers can quickly generate models with various facial expression parameters for the character without having to manually adjust the position and shape of each model element. This not only improves the efficiency of animation production, but also ensures the naturalness and realism of the character's expression parameters.

[0128] In this way, by extracting and analyzing the shape features of each model unit in the third three-dimensional model, the corresponding sub-deformation parameters are calculated. These parameters accurately describe the way and degree of change of each model unit during the deformation process. Then, for each third surface element, based on its corresponding sub-deformation parameter, the shape of the unit is adjusted to obtain the corresponding fourth surface element. Finally, each third surface element in the third three-dimensional model is replaced with the corresponding fourth surface element, thereby obtaining a fourth three-dimensional model with specific expression parameters or morphology. The advantage of this method is that through precise control at the model unit level, fine adjustment of the entire model is achieved while maintaining the overall consistency and integrity of the model. In addition, by pre-calculating the deformation parameters, the model can be deformed automatically, reducing the workload of manual adjustment and improving work efficiency.

[0129] In step 105 , a second image of the second object is generated based on the fourth three-dimensional model.

[0130] In some embodiments, the expression parameters of the second object in the second image are the same as the expression parameters of the first object in the first image. The above-mentioned generation of the second image of the second object based on the fourth three-dimensional model can be achieved in the following manner: based on the third image of the second object, predicting the expected position of the key points in the fourth three-dimensional model; controlling the key points in the fourth three-dimensional model, adjusting the position according to the expected position to obtain the fifth three-dimensional model, and smoothing the fifth facet in the fifth three-dimensional model to obtain the sixth three-dimensional model; generating the second image of the second object based on the sixth three-dimensional model.

[0131] In some embodiments, the fourth three-dimensional model is key-point corrected to ensure that the positions and shapes of key parts of the model (such as eyes, mouth, nose, etc.) match the expected first expression parameters. This step may involve adjusting the geometry of the model to eliminate any inaccurate or unnatural features. Each fifth facet in the fifth three-dimensional model that has undergone key-point correction is smoothed to eliminate rough or sharp edges on the surface of the model, making the model look more natural and realistic. Based on the smoothed sixth three-dimensional model, a second image of the second object with the first expression parameters is generated. This step may involve using rendering technology to convert the three-dimensional model into a two-dimensional image, taking into account lighting, texture and other visual effects to enhance the realism of the image. By performing key-point correction and smoothing on the three-dimensional model, a high-quality image with natural expression parameters can be generated.

[0132] As an example, suppose a movie is being made in which a character needs to display sad expression parameters. A basic expression parameter model of the character (a third 3D model) has been created using 3D modeling software, and deformation parameters have been used to adjust it to a fourth 3D model with sad expression parameters. It is found that the positions of the character's eyes and mouth in the fourth 3D model do not match the expected sad expression parameters. Using the key point editing tool in the modeling software, the key points of the fourth 3D model are adjusted to make the eyes droop and the mouth look like crying. When the fifth 3D model is checked, it is found that there are some rough areas on the model surface, especially around the eyes and mouth. Using the smoothing tool in the modeling software, the fifth facet in the fifth 3D model is smoothed. Through these processes, a sixth 3D model is obtained, the surface of this model is smoother and looks more natural. The model is rendered to generate a second image of the second object with sad expression parameters.

[0133] In this way, key point correction ensures that the key parts of the model (such as eyes, mouth, etc.) match the expected expression parameters. By adjusting the position and shape of these key points, the expression parameters of the model can be made more natural and realistic. Model smoothing eliminates the roughness and sharp edges of the model surface. Through smoothing, the surface of the model is smoother and looks more natural. Finally, based on the sixth three-dimensional model that has undergone key point correction and smoothing, a second image of the second object with the first expression parameters is generated to ensure the quality and realism of the image. Through rigorous key point correction and smoothing processes, the quality and realism of the generated image are ensured, providing users with a more immersive experience.

[0134] In some embodiments, computer vision technology or machine learning algorithms are used to analyze a third image of a second object having second expression parameters. Through image processing, the expected positions of key points in the fourth three-dimensional model in the third image are predicted. These key points may include facial feature points such as eyes, mouth, and nose. The predicted expected positions are mapped back to the spatial coordinate system of the fourth three-dimensional model. Modeling software or a programming interface is used to control the key points in the fourth three-dimensional model so that they are adjusted according to the expected positions. The positions of the key points are adjusted, and the fourth three-dimensional model is updated to a fifth three-dimensional model. This ensures that the key parts of the model match the expected expression parameters, thereby improving the realism and expressiveness of the model's expression parameters. The generated fifth three-dimensional model is visually inspected to ensure that the position adjustment of the key points achieves the expected effect. If necessary, the model can be further optimized and other parameters can be adjusted to improve the overall quality of the model.

[0135] As an example, let's assume we're creating a character for an animated film that needs to display facial expressions of surprise. We've already created a basic model of the character's facial expressions (the fourth 3D model) and now need to adjust the model through keypoint calibration to achieve the desired parameters. First, find a photograph or image of the character displaying the desired parameters of surprise (the third image). Using computer vision techniques, such as feature detection and tracking algorithms, we predict the locations of key points on the character's face in the image, such as those for eyes wide open and mouth open. These predicted keypoint locations will serve as a reference for adjusting the fourth 3D model. We map the predicted keypoint locations into the 3D space of the fourth 3D model. Using 3D modeling software, we select and adjust keypoints in the fourth 3D model so that their positions match the predicted desired positions. For example, we adjust the eyes to widen and the mouth to open. After completing the keypoint position adjustments, the fourth 3D model is updated to the fifth 3D model. The fifth 3D model should now display the desired parameters of surprise, with the key points aligned with the expected parameters in the third image. Visually inspect the fifth 3D model to ensure that the keypoint adjustments have achieved the desired results. If necessary, you can further adjust other parts of the model, such as wrinkles, muscles, etc., to enhance the naturalness of the surprise expression parameters.

[0136] In this way, the process of obtaining the fifth 3D model by performing key point correction on the fourth 3D model can significantly improve the realism and expressiveness of the 3D model's expression parameters. First, based on a third image of a second object with second expression parameters, we use computer vision technology to predict the expected positions of key points in the fourth 3D model. These predicted positions reflect the specific characteristics of the second expression parameters, such as eye opening and mouth shape. We then control the key points in the fourth 3D model and adjust them according to these expected positions. In this way, we can precisely adjust the key parts of the model to match the second expression parameters. This not only improves the realism of the model's expression parameters but also enables the model to more accurately convey emotion and intent. Furthermore, this method is efficient and repeatable, and can be quickly applied to multiple models and expression parameter adjustments. By performing key point correction on the fourth 3D model to obtain the fifth 3D model, we not only improve the realism and expressiveness of the model's expression parameters, but also offer efficiency and repeatability, providing strong support for 3D model creation and animation.

[0137] In some embodiments, the fifth face element in the fifth three-dimensional model corresponds one-to-one to the third face element in the third three-dimensional model, and the fifth face element in the fifth three-dimensional model is smoothed to obtain a sixth three-dimensional model. This can be achieved in the following way: for each fifth face element in the fifth three-dimensional model, based on the third face element corresponding to the fifth face element in the third three-dimensional model, the fifth face element is smoothed to obtain the sixth face element corresponding to the fifth face element; according to the arrangement of each fifth face element in the fifth three-dimensional model, each sixth face element is constructed into the sixth three-dimensional model.

[0138] In some embodiments, it is confirmed that each fifth face element in the fifth three-dimensional model corresponds one-to-one to the third face element in the third three-dimensional model. This step ensures that each unit has a reference model unit when smoothing. Each fifth face element in the fifth three-dimensional model is smoothed based on the corresponding third face element in the third three-dimensional model. This may involve calculating the differences between the two model units and then applying a smoothing algorithm to reduce these differences so that the fifth face element is closer to the shape and surface characteristics of the third face element. By smoothing each fifth face element, a corresponding sixth face element is generated. This process may include adjusting the vertex position, normal direction or other geometric properties to achieve the smoothing effect. According to the arrangement of the fifth face elements in the fifth three-dimensional model, the generated sixth face elements are recombined to construct a sixth three-dimensional model. The generated sixth three-dimensional model is visually inspected to ensure that the smoothing process has achieved the expected effect. If necessary, the model can be further adjusted and the parameters of the smoothing algorithm can be optimized to improve the quality of the model.

[0139] As an example, imagine a 3D animator creating a character model for an animated film. A base model of the character (the third 3D model) has already been created, and a fifth 3D model with specific expression parameters has been generated based on this model. Now, the goal is to improve the visual quality of the fifth 3D model through smoothing, making it appear smoother and more natural. First, ensure that each model element in the fifth 3D model (such as the head, arms, and legs) corresponds one-to-one with the corresponding model element in the third 3D model. This step ensures that each element has a reference model element when smoothing. Each model element in the fifth 3D model is smoothed based on the corresponding model element in the third 3D model. For example, if the surface of the head model element in the fifth 3D model is uneven, a smoothing algorithm can be used to reduce these unevenness and make the head appear smoother. By smoothing each model element, a corresponding sixth facet is generated. For example, the smoothed head model element becomes the head model element in the sixth 3D model. The generated sixth facets are reassembled according to the arrangement of the model elements in the fifth 3D model to construct the sixth 3D model. This step ensures that the geometry and topology of the entire model remain consistent, and that each element has been smoothed. Visually inspect the resulting 3D model to ensure that the smoothing process has achieved the desired results. If necessary, further adjustments can be made to the model and the smoothing algorithm parameters can be optimized to improve the model quality.

[0140] In this way, the one-to-one correspondence of model units ensures that each fifth facet has a clear reference standard, namely the third facet, which provides precise guidance for the smoothing process. Then, by smoothing the fifth facet, the unevenness and sharp edges of the model surface can be reduced, making the model appear smoother and more natural. This smoothing process not only improves the visual effect of the model, but also enhances the model's realism, making it more expressive in applications such as animation and games. Finally, according to the arrangement of the various model units in the fifth three-dimensional model, the smoothed sixth facets are reassembled to construct the sixth three-dimensional model. This process maintains the overall structure and topological relationships of the model while ensuring that the smoothing effect of each unit is reflected. Therefore, by smoothing the fifth facet, the resulting sixth three-dimensional model has significantly improved visual quality and realism, providing strong technical support for the development of the digital media and entertainment industry.

[0141] In some embodiments, there are multiple first images, and the second images correspond one-to-one to the first images. After generating the second image of the second object based on the fourth three-dimensional model, the following processing can also be performed: the audio of the first object corresponding to the first image is fused with the corresponding second image to obtain a video of the second object.

[0142] In some embodiments, audio corresponding to the first subject is extracted from the first image. This may involve audio processing techniques, such as speech recognition and audio separation, to ensure that the extracted audio matches the motion and expression parameters of the first subject. The extracted audio is then processed as necessary, such as noise reduction, equalization, and compression, to improve audio quality and ensure its compatibility with the visual content of the second image. The quality and format of the second image are ensured to be suitable for integration with the audio. This may involve adjusting the image resolution and color correction. The audio is synchronized with the second image to ensure temporal alignment with the motion and expression parameters in the image. This may involve adjusting the timelines of the audio and video to achieve precise synchronization. Using video editing software, the processed audio is combined with the second image to generate a video of the second subject. This step ensures a seamless integration of audio and image, creating a complete multimedia content. The generated video is then post-produced, such as by adding special effects, subtitles, and transitions, to enhance its visual appeal and professionalism. This effectively integrates the audio of the first subject corresponding to the first image with the second image to produce a video of the second subject.

[0143] As an example, imagine an animator working on animating a character for an animated film. A second image of a second object with first expression parameters has been generated using a fourth 3D model. Now, the animator wishes to fuse the audio of the first object corresponding to the first image with the second image to produce a video of the second object. The animator extracts the audio corresponding to the first object from the first image. This may involve using professional audio processing software, such as Adobe Audition, to separate and extract the audio. The extracted audio is then processed, such as removing background noise, adjusting volume, and equalizing, to ensure that the audio quality matches the visual content of the second image. The animator also ensures that the quality and format of the second image are suitable for integration with the audio. This may involve adjusting the image's resolution, color, and contrast to match the film's overall visual style. The animator synchronizes the audio with the second image, ensuring that the audio is temporally aligned with the motion and expression parameters in the image. This may involve adjusting the audio and video timelines in video editing software. Using video editing software, such as Adobe Premiere Pro, the processed audio is composited with the second image to produce a video of the second object. The resulting video is then post-processed, such as adding special effects, subtitles, and transitions, to enhance its visual appeal and professionalism. Perform visual and auditory checks on the resulting video to ensure that the audio and image blend as expected. If necessary, further adjustments can be made to the audio and video parameters to optimize the final result.

[0144] In this way, the second image generated based on the fourth three-dimensional model already has the visual expression of the first expression parameter, while the audio of the first object contains sound information related to the expression parameter, such as conversation, laughter or crying. Fusion of the two can achieve synchronization of sound and image, allowing viewers to experience a more natural and realistic experience when watching the video. This fusion not only enhances the appeal of the video, but also increases the audience's participation and immersion. In addition, in this way, different visual and sound elements can be combined to create richer and more diverse multimedia content to meet the needs of different audiences. Therefore, fusing audio and images to obtain video is a very effective way to create multimedia content, which can significantly improve the quality and viewing experience of the content.

[0145] In this manner, a first 3D model of a first object is determined, a second 3D model of the first object is constructed based on a first image of the first object, deformation parameters are determined based on the first and second 3D models, and a third 3D model is adjusted based on the deformation parameters to obtain a fourth 3D model of the second object. A second image is generated based on the fourth 3D model, and the deformation parameters are determined using the first 3D model of the first object as a reference. The first 3D model and the third 3D model of the second object both have the same expression parameters. By comparing the second 3D model with the first 3D model, deformation patterns based on the expression parameters can be identified. Based on the determined deformation parameters, the third 3D model is adjusted. The third 3D model of the second object is converted to the fourth 3D model. This adjustment is based on the deformation relationship between the second 3D model and the first 3D model, ensuring the accuracy and naturalness of the expression parameter transfer. A second image of the second object is generated based on the fourth 3D model. The adjusted fourth 3D model is converted into a visual image, completing the transfer of expression parameters from the first object to the second object. The entire process is based on the deformation relationship between models, ensuring the natural transition and accurate reproduction of expression parameters. It takes into account the differences in expression parameters between different objects, making the migrated expression parameters more consistent with the characteristics and expression parameter features of the second object, thereby effectively improving the accuracy of expression parameter migration.

[0146] Below, an exemplary application of the embodiment of the present application in an actual application scenario of migrating human facial expression parameters to animal facial labels will be described.

[0147] Training data collection. The collection of 3D models requires: 50 models, each with more than 2 UVs. A 3D model contains a mesh and multiple UV maps, such as Figure 10 As shown, Figure 10This is a schematic diagram of the data processing method provided in this embodiment. The same 3D model can be rendered into different visualizations using different UV maps. The human talking dataset requires: 500 IDs (persons), each containing 30 minutes of audio and video. The 3DMM face reconstruction algorithm requires: 10,000 facial images, 2D landmarks, and 3D landmarks.

[0148] 3DMM reconstruction of human face. Figure 4 , Figure 4 This is a schematic diagram of the principle of the data processing method provided in the embodiment of the present application Figure 1 The reconstruction process uses a 3DMM to reconstruct a single frame of talking image data. The facial image data is passed through a shape encoder and an expression encoder to obtain the portrait's shape and expression parameter hyperparameters, which drive the 3DMM. The driven 3DMM obtains the 2D and 3D landmark coordinates of 68 facial key points. Based on the 2D and 3D landmark coordinates obtained by driving, 68 key point losses are calculated for the GT landmarks corresponding to the image, including 12 eye key point losses and 20 mouth key point losses. loss = L1(pred 2d landmarks, GT 2d landmarks) + L1(pred 2d eyes landmarks, GT2d eyes landmarks) + L1(pred 2d mouth landmarks, GT 2d mouth land marks) + L1(pred3d landmarks, GT 3d landmarks) + L1(pred 3d eyes landmarks, GT 3d eyes landmarks) + L1(pred 3d mouth landmarks, GT 3d mouth landmarks).

[0149] In some embodiments, see Figure 5 , Figure 5 This is a schematic diagram of the principle of the data processing method provided in the embodiment of the present application Figure 2By manually annotating 8 key points of 10 models, the key points are 3 key points of each eye (left eye key points 1, 2, 3 and right eye key points 4, 5, 6) and 2 key points of the mouth (mouth key points 8 and 9), as well as key points 7 of the bridge of the nose, 10, 11, 12 of the left ear, and 13, 14, 15 of the right ear. Using 500 ID data, 25 frames of image data were extracted for each ID, totaling 12,500 face images. The corresponding 3DMM mesh was reconstructed using the face 3DMM, obtaining 12,500 different face meshes.

[0150] In some embodiments, see Figure 6 , Figure 6 This is a schematic diagram of the principle of the data processing method provided in the embodiment of the present application Figure 3 , using deformation transfer technology (3D mesh deformation) to migrate 12,500 human face mesh networks to 10 animal models. Using the deformation gradient of the source model (human face model), it is transferred to the target model (animal model) and the deformation gradient of the source model is calculated:

[0151]

[0152] Among them, the deformation gradient is defined as the linear transformation of each triangle of the mesh, which represents the change of the triangle from the reference shape to the deformed shape. For the source model, the deformation gradient Ti of each triangle of the reference shape and each deformed shape is calculated. Ai is the vertex matrix of the reference shape of triangle i. Bi is the vertex matrix of the deformed triangle i. The deformation gradient is transferred to the target model: for each triangle of the target model, the deformation gradient Ti of the source model is used as the deformation gradient Ti' of the target model. The vertex position is optimized and the vertex position of the target mesh is solved by a linear system so that the deformation gradient is consistent with the deformation gradient of the source model. The optimization equation form is:

[0153] min∑ i ||T′ i -T i || 2 (2)

[0154] Find the vertex positions of the target model so that the deformation gradient Ti' of each triangle of the target is as close as possible to the deformation gradient Ti of the source model. Ti' is the deformation gradient calculated by the vertices of the target mesh. Add the constraints of 8 feature points to ensure that the correspondence of the feature points remains accurate. Solve the linear system, take the vertices of the target model as unknowns, construct a sparse linear system, and solve it using the least squares method through iterative optimization. The positions of the 8 feature points are directly added to the optimization system as hard constraints. For post-processing of the results, the transferred target mesh needs to be locally smoothed to eliminate unreasonable noise or artifacts. The meshes of 125,000 animals are obtained, and through UV rendering, 250,000 animal images and 8 key points corresponding to each image can be obtained.

[0155] As an example, take the opening and closing of the mouth as an example, character A (human) and character B (animal), establish key points and triangulation, and collect data of character A: When character A closes his mouth: key point p i , i=1,…,8. When character A opens his mouth: key point p′ i , i=1,…,8, these key points can be the corners of the mouth, the middle of the upper lip, the middle of the lower lip, etc. Collect data for character B: When character B closes his mouth: key point qi, i=1,…,8i=1,…,8. We hope to eventually get the key point q′ when B opens his mouth i , and the corresponding overall mesh deformation. Triangulation: Triangulate the mouth area of ​​both models A and B using 8 key points, ensuring that the triangle topology of A and B is consistent. Each triangle (or quadrilateral) will undergo local deformation between the closed and open mouth configurations.

[0156] In some embodiments, geometrically, we can regard "from closed mouth to open mouth" as a mapping F: x>x', which can be represented in 3D space by a gradient matrix Or called the local affine transformation matrix) to characterize "how the local part is stretched, rotated, and translated." How to calculate the deformation gradient on the mesh. Taking a triangle as an example, suppose the vertices of a triangle in the closed mouth form are v1, v2, and v3, and in the open mouth form they correspond to v1′, v2′, and v3′. We often use the following method to estimate the affine transformation matrix A of the triangle (ignoring the translation component, which can be handled separately): construct the matrix V = [v2-v1, v3-v1], V′ = [v2′-v1′v3′-v1′]. Assuming A≈v′·v-1, A can be understood as the local deformation gradient of the triangle, reflecting the stretching / rotation from "closed mouth" to "open mouth." The overall deformation gradient of character A can be similarly calculated as a local matrix Ai for all triangles in A's mouth. In this way, we obtain a set of patch-level deformation gradients of A. Transfer the deformation gradient to character B, corresponding to the patch & vertex: The mouth meshes of A and B (or at least the key areas) must maintain the same or mappable topology, that is, "the i-th triangle of A" corresponds to "the i-th triangle of B". The points on the corresponding triangles can also be mapped one-to-one, so that the deformation gradient matrix Ai of A can be migrated to the corresponding patch of B. Preliminary update of B's ​​vertices: Assume that the vertices of a triangle in B's closed mouth form are u1, u2, u3, and we want to get the vertices u1′, u2′, u3 after it opens its mouth. If Ai is directly applied to the patch, it can be written as:

[0157] [u′2-u′1,u′3-u′1]≈A i [u2-u1,u3-u1] (3)

[0158] At this point, we also need to consider the translation vector ti, because the light Ai only includes rotation, scaling, shearing, etc. The general deformation can be expressed as:

[0159] x′=A i (x-u1)+u′1+t i (4)

[0160] In some embodiments, key point constraints are introduced: since we have 8 key points {qi} in the closed mouth form of B, and we hope that they satisfy {q′ i} (or corresponding to the deformation of A), the coordinate constraints of these key points can be incorporated into the overall solution of the B mesh. In other words, for the patches where the key points are located (or around them), it is necessary to force their positions after deformation to be consistent with {q′ i}Fit as closely as possible.

[0161] In some embodiments, the vertex position is optimized: energy function and solution. In practice, it is not necessary to simply perform "deformation matrix replacement" on each triangle independently, but to optimize globally to ensure the smoothness of the mesh, the accuracy of key points, and the reasonable transition between patches. Set an energy function to minimize: deformation gradient consistency term: it is hoped that the gradient of each triangle of B Gradient with A As similar as possible, that is:

[0162]

[0163] This allows B's mouth to "learn" A's mouth opening method. Key point constraint: Ensure that the positions of B's ​​8 key points after deformation {q′ i} is consistent with the expected position (according to A's mouth opening deformation or directly specified target), recorded as

[0164]

[0165] Smoothness: To prevent some faces from being distorted too much, it is necessary to ensure smooth changes between adjacent vertices. You can add constraints such as "Laplacian" or "edge length difference".

[0166]

[0167] Comprehensive energy function:

[0168] E=αE gradient +βE landmark +γE smooth (8)

[0169] Among them, α, β, and γ are weights used to balance "deformation gradient consistency", "key point accurate matching", and "global smoothness".

[0170] By minimizing this energy function (using Newton's method, gradient descent, or a specialized mesh deformation solver), we can obtain a set of solutions {u′ i}, so that: the overall deformation is as consistent as possible with the deformation gradient of A from closed mouth to open mouth, while satisfying the key point constraints of B and maintaining a smooth transformation on the mesh.

[0171] In some embodiments, deformation gradient extraction: first calculate the "local triangle deformation matrix" from closed mouth to open mouth on character A. Transfer deformation: apply these deformation matrices to the corresponding triangles of character B to obtain a preliminary "open mouth shape". Key point correction: use B's 8 mouth key points to tighten or relax them to the required mouth opening position, and use this as a hard or soft constraint. Global optimization: construct an energy function, combine "deformation gradient consistency", "key point constraints" and "smoothness", and iteratively solve the optimal vertex position distribution. In this way, character B can imitate the dynamic changes of character A from closed mouth to open mouth to the greatest extent while retaining the structural characteristics of its own mouth, and achieve realistic "mouth movement migration".

[0172] In some embodiments, see Figure 7 , Figure 7 This is a schematic diagram of the principle of the data processing method provided in the embodiment of the present application Figure 4 , the effect of migration is as follows Figure 7 As shown in the figure, different blinking states of the right eye are transferred. Similarly, the left eye and open mouth are also transformed and transferred.

[0173] In some embodiments, see Figure 8 , Figure 8 This is a schematic diagram of the principle of the data processing method provided in the embodiment of the present application Figure 5 , obtain the data, train a model, see Figure 8 , animal facial key point recognition model (key point detection model and key point mapping 3D model), the model input is an animal image, the output is the 8 2D 3D key points corresponding to the image, training loss = L1 (pred 2d landmarks, GT 2d landmarks) + L1 (pred 3d landmarks, GT 3dlandmarks).

[0174] In some embodiments, see Figure 9 , Figure 9 This is a schematic diagram of the principle of the data processing method provided in the embodiment of the present application Figure 6 , 2D key points are matched to 3D animal models (for animation video production). 50 animal models are each rendered into 50 animal face images through a UV. The trained image animal key point detection model is used to detect 2D 3D landmarks in the 50 animal images. 8 key points of the mesh corresponding to the model are converted from 3D landmarks. The process of migrating human face images to animal images is shown in [1]. Figure 9The human face is then mapped to a corresponding 3DMM model and its eight key points. The 3DMM model is then mapped to the eight key points of 50 animal models. Using the 3-point deformation transfer technique, the 3DMM of the human face is transferred to the meshes of the 50 models. By rendering each model with two UVs, 100 animal images are generated that match the human face's mouth shape and blinking. Using these steps, a 30-minute video of a person ID is transformed into 100 animal talking videos. The audio from the person talking video is then merged with the animal talking video to create an audio and visual video of the animals talking. This results in 500*50*2 = 50,000 animal talking videos.

[0175] In some embodiments, see Figure 10 , automatically produces animal talking videos, reducing the cost of digital animal training data. The generated animal talking videos can be used to retrain various digital human algorithms and turn them into digital animals, such as Figure 10 Video frames of the animal cat are shown.

[0176] It is understandable that in the embodiments of the present application, when data related to the first image is involved and is applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0177] The following continues to describe the exemplary structure of the data processing device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the data processing device 455 of the memory 450 may include: a construction module for determining a first three-dimensional model of the first object, the first three-dimensional model of the first object and the third three-dimensional model of the second object having the same expression parameters; constructing a second three-dimensional model of the first object based on the first image of the first object; a determination module for determining deformation parameters based on the second three-dimensional model and the first three-dimensional model; an adjustment module for adjusting the third three-dimensional model based on the deformation parameters to obtain a fourth three-dimensional model of the second object, the fourth three-dimensional model and the second three-dimensional model having the same expression parameters; a generation module for generating a second image of the second object based on the fourth three-dimensional model, the expression parameters of the second object in the second image being the same as the expression parameters of the first object in the first image.

[0178] In some embodiments, the above-mentioned construction module is also used to perform a first feature extraction on the first image to obtain the facial shape feature of the first object; perform a second feature extraction on the first image to obtain the expression parameter feature of the first object; and construct a second three-dimensional model of the first object based on the facial shape feature and the expression parameter feature.

[0179] In some embodiments, the above-mentioned construction module is also used to predict driving parameters based on the facial shape features and the expression parameter features through a parameter prediction model to obtain driving parameters; and adjust the first three-dimensional model based on the driving parameters to obtain a second three-dimensional model of the first object.

[0180] In some embodiments, the above-mentioned construction module is also used to predict driving parameters based on image samples through an initial parameter prediction model to obtain the driving parameters of the third object in the image samples; based on the driving parameters of the third object and the driving parameter label of the third object, the initial parameter prediction model is trained to obtain the parameter prediction model.

[0181] In some embodiments, the second three-dimensional model includes multiple first surface elements, the first three-dimensional model includes second surface elements corresponding one-to-one to each of the first surface elements, and the deformation parameters include multiple sub-deformation parameters; the above-mentioned determination module is also used to perform the following processing on each of the first surface elements: feature extraction is performed on the first surface element to obtain the first shape feature of the first surface element, and feature extraction is performed on the second surface element corresponding to the first surface element to obtain the second shape feature of the second surface element; the first shape feature includes multiple first sub-shape features, and the first sub-shape feature is used to indicate the coordinate difference between different vertices on the first surface element. The second shape feature includes multiple second sub-shape features corresponding one-to-one to the first sub-shape features; the second shape feature is multiplied by the inverse feature of the first shape feature to obtain the sub-deformation parameter corresponding to the first surface element.

[0182] In some embodiments, the third three-dimensional model includes multiple third surface elements, and the deformation parameters include sub-deformation parameters corresponding to each of the third surface elements; the adjustment module is also used to adjust the shape of each of the third surface elements in the third three-dimensional model based on the sub-deformation parameters corresponding to the third surface elements, to obtain a fourth surface element corresponding to the third surface element; each of the third surface elements in the third three-dimensional model is replaced by the corresponding fourth surface element to obtain the fourth three-dimensional model.

[0183] In some embodiments, the generation module is also used to predict the expected position of the key points on the fourth three-dimensional model based on the third image of the second object; control the key points in the fourth three-dimensional model and adjust the positions according to the expected positions to obtain a fifth three-dimensional model; smooth the fifth facet in the fifth three-dimensional model to obtain a sixth three-dimensional model; and generate the second image of the second object based on the sixth three-dimensional model.

[0184] In some embodiments, the fifth face element in the fifth three-dimensional model corresponds one-to-one to the third face element in the third three-dimensional model, and the generation module is further used to smooth each fifth face element in the fifth three-dimensional model based on the third face element corresponding to the fifth face element in the third three-dimensional model to obtain a sixth face element corresponding to the fifth face element; and construct each sixth face element into the sixth three-dimensional model according to the arrangement of each fifth face element in the fifth three-dimensional model.

[0185] In some embodiments, the fifth element in the fifth three-dimensional model corresponds one-to-one to the third element in the third three-dimensional model, and the generation module is also used to fuse the audio of the first object corresponding to each first image with the corresponding second image to obtain the video of the second object.

[0186] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the data processing method described in the embodiment of the present application.

[0187] The embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor will execute the data processing method provided by the embodiment of the present application, for example, Figure 3 The data processing method is shown.

[0188] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various electronic devices including one or any combination of the above memories.

[0189] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0190] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0191] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0192] In summary, the embodiments of the present application have the following beneficial effects:

[0193] (1) By determining the first three-dimensional name of the first object, a second three-dimensional model of the first object is constructed based on the first image of the first object. Based on the first three-dimensional model and the second three-dimensional model, deformation parameters are determined. Based on the deformation parameters, the third three-dimensional model is adjusted to obtain a fourth three-dimensional model of the second object. Based on the fourth three-dimensional model, a second image is generated. The deformation parameters are determined with reference to the first three-dimensional model of the first object. The first three-dimensional model and the third three-dimensional model of the second object have the same expression parameters. By comparing the second three-dimensional model and the first three-dimensional model, the deformation law between the expression parameters can be found. Based on the determined deformation parameters, the third three-dimensional model is adjusted. The third three-dimensional model of the second object is converted into a fourth three-dimensional model. This adjustment is based on the deformation relationship between the second three-dimensional model and the first three-dimensional model, ensuring the accuracy and naturalness of the expression parameter migration. Based on the fourth three-dimensional model, a second image of the second object is generated. The adjusted fourth three-dimensional model is converted into a visual image, completing the migration of the expression parameters from the first object to the second object. The entire process is based on the deformation relationship between models, ensuring the natural transition and accurate reproduction of expression parameters. It takes into account the differences in expression parameters between different objects, making the migrated expression parameters more consistent with the characteristics and expression parameter features of the second object, thereby effectively improving the accuracy of expression parameter migration.

[0194] (2) By accurately extracting facial shape and expression parameter features, the constructed three-dimensional model can highly restore the facial features and expression parameters of the first subject, making the virtual character appear more realistic and natural. In virtual reality and game development, the virtual character can respond to the user's expression parameters and actions in real time, improving the user's immersion and interactive experience. By extracting the facial features and expression parameters of different individuals, a personalized three-dimensional model can be customized for each person to meet the needs of different users.

[0195] (3) The parameter prediction model is used to predict the driving parameters of the key points on the first three-dimensional model based on the facial shape features and expression parameter features, and the key points are controlled to adjust their positions according to the driving parameters, thereby constructing a second three-dimensional model of the first object. This method has significant beneficial effects. This method first uses machine learning technology to learn the mapping relationship between the facial shape and expression parameter features and the positions of the key points of the three-dimensional model, so that the predicted driving parameters can accurately reflect the facial features and expression parameters of the first object. Secondly, by adjusting the positions of the key points in a quantitative manner, the efficiency and consistency of model construction are greatly improved, and the need for manual adjustment and possible errors are reduced. It can also process large amounts of data and quickly generate multiple personalized three-dimensional models to meet the needs of different application scenarios.

[0196] (4) An initial parameter prediction model was constructed, which predicts the key point driving parameters of the third object based on the facial shape features and expression parameter features of the third object in the image sample. Then, we used these predicted driving parameters to compare with the actual driving parameter labels, and trained and optimized the initial model through supervised learning methods to obtain a parameter prediction model that can accurately predict the key point driving parameters of the first object. The technical derivation of this process is rigorous, and through a large amount of data training and model optimization, the accuracy and generalization ability of the parameter prediction model are ensured. Ultimately, this parameter prediction model can predict the driving parameters of each key point on the first three-dimensional model based on the facial shape features and expression parameter features of the first object in actual applications, thereby constructing a highly realistic and personalized three-dimensional model, improving the efficiency and accuracy of model construction.

[0197] (5) By extracting the shape features of each first element and the shape features of the corresponding second element, and determining the sub-deformation parameters based on these features, efficient and accurate three-dimensional model deformation can be achieved. This method first captures subtle shape changes by quantitatively analyzing the geometric characteristics of the model unit, providing accurate guidance for deformation. Then, by comparing the shape features of the first element and the second element, the required deformation parameters are calculated. These parameters can accurately reflect the transformation process from the second three-dimensional model to the first three-dimensional model. This feature-based deformation method not only improves the accuracy and naturalness of the deformation, but also greatly reduces the workload of manual adjustment and improves work efficiency.

[0198] (6) By extracting and analyzing the shape features of each model unit in the third three-dimensional model, the corresponding sub-deformation parameters are calculated. These parameters accurately describe the way and degree of change of each model unit during the deformation process. Then, for each third face element, based on its corresponding sub-deformation parameter, the shape of the unit is adjusted to obtain the corresponding fourth face element. Finally, each third face element in the third three-dimensional model is replaced by the corresponding fourth face element, thereby obtaining a fourth three-dimensional model with specific expression parameters or morphology. The advantage of this method is that through precise control of the model unit level, fine adjustment of the entire model is achieved while maintaining the overall consistency and integrity of the model. In addition, by pre-calculating the deformation parameters, the model can be deformed automatically, reducing the workload of manual adjustment and improving work efficiency.

[0199] (7) Key point correction ensures that the key parts of the model (such as eyes, mouth, etc.) match the expected expression parameters. By adjusting the position and shape of these key points, the expression parameters of the model can be made more natural and realistic. Model smoothing eliminates the roughness and sharp edges of the model surface. Through smoothing, the surface of the model is smoother and looks more natural. Finally, based on the sixth three-dimensional model that has been corrected and smoothed by key points, a second image of the second object with the first expression parameters is generated, which can ensure the quality and realism of the image. Through the rigorous key point correction and smoothing process, the quality and realism of the generated image are ensured, providing users with a more immersive experience.

[0200] (8) By correcting the key points of the fourth three-dimensional model to obtain the fifth three-dimensional model, the authenticity and expressiveness of the expression parameters of the three-dimensional model can be significantly improved. First, based on the third image of the second object with the second expression parameter, we use computer vision technology to predict the expected positions of the key points in the fourth three-dimensional model. These predicted positions reflect the specific characteristics of the second expression parameter, such as the opening and closing of the eyes, the shape of the mouth, etc. Then, we control the key points in the fourth three-dimensional model and adjust their positions according to these expected positions. In this way, we can accurately adjust the key parts of the model, which not only improves the realism of the expression parameters of the model, but also enables the model to convey emotions and intentions more accurately. In addition, this method is efficient and repeatable and can be quickly applied to the adjustment of multiple models and expression parameters. By correcting the key points of the fourth three-dimensional model to obtain the fifth three-dimensional model, not only can the authenticity and expressiveness of the expression parameters of the model be improved, but it is also efficient and repeatable, providing strong support for the production and animation of three-dimensional models.

[0201] (9) The one-to-one correspondence of the model units ensures that each fifth facet has a clear reference standard, namely the third facet, which provides precise guidance for the smoothing process. Then, by smoothing the fifth facet, the unevenness and sharp edges of the model surface can be reduced, making the model look smoother and more natural. This smoothing process not only improves the visual effect of the model, but also enhances the realism of the model, making it more expressive in applications such as animation and games. Finally, according to the arrangement of each model unit in the fifth three-dimensional model, the smoothed sixth facet is reassembled to construct the sixth three-dimensional model. This process maintains the overall structure and topological relationship of the model while ensuring that the smoothing effect of each unit is reflected. Therefore, by smoothing the fifth facet, the sixth three-dimensional model obtained has significantly improved visual quality and realism, providing strong technical support for the development of the digital media and entertainment industry.

[0202] (10) The second image generated based on the fourth three-dimensional model already has the same visual expression as the expression parameters of the first object, while the audio of the first object contains sound information related to the expression parameters, such as dialogue, laughter or crying. By fusing the two, the synchronization of sound and image can be achieved, so that the audience can feel a more natural and real experience when watching the video. This fusion not only enhances the appeal of the video, but also improves the audience's participation and immersion. In addition, in this way, different visual and sound elements can be combined to create richer and more diverse multimedia content to meet the needs of different audiences. Therefore, fusing audio and image to obtain video is a very effective way to create multimedia content, which can significantly improve the quality and viewing experience of the content.

[0203] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. A data processing method, characterized in that: The method comprises: Determining a first three-dimensional model of a first object, wherein the first three-dimensional model of the first object and the third three-dimensional model of the second object have the same expression parameters; constructing a second three-dimensional model of the first object based on the first image of the first object; determining deformation parameters based on the second three-dimensional model and the first three-dimensional model; adjusting the third three-dimensional model based on the deformation parameters to obtain a fourth three-dimensional model of the second object, wherein the fourth three-dimensional model has the same expression parameters as the second three-dimensional model; A second image of the second object is generated based on the fourth three-dimensional model, wherein expression parameters of the second object in the second image are the same as expression parameters of the first object in the first image.

2. The method according to claim 1, characterized in that The constructing a second three-dimensional model of the first object based on the first image of the first object includes: performing a first feature extraction on the first image to obtain a facial shape feature of the first object; performing a second feature extraction on the first image to obtain an expression parameter feature of the first object; A second three-dimensional model of the first object is constructed based on the facial shape features and the expression parameter features.

3. The method according to claim 2, characterized in that The constructing a second three-dimensional model of the first object based on the facial shape feature and the expression parameter feature includes: Based on the facial shape features and the expression parameter features, driving parameter prediction is performed using a parameter prediction model to obtain driving parameters; The first three-dimensional model is adjusted based on the driving parameters to obtain a second three-dimensional model of the first object.

4. The method according to claim 3, characterized in that Before obtaining the driving parameters by predicting the driving parameters using a parameter prediction model based on the facial shape features and the expression parameter features, the method further comprises: Based on the image sample, predicting driving parameters by using an initial parameter prediction model to obtain driving parameters of a third object in the image sample; The initial parameter prediction model is trained based on the driving parameters of the third object and the driving parameter label of the third object to obtain the parameter prediction model.

5. The method according to claim 1, wherein The second three-dimensional model includes a plurality of first surface elements, the first three-dimensional model includes second surface elements corresponding to each of the first surface elements in a one-to-one manner, and the deformation parameter includes a plurality of sub-deformation parameters; The determining of deformation parameters based on the second three-dimensional model and the first three-dimensional model includes: The following processing is performed on each of the first bins: Performing feature extraction on the first surface element to obtain a first shape feature of the first surface element, and performing feature extraction on a second surface element corresponding to the first surface element to obtain a second shape feature of the second surface element; The first shape feature includes a plurality of first sub-shape features, each of which is used to indicate a coordinate difference between different vertices on the first facet; and the second shape feature includes a plurality of second sub-shape features corresponding to each of the first sub-shape features. The second shape feature is multiplied by the inverse feature of the first shape feature to obtain a sub-deformation parameter corresponding to the first surface element.

6. The method according to claim 1, wherein The third three-dimensional model includes a plurality of third surface elements, and the deformation parameters include sub-deformation parameters corresponding to each of the third surface elements. The adjusting the third three-dimensional model based on the deformation parameter to obtain a fourth three-dimensional model of the second object includes: For each of the third surface elements in the third three-dimensional model, adjusting the shape of the third surface element based on the sub-deformation parameter corresponding to the third surface element to obtain a fourth surface element corresponding to the third surface element; Each of the third surface elements in the third three-dimensional model is replaced by the corresponding fourth surface element to obtain the fourth three-dimensional model.

7. The method according to claim 1, characterized in that Generating a second image of the second object based on the fourth three-dimensional model includes: predicting expected positions of key points on the fourth three-dimensional model based on the third image of the second object; controlling the key points in the fourth three-dimensional model to adjust their positions according to the desired positions to obtain a fifth three-dimensional model; performing smoothing on the fifth surface element in the fifth three-dimensional model to obtain a sixth three-dimensional model; Based on the sixth three-dimensional model, a second image of the second object is generated.

8. The method according to claim 7, characterized in that The fifth surface element in the fifth three-dimensional model corresponds one-to-one to the third surface element in the third three-dimensional model, and the fifth surface element in the fifth three-dimensional model is smoothed to obtain a sixth three-dimensional model, including: For each fifth bin in the fifth three-dimensional model, smoothing the fifth bin based on a third bin corresponding to the fifth bin in the third three-dimensional model to obtain a sixth bin corresponding to the fifth bin; According to the arrangement of the fifth surface elements in the fifth three-dimensional model, the sixth surface elements are constructed into the sixth three-dimensional model.

9. The method according to claim 1, characterized in that There are multiple first images, and the second images correspond to the first images one by one. After generating the second image of the second object based on the fourth three-dimensional model, the method further includes: The audio of the first object corresponding to each first image is fused with the corresponding second image to obtain a video of the second object.

10. A data processing device, characterized in that: The device comprises: A construction module is configured to determine a first three-dimensional model of the first object, wherein the first three-dimensional model of the first object and the third three-dimensional model of the second object have the same expression parameters; and construct a second three-dimensional model of the first object based on the first image of the first object; a determining module, configured to determine a deformation parameter based on the second three-dimensional model and the first three-dimensional model, wherein the deformation parameter is used to indicate a degree of deformation of the second three-dimensional model compared to the first three-dimensional model; an adjustment module, configured to adjust the third three-dimensional model based on the deformation parameters to obtain a fourth three-dimensional model of the second object, wherein the fourth three-dimensional model has the same expression parameters as the second three-dimensional model; A generating module is configured to generate a second image of the second object based on the fourth three-dimensional model, wherein expression parameters of the second object in the second image are the same as expression parameters of the first object in the first image.

11. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; The processor is configured to implement the data processing method according to any one of claims 1 to 9 when executing the computer-executable instructions or computer programs stored in the memory.

12. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the data processing method according to any one of claims 1 to 9 is implemented.

13. A computer program product comprising a computer program or computer executable instructions, characterized in that When the computer program or computer executable instructions are executed by a processor, the data processing method according to any one of claims 1 to 9 is implemented.