Method and device for acquiring multi-expression face homotopy grid data and electronic equipment
By obtaining semantic information from the original face database and deforming the vertices of the expression mesh using dense optical flow and multi-view localization methods, the problem of semantic inconsistency between expressions in multi-expression face mesh data is solved, and the acquisition of strongly semantically consistent multi-expression face mesh data with the same topology is realized.
Patent Information
- Application Number
- CN202211129677.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-16
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2042-09-16
AI Technical Summary
Existing technologies that obtain multi-expression face mesh data through individual registration suffer from semantic inconsistencies between expressions.
By obtaining dense semantic information of faces from the original face database, dense optical flow between different expression view maps under various perspectives is calculated. Based on dense optical flow and multi-view localization method, the vertices of each expression mesh are deformed vertex by vertex. Dense optical flow between UV maps is further calculated, the vertex coordinates of each expression mesh are corrected, and semantic consistency is enhanced.
This method achieves strong semantic consistency in obtaining multi-expression face topology mesh data, enhancing the semantic consistency between expressions and solving the problem of semantic inconsistency between expressions under the separate registration method.
Smart Images

Figure CN115631318B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular to a method, apparatus and electronic device for acquiring multi-expression facial topological mesh data. Background Technology
[0002] High-quality 3D facial topology mesh data plays an important role in both industrial applications such as game and animation production, and scientific research in providing facial datasets.
[0003] In related technologies, most 3D face topology mesh data are obtained by fitting a specified template face mesh onto the original acquired face scan, i.e., by registering separately.
[0004] However, this registration method is independent of facial expressions. During the registration of these grid data, the original 3D face scans and corresponding multi-view images contain dense semantic information of different facial expressions. Therefore, although the separate registration method can achieve the same topology, the same area of the grid may correspond to different semantics of the face, that is, it is impossible to achieve semantic consistency between facial expressions.
[0005] Therefore, it is essential to establish a registration method for multi-expression facial mesh data. Summary of the Invention
[0006] This application provides a method, apparatus, and electronic device for acquiring multi-expression face mesh data with the same topology, in order to solve the problem of semantic inconsistency between expressions caused by obtaining multi-expression face mesh data of the same individual through separate registration.
[0007] The first aspect of this application provides a method for acquiring multi-expression facial topological mesh data, including the following steps:
[0008] Extract dense semantic information about faces from the original face database;
[0009] Based on the semantic information, a first dense optical flow is calculated between different facial expression viewpoints under each viewpoint. Then, based on the first dense optical flow and the multi-viewpoint localization method, the vertices of each facial expression mesh are deformed vertex-by-vertex to obtain a preliminary semantically consistent enhanced facial expression mesh; and
[0010] Based on the preliminarily semantically consistent enhanced facial expression mesh, the second dense optical flow between the UV maps of each facial expression mesh is calculated, and the coordinates of the vertices of each facial expression mesh are corrected based on the second dense optical flow to obtain strongly semantically consistent multi-expression face topological mesh data.
[0011] According to one embodiment of this application, the step of calculating the first dense optical flow between different facial expression viewpoints based on the semantic information, and deforming the vertices of each facial expression mesh vertex by vertex based on the first dense optical flow and the multi-view localization method to obtain a preliminary semantically consistent enhanced facial expression mesh includes:
[0012] Projecting the first expression mesh onto the expression map from multiple perspectives yields the two-dimensional positions of each vertex of the first expression mesh in the first expression map.
[0013] The first dense optical flow from the first expression map to the second expression map is calculated using the first optical flow solver, and the two-dimensional positions of each vertex of the first expression mesh projected onto the second expression map under semantic consistency constraints are obtained based on the first dense optical flow.
[0014] By utilizing the topological consistency between the first and second expression grids, the two-dimensional positions of each vertex of the second expression grid projected onto the second expression graph under semantic consistency constraints are obtained.
[0015] Using the multi-view localization method, each vertex of the second expression mesh is relocated to obtain the deformed second expression mesh, and the first expression mesh and the deformed second expression mesh are used as the expression mesh for preliminary semantic consistency enhancement.
[0016] According to one embodiment of this application, the step of calculating a second dense optical flow between the UV (UV texture map coordinates) maps of each facial expression mesh based on the preliminary semantic consistency enhancement, and correcting the coordinates of the vertices of each facial expression mesh based on the second dense optical flow, includes:
[0017] The first expression mesh and the deformed second expression mesh are projected onto the corresponding original 3D face scan image to obtain the 3D coordinates and material information of each vertex projected onto the original 3D face scan image;
[0018] Utilizing the topological consistency, the first facial expression mesh and the deformed second facial expression mesh are projected into the same UV space, and a UV map and a three-dimensional coordinate distribution field are generated based on the three-dimensional coordinates and the material information;
[0019] The second dense optical flow is calculated from the UV map of the first expression mesh to the UV map of the deformed second expression mesh using a second optical flow solver.
[0020] According to one embodiment of this application, correcting the coordinates of the vertices of each facial expression mesh based on the second dense optical flow includes:
[0021] The UV coordinates of each vertex of the first facial expression mesh on the UV map of the deformed second facial expression mesh under semantic consistency constraints are obtained based on the second dense optical flow.
[0022] The three-dimensional coordinates of the original three-dimensional face scan image of the deformed second expression mesh are obtained by sampling from the three-dimensional coordinate distribution field of the first expression mesh under semantic consistency constraints.
[0023] The deformed second expression mesh is corrected based on the coordinates of each vertex of the first expression mesh in the original 3D face scan image of the deformed second expression mesh.
[0024] According to one embodiment of this application, both the first optical flow solver and the second optical flow solver are trained using a pre-trained neural network for optical flow calculation.
[0025] According to the method for acquiring multi-expression face topological mesh data according to embodiments of this application, dense semantic information of faces is obtained from the original face database, and a first dense optical flow is calculated between different expression view maps under various perspectives. Based on the first dense optical flow and a multi-view localization method, the vertices of each expression mesh are deformed vertex by vertex to obtain an expression mesh with preliminary enhanced semantic consistency. Then, a second dense optical flow is calculated between the UV maps of each expression mesh, and the coordinates of the vertices of each expression mesh are corrected based on the second dense optical flow to obtain multi-expression face topological mesh data with strong semantic consistency. Thus, semantic information is obtained from the original face acquisition data, and semantic relationships between expressions are constructed using optical flow. The individually registered multi-expression mesh data is deformed to enhance its semantic consistency, solving the problem of semantic inconsistency between expressions in multi-expression face mesh data of the same individual obtained through individual registration.
[0026] A second aspect of this application provides an apparatus for acquiring multi-expression facial topological mesh data, comprising:
[0027] The acquisition module is used to obtain dense semantic information about faces from the original face database;
[0028] The first calculation module is used to calculate the first dense optical flow between different facial expression viewpoints based on the semantic information, and to deform the vertices of each facial expression mesh vertex by vertex based on the first dense optical flow and the multi-view localization method to obtain a preliminary semantically consistent enhanced facial expression mesh; and
[0029] The second calculation module is used to calculate the second dense optical flow between the UV maps of each expression grid based on the preliminary semantically consistent enhanced expression grid, and to correct the coordinates of the vertices of each expression grid based on the second dense optical flow, so as to obtain strong semantically consistent multi-expression face topology grid data.
[0030] According to one embodiment of this application, the first computing module is specifically used for:
[0031] Projecting the first expression mesh onto the expression map from multiple perspectives yields the two-dimensional positions of each vertex of the first expression mesh in the first expression map.
[0032] The first dense optical flow from the first expression map to the second expression map is calculated using the first optical flow solver, and the two-dimensional positions of each vertex of the first expression mesh projected onto the second expression map under semantic consistency constraints are obtained based on the first dense optical flow.
[0033] By utilizing the topological consistency between the first and second expression grids, the two-dimensional positions of each vertex of the second expression grid projected onto the second expression graph under semantic consistency constraints are obtained.
[0034] Using the multi-view localization method, each vertex of the second expression mesh is relocated to obtain the deformed second expression mesh, and the first expression mesh and the deformed second expression mesh are used as the expression mesh for preliminary semantic consistency enhancement.
[0035] According to one embodiment of this application, the second computing module is specifically used for:
[0036] The first expression mesh and the deformed second expression mesh are projected onto the corresponding original 3D face scan image to obtain the 3D coordinates and material information of each vertex projected onto the original 3D face scan image;
[0037] Utilizing the topological consistency, the first facial expression mesh and the deformed second facial expression mesh are projected into the same UV space, and a UV map and a three-dimensional coordinate distribution field are generated based on the three-dimensional coordinates and the material information;
[0038] The second dense optical flow is calculated from the UV map of the first expression mesh to the UV map of the deformed second expression mesh using a second optical flow solver.
[0039] According to one embodiment of this application, the second computing module is specifically used for:
[0040] The UV coordinates of each vertex of the first facial expression mesh on the UV map of the deformed second facial expression mesh under semantic consistency constraints are obtained based on the second dense optical flow.
[0041] The three-dimensional coordinates of the original three-dimensional face scan image of the deformed second expression mesh are obtained by sampling from the three-dimensional coordinate distribution field of the first expression mesh under semantic consistency constraints.
[0042] The deformed second expression mesh is corrected based on the coordinates of each vertex of the first expression mesh in the original 3D face scan image of the deformed second expression mesh.
[0043] According to one embodiment of this application, both the first optical flow solver and the second optical flow solver are trained using a pre-trained neural network for optical flow calculation.
[0044] The apparatus for acquiring multi-expression face topological mesh data according to embodiments of this application obtains dense semantic information of faces from an original face database, calculates the first dense optical flow between different expression view maps under various perspectives, and deforms the vertices of each expression mesh vertex by vertex based on the first dense optical flow and a multi-view localization method to obtain an expression mesh with preliminary enhanced semantic consistency. Then, it calculates the second dense optical flow between the UV maps of each expression mesh and corrects the coordinates of the vertices of each expression mesh based on the second dense optical flow to obtain multi-expression face topological mesh data with strong semantic consistency. Thus, by obtaining semantic information from the original face acquisition data, constructing semantic relationships between expressions using optical flow, and deforming individually registered multi-expression mesh data to enhance its semantic consistency, the apparatus solves the problem of semantic inconsistency between expressions in multi-expression face mesh data of the same individual obtained through individual registration.
[0045] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for acquiring multi-expression facial topology mesh data as described in the above embodiments.
[0046] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the method for acquiring multi-expression face topology mesh data as described in the above embodiments.
[0047] The method for obtaining multi-expression face topology mesh data according to the embodiments of this application has the following advantages:
[0048] (1) It can obtain registered, semantically consistent facial mesh data.
[0049] (2) This registration method can be used as an enhancement of the traditional registration method. It can not only directly operate on the face grid data registered using the traditional method, but also enhance the semantic consistency between expressions.
[0050] (3) This registration method can directly utilize the original collected face data without adding any additional data, making it easy to operate.
[0051] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0052] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0053] Figure 1 This is a flowchart illustrating a method for acquiring multi-expression facial topological mesh data according to an embodiment of this application;
[0054] Figure 2 This is a block diagram of a device for acquiring multi-expression face topology mesh data according to an embodiment of this application;
[0055] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0056] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0057] The following describes a method, apparatus, and electronic device for acquiring multi-expression face co-topological mesh data according to embodiments of this application, with reference to the accompanying drawings. Addressing the problem mentioned in the background art of semantic inconsistencies between expressions in multi-expression face mesh data obtained from a single individual through individual registration, this application provides a method for acquiring multi-expression face co-topological mesh data. In this method, dense semantic information of faces is obtained from an original face database, and a first dense optical flow is calculated between different expression viewpoints at various perspectives. Based on the first dense optical flow and a multi-view localization method, the vertices of each expression mesh are deformed vertex-by-vertex to obtain an expression mesh with enhanced initial semantic consistency. Then, a second dense optical flow is calculated between the UV maps of each expression mesh, and the coordinates of the vertices of each expression mesh are corrected based on the second dense optical flow to obtain strongly semantically consistent multi-expression face co-topological mesh data. This solves the problem of semantic inconsistencies between expressions in multi-expression face mesh data obtained from a single individual through individual registration. By constructing the correspondence between semantic information of expressions through optical flow, the individually registered multi-expression mesh data is deformed, enhancing its semantic consistency.
[0058] Specifically, Figure 1 This is a flowchart illustrating a method for acquiring multi-expression facial topology mesh data provided in an embodiment of this application.
[0059] like Figure 1 As shown, the method for obtaining multi-expression face topology mesh data includes the following steps:
[0060] In step S101, dense semantic information of faces is obtained from the original face database.
[0061] It should be understood that high-quality 3D facial mesh data plays a crucial role in both industrial applications such as game and animation production, and in providing facial datasets for scientific research. Generally, multi-expression facial mesh data of the same individual is obtained through individual registration, but this method ignores the semantic consistency between expressions.
[0062] Therefore, in order to obtain highly semantically consistent multi-expression facial mesh data, the embodiments of this application can first use a three-dimensional scanning method to collect dense semantic information of faces from original face databases such as multi-view images.
[0063] In step S102, based on semantic information, the first dense optical flow between different facial expression viewpoints is calculated, and based on the first dense optical flow and the multi-view localization method, the vertices of each facial expression mesh are deformed vertex by vertex to obtain a preliminary semantically consistent enhanced facial expression mesh.
[0064] Further, in some embodiments, a first dense optical flow is calculated between different facial expression view maps under various perspectives based on semantic information, and the vertices of each facial expression mesh are deformed vertex by vertex based on the first dense optical flow and a multi-view localization method to obtain a preliminary semantically consistent enhanced facial expression mesh. This includes: projecting the first facial expression mesh onto facial expression maps under multiple perspectives to obtain the two-dimensional position of each vertex of the first facial expression mesh in the first facial expression map; calculating the first dense optical flow from the first facial expression map to the second facial expression map using a first optical flow solver, and obtaining the two-dimensional position of each vertex of the first facial expression mesh projected onto the second facial expression map under semantic consistency constraints based on the first dense optical flow; using the topological consistency of the first and second facial expression meshes to obtain the two-dimensional position of each vertex of the second facial expression mesh projected onto the second facial expression map under semantic consistency constraints; using the multi-view localization method to relocate each vertex of the second facial expression mesh to obtain a deformed second facial expression mesh, and using the first facial expression mesh and the deformed second facial expression mesh as the preliminary semantically consistent enhanced facial expression mesh.
[0065] In this system, the first expression mesh can be a unique original expression mesh, and the second expression mesh can be a corresponding multi-view expression mesh. For example, the first expression mesh is expression mesh A, and the second expression mesh can represent only expression mesh B, or it can represent both expression meshes B and C, or even expression meshes B, C, and D, or even more. Specifically, by collecting dense semantic information about faces from original face databases such as multi-view images, a streaming method can be used to establish semantic correspondences between expressions. If the strong semantics of the meshes between expressions are consistent, then these meshes will be projected onto different view images under various perspectives, so that the same mesh area corresponds to the same face area; otherwise, there will be a certain offset. Based on the resulting offset, the correct position of each vertex of the mesh in each view image under the constraint of semantic consistency can be obtained. Then, the mesh is deformed vertex-by-vertex through multi-view positioning to obtain a mesh with preliminary enhanced semantic consistency.
[0066] For example, suppose an individual has two expressions, A and B, where expression A is the unique original expression and expression B is the corresponding multi-view expression. In order to achieve more precise semantic consistency between expressions, the embodiments of this application can map the semantics of the expression B mesh onto the expression A mesh.
[0067] Specifically, firstly, in this embodiment, the A-expression mesh can be projected onto the expression map at various viewpoints based on camera parameters (which can be provided in the original acquired data or obtained through registration), thereby obtaining the two-dimensional positions of the vertices of the A-expression mesh on the A-expression map; secondly, a neural network-based optical flow solver is used to calculate the dense optical flow from the A-expression map to the B-expression map at each viewpoint, i.e., the first dense optical flow; thirdly, for each vertex of the A-expression mesh, the two-dimensional position of the vertex on the B-expression map is obtained based on the optical flow calculation results, and by utilizing the topological consistency of the A and B-expression meshes, the correct positions of each vertex of the B-expression mesh projected onto the B-expression map at each viewpoint under the constraint of semantic consistency can be directly obtained; finally, multi-view positioning is used to reposition the positions of each vertex of the B-expression mesh to obtain the deformed B-expression mesh, and the semantics of the obtained mesh should be coarsely aligned with the A-expression mesh.
[0068] In step S103, based on the preliminarily semantically consistent enhanced facial expression mesh, the second dense optical flow between the UV maps of each facial expression mesh is calculated, and the coordinates of the vertices of each facial expression mesh are corrected based on the second dense optical flow to obtain strongly semantically consistent multi-expression face topological mesh data.
[0069] Furthermore, in some embodiments, the second dense optical flow between the UV maps of each expression mesh is calculated based on the preliminary semantic consistency enhancement of the expression mesh, and the coordinates of the vertices of each expression mesh are corrected based on the second dense optical flow, including: projecting the first expression mesh and the deformed second expression mesh onto the corresponding original 3D face scan image to obtain the 3D coordinates and material information of each vertex projected onto the original 3D face scan image; using topological consistency, projecting the first expression mesh and the deformed second expression mesh onto the same UV space, and generating a UV map and a 3D coordinate distribution field based on the 3D coordinates and material information; and calculating the second dense optical flow from the UV map of the first expression mesh to the UV map of the deformed second expression mesh through a second optical flow solver.
[0070] Furthermore, in some embodiments, correcting the coordinates of the vertices of each expression mesh based on the second dense optical flow includes: obtaining the UV coordinates of each vertex of the first expression mesh on the UV map of the deformed second expression mesh under semantic consistency constraints according to the second dense optical flow; sampling the three-dimensional coordinates of the original three-dimensional face scan map of the deformed second expression mesh from the three-dimensional coordinate distribution field of the deformed second expression mesh; and correcting the deformed second expression mesh based on the coordinates of each vertex of the first expression mesh in the original three-dimensional face scan map of the deformed second expression mesh.
[0071] Specifically, in daily life, cameras are mostly used to collect dense semantic information about faces. However, due to issues such as parameter errors in cameras, in order to obtain more semantically consistent expression meshes, a 3D face scanning method can be used. This involves projecting the expression meshes onto their respective original 3D face scans, obtaining the 3D projection position and material of the meshes on the scan, and using topological consistency to project the original expression meshes and the deformed expression meshes onto the same UV space to generate UV maps. In other words, the first expression mesh and the deformed second expression mesh are projected onto the same UV space, and UV maps are generated based on 3D coordinates and material information. Under the condition of semantic consistency, the face regions in the UV maps of each expression mesh should overlap; otherwise, there will be a certain offset.
[0072] Furthermore, under the condition of fine semantic alignment, the offset can also be corrected by using a light-flowing method, thereby obtaining the correct 3D projection position of the expression mesh in the original face scan under the constraint of semantic consistency, thus correcting the vertex position of each expression mesh, and finally obtaining a strongly semantically consistent mesh.
[0073] It should be noted that the first and second optical flow solvers used in this embodiment are both trained using pre-trained neural networks for optical flow calculation. For example, based on the A and B expressions obtained in the above steps, firstly, the A and B expression meshes are projected onto their respective original 3D face scans to obtain the 3D coordinates and material information of each vertex projected onto the scan; secondly, using topological consistency, the A and B expression meshes are projected into the same UV space, and the obtained 3D coordinates and material information are used to generate UV maps and 3D coordinate distribution fields; thirdly, using a neural network-based optical flow solver, the dense optical flow from the A expression UV map to the B expression UV map is calculated; finally, for each vertex of the A expression mesh, the UV coordinates of the vertex on the B expression UV map are obtained according to the optical flow calculation results, and the coordinates of the vertex on the original 3D face scan of the B expression are directly sampled from the 3D coordinate distribution field of the B expression, and topological consistency is used to obtain the correct position of each vertex of the B expression mesh on the face scan.
[0074] Therefore, the correct positions of each vertex obtained through the above steps are used to further correct the B expression mesh, thereby obtaining B expression mesh data that is semantically aligned with the A expression mesh.
[0075] In summary, the method for acquiring multi-expression face topological mesh data in this application embodiment can obtain semantic information from raw face data such as 3D face scans and multi-view images, use optical flow methods to establish semantic correspondences between expressions, determine the correct positions of each vertex of the expression face mesh based on these relationships, and further deform the mesh to enhance the semantic consistency between expressions; that is, this application embodiment can obtain strongly semantically consistent multi-expression face topological mesh data through coarse semantic alignment steps and fine semantic alignment steps.
[0076] According to the method for acquiring multi-expression face topological mesh data according to embodiments of this application, dense semantic information of faces is obtained from the original face database, and a first dense optical flow is calculated between different expression view maps under various perspectives. Based on the first dense optical flow and a multi-view localization method, the vertices of each expression mesh are deformed vertex by vertex to obtain an expression mesh with preliminary enhanced semantic consistency. Then, a second dense optical flow is calculated between the UV maps of each expression mesh, and the coordinates of the vertices of each expression mesh are corrected based on the second dense optical flow to obtain multi-expression face topological mesh data with strong semantic consistency. Thus, semantic information is obtained from the original face acquisition data, and semantic relationships between expressions are constructed using optical flow. The individually registered multi-expression mesh data is deformed to enhance its semantic consistency, solving the problem of semantic inconsistency between expressions in multi-expression face mesh data of the same individual obtained through individual registration.
[0077] Next, with reference to the accompanying drawings, a device for acquiring multi-expression face topology mesh data according to an embodiment of this application is described.
[0078] Figure 2 This is a block diagram of a device for acquiring multi-expression facial topology mesh data according to an embodiment of this application.
[0079] like Figure 2 As shown, the multi-expression face topology mesh data acquisition device 10 includes: acquisition module 100, first calculation module 200 and second calculation module 300.
[0080] The acquisition module 100 is used to acquire dense semantic information about faces from the original face database;
[0081] The first calculation module 200 is used to calculate the first dense optical flow between different facial expression viewpoints based on semantic information, and to deform the vertices of each facial expression mesh vertex by vertex based on the first dense optical flow and a multi-view localization method to obtain a preliminary semantically consistent enhanced facial expression mesh; and
[0082] The second calculation module 300 is used to calculate the second dense optical flow between the UV maps of each expression mesh based on the preliminary semantic consistency enhancement expression mesh, and correct the coordinates of the vertices of each expression mesh based on the second dense optical flow to obtain strong semantic consistency multi-expression face topological mesh data.
[0083] Furthermore, in some embodiments, the first computing module 200 is specifically used for:
[0084] Projecting the first expression mesh onto the expression map from multiple perspectives yields the two-dimensional positions of each vertex of the first expression mesh in the first expression map.
[0085] The first dense optical flow from the first expression map to the second expression map is calculated using the first optical flow solver, and the two-dimensional positions of each vertex of the first expression mesh projected onto the second expression map under semantic consistency constraints are obtained based on the first dense optical flow.
[0086] By utilizing the topological consistency between the first and second expression meshes, the two-dimensional positions of each vertex of the second expression mesh projected onto the second expression graph under semantic consistency constraints are obtained.
[0087] Using a multi-view localization method, each vertex of the second expression mesh is relocated to obtain the deformed second expression mesh. The first expression mesh and the deformed second expression mesh are then used as the initial semantic consistency enhancement expression mesh.
[0088] Furthermore, in some embodiments, the second computing module 300 is specifically used for:
[0089] The first expression mesh and the deformed second expression mesh are projected onto the corresponding original 3D face scan image to obtain the 3D coordinates and material information of each vertex projected onto the original 3D face scan image;
[0090] By utilizing topological consistency, the first facial expression mesh and the deformed second facial expression mesh are projected onto the same UV space, and a UV map and a three-dimensional coordinate distribution field are generated based on the three-dimensional coordinates and material information.
[0091] The second dense optical flow is calculated from the UV map of the first expression mesh to the UV map of the deformed second expression mesh using the second optical flow solver.
[0092] Furthermore, in some embodiments, the second computing module 300 is specifically used for:
[0093] The UV coordinates of each vertex of the first expression mesh on the UV map of the second expression mesh after deformation under semantic consistency constraints are obtained based on the second dense optical flow.
[0094] The three-dimensional coordinates of the original three-dimensional face scan image of the deformed second expression mesh are obtained by sampling the three-dimensional coordinate distribution field of the first expression mesh under semantic consistency constraints.
[0095] The second expression mesh is corrected based on the coordinates of each vertex of the first expression mesh in the original 3D face scan image of the deformed second expression mesh.
[0096] Furthermore, in some embodiments, both the first optical flow solver and the second optical flow solver are trained using a pre-trained neural network for optical flow calculation.
[0097] The apparatus for acquiring multi-expression face topological mesh data according to embodiments of this application obtains dense semantic information of faces from an original face database, calculates the first dense optical flow between different expression view maps under various perspectives, and deforms the vertices of each expression mesh vertex by vertex based on the first dense optical flow and a multi-view localization method to obtain an expression mesh with preliminary enhanced semantic consistency. Then, it calculates the second dense optical flow between the UV maps of each expression mesh and corrects the coordinates of the vertices of each expression mesh based on the second dense optical flow to obtain multi-expression face topological mesh data with strong semantic consistency. Thus, by obtaining semantic information from the original face acquisition data, constructing semantic relationships between expressions using optical flow, and deforming individually registered multi-expression mesh data to enhance its semantic consistency, the apparatus solves the problem of semantic inconsistency between expressions in multi-expression face mesh data of the same individual obtained through individual registration.
[0098] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include:
[0099] The memory 301, the processor 302, and the computer program stored on the memory 301 and capable of running on the processor 302.
[0100] When the processor 302 executes the program, it implements the method for obtaining multi-expression face topology mesh data provided in the above embodiments.
[0101] Furthermore, electronic devices also include:
[0102] Communication interface 303 is used for communication between memory 301 and processor 302.
[0103] The memory 301 is used to store computer programs that can run on the processor 302.
[0104] The memory 301 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0105] If the memory 301, processor 302, and communication interface 303 are implemented independently, then the communication interface 303, memory 301, and processor 302 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 3 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0106] Optionally, in a specific implementation, if the memory 301, processor 302, and communication interface 303 are integrated on a single chip, then the memory 301, processor 302, and communication interface 303 can communicate with each other through an internal interface.
[0107] Processor 302 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.
[0108] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described method for acquiring multi-expression facial topology mesh data.
[0109] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0110] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0111] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0112] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0113] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0114] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by a program instructing related hardware, and the program can be stored in a computer-readable storage medium. When executed, the program includes one or a combination of the steps of the method embodiments.
[0115] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0116] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A method for acquiring multi-expression facial topological mesh data, characterized in that, Includes the following steps: Extract dense semantic information about faces from the original face database; Based on the semantic information, the first dense optical flow between different facial expression view maps under each view is calculated, and based on the first dense optical flow and the multi-view positioning method, the vertices of each facial expression mesh are deformed vertex by vertex to obtain a preliminary semantically consistent enhanced facial expression mesh. as well as Based on the preliminarily semantically consistent enhanced facial expression mesh, the second dense optical flow between the UV maps of each facial expression mesh is calculated, and the coordinates of the vertices of each facial expression mesh are corrected based on the second dense optical flow to obtain strongly semantically consistent multi-expression face topological mesh data. The process of calculating the first dense optical flow between different facial expression viewpoints based on the semantic information, and deforming the vertices of each facial expression mesh vertex by vertex based on the first dense optical flow and the multi-view localization method to obtain a preliminary semantically consistent enhanced facial expression mesh includes: Projecting the first expression mesh onto the expression map from multiple perspectives yields the two-dimensional positions of each vertex of the first expression mesh in the first expression map. The first dense optical flow from the first expression map to the second expression map is calculated using the first optical flow solver, and the two-dimensional positions of each vertex of the first expression mesh projected onto the second expression map under semantic consistency constraints are obtained based on the first dense optical flow. By utilizing the topological consistency between the first and second expression grids, the two-dimensional positions of each vertex of the second expression grid projected onto the second expression graph under semantic consistency constraints are obtained. Using the multi-view localization method, each vertex of the second expression mesh is relocated to obtain the deformed second expression mesh, and the first expression mesh and the deformed second expression mesh are used as the expression mesh for preliminary semantic consistency enhancement. The calculation of the second dense optical flow between the UV maps of each facial expression mesh based on the preliminary semantic consistency enhancement, and the correction of the vertex coordinates of each facial expression mesh based on the second dense optical flow, includes: The first expression mesh and the deformed second expression mesh are projected onto the corresponding original 3D face scan image to obtain the 3D coordinates and material information of each vertex projected onto the original 3D face scan image; Utilizing the topological consistency, the first facial expression mesh and the deformed second facial expression mesh are projected into the same UV space, and a UV map and a three-dimensional coordinate distribution field are generated based on the three-dimensional coordinates and the material information; The second dense optical flow from the UV map of the first expression mesh to the UV map of the deformed second expression mesh is calculated using the second optical flow solver. The step of correcting the coordinates of the vertices of each facial expression mesh based on the second dense optical flow includes: The UV coordinates of each vertex of the first facial expression mesh on the UV map of the deformed second facial expression mesh under semantic consistency constraints are obtained based on the second dense optical flow. The three-dimensional coordinates of the original three-dimensional face scan image of the deformed second expression mesh are obtained by sampling from the three-dimensional coordinate distribution field of the first expression mesh under semantic consistency constraints. The deformed second expression mesh is corrected based on the coordinates of each vertex of the first expression mesh in the original 3D face scan image of the deformed second expression mesh.
2. The method according to claim 1, characterized in that, Both the first optical flow solver and the second optical flow solver are trained using a pre-trained neural network for optical flow calculation.
3. A device for acquiring multi-expression facial topological mesh data, characterized in that, include: The acquisition module is used to obtain dense semantic information about faces from the original face database; The first calculation module is used to calculate the first dense optical flow between different facial expression view maps under each view based on the semantic information, and deform the vertices of each facial expression mesh vertex by vertex based on the first dense optical flow and the multi-view positioning method to obtain a preliminary semantically consistent enhanced facial expression mesh. as well as The second calculation module is used to calculate the second dense optical flow between the UV maps of each expression grid based on the preliminary semantically consistent enhanced expression grid, and correct the coordinates of the vertices of each expression grid based on the second dense optical flow to obtain strong semantically consistent multi-expression face topology grid data. The first calculation module is specifically used for: Projecting the first expression mesh onto the expression map from multiple perspectives yields the two-dimensional positions of each vertex of the first expression mesh in the first expression map. The first dense optical flow from the first expression map to the second expression map is calculated using the first optical flow solver, and the two-dimensional positions of each vertex of the first expression mesh projected onto the second expression map under semantic consistency constraints are obtained based on the first dense optical flow. By utilizing the topological consistency between the first and second expression grids, the two-dimensional positions of each vertex of the second expression grid projected onto the second expression graph under semantic consistency constraints are obtained. Using the multi-view localization method, each vertex of the second expression mesh is relocated to obtain the deformed second expression mesh, and the first expression mesh and the deformed second expression mesh are used as the expression mesh for preliminary semantic consistency enhancement. The second calculation module is specifically used for: The first expression mesh and the deformed second expression mesh are projected onto the corresponding original 3D face scan image to obtain the 3D coordinates and material information of each vertex projected onto the original 3D face scan image; Utilizing the topological consistency, the first facial expression mesh and the deformed second facial expression mesh are projected into the same UV space, and a UV map and a three-dimensional coordinate distribution field are generated based on the three-dimensional coordinates and the material information; The second dense optical flow from the UV map of the first expression mesh to the UV map of the deformed second expression mesh is calculated using the second optical flow solver. The step of correcting the coordinates of the vertices of each facial expression mesh based on the second dense optical flow includes: The UV coordinates of each vertex of the first facial expression mesh on the UV map of the deformed second facial expression mesh under semantic consistency constraints are obtained based on the second dense optical flow. The three-dimensional coordinates of the original three-dimensional face scan image of the deformed second expression mesh are obtained by sampling from the three-dimensional coordinate distribution field of the first expression mesh under semantic consistency constraints. The deformed second expression mesh is corrected based on the coordinates of each vertex of the first expression mesh in the original 3D face scan image of the deformed second expression mesh.
4. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the method for acquiring multi-expression facial topological mesh data as described in any one of claims 1-2.
5. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the method for acquiring multi-expression facial topological mesh data as described in any one of claims 1-2.