Mapping Optimization Method, Device, Equipment and Medium for Agent Facial Expression Imitation
By using a mixed shape substrate for cross-modal mapping in the mapping stage of expression imitation, the problem of semantic distortion in the prior art is solved, and the consistent migration of expression semantics and efficient mapping process are realized.
Patent Information
- Application Number
- CN202510300985.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-14
AI Technical Summary
The prior art has problems with semantic distortion and the inability to directly drive actual physical executors during the mapping migration stage of expression imitation.
By introducing a mixed shape substrate, a first mixed shape substrate for a normalized average face is generated, and dense expression motion information is mapped to it to generate a second mixed shape substrate, passing its coefficients to the third mixed shape substrate for the intelligent face, achieving cross-modal mapping and semantic consistency.
The geometric differences between human faces and intelligent faces are solved, the consistency of expression semantics during the migration process is ensured, and the efficiency and robustness of the mapping stage are improved.
Smart Images

Figure CN119830945B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, specifically to the field of image data processing technology, and more specifically to a mapping optimization method, device, equipment, medium and product for intelligent agent expression imitation. Background Art
[0002] Artificial Intelligence (AI for short) is an important driving force for the new round of scientific and technological revolution and industrial transformation. It is a new key technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence. As an important part of intelligent science, artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine (such as a highly biomimetic expression robot) that can respond in a way similar to human intelligence.
[0003] As an important part of artificial intelligence technology, the technology of human expression recognition and reproduction (such as expression imitation) has received extensive attention in the fields of human-computer interaction, security recognition, robot manufacturing, automation, medical treatment, communication, and driving, and has quickly become a research hotspot in the industry. The expression imitation technology is mainly divided into three stages: expression perception, mapping migration, and driving execution. The aim is to accurately capture the dynamic expression features of the human face in the expression perception stage, realize the face mapping to embodied intelligence such as highly biomimetic expression robots in the mapping migration stage, and finally control the expression actuator to achieve accurate expression reproduction drive in the driving execution stage. However, in the existing traditional expression imitation technology, there are problems such as easy to generate serious semantic distortion and unable to directly drive actual physical actuators in the mapping migration stage. Summary of the Invention
[0004] In view of at least one of the above problems, the embodiments of the present invention aim to provide a mapping optimization method, device, equipment, medium and product for intelligent agent expression imitation, thereby providing an optimization solution for the mapping stage of intelligent agent expression imitation. By using a blend shape basis to accurately depict expression semantics, introducing a unified expression parameter space, and establishing a cross-modal mapping mechanism from a standardized human face expression to a robot expression space, the differences in geometric structures between the human face and the robot face are solved, and the consistency of expression semantics during the migration process is ensured.
[0005] One aspect of the embodiments of the present invention provides a mapping optimization method for intelligent agent expression imitation, which includes: generating a first blend shape basis of a standardized average face based on a preset facial action coding rule; mapping the dense expression motion information of the standardized average face to the first blend shape basis to generate a second blend shape basis; and transferring the blend shape coefficients of the second blend shape basis to a third blend shape basis of the intelligent agent's face to complete the mapping optimization of intelligent agent expression imitation.
[0006] According to an embodiment of the present invention, in the first blend shape basis for generating a normalized average face based on a preset facial action coding rule, it includes: generating a blend shape expression basis matrix according to the preset facial action coding rule; generating the first blend shape basis of the normalized average face based on the preset blend shape weight vector and the blend shape expression basis matrix.
[0007] According to an embodiment of the present invention, before mapping the dense expression motion information of the normalized average face to the first blend shape basis to generate the second blend shape basis, it further includes: generating the dense expression motion information of the normalized average face according to the expression motion information of the extracted target expression image.
[0008] According to an embodiment of the present invention, in mapping the dense expression motion information of the normalized average face to the first blend shape basis to generate the second blend shape basis, it includes: extracting the first dense annotation points on the dense expression grid corresponding to the dense expression motion information and the second dense annotation points on the average face grid; obtaining the dense annotation basis matrix according to the blend shape expression basis matrix; generating a blend shape weight optimization task according to the dense annotation basis matrix, the first dense annotation points and the second dense annotation points.
[0009] According to an embodiment of the present invention, in mapping the dense expression motion information of the normalized average face to the first blend shape basis to generate the second blend shape basis, it further includes: obtaining the blend shape coefficients corresponding to the blend shape weight optimization task by minimizing a preset error vector and a preset coefficient constraint condition, so as to realize mapping the dense expression motion information of the normalized average face to the first blend shape basis to generate the second blend shape basis.
[0010] According to an embodiment of the present invention, before transferring the blend shape coefficients of the second blend shape basis to the third blend shape basis of the agent's face and completing it, it further includes: generating the third blend shape basis of the agent's face based on the preset facial action coding rule, wherein the blend shape semantics of the third blend shape basis is the same as that of the first blend shape basis.
[0011] According to an embodiment of the present invention, in transferring the blend shape coefficients of the second blend shape basis to the third blend shape basis of the agent's face and completing it, it includes: applying the blend shape coefficients to the third blend shape basis to generate the agent expression grid corresponding to the agent.
[0012] Another aspect of an embodiment of the present invention provides a mapping optimization device for agent expression imitation, which includes a basis generation module, an information mapping module, and a coefficient transfer module. The basis generation module is used to generate a first blend shape basis of a standardized average face based on a preset facial action coding rule; the information mapping module is used to map the dense expression motion information of the standardized average face to the first blend shape basis to generate a second blend shape basis; and the coefficient transfer module is used to transfer the blend shape coefficients of the second blend shape basis to a third blend shape basis of the agent's face to complete the mapping optimization of agent expression imitation.
[0013] Another aspect of an embodiment of the present invention provides an electronic device, including one or more processors and a memory, where the memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned mapping optimization method for agent expression imitation.
[0014] Another aspect of an embodiment of the present invention provides a computer-readable storage medium, on which executable instructions are stored. When the instructions are executed by a processor, the processor is caused to execute the above-mentioned mapping optimization method for agent expression imitation.
[0015] Another aspect of an embodiment of the present invention provides a computer program product, including a computer program, which when executed by a processor implements the above-mentioned mapping optimization method for agent expression imitation.
[0016] The mapping optimization method for agent expression imitation provided by the embodiment of the present invention can at least partially solve the problems in the related art and thus can at least achieve one of the following technical effects:
[0017] Therefore, in the mapping stage, the above-mentioned mapping optimization method for agent expression imitation in the embodiment of the present invention solves the problems of differences in topological structure and geometric characteristics between the human face and the agent's face through a blend shape semantic space of a second blend shape basis based on the standardized average face as an intermediate bridge. Specifically, first, a set of first blend shape bases is defined on the standardized average face for dense representation of expression semantics. The first blend shape bases can then map the dense expression motion generated in the perception stage (refer to Embodiment 1 above) to the blend shape coefficient space through dense semantic key points to form a second blend shape basis. During this process, by transferring the blend shape coefficients, the expression motion is mapped from the topological structure of the standardized average face to the topological structure of the agent's face, realizing semantic-consistent expression transfer across topological structures.
[0018] Therefore, through the conversion from the geometric domain to the semantic domain, the semantic distortion problem caused by the topological differences between human faces and robot faces in the traditional solution is solved. In addition, the number of hybrid shape coefficients is limited, which is convenient for real-time transmission, thus greatly improving the efficiency and robustness of the mapping stage. At the same time, the generated hybrid shape coefficients directly correspond to the motion state of the robot's epidermis (such as Figure 5C the skin Skin shown), ensuring the physical executability of the generated expressions. It can be seen that in the cross-entity mapping process of expression imitation, the hybrid shape basis designed based on the above-mentioned preset facial action coding rules can be used to accurately depict expression semantics. By introducing a unified expression parameter space, a cross-modal mapping mechanism from the dense facial expressions of the standardized average face to the robot expression space can be established, thus solving the differences in geometric structures between human faces and robot faces and ensuring the consistency of expression semantics during the migration process.
[0019] It should be understood that the above general description and the following specific embodiments are only exemplary and explanatory, and they do not limit the scope of what the present invention intends to claim. Description of the Drawings
[0020] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features, and advantages of the present invention will become clearer. In the drawings:
[0021] Figure 1 Schematically shows an application scenario diagram of an optimization method, device, equipment, medium, and program product for agent expression imitation according to Embodiments 1-3 of the present invention;
[0022] Figure 2A Schematically shows a flowchart of a perception optimization method for agent expression imitation according to Embodiment 1 of the present invention;
[0023] Figure 2B Schematically shows a structural block diagram of a perception optimization device for agent expression imitation according to Embodiment 1 of the present invention;
[0024] Figure 3A Schematically shows a flowchart of a mapping optimization method for agent expression imitation according to Embodiment 2 of the present invention;
[0025] Figure 3B Schematically shows a structural block diagram of a mapping optimization device for agent expression imitation according to Embodiment 2 of the present invention;
[0026] Figure 4A Schematically shows a flowchart of an execution optimization method for agent expression imitation according to Embodiment 3 of the present invention;
[0027] Figure 4BSchematically shows a structural block diagram of an execution optimization device for agent expression imitation according to Embodiment 3 of the present invention;
[0028] Figure 5A Schematically shows an expression semantic alignment diagram from the human face expression mixing space to the robot expression mixing space according to Embodiments 1-3 of the present invention;
[0029] Figure 5B Schematically shows a robot expression kinematics design diagram according to Embodiments 1-3 of the present invention;
[0030] Figure 5C Schematically shows a robot expression electromechanical design diagram according to Embodiments 1-3 of the present invention; and
[0031] Figure 6 Schematically shows a block diagram of an electronic device according to the method of Embodiments 1-3 of the present invention.
[0032] The above-mentioned drawings are part of the specification of the embodiments of the present invention, which illustrate exemplary embodiments of the present invention. The accompanying drawings, together with the description of the specification, are used to explain the principles of the embodiments of the present invention. It should be understood that the above general description of the drawings and the following detailed description are only exemplary and explanatory, and they do not limit the scope that the present invention intends to claim. Detailed Embodiments
[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the spirit of what the present invention discloses will be clearly described below with reference to the drawings and in detail. After any person skilled in the art understands the embodiments of the content of the present invention, they can make changes and modifications based on the techniques taught by the content of the present invention, and it does not depart from the spirit and scope of the content of the present invention.
[0034] The schematic embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention. Additionally, elements / components using the same or similar reference numerals in the drawings and embodiments are used to represent the same or similar parts.
[0035] Regarding the "first", "second",... etc. used in the present invention, they do not particularly refer to the meaning of order or sequence, nor are they used to limit the present invention. They are only used to distinguish elements or operations described with the same technical terms.
[0036] Regarding the direction terms used in the present invention, such as: up, down, left, right, front or back, etc., they are only references to the directions in the drawings. Therefore, the direction terms used are for explanation and not for limiting this creation.
[0037] The terms "comprising", "including", "having", "containing", etc. used in the present invention are all open-ended terms, meaning including but not limited to.
[0038] The term "and / or" used in the present invention includes any one or all combinations of things.
[0039] The "plurality" in the present invention includes "two" and "more than two"; the "multiple groups" in the present invention includes "two groups" and "more than two groups".
[0040] The terms "substantially", "about", etc. used in the present invention are used to modify any quantity or error that can vary slightly, but these slight variations or errors do not change its essence. Generally, the range of such slight variations or errors modified by such terms can be 20% in some embodiments, 10% in some embodiments, 5% in some embodiments, or other values. Those skilled in the art should understand that the aforementioned values can be adjusted according to actual needs and are not limited thereto.
[0041] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0042] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but not be limited to a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). In the case of using expressions such as "at least one of A, B, or C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, or C" should include but not be limited to a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). Those skilled in the art should also understand that substantially any disjunctive conjunction and / or phrase representing two or more alternative items, whether in the specification, claims, or drawings, should be understood as giving the possibility of including one of these items, either of these items, or both items. For example, the phrase "A or B" should be understood as including the possibility of "A" or "B", or "A and B".
[0043] In addition, all actions of obtaining information, signals, or data in the present invention are carried out on the premise of complying with the corresponding data protection laws, regulations, and policies of the country where it is located and obtaining authorization from the owner of the corresponding device.
[0044] In human interpersonal communication in daily life, by controlling facial expressions, the communication effect can be enhanced. For humans, facial expressions are the result of one or more actions or states of facial muscles. These movements express the emotional state of an individual human being towards the external environment. Facial expressions are a form of non-verbal communication and are the main means of expressing social information between humans.
[0045] For artificial intelligence, in order for intelligent machines to make real human interaction responses in a way similar to human intelligence, it is necessary to create a truly humanoid machine in all aspects, including limbs, appearance, body posture, expressions, and even voice conversations, mobility, and human-like thinking patterns. Among them, the expression imitation technology provides strong technical support for today's humanoid robots to enter the real world and provide emotional and non-verbal power to humans.
[0046] The expression imitation technology aims to accurately capture the dynamic characteristics of human facial expressions and achieve precise mapping and physical drive to highly biomimetic expression robots. Its overall technical framework covers three links: expression perception, mapping migration, and drive execution. No matter which link has problems, it will directly limit the realization of high-precision expression imitation.
[0047] To solve at least one of the existing technical problems in the three links of expression perception, mapping migration, and drive execution in the process of intelligent agent expression imitation in the prior art, the present invention provides Examples 1-3 to respectively optimize the above three links in the process of intelligent agent expression imitation, with the expectation of significantly improving the intelligent level of intelligent agent expression imitation.
[0048] Figure 1 Schematically shows the application scenario diagrams of the optimization methods, devices, equipment, media, and program products for intelligent agent expression imitation according to Embodiments 1-3 of the present invention.
[0049] As Figure 1 shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, network 104, and server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0050] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0051] Terminal devices 101, 102, and 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, desktop computers, and so on.
[0052] Server 105 can be a server that provides various services, such as a background management server that supports the websites browsed by users using terminal devices 101, 102, and 103 (for example only). The background management server can analyze and process data such as received user requests, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0053] It should be noted that the optimization method for agent expression imitation provided by the embodiments of the present invention can generally be executed by server 105. Correspondingly, the optimization device for agent expression imitation provided by the embodiments of the present invention can generally be set in server 105. The optimization method for agent expression imitation provided by the embodiments of the present invention can also be executed by a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the optimization device for agent expression imitation provided by the embodiments of the present invention can also be set in a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.
[0054] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0055] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.
[0056] In order to further reflect the complete technical process of the three key stages of perception, mapping and execution in the imitation of intelligent body expression, and to ensure that the full-link expression imitation from the original face input to the accurate reproduction of the robot expression can be achieved, the following is combined with the embodiments 1-3 of the present invention and Figures 1 - 5C Further explanation is as follows:
[0057] It should be noted that the intelligent agent may be the execution subject of at least one of the perception optimization method, mapping optimization method and execution optimization method provided in the above-mentioned embodiments 1-3 of the present invention, or may be the execution party controlled by these methods, such as a humanoid intelligent robot or other AI device with a simulated human face skin that can imitate human facial expressions. Specifically, Figure 5B and Figure 5C The electromechanical design of the intelligent body with facial expressions shown in the figure covers the surface with "muscle" and "skeleton" with a simulated skin (silicone mask, i.e. skin). Thus, the movement of the simulated skin can be realized through the motor (such as LFD-01M servomotor) and steel cable and Teflon conduit matched with the electrical control system, thereby achieving the imitation of facial expressions.
[0058] The following will be combined Figures 2A - 5C The perception optimization method, mapping optimization method and execution optimization method provided in Examples 1-3 of the present invention are described in detail respectively.
[0059] Example 1
[0060] The expression perception stage involves facial expression recognition. Achieving accurate expression recognition perception, obtaining more complex expression descriptions, and ensuring the accuracy and versatility of expression perception are technical issues that need to be urgently resolved in the embodiments of the present invention.
[0061] Specifically, in the perception stage, traditional methods usually use sparse landmarks to perform structured representation of faces. This method relies on the position and displacement of landmarks to describe the dynamic changes of facial expressions. Although simple and efficient, its main problem is that the complexity of expression description is insufficient. Since sparse landmarks can only capture limited local dynamic features, they cannot fully represent the subtle changes and complex characteristics of human facial expressions, resulting in oversimplification of expression description. In addition, there is a coupling relationship between face shape and expression semantics, which makes it impossible for the same geometric displacement to maintain consistent expression semantics in individuals with different face shapes. This coupling further limits the accuracy and versatility of expression perception.
[0062] Therefore, in view of at least one of the above-mentioned technical problems existing in the prior art, embodiments of the present invention aim to provide a method, apparatus, device, medium, and product for optimizing the perception of agent expression imitation, thereby providing an optimization solution for the perception stage of agent expression imitation, in order to achieve effective decoupling of individual face shape features and expression dynamics, significantly improve the ability to capture subtle expressions through dense expression representations, and overcome the limitations of traditional solutions.
[0063] The following will be based on Figure 1 the described scenario, and through Figures 2A - 5C describe in detail the method for optimizing the perception of agent expression imitation in the disclosed embodiments.
[0064] As Figure 2A shown, one aspect of the embodiments of the present invention provides a method for optimizing the perception of agent expression imitation, which includes operations S201 to S204.
[0065] In operation S201, extract the expression motion information of the target expression image;
[0066] In operation S202, generate a dense face reconstruction task through the expression motion information and preset joint optimization information;
[0067] In operation S203, perform joint optimization on the dense face reconstruction task to generate three-dimensional face reconstruction information, and the three-dimensional face reconstruction information is used to generate the target three-dimensional face; and
[0068] In operation S204, extract the dense expression motion information of the standardized average face according to the target three-dimensional face to complete the perception optimization.
[0069] In the embodiments of the present invention, agent expression imitation can be a process in which an agent recognizes and obtains a face expression image, and reproduces the face expression in the face expression image using the above Figure 5C electromechanical design. Among them, the face expression image as the object to be recognized can be image information directly detected and obtained by the agent through its own image detection device (such as a camera image sensor, etc.), or image information transmitted to the agent through a network transmission device (wireless or wired), and these image information can be video or picture data.
[0070] For example, as a type of intelligent agent, a humanoid robot conducts a video call with a real human user. The humanoid robot can directly or indirectly obtain the video data of the video call, preprocess the video content of the human user included in the video data who is the call object, and extract the expression features from each frame of the image containing the human user's expression in the video content. Among them, the target expression image can be one of these frame images, and this single-frame image contains the facial expression of the expression imitation object (human user) at a certain moment. In addition, the humanoid robot can also directly interact with the human user face to face, and obtain and recognize the real-time expression of the human user in real time.
[0071] The expression motion information can be the facial expression feature parameters extracted for the facial expression in the target expression image. These expression feature parameters can involve the expression basis parameters and the facial expression motion parameters of the target facial expression in the target expression image (such as the facial shape coefficient, dynamic expression motion coefficient, and facial pose coefficient in the camera coordinate system, etc.). Among them, the so-called expression basis parameters can include more (such as 50) different expression bases, and each expression basis coefficient represents the displacement of each corner point on the dense mesh of the facial expression of the current facial image relative to the standardized average face; in addition, the facial expression motion parameters are the motion field parameters relative to the standardized average face in the standardized average face coordinate system.
[0072] The preset joint optimization information can be the information for three-dimensional face reconstruction based on the standardized average face and the expression motion information according to the target three-dimensional face reconstruction requirements, where the target three-dimensional face reconstruction requirements are also related to the target three-dimensional face reconstruction model used for this three-dimensional face reconstruction. The standardized average face is the three-dimensional face average shape obtained through statistical learning (i.e., the three-dimensional average face, Mean Shape). The preset joint optimization information is mainly used to generate a dense face reconstruction task according to the expression motion information of the extracted target expression image. The dense face reconstruction task can be a joint optimization task for restoring the three-dimensional face shape and expression motion parameters according to the target three-dimensional face reconstruction requirements. For example, it can be embodied as a joint optimization problem created for three-dimensional face shape restoration.
[0073] By performing joint optimization processing on the dense face reconstruction task, three-dimensional face reconstruction information can be generated according to the processing result of the dense face reconstruction task, where the three-dimensional face reconstruction information can include the corresponding facial expression basis and the corresponding facial expression coefficients . The facial expression basis is usually used to describe the dynamic expression changes of the face, and the facial expression coefficients are usually used as the parameterized representation of the facial expression. The facial expression basis and the facial expression coefficients can be configured to construct a three-dimensional face expression model based on the standardized average face.
[0074] The target 3D face can be understood as a 3D face expression model based on the standardized average face. For example, a 3D deformation model (such as 3D Morphable Model, abbreviated as 3DMM) can be a statistical model for 3D face modeling. By mainly learning a large amount of 3D face data, any 3D face can be represented as a combination of linear bases, which can be represented by the 3D face shape based on the standardized average face. This 3D face shape can be based on the standardized average face (i.e., the 3D average face, Mean Shape), the face shape basis , the face expression basis , the face shape coefficients , the face expression coefficients etc. for description. Therefore, the facial expressions in the 2D image can be mapped to a parameterized 3D representation. By virtue of its decomposition ability for face shape, expression, and pose spaces, the effective decoupling of individual face shape features and expression dynamics can be achieved.
[0075] Among them, the dense expression motion information can be information directly related to the dynamic changes of facial expressions related to the extracted expression motion information, and can be specifically characterized by the linear combination of the face expression basis and the face expression coefficients. The dense expression motion information can most directly reflect the dense expression motion information of the user extracted by the agent on the standardized average face, realizing the facial expression representation.
[0076] Therefore, in the perception stage, compared with the traditional sparse key-point-based expression representation scheme, the above-mentioned perception optimization method for agent expression imitation in the embodiments of the present invention can represent the facial expression motion information by adopting the target 3D face, map the facial expressions in the 2D image to a parameterized 3D representation, and by virtue of its decomposition ability for face shape, expression, and pose spaces, achieve the effective decoupling of individual face shape features and expression dynamics, generate a dense expression motion under the topology of the standardized average face, thereby overcoming the abstract defects of the traditional sparse key-point scheme, realizing the densification of expression description, and at the same time, by normalizing the expression to the standardized average face, eliminating the coupling problem between the expression semantics and the individual face shape, ensuring the consistency of the expression semantics among different face shape individuals. Therefore, by virtue of the high precision and density of the above-mentioned 3D face expression model based on the standardized average face of the target 3D face, the accuracy and fineness of facial expression perception can be significantly improved, the ability to capture subtle expressions can be enhanced, and thus a high-quality foundation can be laid for the subsequent mapping and execution stages.
[0077] Such as Figures 2A - 5C shown, according to an embodiment of the present invention, in the operation of extracting the expression motion information of the target expression image in S201, it includes: performing expression feature recognition on the target expression image to extract the expression motion information.
[0078] The target expression image can be a single-frame two-dimensional image (such as an RGB image), which can be selected from a set of two-dimensional images based on time series (such as a video), and can be specifically selected sequentially according to the time series. Each time the expression feature recognition is mainly performed on the selected single-frame two-dimensional image for extracting the expression feature parameters of the target human face in the single-frame two-dimensional image. The extracted content mainly includes the feature parameters of the human face expression in the single-frame two-dimensional image. These feature parameters can include expression basis parameters and human face expression motion parameters, etc.
[0079] For a complete video, it can record the human face images at each moment within a certain time period, and at the same time, the human face images at each moment record the instantaneous expression images of the human face at that moment. These instantaneous expression images are the above-mentioned target expression images, serving as the targets for expression feature extraction. Therefore, by completing the expression feature recognition for all the target expression images in the above video one by one, and realizing the dense expression motion representation of each target expression image one by one, the perception and acquisition of the human face expression within the corresponding time period of the entire video can be achieved.
[0080] Therefore, through the above-mentioned expression feature recognition, the accurate expression feature extraction for each frame of two-dimensional image can be realized, thereby laying a foundation for the capture of micro-expressions and improving the accuracy and fineness of expression perception.
[0081] As Figures 2A - 5C shown, according to an embodiment of the present invention, in the operation of generating a dense human face reconstruction task through expression motion information and preset joint optimization information in S201, it includes: obtaining the task reconstruction information corresponding to the dense human face reconstruction task according to the expression parameters of the expression motion information; constructing a dense human face reconstruction task according to the task reconstruction information and the preset joint optimization information.
[0082] The expression parameters of the expression motion information can be the expression feature parameters extracted from the human face expression of the target expression image, specifically involving expression basis parameters and human face expression motion parameters. Among them, the human face expression motion parameters can be the human face shape coefficient, dynamic expression motion coefficient, and human face pose coefficient, etc. in the camera coordinate system.
[0083] To perceive the human face expression motion information of the expression motion information and realize the representation based on the standardized average face, three-dimensional human face reconstruction can be performed based on the standardized average face and the expression motion information. Among them, the task reconstruction information can be used to construct the optimization variable information and objective function information of the joint optimization problem corresponding to the above-mentioned dense human face reconstruction task, and specifically can be determined according to the construction requirements of the dense human face reconstruction task. Among them, the preset joint optimization information can define the construction rules of the dense human face reconstruction task, and through this task construction rule, the dense human face reconstruction task can be generated according to the task reconstruction information.
[0084] Specifically, for the above single-frame target expression image, a dense face reconstruction task is constructed to generate the joint optimization problem required for dense three-dimensional face reconstruction.
[0085] Among them, the single-frame target expression image can be expressed as , is the width of the image, is the height of the image. Among them, the reconstruction target of the dense face reconstruction task is to restore the corresponding three-dimensional face shape and its expression motion parameters .
[0086] According to the task construction rules defined by the preset joint optimization information, the dense face reconstruction task can be formalized into the following joint optimization problem (Formula 1) based on the optimization variable information and objective function information corresponding to the task reconstruction information:
[0087] (1)
[0088] Among them, the optimization variable information includes the face shape coefficient for reconstructing the static face shape features of an individual, the face expression coefficient for reconstructing the dense dynamic expression motion, the face rotation matrix for representing the face pose and the translation vector and the lighting parameter for describing the face lighting condition.
[0089] In addition, the objective function information can include projection error, photometric consistency error, shape and regularization constraint information, etc. Among them, the objective function of the projection error (Landmark Projection Loss) can be expressed as the following formula 2:
[0090] (2)
[0091] Among them, is the camera intrinsic matrix; , is the 3D vertex corresponding to the target three-dimensional face; , is the two-dimensional key point detected in the target expression image (such as sparse corner points like the corners of the eyes and mouth). Among them, can be the index of the two-dimensional key point (Landmarks), ; among them, can be the number of two-dimensional key points of the preset grid, which can be specifically determined according to different application systems of dense grid corner points, such as 146 key points of the Mediapipe standard or 68 key points of Dlib.
[0092] Objective function of photometric consistency error (Photometric Loss) It can be expressed as Equation 3 below:
[0093] (3)
[0094] Where is the face region in the target expression image; is the pixel color in the input target expression image; is the pixel color generated by the rendering model (e.g., based on 3DMM face shape, lighting parameters , material, and camera parameters); can be represented as the pixel position of the target expression image. can be the weight of this objective function , such as = 0.05.
[0095] The regularization constraint information can include regularization terms for face shape and face expression. Among them, can be a face shape coefficient constraint term for constraining the magnitude of the shape coefficient to prevent overfitting; can be a face expression coefficient constraint term for constraining the magnitude of the expression coefficient; can be an additional regularization term (such as the sparsity of the expression coefficient); can be the weight of the additional regularization term , such as = 1.0.
[0096] Therefore, by means of the construction of the above-mentioned dense face reconstruction task, it can lay a foundation for the subsequent reconstruction of the target three-dimensional face based on the standardized average face, thereby replacing the traditional sparse key-point-based facial emotion perception scheme and using the three-dimensional deformation model as the expression representation means in the perception stage, ensuring that the facial expression movement of the face can be decomposed into movements of pose, shape, and expression, and ensuring that the reconstructed target three-dimensional face can achieve refined facial expression description.
[0097] As Figures 2A - 5C shown, according to an embodiment of the present invention, in operation S202 for performing joint optimization on the dense face reconstruction task to generate three-dimensional face reconstruction information, it includes: generating an initial three-dimensional face corresponding to the dense face reconstruction task; performing joint optimization processing on the dense face reconstruction task according to the initial three-dimensional face to generate three-dimensional face reconstruction information.
[0098] Based on the above standardized average face, an initial 3D face is constructed. This initial 3D face can be a standardized 3D face with initial lighting conditions processed on the standardized average face, which can be used as the basis for constructing the target 3D face to ensure that the target 3D face can accurately represent the facial expression information of the target expression image.
[0099] The 3D face reconstruction information is used to construct the target 3D face based on the initial 3D face. It can be mainly generated according to the joint optimization processing results of the dense face reconstruction task, which may include facial expression bases and facial expression coefficients and other basic information for 3D face reconstruction.
[0100] As shown in the aforementioned formula 1, the joint optimization processing process of the dense face reconstruction task can be expressed as a process of solving this formula 1. By solving the joint optimization problem constructed according to the preset joint optimization information, and performing the solution based on the above initial 3D face and expression motion information, the above facial expression bases and facial expression coefficients can be obtained, so as to acquire the 3D face reconstruction information.
[0101] Therefore, by obtaining this 3D face reconstruction information, accurate modeling of the facial geometry and expression motion can be ensured.
[0102] As Figures 2A - 5C shown, according to an embodiment of the present invention, in generating the initial 3D face corresponding to the dense face reconstruction task, it includes: detecting the two-dimensional key point information in the target expression image through a preset key point detection model; obtaining the initial pose parameters of the standardized average face according to the two-dimensional key point information.
[0103] The preset key point detection model can be a pre-trained model for detecting the two-dimensional face key points of the target expression image, such as the MediaPipe model or the Dlib model, etc. Among them, the target expression image can represent the facial expression based on the topological structure of the two-dimensional face. Among them, the two-dimensional key point information can be the two-dimensional information expression of the sparse corner points (such as eyebrows, eyelids, etc.) on the topological structure of the two-dimensional face, and specifically can be represented.
[0104] The two-dimensional key point information (face mesh) can include the position information (such as two-dimensional coordinates) of the above two-dimensional key points in the two-dimensional topological structure. With the help of this position information, the initial pose parameters can be estimated by combining the face pose parameters of the expression motion information. Among them, the initial pose parameters can include the face pose parameters of the target expression image of the current frame relative to the face of the standardized average face, such as the rotation matrix of the face pose relative to the standardized average face and the translation vector .
[0105] Therefore, by obtaining the initial pose parameters, the mapping relationship between the two-dimensional face topology and the standardized average face can be effectively established, so that the extracted two-dimensional key point information can be accurately mapped to the standardized average face, which is beneficial to more comprehensively and accurately representing the subtle changes and complex characteristics of facial expressions in the follow-up.
[0106] As Figures 2A - 5C shown, according to an embodiment of the present invention, in generating the initial three-dimensional face corresponding to the dense face reconstruction task, it further includes: initializing the optimization variable parameters of the dense face reconstruction task based on the initial pose parameters; generating the initial three-dimensional face through the standardized average face and the default lighting condition based on the initialization result of the optimization variable parameters.
[0107] Through the rotation matrix of the face pose of the standardized average face corresponding to the initial pose parameters and the translation vector , the optimization variable parameters corresponding to the dense face reconstruction task can be initialized, where the optimization variable parameters may include the face shape coefficient , the face expression coefficient and the lighting parameter and so on. Among them, the above parameter initialization is mainly used to initialize the above face shape coefficient , the face expression coefficient and the lighting parameter to zero vectors.
[0108] Therefore, the initialization result of the above optimization variable parameters is the face shape coefficient of the zero vector , the face expression coefficient and the lighting parameter . According to the face shape coefficient of the above zero vector , the face expression coefficient and the lighting parameter , the initial three-dimensional face is constructed based on the standardized average face according to the default lighting condition. Among them, the default lighting condition can be used as the default lighting data for the standardized average face during the real-time face alignment process, which can ensure the detection and alignment accuracy of the generated initial three-dimensional face and exclude the influence of the change of lighting conditions. The default lighting condition can be the lighting information of a point light source that is consistent with the camera center. When obtaining the above target expression image, it is generated according to the lighting information of the camera or the lighting condition is estimated according to the target expression image.
[0109] Therefore, it can be ensured that the facial expression of the target expression image can be normalized to the standardized average face, thereby eliminating the coupling problem between the expression semantics and the individual face shape, and ensuring the consistency of the expression semantics among individuals with different face shapes.
[0110] As Figures 2A - 5C shown, according to an embodiment of the present invention, in operation S203, joint optimization processing is performed on the dense face reconstruction task according to the initial three-dimensional face to generate three-dimensional face reconstruction information, including: generating the three-dimensional face reconstruction information after convergence by performing iterative optimization on the projection error, photometric consistency error, and regularization constraint in the dense face reconstruction task.
[0111] The joint optimization process of the dense face reconstruction task can define the objective function of the projection error in formula 1 of the above joint optimization problem according to iterative optimization means (such as the Levenberg-Marquardt algorithm or the gradient descent scheme) , the objective function of the photometric consistency error and the regularization constraint information (such as the face shape coefficient constraint term , the face expression coefficient constraint term , the additional regularization term , etc.) for iterative optimization to minimize, so that the final optimization process can converge to satisfy the following three loss conditions in turn: (1) Key point alignment based on the objective function of the projection error, such as the alignment of the standardized average face and the two-dimensional detected face, to ensure the minimum difference in face key points between the two; (2) Photometric consistency optimization based on the objective function of the photometric consistency error, such as for each pixel point, the lighting direction parameters of the standardized average face and the two-dimensional detected face can conform to a lighting condition model; (3) The regularization term controls the stability and physical rationality of the optimization, such as minimizing the face shape coefficient and the face expression coefficient .
[0112] This can achieve the joint optimization of the above dense face reconstruction task, thereby ensuring that the three-dimensional face reconstruction of the standardized average face can be performed on the input target expression image through the three-dimensional deformation model.
[0113] The target three-dimensional face can be topologically constructed based on the three-dimensional deformation model. The three-dimensional deformation model can represent any three-dimensional face as a combination of linear bases by learning a large amount of three-dimensional face data. Specifically, the three-dimensional deformation model can be expressed as the following formula 4:
[0114] (4)
[0115] Among them, the three-dimensional face shape of the target three-dimensional face can be expressed as , is the number of mesh vertices of the three-dimensional face topology structure, is the real number space. Further, can be used to express the standardized average face (i.e., the three-dimensional average face, Mean Shape), representing the average shape of the face obtained by statistical learning. represents the face shape basis (Identity Basis), which can be used to represent the principal component basis of individual face differences, and s is the number of shape bases. represents the face expression basis (Expression Basis), which can be used to describe the dynamic expression changes of the face, is the number of expression bases. represents the face shape coefficients (Identity Coefficients), which can represent the facial features of a specific individual. represents the face expression coefficients (Expression Coefficients), which can be used as a parametric representation of the face expression.
[0116] Through the above joint optimization process for the dense face reconstruction task, the face expression basis and the face expression coefficients of the three-dimensional face reconstruction information can be generated.
[0117] For the face expression basis and the face shape basis of the three-dimensional face reconstruction information, a linear combination representation is performed based on Equation 4 of the above three-dimensional deformation model, so as to realize the dense expression motion representation of any face shape and form the target three-dimensional face.
[0118] Therefore, through the construction process of the target three-dimensional face, the three-dimensional face reconstruction of the input target expression image can be carried out by the three-dimensional deformation model, realizing the decomposition of the three subspaces of the pose, face shape, and expression of the face expression motion, corresponding respectively to the pose subspace for representing the global rotation and translation of the face expression, the face shape subspace for representing the individual facial features, and the expression subspace for representing the dynamic expression changes of the face, thus realizing the mapping of the face expression of a single-frame two-dimensional expression image into a parametric three-dimensional representation. With the help of the decomposition ability of the face shape, expression, and pose subspaces, the effective decoupling of the individual facial features and expression dynamics is realized, ensuring the significant improvement of the capture ability of subtle expressions through dense expression representation, thereby overcoming the limitations of the traditional sparse feature key point scheme.
[0119] Such as Figures 2A - 5CAs shown, according to an embodiment of the present invention, in operation S204, extracting the dense expression motion information of the standardized average face from the target three-dimensional face to complete perception optimization includes: optimizing the shape parameters of the target three-dimensional face to complete the extraction of the dense expression motion information.
[0120] The three-dimensional face reconstruction of a single-frame target expression image can be achieved based on methods such as global photometric consistency optimization and end-to-end reconstruction schemes based on deep learning (such as using CNN and Transformer models). Specifically, as described above for the face expression basis and face shape basis based on the three-dimensional deformation model, it is possible to restore the dense and high-precision three-dimensional face shape and its dynamic expression parameters, thereby achieving accurate modeling of the face geometry and expression motion.
[0121] According to the above-mentioned face expression basis for the dense face reconstruction task of formula 1 and the face expression coefficients Based on the joint optimization results, the dense expression motion information can be expressed according to formula 4 of the above three-dimensional deformation model, specifically as formula 5 below:
[0122] (5)
[0123] Furthermore, the three-dimensional face shape of the above target three-dimensional face can be further expressed as formula 6 below:
[0124] (6)
[0125] Therefore, after completing the reconstruction of the three-dimensional face shape of the target three-dimensional face, the generated three-dimensional face shape can include the individualized face shape and the face expression motion information These two parts can be specifically expressed as formula 4 above.
[0126] To eliminate the influence of the individualized face shape features and ensure the consistency of the expression semantics, the face shape coefficients can be set to zero, only retaining the part related to the expression, that is, completing the optimization of the shape parameters of the target three-dimensional face, generating a three-dimensional face shape based on the standardized average face corresponding to formula 4 above , this three-dimensional face shape can be used to represent the dense expression grid topology structure, define the vertex positions of the dense expression grid, and its specific expression is as formula 7 below:
[0127] (7)
[0128] Where, is the face shape of the standardized average face; describes dense expression motion information.
[0129] Therefore, the generated 3D face shape is the 3D expression shape under the 3D standardized average face which has eliminated individual differences and only retains dynamic expression features, thus providing a basis representation with consistent semantic mapping for subsequent expression mapping. Among them, the dense expression motion information is represented by and can be expressed as the following formula 8 as the result of dense expression extraction:
[0130] (8)
[0131] Among them, the dense expression motion information is a dense 3D vertex motion field, which can be used to represent the expression offset of each vertex. The significance of extracting dense expression motion information is that it provides a fine representation at the vertex level for the entire expression motion, can reflect extremely subtle expression changes, and at the same time eliminates the interference of individual shapes, providing high-precision input for subsequent expression operations.
[0132] Therefore, in the above-mentioned standardization process and dense expression extraction process, in order to eliminate the influence of individual face shapes on expression description, the shape parameters are fixed to zero in the post-processing stage, and only the expression and pose parameters are retained, thus generating a dense expression motion under the standardized average face. This process ensures the unity of expression semantics. Therefore, a high-resolution dense expression motion information can be generated through the linear combination of the face expression basis and the face expression coefficients of the 3D deformation model, thus providing an accurate and standardized expression description for the subsequent mapping and execution stages.
[0133] Furthermore, in the perception stage, the face motion is decomposed into three independent subspaces of pose, face shape, and expression through the 3D deformation model, generating a standardized and dense expression motion representation, which not only improves the refinement of expression description, but also eliminates the coupling problem between expression semantics and individual face shapes through the standardized average face topology, realizing the consistency of expression semantics among different individuals.
[0134] In summary, based on the above-mentioned perception optimization method for agent expression imitation according to the embodiments of the present invention, in the perception stage of expression imitation, a three-dimensional deformation model can be used as the core technical means to replace the traditional sparse key-point based expression representation method. By decomposing the movement of the human face into three independent subspaces of pose, face shape, and expression, a refined description of human face expressions is achieved. In specific implementation, the above-mentioned three-dimensional human face reconstruction information is further calculated for a single-frame image, and then in the post-processing stage, the shape parameters are fixed, and only the expression and pose parameters are retained, so as to generate dense expression motion information under the standardized average face topology. Therefore, it can significantly overcome the abstract defects of the traditional sparse key-point method, realize the densification of expression description, and at the same time, by normalizing the expression to the standardized average face, the coupling problem between expression semantics and individual face shapes is eliminated, ensuring the consistency of expression semantics among individuals with different face shapes.
[0135] Compared with the traditional coefficient key-point method, the high precision and density of the dense expression motion representation of the above-mentioned three-dimensional deformation model according to the embodiments of the present invention can significantly improve the accuracy and fineness of expression perception, laying a high-quality foundation for the subsequent mapping and execution stages. Specifically, in the perception stage of expression, by using a three-dimensional deformable model, the human face expression in a two-dimensional video can be mapped into a parameterized three-dimensional representation. With its decomposition ability for the face shape, expression, and pose subspaces, the effective decoupling of individual face shape characteristics and expression dynamics is realized. This method significantly improves the ability to capture subtle expressions through dense expression representation, overcoming the limitations of the traditional sparse feature point method.
[0136] Based on the above-mentioned perception optimization method for agent expression imitation, the present invention also provides a perception optimization device for agent expression imitation. The following will be combined with Figure 2B to describe this device in detail.
[0137] Figure 2B The structural block diagram of the perception optimization device for agent expression imitation according to the embodiments of the present invention is schematically shown.
[0138] As Figure 2B shown, the perception optimization device 200 for agent expression imitation in this embodiment includes a motion extraction module 210, a task generation module 220, a joint optimization module 230, and an information extraction module 240.
[0139] The motion extraction module 210 is used to extract the expression motion information of the target expression image. In one embodiment, the motion extraction module 210 can be used to execute the operation S201 described above, which will not be elaborated here.
[0140] The task generation module 220 is used to generate a dense face reconstruction task based on the expression motion information and the preset joint optimization information. In one embodiment, the task generation module 220 can be used to perform the operation S202 described above, which will not be elaborated here.
[0141] The joint optimization module 230 is used to perform joint optimization on the dense face reconstruction task to generate three-dimensional face reconstruction information, and the three-dimensional face reconstruction information is used to generate the target three-dimensional face. In one embodiment, the joint optimization module 230 can be used to perform the operation S203 described above, which will not be elaborated here.
[0142] The information extraction module 240 is used to extract the dense expression motion information of the standardized average face according to the target three-dimensional face to complete the perception optimization. In one embodiment, the information extraction module 240 can be used to perform the operation S204 described above, which will not be elaborated here.
[0143] According to an embodiment of the present invention, any multiple of the motion extraction module 210, the task generation module 220, the joint optimization module 230, and the information extraction module 240 can be combined and implemented in one module, or any one of them can be split into multiple modules. Or, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the motion extraction module 210, the task generation module 220, the joint optimization module 230, and the information extraction module 240 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable way of integrating or packaging the circuit and other hardware or firmware, or implemented in any one of the three implementation ways of software, hardware, and firmware or in any appropriate combination of several of them. Or, at least one of the motion extraction module 210, the task generation module 220, the joint optimization module 230, and the information extraction module 240 can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions can be executed.
[0144] Embodiment 2
[0145] In the expression mapping stage, existing methods mainly rely on two technical paths: geometric displacement migration based on sparse key points and expression semantic mapping based on image generation. Among them, the method of geometric displacement migration using sparse key points attempts to directly map the key point displacements of the human face to the robot face. However, due to the significant differences in topological structure and geometric characteristics between the human face and the robot face, simple displacement copying or regularization methods based on rules such as linear interpolation and affine transformation cannot effectively transfer expression semantics, and may even cause expression features to be mis-mapped or lost, resulting in serious semantic distortion. In some cases, this method is completely unable to transfer specific expressions. Another image generation method based on a generative model can visually generate images that conform to expression semantics, but these images only achieve expression consistency in the pixel space and do not consider the physical constraints of the robot actuators, so they cannot be directly used to drive actual physical actuators.
[0146] Therefore, in view of at least one of the technical problems existing in the above-mentioned prior art, embodiments of the present invention aim to be able to provide a mapping optimization method, device, equipment, medium and product for agent expression imitation, thereby providing an optimization solution for the mapping stage of agent expression imitation, with the expectation of achieving a cross-modal mapping mechanism from standardized human face expressions to the robot expression space by introducing a unified expression parameter space, solving the differences in geometric structures between the human face and the robot face, and ensuring the consistency of expression semantics during the migration process.
[0147] The following will be based on Figure 1 the described scenario, and through Figures 3A - 5C a detailed description of the mapping optimization method for agent expression imitation of the disclosed embodiments will be given.
[0148] As Figure 3A shown, one aspect of the embodiments of the present invention provides a mapping optimization method for agent expression imitation, which includes operations S301 to S303.
[0149] In operation S301, a first blend shape basis of a standardized average face is generated based on a preset facial action coding rule;
[0150] In operation S302, the dense expression motion information of the standardized average face is mapped to the first blend shape basis to generate a second blend shape basis; and
[0151] In operation S303, the blend shape coefficients of the second blend shape basis are transferred to a third blend shape basis of the agent face to complete the mapping optimization.
[0152] The preset facial action coding rule can be a coding rule for facial behaviors formed based on the facial muscle movement state, so as to realize the coding of facial expressions, and can improve the processing accuracy and efficiency of facial expressions in the process of expression recognition and processing. For example, it can be implemented by using the Facial Action Coding System (FACS for short).
[0153] The standardized average face is the three-dimensional average face shape of the static face under the standardized human face topology, which can usually be obtained based on big data statistics. Specifically, reference can be made to the standardized average face in Embodiment 1. 。
[0154] Based on the preset facial action coding rule and on the basis of the human face topology structure of the standardized average face, semantic processing of the expression shape offset is carried out to form the first blend shape basis. Among them, the first blend shape basis can describe the vertex offset of the corresponding semantic expression (such as "open mouth" or "blink left eye", etc.) relative to the standardized average face, and can be used to describe the expression change through the weighted combination of predefined semantic expression bases. Specifically, it can be implemented through the Blendshape expression representation technology.
[0155] The dense expression motion information can be information directly related to the extracted expression motion information and related to the dynamic change of the human face expression. Specifically, it can be characterized by the linear combination of the human face expression basis and the human face expression coefficient. The dense expression motion information can most directly reflect the dense expression motion information of the user extracted by the agent on the standardized average face, and realize the representation of the human face expression. Specifically, reference can be made to the extraction of the dense expression motion information in Embodiment 1.
[0156] Mapping the dense expression motion information to the first blend shape basis can realize mapping the expression motion information of the human face in each frame of the target expression image extracted in the perception stage in Embodiment 1 to the standardized average face. At the same time, the second blend shape basis generated thereby can also exclude the influence of individual human face shape features, ensure the expression semantic consistency in the mapping stage and the execution stage, and thus complete the dense representation of the expression semantics. The second blend shape basis can be a semantic expression basis expressed by the dense expression motion information obtained through perception, and can be used to describe the vertex offset of the semantic expression corresponding to the dense expression motion information relative to the standardized average face. Specifically, it can be implemented through the Blendshape expression representation technology.
[0157] Accordingly, the third blend shape base is a semantic expression base that performs semantic processing of expression shape offsets based on the agent's facial topology, and can be used to describe the vertex offsets of the corresponding semantic expressions relative to the agent's facial topology. Among them, the semantics of the third blend shape base correspond one-to-one and are consistent with the semantics of the above-mentioned second blend shape base.
[0158] The blend shape coefficient can be understood as the blend shape weight of the second blend shape base (such as the Blendshape coefficient), and can be used to express the vertex offsets of the expression actions of the dense expression motion information obtained in the perception stage (refer to the above-mentioned Embodiment 1). By transferring this blend shape coefficient to the third blend shape base of the agent's face, direct cross-topology migration of human dense expression motion information can be achieved.
[0159] Therefore, in the mapping stage, the above-mentioned mapping optimization method for agent expression imitation in the embodiments of the present invention solves the problems of differences in topology and geometric characteristics between the human face and the agent's face through a blend shape semantic space based on the second blend shape base of the standardized average face as an intermediate bridge. Specifically, first, a set of first blend shape bases is defined on the standardized average face for dense representation of expression semantics. The first blend shape base can then map the dense expression motion generated in the perception stage (refer to the above-mentioned Embodiment 1) to the blend shape coefficient space through dense semantic key points, forming the second blend shape base. In this process, by transferring the blend shape coefficient, the expression motion is mapped from the topology of the standardized average face to the topology of the agent's face, realizing cross-topology semantic-consistent expression migration.
[0160] Therefore, through the conversion from the geometric domain to the semantic domain, the problem of semantic distortion caused by the topological differences between the human face and the robot's face in the traditional solution is solved. In addition, the number of blend shape coefficients is limited, which is convenient for real-time transmission, thus greatly improving the efficiency and robustness of the mapping stage. At the same time, the generated blend shape coefficients directly correspond to the motion state of the robot's epidermis (such as Figure 5C the shown skin Skin), ensuring the physical executability of the generated expression. It can be seen that in the cross-entity mapping process of expression imitation, the blend shape base designed based on the above-mentioned preset facial action coding rules can be used to accurately depict expression semantics. By introducing a unified expression parameter space, a cross-modal mapping mechanism from the dense human face expressions of the standardized average face to the robot expression space can be established, thus solving the differences in geometric structures between the human face and the robot's face and ensuring the consistency of expression semantics during the migration process.
[0161] The mixed - shape base is generated mainly based on semantic expression modeling using Blendshape technology. After the normalization process and the extraction of dense expression motion information in the above - mentioned Embodiment 1, in order to further realize semantic expression modeling, an expression representation method based on this mixed - shape base is further introduced.
[0162] The Blendshape technology can be understood as describing expression changes through the weighted combination of a set of predefined semantic expression bases. Therefore, the standardized average face corresponding three - dimensional human face shape can be expressed as the following formula (9):
[0163] (9)
[0164] where can be expressed as a dense grid containing semantic expressions. can be the three - dimensional average face shape, which is consistent with the definition of the standardized average face in Embodiment 1, representing the static face shape under the standardized human face topology. can be the n - th Human Blendshape base, used to represent the semantic expression shape offset, and can describe the vertex offset of the corresponding semantic expression (such as "open mouth" or "blink left eye") relative to the standardized average face. can be the weight of the n - th Human Blendshape base, which can represent the intensity coefficient of this expression. Among them, the weight of this Human Blendshape base can usually be normalized to the interval For example, means no such expression, means fully applying this expression. is the number of defined Blendshape bases, and each Blendshape base can correspond to a semantic expression action.
[0165] Combined with the mathematical representation of the Blendshape base in the above formula (9), it can be combined with Figures 3A - 5C to further describe the generation of the first Blendshape base in Embodiment 2 of the present invention as follows.
[0166] As Figures 3A - 5C shown, according to an embodiment of the present invention, in operation 301 for generating the first Blendshape base of the standardized average face based on a preset facial action coding rule, it includes:
[0167] Generating a Blendshape expression base matrix according to the preset facial action coding rule;
[0168] Generate the first blend shape basis for the standardized average face based on a preset blend shape weight vector and a blend shape expression basis matrix.
[0169] As previously mentioned, the first blend shape basis can be designed based on the three-dimensional average face shape (i.e., the standardized average face ), which can have the same topological structure (connection relationship between vertices and faces) as the standardized average face, and is defined in combination with the expression semantics of a preset facial action coding rule (such as the Facial Action Coding System, abbreviated as FACS), for example, "open mouth" or "blink left eye".
[0170] As defined in the above formula 9, the human face blend shape basis that is part of the first blend shape basis can be designed in a predefined manner, specifically based on the facial topological structure of the standardized average face. Among them, the first blend shape basis is designed according to the expression semantics of the preset facial action coding rule, and can be divided according to the dominant area to ensure that the semantics of each blend shape basis are clear and consistent. For example, the expression action of "right corner of the mouth rising" can correspond to the movement of an actuator in FACS.
[0171] Specifically, the three-dimensional human face deformation represented by the first blend shape basis can be defined based on the standardized average face and generated through a weighted combination of a set of predefined semantic expression bases. Combining the above formula 9, specifically, the first blend shape basis can be expressed in matrix form in the following formula 10 to form an expression combination based on the blend shape basis:
[0172] (10)
[0173] Among them, the blend shape expression basis matrix can form the semantic human face blend shape basis matrix of the first blend shape basis, specifically composed of predefined human face expression bases . In addition, the blend shape weight set can be correspondingly expressed as the preset weight vector of the blend shape corresponding to the blend shape expression basis matrix to represent the intensity of each expression. Therefore, through this preset blend shape weight vector and the blend shape expression basis matrix can form the first blend shape basis corresponding to the standardized average face .
[0174] Therefore, by adjusting the blend shape weights , that is, it can flexibly control the contribution corresponding to each component blend shape basis in the first blend shape basis, so as to generate expression shapes with different intensities and combinations. This predefined modeling method can ensure that the generated expression shapes have clear semantics, dense geometric details, and a topology consistent with the average face.
[0175] It can be seen that through the semantic expression modeling of the blend shape basis, on the facial topology (Mesh) of the standardized average face, a set of blend shape bases corresponding to its dense expression movements can be designed based on the Facial Action Coding System for the compact representation of expression semantics (such as 51 standard Blendshapes of ARKit).
[0176] Such as Figures 3A - 5C As shown, according to an embodiment of the present invention, before mapping the dense expression movement information of the standardized average face to the first blend shape basis to generate the second blend shape basis in operation 302, it further includes:
[0177] Generating the dense expression movement information of the standardized average face according to the expression movement information of the extracted target expression image.
[0178] As mentioned in Embodiment 1, the target expression image can be a single-frame image containing a human facial expression, and the expression movement information can include the facial expression feature parameters extracted for the facial expression in the target expression image. Correspondingly, the dense expression movement information can be information related to the dynamic change of the facial expression directly related to the extracted expression movement information, and can be specifically characterized by the linear combination of the facial expression basis and the facial expression coefficient. See Formula 5 in Embodiment 1 for details. The dense expression movement information can most directly reflect the dense expression movement information of the user extracted by the agent on the standardized average face, realizing the representation of the facial expression.
[0179] Therefore, based on the extraction of the dense expression movement, a facial expression mesh on the basis of the standardized average face corresponding to the dense expression movement information can be constructed (as shown in Formula 7 in Embodiment 1), and this facial expression mesh can be used as a three-dimensional human face shape with dense expression movement information to represent the dense expression mesh topology and define the vertex positions of the dense expression mesh.
[0180] By means of the above process of extracting the dense expression movement information in the perception stage, the effective decoupling of the individual face shape characteristics and the expression dynamic information can be realized, and the ability to capture subtle expressions can be significantly improved through dense expression representation.
[0181] In the mapping stage of expression imitation in Embodiment 2 of the present invention, in order to transfer the dense expression movements generated in the perception stage of Embodiment 1 to the topological structure of the agent's face, achieve cross-topological semantic transfer, and ensure cross-topological semantic consistency, it is necessary to consider the mapping from dense expression movements to blend shape coefficients.
[0182] As Figures 3A - 5C shown, according to an embodiment of the present invention, in operation 302 of mapping the dense expression movement information of the standardized average face to the first blend shape basis to generate the second blend shape basis, it includes:
[0183] Extract the first dense annotation points on the dense expression grid corresponding to the dense expression movement information and the second dense annotation points on the average face grid;
[0184] Obtain the dense annotation basis matrix according to the blend shape expression basis matrix;
[0185] Generate a blend shape weight optimization task according to the dense annotation basis matrix, the first dense annotation points, and the second dense annotation points.
[0186] The dense expression grid can be understood as the human face expression grid based on the standardized average face corresponding to the above dense expression movement information (as shown in Equation 7 in Embodiment 1), which is used to represent the vertex positions of the standardized dense expression grid, including vertices, and each vertex has three-dimensional coordinates (a total of degrees of freedom), and can be understood as the three-dimensional geometric representation of the entire facial network.
[0187] The average face grid can be understood as the standardized human face grid of the standardized average face and can represent the three-dimensional vertex positions of the corresponding human face grid.
[0188] The first dense annotation points can be represented as the positions of the standardized dense annotation points , which are usually a specific set of points selected from the above dense expression grid in dense expression modeling, such as distributed in key regions of the face (such as the mouth, eyes, and nose, etc.), and specifically can be represented as the positions of dense annotation points selected from the dense expression grid (a total of degrees of freedom).
[0189] Among them, the positions of the first dense annotation points can be extracted from the dense expression grid through the dense annotation matrix , and specifically can be expressed as the following Equation 11:
[0190] (11)
[0191] Among them, the dense annotation matrix is a sparse selection matrix, which can be used to extract the first dense annotation points of GIA from the dense expression grid .
[0192] Correspondingly, the second dense annotation point can be expressed as the average face position of the dense annotation points , defined as the dense annotation points extracted from the average face grid through the dense annotation matrix , and can be specifically expressed as formula 12 below:
[0193] (12)
[0194] Among them, the average face grid can represent the three-dimensional vertex positions of the face mesh (Mesh) defined by the normalized average face.
[0195] Therefore, with the help of the above formula (11) and formula (12), the correspondence between the dense annotation points and the face mesh can be constructed. Among them, the position of the first dense annotation point and the position of the second dense annotation point are directly related to the dense expression grid (full grid expression) and the average face grid .
[0196] As mentioned above, the mixed shape expression basis matrix can be used as the full grid basis matrix to form the semantic face mixed shape basis matrix of the first mixed shape basis, which can be specifically composed of K predefined expression bases.
[0197] The dense annotation basis matrix can be understood as the mixed shape basis matrix of the dense annotation , which is mainly used to represent the mixed shape basis on the dense annotation points, and can be specifically extracted from the full grid basis matrix , and can be specifically expressed as formula 13 below:
[0198] (13)
[0199] Therefore, according to the dense annotation basis matrix defined by the above formula 13, the first dense annotation point defined by formula 11, and the second dense annotation point defined by formula 12, the mixed shape weight corresponding to the above mixed shape weight optimization task can be defined as the optimization problem represented by formula 14 as follows:
[0200] (14)
[0201] Among them, it can be understood as a mixed shape weight, which can make the position of the reconstructed dense annotation points satisfy the relationship defined by the following formula (15):
[0202] (15)
[0203] Therefore, modeling based on the mixed shape technology can utilize the geometric information of the dense annotation points to solve the optimization problem corresponding to the mixed shape weight optimization task of the above formula (14) and obtain the mixed shape weight , so as to ensure that the position of the reconstructed dense annotation points is as close as possible to the input dense annotation points.
[0204] As Figures 3A - 5C shown, according to an embodiment of the present invention, in the operation of mapping the dense expression motion information of the standardized average face to the first mixed shape basis to generate the second mixed shape basis in operation 302, it further includes:
[0205] Obtaining the mixed shape coefficients corresponding to the mixed shape weight optimization task by minimizing the preset error vector and the preset coefficient constraint condition, so as to map the dense expression motion information of the standardized average face to the first mixed shape basis to generate the second mixed shape basis.
[0206] The preset error vector can be the error vector objective function for performing optimization processing on the above mixed shape weight optimization task , and specifically can be expressed as the following formula (16):
[0207] (16)
[0208] The preset error vector can be the error vector of the first dense annotation point and the second dense annotation point based on the dense annotation basis matrix . By setting the sum of the squares of the minimum errors of the preset error vector as the optimization objective, the following formula (17) can be defined:
[0209] (17)
[0210] Expanding the above formula (17), it can be expressed as the following formula (18):
[0211] (18)
[0212] The preset coefficient constraint condition can be used as the constraint condition for performing optimization processing on the above mixed shape weight optimization task. For example, the value of the mixed shape weight can be limited to within the range, so that the blend shape coefficient corresponding to the final blend shape weight optimization task is . Therefore, the above blend shape weight optimization task corresponds to a quadratic programming problem with boundary constraints (preset coefficient constraint conditions).
[0213] Therefore, by means of the solution process of the least squares problem of the above blend shape coefficient, dense expression motion information can be mapped to a low-dimensional first blend shape basis in real time to form a second blend shape basis, that is, by selecting dense semantic key points (based on barycentric coordinates, rather than vertex coordinates) on the surface of the average face mesh, the dense expression motion generated in the perception stage is mapped to the blend shape coefficient space, thus completing the coefficient mapping from the dense expression to the blend shape space.
[0214] In summary, through the least squares problem with an overdetermined constraint of the above blend shape weight optimization task, the dense expression motion information can be compressed into a low-dimensional blend shape space. This process can generate blend shape coefficients in real time while ensuring the density and accuracy of the expression.
[0215] Such as Figures 3A - 5C shown, according to an embodiment of the present invention, before transferring the blend shape coefficients of the second blend shape basis to the third blend shape basis of the agent's face in operation 303 to complete the mapping optimization, it further includes:
[0216] Generating a third blend shape basis of the agent's face based on a preset facial action encoding rule, where the third blend shape basis has the same blend shape semantics as the first blend shape basis.
[0217] The third blend shape basis is actually built for the blend shape space of the agent's face. Among them, the agent's blend shape space needs to ensure semantic consistency with the first blend shape basis space corresponding to the dense expression grid of the human face so that the agent can accurately express the same basic expression semantics as humans (such as "opening the mouth", "blinking", etc.).
[0218] Similar to the human first blend shape basis, according to the semantic consistency principle, the third blend shape basis space of each agent corresponds one-to-one and is consistent with the semantic of the first blend shape basis space of the standardized average face (for example, "opening the mouth" should express the same semantics on the facial topology of the agent and in the blend basis space of the human standardized average face model).
[0219] Therefore, according to the semantic set of the first blend shape basis space of the human face (such as "open mouth", "blink left eye", etc.). Based on the above preset facial action coding rules, a corresponding set of third blend shape bases can also be defined for the agent based on the agent's facial topology, and the semantic set in the space of the third blend shape bases is , and the agent blend shape bases that conform to each semantics , where is the number of mesh vertices in the agent's facial topology, can be the semantic number.
[0220] Therefore, for the topology of the agent's face, based on the first blend shape bases of the standardized average face corresponding to the dense expression motion information, the third blend shape bases of the agent's face with consistent semantics can be constructed to ensure that the semantics of the two blend shape base spaces are completely consistent, thus ensuring efficient and accurate dense expression transfer from the standardized average face to the agent's facial topology.
[0221] For example Figures 3A - 5C As shown, according to an embodiment of the present invention, in operation 303, transferring the blend shape coefficients of the second blend shape bases to the third blend shape bases of the agent's face to complete the mapping optimization includes:
[0222] Applying the blend shape coefficients to the third blend shape bases to generate the agent's corresponding agent expression mesh.
[0223] Directly transferring the blend shape coefficients generated from the standardized average face to the third blend shape bases corresponding to the agent's facial topology can achieve dense expression cross-topology semantic transfer, generate the agent expression mesh, and thus complete the mapping process of the blend shape coefficients.
[0224] Among them, the core of transferring the blend shape coefficients is to directly apply the blend shape coefficients generated from the human standardized average face to the third blend shape bases of the agent to achieve dense expression cross-topology semantic transfer.
[0225] Specifically, through the third blend shape bases of the agent's facial topology with semantic consistency , the semantics of each basic expression are completely aligned with the first blend shape bases of the human standardized average face topology. Therefore, the blend shape coefficients of the first blend shape bases can be directly transferred to the third blend shape bases without any conversion, satisfying the following formula 19:
[0226] (19)
[0227] Therefore, transferring the blend shape coefficients of the first blend shape bases Applied to the agent's facial topology and the agent's third blend shape basis , an agent expression mesh corresponding to the agent can be generated Expressed as formula 20 below:
[0228] (20)
[0229] Therefore, it can ensure that the semantics of the robot's expression are consistent with those of human expressions, while adapting to the geometric characteristics of the agent's facial topology. Specifically, refer to the semantic alignment operation from the second blend shape basis to the third blend shape basis as shown Figure 5A in the semantic alignment operation from the second blend shape basis to the third blend shape basis as shown
[0230] In summary, based on the mapping optimization method for agent expression imitation provided in Embodiment 2 of the present invention, a blend shape semantic space based on a standardized average face can be provided as an intermediate bridge in the mapping stage to solve the differences in topological structure and geometric characteristics between the human face and the agent's face. Specifically, first, a set of blend shape bases are defined on the standardized average face for dense representation of expression semantics. These blend shape bases map the dense expression movements generated in the perception stage to the blend shape coefficient space through dense semantic key points. This process maps the dense expression movements to low-dimensional blend shape coefficients in real time by solving a least squares problem
[0231] After that, a set of blend shape bases consistent with the average face semantics are provided based on the agent's face to ensure that they have the same semantic definition. By transferring the blend shape coefficients, the dense expression movements can be mapped from the standardized average face to the agent's face, thus achieving efficient and accurate expression transfer across topological structures with consistent semantics
[0232] Therefore, by converting from the geometric domain to the semantic domain, the problem of semantic distortion caused by the topological differences between the human face and the agent's face in traditional methods is solved. In addition, the number of blend shape coefficients is limited, which is convenient for real-time transmission, greatly improving the efficiency and robustness of the mapping stage. At the same time, the generated blend shape coefficients directly correspond to the motion state of the agent's epidermis, ensuring the physical executability of the generated expressions
[0233] In the cross-entity mapping of expressions, the blend shape bases designed based on the preset facial action coding rules are used to accurately depict expression semantics. By introducing a unified expression parameter space, a cross-modal mapping mechanism from the standardized human face expression to the agent expression space is established, solving the differences in geometric structures between the human face and the agent's face and ensuring the consistency of expression semantics during the migration process
[0234] Therefore, the mapping optimization method for agent expression imitation provided in Embodiment 2 of the present invention can at least achieve the following technical effects:
[0235] (1) Directness: No additional conversion or mapping is required for the human blend shape coefficients.
[0236] (2) Efficiency: Utilize the blend shape bases with semantic consistency to quickly achieve cross-topology migration of expressions.
[0237] (3) Versatility: Applicable to any agent blend shape bases that conform to the semantic consistency design.
[0238] Based on the above mapping optimization method for agent expression imitation, the present invention also provides a mapping optimization device for agent expression imitation. The following will be combined with Figure 3B to describe this device in detail.
[0239] Figure 3B The structural block diagram of the mapping optimization device for agent expression imitation according to an embodiment of the present invention is schematically shown.
[0240] As Figure 3B shown, the mapping optimization device 300 for agent expression imitation in this embodiment includes a base generation module 310, an information mapping module 320, and a coefficient transfer module 330.
[0241] The base generation module 310 is used to generate a first blend shape base of a standardized average face based on a preset facial action coding rule. In one embodiment, the base generation module 310 can be used to perform the operation S301 described above, which will not be elaborated here.
[0242] The information mapping module 320 is used to map the dense expression motion information of the standardized average face to the first blend shape base to generate a second blend shape base. In one embodiment, the information mapping module 320 can be used to perform the operation S302 described above, which will not be elaborated here.
[0243] The coefficient transfer module 330 is used to transfer the blend shape coefficients of the second blend shape base to a third blend shape base of the agent's face to complete the mapping optimization. In one embodiment, the coefficient transfer module 330 can be used to perform the operation S303 described above, which will not be elaborated here.
[0244] According to an embodiment of the present invention, any plurality of modules among the substrate generation module 310, the information mapping module 320, and the coefficient transfer module 330 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the substrate generation module 310, the information mapping module 320, and the coefficient transfer module 330 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the substrate generation module 310, the information mapping module 320, and the coefficient transfer module 330 may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions may be executed.
[0245] Embodiment 3
[0246] In the execution stage, the traditional method converts the epidermal motion representation into actuator instructions in two ways. One is the model-based inverse kinematics method, which converts the motion of the epidermal traction point into specific instructions of the actuator through inverse kinematics solution. However, this method has extremely high requirements for the accuracy of the geometric model and needs to accurately describe the topological structure of the robot face and the actuator position, otherwise it is difficult to ensure the accuracy of the inverse kinematics solution. The other method is the model-free neural networks method based on deep learning, which directly calculates the actuator instructions from the epidermal motion (such as pixel space) using a neural network. Although this method has higher flexibility, the generated execution instructions often do not fully consider the physical constraints of the robot actuator, resulting in the difficulty of implementing the instructions in actual applications and causing problems in the control accuracy and stability of the expression.
[0247] In view of at least one of the technical problems existing in the above-mentioned execution schemes for expression imitation, embodiments of the present invention provide an execution optimization method, device, equipment, medium, and product for agent expression imitation, in order to optimize the expression driving parameters by explicitly modeling the physical constraints of the agent actuator and ensure that the generated control instructions are both physically executable and meet the requirements of high precision and real-time performance.
[0248] The following will be based on Figure 1 the described scenario, through Figures 4A - 5CA detailed description is given of the method for optimizing the execution of agent expression imitation in the disclosed embodiments.
[0249] As Figure 4A shown, one aspect of an embodiment of the present invention provides a method for optimizing the execution of agent expression imitation, which includes operations S401 to S403.
[0250] In operation S401, an inverse kinematics target optimization task is generated based on the control point simulation information of the agent expression grid and the target epidermal vertex information;
[0251] In operation S402, optimization processing is performed on the inverse kinematics target optimization task to obtain the optimal joint angle parameters corresponding to the control point simulation information; and
[0252] In operation S403, the execution of agent expression imitation is optimized according to the optimal joint angle parameters.
[0253] The agent expression grid may include the vertex relationship of the facial topology grid (Mesh) structure of the agent mentioned in the above embodiment 2, where the agent expression grid is the robot facial topology that has completed the cross-topology migration of dense expression motion information in the perception stage in embodiment 1, and specifically may be embodied as the expression grid in formula 20 in embodiment 2 .
[0254] As Figure 5B shown, there are many key control points on the agent expression grid, such as eyebrows, eyelids, eyeballs, nose, cheeks, mouth, and jaw, etc. These control points can be driven by cables / PTFE conduits as Figure 5C shown through motors (such as LFD-01M servo motors) to achieve traction, thereby driving the deformation of the epidermis in the surrounding area of the control points (such as Figure 5C shown Silicone Mask). In the traditional technical solution, it is very difficult to solve the deformation of the surrounding area of the control points, resulting in a rigid effect of agent expression imitation and unable to achieve a delicate and natural "skin" performance.
[0255] The control point simulation information can be the position information formed by performing control point position simulation calculations on the agent expression grid containing dense expression motion information through a simulation tool (such as Blender), for example, the simulation positions of each control point in the topological network of the agent expression grid with dense expression motion information. The control point simulation information can be used to define the drive angle parameters of the drive arm of the drive motor of the cable / PTFE conduit as Figure 5C shown, that is, the joint angle parameters.
[0256] The target epidermal vertex information can be the target epidermal mesh vertex positions that each control point on the agent's expression mesh should reach when achieving the facial expression reflected by the dense expression motion information. The target epidermal vertex information can include the simulation control point positions of the control point simulation information calculated by simulation and the relative displacement distances and displacement direction information of the epidermal mesh vertices in its surrounding area, and can be used to define the agent's expression mesh relative to the dense expression motion information and the epidermal shape (i.e., the target epidermal shape) that should be achieved when the agent's facial mesh is driven.
[0257] The inverse kinematics target optimization task can be an optimization task that continuously optimizes the joint angle parameters corresponding to the control point simulation information according to the target epidermal vertex information. Through the continuous optimization of this inverse kinematics target optimization task, during actual driving execution, according to the optimization result of this inverse kinematics target optimization task, the motion corresponding to the agent's facial epidermal mesh can be made as close as possible to the target epidermal shape defined by the dense expression motion information. This inverse kinematics target optimization task can be expressed in the form of an inverse kinematics objective function.
[0258] Performing the optimization of the above joint angle parameters on the inverse kinematics target optimization task can make the agent's epidermal mesh as close as possible to the target epidermal shape until the optimal joint angle parameters are obtained. When the optimal joint angle parameters are used as driving parameters by the corresponding driving motors of the agent's facial structure as shown in Figure 5C it can make the corresponding driven control points and their surrounding epidermal areas present the most fitting expression imitation effect with the dense expression motion information during epidermal driving, the expression is more delicate and natural, and at the same time, it can ensure that the epidermal skin around the control points presents the most natural epidermal changes.
[0259] Therefore, the optimal joint angle parameters can be the motor rocker joint angle parameters that the corresponding control points and their surrounding epidermal areas can reach the closest to the target epidermal deformation when driven by the motor. As shown in Figure 5CAs shown in the figure, a "Skin", "Muscle" structure and an Electrical Control drive structure of a silicone mask are supported by a "Skeleton" structure. The Electrical Control drive structure includes various drive motors or servo motors, which can perform rotation or even pulling actions according to the optimal joint angle parameters under the control of a controller, thereby driving the connected Steel Cable / Teflon Conduit to generate an optimal displacement in a specified direction at a specified position (control point) on the connected silicone mask ("Skin"), reaching the optimal position of the control point under the target expression, thus generating epidermal deformation on the face of the agent, completing the optimization of the execution of the expression action, and realizing the natural deformation of the facial expression of the agent. Among them, the control point can be the fixed point (such as adhesive fixation) of the above-mentioned Steel Cable / Teflon Conduit on the inner surface of the silicone mask. In addition, the above-mentioned "Skeleton" structure, "Skin", "Muscle" structure and Electrical Control drive structure can constitute the actuator of the agent.
[0260] Among them, since the optimal joint angle parameters are obtained by optimizing the inverse kinematics target optimization task according to the agent expression grid of dense motion expression information, the above-mentioned execution process of the expression action corresponds to the execution process of the agent expression imitation. In short, the generation of the agent's facial expression depends on the kinematic modeling of the subcutaneous control points, and at the same time, the joint angle parameters obtained by inverse kinematics solution are used to drive the movement of the control points, thereby realizing the transmission of the dense expression motion information obtained in the perception stage to the vertices of the agent's facial epidermal grid, ensuring the natural deformation of the expression restoration process.
[0261] Therefore, through the above-mentioned execution optimization method of agent expression imitation in Embodiment 3 of the present invention, by defining the control point area, the optimal position of the control point under the target expression can be obtained by using inverse kinematics, and accurate actuator motion instructions can be generated, so that the motion state of the control point area approaches the target epidermal motion. In this case, only by increasing the density of the control points can the fineness and naturalness of the expression restoration be further improved without changing the algorithms in the perception and mapping stages. Therefore, not only the accuracy of expression imitation is greatly improved, but also the physical executability of the generated instructions is ensured through the explicit modeling of the physical constraints of the actuator, avoiding the execution failure problem caused by ignoring the physical constraints in the traditional method. In addition, the independence of the control point optimization also enables the execution stage to flexibly adapt to different types of agent facial structures, providing technical guarantees for the scalability and modular design of the system, and having extremely high engineering application value and commercial application value.
[0262] As Figures 4A - 5C shown, according to an embodiment of the present invention, before generating the inverse kinematics target optimization task based on the control point simulation information of the agent expression grid and the target epidermal vertex information in operation S401, it further includes:
[0263] Generating an agent expression grid corresponding to the agent according to the blend shape basis of the agent's face;
[0264] Generating a distance weight matrix based on the control point simulation information of the agent expression grid and the target epidermal vertex information.
[0265] The blend shape basis of the agent's face can be a semantic expression space for semantic deformation offset of the surface according to the agent's face topology, used to describe the vertex offset of the semantic expression corresponding to the dense expression motion information relative to the agent's face topology. Specifically, reference can be made to the third blend shape basis provided in the above-mentioned embodiment 2 . The dense expression motion information extracted in the perception stage can be cross-topologically semantically migrated to the above-mentioned blend shape basis of the agent's face through the mapping process of the blend shape coefficients in the mapping stage.
[0266] Furthermore, an agent expression grid can be formed based on the blend shape basis of the agent's face , which can be specifically reflected in Formula 20 in the above-mentioned embodiment 2. In this way, it can ensure that the semantics of the robot's expression are consistent with human expressions, and at the same time can adapt to the geometric characteristics of the agent's face topology, ensuring that it can be aligned with the semantics of the perceived human face expression, thus ensuring the accuracy of subsequent physical drive execution, guaranteeing the natural deformation during the expression restoration process, and at the same time ensuring a better restoration degree of the perceived human face expression.
[0267] Each constituent matrix element in the distance weight matrix can be used to represent the distance assignment between a single control point on the agent expression grid and the corresponding epidermal vertex . Therefore, the distance weight matrix can define the distance-based weight assignment between the epidermal control points and the epidermal grid vertices. Among them, these matrix elements can be specifically expressed as the following Formula 21:
[0268] (21)
[0269] Wherein, is the Euclidean distance between the epidermal vertex and the control point ; is the distance from the epidermal vertex to the control point; are the three-dimensional position coordinates of the subcutaneous controller; is the control parameter of the distance influence, used to adjust the attenuation range of the weight; Weight normalization ensures that (the sum of the weights of each epidermal vertex is 1), is the number of control points.
[0270] Among them, a single control point can be the position information formed by performing position simulation calculations on the agent's expression mesh through a simulation tool. The set of these control points can constitute the above-mentioned control point simulation information. Among them, each control point can be used to define the three-dimensional coordinates of the subcutaneous traction position. Correspondingly, a single epidermal vertex can be the position expected by the dense expression motion information reached by the vertex of the corresponding mesh position. The set of these epidermal vertices can constitute the above-mentioned target epidermal vertex information.
[0271] Therefore, based on the control point simulation information and the target epidermal vertex information, the dense expression motion information of the agent's expression mesh can be mapped onto the topological structure of the agent's facial structure, and the accurate description of the agent's facial topology and actuator position can be achieved through the distance weight matrix.
[0272] Such as Figures 4A - 5C shown, according to an embodiment of the present invention, before generating the distance weight matrix based on the control point simulation information and the target epidermal vertex information of the agent's expression mesh, it further includes:
[0273] Generating control point simulation information according to the preset forward kinematics optimization task of the agent's expression mesh and the joint angle parameter vector corresponding to the simulation control points;
[0274] Generating target epidermal vertex information according to the deformation weight matrix of the epidermal mesh vertices of the agent's expression mesh and the control point simulation information.
[0275] On the basis of constructing the geometric model and kinematic structure of the agent's face corresponding to the agent's expression mesh (such as joint control points and their associated relationships) through a simulation tool, the preset forward kinematics optimization task can be expressed as a forward kinematics function based on the motor kinematic joint angle parameters corresponding to the agent's kinematic structure.
[0276] The simulation control points can be the subcutaneous control points defined by the above-mentioned geometric model and kinematic structure of the agent's face. The set of subcutaneous control points is defined as , each control point represents the three-dimensional coordinates of a certain position under the skin, is the number of control points. Each simulation control point corresponds to different joint angle parameters, which can be used as sub-elements of the above joint angle parameter vector.
[0277] Therefore, according to the input parameter vector such as joint angles and combined with the kinematic relationship and geometric constraints through a simulation tool, the forward kinematic information of the control point position can be generated as the control point simulation information.
[0278] Specifically, for the implementation operation of the forward kinematics of the control point simulation information, first, a geometric model and kinematic structure of the intelligent agent's face can be constructed through a simulator (such as Blender), including joints, control points and their associated relationships, and then the simulator is used according to the input joint angle parameters , combined with the kinematic relationship and geometric constraints, to directly calculate the position of the control point .
[0279] Furthermore, the control point position corresponding to the control point simulation information is driven by the kinematic joint angle parameters , and the forward kinematic model of the control point simulation information is defined as the following formula 22:
[0280] (22)
[0281] Where, is the forward kinematic function; is the joint angle parameter vector; is the number of degrees of freedom of kinematics; is the number of control points. Through this, the control point simulation information can complete the construction of the control point model.
[0282] The skin mesh vertices of the intelligent agent's expression mesh can be the vertices of the face mesh based on the geometric model and kinematic structure of the intelligent agent's face corresponding to the intelligent agent's expression mesh, and can specifically be expressed as a set of skin mesh vertices as , where each skin mesh vertex represents the three-dimensional coordinates of a certain point on the skin, and is the number of skin mesh vertices.
[0283] Correspondingly, the control point movement of the control point model corresponding to the control point simulation information can drive the deformation of the corresponding skin mesh vertices through the deformation weight matrix . Among them, the deformation weight matrix can represent the degree of deformation influence of the control point on the skin mesh vertices. Specifically, the target skin vertex information can be expressed as a linear relationship through the deformation weight matrix and the control point simulation information as the following formula 23: drives the deformation of the corresponding skin mesh vertices . Among them, the deformation weight matrix can represent the degree of deformation influence of the control point on the skin mesh vertices. Specifically, the target skin vertex information can be expressed as a linear relationship through the deformation weight matrix and the control point simulation information as the following formula 23:
[0284] (23)
[0285] Among them, is the deformation weight matrix, representing the influence relationship of the control points on the skin vertices ; is the global position of the control points, stacked in order as ; is the global position of the skin mesh vertices, stacked in order as . Therefore, the construction of the skin mesh vertex model can be completed. Among them, the deformation weight matrix can be a weight construction model of the distance weight matrix, used to generate the distance weight matrix according to the agent expression mesh production.
[0286] Thereby, the geometric models of the control points and their surrounding areas and the mesh skin vertices can be accurately defined, so that the optimal position of the control points under the target expression can be calculated accordingly, and accurate actuator motion instructions can be generated to ensure that the motion state of the control point area is close to the target skin motion. Moreover, it is also beneficial to further improve the fineness and naturalness of expression restoration by increasing the control point density without changing the algorithms in the perception stage and the mapping stage.
[0287] As Figures 4A - 5C shown, according to an embodiment of the present invention, in the operation S401 of generating the inverse kinematics target optimization task based on the control point simulation information and the target skin vertex information of the agent expression mesh, it includes:
[0288] Expanding the distance weight matrix into a weight amplification matrix;
[0289] Generating control points to drive the skin mesh vertices according to the control point simulation information and the weight amplification matrix;
[0290] Generating the inverse kinematics target optimization task according to the control points driving the skin mesh vertices and the target skin vertex information.
[0291] The weight amplification matrix can be formed by three-dimensional expansion of each sub-element in the aforementioned distance weight matrix. Among them, since each control point and the skin mesh vertex both have three-dimensional coordinates ( ), the dimension of this weight amplification matrix can be . The specific form is as formula 24 below:
[0292] (24)
[0293] Among them, each matrix element Indicates the corresponding control point For the epidermal mesh vertices The degree of influence extends to each coordinate component ( ).
[0294] The inverse kinematics target optimization task can be embodied as an objective optimization function for optimizing the joint angle parameters corresponding to the control points by combining the position information of the target epidermal vertices, so as to realize the inverse kinematics solution of the targeted joint angle parameters and obtain the optimal joint angle parameters required for the control points to control the epidermis.
[0295] Specifically, the goal of the inverse kinematics target optimization task is to optimize the joint angle parameters according to the target epidermal shape so that the epidermal mesh is as close as possible to the target shape. Among them, the inverse kinematics target optimization task can be defined in the form of an objective function as the following formula 25:
[0296] (25)
[0297] Among them, is the position of the target epidermal mesh vertex defined by the target epidermal vertex information; is the position of the control point calculated by simulation defined by the foregoing control point simulation information; is the epidermal mesh vertex generated by the drive of the control point, which can be linearly combined by the control point simulation information and the weight amplification matrix.
[0298] As Figures 4A - 5C shown, according to an embodiment of the present invention, in the operation S402 of performing optimization processing on the inverse kinematics target optimization task to obtain the optimal joint angle parameters corresponding to the control point simulation information, it includes:
[0299] Performing a minimization objective optimization iteration process on the inverse kinematics target optimization task according to the joint angle constraint conditions to generate the optimal joint angle parameters.
[0300] The joint angle constraint conditions can include joint angle limit constraints and geometric constraints of the kinematic structure of the agent's face. Among them, the joint angle limit constraint can define the activity range of the joint angle between the minimum joint angle and the maximum joint angle of the motor rocker arm, and the geometric constraint of the kinematic structure can include, for example, the pulling range of the steel cable / PTFE catheter, the deformation range of the "skin" of the silicone mask, etc. Through the minimization objective iteration process (such as nonlinear optimization) of the inverse kinematics target optimization task for the above formula 25, the control deformation position of the control point is continuously close to the optimal control point position under the target expression, and the joint angle of the motor rocker arm corresponding to the optimal control point position (for example, relative to the starting joint angle) can be used as the optimal joint angle parameter 。
[0301] Therefore, by combining the above kinematic modeling and inverse kinematic solution, for the control point area defined on the agent's face, the actuator command solution can be achieved through inverse kinematics under physical constraints, so that it approximates the target epidermal motion generated in the mapping phase.
[0302] As Figures 4A - 5C shown, according to an embodiment of the present invention, in the execution optimization of agent expression imitation according to the optimal joint angle parameters in operation S403, it includes:
[0303] Obtaining the control point drive information of agent expression imitation through the control point simulation information and the optimal joint angle parameters;
[0304] Performing the drive of agent expression imitation according to the control point drive information to complete the execution optimization.
[0305] The forward kinematic function corresponding to the control point simulation information obtained through the simulation tool , further combined with the optimal joint angle parameters obtained after optimizing and processing the above inverse kinematic target optimization task , can define the control point drive information for the agent's facial motor control system to achieve agent expression imitation as formula 26 below:
[0306] (26)
[0307] Among them, when the control point drive information can define the motor rocker of the motor control system to be controlled to be driven according to the above optimal joint angle parameters, the deformation position of the corresponding control point on the facial epidermis and the deformation position information of the epidermal mesh in the surrounding area can be obtained.
[0308] Specifically, the epidermal mesh shape can be generated by the drive of the corresponding control point, and the epidermal mesh shape can be expressed as formula 27 below:
[0309] (27)
[0310] Therefore, as Figure 5C shown, the control point drive information defined by the above formulas 26 and 27 can be converted into executable instructions for physical actuators (such as motors), so as to achieve epidermal drive and control point fitting, so that the mixed shape coefficients based on dense expression motion information generated in the mapping phase can be converted into physical executable agent facial epidermal motion, and the precise expression imitation can be achieved through the actuator driving the epidermal motion, realizing the high-precision physical drive optimization for expression imitation.
[0311] In summary, according to the execution optimization method for agent expression imitation in the embodiments of the present invention, during the execution phase, an optimization driving means based on control points can be constructed for the physical characteristics of the agent actuator. First, the control point area is defined, and the inverse kinematics is used to solve the optimal positions of the control points under the target expression, and precise actuator motion instructions are generated, so that the motion state of the control point area approximates the target skin motion. By increasing the density of the control points, the fineness and naturalness of expression restoration can be further improved without changing the algorithms in the perception and mapping phases.
[0312] This method not only improves the accuracy of expression imitation, but also ensures the physical executability of the generated instructions through explicit modeling of the physical constraints of the actuator, avoiding the execution failure problems caused by ignoring physical constraints in traditional methods. In addition, the independence of control point optimization enables the execution phase to flexibly adapt to different types of robot facial structures, providing technical guarantees for the scalability and modular design of the system.
[0313] In summary, in the expression driving phase, by explicitly modeling the physical constraints of the robot actuator and optimizing the expression driving parameters, it is ensured that the generated control instructions are both physically executable and meet the requirements of high precision and real-time performance.
[0314] Based on the above execution optimization method for agent expression imitation, the present invention also provides an execution optimization device for agent expression imitation. The following will be combined with Figure 4B to describe this device in detail.
[0315] Figure 4B The structural block diagram of the execution optimization device for agent expression imitation according to Embodiment 3 of the present invention is schematically shown.
[0316] As Figure 4B shown, the execution optimization device 400 for agent expression imitation in this Embodiment 3 includes a target task generation module 410, a task optimization module 420, and an imitation execution module 430.
[0317] The target task generation module 410 is used to generate an inverse kinematics target optimization task based on the control point simulation information of the agent expression grid and the target skin vertex information. In one embodiment, the target task generation module 410 can be used to execute the operation S401 described above, which will not be elaborated here.
[0318] The task optimization module 420 is used to perform optimization processing on the inverse kinematics target optimization task to obtain the optimal joint angle parameters corresponding to the control point simulation information. In one embodiment, the task optimization module 420 can be used to execute the operation S402 described above, which will not be elaborated here.
[0319] The imitation execution module 430 is used to perform execution optimization for the intelligent agent's expression imitation according to the optimal joint angle parameters. In one embodiment, the imitation execution module 430 can be used to perform the operation S403 described above, which will not be elaborated here.
[0320] According to an embodiment of the present invention, any multiple of the target task generation module 410, the task optimization module 420, and the imitation execution module 430 can be combined and implemented in one module, or any one of them can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the target task generation module 410, the task optimization module 420, and the imitation execution module 430 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable way of integrating or packaging circuits, etc., in hardware or firmware, or in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the target task generation module 410, the task optimization module 420, and the imitation execution module 430 can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions can be executed.
[0321] Combining the above embodiments 1-3, it can be seen that: the existing intelligent agent expression imitation technology has significant technical limitations in the three key stages of perception, mapping, and execution. These problems directly limit the realization of high-precision expression imitation, especially in cross-entity expression migration and physical drive implementation. Among them, the perception stage faces the problems of sparse expression representation and coupling between expression and face shape; it is difficult to solve the problem of cross-topology expression migration in the mapping stage, and the expression semantics are distorted or even lost during migration; there is a problem that the generated instructions are physically unexecutable in the execution stage. These limitations restrict the fineness and robustness of expression imitation.
[0322] Therefore, in order to overcome the technical bottlenecks existing in the three key stages of perception, mapping, and execution of existing facial expression imitation technologies, the present invention provides a brand-new facial expression imitation technical solution, covering three aspects of perception optimization, mapping optimization, and execution optimization respectively, aiming to achieve higher-precision and more robust facial expression perception and transfer, while ensuring that the generated control instructions are physically executable, laying a technical foundation for the development of highly biomimetic facial expression robots. Among them, through the densification of facial expression representation in the perception stage, cross-topology semantic consistency transfer in the mapping stage, and high-precision physical drive optimization in the execution stage, the core problems such as sparse facial expression description, semantic transfer distortion, and unachievable execution in traditional facial expression imitation methods are systematically solved. The following is the core design of the new technical solution and its corresponding relationship with existing problems.
[0323] In the perception stage, a three-dimensional deformation model (3DMM) is mainly used to replace the traditional sparse key-point representation method. The facial movement is decomposed into three independent subspaces of pose, face shape, and facial expression, generating a standardized and dense facial expression movement representation. That is, dense facial expression movements are extracted from the input facial images, and high-precision facial expression representations under the standardized average face are generated. This method not only improves the refinement degree of facial expression description but also eliminates the coupling problem between facial expression semantics and individual face shapes through the standardized average face topology, achieving the consistency of facial expression semantics among different individuals. For specific details, refer to Embodiment 1.
[0324] In the mapping stage, a set of Blendshape semantic spaces based on the standardized average face is provided as an intermediate bridge between the human face and the robot's face, solving the problems of differences in topological structure and geometric characteristics between the two. The dense facial expression movements generated in the perception stage are real-time mapped to the low-dimensional Blendshape coefficient space, and distortion-free transfer of facial expressions from the human face to the robot's face is achieved through the semantically consistent Blendshape basis. In other words, through the semantically consistent Blendshape basis, the dense facial expression movements are mapped to the robot's face topological structure to achieve cross-topology facial expression transfer. This method not only ensures the consistency of facial expression semantics but also greatly improves the efficiency and robustness of mapping. For specific details, refer to Embodiment 2.
[0325] In the execution stage, through the high-precision physical drive optimization based on control points, the inverse kinematics is used to solve the optimal positions of the control points under the target facial expression, generating accurate actuator movement instructions. That is, the mapping result is converted into physically executable robot skin movement instructions, and accurate facial expression imitation is achieved through inverse kinematics and physical constraint optimization. This method explicitly models the physical constraints of the actuator, ensuring the physical executability of the generated instructions. For specific details, refer to Embodiment 3.
[0326] The process provided in Embodiments 1-3 of the present invention is optimized layer by layer from perception to execution, ensuring the fineness, semantic consistency, and physical executability of facial expression imitation, providing a complete solution for the realization of highly biomimetic facial expression robots. Through systematic design and optimization, this process can achieve full-link facial expression imitation from the original human face input to the precise reproduction of robot facial expressions.
[0327] In short, Embodiments 1-3 of the present invention focus on a real-time imitation system for highly biomimetic facial expression robots based on two-dimensional videos, aiming to achieve precise mapping and physical driving to highly biomimetic facial expression robots by accurately capturing the dynamic features of human facial expressions. The technical framework covers three links: perception, migration, and execution, breaking through the technical bottlenecks in subtle facial expression capture, cross-modal facial expression migration from human facial expressions to robot facial expressions, and physical driving in traditional methods. This imitation system shows broad prospects in application scenarios. For example, in a remote embodied interaction system, an operator can remotely control a highly biomimetic facial expression robot through video input to achieve cross-space embodied emotional interaction, which is applicable to scenarios that require emotional expression such as remote medical escort and psychological counseling. In the basic research of social emotion computing, this technology provides strong technical support for the autonomous emotional expression of robots and human-robot emotional interaction, promoting the further development of social emotion robot research. Generally speaking, through an innovative technical optimization framework for facial expression perception, mapping, and execution, the key technical problems in the facial expression migration process are systematically solved, and the real-time precise driving of highly biomimetic facial expression robots is successfully achieved. This technology not only provides reliable technical support for remote emotional interaction but also lays an important foundation for the development of future autonomous emotional interaction robots, with extremely high commercial application value.
[0328] Figure 6 A block diagram of an electronic device suitable for implementing a perception optimization method, a mapping optimization method, and an execution optimization method for agent facial expression imitation according to Embodiments 1-3 of the present invention is schematically shown.
[0329] The above-mentioned electronic device provided by the embodiment of the present invention includes one or more processors and a memory. The memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned perception optimization method, mapping optimization method, and execution optimization method for agent facial expression imitation.
[0330] As Figure 6As shown, an electronic device 600 according to an embodiment of the present invention includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage section 608 into a random access memory (RAM) 603. The processor 601 can include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), etc. The processor 601 can also include on-board memory for caching purposes. The processor 601 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0331] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method flow according to an embodiment of the present invention by executing the programs in the ROM 602 and / or RAM 603. It should be noted that the program can also be stored in one or more memories other than the ROM 602 and RAM 603. The processor 601 can also perform various operations of the method flow according to an embodiment of the present invention by executing the programs stored in the one or more memories.
[0332] According to an embodiment of the present invention, the electronic device 600 may further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed into the storage section 608 as needed.
[0333] The present invention also provides a computer-readable storage medium having executable instructions stored thereon, and when the instructions are executed by a processor, the processor is caused to execute the above-mentioned perception optimization method, mapping optimization method, and execution optimization method for agent expression imitation.
[0334] Among them, the computer-readable storage medium may be included in the device / apparatus / system described in the above embodiments; or it may exist alone without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the above-mentioned perception optimization method, mapping optimization method, and execution optimization method for agent expression imitation according to Embodiments 1-3 of the present invention are realized.
[0335] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-mentioned ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603.
[0336] An embodiment of the present invention further includes a computer program product, which includes a computer program. When the computer program is executed by a processor, the above-mentioned perception optimization method, mapping optimization method, and execution optimization method for agent expression imitation are realized.
[0337] Among them, the computer program contains program codes for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program codes are used to enable the computer system to implement the above-mentioned perception optimization method, mapping optimization method, and execution optimization method for agent expression imitation provided in Embodiments 1-3 of the present invention.
[0338] When the computer program is executed by the processor 601, the above-mentioned functions defined in the system / apparatus of the embodiment of the present invention are executed. According to an embodiment of the present invention, the above-mentioned systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0339] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 609, and / or be installed from the removable medium 611. The program codes included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0340] In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, the above functions defined in the system of the embodiments of the present invention are executed. According to the embodiments of the present invention, the systems, devices, apparatuses, modules, units, etc. described above can be implemented by computer program modules.
[0341] According to the embodiments of the present invention, the program code for executing the computer program provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include but are not limited to, such as Java, C++, python, the "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by connecting through the Internet using an Internet service provider).
[0342] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0343] In addition, all actions of obtaining information, signals, or data in the present invention are carried out on the premise of complying with the corresponding data protection laws, regulations, and policies of the country where it is located, and obtaining the authorization given by the owner of the corresponding device.
[0344] Those skilled in the art will understand that the features recited in the various embodiments and / or claims of the present invention can be combined or / and combined in various ways, even if such combinations or combinations are not explicitly recited in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features recited in the various embodiments and / or claims of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0345] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in the respective embodiments cannot be used advantageously in combination. The scope of the present invention is defined by the appended claims and their equivalents. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. A mapping optimization method for agent expression imitation, characterized in that: include: generating a first blend shape basis of a standardized average face based on a preset facial action coding rule; Mapping the dense expression motion information of the standardized average face to the first mixed shape basis to generate a second mixed shape basis, including: extracting first dense annotation points on the dense expression grid and second dense annotation points on the average face grid corresponding to the dense expression motion information; obtaining a dense annotation basis matrix according to the mixed shape expression basis matrix corresponding to the first mixed shape basis; generating a mixed shape weight optimization task according to the dense annotation basis matrix, the first dense annotation points and the second dense annotation points; obtaining the mixed shape coefficients corresponding to the mixed shape weight optimization task by minimizing a preset error vector and a preset coefficient constraint; and transferring the blend shape coefficients of the second blend shape basis to a third blend shape basis of the face of the agent to complete the mapping optimization; The optimization problem of the mixed shape coefficient corresponding to the mixed shape weight optimization task satisfies: ; in, Used to represent the first dense annotation point, Used to represent the second dense annotation point, Used to represent dense labeled basis matrices, Used to represent the blend shape coefficients, K is the number of predefined expression bases of the blend shape expression base matrix.
2. The method according to claim 1, characterized in that In the first mixed shape basis for generating a standardized average face based on a preset facial action coding rule, the method includes: Generate a mixed shape expression basis matrix according to a preset facial action coding rule; A first blend shape basis for a normalized average face is generated based on a preset blend shape weight vector and the blend shape expression basis matrix.
3. The method according to claim 1, characterized in that Before mapping the dense expression motion information of the standardized average face to the first mixed shape basis to generate the second mixed shape basis, the method further includes: The dense expression motion information of the standardized average face is generated according to the extracted expression motion information of the target expression image.
4. The method according to claim 1, characterized in that Before transferring the blended shape coefficients of the second blended shape basis to the third blended shape basis of the face of the agent to complete the mapping optimization, the method further includes: A third blended shape basis of the agent's face is generated based on a preset facial action coding rule, wherein the blended shape semantics of the third blended shape basis is the same as that of the first blended shape basis.
5. The method according to claim 1, characterized in that: In the transferring of the blended shape coefficients of the second blended shape basis to the third blended shape basis of the face of the agent, completing the mapping optimization, comprising: The blend shape coefficients are applied to the third blend shape base to generate an agent expression mesh corresponding to the agent.
6. A mapping optimization device for imitating the facial expressions of an intelligent agent, characterized in that: include: A base generation module, for generating a first blend shape base of a standardized average face based on a preset facial action coding rule; An information mapping module is used to map the dense expression motion information of the standardized average face to the first mixed shape basis to generate a second mixed shape basis, which includes: extracting first dense annotation points on the dense expression grid and second dense annotation points on the average face grid corresponding to the dense expression motion information; obtaining a dense annotation basis matrix according to the mixed shape expression basis matrix corresponding to the first mixed shape basis; generating a mixed shape weight optimization task according to the dense annotation basis matrix, the first dense annotation points and the second dense annotation points; obtaining the mixed shape coefficient corresponding to the mixed shape weight optimization task by minimizing a preset error vector and a preset coefficient constraint condition; and A coefficient transfer module, used to transfer the blended shape coefficients of the second blended shape basis to the third blended shape basis of the face of the agent to complete the mapping optimization; The optimization problem of the mixed shape coefficient corresponding to the mixed shape weight optimization task satisfies: ; in, Used to represent the first dense annotation point, Used to represent the second dense annotation point, Used to represent dense labeled basis matrices, is used to represent the blend shape coefficients, The number of predefined expression bases for the blend shape's expression base matrix.
7. An electronic device comprising: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors execute the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the method according to any one of claims 1 to 5.
9. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Image processing method and device, equipment and computer storage medium
CN113808249A