Execution Optimization Method, Device, Equipment and Medium for Agent Expression Imitation

By explicitly modeling the physical constraints of the agent actuator, and using inverse kinematics to optimize expression driver parameters, the problem of poor control accuracy and stability in expression imitation is solved, and high-precision and natural expression imitation effect is achieved.

CN119830944BActive Publication Date: 2025-07-11BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510300928.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-11
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

In the driving execution stage of expression imitation, the prior art has problems such as inaccurate inverse kinematic solution and insufficient physical constraints of robotic actuators, resulting in poor expression control accuracy and stability.

Method used

By explicitly modeling the physical constraints of the agent actuator, the expression driver parameters are optimized, and the inverse kinematics solution is used to generate precise control instructions to ensure that the generated instructions are physically executable and high precision.

Benefits of technology

It improves the accuracy and nature of expression imitation, avoids execution failure caused by ignoring physical constraints, and has high engineering application value and commercial value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119830944B_ABST
    Figure CN119830944B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides an execution optimization method for agent expression imitation, which can be applied to the field of artificial intelligence technology. The execution optimization method for agent expression imitation includes: generating an inverse kinematics target optimization task based on the control point simulation information of the agent expression grid and the target skin vertex information; performing optimization processing on the inverse kinematics target optimization task to obtain the optimal joint angle parameters corresponding to the control point simulation information; and performing execution optimization of agent expression imitation according to the optimal joint angle parameters. An embodiment of the present invention also provides an execution optimization device, device, storage medium, and program product for agent expression imitation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, specifically to the field of image processing technology, and more specifically to an execution optimization method, device, equipment, medium and product for intelligent agent expression imitation. Background Art

[0002] Artificial Intelligence (AI for short) is an important driving force for the new round of scientific and technological revolution and industrial transformation. It is a new key technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence. As an important part of intelligent science, artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence (such as a highly biomimetic expression robot).

[0003] As an important part of artificial intelligence technology, the technology of human expression recognition and reproduction (such as expression imitation) has received extensive attention in the fields of human-computer interaction, security recognition, robot manufacturing, automation, medicine, communication and driving, and has quickly become a research hotspot in the industry. The expression imitation technology is mainly divided into three stages: expression perception, mapping migration and driving execution. The aim is to accurately capture the dynamic expression features of the human face in the expression perception stage, realize the face mapping to embodied intelligence such as a highly biomimetic expression robot in the mapping migration stage, and finally control the expression actuator to achieve accurate expression reproduction driving in the driving execution stage. However, in the existing traditional expression imitation technology, there are problems in the driving execution stage, such as it is difficult to ensure the accuracy of inverse kinematics solution, the physical constraints of the robot actuator cannot be fully considered, and the control accuracy and stability of the expression are poor. Summary of the Invention

[0004] In view of at least one of the technical problems existing in the above-mentioned execution scheme of expression imitation, embodiments of the present invention provide an execution optimization method, device, equipment, medium and product for intelligent agent expression imitation, so as to explicitly model the physical constraints of the intelligent agent actuator in the expression driving execution stage, optimize the expression driving parameters, and ensure that the generated control instructions are both physically executable and meet the requirements of high precision and real-time performance.

[0005] One aspect of the embodiments of the present invention provides an execution optimization method for intelligent agent expression imitation, which includes: generating an inverse kinematics target optimization task based on the control point simulation information of the intelligent agent expression grid and the target skin vertex information; performing optimization processing on the inverse kinematics target optimization task to obtain the optimal joint angle parameters corresponding to the control point simulation information; and performing execution optimization of the intelligent agent expression imitation according to the optimal joint angle parameters.

[0006] According to an embodiment of the present invention, before generating an inverse kinematics target optimization task based on the control point simulation information and target skin vertex information of the agent expression mesh, it further includes: generating an agent expression mesh corresponding to the agent according to the blend shape basis of the agent's face; generating a distance weight matrix based on the control point simulation information of the agent expression mesh and the target skin vertex information.

[0007] According to an embodiment of the present invention, before generating a distance weight matrix based on the control point simulation information of the agent expression mesh and the target skin vertex information, it further includes: generating control point simulation information according to the preset forward kinematics optimization task of the agent expression mesh and the joint angle parameter vector corresponding to the simulation control points; generating target skin vertex information according to the deformation weight matrix of the skin mesh vertices of the agent expression mesh and the control point simulation information.

[0008] According to an embodiment of the present invention, in generating an inverse kinematics target optimization task based on the control point simulation information and target skin vertex information of the agent expression mesh, it includes: expanding the distance weight matrix into a weight amplification matrix; generating control point-driven skin mesh vertices according to the control point simulation information and the weight amplification matrix; generating an inverse kinematics target optimization task according to the control point-driven skin mesh vertices and the target skin vertex information.

[0009] According to an embodiment of the present invention, in performing optimization processing on the inverse kinematics target optimization task to obtain the optimal joint angle parameters corresponding to the control point simulation information, it includes: performing a minimization target optimization iteration process on the inverse kinematics target optimization task according to the joint angle constraint conditions to generate the optimal joint angle parameters.

[0010] According to an embodiment of the present invention, in performing optimization on the agent expression imitation according to the optimal joint angle parameters, it includes: obtaining the control point driving information of the agent expression imitation through the control point simulation information and the optimal joint angle parameters; performing the driving of the agent expression imitation according to the control point driving information to complete the execution optimization.

[0011] Another aspect of the embodiments of the present invention provides an execution optimization device for agent expression imitation, which includes a target task generation module, a task optimization module, and an imitation execution module. The target task generation module is used to generate an inverse kinematics target optimization task based on the control point simulation information and target skin vertex information of the agent expression mesh; the task optimization module is used to perform optimization processing on the inverse kinematics target optimization task to obtain the optimal joint angle parameters corresponding to the control point simulation information; the imitation execution module is used to perform optimization on the agent expression imitation according to the optimal joint angle parameters.

[0012] Another aspect of an embodiment of the present invention provides an electronic device, including one or more processors and a memory. The memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned execution optimization method for agent expression imitation.

[0013] Another aspect of an embodiment of the present invention provides a computer-readable storage medium, on which executable instructions are stored. When the instructions are executed by a processor, the processor is caused to execute the above-mentioned execution optimization method for agent expression imitation.

[0014] Another aspect of an embodiment of the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the above-mentioned execution optimization method for agent expression imitation is implemented.

[0015] The execution optimization method for agent expression imitation provided by the embodiment of the present invention can at least partially solve the problems in the related art, and thus can at least achieve one of the following technical effects:

[0016] Through the above-mentioned execution optimization method for agent expression imitation in Embodiment 3 of the present invention, by defining a control point area, the optimal positions of the control points under the target expression can be solved using inverse kinematics, and precise actuator motion instructions can be generated, so that the motion state of the control point area approaches the target skin motion. In this case, by simply increasing the density of the control points, the fineness and naturalness of expression restoration can be further improved without changing the algorithms in the perception and mapping stages. Therefore, not only is the accuracy of expression imitation greatly improved, but also the physical executability of the generated instructions is ensured through explicit modeling of the physical constraints of the actuators, avoiding the execution failure problems caused by ignoring physical constraints in traditional methods. In addition, the independence of control point optimization also enables the execution stage to flexibly adapt to different types of agent facial structures, providing technical guarantees for the scalability and modular design of the system, and having extremely high engineering application value and commercial application value.

[0017] It should be understood that the above general description and the following specific embodiments are only exemplary and explanatory, and cannot limit the scope claimed by the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features, and advantages of the present invention will become clearer. In the drawings:

[0019] Figure 1 Schematically shows an application scenario diagram of the optimization method, device, equipment, medium, and program product for agent expression imitation according to Embodiments 1-3 of the present invention;

[0020] Figure 2A Schematically shows a flowchart of a perception optimization method for agent expression imitation according to Embodiment 1 of the present invention;

[0021] Figure 2B Schematically shows a structural block diagram of a perception optimization device for agent expression imitation according to Embodiment 1 of the present invention;

[0022] Figure 3A Schematically shows a flowchart of a mapping optimization method for agent expression imitation according to Embodiment 2 of the present invention;

[0023] Figure 3B Schematically shows a structural block diagram of a mapping optimization device for agent expression imitation according to Embodiment 2 of the present invention;

[0024] Figure 4A Schematically shows a flowchart of an execution optimization method for agent expression imitation according to Embodiment 3 of the present invention;

[0025] Figure 4B Schematically shows a structural block diagram of an execution optimization device for agent expression imitation according to Embodiment 3 of the present invention;

[0026] Figure 5A Schematically shows an expression semantic alignment diagram from the human face expression mixing space to the robot expression mixing space according to Embodiments 1 - 3 of the present invention;

[0027] Figure 5B Schematically shows a robot expression kinematics design diagram according to Embodiments 1 - 3 of the present invention;

[0028] Figure 5C Schematically shows a robot expression electromechanical design diagram according to Embodiments 1 - 3 of the present invention; and

[0029] Figure 6 Schematically shows a block diagram of an electronic device according to the method of Embodiments 1 - 3 of the present invention.

[0030] The above - mentioned drawings are a part of the specification of the embodiments of the present invention, which illustrate the exemplary embodiments of the present invention. The attached drawings, together with the description of the specification, are used to explain the principles of the embodiments of the present invention. It should be understood that the above general description of the drawings and the following detailed description are only exemplary and explanatory, and they do not limit the scope that the present invention intends to claim. Detailed Description of the Invention

[0031] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer and more understandable, the following will clearly explain the spirit of the content disclosed by the present invention with reference to the drawings and detailed descriptions. After any person skilled in the relevant technical field understands the embodiments of the content of the present invention, they can make changes and modifications based on the technologies taught by the content of the present invention, and such changes and modifications do not depart from the spirit and scope of the content of the present invention.

[0032] The exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but not to limit the present invention. In addition, elements / components with the same or similar reference numerals used in the drawings and embodiments represent the same or similar parts.

[0033] Regarding the use of "first", "second",... etc. in the present invention, it does not particularly refer to the order or sequence, nor is it used to limit the present invention. It is only used to distinguish elements or operations described with the same technical terms.

[0034] Regarding the directional terms used in the present invention, such as: up, down, left, right, front or back, etc., they are only references to the directions in the drawings. Therefore, the directional terms used are for explanation and not for limiting the present creation.

[0035] Regarding the use of "comprising", "including", "having", "containing", etc. in the present invention, they are all open-ended terms, that is, they mean including but not limited to.

[0036] Regarding the use of "and / or" in the present invention, it includes any one or all combinations of things.

[0037] Regarding "a plurality of" in the present invention, it includes "two" and "more than two"; regarding "a plurality of groups" in the present invention, it includes "two groups" and "more than two groups".

[0038] Regarding the terms "substantially", "about", etc. used in the present invention, they are used to modify any quantity or error that can vary slightly, but these slight variations or errors do not change their essence. Generally, the range of such slight variations or errors modified by such terms can be 20% in some embodiments, 10% in some embodiments, 5% in some embodiments or other values. Those skilled in the art should understand that the aforementioned values can be adjusted according to actual needs and are not limited thereto.

[0039] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used here should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0040] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning that those skilled in the art usually understand this expression (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). In the case of using expressions such as "at least one of A, B, or C, etc.", generally, it should be interpreted according to the meaning that those skilled in the art usually understand this expression (for example, "a system having at least one of A, B, or C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). Those skilled in the art should also understand that substantially any disjunctive conjunction and / or phrase representing two or more alternative items, whether in the specification, claims, or drawings, should be understood as giving the possibility of including one of these items, either side of these items, or both items. For example, the phrase "A or B" should be understood as including the possibility of "A" or "B", or "A and B".

[0041] In addition, all actions of obtaining information, signals, or data in the present invention are carried out on the premise of complying with the corresponding data protection laws, regulations, and policies of the country where it is located and obtaining authorization from the owner of the corresponding device.

[0042] In human interpersonal communication in daily life, by controlling facial expressions, the communication effect can be enhanced. For humans, facial expressions are the result of one or more actions or states of facial muscles. These movements express the emotional state of an individual human being towards the external environment. Facial expressions are a form of non-verbal communication and are the main means of expressing social information among humans.

[0043] For artificial intelligence, in order for intelligent machines to make real human interaction responses in a way similar to human intelligence, it is necessary to consider creating a truly humanoid machine in all aspects, such as limbs, appearance, posture, expression, even speech dialogue, action ability, and human-like thinking patterns. Among them, the expression imitation technology provides strong technical support for today's humanoid robots to enter the real world and provide emotional value and non-verbal power for humans.

[0044] The expression imitation technology aims to accurately capture the dynamic features of human facial expressions and achieve precise mapping and physical drive to highly biomimetic expression robots. Its overall technical framework covers three links: expression perception, mapping migration, and drive execution. No matter which link has problems, it will directly limit the realization of high-precision expression imitation.

[0045] To solve at least one of the existing technical problems in the three links of emotional perception, mapping migration, and driving execution in the process of intelligent agent expression imitation in the existing technology, the present invention provides Embodiments 1-3 to respectively optimize the above three links in the process of intelligent agent expression imitation, with the expectation of significantly improving the intelligent level of intelligent agent expression imitation.

[0046] Figure 1 Schematically shows an application scenario diagram of an optimization method, device, device, medium, and program product for intelligent agent expression imitation according to Embodiments 1-3 of the present invention.

[0047] As Figure 1 shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0048] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).

[0049] The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0050] The server 105 may be a server providing various services, such as a background management server that supports the websites browsed by users using the terminal devices 101, 102, 103 (only for example). The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0051] It should be noted that the optimization method for agent expression imitation provided by the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the optimization device for agent expression imitation provided by the embodiments of the present invention can generally be set in the server 105. The optimization method for agent expression imitation provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103, and / or the server 105. Correspondingly, the optimization device for agent expression imitation provided by the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103, and / or the server 105.

[0052] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0053] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.

[0054] To further illustrate the complete technical process of the three key stages of perception, mapping, and execution in agent expression imitation and ensure the realization of full-link expression imitation from the original human face input to the precise reproduction of the robot's expression, the following further describes in combination with Embodiments 1-3 of the present invention and Figures 1 - 5C as follows:

[0055] It should be noted that the agent can be the execution subject of at least one of the perception optimization method, mapping optimization method, and execution optimization method provided in Embodiments 1-3 of the present invention, or it can also be the execution party controlled by these methods. For example, it can be a humanoid intelligent robot or other AI devices with a simulated human face skin that can achieve human face expression imitation. Specifically, as Figure 5B and Figure 5C shown in the electromechanical design of the agent's facial expression, a simulated skin (silicone mask, i.e., skin) is covered on the surface with "muscle" and "skeleton", so that the action of the simulated skin can be achieved through a motor (such as the LFD-01M servo motor, i.e., LFD-01M servomotor) and a steel cable and a Teflon conduit matched by an electrical control system, thereby achieving the realization of expression imitation.

[0056] The following will be combined Figures 2A - 5C The perception optimization method, mapping optimization method and execution optimization method provided in Examples 1-3 of the present invention are described in detail respectively.

[0057] Example 1

[0058] The expression perception stage involves facial expression recognition. Achieving accurate expression recognition perception, obtaining more complex expression descriptions, and ensuring the accuracy and versatility of expression perception are technical issues that need to be urgently resolved in the embodiments of the present invention.

[0059] Specifically, in the perception stage, traditional methods usually use sparse landmarks to perform structured representation of faces. This method relies on the position and displacement of landmarks to describe the dynamic changes of facial expressions. Although simple and efficient, its main problem is that the complexity of expression description is insufficient. Since sparse landmarks can only capture limited local dynamic features, they cannot fully represent the subtle changes and complex characteristics of human facial expressions, resulting in oversimplification of expression description. In addition, there is a coupling relationship between face shape and expression semantics, which makes it impossible for the same geometric displacement to maintain consistent expression semantics in individuals with different face shapes. This coupling further limits the accuracy and versatility of expression perception.

[0060] Therefore, in view of at least one of the technical problems existing in the above-mentioned prior art, the embodiments of the present invention aim to provide a perception optimization method, device, equipment, medium and product for intelligent body expression imitation, thereby providing a perception stage optimization scheme for intelligent body expression imitation, in order to achieve effective decoupling of individual facial features and expression dynamics, significantly improve the ability to capture subtle expressions through dense expression representation, and overcome the limitations of traditional schemes.

[0061] The following will be based on Figure 1 The scene described by Figures 2A - 5C The perceptual optimization method for intelligent agent expression imitation in the disclosed embodiment is described in detail.

[0062] like Figure 2A As shown, one aspect of an embodiment of the present invention provides a perceptual optimization method for intelligent agent expression imitation, which includes operations S201 to S204.

[0063] In operation S201, facial expression motion information of a target facial expression image is extracted;

[0064] In operation S202, a dense face reconstruction task is generated by using the expression motion information and the preset joint optimization information;

[0065] In operation S203, joint optimization is performed on the dense face reconstruction task to generate three-dimensional face reconstruction information, and the three-dimensional face reconstruction information is used to generate the target three-dimensional face; and

[0066] In operation S204, dense expression motion information of the standardized average face is extracted according to the target three-dimensional face to complete perceptual optimization.

[0067] In an embodiment of the present invention, agent expression imitation may be a process in which an agent obtains a face expression image through recognition and reproduces the face expression in the face expression image by using the above Figure 5C electromechanical design. Among them, the face expression image as the object to be recognized may be image information directly detected by the agent through its own image detection device (such as a camera image sensor, etc.), or image information transmitted to the agent through a network transmission device (wireless or wired). These image information may be video or picture data.

[0068] For example, a humanoid robot, as a type of agent, conducts a video call with a real human user. The humanoid robot can directly or indirectly obtain the video data of the video call, preprocess the video content of the human user including the call object in the video data, and extract the expression features of each frame of the image containing the human user's expression in the video content. Among them, the target expression image can be one of these frame images, and the single-frame image contains the face expression of the expression imitation object (human user) at a certain moment. In addition, the humanoid robot can also directly interact with the human user face to face and obtain and recognize the real-time expression of the human user in real time.

[0069] The expression motion information may be face expression feature parameters extracted for the face expression in the target expression image. These expression feature parameters may involve the expression basis parameters and face expression motion parameters of the target face expression in the target expression image (such as the face shape coefficient, dynamic expression motion coefficient, and face pose coefficient in the camera coordinate system, etc.). Among them, the so-called expression basis parameters may include more than one (such as 50) different expression bases, and each expression basis coefficient represents the displacement of the face expression of the current face image relative to each corner point on the dense mesh of the standardized average face; in addition, the face expression motion parameters are the motion field parameters relative to the standardized average face in the standardized average face coordinate system.

[0070] The preset joint optimization information can be information for three-dimensional face reconstruction based on a standardized average face and expression motion information according to the target three-dimensional face reconstruction requirements, where the target three-dimensional face reconstruction requirements are also related to the target three-dimensional face reconstruction model used for the three-dimensional face reconstruction. The standardized average face is the three-dimensional face average shape obtained through statistical learning (i.e., the three-dimensional average face, Mean Shape). The preset joint optimization information is mainly used to generate a dense face reconstruction task according to the expression motion information of the extracted target expression image. The dense face reconstruction task can be a joint optimization task for restoring the three-dimensional face shape and expression motion parameters according to the target three-dimensional face reconstruction requirements. For example, it can be embodied as a joint optimization problem created for three-dimensional face shape restoration.

[0071] By performing joint optimization processing on the dense face reconstruction task, three-dimensional face reconstruction information can be generated according to the processing result of the dense face reconstruction task, where the three-dimensional face reconstruction information can include the corresponding face expression basis and the corresponding face expression coefficients . The face expression basis is usually used to describe the dynamic expression changes of the face, and the face expression coefficients are usually used as a parametric representation of the face expression. The face expression basis and the face expression coefficients can be configured to construct a three-dimensional face expression model based on the standardized average face.

[0072] The target three-dimensional face can be understood as a three-dimensional face expression model based on the standardized average face. For example, a three-dimensional deformation model (such as 3D Morphable Model, abbreviated as 3DMM) can be a statistical model for three-dimensional face modeling. It mainly learns a large amount of three-dimensional face data and represents any three-dimensional face as a combination of linear bases. It can be represented by the three-dimensional face shape based on the standardized average face, and the three-dimensional face shape can be based on the standardized average face (i.e., the three-dimensional average face, Mean Shape), the face shape basis , the face expression basis , the face shape coefficients , the face expression coefficients etc. for description. Therefore, the face expression in the two-dimensional image can be mapped to a parametric three-dimensional representation, and with its ability to decompose the face shape, expression, and pose space, the effective decoupling of individual face shape features and expression dynamics can be achieved.

[0073] Among them, the dense expression motion information can be information related to the dynamic changes of the face expression directly related to the extracted expression motion information, and can be specifically characterized by a linear combination of the face expression basis and the face expression coefficients. The dense expression motion information can most directly reflect the dense expression motion information of the user extracted by the agent on the standardized average face, realizing face expression representation.

[0074] Therefore, in the perception stage, compared with the traditional sparse key-point-based facial expression representation scheme, the above-mentioned perception optimization method for intelligent agent facial expression imitation according to the embodiments of the present invention can represent the facial expression motion information by using the target three-dimensional face, map the facial expression in the two-dimensional image into a parameterized three-dimensional representation, and by virtue of its decomposition ability for face shape, expression, and pose spaces, effectively decouple the individual face shape characteristics from the expression dynamics, generate a dense expression motion under the standardized average face topology, thereby overcoming the abstract defects of the traditional sparse key-point scheme, realizing the densification of expression description, and at the same time, by normalizing the expression to the standardized average face, eliminating the coupling problem between the expression semantics and the individual face shape, ensuring the consistency of the expression semantics among individuals with different face shapes. Therefore, by virtue of the high accuracy and density of the three-dimensional facial expression model based on the standardized average face of the above-mentioned target three-dimensional face, the accuracy and fineness of facial expression perception can be significantly improved, and the ability to capture subtle expressions can be enhanced, thus laying a high-quality foundation for the subsequent mapping and execution stages.

[0075] As Figures 2A - 5C shown, according to an embodiment of the present invention, in the operation of extracting the facial expression motion information of the target expression image in S201, it includes: performing facial expression feature recognition on the target expression image to extract the facial expression motion information.

[0076] The target expression image can be a single-frame two-dimensional image (such as an RGB image), which can be selected from a time-sequence-based two-dimensional image set (such as a video), and specifically can be selected sequentially according to the time sequence. Each time the facial expression feature recognition is mainly to extract the facial expression features of the target face in the selected single-frame two-dimensional image, and the extracted content mainly includes the feature parameters of the facial expression of the target face in the above single-frame two-dimensional image. These feature parameters can include expression basis parameters, facial expression motion parameters, etc.

[0077] For a complete video, it can record the face images at each moment within a certain time period, and at the same time, the face images at each moment record the instantaneous facial expression images at that moment, and these instantaneous facial expression images are the above-mentioned target expression images, serving as the targets for facial expression feature extraction. Therefore, by performing facial expression feature recognition on all the target expression images in the above video one by one and realizing the representation of the dense facial expression motion of each target expression image one by one, the perception and acquisition of the facial expressions within the corresponding time period of the entire video can be achieved.

[0078] Therefore, through the above facial expression feature recognition, the accurate facial expression feature extraction of each frame of two-dimensional image can be realized, thus laying a foundation for the capture of subtle expressions and the improvement of the accuracy and fineness of facial expression perception.

[0079] As Figures 2A - 5CAs shown, according to an embodiment of the present invention, in operation S201 for generating a dense face reconstruction task through expression motion information and preset joint optimization information, it includes: obtaining task reconstruction information corresponding to the dense face reconstruction task according to the expression parameters of the expression motion information; constructing the dense face reconstruction task according to the task reconstruction information and the preset joint optimization information.

[0080] The expression parameters of the expression motion information can be expression feature parameters extracted from the facial expressions of the target expression image, specifically involving expression basis parameters and facial expression motion parameters. Among them, the facial expression motion parameters can be the facial shape coefficients, dynamic expression motion coefficients, and facial pose coefficients in the camera coordinate system, etc.

[0081] To perceive the facial expression motion information of the expression motion information and achieve representation based on the standardized average face, three-dimensional face reconstruction can be performed based on the standardized average face and the expression motion information. Among them, the task reconstruction information can be used to construct the optimization variable information and objective function information of the joint optimization problem corresponding to the above-mentioned dense face reconstruction task, specifically determined according to the construction requirements of the dense face reconstruction task. Among them, the preset joint optimization information can define the construction rules of the dense face reconstruction task, and through this task construction rule, the dense face reconstruction task can be generated according to the task reconstruction information.

[0082] Specifically, for the above-mentioned single-frame target expression image, a dense face reconstruction task is constructed to generate the joint optimization problem required for dense three-dimensional face reconstruction.

[0083] Among them, the single-frame target expression image can be expressed as , is the width of the image, is the height of the image. Among them, the reconstruction target of the dense face reconstruction task is to restore the corresponding three-dimensional face shape and its expression motion parameters .

[0084] According to the task construction rule defined by the preset joint optimization information, the dense face reconstruction task can be formalized into the following joint optimization problem (Formula 1) based on the optimization variable information and objective function information corresponding to the task reconstruction information:

[0085] (1)

[0086] Among them, the optimization variable information includes the facial shape coefficients for reconstructing the individual static facial shape features, the facial expression coefficients for reconstructing the dense dynamic facial expressions, the facial rotation matrix and translation vector for representing the facial pose, and the lighting parameters for describing the facial lighting conditions.

[0087] In addition, the objective function information may include projection error, photometric consistency error, shape and regularization constraint information, etc. Among them, the objective function of the projection error (Landmark Projection Loss) can be expressed as Formula 2 below:

[0088] (2)

[0089] Among them, is the camera intrinsic matrix; , is the 3D vertex corresponding to the target 3D face; , is the two-dimensional key point detected in the target expression image (such as sparse corner points such as the corners of the eyes and mouth). Among them, can be the index of the two-dimensional key point (Landmarks), ; among them, can be the number of two-dimensional key points of the preset grid, which can be specifically determined according to different application systems of dense grid corner points, such as 146 key points of the Mediapipe standard or 68 key points of Dlib.

[0090] The objective function of the photometric consistency error (Photometric Loss) can be expressed as Formula 3 below:

[0091] (3)

[0092] Among them, is the face region in the target expression image; is the pixel color in the input target expression image; is the pixel color generated by the rendering model (such as the face shape, lighting parameters , material, and camera parameters based on 3DMM); can be represented as the pixel position of the target expression image. can be the weight of this objective function , such as =0.05.

[0093] The regularization constraint information may include regularization terms for face shape and face expression. Among them, can be a face shape coefficient constraint term used to constrain the magnitude of the shape coefficient to prevent overfitting; can be a face expression coefficient constraint term for constraining the magnitude of the expression coefficient; can be an additional regularization term (such as the sparsity of the expression coefficient); can be the weight of the additional regularization term , such as = 1.0.

[0094] Therefore, by means of the construction of the above-mentioned dense face reconstruction task, it can lay a foundation for the subsequent reconstruction of the target three-dimensional face based on the standardized average face, thereby replacing the traditional sparse key-point-based emotion perception scheme and using the three-dimensional deformation model as the means of emotion representation in the perception stage to ensure that the expression movement of the human face can be decomposed into movements of pose, shape, and expression, and ensuring that the reconstructed target three-dimensional face can achieve a refined description of human face expressions.

[0095] As Figures 2A - 5C shown, according to an embodiment of the present invention, in the operation of performing joint optimization on the dense face reconstruction task in S202 to generate three-dimensional face reconstruction information, it includes: generating an initial three-dimensional face corresponding to the dense face reconstruction task; and performing joint optimization processing on the dense face reconstruction task according to the initial three-dimensional face to generate three-dimensional face reconstruction information.

[0096] When constructing the initial three-dimensional face based on the above-mentioned standardized average face, the initial three-dimensional face can be a standardized three-dimensional face obtained by performing initial lighting condition processing on the standardized average face, which can be used as the basis for constructing the target three-dimensional face to ensure that the target three-dimensional face can accurately represent the human face expression information of the target expression image.

[0097] The three-dimensional face reconstruction information is used to construct the target three-dimensional face based on the initial three-dimensional face, and it can be mainly generated according to the result of the joint optimization processing of the dense face reconstruction task, which may include basic information for three-dimensional face reconstruction such as the human face expression basis and the human face expression coefficient.

[0098] As shown in the foregoing formula 1, the process of joint optimization processing of the dense face reconstruction task can be expressed as a process of solving this formula 1. By solving the joint optimization problem constructed according to the preset joint optimization information on the basis of the above-mentioned initial three-dimensional face and expression movement information, the above-mentioned human face expression basis and the human face expression coefficient are solved, thereby enabling the acquisition of three-dimensional face reconstruction information.

[0099] Therefore, by means of the acquisition of this three-dimensional face reconstruction information, it can ensure the accurate modeling of the human face geometry and expression movement.

[0100] As Figures 2A - 5CAs shown, according to an embodiment of the present invention, in generating an initial three-dimensional human face corresponding to a dense human face reconstruction task, it includes: detecting two-dimensional key point information in a target expression image through a preset key point detection model; obtaining initial pose parameters of a standardized average face according to the two-dimensional key point information.

[0101] The preset key point detection model can be a pre-trained model for detecting two-dimensional human face key points of a target expression image, such as the MediaPipe model or the Dlib model, etc. Among them, the target expression image can represent human face expressions based on the topological structure of a two-dimensional human face. Among them, the two-dimensional key point information can be the two-dimensional information expression of sparse corner points (such as eyebrows, eyelids, etc.) on the topological structure of the two-dimensional human face, and specifically can be represented.

[0102] The two-dimensional key point information (face mesh) can include the position information (such as two-dimensional coordinates) of the above two-dimensional key points in the two-dimensional topological structure. With the help of this position information, the initial pose parameters of the human face pose combined with expression motion information can be estimated. Among them, the initial pose parameters can include the human face pose parameters of the target expression image of the current frame relative to the human face of the standardized average face, such as the rotation matrix of the human face pose relative to the standardized average face and the translation vector .

[0103] Therefore, by obtaining the initial pose parameters, the mapping relationship between the two-dimensional human face topology and the standardized average face can be effectively established, so that the extracted two-dimensional key point information can be accurately mapped to the standardized average face, which is beneficial to more comprehensively and accurately representing the subtle changes and complex characteristics of human face expressions in the subsequent process.

[0104] As Figures 2A - 5C shown, according to an embodiment of the present invention, in generating an initial three-dimensional human face corresponding to a dense human face reconstruction task, it further includes: initializing the optimization variable parameters of the dense human face reconstruction task based on the initial pose parameters; generating an initial three-dimensional human face through the standardized average face and the default lighting conditions based on the initialization result of the optimization variable parameters.

[0105] Through the rotation matrix and the translation vector of the human face pose of the standardized average face corresponding to the initial pose parameters, the optimization variable parameters corresponding to the dense human face reconstruction task can be initialized. Among them, the optimization variable parameters can include the human face shape coefficient , the human face expression coefficient and the lighting parameters , etc. Among them, the above parameter initialization is mainly used to initialize the above human face shape coefficient , the human face expression coefficient and lighting parameters are initialized to zero vectors.

[0106] Therefore, the initialization result of the above optimization variable parameters is the face shape coefficient of the zero vector , face expression coefficient and lighting parameters . According to the face shape coefficient of the above zero vector , face expression coefficient and lighting parameters , an initial 3D face is constructed based on the default lighting condition on the basis of the normalized average face. Among them, the default lighting condition can be used as the default lighting data for the normalized average face during the real-time face alignment process, which can ensure the detection and alignment accuracy of the generated initial 3D face and exclude the influence of lighting condition changes. The default lighting condition can be the lighting information of a point light source that is consistent with the camera center. When obtaining the above target expression image, it is generated according to the lighting information of the camera or the lighting condition is estimated according to the target expression image.

[0107] Therefore, it can be ensured that the face expression of the target expression image can be normalized to the normalized average face, thereby eliminating the coupling problem between the expression semantics and the individual face shape, and ensuring the consistency of the expression semantics among different face shape individuals.

[0108] As Figures 2A - 5C shown, according to an embodiment of the present invention, in operation S203, joint optimization processing is performed on the dense face reconstruction task according to the initial 3D face to generate 3D face reconstruction information, including: by minimizing the projection error, photometric consistency error, and regularization constraint in the dense face reconstruction task through iterative optimization to generate the 3D face reconstruction information after convergence.

[0109] The joint optimization process of the dense face reconstruction task can use iterative optimization means (such as the Levenberg-Marquardt algorithm or the gradient descent scheme) to optimize the objective function of the projection error in formula 1 defined by the above joint optimization problem , the objective function of the photometric consistency error and the regularization constraint information (such as the face shape coefficient constraint term , the face expression coefficient constraint term , the additional regularization term , etc.) for minimizing iterative optimization processing, so that the final optimization process can converge to meet the following three loss conditions in turn: (1) Key point alignment based on the objective function of the projection error , such as the alignment of the normalized average face and the 2D detected face, to ensure the minimum difference in face key points between the two; (2) The objective function based on the photometric consistency error Optimization of photometric consistency. For each pixel, the illumination direction parameters of the normalized average face and the two-dimensional detected face can conform to an illumination condition model; (3) The regularization term controls the stability and physical rationality of the optimization. For example, the face shape coefficient and the face expression coefficient are minimized.

[0110] Through this, the joint optimization of the above-mentioned dense face reconstruction task can be achieved, thereby ensuring that the three-dimensional face reconstruction of the normalized average face can be performed on the input target expression image through the three-dimensional deformation model.

[0111] The target three-dimensional face can be topologically constructed based on the three-dimensional deformation model. The three-dimensional deformation model can represent any three-dimensional face as a combination of linear bases by learning a large amount of three-dimensional face data. Specifically, the three-dimensional deformation model can be expressed as the following formula 4:

[0112] (4)

[0113] Among them, the three-dimensional face shape of the target three-dimensional face can be expressed as , is the number of mesh vertices of the three-dimensional face topology structure, is the real number space. Further, can be used to represent the normalized average face (i.e., the three-dimensional average face, Mean Shape), representing the average shape of the face obtained by statistical learning. represents the face shape basis (Identity Basis), which can be used to represent the principal component basis of individual differences in the face, and s is the number of shape bases. represents the face expression basis (Expression Basis), which can be used to describe the dynamic expression changes of the face, is the number of expression bases. represents the face shape coefficients (Identity Coefficients), which can represent the facial features of a specific individual. represents the face expression coefficients (Expression Coefficients), which can be used as a parametric representation of the face expression.

[0114] Through the above joint optimization processing for the dense face reconstruction task, the face expression basis and the face expression coefficients of the three-dimensional face reconstruction information can be generated.

[0115] For the face expression basis and the face shape basis Based on the linear combination representation of Formula 4 of the above three-dimensional deformation model, it is possible to achieve a dense expression motion representation of any face shape, thus constituting the target three-dimensional face.

[0116] Therefore, through the construction process of the target three-dimensional face, the three-dimensional face reconstruction of the standardized average face can be performed on the input target expression image by the three-dimensional deformation model, realizing the decomposition of the three subspaces of the pose, face shape, and expression of the face expression motion, corresponding respectively to the pose subspace for representing the global rotation and translation of the face expression, the face shape subspace for representing the individual face shape features, and the expression subspace for representing the dynamic expression changes of the face. Thus, the face expression of a single-frame two-dimensional expression image is mapped to a parameterized three-dimensional representation. With the help of the decomposition ability of the face shape, expression, and pose subspaces, the effective decoupling of the individual face shape features and expression dynamics is realized, ensuring the significant improvement of the capture ability of subtle expressions through dense expression representation, thereby overcoming the limitations of the traditional sparse feature key point scheme.

[0117] Such as Figures 2A - 5C shown, according to an embodiment of the present invention, in operation S204, extracting the dense expression motion information of the standardized average face from the target three-dimensional face to complete the perception optimization includes: optimizing the shape parameters of the target three-dimensional face to complete the extraction of the dense expression motion information.

[0118] The three-dimensional face reconstruction of a single-frame target expression image can be realized based on methods such as global photometric consistency optimization and end-to-end reconstruction schemes based on deep learning (such as using CNN and Transformer models). Specifically, as described above for the face expression basis and face shape basis of the three-dimensional deformation model, it is possible to recover the dense and high-precision three-dimensional face shape and its dynamic expression parameters, thus realizing the precise modeling of the face geometry and expression motion.

[0119] According to the above joint optimization results of the face expression basis and face expression coefficients for the dense face reconstruction task of Formula 1, the dense expression motion information can be expressed based on Formula 4 of the above three-dimensional deformation model, specifically as the following Formula 5:

[0120] (5)

[0121] Furthermore, the three-dimensional face shape of the above target three-dimensional face can be further expressed as the following Formula 6:

[0122] (6)

[0123] Therefore, after the reconstruction of the three-dimensional face shape of the target three-dimensional face is completed, the generated three-dimensional face shape can include an individualized face shape and facial expression motion information These two parts can be specifically expressed as Formula 4 above.

[0124] To eliminate the influence of individual facial shape features and ensure the consistency of expression semantics, the facial shape coefficient can be set to zero, only retaining the part related to expressions, that is, the optimization of the shape parameters of the target 3D face is completed, and the 3D face shape based on the standardized average face corresponding to the above Formula 4 is generated . This 3D face shape can be used to represent the dense expression grid topology and define the vertex positions of the dense expression grid. Its specific expression is as Formula 7 below:

[0125] (7)

[0126] where is the facial shape of the standardized average face; then describes the dense expression motion information.

[0127] Therefore, the generated 3D face shape is the 3D expression shape under the 3D standardized average face . It has eliminated individual differences and only retained dynamic expression features, so it can be used as a basis representation with consistent semantic mapping for subsequent expression mapping. Among them, the dense expression motion information is represented by , so it can be expressed as the following Formula 8 as the result of dense expression extraction:

[0128] (8)

[0129] where the dense expression motion information is a dense 3D vertex motion field, which can be used to represent the expression offset of each vertex. The significance of extracting the dense expression motion information is that it provides a fine representation at the vertex level for the entire expression motion, can reflect extremely subtle expression changes, and at the same time eliminates the interference of individual shapes, providing high-precision input for subsequent expression operations.

[0130] Therefore, in the above-mentioned standardization process and dense expression extraction process, in order to eliminate the influence of individual face shapes on expression description, the shape parameters are fixed to zero in the post-processing stage, and only the expression and pose parameters are retained, thereby generating a dense expression motion under the standardized average face. This process ensures the unity of expression semantics. Therefore, a high-resolution dense expression motion (Dense Expression Motion) information can be generated through the linear combination of the face expression basis and the face expression coefficients of the three-dimensional deformation model, so as to provide accurate and standardized expression descriptions for the subsequent mapping and execution stages.

[0131] Furthermore, in the perception stage, the face motion is decomposed into three independent subspaces of pose, face shape, and expression through the three-dimensional deformation model, generating a standardized and dense expression motion representation, which not only improves the refinement degree of expression description, but also eliminates the coupling problem between expression semantics and individual face shapes through the standardized average face topology, achieving the consistency of expression semantics among different individuals.

[0132] In summary, based on the above-mentioned perception optimization method for agent expression imitation in the embodiments of the present invention, in the perception stage of expression imitation, the three-dimensional deformation model can be used as the core technical means to replace the traditional sparse key-point-based expression representation method. By decomposing the face motion into three independent subspaces of pose, face shape, and expression, a refined description of the face expression is achieved. In specific implementation, the above-mentioned three-dimensional face reconstruction information is further calculated for a single-frame image, and then the shape parameters are fixed in the post-processing stage, and only the expression and pose parameters are retained, thereby generating a dense expression motion information under the standardized average face topology. Therefore, it can significantly overcome the abstract defects of the traditional sparse key-point method, realize the densification of expression description, and at the same time eliminate the coupling problem between expression semantics and individual face shapes by normalizing the expression to the standardized average face, ensuring the consistency of expression semantics among individuals with different face shapes.

[0133] Compared with the traditional coefficient key-point method, the high accuracy and density of the dense expression motion representation of the three-dimensional deformation model in the embodiments of the present invention can significantly improve the accuracy and fineness of expression perception, laying a high-quality foundation for the subsequent mapping and execution stages. Specifically, in the perception stage of expression, the face expression in the two-dimensional video can be mapped into a parameterized three-dimensional representation by using the three-dimensional deformable model. With its decomposition ability for the face shape, expression, and pose subspaces, the effective decoupling of individual face shape features and expression dynamics is realized. This method significantly improves the ability to capture subtle expressions through dense expression representation, overcoming the limitations of the traditional sparse feature point method.

[0134] Based on the above-mentioned perception optimization method for agent expression imitation, the present invention also provides a perception optimization device for agent expression imitation. The following will be combined withFigure 2B Describe the device in detail.

[0135] Figure 2B The structural block diagram of the perception optimization device for agent expression imitation according to an embodiment of the present invention is schematically shown.

[0136] As Figure 2B shown, the perception optimization device 200 for agent expression imitation in this embodiment includes a motion extraction module 210, a task generation module 220, a joint optimization module 230, and an information extraction module 240.

[0137] The motion extraction module 210 is used to extract the expression motion information of the target expression image. In one embodiment, the motion extraction module 210 can be used to perform the operation S201 described above, which will not be elaborated here.

[0138] The task generation module 220 is used to generate a dense face reconstruction task through the expression motion information and the preset joint optimization information. In one embodiment, the task generation module 220 can be used to perform the operation S202 described above, which will not be elaborated here.

[0139] The joint optimization module 230 is used to perform joint optimization on the dense face reconstruction task to generate three-dimensional face reconstruction information, and the three-dimensional face reconstruction information is used to generate the target three-dimensional face. In one embodiment, the joint optimization module 230 can be used to perform the operation S203 described above, which will not be elaborated here.

[0140] The information extraction module 240 is used to extract the dense expression motion information of the standardized average face according to the target three-dimensional face to complete the perception optimization. In one embodiment, the information extraction module 240 can be used to perform the operation S204 described above, which will not be elaborated here.

[0141] According to an embodiment of the present invention, any plurality of modules among the motion extraction module 210, the task generation module 220, the joint optimization module 230, and the information extraction module 240 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the motion extraction module 210, the task generation module 220, the joint optimization module 230, and the information extraction module 240 may be at least partially implemented as a hardware circuit, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented by any other reasonable means such as hardware or firmware through circuit integration or packaging, or may be implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the motion extraction module 210, the task generation module 220, the joint optimization module 230, and the information extraction module 240 may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions may be executed.

[0142] Embodiment 2

[0143] In the expression mapping stage, existing methods mainly rely on two technical paths: geometric displacement migration based on sparse key points and expression semantic mapping based on image generation. Among them, the method of using geometric displacement migration of sparse key points attempts to directly map the key point displacements of the human face to the robot face. However, due to the significant differences in topological structure and geometric characteristics between the human face and the robot face, simple displacement copying or regularization methods based on linear interpolation, affine transformation, etc. cannot effectively transfer expression semantics, and may even cause expression features to be mis-mapped or lost, resulting in serious semantic distortion. In some cases, this method cannot migrate specific expressions at all. Another image generation method based on a generative model can visually generate images that conform to expression semantics, but these images only achieve expression consistency in the pixel space and do not consider the physical constraints of the robot actuators, so they cannot be directly used to drive actual physical actuators.

[0144] Therefore, in view of at least one of the technical problems existing in the above-mentioned prior art, embodiments of the present invention aim to provide a mapping optimization method, device, equipment, medium and product for agent expression imitation, thereby providing an optimization solution for the mapping stage of agent expression imitation, with the expectation of achieving a cross-modal mapping mechanism from standardized human facial expressions to the robot expression space by introducing a unified expression parameter space, solving the differences in geometric structures between the human face and the robot face, and ensuring the consistency of expression semantics during the migration process.

[0145] Based on the Figure 1 described scenario, through Figures 3A - 5C a detailed description of the mapping optimization method for agent expression imitation of the disclosed embodiments will be given.

[0146] As Figure 3A shown, one aspect of the embodiments of the present invention provides a mapping optimization method for agent expression imitation, which includes operations S301 to S303.

[0147] In operation S301, a first blend shape basis of a standardized average face is generated based on a preset facial action coding rule;

[0148] In operation S302, the dense expression motion information of the standardized average face is mapped to the first blend shape basis to generate a second blend shape basis; and

[0149] In operation S303, the blend shape coefficients of the second blend shape basis are transferred to the third blend shape basis of the agent's face to complete the mapping optimization.

[0150] The preset facial action coding rule can be a coding rule for facial behaviors formed by the facial muscle movement state, so as to realize the coding of facial expressions, which can improve the processing accuracy and efficiency of facial expressions during the expression recognition and processing process. For example, the Facial Action Coding System (FACS) can be used to implement it.

[0151] The standardized average face is the three-dimensional average face shape of the static face under the standardized human face topology, which can usually be obtained based on big data statistics. Specifically, reference can be made to the standardized average face in Embodiment 1 .

[0152] Based on the face topology of the standardized average face, semantic processing of the expression shape offset is performed through preset facial action coding rules, thereby forming the first blend shape basis. Among them, the first blend shape basis can describe the vertex offset of the corresponding semantic expression (such as "open mouth" or "blink left eye", etc.) relative to the standardized average face, and can be used to describe the expression change through the weighted combination of predefined semantic expression bases. Specifically, it can be implemented through the Blendshape expression representation technology.

[0153] The dense expression motion information can be information directly related to the dynamic change of the facial expression related to the extracted expression motion information, and can be specifically characterized by the linear combination of the facial expression basis and the facial expression coefficient. The dense expression motion information can most directly reflect the dense expression motion information of the user extracted by the agent on the standardized average face, realizing facial expression representation. Specifically, reference can be made to the extraction of dense expression motion information in Embodiment 1.

[0154] Mapping the dense expression motion information onto the first blend shape basis can achieve mapping the expression motion information of the face of each frame of the target expression image extracted in the perception stage in Embodiment 1 onto the standardized average face. At the same time, the second blend shape basis generated thereby can also exclude the influence of individual face shape features, ensuring the expression semantic consistency in the mapping stage and the execution stage, thereby completing the dense representation of the expression semantics. The second blend shape basis can be a semantic expression basis expressed by the dense expression motion information obtained through perception, and can be used to describe the vertex offset of the semantic expression corresponding to the dense expression motion information relative to the standardized average face. Specifically, it can be implemented through the Blendshape expression representation technology.

[0155] Correspondingly, the third blend shape basis is a semantic expression basis obtained by performing semantic processing of the expression shape offset based on the agent's face topology, and can be used to describe the vertex offset of the corresponding semantic expression relative to the agent's face topology. Among them, the semantics of the third blend shape basis correspond one-to-one and are consistent with the semantics of the above-mentioned second blend shape basis.

[0156] The blend shape coefficient can be understood as the blend shape weight (such as the Blendshape coefficient) of the second blend shape basis, and can be used to express the vertex offset of the expression action of the dense expression motion information obtained in the perception stage (refer to Embodiment 1 above). By transferring the blend shape coefficient to the third blend shape basis of the agent's face, direct cross-topology migration of human dense expression motion information can be achieved.

[0157] Therefore, in the mapping stage, the above mapping optimization method for agent expression imitation in the embodiments of the present invention uses a hybrid shape semantic space based on a set of second hybrid shape bases of a standardized average face as an intermediate bridge to solve the differences in topological structure and geometric characteristics between the human face and the agent's face. Specifically, first, a set of first hybrid shape bases is defined on the standardized average face for dense representation of expression semantics. The first hybrid shape bases can then map the dense expression movements generated in the perception stage (refer to the above Embodiment 1) to the hybrid shape coefficient space through dense semantic key points, forming the second hybrid shape bases. In this process, by transferring the hybrid shape coefficients, the expression movements are mapped from the topological structure of the standardized average face to the topological structure of the agent's face, realizing semantic-consistent expression transfer across topological structures.

[0158] Therefore, through this conversion from the geometric domain to the semantic domain, the problem of semantic distortion caused by the topological differences between the human face and the robot's face in the traditional solution is solved. In addition, the number of hybrid shape coefficients is limited, which is convenient for real-time transmission, thus greatly improving the efficiency and robustness of the mapping stage. At the same time, the generated hybrid shape coefficients directly correspond to the motion state of the robot's epidermis (such as Figure 5C the shown skin Skin), ensuring the physical executability of the generated expressions. It can be seen that in the cross-entity mapping process of expression imitation, the hybrid shape bases designed based on the above preset facial action coding rules can be used to accurately depict expression semantics. By introducing a unified expression parameter space, a cross-modal mapping mechanism from the dense human face expressions of the standardized average face to the robot expression space can be established, thus solving the differences in geometric structures between the human face and the robot's face and ensuring the consistency of expression semantics during the migration process.

[0159] The hybrid shape bases are mainly generated based on semantic expression modeling of the Blendshape technology. After the standardization processing and extraction of dense expression motion information in the above Embodiment 1, in order to further realize semantic expression modeling, an expression representation method based on this hybrid shape base is further introduced.

[0160] The Blendshape technology can be understood as describing expression changes through the weighted combination of a set of predefined semantic expression bases. Therefore, the three-dimensional human face shape corresponding to the standardized average face represented by the hybrid shape base can be expressed as the following formula 9:

[0161] (9)

[0162] Wherein, can be expressed as a dense grid containing semantic expressions. It can be a three-dimensional average face shape, which is consistent with the definition of the standardized average face in Embodiment 1, representing the static face shape under the standardized human face topology. It can be the nth Human Blendshape basis, which is used to represent the semantic expression shape offset and can describe the vertex offset of the corresponding semantic expression (such as "open mouth" or "blink left eye") relative to the standardized average face. It can be the weight of the nth Human Blendshape basis, which can represent the intensity coefficient of this expression. Among them, the weight of this Human Blendshape basis can usually be normalized to the interval For example, indicates no such expression, while indicates fully applying this expression. n is the number of defined blendshape bases, and each blendshape basis can correspond to a semantic expression action.

[0163] Combined with the mathematical representation of the blendshape basis in Formula 9 above, it can be combined with Figures 3A - 5C to further describe the generation of the first blendshape basis in Embodiment 2 of the present invention as follows.

[0164] As Figures 3A - 5C shown, according to an embodiment of the present invention, in generating the first blendshape basis of the standardized average face based on the preset facial action coding rule in Operation 301, it includes:

[0165] Generating a blendshape expression basis matrix according to the preset facial action coding rule;

[0166] Generating the first blendshape basis of the standardized average face based on the preset blendshape weight vector and the blendshape expression basis matrix.

[0167] As mentioned above, the first blendshape basis can be designed based on the three-dimensional average face shape (i.e., the standardized average face ), which can have the same topology (connection relationship between vertices and faces) as the standardized average face, and combine the expression semantics of the preset facial action coding rule (such as the Facial Action Coding System, abbreviated as FACS), such as "open mouth" or "blink left eye".

[0168] As defined in the above Formula 9, the Human Blendshape basis that is a component of the first blendshape basis can be designed by a predefined method, specifically based on the standardized average face Construct the facial topology. Among them, the first blend shape basis is designed according to the expression semantics of the preset facial action coding rules. It can be divided according to the dominant area to ensure that the semantics of each blend shape basis are clear and consistent. For example, the facial expression action of "the right corner of the mouth is raised" can correspond to the movement of an actuator in FACS.

[0169] Specifically, the three-dimensional human face deformation represented by the first blend shape basis can be defined on the basis of the standardized average face and generated by the weighted combination of a set of predefined semantic expression bases. Combining the above formula 9, specifically, the first blend shape basis Can be expressed in matrix form in the following formula 10 to form an expression combination based on the blend shape basis:

[0170] (10)

[0171] Among them, the blend shape expression basis matrix Can form the semantic human face blend shape basis matrix of the first blend shape basis, which can specifically be composed of Predefined human face expression bases Composed. In addition, the blend shape weight set Can be correspondingly expressed as the preset weight vector of the blend shape corresponding to the blend shape expression basis matrix to represent the intensity of each expression. Therefore, through this preset blend shape weight vector And the blend shape expression basis matrix Can form the first blend shape basis corresponding to the standardized average face .

[0172] Therefore, by adjusting the blend shape weight , that is, the contribution of each component blend shape basis in the first blend shape basis can be flexibly controlled, so as to generate expression shapes with different intensities and combinations. This predefined modeling method can ensure that the generated expression shapes have clear semantics, dense geometric details and a topological structure consistent with the average face.

[0173] It can be seen that through the semantic expression modeling of the blend shape basis, on the facial topology (Mesh) of the standardized average face, a set of blend shape bases corresponding to its dense expression movements can be designed based on the facial action coding system for the compact representation of expression semantics (such as 51 standard Blendshapes of ARKit).

[0174] Such as Figures 3A - 5C Shown, according to an embodiment of the present invention, before mapping the dense expression movement information of the standardized average face to the first blend shape basis to generate the second blend shape basis in operation 302, it further includes:

[0175] Generate the dense expression motion information of the standardized average face based on the expression motion information of the extracted target expression image.

[0176] As mentioned in Embodiment 1, the target expression image can be a single-frame image containing a human facial expression, and the expression motion information can include the facial expression feature parameters extracted for the facial expression in the target expression image. Correspondingly, the dense expression motion information can be the information related to the dynamic changes of the facial expression directly related to the extracted expression motion information, and can be specifically characterized by the linear combination of the facial expression basis and the facial expression coefficients. See Formula 5 in Embodiment 1 for details. The dense expression motion information can most directly reflect the dense expression motion information of the user extracted by the agent on the standardized average face, realizing facial expression representation.

[0177] Therefore, based on the extraction of the dense expression motion, a facial expression grid on the standardized average face corresponding to the dense expression motion information can be constructed. (as shown in Formula 7 in Embodiment 1), this facial expression grid can be used to represent the dense expression grid topological structure as a three-dimensional human face shape with dense expression motion information, and define the vertex positions of the dense expression grid.

[0178] By means of the above extraction process of the dense expression motion information in the perception stage, the effective decoupling of the individual facial shape features and the expression dynamic information can be realized, and the ability to capture subtle expressions can be significantly improved through dense expression representation.

[0179] In the expression imitation mapping stage of Embodiment 2 of the present invention, in order to transfer the dense expression motion generated in the perception stage of Embodiment 1 to the topological structure of the agent's face, realize cross-topological semantic transfer, and ensure cross-topological semantic consistency, it is necessary to consider the mapping from the dense expression motion to the blend shape coefficients.

[0180] As Figures 3A - 5C shown, according to an embodiment of the present invention, in the operation of mapping the dense expression motion information of the standardized average face to the first blend shape basis in Operation 302 to generate the second blend shape basis, it includes:

[0181] Extract the first dense annotation points on the dense expression grid corresponding to the dense expression motion information and the second dense annotation points on the average face grid;

[0182] Obtain the dense annotation basis matrix according to the blend shape expression basis matrix;

[0183] Generate a blend shape weight optimization task according to the dense annotation basis matrix, the first dense annotation points, and the second dense annotation points.

[0184] The dense expression grid can be understood as a face expression grid based on the standardized average face corresponding to the above-mentioned dense expression motion information. (as shown in Equation 7 in Embodiment 1), which is used to represent the vertex positions of the dense expression grid after standardization, and contains vertices, and each vertex has three-dimensional coordinates (with a total of degrees of freedom), and can be understood as a three-dimensional geometric representation of the entire facial network.

[0185] The average face grid can be understood as the standardized face grid of the standardized average face and can represent the three-dimensional vertex positions of the corresponding face grid.

[0186] The first dense annotation point can be expressed as the position of the standardized dense annotation point , which is usually a specific set of points selected from the above-mentioned dense expression grid in dense expression modeling, for example, distributed in key areas of the face (such as the mouth, eyes, and nose, etc.), and can specifically be expressed as the positions of selected from the dense expression grid dense annotation points (with a total of degrees of freedom).

[0187] Among them, the position of the first dense annotation point can be extracted from the dense expression grid through the dense annotation matrix and can specifically be expressed as Equation 11 below:

[0188] (11)

[0189] where the dense annotation matrix is a sparse selection matrix that can be used to extract the first dense annotation point of GIA from the dense expression grid .

[0190] Correspondingly, the second dense annotation point can be expressed as the average face position of the dense annotation point , which is defined as the dense annotation point extracted from the average face grid through the dense annotation matrix and can specifically be expressed as Equation 12 below:

[0191] (12)

[0192] where the average face grid can represent the three-dimensional vertex positions of the face grid (Mesh) defined by the standardized average face.

[0193] Therefore, with the help of the above formulas (11) and (12), the correspondence between the dense annotation points and the face mesh can be constructed, where the position of the first dense annotation point and the position of the second dense annotation point are directly associated with the dense expression mesh (full mesh expression) and the average face mesh .

[0194] As mentioned above, the mixed shape expression basis matrix can be used as the full mesh basis matrix to form the semantic face mixed shape basis matrix of the first mixed shape basis, which can be specifically composed of K predefined expression bases.

[0195] The dense annotation basis matrix can be understood as the mixed shape basis matrix of the dense annotation , which is mainly used to represent the mixed shape basis on the dense annotation points, and can be specifically extracted from the full mesh basis matrix , and can be specifically expressed as the following formula 13:[[]]

[0196] (13)

[0197] Therefore, according to the dense annotation basis matrix defined by the above formula 13 , the first dense annotation point defined by formula 11 and the second dense annotation point defined by formula 12 , the mixed shape weight corresponding to the above mixed shape weight optimization task can be defined as the optimization problem represented by formula 14 as follows:[[]]

[0198] (14)

[0199] Wherein,[[]] can be understood as the mixed shape weight, which can make the position of the reconstructed dense annotation point satisfy the relationship defined by the following formula 15:[[]]

[0200] (15)

[0201] Therefore, based on the mixed shape technology modeling, the geometric information of the dense annotation points can be utilized to solve the optimization problem corresponding to the mixed shape weight optimization task of the above formula 14, and the mixed shape weight can be obtained, so as to ensure that the position of the reconstructed dense annotation point is as close as possible to the input dense annotation point.

[0202] As Figures 3A - 5CAs shown, according to an embodiment of the present invention, in operation 302 where the dense expression motion information of the normalized average face is mapped onto the first blend shape basis to generate the second blend shape basis, it further includes:

[0203] Obtain the blend shape coefficients corresponding to the blend shape weight optimization task by minimizing the preset error vector and the preset coefficient constraint conditions, so as to map the dense expression motion information of the normalized average face onto the first blend shape basis to generate the second blend shape basis.

[0204] The preset error vector can be the error vector objective function for performing optimization processing on the above blend shape weight optimization task , and specifically can be expressed as the following formula 16:

[0205] (16)

[0206] The preset error vector can be the first dense annotation point and the second dense annotation point Based on the dense annotation basis matrix of the error vector, by setting the sum of the squares of the minimized errors of the preset error vector as the optimization objective, the following formula 17 can be defined:

[0207] (17)

[0208] Expanding the above formula 17, it can be expressed as the following formula 18:

[0209] (18)

[0210] The preset coefficient constraint conditions can be used as the constraint conditions for performing optimization processing on the above blend shape weight optimization task. For example, the value of the blend shape weight can be limited within the range of , so as to satisfy that the blend shape coefficients corresponding to the final blend shape weight optimization task are . Therefore, the above blend shape weight optimization task corresponds to a quadratic programming problem (quadratic programming) with boundary constraints (preset coefficient constraint conditions).

[0211] Therefore, by means of the solution process of the least squares problem of the above blend shape coefficients, the dense expression motion information can be mapped onto the low-dimensional first blend shape basis in real time to form the second blend shape basis, that is, by selecting dense semantic key points (based on barycentric coordinates, rather than vertex coordinates) on the surface of the average face mesh, the dense expression motion generated in the perception stage is mapped into the blend shape coefficient space, thereby completing the coefficient mapping from the dense expression to the blend shape space.

[0212] In summary, through the least squares problem of an overdetermined constraint in the above-mentioned mixed shape weight optimization task, the dense expression motion information can be compressed into a low-dimensional Blendshape space. This process can generate blendshape coefficients in real time while ensuring the density and accuracy of expressions.

[0213] As Figures 3A - 5C shown, according to an embodiment of the present invention, before transferring the blendshape coefficients of the second blendshape basis to the third blendshape basis of the agent's face in operation 303 to complete the mapping optimization, it further includes:

[0214] Generating a third blendshape basis of the agent's face based on a preset facial action encoding rule, where the third blendshape basis has the same blendshape semantics as the first blendshape basis.

[0215] The third blendshape basis is actually built for the blendshape space of the agent's face. Among them, the agent's blendshape space needs to ensure semantic consistency with the first blendshape basis space corresponding to the dense expression grid of the human face so that the agent can accurately express the same basic expression semantics as humans (such as "opening the mouth", "blinking", etc.).

[0216] Similar to the human first blendshape basis, according to the semantic consistency principle, the third blendshape basis space of each agent corresponds one-to-one and is consistent with the semantic of the first blendshape basis space of the standardized average face (for example, "opening the mouth" should express the same semantic on the facial topology of the agent and in the blendshape basis space of the human standardized average face model).

[0217] Therefore, according to the semantic set of the first blendshape basis space of the human face (such as "opening the mouth", "blinking the left eye", etc.), based on the above-mentioned preset facial action encoding rule, a corresponding set of third blendshape bases can also be defined for the agent based on the facial topology of the agent, and the semantic set in the third blendshape basis space is , and its blendshape basis for each semantic of the agent where is the number of mesh vertices in the facial topology of the agent, can be the semantic number.

[0218] Therefore, for the facial topology of the agent, based on the first blendshape basis of the standardized average face corresponding to the dense expression motion information, a third blendshape basis of the agent's face with consistent semantics can be constructed to ensure that the semantic of the two blendshape basis spaces is completely consistent, thereby ensuring efficient and accurate dense expression transfer from the standardized average face to the facial topology of the agent.

[0219] As Figures 3A - 5C shown, according to an embodiment of the present invention, in the operation of transferring the blend shape coefficients of the second blend shape base to the third blend shape base of the agent's face to complete the mapping optimization in operation 303, it includes:

[0220] Applying the blend shape coefficients to the third blend shape base to generate the agent's corresponding agent expression mesh.

[0221] Directly transferring the blend shape coefficients generated from the standardized average face to the third blend shape base corresponding to the agent's face topology can achieve dense expression cross-topology semantic transfer, generate the agent expression mesh, and thus complete the mapping process of the blend shape coefficients.

[0222] Among them, the core of transferring the blend shape coefficients is to directly apply the blend shape coefficients generated from the human standardized average face to the third blend shape base of the agent to achieve dense expression cross-topology semantic transfer.

[0223] Specifically, through the third blend shape base of the agent's face topology with semantic consistency , the semantics of each basic expression are completely aligned with the first blend shape base of the human standardized average face topology. Therefore, the blend shape coefficients of the first blend shape base can be directly transferred to the third blend shape base without any conversion, satisfying the following formula 19:

[0224] (19)

[0225] Therefore, applying the blend shape coefficients of the first blend shape base to the agent's face topology and the agent's third blend shape base can generate the agent's corresponding agent expression mesh expressed as the following formula 20:

[0226] (20)

[0227] Therefore, it can ensure that the semantics of the robot's expression are consistent with the human expression, and at the same time adapt to the geometric characteristics of the agent's face topology. Specifically, refer to the semantic alignment operation from the second blend shape base to the third blend shape base as Figure 5A shown.

[0228] In summary, based on the mapping optimization method for agent expression imitation provided in Embodiment 2 of the present invention, a hybrid shape semantic space based on a standardized average face can be provided as an intermediate bridge during the mapping stage to address the differences in topological structure and geometric characteristics between the human face and the agent's face. Specifically, a set of hybrid shape bases is first defined on the standardized average face for dense representation of expression semantics. These hybrid shape bases map the dense expression movements generated in the perception stage to the hybrid shape coefficient space through dense semantic key points. This process maps the dense expression movements to low-dimensional hybrid shape coefficients in real time by solving a least squares problem.

[0229] After that, a set of hybrid shape bases that are semantically consistent with the average face is provided based on the agent's face to ensure that both have the same semantic definition. By transmitting the hybrid shape coefficients, the dense expression movements can be mapped from the standardized average face to the agent's face, thus achieving efficient and accurate cross-topological structure and semantically consistent expression transfer.

[0230] Therefore, by converting from the geometric domain to the semantic domain, the problem of semantic distortion caused by the topological differences between the human face and the agent's face in traditional methods is solved. In addition, the number of hybrid shape coefficients is limited, which is convenient for real-time transmission and greatly improves the efficiency and robustness of the mapping stage. At the same time, the generated hybrid shape coefficients directly correspond to the motion state of the agent's epidermis, ensuring the physical executability of the generated expressions.

[0231] In the cross-entity mapping of expressions, the hybrid shape bases designed based on the preset facial action coding rules are used to accurately depict expression semantics. By introducing a unified expression parameter space, a cross-modal mapping mechanism from the standardized human face expression to the agent expression space is established, which solves the differences in geometric structure between the human face and the agent's face and ensures the consistency of expression semantics during the transfer process.

[0232] Therefore, the mapping optimization method for agent expression imitation provided in Embodiment 2 of the present invention can at least achieve the following technical effects:

[0233] (1) Directness: No additional conversion or mapping is required for human hybrid shape coefficients .

[0234] (2) Efficiency: Utilizing semantically consistent hybrid shape bases, quickly realizing cross-topological transfer of expressions.

[0235] (3) Versatility: Applicable to any agent hybrid shape bases that meet the semantic consistency design.

[0236] Based on the above mapping optimization method for agent expression imitation, the present invention also provides a mapping optimization device for agent expression imitation. The following will be combined withFigure 3B Describe the device in detail.

[0237] Figure 3B The structural block diagram of the mapping optimization device for agent expression imitation according to an embodiment of the present invention is schematically shown.

[0238] As Figure 3B shown, the mapping optimization device 300 for agent expression imitation in this embodiment includes a basis generation module 310, an information mapping module 320, and a coefficient transfer module 330.

[0239] The basis generation module 310 is used to generate a first blend shape basis of a standardized average face based on a preset facial action coding rule. In one embodiment, the basis generation module 310 can be used to perform the operation S301 described above, which will not be elaborated here.

[0240] The information mapping module 320 is used to map the dense expression motion information of the standardized average face to the first blend shape basis to generate a second blend shape basis. In one embodiment, the information mapping module 320 can be used to perform the operation S302 described above, which will not be elaborated here.

[0241] The coefficient transfer module 330 is used to transfer the blend shape coefficients of the second blend shape basis to a third blend shape basis of the agent's face to complete the mapping optimization. In one embodiment, the coefficient transfer module 330 can be used to perform the operation S303 described above, which will not be elaborated here.

[0242] According to an embodiment of the present invention, any multiple of the basis generation module 310, the information mapping module 320, and the coefficient transfer module 330 can be combined and implemented in one module, or any one of them can be split into multiple modules. Or, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the basis generation module 310, the information mapping module 320, and the coefficient transfer module 330 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable means such as hardware or firmware through circuit integration or packaging, or implemented in any one of the three implementation manners of software, hardware, and firmware or in any appropriate combination of several of them. Or, at least one of the basis generation module 310, the information mapping module 320, and the coefficient transfer module 330 can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions can be executed.

[0243] Embodiment 3

[0244] During the execution phase, traditional methods convert the epidermal motion representation into actuator commands in two ways. One is the model-based inverse kinematics method, which converts the motion of the epidermal traction points into specific actuator commands through inverse kinematics solving. However, this method has extremely high requirements for the accuracy of the geometric model and needs to accurately describe the topological structure of the robot's face and the actuator positions. Otherwise, it is difficult to ensure the accuracy of the inverse kinematics solving. The other method is the model-free neural networks method based on deep learning, which directly calculates the actuator commands from the epidermal motion (such as pixel space) using neural networks. Although this method has higher flexibility, the generated execution commands often do not fully consider the physical constraints of the robot actuators, resulting in difficulties in implementing the commands in practical applications and causing problems with the control accuracy and stability of expressions.

[0245] In view of at least one of the technical problems existing in the above-mentioned execution schemes for expression imitation, embodiments of the present invention provide an execution optimization method, device, equipment, medium, and product for agent expression imitation, aiming to optimize the expression driving parameters by explicitly modeling the physical constraints of the agent actuators and ensure that the generated control commands are both physically executable and meet the requirements of high precision and real-time performance.

[0246] The following will be based on Figure 1 the described scenario, and through Figures 4A - 5C a detailed description of the execution optimization method for agent expression imitation of the disclosed embodiments will be given.

[0247] As Figure 4A shown, one aspect of the embodiments of the present invention provides an execution optimization method for agent expression imitation, which includes operations S401 to S403.

[0248] In operation S401, an inverse kinematics target optimization task is generated based on the control point simulation information of the agent expression grid and the target epidermal vertex information;

[0249] In operation S402, an optimization process is performed on the inverse kinematics target optimization task to obtain the optimal joint angle parameters corresponding to the control point simulation information; and

[0250] In operation S403, the execution optimization of agent expression imitation is performed according to the optimal joint angle parameters.

[0251] The agent expression grid may include the vertex relationships of the facial topology grid (Mesh) structure of the agent mentioned in Embodiment 2 above. Among them, the agent expression grid is the robot facial topology that has completed the cross-topology migration of dense expression motion information in the perception stage in Embodiment 1, and specifically can be embodied as the expression grid in Formula 20 in Embodiment 2. .

[0252] As Figure 5B shown, there will be many key control points on the agent expression grid, such as eyebrows, eyelids, eyeballs, nose, cheeks, mouth, and jaw. These control points can be driven by cables / PTFE conduits as Figure 5C shown through motors (such as LFD-01M servo motors) to achieve pulling, thereby driving the deformation of the epidermis in the surrounding area of the control points (such as Figure 5C shown Silicone Mask). In the traditional technical solution, it is very difficult to solve the deformation of the surrounding area of the control points, resulting in a rigid imitation effect of the agent's expression and unable to achieve a delicate and natural "skin" performance.

[0253] The control point simulation information can be the position information formed by performing control point position simulation calculations on the agent expression grid containing dense expression motion information through a simulation tool (such as Blender), for example, the simulation positions of each control point in the topological network of the agent expression grid with dense expression motion information. The control point simulation information can be used to define the drive angle parameters of the drive motor rocker arm of the cable / PTFE conduit as Figure 5C shown, that is, the joint angle parameters.

[0254] The target epidermis vertex information can be the target epidermis grid vertex positions that each control point on the agent expression grid should reach when achieving the human face expression reflected by the dense expression motion information. The target epidermis vertex information can include the simulation control point positions of the control point simulation information obtained through simulation calculations and the information on the relative displacement distance and displacement direction of the epidermis grid vertices in the surrounding area, and can be used to define the epidermis shape (i.e., the target epidermis shape) that should be achieved when driving the agent facial grid relative to the dense expression motion information.

[0255] The inverse kinematics target optimization task can be an optimization task that continuously optimizes the joint angle parameters corresponding to the control point simulation information based on the target epidermis vertex information. Through the continuous optimization of this inverse kinematics target optimization task, during actual driving execution, according to the optimization result of this inverse kinematics target optimization task, the movement of the intelligent agent's facial epidermis mesh can be made as close as possible to the target epidermis shape defined by the dense expression movement information. This inverse kinematics target optimization task can be expressed in the form of an inverse kinematics objective function.

[0256] Performing the optimization of the above joint angle parameters on the inverse kinematics target optimization task can make the intelligent agent's epidermis mesh as close as possible to the target epidermis shape until the optimal joint angle parameters are obtained. When the optimal joint angle parameters are used as driving parameters by the corresponding driving motors in the intelligent agent's facial structure as shown in Figure 5C it can make the corresponding driven control points and their surrounding epidermis areas present the most fitting expression imitation effect with the dense expression movement information, with a more delicate and natural expression. At the same time, it can ensure that the skin around the control points presents the most natural epidermal changes.

[0257] Therefore, the optimal joint angle parameters can be the motor rocker joint angle parameters that can make the control point and its surrounding epidermis area reach the closest state to the target epidermis deformation when the control point is driven by the motor. As shown in Figure 5C through the "Skeleton" architecture to support the "Skin", "Muscle" architecture of the silicone mask and the driving architecture of the electrical control. The electrical control architecture includes various driving motors or servo motors, and these motors can perform rotational or even pulling actions according to the optimal joint angle parameters under the control of the controller, thereby driving the connected steel cable / Teflon conduit to generate the optimal displacement in the specified direction at the specified position (control point) on the connected silicone mask ("skin"), reaching the optimal position of the control point under the target expression, thereby generating the epidermal deformation of the intelligent agent's face, completing the execution optimization of the expression action, and realizing the natural deformation of the intelligent agent's facial expression. Among them, the control point can be the fixed point (such as adhesive fixation) of the above steel cable / Teflon conduit on the inner surface of the silicone mask. In addition, the above "Skeleton" architecture, "Skin", "Muscle" architecture, and electrical control driving architecture can constitute the actuator of the intelligent agent.

[0258] Among them, since the optimal joint angle parameters are obtained by optimizing the inverse kinematics target optimization task based on the agent's expression grid of dense motion expression information, the execution process of the above expression actions corresponds to the execution process of the agent's expression imitation. In short, the generation of the agent's facial expression depends on the kinematic modeling of the subcutaneous control points. At the same time, the joint angle parameters obtained by inverse kinematics solution are used to drive the movement of the control points, thereby realizing the transfer of the dense expression motion information obtained in the perception stage to the vertices of the agent's facial epidermal grid, ensuring the natural deformation of the expression restoration process.

[0259] Therefore, through the above-mentioned execution optimization method for agent expression imitation in Embodiment 3 of the present invention, by defining the control point area, the inverse kinematics (Inverse Kinematics) can be used to solve the optimal position of the control points under the target expression and generate accurate actuator motion instructions, so that the motion state of the control point area approaches the target epidermal motion. In this case, only by increasing the density of the control points can the fineness and naturalness of the expression restoration be further improved without changing the algorithms in the perception and mapping stages. Thus, not only is the accuracy of expression imitation greatly improved, but also the physical executability of the generated instructions is ensured through explicit modeling of the physical constraints of the actuator, avoiding the execution failure problem caused by ignoring physical constraints in traditional methods. In addition, the independence of control point optimization also enables the execution stage to flexibly adapt to different types of agent facial structures, providing technical guarantees for the scalability and modular design of the system, and having extremely high engineering application value and commercial application value.

[0260] Such as Figures 4A - 5C shown, according to an embodiment of the present invention, before generating the inverse kinematics target optimization task based on the control point simulation information of the agent expression grid and the target epidermal vertex information in operation S401, it further includes:

[0261] Generating an agent expression grid corresponding to the agent according to the blend shape basis of the agent's face;

[0262] Generating a distance weight matrix based on the control point simulation information of the agent expression grid and the target epidermal vertex information.

[0263] The blend shape basis of the agent's face can be a semantic expression space for surface shape offset semantic processing based on the agent's face topology, used to describe the vertex offset of the semantic expression corresponding to the dense expression motion information relative to the agent's face topology. Specifically, it can refer to the third blend shape basis provided in the above-mentioned Embodiment 2 . The dense expression motion information extracted in the perception stage can be transferred to the cross-topology semantics of the above-mentioned blend shape basis of the agent's face through the mapping process of the blend shape coefficients in the mapping stage.

[0264] Furthermore, an agent expression mesh can be constructed based on the blend shape basis of the agent's face , which can be specifically reflected in Formula 20 in the above-mentioned Embodiment 2. In this way, it can ensure that the semantics of the robot's expression are consistent with those of human expressions, and at the same time, it can adapt to the geometric characteristics of the agent's facial topology, ensure that it can be aligned with the semantics of the perceived human face expression, so as to ensure the accuracy of subsequent physical drive execution, guarantee the natural deformation during the expression restoration process, and at the same time ensure a better restoration degree of the perceived human face expression.

[0265] Each component matrix element in the distance weight matrix can be used to represent the distance distribution between a single control point on the agent expression mesh and the corresponding epidermal vertex . Therefore, the distance weight matrix can define the distance-based weight assignment between the epidermal control points and the epidermal mesh vertices. Among them, these matrix elements can be specifically expressed as the following Formula 21:

[0266] (21)

[0267] Among them, is the Euclidean distance between the epidermal vertex and the control point ; is the distance from the epidermal vertex to the control point; is the three-dimensional position coordinate of the subcutaneous controller; is the control parameter of the distance influence, which is used to adjust the attenuation range of the weight; the weight normalization ensures (the sum of the weights of each epidermal vertex is 1), is the number of control points.

[0268] Among them, a single control point can be the position information formed by performing position simulation calculations on the agent expression mesh through a simulation tool, and the set of these control points can constitute the above-mentioned control point simulation information. Among them, each control point can be used to define the three-dimensional coordinates of the subcutaneous traction position. Correspondingly, a single epidermal vertex can be the position expected by the dense expression motion information corresponding to the vertex of the mesh position, and the set of these epidermal vertices can constitute the above-mentioned target epidermal vertex information.

[0269] Therefore, based on the control point simulation information and the target epidermal vertex information, the dense expression motion information of the agent expression mesh can be mapped onto the topological structure of the agent's facial structure, and the accurate description of the agent's facial topology and the actuator position can be realized through the distance weight matrix.

[0270] As Figures 4A - 5C shown, according to an embodiment of the present invention, before generating the distance weight matrix based on the control point simulation information and the target skin vertex information of the agent expression mesh, it further includes:

[0271] Generating control point simulation information according to the preset forward kinematics optimization task of the agent expression mesh and the joint angle parameter vector corresponding to the simulation control points;

[0272] Generating target skin vertex information according to the deformation weight matrix of the skin mesh vertices of the agent expression mesh and the control point simulation information.

[0273] Based on constructing the geometric model and kinematic structure (such as joint control points and their associated relationships) of the agent's face corresponding to the agent expression mesh through a simulation tool, the preset forward kinematics optimization task can be expressed as a forward kinematics function based on the motor kinematic joint angle parameters corresponding to the agent's kinematic structure. The simulation control points can be the subcutaneous control points defined by the above-mentioned geometric model and kinematic structure of the agent's face. The subcutaneous control point set is defined as

[0274] , and each control point represents the three-dimensional coordinates of a certain position under the skin, where is the number of control points. Each simulation control point corresponds to different joint angle parameters, and these joint angle parameters can be used as sub-elements of the above joint angle parameter vector.

[0275] Therefore, through the simulation tool, according to the input parameter vector such as joint angles, etc., combined with the kinematic relationship and geometric constraints, the forward kinematic information of the control point position can be generated as the control point simulation information.

[0276] Specifically, for the implementation operation of the forward kinematics of the above control point simulation information, the geometric model and kinematic structure of the agent's face, including joints, control points and their associated relationships, can be constructed first through a simulator (such as Blender), and then the simulator is used to directly calculate the position of the control points according to the input joint angle parameters , combined with the kinematic relationship and geometric constraints. .

[0277] Furthermore, the control point position corresponding to the control point simulation information is driven by the kinematic joint angle parameter , and the forward kinematic model of the control point simulation information is defined as the following formula 22:

[0278] (22)

[0279] Where is the forward kinematics function; is the joint angle parameter vector; is the number of kinematic degrees of freedom; is the number of control points. The control point simulation information can be used to complete the construction of the control point model.

[0280] The skin mesh vertices of the agent expression mesh can be based on the agent expression mesh The corresponding facial geometry model and the vertices of the facial mesh of the kinematic structure of the agent can be expressed as a set of epidermal mesh vertices: , where each skin mesh vertex Represents the three-dimensional coordinates of a point on the epidermis, is the number of vertices in the skin mesh.

[0281] Correspondingly, the control point model of the control point simulation information The corresponding control point movement can be obtained through the deformation weight matrix Driver corresponding to the skin mesh vertex The deformation weight matrix can represent the influence of the control point on the deformation of the skin mesh vertex. Specifically, the target skin vertex information can be expressed as a linear relationship through the deformation weight matrix and the control point simulation information as follows:

[0282] (twenty three)

[0283] in, is the deformation weight matrix, representing the control points Vertex of epidermis The influence relationship; is the global position of the control points, stacked in order as ; is the global position of the skin mesh vertices, stacked in order as Therefore, the construction of the skin mesh vertex model can be completed. Among them, the deformation weight matrix can be a weight construction model of the distance weight matrix, which is used to generate the distance weight matrix according to the expression mesh of the intelligent body.

[0284] This allows us to accurately define the geometric model of the control points and their surrounding areas, as well as the mesh skin vertices, so that we can subsequently calculate the optimal position of the control points under the target expression, generate accurate actuator motion instructions, and ensure that the motion state of the control point area is close to the target skin motion. Moreover, it is also helpful to further improve the precision and naturalness of expression restoration by increasing the density of control points without changing the algorithms in the perception and mapping stages.

[0285] like Figures 4A - 5CAs shown, in the operation S401 of generating an inverse kinematics target optimization task based on the control point simulation information of the agent expression grid and the target epidermal vertex information, it includes:

[0286] The extended distance weight matrix is a weight amplification matrix;

[0287] Generate control point-driven epidermal grid vertices according to the control point simulation information and the weight amplification matrix;

[0288] Generate an inverse kinematics target optimization task according to the control point-driven epidermal grid vertices and the target epidermal vertex information.

[0289] The weight amplification matrix can be formed by three-dimensionally expanding each sub-element in the aforementioned distance weight matrix. Among them, since each control point and epidermal grid vertex both have three-dimensional coordinates ( ), the dimension of this weight amplification matrix can be . The specific form is as formula 24 below:

[0290] (24)

[0291] Among them, each matrix element represents the influence degree of the corresponding control point on the epidermal grid vertex , extended to each coordinate component ( ).

[0292] The inverse kinematics target optimization task can be embodied as an objective optimization function that optimizes the joint angle parameters corresponding to the control points in combination with the position information of the target epidermal vertices, so as to realize the inverse kinematics solution of the targeted joint angle parameters and obtain the optimal joint angle parameters required for the control points to control the epidermis.

[0293] Specifically, the goal of the inverse kinematics target optimization task is to optimize the joint angle parameters according to the target epidermal shape , so that the epidermal grid is as close as possible to the target shape. Among them, this inverse kinematics target optimization task can be defined in the form of an objective function as formula 25 below:

[0294] (25)

[0295] Among them, is the position of the target epidermal grid vertex defined by the target epidermal vertex information; is the position of the control point calculated by simulation defined by the aforementioned control point simulation information; They are the skin mesh vertices generated by the control points and can be linearly combined through the control point simulation information and the weight amplification matrix.

[0296] Such as Figures 4A - 5C As shown, according to an embodiment of the present invention, in operation S402 for performing optimization processing on the inverse kinematics target optimization task to obtain the optimal joint angle parameters corresponding to the control point simulation information, it includes:

[0297] Performing a minimization target optimization iteration process on the inverse kinematics target optimization task according to the joint angle constraint conditions to generate the optimal joint angle parameters.

[0298] The joint angle constraint conditions can include joint angle limit constraints and geometric constraints of the kinematic structure of the agent's face. Among them, the joint angle limit constraints can define the range of joint angle activities between the minimum joint angle and the maximum joint angle of the motor rocker arm, and the geometric constraints of the kinematic structure can include, for example, the pulling range of the steel cable / PTFE catheter, the deformation range of the "skin" of the silicone mask, etc. Through the minimization target iteration process (such as non-linear optimization) of the inverse kinematics target optimization task for the above formula 25, the control deformation position of the control point is continuously approximated to the optimal control point position under the target expression, and the joint angle of the motor rocker arm corresponding to this optimal control point position (for example, relative to the starting joint angle) can be used as the optimal joint angle parameter .

[0299] Therefore, by combining the above kinematic modeling and inverse kinematics solution, for the control point area defined by the agent's face, the actuator command solution can be realized through inverse kinematics under physical constraints, so that it approximates the target skin movement generated in the mapping stage.

[0300] Such as Figures 4A - 5C As shown, according to an embodiment of the present invention, in operation S403 for performing execution optimization of the agent's expression imitation according to the optimal joint angle parameters, it includes:

[0301] Obtaining the control point drive information for the agent's expression imitation through the control point simulation information and the optimal joint angle parameters;

[0302] Performing the drive of the agent's expression imitation according to the control point drive information to complete the execution optimization.

[0303] The forward kinematics function corresponding to the control point simulation information obtained through the simulation tool , further combined with the optimal joint angle parameters obtained after the optimization processing of the above inverse kinematics target optimization task , can define the control point drive information for the agent's face motor control system to realize the agent's expression imitation as formula 26 below:

[0304] (26)

[0305] Among them, when the control point driving information can define that the motor rocker arm to be controlled by the motor control system is driven according to the above optimal joint angle parameters, the deformation position of the corresponding control point on the facial epidermis and the deformation position information of the epidermal grid in the surrounding area can be obtained.

[0306] Specifically, the epidermal grid shape can be generated by driving the corresponding control points, and the epidermal grid shape can be expressed by the following formula (27):

[0307] (27)

[0308] Therefore, as Figure 5C shown, the control point driving information defined by the above formulas (26) and (27) can be converted into executable instructions for a physical actuator (such as a motor), so as to realize the fitting of the epidermal drive and the control points, so that the blend shape coefficients based on dense expression motion information generated in the mapping stage can be converted into physically executable intelligent agent facial epidermal motion, and the precise expression imitation can be realized by driving the epidermal motion through the actuator, and the high-precision physical drive optimization for expression imitation can be realized.

[0309] In summary, according to the execution optimization method for intelligent agent expression imitation in the embodiments of the present invention, in the execution stage, an optimization driving means based on control points can be constructed for the physical characteristics of the intelligent agent actuator. First, the control point area is defined, and the inverse kinematics is used to solve the optimal position of the control points under the target expression, and precise actuator motion instructions are generated, so that the motion state of the control point area approaches the target epidermal motion. By increasing the density of the control points, the fineness and naturalness of the expression restoration can be further improved without changing the algorithms in the perception and mapping stages.

[0310] This method not only improves the accuracy of expression imitation, but also ensures the physical executability of the generated instructions through explicit modeling of the physical constraints of the actuator, avoiding the execution failure problem caused by ignoring physical constraints in traditional methods. In addition, the independence of the control point optimization enables the execution stage to flexibly adapt to different types of robot facial structures, providing technical guarantee for the scalability and modular design of the system.

[0311] In summary, in the expression driving stage, by explicitly modeling the physical constraints of the robot actuator and optimizing the expression driving parameters, it is ensured that the generated control instructions are both physically executable and meet the requirements of high precision and real-time performance.

[0312] Based on the above execution optimization method for agent expression imitation, the present invention also provides an execution optimization device for agent expression imitation. The following will be combined with Figure 4B to describe this device in detail.

[0313] Figure 4B The structural block diagram of the execution optimization device for agent expression imitation according to Embodiment 3 of the present invention is schematically shown.

[0314] As Figure 4B shown, the execution optimization device 400 for agent expression imitation in this Embodiment 3 includes a target task generation module 410, a task optimization module 420, and an imitation execution module 430.

[0315] The target task generation module 410 is used to generate an inverse kinematics target optimization task based on the control point simulation information of the agent expression grid and the target skin vertex information. In one embodiment, the target task generation module 410 can be used to perform the operation S401 described above, which will not be elaborated here.

[0316] The task optimization module 420 is used to perform optimization processing on the inverse kinematics target optimization task to obtain the optimal joint angle parameters corresponding to the control point simulation information. In one embodiment, the task optimization module 420 can be used to perform the operation S402 described above, which will not be elaborated here.

[0317] The imitation execution module 430 is used to perform execution optimization of agent expression imitation according to the optimal joint angle parameters. In one embodiment, the imitation execution module 430 can be used to perform the operation S403 described above, which will not be elaborated here.

[0318] According to an embodiment of the present invention, any multiple of the target task generation module 410, the task optimization module 420, and the imitation execution module 430 can be combined and implemented in one module, or any one of them can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the target task generation module 410, the task optimization module 420, and the imitation execution module 430 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable means such as integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the target task generation module 410, the task optimization module 420, and the imitation execution module 430 can be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.

[0319] Combining the above embodiments 1-3, it can be seen that: there are significant technical limitations in the three key stages of perception, mapping, and execution in the existing agent expression imitation technology. These problems directly limit the realization of high-precision expression imitation, especially in the cross-entity expression migration and physical drive implementation. Among them, the perception stage faces the problems of sparse expression representation and coupling between expression and face shape; it is difficult to solve the difficulty of cross-topology expression migration in the mapping stage, and the expression semantics are distorted or even lost during migration; there is a problem that the generated instruction is physically unexecutable in the execution stage. These limitations restrict the fineness and robustness of expression imitation.

[0320] Therefore, in order to overcome the technical bottlenecks existing in the three key stages of perception, mapping, and execution in the existing expression imitation technology, the present invention provides a set of brand-new expression imitation technical solutions, covering three links of perception optimization, mapping optimization, and execution optimization respectively, aiming to achieve higher-precision and more robust expression perception and migration, while ensuring that the generated control instructions are physically executable, laying a technical foundation for the development of highly biomimetic expression robots. Among them, by densifying the expression representation in the perception stage, cross-topology semantic consistency migration in the mapping stage, and high-precision physical drive optimization in the execution stage, the core problems such as sparse expression description, semantic migration distortion, and unimplementable execution in traditional expression imitation methods are systematically solved. The following is the core design of the new technical solution and its corresponding relationship with the existing problems.

[0321] In the perception stage, a three-dimensional deformation model (3DMM) is mainly used to replace the traditional sparse key-point representation method. The facial motion is decomposed into three independent subspaces: pose, face shape, and expression, generating a standardized and dense expression motion representation. That is, dense expression motions are extracted from the input facial images to generate high-precision expression representations under the standardized average face. This method not only improves the refinement of expression description but also eliminates the coupling problem between expression semantics and individual face shapes through the standardized average face topology, achieving the consistency of expression semantics among different individuals. For details, see Embodiment 1.

[0322] In the mapping stage, a set of Blendshape semantic spaces based on the standardized average face is provided as an intermediate bridge between the human face and the robot's face, solving the differences in topological structure and geometric characteristics between the two. The dense expression motions generated in the perception stage are real-time mapped to the low-dimensional Blendshape coefficient space, and distortion-free transfer of expressions from the human face to the robot's face is achieved through the semantically consistent Blendshape basis. In other words, through the semantically consistent Blendshape basis, the dense expression motions are mapped to the robot's face topological structure to achieve cross-topology expression transfer. This method not only ensures the consistency of expression semantics but also greatly improves the efficiency and robustness of the mapping. For details, see Embodiment 2.

[0323] In the execution stage, through high-precision physical drive optimization based on control points, the inverse kinematics is used to solve the optimal positions of the control points under the target expression, generating accurate actuator motion instructions. That is, the mapping results are converted into physically executable robot skin motion instructions, and accurate expression imitation is achieved through inverse kinematics and physical constraint optimization. This method explicitly models the physical constraints of the actuators to ensure the physical executability of the generated instructions. For details, see Embodiment 3.

[0324] The process provided in Embodiments 1-3 of the present invention is optimized layer by layer from perception to execution, ensuring the fineness, semantic consistency, and physical executability of expression imitation, providing a complete solution for the realization of highly biomimetic expression robots. Through systematic design and optimization, this process can achieve full-link expression imitation from the original human face input to the accurate reproduction of robot expressions.

[0325] In short, Embodiments 1-3 of the present invention focus on a high-bionic expression robot real-time imitation system based on two-dimensional video, aiming to accurately capture the dynamic features of human facial expressions and achieve precise mapping and physical driving to a high-bionic expression robot. The technical framework covers three links: perception, migration, and execution, breaking through the technical bottlenecks in subtle expression capture, cross-modal expression migration from human facial expressions to robot facial expressions, and physical driving in traditional methods. This real-time imitation system shows broad prospects in application scenarios. For example, in a remote embodied interaction system, an operator can remotely control a high-bionic expression robot through video input to achieve embodied emotional interaction across spaces, which is applicable to scenarios that require emotional expression such as remote medical escort and psychological counseling. In the basic research of social emotion computing, this technology provides solid technical support for robot autonomous emotional expression and human-robot emotional interaction, promoting the further development of social emotion robot research. Generally speaking, through an innovative technical optimization framework for expression perception, mapping, and execution, the key technical problems in the expression migration process are systematically solved, and the real-time and precise driving of a high-bionic expression robot is successfully achieved. This technology not only provides reliable technical support for remote emotional interaction but also lays an important foundation for the development of future autonomous emotional interaction robots, with extremely high commercial application value.

[0326] Figure 6 A block diagram of an electronic device suitable for implementing a perception optimization method, a mapping optimization method, and an execution optimization method for agent expression imitation according to Embodiments 1-3 of the present invention is schematically shown.

[0327] The above-mentioned electronic device provided by the embodiment of the present invention includes one or more processors and a memory. The memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned perception optimization method, mapping optimization method, and execution optimization method for agent expression imitation.

[0328] As Figure 6 shown, an electronic device 600 according to an embodiment of the present invention includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage section 608 into a random access memory (RAM) 603. The processor 601 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), etc. The processor 601 can also include on-board memory for caching purposes. The processor 601 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0329] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method flow according to the embodiments of the present invention by executing programs in the ROM 602 and / or the RAM 603. It should be noted that the program may also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 may also perform various operations of the method flow according to the embodiments of the present invention by executing programs stored in the one or more memories.

[0330] According to an embodiment of the present invention, the electronic device 600 may further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the I / O interface 605: an input portion 606 including a keyboard, a mouse, etc.; an output portion 607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 608 including a hard disk, etc.; and a communication portion 609 including a network interface card such as a LAN card, a modem, etc. The communication portion 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 610 as needed so that a computer program read from it can be installed into the storage portion 608 as needed.

[0331] The present invention also provides a computer-readable storage medium on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the above-mentioned perception optimization method, mapping optimization method, and execution optimization method for agent expression imitation.

[0332] Among them, the computer-readable storage medium may be included in the device / apparatus / system described in the above embodiments; or it may exist alone and not be assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the above-mentioned perception optimization method, mapping optimization method, and execution optimization method for agent expression imitation according to Embodiments 1-3 of the present invention are implemented.

[0333] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-mentioned ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603.

[0334] An embodiment of the present invention further includes a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the above-mentioned perception optimization method, mapping optimization method, and execution optimization method for agent expression imitation.

[0335] Wherein, the computer program contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the above-mentioned perception optimization method, mapping optimization method, and execution optimization method provided in Embodiments 1-3 of the present invention.

[0336] When the computer program is executed by the processor 601, it executes the above-mentioned functions defined in the system / apparatus of the embodiment of the present invention. According to an embodiment of the present invention, the above-mentioned systems, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0337] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 609, and / or be installed from the removable medium 611. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0338] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or be installed from the removable medium 611. When the computer program is executed by the processor 601, it executes the above-mentioned functions defined in the system of the embodiment of the present invention. According to an embodiment of the present invention, the above-mentioned systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0339] According to embodiments of the present invention, program code for executing the computer programs provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0340] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0341] In addition, all actions of obtaining information, signals, or data in the present invention are carried out on the premise of complying with the corresponding data protection laws, regulations, and policies of the country where it is located, and obtaining authorization from the owner of the corresponding device.

[0342] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present invention can be combined or combined in various ways, even if such combinations or combinations are not explicitly recited in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features recited in the various embodiments and / or claims of the present invention can be combined and combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0343] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present invention is defined by the appended claims and their equivalents. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should fall within the scope of the present invention.

Claims

1. An execution optimization method for agent expression imitation, characterized in that including: generating an inverse kinematics target optimization task based on the control point simulation information of the agent expression grid and the target epidermal vertex information, where the target epidermal vertex information includes the relative displacement distance and displacement direction information of the epidermal grid vertices in the surrounding area of the simulation control point position of the control point simulation information; Perform optimization processing on the inverse kinematics target optimization task to obtain the optimal joint angle parameters corresponding to the control point simulation information, including: performing minimization target optimization iteration processing on the inverse kinematics target optimization task according to the joint angle constraint conditions to generate the optimal joint angle parameters; wherein, the joint angle parameters are the driving angle parameters of the driving motor rocker corresponding to the control points defined by the control point simulation information; and performing execution optimization of the agent expression imitation according to the optimal joint angle parameters; wherein, the inverse kinematics target optimization task is defined in the form of an objective function as: ; Among them, is the position of the target skin mesh vertex defined by the target skin vertex information; is the position of the control point calculated through simulation defined by the control point simulation information; is the skin mesh vertex generated by driving with the control point, and is linearly combined through the control point simulation information and the weight amplification matrix; wherein the weight amplification matrix is obtained by expanding the distance weight matrix, and the distance weight matrix is generated based on the control point simulation information and the target skin vertex information; Among them, each constituent matrix element in the distance weight matrix is used to represent a single control point on the agent's expression grid and the corresponding epidermal vertex The distance distribution between them. Therefore, the distance weight matrix can define the distance-based weight distribution between the epidermal control points and the epidermal grid vertices; among them, these matrix elements are specifically expressed by the following formula: ; Among them, is the Euclidean distance between the epidermal vertex and the control point ; is the distance from the epidermal vertex to the control point; is the three-dimensional position coordinates of the subcutaneous controller; is the control parameter of the distance influence, used to adjust the attenuation range of the weight; the weight normalization ensures , that is, the sum of the weights of each epidermal vertex is 1, where is the number of control points.

2. The method according to claim 1, characterized in that, before generating the inverse kinematics target optimization task based on the control point simulation information of the agent expression grid and the target epidermal vertex information, it further includes: generating an agent expression grid corresponding to the agent according to the blend shape basis of the agent's face.

3. The method according to claim 2, wherein before generating the distance weight matrix based on the control point simulation information and the target epidermal vertex information of the agent expression grid, it further includes: generating the control point simulation information according to the preset forward kinematics optimization task of the agent expression grid and the joint angle parameter vector corresponding to the simulation control point; generating the target epidermal vertex information according to the deformation weight matrix of the epidermal grid vertices of the agent expression grid and the control point simulation information.

4. The method according to claim 1, wherein in the execution optimization of the agent expression imitation according to the optimal joint angle parameters, it includes: obtaining the control point driving information of the agent expression imitation through the control point simulation information and the optimal joint angle parameters; performing the driving of the agent expression imitation according to the control point driving information to complete the execution optimization.

5. An execution optimization device for agent expression imitation, characterized in that including: a target task generation module, configured to generate an inverse kinematics target optimization task based on the control point simulation information of the agent expression grid and the target epidermal vertex information, where the target epidermal vertex information includes the relative displacement distance and displacement direction information of the epidermal grid vertices in the surrounding area of the simulation control point position of the control point simulation information; The task optimization module is used to perform optimization processing on the inverse kinematics target optimization task to obtain the optimal joint angle parameters corresponding to the control point simulation information, including: performing minimization target optimization iteration processing on the inverse kinematics target optimization task according to the joint angle constraint conditions to generate the optimal joint angle parameters; wherein, the joint angle parameters are the driving angle parameters of the driving motor rocker corresponding to the control points defined by the control point simulation information; and an imitation execution module, configured to perform execution optimization of the agent expression imitation according to the optimal joint angle parameters; wherein, the inverse kinematics target optimization task is defined in the form of an objective function as: ; Among them, is the position of the target skin mesh vertex defined by the target skin vertex information; is the position of the control point calculated by simulation defined by the control point simulation information; is the skin mesh vertex generated by driving the control point, and is linearly combined by the control point simulation information and the weight amplification matrix; wherein the weight amplification matrix is obtained by expanding the distance weight matrix, and the distance weight matrix is generated based on the control point simulation information and the target skin vertex information; Among them, each component matrix element in the distance weight matrix is used to represent a single control point on the expression grid of the agent and the corresponding epidermal vertex The distance distribution between them. Therefore, the distance weight matrix can define the distance-based weight distribution between the epidermal control points and the epidermal grid vertices; among them, these matrix elements are specifically expressed by the following formula: ; wherein, is the Euclidean distance between the epidermal vertex and the control point ; is the distance from the epidermal vertex to the control point; is the three-dimensional position coordinates of the subcutaneous controller; is the control parameter of the distance influence, used to adjust the attenuation range of the weight; the weight normalization ensures , and the sum of the weights of each epidermal vertex is 1, where is the number of control points.

6. An electronic device, including: one or more processors; a memory, configured to store one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 4.

7. A computer-readable storage medium, having executable instructions stored thereon, and when the instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 4.

8. A computer program product, including a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Tandem robot, inverse kinematics solving method thereof, related medium and equipment

    CN118636129A