Perception optimization method, device and equipment for agent expression simulation and medium
By generating and optimizing the reconstruction task of dense faces, decoupling individual face shape characteristics and expression dynamics, high-precision and dense expression representation are achieved, and the problems of insufficient complexity and poor accuracy of expression description in the prior art are solved.
Patent Information
- Application Number
- CN202510307489.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-13
AI Technical Summary
In the emoticon perception stage, the prior art has problems such as difficult to capture subtle expressions, insufficient complexity and oversimplification of expression description, and poor accuracy and versatility.
By extracting the expression motion information of the target expression image, a dense face reconstruction task is generated, and three-dimensional face reconstruction information is generated through joint optimization to achieve the decoupling of individual face characteristics and expression dynamics, and dense expression motion under standardized average faces is generated.
It significantly improves the accuracy and delicateness of expression perception, improves the ability to capture subtle expressions, and ensures the consistency of expression semantics among individuals with different face shapes.
Smart Images

Figure CN120148089A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, specifically to the field of image data processing technology, and more specifically to a perception optimization method, device, equipment, medium and product for intelligent agent expression imitation. Background Art
[0002] Artificial Intelligence (AI for short) is an important driving force for the new round of scientific and technological revolution and industrial transformation. It is a new key technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence. As an important part of intelligent science, artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine (such as a highly biomimetic expression robot) that can react in a way similar to human intelligence.
[0003] As an important part of artificial intelligence technology, the technology of human expression recognition and reproduction (such as expression imitation) has received extensive attention in the fields of human-computer interaction, security recognition, robot manufacturing, automation, medical treatment, communication and driving, and has quickly become a research hotspot in the industry. The expression imitation technology is mainly divided into three stages: expression perception, mapping migration and driving execution. The aim is to accurately capture the dynamic expression features of the human face in the expression perception stage, realize the face mapping to embodied intelligence such as highly biomimetic expression robots in the mapping migration stage, and finally control the expression actuator to achieve accurate expression reproduction drive in the driving execution stage. However, in the existing traditional expression imitation technology, there are problems such as difficult capture of subtle expressions, insufficient complexity and over-simplification of expression description, and poor accuracy and generality in the expression perception stage. Summary of the Invention
[0004] In view of at least one of the technical problems existing in the above-mentioned prior art, the embodiments of the present invention aim to provide a perception optimization method, device, equipment, medium and product for intelligent agent expression imitation, thereby providing an optimization scheme for the perception stage of intelligent agent expression imitation, in order to achieve effective decoupling of individual face shape features and expression dynamics, significantly improve the capture ability of subtle expressions through dense expression representation, and overcome the limitations of traditional solutions.
[0005] One aspect of the embodiments of the present invention provides a perception optimization method for intelligent agent expression imitation, which includes: extracting the expression motion information of the target expression image; generating a dense face reconstruction task through the expression motion information and preset joint optimization information; performing joint optimization on the dense face reconstruction task to generate three-dimensional face reconstruction information, and the three-dimensional face reconstruction information is used to generate the target three-dimensional face; and extracting the dense expression motion information of the standardized average face according to the target three-dimensional face to complete the perception optimization.
[0006] According to an embodiment of the present invention, in extracting the expression motion information of the target expression image, it includes: performing expression feature recognition on the target expression image to extract the expression motion information.
[0007] According to an embodiment of the present invention, in generating a dense face reconstruction task through the expression motion information and the preset joint optimization information, it includes: obtaining the task reconstruction information of the corresponding dense face reconstruction task according to the expression parameters of the expression motion information; constructing a dense face reconstruction task according to the task reconstruction information and the preset joint optimization information.
[0008] According to an embodiment of the present invention, in performing joint optimization on the dense face reconstruction task to generate three-dimensional face reconstruction information, it includes: generating an initial three-dimensional face corresponding to the dense face reconstruction task; performing joint optimization processing on the dense face reconstruction task according to the initial three-dimensional face to generate three-dimensional face reconstruction information.
[0009] According to an embodiment of the present invention, in generating an initial three-dimensional face corresponding to the dense face reconstruction task, it includes: detecting two-dimensional key point information in the target expression image through a preset key point detection model; obtaining the initial pose parameters of the standardized average face according to the two-dimensional key point information.
[0010] According to an embodiment of the present invention, in generating an initial three-dimensional face corresponding to the dense face reconstruction task, it further includes: initializing the optimization variable parameters of the dense face reconstruction task based on the initial pose parameters; generating an initial three-dimensional face through the standardized average face and the default lighting conditions based on the initialization result of the optimization variable parameters.
[0011] According to an embodiment of the present invention, in performing joint optimization processing on the dense face reconstruction task according to the initial three-dimensional face to generate three-dimensional face reconstruction information, it includes: generating the three-dimensional face reconstruction information after convergence by performing iterative minimization optimization on the projection error, photometric consistency error, and regularization constraint in the dense face reconstruction task.
[0012] According to an embodiment of the present invention, in extracting the dense expression motion information of the standardized average face from the target three-dimensional face to complete the perception optimization, it includes: optimizing the shape parameters of the target three-dimensional face to complete the extraction of the dense expression motion information.
[0013] Another aspect of an embodiment of the present invention provides a perception optimization device for agent expression imitation, which includes a motion extraction module, a task generation module, a joint optimization module, and an information extraction module. The motion extraction module is used to extract the expression motion information of the target expression image; the task generation module is used to generate a dense face reconstruction task through the expression motion information and preset joint optimization information; the joint optimization module is used to perform joint optimization on the dense face reconstruction task to generate three-dimensional face reconstruction information, and the three-dimensional face reconstruction information is used to generate the target three-dimensional face; the information extraction module is used to extract the dense expression motion information of the standardized average face according to the target three-dimensional face to complete the perception optimization.
[0014] Another aspect of an embodiment of the present invention provides an electronic device, including one or more processors and a memory, where the memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned perception optimization method for agent expression imitation.
[0015] Another aspect of an embodiment of the present invention provides a computer-readable storage medium, on which executable instructions are stored. When the instructions are executed by a processor, the processor is caused to execute the above-mentioned perception optimization method for agent expression imitation.
[0016] Another aspect of an embodiment of the present invention provides a computer program product, including a computer program, which when executed by a processor implements the above-mentioned perception optimization method for agent expression imitation.
[0017] The perception optimization method for agent expression imitation provided by the embodiments of the present invention can at least partially solve the problems in the related art, and thus can at least achieve one of the following technical effects:
[0018] Therefore, in the perception stage, compared with the traditional sparse key-point based facial expression representation scheme, the above-mentioned perception optimization method for agent facial expression imitation in the embodiments of the present invention can represent the facial expression motion information by using the target three-dimensional face, map the facial expressions in the two-dimensional image into parameterized three-dimensional representations, and decouple the individual facial shape features and expression dynamics effectively by virtue of its decomposition ability for facial shape, expression, and pose spaces, generating a dense expression motion under the standardized average face topology, thereby overcoming the abstract defect of the traditional sparse key-point scheme, achieving the densification of expression description, and at the same time eliminating the coupling problem between expression semantics and individual facial shapes by normalizing the expressions to the standardized average face, ensuring the consistency of expression semantics among individuals with different facial shapes. Therefore, by virtue of the high precision and density of the above-mentioned three-dimensional facial expression model based on the standardized average face of the target three-dimensional face, the accuracy and fineness of facial expression perception can be significantly improved, the ability to capture subtle expressions can be enhanced, thus laying a high-quality foundation for the subsequent mapping and execution stages.
[0019] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the scope of what the present invention claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features, and advantages of the present invention will become clearer. In the drawings:
[0021] Figure 1 Schematically shows an application scenario diagram of an optimization method, device, equipment, medium, and program product for agent facial expression imitation according to Embodiments 1-3 of the present invention;
[0022] Figure 2A Schematically shows a flowchart of a perception optimization method for agent facial expression imitation according to Embodiment 1 of the present invention;
[0023] Figure 2B Schematically shows a structural block diagram of a perception optimization device for agent facial expression imitation according to Embodiment 1 of the present invention;
[0024] Figure 3A Schematically shows a flowchart of a mapping optimization method for agent facial expression imitation according to Embodiment 2 of the present invention;
[0025] Figure 3B Schematically shows a structural block diagram of a mapping optimization device for agent facial expression imitation according to Embodiment 2 of the present invention;
[0026] Figure 4A Schematically shows a flowchart of an execution optimization method for agent facial expression imitation according to Embodiment 3 of the present invention;
[0027] Figure 4B Schematically shows a structural block diagram of an execution optimization device for agent expression imitation according to Embodiment 3 of the present invention;
[0028] Figure 5A Schematically shows an expression semantic alignment diagram from the human face expression mixing space to the robot expression mixing space according to Embodiments 1-3 of the present invention;
[0029] Figure 5B Schematically shows a robot expression kinematics design diagram according to Embodiments 1-3 of the present invention;
[0030] Figure 5C Schematically shows a robot expression electromechanical design diagram according to Embodiments 1-3 of the present invention; and
[0031] Figure 6 Schematically shows a block diagram of an electronic device according to the method of Embodiments 1-3 of the present invention.
[0032] The above-mentioned drawings are a part of the specification of the embodiments of the present invention, which illustrate the exemplary embodiments of the present invention. The accompanying drawings, together with the description of the specification, are used to explain the principles of the embodiments of the present invention. It should be understood that the above general description of the drawings and the following specific embodiments are only exemplary and explanatory, and they do not limit the scope that the present invention intends to claim. Detailed Embodiments
[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the spirit of the content disclosed by the present invention will be clearly described below with reference to the drawings and in detail. After any person skilled in the art understands the embodiments of the content of the present invention, they can make changes and modifications based on the techniques taught by the content of the present invention, and these changes and modifications do not depart from the spirit and scope of the content of the present invention.
[0034] The schematic embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention. In addition, elements / components using the same or similar reference numerals in the drawings and embodiments are used to represent the same or similar parts.
[0035] Regarding the "first", "second",... used in the present invention, they do not particularly refer to the meaning of order or sequence, nor are they used to limit the present invention. They are only used to distinguish elements or operations described with the same technical terms.
[0036] Regarding the directional terms used in the present invention, such as: up, down, left, right, front or back, etc., they are only references to the directions in the drawings. Therefore, the directional terms used are for explanation and not for limiting this creation.
[0037] The terms "comprising", "including", "having", "containing", etc. used in the present invention are all open-ended terms, meaning including but not limited to.
[0038] The "and / or" used in the present invention includes any one or all combinations of things.
[0039] The "plurality" in the present invention includes "two" and "more than two"; the "multiple groups" in the present invention includes "two groups" and "more than two groups".
[0040] The terms "substantially", "about", etc. used in the present invention are used to modify any quantity or error that can vary slightly, but these slight variations or errors do not change its essence. Generally, the range of such slight variations or errors modified by such terms can be 20% in some embodiments, 10% in some embodiments, 5% in some embodiments, or other values. Those skilled in the art should understand that the aforementioned values can be adjusted according to actual needs and are not limited thereto.
[0041] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0042] In cases where expressions such as "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but not be limited to a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). In cases where expressions such as "at least one of A, B, or C, etc." are used, generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, or C" should include but not be limited to a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). Those skilled in the art should also understand that substantially any disjunctive conjunction and / or phrase representing two or more alternative items, whether in the specification, claims, or drawings, should be understood as giving the possibility of including one of these items, either of these items, or both items. For example, the phrase "A or B" should be understood as including the possibility of "A" or "B", or "A and B".
[0043] In addition, all actions of obtaining information, signals or data in the present invention are carried out on the premise of complying with the corresponding data protection laws, regulations and policies of the country where it is located and obtaining the authorization given by the owner of the corresponding device.
[0044] In human interpersonal communication in daily life, by controlling facial expressions, the communication effect can be enhanced. For humans, facial expressions are the result of one or more actions or states of facial muscles. These movements express the emotional state of an individual human being towards the external environment. Facial expressions are a form of non-verbal communication and are the main means of expressing social information between humans.
[0045] For artificial intelligence, in order to enable intelligent machines to make real human interaction responses in a way similar to human intelligence, it is necessary to create a truly humanoid machine in all aspects, including limbs, appearance, body posture, expressions, and even voice conversations, mobility, and human-like thinking patterns. Among them, the expression imitation technology provides strong technical support for today's humanoid robots to enter the real world and provide emotional value and non-verbal power for humans.
[0046] The expression imitation technology aims to accurately capture the dynamic characteristics of human facial expressions and achieve precise mapping and physical drive to highly biomimetic expression robots. Its overall technical framework covers three links: expression perception, mapping migration, and drive execution. No matter which link has problems, it will directly limit the realization of high-precision expression imitation.
[0047] To solve at least one of the existing technical problems in the three links of expression perception, mapping migration, and drive execution in the process of intelligent agent expression imitation in the existing technology, the present invention provides Examples 1-3 to respectively optimize the above three links in the process of intelligent agent expression imitation, with the expectation of significantly improving the intelligent level of intelligent agent expression imitation.
[0048] Figure 1 Schematically shows the application scenario diagrams of the optimization methods, devices, equipment, media, and program products for intelligent agent expression imitation according to Embodiments 1-3 of the present invention.
[0049] As Figure 1 shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0050] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0051] Terminal devices 101, 102, and 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0052] Server 105 can be a server that provides various services, such as a background management server that supports the websites browsed by users using terminal devices 101, 102, and 103 (for example only). The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0053] It should be noted that the optimization method for agent expression imitation provided by the embodiments of the present invention can generally be executed by server 105. Correspondingly, the optimization device for agent expression imitation provided by the embodiments of the present invention can generally be set in server 105. The optimization method for agent expression imitation provided by the embodiments of the present invention can also be executed by a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the optimization device for agent expression imitation provided by the embodiments of the present invention can also be set in a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.
[0054] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0055] are merely illustrative. According to actual needs, there can be any number of terminal devices, networks, and servers.
[0056] In order to further reflect the complete technical process of the three key stages of perception, mapping and execution in the imitation of intelligent body expression, and to ensure that the full-link expression imitation from the original face input to the accurate reproduction of the robot expression can be achieved, the following is combined with the embodiments 1-3 of the present invention and Figures 1 - 5C Further explanation is as follows:
[0057] It should be noted that the intelligent agent may be the execution subject of at least one of the perception optimization method, mapping optimization method and execution optimization method provided in the above-mentioned embodiments 1-3 of the present invention, or may be the execution party controlled by these methods, such as a humanoid intelligent robot or other AI device with a simulated human face skin that can imitate human facial expressions. Specifically, Figure 5B and Figure 5C The electromechanical design of the intelligent body of facial expression shown in the figure covers the surface with "muscle" and "skeleton" with simulated skin (silicone mask, i.e. skin), so that the movement of the simulated skin can be realized through the motor (such as LFD-01M servomotor) and steel cable and Teflon conduit matched by the electrical control system, thereby achieving the imitation of facial expressions.
[0058] The following will be combined Figures 2A - 5C The perception optimization method, mapping optimization method and execution optimization method provided in Examples 1-3 of the present invention are described in detail respectively.
[0059] Embodiment 1:
[0060] The expression perception stage involves facial expression recognition. Achieving accurate expression recognition perception, obtaining more complex expression descriptions, and ensuring the accuracy and versatility of expression perception are technical issues that need to be urgently resolved in the embodiments of the present invention.
[0061] Specifically, in the perception stage, traditional methods usually use sparse landmarks to structurally represent faces. This method relies on the position and displacement of landmarks to describe the dynamic changes of facial expressions. Although simple and efficient, its main problem is that the complexity of expression description is insufficient. Since sparse landmarks can only capture limited local dynamic features, they cannot fully represent the subtle changes and complex characteristics of human facial expressions, resulting in oversimplification of expression description. In addition, there is a coupling relationship between face shape and expression semantics, which makes it impossible for the same geometric displacement to maintain consistent expression semantics in individuals with different face shapes. This coupling further limits the accuracy and versatility of expression perception.
[0062] Therefore, in view of at least one of the above technical problems existing in the prior art, embodiments of the present invention aim to provide a method, apparatus, device, medium and product for optimizing the perception of agent expression imitation, thereby providing an optimization scheme for the perception stage of agent expression imitation, in order to achieve effective decoupling of individual face shape features and expression dynamics, significantly improve the ability to capture subtle expressions through dense expression representation, and overcome the limitations of traditional solutions.
[0063] The following will be based on Figure 1 the described scenario, and through Figures 2A - 5C a detailed description of the method for optimizing the perception of agent expression imitation in the disclosed embodiments will be given.
[0064] As Figure 2A shown, one aspect of the embodiments of the present invention provides a method for optimizing the perception of agent expression imitation, which includes operations S201 to S204.
[0065] In operation S201, extract the expression motion information of the target expression image;
[0066] In operation S202, generate a dense face reconstruction task through the expression motion information and preset joint optimization information;
[0067] In operation S203, perform joint optimization on the dense face reconstruction task to generate three-dimensional face reconstruction information, and the three-dimensional face reconstruction information is used to generate the target three-dimensional face; and
[0068] In operation S204, extract the dense expression motion information of the standardized average face according to the target three-dimensional face to complete the perception optimization.
[0069] In the embodiments of the present invention, agent expression imitation can be a process in which an agent recognizes and obtains a face expression image, and reproduces the face expression in the face expression image using the above Figure 5C electromechanical design. Among them, the face expression image as the object to be recognized can be the image information directly detected by the agent through its own image detection device (such as a camera image sensor, etc.), or the image information transmitted to the agent through a network transmission device (wireless or wired), and these image information can be video or picture data.
[0070] For example, as a type of intelligent agent, a humanoid robot conducts a video call with a real human user. The humanoid robot can directly or indirectly obtain the video data of the video call and preprocess the video content of the human user included in the video data, and extract the expression features from each frame of the image containing the human user's expression in the video content. Among them, the target expression image can be one of these frame images, and the single-frame image contains the facial expression of the expression imitation object (human user) at a certain moment. In addition, the humanoid robot can also directly interact with the human user face to face, and obtain and recognize the real-time expression of the human user in real time.
[0071] The expression motion information can be the facial expression feature parameters extracted for the facial expression in the target expression image. These expression feature parameters can involve the expression basis parameters and the facial expression motion parameters of the target facial expression in the target expression image (such as the facial shape coefficient, the dynamic expression motion coefficient, and the facial pose coefficient in the camera coordinate system, etc.). Among them, the so-called expression basis parameters can include more (such as 50) different expression bases, and each expression basis coefficient represents the displacement of each corner point on the dense mesh of the facial expression of the current facial image relative to the standardized average face; in addition, the facial expression motion parameters are the motion field parameters relative to the standardized average face in the standardized average face coordinate system.
[0072] The preset joint optimization information can be the information for three-dimensional human face reconstruction based on the standardized average face and the expression motion information according to the target three-dimensional human face reconstruction requirements, and the target three-dimensional human face reconstruction requirements are also related to the target three-dimensional human face reconstruction model used for this three-dimensional human face reconstruction. The standardized average face is the three-dimensional human face average shape obtained through statistical learning (i.e., the three-dimensional average face, Mean Shape). The preset joint optimization information is mainly used to generate a dense human face reconstruction task according to the expression motion information of the extracted target expression image. The dense human face reconstruction task can be a joint optimization task for restoring the three-dimensional human face shape and the expression motion parameters according to the target three-dimensional human face reconstruction requirements, for example, it can be embodied as a joint optimization problem created for the restoration of the three-dimensional human face shape.
[0073] By performing joint optimization processing on the dense human face reconstruction task, three-dimensional human face reconstruction information can be generated according to the processing results of the dense human face reconstruction task, and the three-dimensional human face reconstruction information can include the corresponding facial expression basis B exp and the corresponding facial expression coefficient β. The facial expression basis is usually used to describe the dynamic expression changes of the human face, and the facial expression coefficient is usually used as a parameterized representation of the facial expression. The facial expression basis and the facial expression coefficient can be configured to construct a three-dimensional human face expression model based on the standardized average face.
[0074] The target 3D face can be understood as a 3D face expression model based on a standardized average face. For example, a 3D deformation model (such as the 3D Morphable Model, abbreviated as 3DMM) can be a statistical model for 3D face modeling. By mainly learning a large amount of 3D face data, any 3D face can be represented as a combination of linear bases, which can be represented by the 3D face shape based on the standardized average face. This 3D face shape can be based on the standardized average face (i.e., the 3D average face, Mean Shape), the face shape basis B id , the face expression basis B exp , the face shape coefficient α, the face expression coefficient β, etc. Therefore, the face expression in the 2D image can be mapped to a parameterized 3D representation. With its ability to decompose the face shape, expression, and pose spaces, the effective decoupling of individual face shape features and expression dynamics can be achieved.
[0075] Among them, the dense expression motion information can be information directly related to the dynamic changes of the face expression related to the extracted expression motion information, and can specifically be characterized by the linear combination of the face expression basis and the face expression coefficient. The dense expression motion information can most directly reflect the dense expression motion information of the user extracted by the agent on the standardized average face, realizing the face expression representation.
[0076] Therefore, in the perception stage, compared with the traditional sparse key-point based expression representation scheme, the above-mentioned perception optimization method for agent expression imitation in the embodiments of the present invention can represent the face expression motion information by adopting the target 3D face, map the face expression in the 2D image to a parameterized 3D representation, and with its ability to decompose the face shape, expression, and pose spaces, achieve the effective decoupling of individual face shape features and expression dynamics, generate a dense expression motion under the topology of the standardized average face, thus overcoming the abstract defect of the traditional sparse key-point scheme, realizing the densification of expression description. At the same time, by normalizing the expression to the standardized average face, the coupling problem between the expression semantics and the individual face shape is eliminated, ensuring the consistency of the expression semantics among different face shape individuals. Therefore, with the high accuracy and density of the above-mentioned 3D face expression model based on the standardized average face of the target 3D face, the accuracy and fineness of face expression perception can be significantly improved, and the ability to capture subtle expressions can be enhanced, thereby laying a high-quality foundation for the subsequent mapping and execution stages.
[0077] As Figures 2A - 5C shown, according to an embodiment of the present invention, in the operation of extracting the expression motion information of the target expression image in S201, it includes: performing expression feature recognition on the target expression image to extract the expression motion information.
[0078] The target expression image can be a single-frame two-dimensional image (such as an RGB image), which can be selected from a set of two-dimensional images based on time series (such as a video), and can be specifically selected successively according to the time series. Each time the expression feature recognition is mainly carried out for the selected single-frame two-dimensional image to extract the expression feature parameters of the target human face in the single-frame two-dimensional image. These feature parameters can include expression basis parameters and human face expression motion parameters, etc.
[0079] For a complete video, it can record the human face images at each moment within a certain time period, and at the same time, the human face images at each moment record the instantaneous expression images of the human face at that moment. These instantaneous expression images are the above-mentioned target expression images and serve as the targets for expression feature extraction. Therefore, by completing the expression feature recognition for all the target expression images in the above video one by one and realizing the dense expression motion representation of each target expression image one by one, the perception and acquisition of the human face expressions within the corresponding time period of the entire video can be achieved.
[0080] Therefore, through the above-mentioned expression feature recognition, the accurate expression feature extraction of each frame of two-dimensional image can be realized, thereby laying a foundation for the capture of micro-expressions and improving the accuracy and fineness of expression perception.
[0081] As Figures 2A - 5C shown, according to an embodiment of the present invention, in the operation of generating a dense human face reconstruction task through expression motion information and preset joint optimization information in S201, it includes: obtaining the task reconstruction information corresponding to the dense human face reconstruction task according to the expression parameters of the expression motion information; constructing the dense human face reconstruction task according to the task reconstruction information and the preset joint optimization information.
[0082] The expression parameters of the expression motion information can be the expression feature parameters extracted from the human face expression of the target expression image, specifically involving expression basis parameters and human face expression motion parameters. Among them, the human face expression motion parameters can be the human face shape coefficient, dynamic expression motion coefficient, and human face pose coefficient, etc. in the camera coordinate system.
[0083] In order to perceive the human face expression motion information of the expression motion information and realize the representation based on the standardized average face, three-dimensional human face reconstruction can be carried out based on the standardized average face and the expression motion information. Among them, the task reconstruction information can be used to construct the optimization variable information and objective function information of the joint optimization problem corresponding to the above-mentioned dense human face reconstruction task, and specifically can be determined according to the construction requirements of the dense human face reconstruction task. Among them, the preset joint optimization information can define the construction rules of the dense human face reconstruction task, and through this task construction rule, the dense human face reconstruction task can be generated according to the task reconstruction information.
[0084] Specifically, for the above single-frame target expression image, a dense face reconstruction task is constructed to generate the joint optimization problem required for dense three-dimensional face reconstruction.
[0085] Among them, the single-frame target expression image can be expressed as where w is the width of the image and h is the height of the image. Among them, the reconstruction target of the dense face reconstruction task is to restore the corresponding three-dimensional face shape S and its expression motion parameters β.
[0086] According to the task construction rules defined by the preset joint optimization information, the dense face reconstruction task can be formalized into the following joint optimization problem (Formula 1) based on the optimization variable information and objective function information corresponding to the task reconstruction information:
[0087]
[0088] Among them, the optimization variable information includes the face shape coefficient α for reconstructing the individual static face shape features, the face expression coefficient β for reconstructing the dense dynamic expression motion, the face rotation matrix R and translation vector t for representing the face pose, and the lighting parameter γ for describing the face lighting conditions.
[0089] In addition, the objective function information can include projection error, photometric consistency error, shape and regularization constraint information, etc. Among them, the objective function of the projection error (Landmark Projection Loss) can be expressed as Formula 2 below:
[0090]
[0091] where K is the camera intrinsic matrix; is the 3D vertex corresponding to the target three-dimensional face; is the two-dimensional key point detected in the target expression image (such as sparse corner points like the corners of the eyes and mouth). Among them, i can be the index of the two-dimensional key point (Landmarks), i ∈ [1, n k ; where n k can be the number of two-dimensional key points of the preset grid, which can be specifically determined according to different application systems of dense grid corner points, such as 146 key points of the Mediapipe standard or 68 key points of Dlib.
[0092] The objective function of the photometric consistency error (Photometric Loss) can be expressed as Formula 3 below:
[0093]
[0094] Wherein, Ω is the face region in the target expression image; I(u) is the pixel color in the input target expression image; is the pixel color generated by the rendering model (such as the face shape based on 3DMM, the lighting parameter γ, the material, and the camera parameters); u can represent the pixel position of the target expression image. λ photo can be the objective function weight, such as λ photo = 0.05.
[0095] The regularization constraint information can include regularization terms for face shape and face expression. Among them, λ id ∥α∥ 2 can be the face shape coefficient constraint term for constraining the magnitude of the shape coefficient to prevent overfitting; λ exp ∥β∥ 2 can be the face expression coefficient constraint term for constraining the magnitude of the expression coefficient; can be an additional regularization term (such as the sparsity of the expression coefficient); λ reg can be an additional regularization term weight, such as λ reg = 1.0.
[0096] Therefore, by means of the construction of the above-mentioned dense face reconstruction task, it can lay a foundation for the subsequent reconstruction of the target 3D face based on the standardized average face, thereby replacing the traditional sparse key-point-based facial expression perception scheme, using the 3D deformation model as the facial expression representation means in the perception stage, ensuring that the facial expression movement of the face can be decomposed into movements of pose, shape, and expression, and ensuring that the reconstructed target 3D face can achieve refined facial expression description.
[0097] Such as Figures 2A - 5C shown, according to an embodiment of the present invention, in operation S202 for performing joint optimization on the dense face reconstruction task to generate 3D face reconstruction information, it includes: generating an initial 3D face corresponding to the dense face reconstruction task; performing joint optimization processing on the dense face reconstruction task according to the initial 3D face to generate 3D face reconstruction information.
[0098] When constructing the initial 3D face based on the above-mentioned standardized average face, the initial 3D face can be a standardized 3D face with initial lighting conditions processed on the standardized average face, which can be used as the construction basis for the target 3D face to ensure that the target 3D face can accurately represent the facial expression information of the target expression image.
[0099] The three-dimensional face reconstruction information is used to construct the target three-dimensional face based on the initial three-dimensional face, and can be mainly generated according to the joint optimization processing result of the dense face reconstruction task, which may include the face expression basis B exp and basic information for three-dimensional face reconstruction such as the face expression coefficient β.
[0100] As shown in the foregoing formula 1, the joint optimization processing process of the dense face reconstruction task can be expressed as a process of solving this formula 1. By solving the joint optimization problem constructed according to the preset joint optimization information, on the basis of the above initial three-dimensional face and expression motion information, the above face expression basis B exp and the face expression coefficient β are solved, so that the acquisition of three-dimensional face reconstruction information can be realized.
[0101] Therefore, by means of the acquisition of this three-dimensional face reconstruction information, the accurate modeling of the face geometry and expression motion can be ensured.
[0102] As Figures 2A - 5C shown, according to an embodiment of the present invention, in generating the initial three-dimensional face corresponding to the dense face reconstruction task, it includes: detecting the two-dimensional key point information in the target expression image through a preset key point detection model; obtaining the initial pose parameters of the standardized average face according to the two-dimensional key point information.
[0103] The preset key point detection model can be a pre-trained model for detecting the two-dimensional face key points of the target expression image, such as the MediaPipe model or the Dlib model, etc. Among them, the target expression image can represent the face expression based on the topological structure of the two-dimensional face. Among them, the two-dimensional key point information can be the two-dimensional information expression of the sparse corner points (such as eyebrows, eyelids, etc.) on the topological structure of the two-dimensional face, and can be specifically characterized.
[0104] The two-dimensional key point information (face mesh) can include the position information (such as two-dimensional coordinates) of the above two-dimensional key points in the two-dimensional topological structure. With the help of this position information, the initial pose parameters of the face pose parameters of the expression motion information can be estimated. Among them, the initial pose parameters can include the face pose parameters of the target expression image of the current frame relative to the standardized average face, such as the rotation matrix R and the translation vector t of the face pose relative to the standardized average face.
[0105] Therefore, by means of the acquisition of the initial pose parameters, the mapping relationship between the two-dimensional face topology and the standardized average face can be effectively established, so that the extracted two-dimensional key point information can be accurately mapped to the standardized average face, which is beneficial to more comprehensively and accurately representing the subtle changes and complex characteristics of the face expression subsequently.
[0106] As Figures 2A - 5C shown, according to an embodiment of the present invention, in generating an initial three-dimensional human face corresponding to a dense human face reconstruction task, it further includes: initializing the optimization variable parameters of the dense human face reconstruction task based on the initial pose parameters; and generating the initial three-dimensional human face by means of a standardized average face and a default lighting condition based on the initialization result of the optimization variable parameters.
[0107] The rotation matrix R and the translation vector t of the human face pose of the standardized average face corresponding to the initial pose parameters can be used to initialize the optimization variable parameters corresponding to the dense human face reconstruction task, where the optimization variable parameters may include a human face shape coefficient α, a human face expression coefficient β, a lighting parameter γ, etc. Among them, the above parameter initialization is mainly used to initialize the human face shape coefficient α, the human face expression coefficient β, and the lighting parameter γ to zero vectors.
[0108] Therefore, the initialization results of the above optimization variable parameters are the human face shape coefficient α, the human face expression coefficient β, and the lighting parameter γ that are zero vectors. Based on the above zero-vector human face shape coefficient α, human face expression coefficient β, and lighting parameter γ, an initial three-dimensional human face is constructed according to the default lighting condition on the basis of the standardized average face. Among them, the default lighting condition can be used as the default lighting data for the standardized average face during the real-time human face alignment process, which can ensure the detection and alignment accuracy of the generated initial three-dimensional human face and exclude the influence of changes in lighting conditions. The default lighting condition can be the lighting information of a point light source that is consistent with the camera center. When acquiring the above target expression image, it is generated according to the lighting information of the camera or the lighting condition is estimated according to the target expression image.
[0109] Thus, it can be ensured that the human face expression of the target expression image can be normalized to the standardized average face, thereby eliminating the coupling problem between the expression semantics and the individual face shape, and ensuring the consistency of the expression semantics among different face shape individuals.
[0110] As Figures 2A - 5C shown, according to an embodiment of the present invention, in operation S203, when performing joint optimization processing on the dense human face reconstruction task according to the initial three-dimensional human face to generate three-dimensional human face reconstruction information, it includes: generating the three-dimensional human face reconstruction information after convergence by performing iterative optimization on the projection error, photometric consistency error, and regularization constraint in the dense human face reconstruction task.
[0111] In the joint optimization process of the dense human face reconstruction task, the objective function of the projection error in formula 1 defined for the above joint optimization problem can be based on iterative optimization means (such as the Levenberg-Marquardt algorithm or the gradient descent scheme) The objective function of the photometric consistency error and regularization constraint information (such as the face shape coefficient constraint term λ id ∥α∥ 2 and the face expression coefficient constraint term λ exp ∥β∥ 2 , additional regularization terms , etc.) are used for minimizing iterative optimization processing, so that the final optimization process can converge to meet the following three loss conditions in sequence: (1) Key point alignment of the objective function based on projection error, such as the alignment of the standardized average face and the two-dimensional detected face, to ensure the minimum difference in face key points between the two; (2) Photometric consistency optimization of the objective function based on photometric consistency error, such as for each pixel point, the lighting direction parameters of the standardized average face and the two-dimensional detected face can conform to a lighting condition model; (3) The regularization term controls the stability and physical rationality of the optimization, such as minimizing the face shape coefficient α and the face expression coefficient β.
[0112] In this way, the joint optimization of the above-mentioned dense face reconstruction task can be realized, so as to ensure that the three-dimensional face reconstruction of the standardized average face can be performed on the input target expression image through the three-dimensional deformation model.
[0113] The target three-dimensional face can be topologically constructed based on the three-dimensional deformation model. This three-dimensional deformation model can learn a large amount of three-dimensional face data and represent any three-dimensional face as a combination of linear bases. Specifically, this three-dimensional deformation model can be expressed as the following formula 4:
[0114]
[0115] Among them, the three-dimensional face shape of the target three-dimensional face can be expressed as n is the number of grid vertices of the three-dimensional face topology structure, is the real number space. Further, can be used to represent the standardized average face (i.e., the three-dimensional average face, Mean Shape), which represents the average shape of the face obtained by statistical learning. represents the face shape basis (Identity Basis), which can be used to represent the principal component basis of individual differences in the face, and s is the number of shape bases. represents the face expression basis (Expression Basis), which can be used to describe the dynamic expression changes of the face, and e is the number of expression bases. represents the face shape coefficient (Identity Coefficients), which can represent the facial features of a specific individual. represents the face expression coefficient (Expression Coefficients), which can be used as a parametric representation of the face expression.
[0116] Through the above joint optimization process for the dense face reconstruction task, the facial expression basis B can be generated. exp And the three-dimensional face reconstruction information of the facial expression coefficient β.
[0117] For the facial expression basis B of the three-dimensional face reconstruction information. exp And the facial shape basis B. id Based on Formula 4 of the above three-dimensional deformation model, a linear combination representation is performed, so that the dense expression motion representation of any facial shape can be realized, and the target three-dimensional face is constructed.
[0118] Therefore, through the construction process of the target three-dimensional face, the three-dimensional face reconstruction of the standardized average face can be performed on the input target expression image through the three-dimensional deformation model, and the decomposition of the three subspaces of the pose, facial shape, and expression of the facial expression motion is realized, corresponding respectively to the pose subspace for representing the global rotation and translation of the facial expression, the facial shape subspace for representing the individual facial shape characteristics, and the expression subspace for representing the dynamic expression changes of the face. Thus, the facial expression of the single-frame two-dimensional expression image is mapped to a parameterized three-dimensional representation. By virtue of the decomposition ability of the facial shape, expression, and pose subspaces, the effective decoupling of the individual facial shape characteristics and expression dynamics is realized, ensuring that the capture ability of subtle expressions is significantly improved through dense expression representation, thereby overcoming the limitations of the traditional sparse feature key point scheme.
[0119] As Figures 2A - 5C shown, according to an embodiment of the present invention, in operation S204, extracting the dense expression motion information of the standardized average face from the target three-dimensional face to complete the perception optimization includes: optimizing the shape parameters of the target three-dimensional face to complete the extraction of the dense expression motion information.
[0120] The three-dimensional face reconstruction of the single-frame target expression image can be realized based on methods such as global photometric consistency optimization and end-to-end reconstruction schemes based on deep learning (such as using CNN and Transformer models). Specifically, like the facial expression basis and facial shape basis of the three-dimensional deformation model mentioned above, the dense and high-precision three-dimensional face shape and its dynamic expression parameters can be restored, so as to realize the accurate modeling of the facial geometry and expression motion.
[0121] According to the joint optimization result of the facial expression basis B exp and the facial expression coefficient β for the dense face reconstruction task of Formula 1 above, the dense expression motion information can be expressed based on Formula 4 of the above three-dimensional deformation model, specifically as the following Formula 5:
[0122] Δv exp = B exp β (5)
[0123] Furthermore, the three-dimensional face shape of the above-mentioned target three-dimensional face can be further expressed by the following formula 6:
[0124]
[0125] Therefore, after the reconstruction of the three-dimensional face shape of the target three-dimensional face is completed, the generated three-dimensional face shape can include the individualized face shape B id α and the facial expression motion information B exp β, and can be specifically expressed by the above formula 4.
[0126] In order to eliminate the influence of the individualized face shape feature B id α and ensure the consistency of the expression semantics, the face shape coefficient α can be set to zero, and only the expression-related part is retained, that is, the optimization of the shape parameters of the target three-dimensional face is completed, and the three-dimensional face shape S on the basis of the standardized average face corresponding to the above formula 4 is generated norm , and this three-dimensional face shape can be used to represent the dense expression grid topology and define the vertex positions of the dense expression grid, and its specific expression is as follows formula 7:
[0127]
[0128] Among them, is the face shape of the standardized average face; B exp β describes the dense expression motion information.
[0129] Therefore, the generated three-dimensional face shape S norm is the three-dimensional expression shape under the three-dimensional standardized average face , and its individual differences have been eliminated, and only the dynamic expression features are retained, so that it can be used as a basis representation with consistent semantic mapping for subsequent expression mapping. Among them, the dense expression motion information is represented by B exp β, so it can be expressed as the following formula 8 as the result of dense expression extraction:
[0130] ΔS exp = B exp β (8)
[0131] Among them, the dense expression motion information is a dense three-dimensional vertex motion field, which can be used to represent the expression offset of each vertex. The significance of extracting the dense expression motion information is that it can provide a fine representation at the vertex level for the entire expression motion, can reflect extremely subtle expression changes, and at the same time eliminates the interference of individual shapes, providing a high-precision input for subsequent expression operations.
[0132] Therefore, in the above-mentioned standardization process and dense expression extraction process, in order to eliminate the influence of individual face shapes on expression description, the shape parameters are fixed to zero in the post-processing stage, and only the expression and pose parameters are retained, thereby generating a dense expression motion under the standardized average face. This process ensures the unity of expression semantics. Therefore, a high-resolution dense expression motion (Dense Expression Motion) information can be generated through the linear combination of the face expression basis and the face expression coefficients of the three-dimensional deformation model, so as to provide an accurate and standardized expression description for the subsequent mapping and execution stages.
[0133] Furthermore, in the perception stage, the face motion is decomposed into three independent subspaces of pose, face shape, and expression through the three-dimensional deformation model, generating a standardized and dense expression motion representation, which not only improves the refinement of expression description, but also eliminates the coupling problem between expression semantics and individual face shapes through the standardized average face topology, achieving the consistency of expression semantics among different individuals.
[0134] In summary, based on the above-mentioned perception optimization method for agent expression imitation in the embodiments of the present invention, in the perception stage of expression imitation, the three-dimensional deformation model can be used as the core technical means to replace the traditional sparse key point-based expression representation method. By decomposing the face motion into three independent subspaces of pose, face shape, and expression, a refined description of the face expression is achieved. In the specific implementation, the above-mentioned three-dimensional face reconstruction information is further calculated for a single-frame image, and then the shape parameters are fixed in the post-processing stage, and only the expression and pose parameters are retained, thereby generating a dense expression motion information under the standardized average face topology. Therefore, it can significantly overcome the abstract defects of the traditional sparse key point method, realize the densification of expression description, and at the same time eliminate the coupling problem between expression semantics and individual face shapes by normalizing the expression to the standardized average face, ensuring the consistency of expression semantics among individuals with different face shapes.
[0135] Compared with the traditional coefficient key point method, the high precision and density of the dense expression motion representation of the three-dimensional deformation model in the embodiments of the present invention can significantly improve the accuracy and fineness of expression perception, laying a high-quality foundation for the subsequent mapping and execution stages. Specifically, in the perception stage of expression, the face expression in the two-dimensional video can be mapped into a parameterized three-dimensional representation by using the three-dimensional deformable model. With its decomposition ability for the face shape, expression, and pose subspaces, the effective decoupling of individual face shape features and expression dynamics is realized. This method significantly improves the ability to capture subtle expressions through dense expression representation, overcoming the limitations of the traditional sparse feature point method.
[0136] Based on the above-mentioned perception optimization method for agent expression imitation, the present invention also provides a perception optimization device for agent expression imitation. The following will be combined withFigure 2B Describe the device in detail.
[0137] Figure 2B A structural block diagram of a perception optimization device for agent expression imitation according to an embodiment of the present invention is schematically shown.
[0138] As Figure 2B shown, the perception optimization device 200 for agent expression imitation in this embodiment includes a motion extraction module 210, a task generation module 220, a joint optimization module 230, and an information extraction module 240.
[0139] The motion extraction module 210 is used to extract the expression motion information of the target expression image. In one embodiment, the motion extraction module 210 can be used to perform the operation S201 described above, which will not be elaborated here.
[0140] The task generation module 220 is used to generate a dense face reconstruction task based on the expression motion information and the preset joint optimization information. In one embodiment, the task generation module 220 can be used to perform the operation S202 described above, which will not be elaborated here.
[0141] The joint optimization module 230 is used to perform joint optimization on the dense face reconstruction task to generate three-dimensional face reconstruction information, and the three-dimensional face reconstruction information is used to generate the target three-dimensional face. In one embodiment, the joint optimization module 230 can be used to perform the operation S203 described above, which will not be elaborated here.
[0142] The information extraction module 240 is used to extract the dense expression motion information of the standardized average face according to the target three-dimensional face to complete the perception optimization. In one embodiment, the information extraction module 240 can be used to perform the operation S204 described above, which will not be elaborated here.
[0143] According to an embodiment of the present invention, any multiple of the motion extraction module 210, the task generation module 220, the joint optimization module 230, and the information extraction module 240 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the motion extraction module 210, the task generation module 220, the joint optimization module 230, and the information extraction module 240 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable means such as hardware or firmware through circuit integration or packaging, or may be implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the motion extraction module 210, the task generation module 220, the joint optimization module 230, and the information extraction module 240 may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions may be executed.
[0144] Embodiment 2
[0145] In the expression mapping stage, existing methods mainly rely on two technical paths: geometric displacement migration based on sparse key points and expression semantic mapping based on image generation. Among them, the method of geometric displacement migration using sparse key points attempts to directly map the key point displacements of the human face to the robot face. However, due to the significant differences in topological structure and geometric characteristics between the human face and the robot face, simple displacement copying or regularization methods based on linear interpolation, affine transformation, etc. cannot effectively transfer expression semantics, and may even cause expression features to be mis-mapped or lost, resulting in serious semantic distortion. In some cases, this method is completely unable to migrate specific expressions. The other image generation method based on a generative model, although visually capable of generating images that conform to expression semantics, only achieves expression consistency in the pixel space and does not consider the physical constraints of the robot actuators, so it cannot be directly used to drive actual physical actuators.
[0146] Therefore, in view of at least one of the technical problems existing in the above-mentioned prior art, embodiments of the present invention aim to provide a mapping optimization method, device, equipment, medium and product for agent expression imitation, thereby providing an optimization scheme for the mapping stage of agent expression imitation, with the expectation of achieving a cross-modal mapping mechanism from standardized human facial expressions to the robot expression space by introducing a unified expression parameter space, solving the differences in geometric structures between human faces and robot faces, and ensuring the consistency of expression semantics during the migration process.
[0147] Based on the Figure 1 described scenario, through Figures 3A - 5C a detailed description of the mapping optimization method for agent expression imitation in the disclosed embodiments will be given.
[0148] As Figure 3A shown, one aspect of the embodiments of the present invention provides a mapping optimization method for agent expression imitation, which includes operations S301 to S303.
[0149] In operation S301, a first blend shape basis of a standardized average face is generated based on a preset facial action coding rule;
[0150] In operation S302, the dense expression motion information of the standardized average face is mapped to the first blend shape basis to generate a second blend shape basis; and
[0151] In operation S303, the blend shape coefficients of the second blend shape basis are transferred to a third blend shape basis of the agent's face to complete the mapping optimization.
[0152] The preset facial action coding rule can be a coding rule for facial behaviors formed based on the facial muscle movement state, so as to realize the coding of facial expressions, which can improve the processing accuracy and efficiency of facial expressions during the expression recognition and processing process. For example, the Facial Action Coding System (FACS for short) can be used to implement it.
[0153] The standardized average face is the three-dimensional average face shape of the static face under the standardized human face topology, which can usually be obtained based on big data statistics. Specifically, reference can be made to the standardized average face in Embodiment 1.
[0154] Based on the face topology of the standardized average face, semantic processing of the expression shape offset is performed through preset facial action coding rules, thereby forming the first blend shape basis. Among them, the first blend shape basis can describe the vertex offset of the corresponding semantic expression (such as "open mouth" or "blink left eye", etc.) relative to the standardized average face, and can be used to describe expression changes through the weighted combination of predefined semantic expression bases. Specifically, it can be achieved through the Blendshape expression representation technology.
[0155] The dense expression motion information can be information directly related to the dynamic changes of the facial expression related to the extracted expression motion information, and can be specifically characterized by the linear combination of the facial expression basis and the facial expression coefficient. The dense expression motion information can most directly reflect the dense expression motion information of the user extracted by the agent on the standardized average face, realizing facial expression representation. Specifically, reference can be made to the extraction of dense expression motion information in Embodiment 1.
[0156] Mapping the dense expression motion information onto the first blend shape basis can achieve mapping the expression motion information of the face of each frame of the target expression image extracted in the perception stage in Embodiment 1 onto the standardized average face. At the same time, the second blend shape basis generated thereby can also exclude the influence of individual face shape features, ensuring the expression semantic consistency in the mapping stage and the execution stage, thereby completing the dense representation of the expression semantics. The second blend shape basis can be a semantic expression basis expressed by the dense expression motion information obtained through perception, and can be used to describe the vertex offset of the semantic expression corresponding to the dense expression motion information relative to the standardized average face. Specifically, it can be achieved through the Blendshape expression representation technology.
[0157] Correspondingly, the third blend shape basis is a semantic expression basis obtained by performing semantic processing of the expression shape offset based on the agent's face topology, and can be used to describe the vertex offset of the corresponding semantic expression relative to the agent's face topology. Among them, the semantics of the third blend shape basis correspond one-to-one and are consistent with the semantics of the above-mentioned second blend shape basis.
[0158] The blend shape coefficient can be understood as the blend shape weight of the second blend shape basis (such as the Blendshape coefficient), and can be used to express the vertex offset of the expression action of the dense expression motion information obtained in the perception stage (refer to Embodiment 1 above). By transferring this blend shape coefficient to the third blend shape basis of the agent's face, direct cross-topology migration of human dense expression motion information can be achieved.
[0159] Therefore, in the mapping stage, the above mapping optimization method for agent expression imitation in the embodiments of the present invention uses a blend shape semantic space based on a set of second blend shape bases of a standardized average face as an intermediate bridge to solve the problems of differences in topological structure and geometric characteristics between human faces and agent faces. Specifically, first, a set of first blend shape bases is defined on the standardized average face for dense representation of expression semantics. The first blend shape bases can then map the dense expression movements generated in the perception stage (refer to Embodiment 1 above) to the blend shape coefficient space through dense semantic key points, forming the second blend shape bases. In this process, by transferring the blend shape coefficients, the expression movements are mapped from the topological structure of the standardized average face to the topological structure of the agent face, realizing semantic-consistent expression transfer across topological structures.
[0160] Therefore, through this conversion from the geometric domain to the semantic domain, the problem of semantic distortion caused by topological differences between human faces and robot faces in traditional solutions is solved. In addition, the number of blend shape coefficients is limited, which is convenient for real-time transmission, thus greatly improving the efficiency and robustness of the mapping stage. At the same time, the generated blend shape coefficients directly correspond to the movement state of the robot's epidermis (such as Figure 5C the shown skin Skin), ensuring the physical executability of the generated expressions. It can be seen that in the cross-entity mapping process of expression imitation, the blend shape bases designed based on the above preset facial action coding rules can be used to accurately depict expression semantics. By introducing a unified expression parameter space, a cross-modal mapping mechanism from the dense human face expressions of the standardized average face to the robot expression space can be established, thus solving the differences in geometric structures between human faces and robot faces and ensuring the consistency of expression semantics during the migration process.
[0161] The blend shape bases are mainly generated based on semantic expression modeling of the blend shape (Blendshape) technology. After the standardization processing and extraction of dense expression movement information in Embodiment 1 above, in order to further realize semantic expression modeling, an expression representation method based on this blend shape base is further introduced.
[0162] The blend shape (Blendshape) technology can be understood as describing expression changes through a weighted combination of a set of predefined semantic expression bases. Therefore, the three-dimensional human face shape S corresponding to the standardized average face represented by the blend shape human can be expressed as the following formula 9:
[0163]
[0164] where S human ∈ can be expressed as a dense grid containing semantic expressions. It can be the three-dimensional average face shape, which is consistent with the definition of the standardized average face in Embodiment 1, representing the static face shape under the standardized human face topology. It can be the i-th Human Blendshape basis, which is used to represent the semantic expression shape offset and can describe the vertex offset of the corresponding semantic expression (such as "open mouth" or "blink left eye") relative to the standardized average face. w i w ∈ [0, 1] can be the weight of the i-th Human Blendshape basis, which can represent the intensity coefficient of this expression. Among them, the weight of this Human Blendshape basis can usually be normalized to the interval [0, 1]. For example, w i = 0 indicates no such expression, and w i = 1 indicates that this expression is fully applied. K is the number of defined blendshape bases, and each blendshape basis B human,i can correspond to a semantic expression action.
[0165] Combined with the mathematical representation of the blendshape basis in Formula 9 above, it can be combined with Figures 3A - 5C to further describe the generation of the first blendshape basis in Embodiment 2 of the present invention as follows.
[0166] As Figures 3A - 5C shown, according to an embodiment of the present invention, in generating the first blendshape basis of the standardized average face based on the preset facial action coding rule in Operation 301, it includes:
[0167] Generating a blendshape expression basis matrix according to the preset facial action coding rule;
[0168] Generating the first blendshape basis of the standardized average face based on the preset blendshape weight vector and the blendshape expression basis matrix.
[0169] As mentioned above, the first blendshape basis can be designed based on the three-dimensional average face shape (i.e., the standardized average face ), which can have the same topology (the connection relationship between vertices and faces) as the standardized average face, and combine the expression semantics defined by the preset facial action coding rule (such as the Facial Action Coding System, abbreviated as FACS), such as "open mouth" or "blink left eye".
[0170] As defined in Formula 9 above, the Human Blendshape basis B that is a component of the first blendshape basis human,i can be designed by a predefined method, specifically based on the standardized average face Construct the facial topology structure. Among them, the first blend shape basis is designed according to the expression semantics of the preset facial action coding rules, and can be divided according to the dominant area to ensure that the semantics of each blend shape basis are clear and consistent. For example, the facial expression action of "raising the right corner of the mouth" can correspond to the movement of an actuator in FACS.
[0171] Specifically, the three-dimensional human face deformation represented by the first blend shape basis can be defined based on the standardized average face and generated by the weighted combination of a set of predefined semantic expression bases. Combining the above formula 9, specifically, the first blend shape basis B human w can be expressed in matrix form in the following formula 10 to form an expression combination based on the blend shape basis:
[0172]
[0173] Among them, the blend shape expression basis matrix can form the semantic human face blend shape basis matrix of the first blend shape basis, which can be specifically composed of K predefined human face expression bases B human,i In addition, the blend shape weight set can be correspondingly expressed as the preset weight vector of the blend shape corresponding to the blend shape expression basis matrix to represent the intensity of each expression. Therefore, through the preset blend shape weight vector w and the blend shape expression basis matrix B human can form the first blend shape basis B corresponding to the standardized average face human w.
[0174] Therefore, by adjusting the blend shape weight w, the contribution of each component blend shape basis in the first blend shape basis can be flexibly controlled, so as to generate expression shapes with different intensities and combinations. This predefined modeling method can ensure that the generated expression shapes have clear semantics, dense geometric details, and a topology structure consistent with the average face.
[0175] It can be seen that through the semantic expression modeling of the blend shape basis, on the facial topology structure (Mesh) of the standardized average face, a set of blend shape bases corresponding to its dense expression movements can be designed based on the facial action coding system for the compact representation of expression semantics (such as 51 standard Blendshapes of ARKit).
[0176] Such as Figures 3A - 5C shown, according to an embodiment of the present invention, before mapping the dense expression movement information of the standardized average face to the first blend shape basis to generate the second blend shape basis in operation 302, it further includes:
[0177] Generate the dense expression motion information of the standardized average face based on the expression motion information of the extracted target expression image.
[0178] As mentioned in Embodiment 1, the target expression image can be a single-frame image containing a human facial expression, and the expression motion information can include the facial expression feature parameters extracted for the facial expression in the target expression image. Correspondingly, the dense expression motion information can be the information related to the dynamic changes of the facial expression directly related to the extracted expression motion information, and can be specifically characterized by the linear combination of the facial expression basis and the facial expression coefficients. Specifically, see Equation 5 in Embodiment 1. The dense expression motion information can most directly reflect the dense expression motion information of the user extracted by the agent on the standardized average face, realizing the facial expression representation.
[0179] Therefore, based on the extraction of the dense expression motion, a facial expression grid S on the basis of the standardized average face corresponding to the dense expression motion information can be constructed norm (as shown in Equation 7 in Embodiment 1), and this facial expression grid can be used as a three-dimensional human face shape with dense expression motion information to represent the dense expression grid topological structure and define the vertex positions of the dense expression grid.
[0180] By means of the above extraction process of the dense expression motion information in the perception stage, the effective decoupling of the individual face shape characteristics and the expression dynamic information can be realized, and the capture ability of the micro-expression can be significantly improved through the dense expression representation.
[0181] In the mapping stage of the expression imitation in Embodiment 2 of the present invention, in order to transfer the dense expression motion generated in the perception stage of Embodiment 1 to the topological structure of the agent's face, realize the cross-topological semantic transfer, and ensure the cross-topological semantic consistency, it is necessary to consider the mapping from the dense expression motion to the blend shape coefficients.
[0182] As Figures 3A - 5C shown, according to an embodiment of the present invention, in operation 302 of mapping the dense expression motion information of the standardized average face to the first blend shape basis to generate the second blend shape basis, it includes:
[0183] Extract the first dense annotation points on the dense expression grid corresponding to the dense expression motion information and the second dense annotation points on the average face grid;
[0184] Obtain the dense annotation basis matrix according to the blend shape expression basis matrix;
[0185] Generate a blend shape weight optimization task according to the dense annotation basis matrix, the first dense annotation points and the second dense annotation points.
[0186] The dense expression grid can be understood as a facial expression grid based on the standardized average face corresponding to the above-mentioned dense expression motion information. (As shown in Equation 7 in Embodiment 1), it is used to represent the vertex positions of the standardized dense expression grid, including n vertices, and each vertex has three-dimensional coordinates (a total of 3n degrees of freedom), which can be understood as the three-dimensional geometric representation of the entire facial network.
[0187] The average face grid can be understood as the standardized average face. The standardized human face grid can represent the three-dimensional vertex positions of the corresponding human face grid.
[0188] The first dense annotation point can be expressed as the position of the standardized dense annotation point. In dense expression modeling, it is usually a specific set of points selected from the above-mentioned dense expression grid S. norm For example, it is distributed in key areas of the face (such as the mouth, eyes, and nose, etc.), and specifically can be expressed as the positions of m dense annotation points selected from the dense expression grid S. norm (A total of 3m degrees of freedom).
[0189] Among them, the position of the first dense annotation point can be extracted from the dense expression grid S through the dense annotation matrix. Specifically, it can be expressed as Equation 11 below: norm L
[0190] L norm = M dense S norm (11)
[0191] Among them, the dense annotation matrix M dense is a sparse selection matrix, which can be used to extract the first dense annotation point of GIA from the dense expression grid S. norm
[0192] Correspondingly, the second dense annotation point can be expressed as the average face position of the dense annotation point. Defined as the dense annotation point extracted from the average face grid. through the dense annotation matrix M. dense Specifically, it can be expressed as Equation 12 below:
[0193]
[0194] Among them, the average face grid can represent the three-dimensional vertex positions of the human face grid (Mesh) defined by the standardized average face.
[0195] Therefore, the correspondence between the dense annotation points and the face mesh can be constructed with the help of the above formulas (11) and (12), where the position L of the first dense annotation point norm and the position of the second dense annotation point are directly related to the dense expression mesh (full mesh expression) S norm and the average face mesh .
[0196] As mentioned above, the mixed shape expression basis matrix B human can be used as the full mesh basis matrix to form the semantic face mixed shape basis matrix of the first mixed shape basis, which can be specifically composed of K predefined expression bases.
[0197] The dense annotation basis matrix can be understood as the mixed shape basis matrix of the dense annotation mainly used to represent the mixed shape basis on the dense annotation points, which can be specifically extracted from the full mesh basis matrix and can be specifically expressed as the following formula 13:
[0198] B dense = M dense B human (13)
[0199] Therefore, according to the dense annotation basis matrix B defined by the above formula 13 dense , the first dense annotation point L defined by formula 11 norm and the second dense annotation point defined by formula 12 , the mixed shape weight corresponding to the above mixed shape weight optimization task can be defined as the optimization problem represented by formula 14 as follows:
[0200]
[0201] where can be understood as the mixed shape weight, which can make the position L of the reconstructed dense annotation points recon satisfy the relationship defined by the following formula 15:
[0202]
[0203] Therefore, based on the mixed shape technology modeling, the geometric information of the dense annotation points can be utilized to solve the optimization problem corresponding to the mixed shape weight optimization task of the above formula 14, obtain the mixed shape weight w, so as to ensure that the position of the reconstructed dense annotation points is as close as possible to the input dense annotation points.
[0204] As Figures 3A - 5CAs shown, according to an embodiment of the present invention, in the operation of mapping the dense expression motion information of the standardized average face to the first blend shape basis to generate the second blend shape basis in operation 302, it further includes:
[0205] Obtaining the blend shape coefficients corresponding to the blend shape weight optimization task by minimizing the preset error vector and the preset coefficient constraint conditions, so as to map the dense expression motion information of the standardized average face to the first blend shape basis to generate the second blend shape basis.
[0206] The preset error vector can be the error vector objective function r(w) for performing optimization processing on the above blend shape weight optimization task, and can be specifically expressed as the following formula 16:
[0207]
[0208] The preset error vector can be the first dense annotation point L norm and the second dense annotation point The error vector based on the dense annotation basis matrix B dense By setting the sum of the squares of the minimum errors of the preset error vector as the optimization objective, the following formula 17 can be defined:
[0209]
[0210] Expanding the above formula 17, it can be expressed as the following formula 18:
[0211]
[0212] The preset coefficient constraint condition can be used as the constraint condition for performing optimization processing on the above blend shape weight optimization task. For example, the value of the blend shape weight can be limited to the range of 0 ≤ w ≤ 1, so as to satisfy that the blend shape coefficient corresponding to the final blend shape weight optimization task is 0 ≤ w i ≤ 1, f Therefore, the above blend shape weight optimization task corresponds to a quadratic programming problem (quadratic programming) with boundary constraints (preset coefficient constraint conditions).
[0213] Therefore, by means of the solution process of the least squares problem of the above blend shape coefficients, the dense expression motion information can be mapped to the low-dimensional first blend shape basis in real time to form the second blend shape basis, that is, by selecting dense semantic key points (based on barycentric coordinates, rather than vertex coordinates) on the surface of the average face mesh, the dense expression motion generated in the perception stage is mapped to the blend shape coefficient space, thereby completing the coefficient mapping from the dense expression to the blend shape space.
[0214] In summary, through the least squares problem of an overdetermined constraint in the above-mentioned hybrid shape weight optimization task, the dense expression motion information can be compressed into a low-dimensional blendshape space. This process can generate blendshape coefficients in real time while ensuring the density and accuracy of expressions.
[0215] As Figures 3A - 5C shown, according to an embodiment of the present invention, before passing the blendshape coefficients of the second blendshape basis to the third blendshape basis of the agent's face in operation 303 to complete the mapping optimization, it further includes:
[0216] Generating a third blendshape basis of the agent's face based on a preset facial action coding rule, where the third blendshape basis has the same blendshape semantics as the first blendshape basis.
[0217] The third blendshape basis is actually built for the blendshape space of the agent's face. Among them, the agent's blendshape space needs to ensure semantic consistency with the first blendshape basis space corresponding to the dense expression grid S of the human face norm so that the agent can accurately express the same basic expression semantics as humans (such as "opening the mouth", "blinking", etc.).
[0218] Similar to the first blendshape basis of humans, according to the principle of semantic consistency, the third blendshape basis space of each agent corresponds one-to-one and is consistent in semantics with the first blendshape basis space of the standardized average face (for example, "opening the mouth" should express the same semantics on the facial topology of the agent and in the blendshape basis space of the human standardized average face model).
[0219] Therefore, according to the semantic set of the first blendshape basis space of the human face (such as "opening the mouth", "blinking the left eye", etc.), based on the above-mentioned preset facial action coding rule, a corresponding set of third blendshape bases can also be defined for the agent based on the facial topology of the agent, and the semantic set in the third blendshape basis space is its blendshape basis of the agent that conforms to each semantics where r is the number of mesh vertices in the facial topology of the agent, and k can be the semantic number.
[0220] Therefore, for the facial topology of the agent, based on the first blendshape basis of the standardized average face corresponding to the dense expression motion information, a third blendshape basis of the agent's face with consistent semantics can be constructed to ensure that the semantic of the two blendshape basis spaces is completely consistent, thereby ensuring efficient and accurate dense expression transfer from the standardized average face to the facial topology of the agent.
[0221] AsFigures 3A - 5C As shown, according to an embodiment of the present invention, in operation 303, transmitting the blend shape coefficients of the second blend shape base to the third blend shape base of the agent's face to complete the mapping optimization includes:
[0222] Applying the blend shape coefficients to the third blend shape base to generate the agent expression mesh corresponding to the agent.
[0223] Directly transmitting the blend shape coefficients generated from the standardized average face to the third blend shape base corresponding to the agent's face topology can achieve dense expression cross-topology semantic transfer, generate the agent expression mesh, and thus complete the mapping process of the blend shape coefficients.
[0224] Among them, the core of transmitting the blend shape coefficients is to directly apply the blend shape coefficients w human generated from the human standardized average face to the third blend shape base B robot of the agent to achieve dense expression cross-topology semantic transfer.
[0225] Specifically, through the third blend shape base B robot of the agent's face topology with semantic consistency, the semantics of each basic expression are completely aligned with the first blend shape base of the human standardized average face topology. Therefore, the blend shape coefficients w human of the first blend shape base can be directly transmitted to the third blend shape base without any conversion, satisfying the following formula 19:
[0226] w robot = w human (19)
[0227] Therefore, applying the blend shape coefficients w human of the first blend shape base to the agent's face topology and the agent's third blend shape base B robot can generate the agent expression mesh S robot corresponding to the agent, expressed as the following formula 20:
[0228]
[0229] Therefore, it can ensure that the semantics of the robot's expression are consistent with those of the human expression, and at the same time adapt to the geometric characteristics of the agent's face topology. Specifically, reference can be made to the semantic alignment operation from the second blend shape base to the third blend shape base as shown in Figure 5A the figure.
[0230] In summary, based on the above mapping optimization method for agent expression imitation provided in Embodiment 2 of the present invention, a hybrid shape semantic space based on a standardized average face can be provided as an intermediate bridge during the mapping stage to address the differences in topological structure and geometric characteristics between the human face and the agent's face. Specifically, a set of hybrid shape bases is first defined on the standardized average face for dense representation of expression semantics. These hybrid shape bases map the dense expression movements generated during the perception stage to the hybrid shape coefficient space through dense semantic key points. This process maps the dense expression movements to low-dimensional hybrid shape coefficients in real time by solving a least squares problem.
[0231] After that, a set of hybrid shape bases consistent with the average face semantics is provided based on the agent's face to ensure the same semantic definition for both. By transmitting the hybrid shape coefficients, the dense expression movements can be mapped from the standardized average face to the agent's face, thus achieving efficient and accurate cross-topological structure and semantic-consistent expression transfer.
[0232] Therefore, by converting from the geometric domain to the semantic domain, the problem of semantic distortion caused by the topological differences between the human face and the agent's face in traditional methods is solved. In addition, the number of hybrid shape coefficients is limited, facilitating real-time transmission and greatly improving the efficiency and robustness of the mapping stage. At the same time, the generated hybrid shape coefficients directly correspond to the motion state of the agent's epidermis, ensuring the physical executability of the generated expressions.
[0233] In the cross-entity mapping of expressions, the hybrid shape bases designed based on the preset facial action coding rules are used to accurately depict expression semantics. By introducing a unified expression parameter space, a cross-modal mapping mechanism from the standardized human face expression to the agent expression space is established, solving the differences in geometric structure between the human face and the agent's face and ensuring the consistency of expression semantics during the transfer process.
[0234] Therefore, the above mapping optimization method for agent expression imitation provided in Embodiment 2 of the present invention can at least achieve the following technical effects:
[0235] (1) Directness: No additional conversion or mapping is required for the human hybrid shape coefficient w human to be carried out.
[0236] (2) Efficiency: Utilizing the semantic-consistent hybrid shape bases, the cross-topological transfer of expressions can be quickly realized.
[0237] (3) Versatility: Applicable to any agent hybrid shape bases that conform to the semantic consistency design.
[0238] Based on the above mapping optimization method for agent expression imitation, the present invention also provides a mapping optimization device for agent expression imitation. The following will be combined withFigure 3B Describe the device in detail.
[0239] Figure 3B A structural block diagram of a mapping optimization device for agent expression imitation according to an embodiment of the present invention is schematically shown.
[0240] As Figure 3B shown, the mapping optimization device 300 for agent expression imitation in this embodiment includes a base generation module 310, an information mapping module 320, and a coefficient transfer module 330.
[0241] The base generation module 310 is used to generate a first blend shape base of a standardized average face based on a preset facial action coding rule. In one embodiment, the base generation module 310 can be used to perform the operation S301 described above, which will not be elaborated here.
[0242] The information mapping module 320 is used to map the dense expression motion information of the standardized average face to the first blend shape base to generate a second blend shape base. In one embodiment, the information mapping module 320 can be used to perform the operation S302 described above, which will not be elaborated here.
[0243] The coefficient transfer module 330 is used to transfer the blend shape coefficients of the second blend shape base to a third blend shape base of the agent's face to complete the mapping optimization. In one embodiment, the coefficient transfer module 330 can be used to perform the operation S303 described above, which will not be elaborated here.
[0244] According to an embodiment of the present invention, any multiple of the base generation module 310, the information mapping module 320, and the coefficient transfer module 330 can be combined and implemented in one module, or any one of them can be split into multiple modules. Or, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the base generation module 310, the information mapping module 320, and the coefficient transfer module 330 can be at least partially implemented as a hardware circuit, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application-specific integrated circuit (ASIC), or can be implemented by any other reasonable means such as integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Or, at least one of the base generation module 310, the information mapping module 320, and the coefficient transfer module 330 can be at least partially implemented as a computer program module, and when the computer program module is run, it can perform the corresponding functions.
[0245] Embodiment 3
[0246] During the execution phase, traditional methods convert the epidermal motion representation into actuator commands in two ways. One is the model-based inverse kinematics method, which converts the motion of the epidermal traction points into specific actuator commands through inverse kinematics solving. However, this method has extremely high requirements for the accuracy of the geometric model and needs to accurately describe the topological structure of the robot's face and the actuator positions. Otherwise, it is difficult to ensure the accuracy of inverse kinematics solving. The other method is the model-free neural networks method based on deep learning, which directly calculates the actuator commands from the epidermal motion (such as pixel space) using neural networks. Although this method has higher flexibility, the generated execution commands often do not fully consider the physical constraints of the robot actuators, resulting in difficulties in implementing the commands in practical applications and causing problems with the control accuracy and stability of expressions.
[0247] In view of at least one of the technical problems existing in the above-mentioned execution schemes for expression imitation, embodiments of the present invention provide an execution optimization method, device, equipment, medium, and product for agent expression imitation, aiming to optimize the expression driving parameters by explicitly modeling the physical constraints of the agent actuators and ensure that the generated control commands are both physically executable and meet the requirements of high precision and real-time performance.
[0248] The following will be based on Figure 1 the described scenario, and through Figures 4A - 5C detailed description of the execution optimization method for agent expression imitation of the disclosed embodiments.
[0249] As Figure 4A shown, one aspect of the embodiments of the present invention provides an execution optimization method for agent expression imitation, which includes operations S401 to S403.
[0250] In operation S401, an inverse kinematics target optimization task is generated based on the control point simulation information of the agent expression grid and the target epidermal vertex information;
[0251] In operation S402, the inverse kinematics target optimization task is optimized to obtain the optimal joint angle parameters corresponding to the control point simulation information; and
[0252] In operation S403, the execution optimization of agent expression imitation is performed according to the optimal joint angle parameters.
[0253] The agent expression grid may include the vertex relationships of the facial topology grid (Mesh) structure of the agent mentioned in Embodiment 2 above. Among them, the agent expression grid is the robot facial topology that has completed the cross-topology migration of dense expression motion information in the perception stage in Embodiment 1, and specifically can be embodied as the expression grid S in Formula 20 in Embodiment 2. robot .
[0254] As Figure 5B shown, there will be many key control points on the agent expression grid, such as eyebrows, eyelids, eyeballs, nose, cheeks, mouth, and jaw, etc. These control points can be driven by cables / PTFE ducts as Figure 5C shown through motors (such as LFD-01M servo motors) to achieve traction, so as to drive the deformation of the epidermis in the surrounding area of the control points (such as Figure 5C shown Silicone Mask). In the traditional technical solution, it is very difficult to calculate the deformation of the surrounding area of the control points, resulting in a rigid imitation effect of the agent's expression and unable to achieve a delicate and natural "skin" performance.
[0255] The control point simulation information can be the position information formed by performing control point position simulation calculations on the agent expression grid containing dense expression motion information through a simulation tool (such as Blender), for example, the simulation positions of each control point in the topological network in the agent expression grid with dense expression motion information. The control point simulation information can be used to define the drive angle parameters of the drive motor rocker arms of the cables / PTFE ducts as Figure 5C shown, that is, the joint angle parameters.
[0256] The target epidermis vertex information can be the target epidermis grid vertex positions that each control point on the agent expression grid should reach when achieving the human face expression reflected by the dense expression motion information. The target epidermis vertex information can include the simulation control point positions of the control point simulation information obtained through simulation calculations and the information on the relative displacement distance and displacement direction of the epidermis grid vertices in the surrounding area, and can be used to define the epidermis shape (i.e., the target epidermis shape) that should be achieved when driving the agent facial grid relative to the dense expression motion information.
[0257] The inverse kinematics target optimization task can be an optimization task that continuously optimizes the joint angle parameters corresponding to the control point simulation information according to the target epidermis vertex information. Through the continuous optimization of this inverse kinematics target optimization task, during actual driving execution, based on the optimization result of this inverse kinematics target optimization task, the movement of the intelligent agent's facial epidermis mesh can be made as close as possible to the target epidermis shape defined by the dense expression movement information. This inverse kinematics target optimization task can be expressed in the form of an inverse kinematics objective function.
[0258] Performing the optimization of the above joint angle parameters on the inverse kinematics target optimization task can make the intelligent agent's epidermis mesh as close as possible to the target epidermis shape until the optimal joint angle parameters are obtained. When these optimal joint angle parameters are used as driving parameters by the corresponding driving motors in the intelligent agent's facial structure as shown Figure 5C below to perform epidermis driving, it can make the corresponding driven control points and their surrounding epidermis areas present the most fitting expression imitation effect with the dense expression movement information, with the expression being more delicate and natural, and at the same time ensuring that the skin around the control points presents the most natural epidermis changes.
[0259] Therefore, the optimal joint angle parameters can be the motor rocker joint angle parameters that can make the control point and its surrounding epidermis area reach the closest state to the target epidermis deformation when the control point is driven by the motor. As shown Figure 5C below, through the "Skeleton" architecture to support the "Skin", "Muscle" architecture of the silicone mask and the driving architecture of the Electrical Control, where the electrical control architecture includes various driving motors or servo motors, these motors can perform rotational or even pulling actions according to the optimal joint angle parameters under the control of the controller, thereby driving the connected steel cable / Teflon conduit to generate the optimal displacement in the specified direction at the specified position (control point) on the connected silicone mask ("skin"), reaching the optimal position of the control point under the target expression, thereby generating the epidermis deformation of the intelligent agent's face, completing the execution optimization of the expression action, and realizing the natural deformation of the intelligent agent's facial expression. Among them, the control point can be the fixed point (such as adhesive fixation) of the above steel cable / Teflon conduit on the inner surface of the silicone mask. In addition, the above "Skeleton" architecture, "Skin", "Muscle" architecture and the driving architecture of the Electrical Control can constitute the actuator of the intelligent agent.
[0260] Among them, since the optimal joint angle parameters are obtained by optimizing the inverse kinematics target optimization task based on the agent's expression grid of dense motion expression information, the execution process of the above expression actions corresponds to the execution process of the agent's expression imitation. In short, the generation of the agent's facial expression depends on the kinematic modeling of the subcutaneous control points. At the same time, the joint angle parameters obtained by solving the inverse kinematics are used to drive the movement of the control points, so as to transmit the dense expression motion information obtained in the perception stage to the vertices of the agent's facial epidermis grid, ensuring the natural deformation of the expression restoration process.
[0261] Therefore, through the above-mentioned execution optimization method of agent expression imitation in Embodiment 3 of the present invention, by defining the control point area, the inverse kinematics can be used to solve the optimal position of the control points under the target expression and generate accurate actuator motion instructions, so that the motion state of the control point area approaches the target epidermis motion. In this case, only by increasing the density of the control points can the fineness and naturalness of the expression restoration be further improved without changing the algorithms in the perception and mapping stages. Thus, not only is the accuracy of expression imitation greatly improved, but also the physical executability of the generated instructions is ensured by explicitly modeling the physical constraints of the actuator, avoiding the execution failure problem caused by ignoring the physical constraints in the traditional method. In addition, the independence of the control point optimization also enables the execution stage to flexibly adapt to different types of agent facial structures, providing technical support for the scalability and modular design of the system, and having extremely high engineering application value and commercial application value.
[0262] Such as Figures 4A - 5C shown, according to an embodiment of the present invention, before generating the inverse kinematics target optimization task based on the control point simulation information and the target epidermis vertex information of the agent expression grid in operation S401, it further includes:
[0263] Generating an agent expression grid corresponding to the agent according to the blend shape basis of the agent's face;
[0264] Generating a distance weight matrix based on the control point simulation information and the target epidermis vertex information of the agent expression grid.
[0265] The blend shape basis of the agent's face can be a semantic expression space for surface shape offset semantic processing based on the agent's facial topology structure, which is used to describe the vertex offset of the semantic expression corresponding to the dense expression motion information relative to the agent's facial topology structure. Specifically, it can refer to the third blend shape basis B provided in the above-mentioned Embodiment 2. robot In the dense expression motion information extracted in the perception stage, through the mapping process of the blend shape coefficients in the mapping stage, it can be cross-topologically semantically migrated to the blend shape basis of the above-mentioned agent's face.
[0266] Furthermore, an agent expression grid S can be constructed based on the blend shape basis of the agent's face robot , which can be specifically reflected in Formula 20 in the above-mentioned Embodiment 2. In this way, it can ensure that the semantics of the robot's expression are consistent with human expressions, and at the same time, it can adapt to the geometric characteristics of the agent's facial topology, ensuring that it can be aligned with the semantics of the perceived human face expression, so as to ensure the accuracy of subsequent physical drive execution, guarantee the natural deformation during the expression restoration process, and at the same time ensure a better restoration degree of the perceived human face expression.
[0267] Each constituent matrix element w in the distance weight matrix ji can be used to represent the distance distribution between a single control point p on the agent expression grid i and the corresponding epidermal vertex v j . Therefore, the distance weight matrix can define the distance-based weight distribution between the epidermal control points and the epidermal grid vertices. Among them, these matrix elements w ji can be specifically expressed as the following Formula 21:
[0268]
[0269] where, ∥v j -p i ∥ is the Euclidean distance between the epidermal vertex v j and the control point p i ; ∥v j -p k ∥ is the distance from the epidermal vertex to the control point; p k is the three-dimensional position coordinate of the subcutaneous controller; σ is the control parameter of the distance influence, which is used to adjust the attenuation range of the weight; weight normalization ensures (the sum of the weights of each epidermal vertex is 1), and c is the number of control points.
[0270] Among them, a single control point p i can be the position information formed by performing position simulation calculation on the agent expression grid through a simulation tool, and the set of these control points can constitute the above-mentioned control point simulation information. Among them, each control point p i can be used to define the three-dimensional coordinates of the subcutaneous traction position. Correspondingly, a single epidermal vertex v j can be the position expected by the dense expression motion information reached by the corresponding grid position vertex, and the set of these epidermal vertices can constitute the above-mentioned target epidermal vertex information.
[0271] Therefore, based on the control point simulation information and the target skin vertex information, the dense expression motion information of the agent's expression mesh can be mapped onto the topological structure of the agent's facial structure, and the accurate description of the agent's facial topology and the actuator positions can be achieved through the distance weight matrix.
[0272] As Figures 4A - 5C shown, according to an embodiment of the present invention, before generating the distance weight matrix based on the control point simulation information and the target skin vertex information of the agent's expression mesh, it further includes:
[0273] Generating control point simulation information according to the preset forward kinematics optimization task of the agent's expression mesh and the joint angle parameter vector corresponding to the simulation control points;
[0274] Generating target skin vertex information according to the deformation weight matrix of the skin mesh vertices of the agent's expression mesh and the control point simulation information.
[0275] Based on constructing the geometric model and kinematic structure (such as joint control points and their associated relationships) of the agent's face corresponding to the agent's expression mesh S robot by means of a simulation tool, the preset forward kinematics optimization task can be expressed as a forward kinematics function based on the motor kinematic joint angle parameters corresponding to the agent's kinematic structure.
[0276] The simulation control points can be subcutaneous control points defined by the above-mentioned geometric model and kinematic structure of the agent's face. The set of subcutaneous control points is defined as P = {p 1 , p 2 , …, p c}, and each control point represents the three-dimensional coordinates of a certain position under the skin, and c is the number of control points. Each simulation control point corresponds to different joint angle parameters, and these joint angle parameters can be used as sub-elements of the above-mentioned joint angle parameter vector.
[0277] Therefore, through the simulation tool, according to the input parameter vector such as joint angles, etc., combined with the kinematic relationship and geometric constraints, the forward kinematic information of the control point positions can be generated as the control point simulation information.
[0278] Specifically, for the implementation operation of the forward kinematics of the above-mentioned control point simulation information, the geometric model and kinematic structure of the agent's face, including joints, control points and their associated relationships, can be first constructed through a simulator (such as Blender), and then the simulator can directly calculate the position P of the control points according to the input joint angle parameter θ, combined with the kinematic relationship and geometric constraints.
[0279] Furthermore, the position of the control point corresponding to the control point simulation information is driven by the kinematic joint angle parameter θ, and the forward kinematic model of the control point simulation information is defined by the following formula 22:
[0280] P = f(θ) (22)
[0281] where, is the forward kinematic function; is the joint angle parameter vector; d is the number of kinematic degrees of freedom; c is the number of control points. Thus, the construction of the control point model can be completed based on the control point simulation information.
[0282] The skin mesh vertices of the agent expression mesh can be the vertices of the face mesh based on the geometric model and kinematic structure of the agent's face corresponding to the agent expression mesh S robot Specifically, they can be expressed as a set of skin mesh vertices V = {v 1 , v 2 , …, v r}, where each skin mesh vertex represents the three-dimensional coordinates of a certain point on the skin, and r is the number of skin mesh vertices.
[0283] Correspondingly, the movement of the control points corresponding to the control point model P of the control point simulation information can drive the deformation of the corresponding skin mesh vertices v through the deformation weight matrix j . Among them, the deformation weight matrix can represent the degree of deformation influence of the control points on the skin mesh vertices. Specifically, the target skin vertex information can be expressed as a linear relationship through the deformation weight matrix and the control point simulation information as the following formula 23:
[0284] V = WP (23)
[0285] where, W is the deformation weight matrix, representing the influence relationship of the control points P on the skin vertices V; is the global position of the control points, stacked in sequence as is the global position of the skin mesh vertices, stacked in sequence as Therefore, the construction of the skin mesh vertex model can be completed. Among them, the deformation weight matrix can be a weight construction model of the distance weight matrix, used to generate the distance weight matrix according to the agent expression mesh.
[0286] Thereby, the geometric models of the control points and their surrounding areas, as well as the vertices of the grid skin, can be accurately defined, ensuring that the optimal positions of the control points under the target expression can be calculated subsequently, generating precise actuator motion instructions, and ensuring that the motion state of the control point area is close to the target skin motion. Moreover, it is also beneficial to further improve the fineness and naturalness of expression restoration by increasing the density of control points, without changing the algorithms in the perception stage and the mapping stage.
[0287] As Figures 4A - 5C shown, according to an embodiment of the present invention, in the operation S401 of generating an inverse kinematics target optimization task based on the control point simulation information of the agent expression grid and the target skin vertex information, it includes:
[0288] The extended distance weight matrix is the weight amplification matrix;
[0289] Generating control point-driven skin grid vertices according to the control point simulation information and the weight amplification matrix;
[0290] Generating an inverse kinematics target optimization task according to the control point-driven skin grid vertices and the target skin vertex information.
[0291] The weight amplification matrix can be formed by three-dimensionally expanding each sub-element in the aforementioned distance weight matrix. Among them, since each control point p i and the skin grid vertex v j both have three-dimensional coordinates (x, y, z), the dimension of the weight amplification matrix W can be 3r×3c. The specific form is as formula 24 below:
[0292]
[0293] Among them, each matrix element w ji represents the influence degree of the corresponding control point p i on the skin grid vertex v j extended to each coordinate component (x, y, z).
[0294] The inverse kinematics target optimization task can be embodied as an objective optimization function for optimizing the joint angle parameters corresponding to the control points in combination with the position information of the target skin vertices, so as to realize the inverse kinematics solution of the targeted joint angle parameters to obtain the optimal joint angle parameters required for the control points to control the skin.
[0295] Specifically, the goal of the inverse kinematics target optimization task is to optimize the joint angle parameter θ according to the target skin shape V target so that the skin grid is as close as possible to the target shape. Among them, the inverse kinematics target optimization task can be defined in the form of an objective function as formula 25 below:
[0296]
[0297] Among them, is the position of the target skin mesh vertex defined by the target skin vertex information; is the position of the control point calculated by simulation defined by the foregoing control point simulation information; Wf(θ) is the skin mesh vertex generated by driving the control point, and can be linearly combined by the control point simulation information and the weight amplification matrix.
[0298] For example, Figures 4A - 5C as shown, according to an embodiment of the present invention, in operation S402, when performing optimization processing on the inverse kinematics target optimization task to obtain the optimal joint angle parameter corresponding to the control point simulation information, it includes:
[0299] Performing a minimization target optimization iteration process on the inverse kinematics target optimization task according to the joint angle constraint conditions to generate the optimal joint angle parameter.
[0300] The joint angle constraint conditions may include joint angle limit constraints and geometric constraints of the kinematic structure of the agent's face. Among them, the joint angle limit constraints can define the activity range of the joint angle between the minimum joint angle and the maximum joint angle of the motor rocker arm, and the geometric constraints of the kinematic structure may include, for example, the pulling range of the steel cable / PTFE catheter, the deformation range of the "skin" of the silicone mask, etc. By performing a minimization target iteration process (such as non-linear optimization) on the inverse kinematics target optimization task of the above formula 25, it is realized that the controlled deformation position of the control point continuously approaches the optimal control point position under the target expression, where the joint angle of the motor rocker arm corresponding to this optimal control point position (for example, relative to the starting joint angle) can be used as the optimal joint angle parameter θ * .
[0301] Therefore, by combining the above kinematic modeling and inverse kinematics calculation, for the control point area defined by the agent's face, the actuator command can be solved through inverse kinematics under physical constraints to approximate the target skin movement generated in the mapping stage.
[0302] For example, Figures 4A - 5C as shown, according to an embodiment of the present invention, in operation S403, when performing execution optimization of the agent's expression imitation according to the optimal joint angle parameter, it includes:
[0303] Obtaining the control point driving information for the agent's expression imitation through the control point simulation information and the optimal joint angle parameter;
[0304] Performing the driving of the agent's expression imitation according to the control point driving information to complete the execution optimization.
[0305] The forward kinematic function f(θ) corresponding to the control point simulation information obtained by the simulation tool, and the optimal joint angle parameter θ obtained after further optimizing the above inverse kinematic target optimization task * , the driving information of the control points for the intelligent agent's facial motor control system to achieve the intelligent agent's expression imitation can be defined as Formula 26 below:
[0306] P * = f(θ * ) (26)
[0307] Among them, when the driving information of the control points can define the deformation position of the corresponding control points on the facial epidermis and the deformation position information of the epidermal grid in the surrounding area when the motor rocker arm to be controlled by the motor control system is driven according to the above optimal joint angle parameter.
[0308] Specifically, the epidermal grid shape can be generated by driving the corresponding control points, and the epidermal grid shape can be expressed as Formula 27 below:
[0309] V = WP * (27)
[0310] Therefore, as Figure 5C shown, the driving information of the control points defined by the above Formulas 26 and 27 can be converted into instructions executable by a physical actuator (such as a motor), so as to realize the fitting of the epidermal drive and the control points, so that the blend shape coefficients based on the dense expression motion information generated in the mapping stage can be converted into physically executable intelligent agent facial epidermal motions, and the precise expression imitation can be realized by driving the epidermal motion by the actuator, and the high-precision physical drive optimization for the expression imitation can be realized.
[0311] In summary, according to the execution optimization method for intelligent agent expression imitation in the embodiments of the present invention, in the execution stage, an optimization driving means based on control points can be constructed for the physical characteristics of the intelligent agent actuator. First, the control point area is defined, and the inverse kinematics is used to solve the optimal position of the control points under the target expression, and precise actuator motion instructions are generated to make the motion state of the control point area approach the target epidermal motion. By increasing the density of the control points, the fineness and naturalness of the expression restoration can be further improved without changing the algorithms in the perception and mapping stages.
[0312] This method not only improves the accuracy of expression imitation, but also ensures the physical executability of the generated instructions by explicitly modeling the physical constraints of the actuator, avoiding the execution failure problem caused by ignoring the physical constraints in the traditional method. In addition, the independence of the control point optimization enables the execution stage to flexibly adapt to different types of robot facial structures, providing technical guarantees for the scalability and modular design of the system.
[0313] In summary, in the expression driving stage, by explicitly modeling the physical constraints of the robot actuator and optimizing the expression driving parameters, it is ensured that the generated control instructions are both physically executable and meet the requirements of high precision and real-time performance.
[0314] Based on the above execution optimization method for agent expression imitation, the present invention also provides an execution optimization device for agent expression imitation. The following will be combined with Figure 4B to describe this device in detail.
[0315] Figure 4B The structural block diagram of the execution optimization device for agent expression imitation according to Embodiment 3 of the present invention is schematically shown.
[0316] As Figure 4B shown, the execution optimization device 400 for agent expression imitation in this Embodiment 3 includes a target task generation module 410, a task optimization module 420, and an imitation execution module 430.
[0317] The target task generation module 410 is used to generate an inverse kinematics target optimization task based on the control point simulation information of the agent expression grid and the target epidermal vertex information. In one embodiment, the target task generation module 410 can be used to perform the operation S401 described above, which will not be elaborated here.
[0318] The task optimization module 420 is used to perform optimization processing on the inverse kinematics target optimization task to obtain the optimal joint angle parameters corresponding to the control point simulation information. In one embodiment, the task optimization module 420 can be used to perform the operation S402 described above, which will not be elaborated here.
[0319] The imitation execution module 430 is used to perform execution optimization of agent expression imitation according to the optimal joint angle parameters. In one embodiment, the imitation execution module 430 can be used to perform the operation S403 described above, which will not be elaborated here.
[0320] According to an embodiment of the present invention, any multiple of the target task generation module 410, the task optimization module 420, and the imitation execution module 430 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the target task generation module 410, the task optimization module 420, and the imitation execution module 430 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the target task generation module 410, the task optimization module 420, and the imitation execution module 430 may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions may be executed.
[0321] Combining the above embodiments 1-3, it can be seen that: the existing intelligent agent expression imitation technology has significant technical limitations in the three key stages of perception, mapping, and execution. These problems directly limit the realization of high-precision expression imitation, especially in cross-entity expression migration and physical drive implementation. Among them, the perception stage faces the problems of sparse expression representation and coupling between expression and face shape; it is difficult to solve the problem of cross-topology expression migration in the mapping stage, and the expression semantics are distorted or even lost during migration; there is a problem that the generated instruction is physically unexecutable in the execution stage. These limitations restrict the fineness and robustness of expression imitation.
[0322] Therefore, in order to overcome the technical bottlenecks existing in the three key stages of perception, mapping, and execution of the existing expression imitation technology, the present invention provides a brand-new expression imitation technical solution, covering three links of perception optimization, mapping optimization, and execution optimization respectively, aiming to achieve higher-precision and more robust expression perception and migration, while ensuring that the generated control instructions are physically executable, laying a technical foundation for the development of highly biomimetic expression robots. Among them, through the densification of expression representation in the perception stage, cross-topology semantic consistency migration in the mapping stage, and high-precision physical drive optimization in the execution stage, the core problems such as sparse expression description, semantic migration distortion, and unexecutable execution in traditional expression imitation methods are systematically solved. The following is the core design of the new technical solution and its corresponding relationship with the existing problems.
[0323] In the perception stage, a three-dimensional deformation model (3DMM) is mainly used to replace the traditional sparse key-point representation method. The facial motion is decomposed into three independent sub-spaces: pose, face shape, and expression, generating a standardized and dense expression motion representation. That is, dense expression motions are extracted from the input facial images to generate high-precision expression representations under the standardized average face. This method not only improves the refinement degree of expression description but also eliminates the coupling problem between expression semantics and individual face shapes through the standardized average face topology, achieving the consistency of expression semantics among different individuals. For details, see Embodiment 1.
[0324] In the mapping stage, a set of Blendshape semantic spaces based on the standardized average face is provided as an intermediate bridge between the human face and the robot's face, solving the problems of differences in topological structure and geometric characteristics between the two. The dense expression motions generated in the perception stage are real-time mapped to the low-dimensional Blendshape coefficient space, and distortion-free transfer of expressions from the human face to the robot's face is achieved through the semantically consistent Blendshape basis. In other words, through the semantically consistent Blendshape basis, the dense expression motions are mapped to the robot's face topological structure to achieve cross-topology expression transfer. This method not only ensures the consistency of expression semantics but also greatly improves the efficiency and robustness of the mapping. For details, see Embodiment 2.
[0325] In the execution stage, through high-precision physical drive optimization based on control points, the inverse kinematics is used to solve the optimal positions of the control points under the target expression, generating precise actuator motion instructions. That is, the mapping results are converted into physically executable robot skin motion instructions, and precise expression imitation is achieved through inverse kinematics and physical constraint optimization. This method explicitly models the physical constraints of the actuators to ensure the physical executability of the generated instructions. For details, see Embodiment 3.
[0326] The process provided in Embodiments 1-3 of the present invention is optimized layer by layer from perception to execution, ensuring the fineness, semantic consistency, and physical executability of expression imitation, providing a complete solution for the realization of highly biomimetic expression robots. Through systematic design and optimization, this process can achieve full-link expression imitation from the original human face input to the precise reproduction of robot expressions.
[0327] In short, Embodiments 1-3 of the present invention focus on a high-bionic expression robot real-time imitation system based on two-dimensional video, aiming to accurately capture the dynamic features of human facial expressions and achieve precise mapping and physical driving to a high-bionic expression robot. The technical framework covers three links: perception, migration, and execution, breaking through the technical bottlenecks in subtle expression capture, cross-modal expression migration from human facial expressions to robot facial expressions, and physical driving in traditional methods. This real-time imitation system shows broad prospects in application scenarios. For example, in a remote embodied interaction system, an operator can remotely control a high-bionic expression robot through video input to achieve embodied emotional interaction across spaces, which is applicable to scenarios that require emotional expression such as remote medical escort and psychological counseling. In the basic research of social emotion computing, this technology provides solid technical support for robot autonomous emotional expression and human-robot emotional interaction, promoting the further development of social emotion robot research. Generally speaking, through an innovative technical optimization framework for expression perception, mapping, and execution, the key technical problems in the expression migration process are systematically solved, and the real-time and precise driving of a high-bionic expression robot is successfully achieved. This technology not only provides reliable technical support for remote emotional interaction but also lays an important foundation for the development of future autonomous emotional interaction robots, with extremely high commercial application value.
[0328] Figure 6 A block diagram of an electronic device suitable for implementing a perception optimization method, a mapping optimization method, and an execution optimization method for agent expression imitation according to Embodiments 1-3 of the present invention is schematically shown.
[0329] The above-mentioned electronic device provided by the embodiment of the present invention includes one or more processors and a memory. The memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned perception optimization method, mapping optimization method, and execution optimization method for agent expression imitation.
[0330] As Figure 6 shown, the electronic device 600 according to the embodiment of the present invention includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage section 608 into a random access memory (RAM) 603. The processor 601 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), etc. The processor 601 can also include on-board memory for caching purposes. The processor 601 can include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of the present invention.
[0331] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method flow according to the embodiments of the present invention by executing programs in the ROM 602 and / or the RAM 603. It should be noted that the programs may also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 may also perform various operations of the method flow according to the embodiments of the present invention by executing programs stored in the one or more memories.
[0332] According to an embodiment of the present invention, the electronic device 600 may further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the I / O interface 605: an input portion 606 including a keyboard, a mouse, etc.; an output portion 607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 608 including a hard disk, etc.; and a communication portion 609 including a network interface card such as a LAN card, a modem, etc. The communication portion 609 performs communication processing via a network such as the Internet. A driver 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the driver 610 as needed so that a computer program read from it can be installed into the storage portion 608 as needed.
[0333] The present invention also provides a computer-readable storage medium on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the above-mentioned perception optimization method, mapping optimization method, and execution optimization method for agent expression imitation.
[0334] Among them, the computer-readable storage medium may be included in the device / device / system described in the above embodiments; it may also exist separately without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the above-mentioned perception optimization method, mapping optimization method, and execution optimization method for agent expression imitation according to Embodiments 1-3 of the present invention are implemented.
[0335] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-described ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603.
[0336] An embodiment of the present invention further includes a computer program product, which includes a computer program that, when executed by a processor, implements the above-mentioned perception optimization method, mapping optimization method, and execution optimization method for agent expression imitation.
[0337] Among them, the computer program contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the above-mentioned perception optimization method, mapping optimization method, and execution optimization method for agent expression imitation provided in Embodiments 1-3 of the present invention.
[0338] When the computer program is executed by the processor 601, it executes the above-mentioned functions defined in the system / apparatus of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0339] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and is downloaded and installed through the communication part 609, and / or installed from the removable medium 611. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0340] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, it executes the above-mentioned functions defined in the system of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0341] According to an embodiment of the present invention, program code for executing the computer programs provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0342] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0343] In addition, all actions of obtaining information, signals, or data in the present invention are carried out on the premise of complying with the corresponding data protection laws, regulations, and policies of the country where it is located, and obtaining authorization from the owner of the corresponding device.
[0344] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present invention can be combined or / and combined in various ways, even if such combinations or combinations are not explicitly recited in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features recited in the various embodiments and / or claims of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0345] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present invention is defined by the appended claims and their equivalents. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. A perceptual optimization method for imitating facial expressions of an intelligent agent, characterized in that: include: Extracting expression motion information of the target expression image; Generate a dense face reconstruction task through the expression motion information and the preset joint optimization information; performing joint optimization on the dense face reconstruction task to generate three-dimensional face reconstruction information, wherein the three-dimensional face reconstruction information is used to generate a target three-dimensional face; and Dense expression motion information of a standardized average face is extracted according to the target three-dimensional human face to complete the perceptual optimization.
2. The method according to claim 1, characterized in that The step of extracting the facial expression motion information of the target facial expression image includes: Performing expression feature recognition on the target expression image to extract the expression motion information.
3. The method according to claim 1, characterized in that In the generating of the dense face reconstruction task by using the expression motion information and the preset joint optimization information, it includes: Acquire task reconstruction information corresponding to the dense face reconstruction task according to the expression parameters of the expression motion information; The dense face reconstruction task is constructed according to the task reconstruction information and preset joint optimization information.
4. The method according to claim 1, characterized in that: In performing joint optimization on the dense face reconstruction task to generate three-dimensional face reconstruction information, the method includes: Generating an initial three-dimensional face corresponding to the dense face reconstruction task; A joint optimization process is performed on the dense face reconstruction task according to the initial three-dimensional face to generate three-dimensional face reconstruction information.
5. The method according to claim 4, characterized in that In the generating the initial three-dimensional face corresponding to the dense face reconstruction task, the method includes: Detecting two-dimensional key point information in the target expression image by using a preset key point detection model; The initial posture parameters of the standardized average face are obtained according to the two-dimensional key point information.
6. The method according to claim 5, characterized in that In the step of generating an initial three-dimensional face corresponding to the dense face reconstruction task, the step further includes: Initializing the optimized variable parameters of the dense face reconstruction task based on the initial posture parameters; Based on the initialization result of the optimization variable parameters, the initial three-dimensional human face is generated by using the standardized average face and default lighting conditions.
7. The method according to claim 4, characterized in that In performing joint optimization processing on the dense face reconstruction task according to the initial three-dimensional face to generate three-dimensional face reconstruction information, the method includes: The projection error, photometric consistency error and regularization constraint in the dense face reconstruction task are minimized and iteratively optimized to generate converged three-dimensional face reconstruction information.
8. The method according to claim 1, characterized in that The step of extracting dense expression motion information of a standardized average face according to the target three-dimensional human face to complete the perceptual optimization includes: The shape parameters of the target three-dimensional human face are optimized to complete the extraction of the dense expression motion information.
9. A perception optimization device for imitating facial expressions of an intelligent agent, characterized in that: include: A motion extraction module, used to extract the expression motion information of the target expression image; A task generation module, used for generating a dense face reconstruction task through the expression motion information and preset joint optimization information; a joint optimization module, configured to perform joint optimization on the dense face reconstruction task to generate three-dimensional face reconstruction information, wherein the three-dimensional face reconstruction information is used to generate a target three-dimensional face; and The information extraction module is used to extract dense expression motion information of the standardized average face according to the target three-dimensional human face to complete the perception optimization.
10. An electronic device comprising: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to execute the method according to any one of claims 1 to 8.
11. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the method according to any one of claims 1 to 8.
12. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.