A robot expression and motion cooperative control method and device, and a robot
By generating a unified vector and combining it with attention and emotion vectors for collaborative control of robot facial expressions and movements, the problem of inconsistent cross-modal styles in robot facial expressions and movements is solved, resulting in a more natural expressive effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING KEYI TECH CO LTD
- Filing Date
- 2026-01-29
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies, the cross-modal style of robot facial expressions and movements is inconsistent, which affects the expressive effect.
By generating a unified vector and combining it with attention vectors and emotion vectors, we can achieve coordinated control of facial expressions and movements, ensuring the natural consistency between them.
It improves cross-modal style consistency between robot facial expressions and movements, enhancing the expressive effect.
Smart Images

Figure CN122113993A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics, and in particular to a method, apparatus, and robot for coordinated control of robot facial expressions and movements. Background Technology
[0002] With the continuous advancement of technology, robots have been widely applied in various scenarios. Robots can express themselves through two modalities: on-screen facial expressions and physical movements, thereby enhancing their anthropomorphism and improving user experience.
[0003] In related technologies, the robot's facial expressions and movements are controlled independently. For example, facial expressions are controlled based on voice rhythm, stickers, or animation template changes; movements are controlled based on task planning and motion status.
[0004] However, in many scenarios, facial expressions and actions are related. The above approach can easily lead to cross-modal style inconsistencies between facial expressions and actions, affecting the robot's expressive ability. For example, an excited expression may be accompanied by slow movements, or vigorous movements may be accompanied by a cold gaze. Summary of the Invention
[0005] This application provides a method, apparatus, and robot for coordinated control of robot facial expressions and movements, in order to improve cross-modal style consistency between facial expressions and movements.
[0006] In a first aspect, embodiments of this application provide a first method for coordinated control of robot facial expressions and movements, the method comprising: Obtain an attention vector and an emotion vector; wherein, the attention vector includes a first initial component corresponding to each of multiple attention information for the target object, and the emotion vector includes a second initial component corresponding to each of multiple emotion information for the target object; A unified vector is generated based on the attention vector and the emotion vector; wherein the unified vector includes a first target component corresponding to each of the multiple attention information and a second target component corresponding to each of the multiple emotion information. The robot's facial expression and motion are controlled based on the unified vector.
[0007] In some alternative implementations, facial expression control is performed in the following ways: For any expression control point, based on the preset vector corresponding to the expression control point and combined with the first component, the deformation offset of the expression control point is obtained; wherein, the first component is one or more of a plurality of second target components; Based on the second component, the translation offset is obtained; wherein the second component is one or more of a plurality of first target components; For any expression control point, a target offset is obtained based on the deformation offset and the translation offset of the expression control point, and the expression control point is adjusted based on the target offset.
[0008] In some alternative implementations, when a blink is triggered, the expression control point is adjusted based on the target offset, including: Based on the target offset and combined with the blinking model, the expression control points are dynamically adjusted. The blinking model is generated based on the target blink interval, the target blink period, and the target blink direction; the target blink interval and the target blink period are both determined based on a third component, which is one or more of the plurality of second target components.
[0009] In some optional implementations, before adjusting the facial expression control points based on the target offset, the method further includes: The target object is determined not to meet the out-of-focus condition; wherein the out-of-focus condition includes the target object being lost, and / or the target object undergoing a sudden change.
[0010] Some optional implementations also include: when the target object meets the out-of-focus condition, controlling the facial expression in the following manner: The target movement speed is obtained based on the current position of the target expression control point; wherein, the target expression control point is the center point of the pupil or the center point of the eye socket; Based on the target movement speed, the control points for each facial expression are dynamically adjusted.
[0011] In some optional implementations, adjusting the facial expression control points based on the target offset includes: Based on the target offset, the facial expression control points are adjusted in conjunction with the target compensation amount; The target compensation amount is determined based on some or all of the first compensation amount, the second compensation amount, and the third compensation amount; the first compensation amount is used to enhance the visual tension of the mechanical action; the second compensation amount is used to simulate changes in spatial depth; and the third compensation amount is used to compensate for the continuation of facial expression at the end of the action.
[0012] In some alternative implementations, the first compensation amount is obtained in the following manner: The first compensation amount is obtained by adjusting the direction mapping vector based on the first coefficient. Wherein, the direction mapping vector is obtained based on the target offset; the first coefficient is positively correlated with the fourth component, the fifth component and the target angle variable respectively; the fourth component is one or more of the plurality of second target components; the fifth component is one or more of the plurality of first target components; the target angle variable is the difference between the robot's current target angle and the initial target angle.
[0013] In some alternative implementations, the second compensation amount is obtained in the following manner: Obtain the depth factor corresponding to the robot's current target angle; Based on the depth factor and the second coefficient, a target scaling factor is obtained; wherein the second coefficient is positively correlated with the fourth component and the fifth component respectively; the fourth component is one or more of the plurality of second target components; and the fifth component is one or more of the plurality of first target components. The second compensation amount is generated based on the target scaling factor.
[0014] In some alternative implementations, the third compensation amount is obtained in the following manner: Obtain the terminal weight corresponding to the current angular velocity; wherein the terminal weight is negatively correlated with the current angular velocity; The third compensation amount is obtained based on the orientation mapping vector, the target angle variable, the end-effector weight, and the third coefficient; wherein the orientation mapping vector is obtained based on the target offset; the third coefficient is positively correlated with the fourth component and the fifth component respectively; the fourth component is one or more of the plurality of second target components; the fifth component is one or more of the plurality of first target components; and the target angle variable is the difference between the robot's current target angle and the initial target angle.
[0015] In some alternative implementations, motion control is performed in the following ways: The robot's position trajectory is adjusted based on a first loss; and the robot's posture is adjusted based on a second loss. The first loss and the second loss are both determined based on the sixth component and the seventh component; the sixth component is one or more of a plurality of second target components; and the seventh component is one or more of a plurality of first target components.
[0016] In some alternative implementations, the first loss is obtained in the following manner: A first loss component is obtained based on the difference between the robot's current trajectory speed and the target trajectory speed; wherein the target trajectory speed is determined based on a preset trajectory speed, the sixth component, and the seventh component; A second loss component is obtained based on the difference between the robot's current trajectory acceleration and the target trajectory acceleration; wherein the target trajectory acceleration is determined based on a preset trajectory acceleration and the sixth component. The first loss is obtained based on the first loss component and the second loss component.
[0017] In some alternative implementations, the second loss is obtained in the following ways: A third loss component is obtained based on the difference between the robot's current angular velocity and the target angular velocity; wherein the target angular velocity is determined based on a preset angular velocity, the sixth component, and the seventh component; A fourth loss component is obtained based on the difference between the robot's current angular acceleration and the target angular acceleration; wherein the target angular acceleration is determined based on a preset angular acceleration and the sixth component; The second loss is obtained based on the third loss component and the fourth loss component.
[0018] Some optional implementations also include: In response to the target action, the corresponding additional facial expression is determined, and the robot's facial expression is controlled based on the additional facial expression.
[0019] Secondly, embodiments of this application provide a robot facial expression and motion coordinated control device, the device comprising: A vector acquisition module is used to acquire an attention vector and an emotion vector; wherein, the attention vector includes a first initial component corresponding to each of multiple attention information for the target object, and the emotion vector includes a second initial component corresponding to each of multiple emotion information for the target object. A vector generation module is used to generate a unified vector based on the attention vector and the emotion vector; wherein the unified vector includes a first target component corresponding to each of the plurality of attention information and a second target component corresponding to each of the plurality of emotion information. The control module is used to control the robot's facial expressions and movements based on the unified vector.
[0020] In some alternative implementations, the control module is specifically used to control facial expressions in the following ways: For any expression control point, based on the preset vector corresponding to the expression control point and combined with the first component, the deformation offset of the expression control point is obtained; wherein, the first component is one or more of a plurality of second target components; Based on the second component, the translation offset is obtained; wherein the second component is one or more of a plurality of first target components; For any expression control point, a target offset is obtained based on the deformation offset and the translation offset of the expression control point, and the expression control point is adjusted based on the target offset.
[0021] In some alternative implementations, when a blink is triggered, the control module is specifically used to: Based on the target offset and combined with the blinking model, the expression control points are dynamically adjusted. The blinking model is generated based on the target blink interval, the target blink period, and the target blink direction; the target blink interval and the target blink period are both determined based on a third component, which is one or more of the plurality of second target components.
[0022] In some optional implementations, before adjusting the expression control points based on the target offset, the control module is further configured to: The target object is determined not to meet the out-of-focus condition; wherein the out-of-focus condition includes the target object being lost, and / or the target object undergoing a sudden change.
[0023] In some optional implementations, the control module is further configured to: control facial expressions when the target object meets the out-of-focus condition, by means of: The target movement speed is obtained based on the current position of the target expression control point; wherein, the target expression control point is the center point of the pupil or the center point of the eye socket; Based on the target movement speed, the control points for each facial expression are dynamically adjusted.
[0024] In some optional implementations, the control module is specifically used for: Based on the target offset, the facial expression control points are adjusted in conjunction with the target compensation amount; The target compensation amount is determined based on some or all of the first compensation amount, the second compensation amount, and the third compensation amount; the first compensation amount is used to enhance the visual tension of the mechanical action; the second compensation amount is used to simulate changes in spatial depth; and the third compensation amount is used to compensate for the continuation of facial expression at the end of the action.
[0025] In some optional implementations, the control module is specifically configured to obtain the first compensation amount in the following manner: The first compensation amount is obtained by adjusting the direction mapping vector based on the first coefficient. Wherein, the direction mapping vector is obtained based on the target offset; the first coefficient is positively correlated with the fourth component, the fifth component and the target angle variable respectively; the fourth component is one or more of the plurality of second target components; the fifth component is one or more of the plurality of first target components; the target angle variable is the difference between the robot's current target angle and the initial target angle.
[0026] In some optional implementations, the control module is specifically configured to obtain the second compensation amount in the following manner: Obtain the depth factor corresponding to the robot's current target angle; Based on the depth factor and the second coefficient, a target scaling factor is obtained; wherein the second coefficient is positively correlated with the fourth component and the fifth component respectively; the fourth component is one or more of the plurality of second target components; and the fifth component is one or more of the plurality of first target components. The second compensation amount is generated based on the target scaling factor.
[0027] In some optional implementations, the control module is specifically configured to obtain the third compensation amount in the following manner: Obtain the terminal weight corresponding to the current angular velocity; wherein the terminal weight is negatively correlated with the current angular velocity; The third compensation amount is obtained based on the orientation mapping vector, the target angle variable, the end-effector weight, and the third coefficient; wherein the orientation mapping vector is obtained based on the target offset; the third coefficient is positively correlated with the fourth component and the fifth component respectively; the fourth component is one or more of the plurality of second target components; the fifth component is one or more of the plurality of first target components; and the target angle variable is the difference between the robot's current target angle and the initial target angle.
[0028] In some alternative implementations, the control module is specifically used to control actions in the following ways: The robot's position trajectory is adjusted based on a first loss; and the robot's posture is adjusted based on a second loss. The first loss and the second loss are both determined based on the sixth component and the seventh component; the sixth component is one or more of a plurality of second target components; and the seventh component is one or more of a plurality of first target components.
[0029] In some alternative implementations, the control module is specifically configured to obtain the first loss in the following manner: A first loss component is obtained based on the difference between the robot's current trajectory speed and the target trajectory speed; wherein the target trajectory speed is determined based on a preset trajectory speed, the sixth component, and the seventh component; A second loss component is obtained based on the difference between the robot's current trajectory acceleration and the target trajectory acceleration; wherein the target trajectory acceleration is determined based on a preset trajectory acceleration and the sixth component. The first loss is obtained based on the first loss component and the second loss component.
[0030] In some alternative implementations, the control module is specifically configured to obtain the second loss in the following manner: A third loss component is obtained based on the difference between the robot's current angular velocity and the target angular velocity; wherein the target angular velocity is determined based on a preset angular velocity, the sixth component, and the seventh component; A fourth loss component is obtained based on the difference between the robot's current angular acceleration and the target angular acceleration; wherein the target angular acceleration is determined based on a preset angular acceleration and the sixth component; The second loss is obtained based on the third loss component and the fourth loss component.
[0031] In some optional implementations, the control module is also used for: In response to the target action, the corresponding additional facial expression is determined, and the robot's facial expression is controlled based on the additional facial expression.
[0032] Thirdly, embodiments of this application provide a robot, including at least one processor and at least one memory, wherein the memory stores a computer program, and when the program is executed by the processor, the processor performs the robot expression and movement collaborative control method described in any of the first aspects above.
[0033] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program executable by a processor, which, when run on the processor, causes the processor to execute the robot facial expression and motion coordinated control method described in any of the first aspects above.
[0034] In this embodiment, by combining attention vectors and emotion vectors, a unified vector for co-driven operation is generated. Based on this unified vector, facial expression control and motion control of the robot are performed, achieving natural consistency between facial expressions and actions across modalities. In other words, attention vectors not only affect the robot's actions but also its facial expressions; similarly, emotion vectors not only affect the robot's facial expressions but also its actions, enabling a strong correlation between facial expressions and actions and enhancing the robot's expressive effect. Attached Figure Description
[0035] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0036] Figure 1 A flowchart illustrating a robot facial expression and motion coordinated control method provided in an embodiment of this application; Figure 2 A flowchart illustrating the first facial expression control process provided in this application embodiment; Figure 3 A flowchart illustrating the second facial expression control process provided in this application embodiment; Figure 4 A flowchart illustrating the motion control process provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of the robot facial expression and motion coordination control device provided in the embodiments of this application; Figure 6 This is a schematic diagram of the robot provided in an embodiment of this application. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0038] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0039] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the term "connection" should be interpreted broadly. For example, it can refer to a direct connection, an indirect connection through an intermediate medium, or a connection within two devices. Those skilled in the art can understand the specific meaning of the above term in this application based on the specific circumstances.
[0040] With the continuous advancement of technology, robots have been widely applied in various scenarios. Robots can express themselves through two modalities: on-screen facial expressions and physical movements, thereby enhancing their anthropomorphism and improving user experience.
[0041] In related technologies, the robot's facial expressions and movements are controlled independently. For example, facial expressions are controlled based on voice rhythm, stickers, or animation template changes; movements are controlled based on task planning and motion status.
[0042] However, in many scenarios, facial expressions and actions are related. The above approach can easily lead to cross-modal style inconsistencies between facial expressions and actions, affecting the robot's expressive ability. For example, an excited expression may be accompanied by slow movements, or vigorous movements may be accompanied by a cold gaze.
[0043] In view of this, embodiments of this application propose a method, device, and robot for collaborative control of robot facial expressions and movements, in order to improve cross-modal style consistency between facial expressions and movements.
[0044] The technical solution of this application and how it solves the above-mentioned technical problems will be described in detail below with reference to the accompanying drawings and specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0045] Figure 1 This is a flowchart illustrating a robot facial expression and motion coordinated control method provided in an embodiment of this application, as shown below. Figure 1 As shown, it includes the following steps: Step S101: Obtain attention vector and emotion vector; wherein, the attention vector includes a first initial component corresponding to each of the multiple attention information for the target object, and the emotion vector includes a second initial component corresponding to each of the multiple emotion information for the target object.
[0046] This application's embodiments sample attention vectors and emotion vectors in real time, obtaining attention vectors and emotion vectors at the same time. In other words, a unified vector is generated based on the attention vectors and emotion vectors at the same time.
[0047] In practice, the aforementioned attention vector includes a first initial component corresponding to each of the multiple attention information for the target object, describing the attention situation for the target object from different perspectives.
[0048] The target object here refers to the most important object for the robot, such as the person who is having a conversation, the person waving, or the person making a sound.
[0049] For example, the attention vector a(t) = [g(t), g (t),p g (t)]; Where g(t) represents the target object at time t; g (t) represents the attention intensity for g(t), which is a continuous value between 0 and 1. g A larger value for (t) indicates a greater focus on the target object; p g (t) represents the position of the target object in the robot coordinate system, for example, p g (t)=[x(t),y(t),z(t) x(t) is the X-axis coordinate of the target object in the robot coordinate system, representing the horizontal distance of the target object in front of the robot; y(t) is the Y-axis coordinate of the target object in the robot coordinate system, representing the left and right distance of the target object, with left being negative and right being positive; z(t) is the Z-axis coordinate of the target object in the robot coordinate system, representing the vertical height or depth direction of the target object.
[0050] As can be seen, based on the attention vector, we can obtain not only the attention intensity of the target object (object salience), but also the position of the target object relative to the robot, as well as the distance trend between the robot and the target object (such as determining whether to get closer to the target object by combining the position at different times).
[0051] In practice, the aforementioned emotion vector includes a second initial component corresponding to each of the multiple emotional information of the target object, describing the target object's emotional state from different perspectives.
[0052] This application does not specifically limit the multiple emotional information. For example, it may include some or all of the following: pleasure level (representing the degree of enjoyment, the higher the pleasure level, the more enjoyable), arousal level (representing the degree of excitement / tension, the higher the arousal level, the more excited / tensionful), and sense of control (representing the degree of confidence / control, the higher the sense of control, the more confident / controlled).
[0053] For example, the emotion vector e t =[V(t),A(t),D(t) V, A, D ∈ [-1, 1]; where V is the level of pleasure, A is the level of arousal, and D is the level of dominance.
[0054] The attention vector and emotion vector described above are illustrative examples. In practice, the attention vector and emotion vector may include other components.
[0055] Step S102: Generate a unified vector based on the attention vector and the emotion vector; wherein the unified vector includes a first target component corresponding to each of the multiple attention information and a second target component corresponding to each of the multiple emotion information.
[0056] In implementation, a unified vector is generated by obtaining the attention vector and emotion vector at the same time point. For example, a unified vector U1 is generated based on the attention vector a1 and emotion vector S1 at time 1; a unified vector U2 is generated based on the attention vector a2 and emotion vector S2 at time 2; ... based on the attention vector a1 at time N... N And the emotion vector S N Generate a set of unified vectors U N .
[0057] This application does not specifically limit the method of generating the unified vector. For example, the attention vector and the emotion vector can be directly concatenated to obtain the unified vector; or the attention vector and the emotion vector can be aligned in dimensions and then weighted to obtain the unified vector. Here, dimension alignment means converting them into vectors that can be added together. For example, if the attention vector is a 3-dimensional vector and the emotion vector is a 3-dimensional vector, the attention vector can be expanded into a 6-dimensional vector (the first 3 dimensions are the first initial component, and the last 3 dimensions are 0), and the emotion vector can be aligned into a 6-dimensional vector (the first 3 dimensions are 0, and the last 3 dimensions are the second initial component). In the unified vector obtained in this way, the first 3 dimensions are the first target component corresponding to attention, and the last 3 dimensions are the second target component corresponding to emotion.
[0058] For example, the unified vector u at time t t =σ(Wa a(t)+We e(t)+b); where Wa is the attention weight, a(t) is the attention vector after alignment at time t, We is the emotion weight, e(t) is the emotion vector after alignment, b is the bias term to give the default expression / action a basic posture, and σ is the compression function (tanh / sigmoid) to prevent the value from going out of range and to ensure the curve is smooth.
[0059] Step S103: Perform facial expression control and motion control on the robot based on the unified vector.
[0060] In practice, after obtaining a unified vector, the robot's facial expressions and movements are controlled collaboratively. For example, when controlling facial expressions, the second target vector is mainly referenced, but the first target vector is also considered; when controlling movements, the first target vector is mainly referenced, but the second target vector is also considered.
[0061] The above scheme combines attention vectors and emotion vectors to generate a unified vector for co-driven operation. Based on this unified vector, the robot's facial expression and motion control are performed, achieving natural consistency between facial expressions and actions across modalities. In other words, the attention vector affects not only the robot's actions but also its facial expressions; similarly, the emotion vector affects not only the robot's facial expressions but also its actions, enabling a strong correlation between facial expressions and actions and enhancing the robot's expressive capabilities.
[0062] In practice, to achieve better interactive effects, the speed of facial expression control is greater than the speed of motion control, thereby achieving facial expression preheating to enhance the mechanical active motion.
[0063] The following is a specific example to illustrate this: The action is controlled according to the first model, where the first model is... ; where t start Action start time (in seconds or frame number), T duration To determine the duration of the action, the first model normalizes the action process into a "progress bar" of 0 to 1, where t=0 represents the start time of the action, t=0 represents the middle time of the action, and t=0 represents the end time of the action.
[0064] Phenotypic control is performed according to the second model, where the second model is... Where γ is a distortion parameter between 0 and 1. Since F2(t) > F1(t) at the same time, the speed of facial expression control is greater than the speed of motion control.
[0065] This application provides a flowchart illustrating a first type of facial expression control process, as shown in the embodiments below. Figure 2 As shown, it includes the following steps: Step S201: Obtain attention vector and emotion vector; wherein, the attention vector includes a first initial component corresponding to each of multiple attention information for the target object, and the emotion vector includes a second initial component corresponding to each of multiple emotion information for the target object.
[0066] Step S202: Generate a unified vector based on the attention vector and the emotion vector; wherein the unified vector includes a first target component corresponding to each of the multiple attention information and a second target component corresponding to each of the multiple emotion information.
[0067] The specific implementation of steps S201 to S202 can be found in other embodiments, and will not be repeated here.
[0068] Step S203: For any expression control point, based on the preset vector corresponding to the expression control point and combined with the first component, obtain the deformation offset of the expression control point; wherein, the first component is one or more of a plurality of second target components.
[0069] During implementation, all facial expression control points of the robot are obtained, which may include the set of eye expression control points, the set of mouth expression control points, etc.
[0070] Taking eye control points as an example, the set of eye expression control points P at time t eye (t)= Where N represents the total number of eye expression control points, and R... 2 It refers to a two-dimensional real vector space, representing each pi = [X] i , Y i [] is a point in the screen coordinate system. In this application embodiment, the total number of eye expression control points is not specifically limited. For example, setting it to 16~32 can obtain a smooth and controllable eye curve.
[0071] The set of control points for eye expressions can be further divided into the left eye orbital point set P. orbL Right eye orbital point set P orbR Left eye pupil dot set P pupL Right eye pupil point set P pupR .
[0072] In practice, Catmull-Rom splines (a commonly used interpolation spline) can be used to generate curves for expression control points, so that wherever the control points go, the curve will smoothly follow.
[0073] Since emotions primarily affect the shape of facial expressions, this application embodiment obtains the corresponding deformation offset for each facial expression control point by combining a preset vector corresponding to the facial expression control point with one or more second target components.
[0074] For example, the deformation offset Δp of expression control point i i(t) =f p (u t ,wi)];where, f p u is the deformation calculation function. t Let wi be the unified vector mentioned above, and wi be the preset vector corresponding to the expression control point i. This preset vector can contain one or more vectors. This formula represents applying a small displacement to each expression control point using the unified vector.
[0075] Since different expression control points have different functions, such as the upper eyelid point mainly affecting vertical opening and closing, and the corner of the eye point mainly affecting horizontal curvature, a corresponding preset vector is set for each expression control point.
[0076] Furthermore, different emotions have different effects. For example, arousal level has a significant impact on vertical eye movement, while pleasure level has a significant impact on horizontal eye movement. Therefore, multiple preset vectors can be set for each expression control point. This allows for eye movements such as widening the eyes, flattening the eyes, or raising the eyebrows to be implemented for different emotions.
[0077] The following two specific examples illustrate this.
[0078] The first method: In this example, the first component includes the second target components corresponding to arousal and pleasure, respectively, and the deformation offset Δp of the expression control point i. i(t) =wi1 u ar ni o +wi2 u val ni a Where wi1 is the first preset vector of expression control point i, u ar For wakefulness, ni o Let wi2 be the unit vector in the vertical opening and closing direction, and wi2 be the second preset vector of the expression control point i. val For pleasure, ni a It is a unit vector in the horizontal arc direction.
[0079] The second method: In this example, the first component includes the second target components corresponding to arousal, pleasure, and sense of control, and the deformation offset Δp of the facial expression control point i. i(t) =wi o u ar ni o +wi a u v ni a +wi3 u dom ni o Where wi1 is the first preset vector of expression control point i, u ar For wakefulness, ni o Let wi2 be the unit vector in the vertical opening and closing direction, and wi2 be the second preset vector of the expression control point i. val For pleasure, nia Let wi3 be the unit vector in the horizontal arc direction, and wi3 be the third preset vector of the expression control point i. dom For a sense of control, ni o It is a unit vector in the vertical opening and closing direction.
[0080] The above example is just for illustration. In practice, other first components can be selected and other methods can be used to obtain the deformation offset.
[0081] Step S204: Obtain the translation offset based on the second component; wherein the second component is one or more of a plurality of first target components.
[0082] In practice, attention mainly affects the direction of facial expressions (i.e., overall translation, where the facial expression is facing). Based on this, the embodiments of this application obtain the translation offset based on one or more first target components.
[0083] For example, the translation offset Δq (t) =f q (u t ); where f p For translation calculation function, u t Let be the unified vector mentioned above. This formula represents moving all facial expression control points as a whole using the unified vector.
[0084] The following is a specific example to illustrate this.
[0085] The second component is the position information of the target object in the robot coordinate system, which may include the X-axis coordinate x(t) of the target object in the robot coordinate system and the Y-axis coordinate y(t) of the target object in the robot coordinate system.
[0086] By mapping x(t) and y(t) to coordinates, the position of the target object is mapped to the screen offset, thus obtaining the translation offset.
[0087] The formula is expressed as Δq(t) = k q ugaze(t); where ugaze(t) is the position information of the target object in the robot coordinate system, k q This is the conversion factor.
[0088] This allows the facial expression to follow the movement of the target object; for example, if the target object is on the right, the facial expression will move to the right.
[0089] Step S205: For any expression control point, based on the deformation offset and the translation offset of the expression control point, obtain the target offset of the expression control point, and adjust the expression control point based on the target offset.
[0090] In practice, after obtaining the deformation offset of each expression control point and the overall translation offset, the target offset of each expression control point can be obtained based on these two offsets, and then each expression control point can be adjusted.
[0091] For example, the sum of the deformation offset and translation offset of an expression control point can be used as the target offset of that expression control point; then, the target offset can be added to the current target position of that expression control point to form a new target position.
[0092] The formula is expressed as: the target position A of expression control point i at time t. i,t =A i,t-1 +ΔA i,t Among them, A i,t-1 Let ΔA be the target position of expression control point i at time t-1. i,t Let A be the target offset of expression control point i at time t. i,0 Let i be the initial position of the expression control point i. In other words, each expression control point is at its initial position at the initial moment.
[0093] In practice, the target offset can be divided into a horizontal component ΔX and a vertical component, such as the target offset ΔA of expression control point i at time t. i,t Including the horizontal component ΔX i,t and the vertical component ΔY i,t .
[0094] The above scheme, for each expression control point, based on the preset vector corresponding to the expression control point and combined with one or more second target components (third components), accurately obtains the deformation offset of each expression control point under the influence of emotion; based on one or more first target components (fourth components), accurately obtains the overall translation offset under the influence of attention; and then combines the deformation offset and translation offset to accurately obtain the target offset.
[0095] It's understandable that, besides using target offsets to control the position of facial expression control points, one can also distort or scale these control points. For example, based on the target control vector [ΔX...] i,t ΔY i,t sx t sy t , t Adjust the facial expression control point i; where ΔX i,t Let ΔY be the horizontal component of the target offset of expression control point i at time t. i,t Let sx be the vertical component of the target offset of expression control point i at time t. t Let sy be the horizontal scaling factor at time t. t Let be the vertical scaling factor at time t. t The deflection angle at time t is the relative deflection angle between the left and right sides, used to achieve additional facial expressions such as tilting the head.
[0096] In practice, the aforementioned horizontal and vertical scaling factors can be determined based on one or more second target components, thereby achieving scaling of parts such as the pupils under specific emotions. The aforementioned deflection angle can be determined based on whether additional facial expressions are triggered, thereby achieving specific additional facial expressions.
[0097] In some alternative implementations, when the additional facial expression of blinking is triggered, step S205 above can be implemented in, but is not limited to, the following ways: Based on the target offset and combined with the blinking model, the expression control points are dynamically adjusted. The blinking model is generated based on the target blink interval, the target blink period, and the target blink direction; the target blink interval and the target blink period are both determined based on a third component, which is one or more of the plurality of second target components.
[0098] In practice, when the additional expression of blinking is triggered, the expression control points are dynamically adjusted based on the target offset and the blinking model. In other words, the blinking model does not change the target offset, but only adds the dynamics of time, thereby realizing the time curve of 0→1→0, where 0 represents the eyes are open and 1 represents the eyes are closed.
[0099] For example, the position A of expression control point i B i,t =A i,t +β i h b (t;u(t)); where, A i,t Let β be the target position of expression control point i at time t. i Used to indicate the direction of the target's blink, showing which direction and how much the eye contracts when blinking at each point, h b It is a periodic function used to indicate the target blink interval and the target blink cycle.
[0100] The target blink interval and target blink period mentioned above are both determined based on one or more second target components. For example: target blink interval T b (u)=T0 / (1+α1 u ar ); Target blink cycle ξ b (u)=ξ0+α2 (1-u ar ); where T0 is the preset blink interval, α1 is the fourth coefficient, and u arThe above-mentioned arousal level; ξ0 is the preset blink cycle, and α2 is the fifth coefficient.
[0101] In some optional implementations, the following steps are performed before step S205: The target object is determined not to meet the out-of-focus condition; wherein the out-of-focus condition includes the target object being lost, and / or the target object undergoing a sudden change.
[0102] When the target object is lost, a complete attention vector cannot be obtained, therefore, facial expression control cannot be performed in the above manner; When the target object undergoes a sudden change (such as the target object suddenly moving away, the target object switching, or a problem with attention), if the facial expression is still made to follow the target object, it will create a sense of camera noise.
[0103] Therefore, when the image is out of focus, the above methods will no longer be used to control facial expressions.
[0104] The above solution improves the expressive effect by coordinating facial expressions and movements when the defocusing condition is not met.
[0105] For cases where the target object meets the out-of-focus condition, this application provides a flowchart of a second expression control process, such as... Figure 3 As shown, it includes the following steps: Step S301: Based on the current position of the target expression control point, obtain the target movement speed; wherein, the target expression control point is the center point of the pupil or the center point of the eye socket.
[0106] In practice, when the target object meets the out-of-focus condition, it is no longer possible or suitable to have the expression follow the target object. Based on this, this application embodiment combines the center point of the pupil or the center point of the eye socket to allow the expression to continue gliding in the direction between them, in order to simulate the inertia of human vision.
[0107] Since the current position of the target expression control point (that is, the coordinates of the target expression control point on the screen) is a vector, it reflects the motion before it went out of focus; based on this, the target's motion speed is obtained according to the current position of the target expression control point.
[0108] For example, target speed of motion ; where p lost The current position of the target facial expression control point, |p lost | is the modulus of the current position of the target facial expression control point.
[0109] Step S302: Based on the target motion speed, dynamically adjust each facial expression control point.
[0110] During implementation, after obtaining the target motion speed, inertial motion is performed on each facial expression control point based on the target motion speed.
[0111] For example, regarding the position A of expression control point i at time j... i, j =A i,lost +k inertia T A i,lost Let k be the position of the expression control point i at the moment of defocus. inertia k is a preset coefficient; when different target facial expression control points are used... inertia They can be different, where T is the time interval between time j and the time of defocusing. The target speed is denoted as .
[0112] In some optional implementations, step S205 above can be implemented in, but is not limited to, the following ways: Based on the target offset, the facial expression control points are adjusted in conjunction with the target compensation amount; The target compensation amount is determined based on some or all of the first compensation amount, the second compensation amount, and the third compensation amount; the first compensation amount is used to enhance the visual tension of the mechanical action; the second compensation amount is used to simulate changes in spatial depth; and the third compensation amount is used to compensate for the continuation of facial expression at the end of the action.
[0113] In practice, style consistency is achieved to a certain extent through the joint generation of homologous vectors and the coordinated control of facial expressions and movements. However, problems such as insufficient semantic perceptibility of movements, insufficient expression of emotional tension, and lack of visual focus may occur in weak movements, short movements, and complex rhythmic movements.
[0114] Based on this, the embodiments of this application compensate for facial expression control by combining the target compensation amount.
[0115] It is understandable that the target compensation amount is not an independent facial expression control parameter, but a parameter superimposed on the basic facial expression control parameters (the target control vector obtained based on the target offset), used to adjust the screen rendering so as to better match the robot's movements.
[0116] In some optional implementations, the rendering control vector c final =c base +Δc; where c base The target control vector is given above (obtained based on the target offset), and Δc is the target compensation amount.
[0117] During implementation, compensation can be made by selecting part or all of the first, second, and third compensation amounts, depending on actual needs.
[0118] For example, the first compensation amount is used to enhance the visual tension of the mechanical action, that is, the first compensation amount is an amplitude amplification compensation component; The second compensation amount is used to simulate changes in spatial depth. In other words, the second compensation amount is a depth scaling simulation component. The scaling of facial expressions is achieved through the second compensation amount to simulate "depth changes" (depth illusion) and enhance the spatial sense of mechanical movements. The third compensation amount is used to compensate for the continuation of facial expression at the end of the action. In other words, the third compensation amount is the end-effector compensation component. Through the third compensation amount, the facial expression can continue to move at the end of the action, simulating the residual momentum of a person's facial expression.
[0119] The above scheme, by combining some or all of the first, second, and third compensation amounts, can enhance the visual tension of mechanical movements, strengthen spatial depth changes, or compensate for the continuation of facial expressions at the end of movements. Facial expressions and movements reinforce each other under the same semantics and rhythm, resulting in more natural, powerful, and emotionally expressive robot expressions.
[0120] In some alternative implementations, the first compensation amount can be obtained in, but is not limited to, the following ways: The first compensation amount is obtained by adjusting the direction mapping vector based on the first coefficient. Wherein, the direction mapping vector is obtained based on the target offset; the first coefficient is positively correlated with the fourth component, the fifth component and the target angle variable respectively; the fourth component is one or more of the plurality of second target components; the fifth component is one or more of the plurality of first target components; the target angle variable is the difference between the robot's current target angle and the initial target angle.
[0121] In implementation, the direction mapping vector is obtained based on the target offset. For example, the target control vector is [ΔX]. i,t ΔY i,t sx t sy t , t For example (the meaning of these parameters can be found in the above embodiments, and will not be repeated here), the final target control vector used for facial expression control contains five pieces of information, and the direction mapping vector also needs to contain these five pieces of information.
[0122] For example, the direction mapping vector [ΔX] i,t ΔY i,t sx0, sy0 0]; where ΔX i,t Let ΔY be the horizontal component of the target offset of expression control point i at time t. i,tLet be the vertical component of the target offset of expression control point i at time t, where sx0 is the preset horizontal scaling factor and sy0 is the preset vertical scaling factor. 0 represents the preset deflection angle.
[0123] Furthermore, based on the fourth component, the fifth component, and the target angle variable, a first coefficient is obtained. The first coefficient is then used to adjust (e.g., multiply) the terms in the direction mapping vector to obtain the first compensation amount.
[0124] For example, the first compensation amount Δc amp =M k amp Where M is the direction mapping vector, k amp It is the first coefficient.
[0125] During implementation, a first compensation amount is generated when a specific action is triggered (such as when the target angle variable is not zero). The aforementioned target angle can be a pitch angle, thus generating the first compensation amount when tilting the head up or down in the longitudinal direction.
[0126] The aforementioned first coefficient is positively correlated with the fourth component, the fifth component, the robot's action stage, and the target angle variable, respectively. This application does not specifically limit the fourth and fifth components; for example, the fourth component may include the aforementioned arousal level, and the fifth component may include the aforementioned attention intensity. When there is a more focused target object, the target object is more excited, and the angle rotation is greater, the compensation force of the first compensation amount is greater.
[0127] In some alternative implementations, the first coefficient is also affected by the robot's motion phase, such as the first coefficient during the main motion phase being greater than the first coefficient during the preparatory / end phase.
[0128] Taking the target angle as the pitch angle as an example, when looking down, attention is drawn to the upper part of the head, and the range of eye movement is greater than in a normal posture. This is accompanied by a certain degree of scaling, which compensates for the lowered facial expression and the feeling of pressure when looking down and looking up.
[0129] In some alternative implementations, the second compensation amount can be obtained in, but is not limited to, the following ways: Obtain the depth factor corresponding to the robot's current target angle; Based on the depth factor and the second coefficient, a target scaling factor is obtained; wherein the second coefficient is positively correlated with the fourth component and the fifth component respectively; the fourth component is one or more of the plurality of second target components; and the fifth component is one or more of the plurality of first target components. The second compensation amount is generated based on the target scaling factor.
[0130] First, for the robot's current target angle, obtain the corresponding depth factor. The target angle can be the pitch angle, and establish the relationship between the pitch angle and the depth factor. For example, when the robot is looking down, the target angle is negative, and the smaller the target angle when looking down (the larger the absolute value), the smaller the depth factor; the normal target angle is 0, and the depth factor is also 0; when the robot is looking up, the target angle is positive, and the larger the target angle when looking up, the larger the depth factor.
[0131] For example, depth factor Where θ(t) is the target angle at time t, σ is the preset angle, and tanh is the hyperbolic tangent function to ensure smooth and saturated output (avoiding extreme values). σ can be set according to actual needs (e.g., set to 10°). σ is used to control sensitivity; the smaller σ is, the more sensitive it is to small angles, and the larger σ is, the more sensitive it is to large angles.
[0132] According to the above formula, when the head is raised (θ(t)>0), f depth >0, and the larger θ(t) is, the better f depth The closer to 1; when the head is lowered (θ(t)<0), f depth <0, and the smaller θ(t) is, the better f depth The closer it is to -1.
[0133] Next, based on the depth factor and the second coefficient, the target scaling factor is obtained.
[0134] The second coefficient mentioned above is positively correlated with both the fourth and fifth components. For example, the fourth component includes the aforementioned arousal level, and the fifth component includes the aforementioned attention intensity. When there is a more focused target object, and the target object is more excited, the compensation strength of the second compensation is greater.
[0135] For example, the product of the depth factor and the second coefficient can be used as the target scaling factor.
[0136] The formula is expressed as: target scaling factor ; where f depth k is the depth factor. depth This is the second coefficient.
[0137] Finally, a second compensation amount is generated based on the target scaling factor.
[0138] As mentioned above, the final target control vector used for facial expression control contains five pieces of information, and the second compensation vector also needs to contain these five pieces of information accordingly. In implementation, the target scaling factor can be used as the horizontal and vertical scaling factors in the second compensation vector, and the other information can be padded with 0 to reduce interference with this information.
[0139] With the target control vector as [ΔX] i,t ΔY i,tsx t sy t , t For example (the meanings of these parameters can be found in the above embodiments, and will not be repeated here), the corresponding second compensation amount is [0, 0, s]. (x,y) s (x,y) 、0],s (x,y) The scaling factor for the above target is .
[0140] In this way, the second compensation amount is dominated by the mechanical angle, but the intensity of attention and emotion modulation ensures that the screen expression does not change independently, but rather enhances the spatial force in conjunction with the mechanical action.
[0141] In some alternative implementations, the third compensation amount can be obtained in, but is not limited to, the following ways: Obtain the terminal weight corresponding to the current angular velocity; wherein the terminal weight is negatively correlated with the current angular velocity; The third compensation amount is obtained based on the orientation mapping vector, the target angle variable, the end-effector weight, and the third coefficient; wherein the orientation mapping vector is obtained based on the target offset; the third coefficient is positively correlated with the fourth component and the fifth component respectively; the fourth component is one or more of the plurality of second target components; the fifth component is one or more of the plurality of first target components; and the target angle variable is the difference between the robot's current target angle and the initial target angle.
[0142] First, obtain the corresponding end effector weight for the robot's current angular velocity.
[0143] Since the angular velocity decreases as the action is about to end, the end-effector weight is obtained based on the current angular velocity to adjust the compensation level of the third compensation amount.
[0144] For example, end weight Where ω(t) is the angular velocity, v scale This is the speed threshold. The speed threshold can be set according to actual needs (e.g., 5° / s) to define the boundary between fast and slow. scale The smaller the value, the more sensitive it is to low speeds. In this formula, w end It is an exponentially decreasing function, ranging from 0 to 1. The greater the velocity (|ω(t)|) (high speed in the middle of the action), the greater w... end The smaller w is end When the velocity is close to 0, end-effector compensation is almost ineffective, avoiding interference with the active motion; the smaller the velocity (|ω(t)|) (quick stop, slow recovery), the more effective the end-effector compensation becomes. end The larger w is end The end-compensation intensity is high when it approaches 1.
[0145] In other words, the end weight is equivalent to a detection switch and intensity regulator for end compensation. In the middle of the movement, the end weight is very small, and the second compensation amount is almost negligible. When the movement ends, the movement continues, and the end weight is larger, realizing the relay of on-screen expressions and prolonging the expression.
[0146] Furthermore, a third compensation amount is obtained based on the direction mapping vector, the target angle variable, the end weight, and the third coefficient.
[0147] For example, the product of these items can be used as the third compensation amount.
[0148] The formula is expressed as the third compensation amount Δc end =M Δθ w end k end Where M is the direction mapping vector, Δθ is the target angle variable, and w end k represents the end weight. end This is the third coefficient. The direction mapping vector and target angle variable can be referred to the above embodiment, and will not be repeated here.
[0149] The third coefficient mentioned above is positively correlated with both the fourth and fifth components. For example, the fourth component includes the aforementioned arousal level, and the fifth component includes the aforementioned attention intensity. When there is a more focused target object, and the target object is more excited, the compensation force of the third compensation factor is greater and more sustained.
[0150] In some alternative implementations, only a portion of the information in M is adjusted, such as performing only vertical translation and scaling. The direction mapping vector is [ΔX]. i,t ΔY i,t sx0, sy0 Taking 0 as an example, only ΔY is compensated. i,t And sy0.
[0151] Taking the target angle as the pitch angle as an example, when looking up (Δθ>0), the follow-up action is upward movement and pupil dilation; when looking down (Δθ<0), the follow-up action is downward movement and pupil constriction.
[0152] Through the above methods, the third compensation amount is controlled by mechanical displacement and speed (mechanical triggering and direction for facial expression compensation), thereby compensating for the semantic weakening when the machine stops, allowing the on-screen expression to "continue" the action and form a more coherent anthropomorphic expression. For example, when attention is focused on the far left, the robot turns its head to the left, the expression moves first, and then the action catches up. The amplitude compensation of the eyes is affected by the target angle of the action. When approaching the target point, the speed will slow down. End-effector compensation avoids the abruptness of the expression stopping abruptly.
[0153] This application provides a flowchart illustrating a motion control process, as shown in the embodiments below. Figure 4 As shown, it includes the following steps: Step S401: Obtain attention vector and emotion vector; wherein, the attention vector includes a first initial component corresponding to each of the multiple attention information for the target object, and the emotion vector includes a second initial component corresponding to each of the multiple emotion information for the target object.
[0154] Step S402: Generate a unified vector based on the attention vector and the emotion vector; wherein the unified vector includes a first target component corresponding to each of the multiple attention information and a second target component corresponding to each of the multiple emotion information.
[0155] The specific implementation of steps S401 to S402 can be found in other embodiments, and will not be repeated here.
[0156] Step S403: Adjust the robot's position trajectory based on the first loss; and adjust the robot's posture based on the second loss; wherein the first loss and the second loss are both determined based on the sixth component and the seventh component; the sixth component is one or more of a plurality of second target components; and the seventh component is one or more of a plurality of first target components.
[0157] In practice, after obtaining the position of the target object in the robot coordinate system, it is necessary to determine the target trajectory of the robot and the target attitude of the robot (including the target pitch angle and the target yaw angle); then, based on the sixth and seventh components, the first loss and the second loss are obtained; the position trajectory of the robot is adjusted based on the first loss so that the robot moves along the target trajectory with a specific trajectory velocity and trajectory acceleration; and the attitude of the robot is adjusted based on the second loss so that the robot achieves the target attitude with a specific angular velocity and angular acceleration.
[0158] The sixth component mentioned above is one or more second target components, such as arousal level; the seventh component mentioned above is one or more first target components, such as attention intensity.
[0159] In practice, the attitude trajectory can be obtained using a cubic polynomial: The angle (target yaw angle or target pitch angle) θ(t) = a0 + a1t + a2t2 + a3t3; where a0, a1, a2, and a3 are undetermined coefficients, and t is the time t.
[0160] By specifying the position and velocity at the start and end times, a smooth rotation curve with continuous velocity and acceleration can be obtained.
[0161] The boundary constraints of the attitude trajectory are: [θ(0)=θ cur ,θ(T)=θ p ,θ y Where, θ(0) = θ cur For the corresponding start time, θ(T) = θp, and for the corresponding end time, θy = θp. cur θ is the current angle. p Let θ be the target pitch angle. y This represents the target yaw angle. It indicates that the starting point must be from the current angle, and the ending point must be the target attitude to align with the target object.
[0162] The following is an example illustrating how to obtain the robot's target trajectory: Continuous target trajectory of the robot on the ground Among them, C i For each trajectory control point; m is the total number of trajectory control points, the specific number of which can be set according to the actual application (e.g., 5~7); N i These are B-spline basis functions used to connect trajectory control points into a smooth curve.
[0163] The following is an example illustrating how to obtain the target's pitch angle: Target pitch angle Where x(t) is the X-axis coordinate of the target object in the robot coordinate system; y(t) is the Y-axis coordinate of the target object in the robot coordinate system; z(t) is the Z-axis coordinate of the target object in the robot coordinate system; and atan2 is the arctangent function in the four quadrants.
[0164] The target pitch angle reflects whether the target object is high or low in the robot's field of view. For example, when the target object becomes higher, z(t) > 0, and the corresponding -z(t) < 0, θ p When z(t) > 0, the robot adopts a head-up posture; when the target object lowers, z(t) < 0, and the corresponding -z(t) > 0, θ p (t) < 0, the robot adopts a head-down posture.
[0165] The following is an example illustrating how to obtain the target yaw angle: The target yaw angle is the "heading angle", and the target yaw angle θ y (t) = atan2(y(t),x(t)); where x(t) is the X-axis coordinate of the target object in the robot coordinate system; y(t) is the Y-axis coordinate of the target object in the robot coordinate system.
[0166] For example, y(t) corresponding to the front is 0, y(t) corresponding to the right is positive, and y(t) corresponding to the left is negative. When the target is on the right, y(t) > 0, and the robot will turn its head / turn to the right.
[0167] In some alternative implementations, adjusting the robot's position trajectory based on the first loss can be achieved, but is not limited to, in the following ways: First, based on the first loss, trajectory loss, and collision loss, the first target loss is obtained.
[0168] In practice, the first loss reflects the loss caused by emotions and attention, the trajectory loss is the loss due to the actual trajectory deviating from the target trajectory, and the collision loss is the loss due to approaching obstacles. Combining these three losses ensures that the position trajectory not only approaches the target trajectory and moves away from obstacles, but also aligns with the style of emotions and attention.
[0169] For example, the weighted sum of the three can be used as the first target loss.
[0170] The formula is expressed as: J(C) = λ path J path +λ obs J obs +λ em J em ; where λ path J is the weight of the trajectory loss. path For trajectory loss, λ obs J is the weight of the collision loss. obs For collision loss, λ em J is the weight of the first loss. em This is the first loss.
[0171] Then, the robot's position trajectory is adjusted based on the first target loss.
[0172] For example, the robot's position trajectory is adjusted by minimizing the loss of the first objective.
[0173] In practice, the robot's posture can be adjusted directly by minimizing the second loss.
[0174] In some alternative implementations, the first loss can be obtained in, but is not limited to, the following ways: A first loss component is obtained based on the difference between the robot's current trajectory speed and the target trajectory speed; wherein the target trajectory speed is determined based on a preset trajectory speed, the sixth component, and the seventh component; A second loss component is obtained based on the difference between the robot's current trajectory acceleration and the target trajectory acceleration; wherein the target trajectory acceleration is determined based on a preset trajectory acceleration and the sixth component. The first loss is obtained based on the first loss component and the second loss component.
[0175] First, the target trajectory velocity is determined based on the preset trajectory velocity, the sixth component, and the seventh component; the target trajectory acceleration is determined based on the preset trajectory acceleration and the sixth component.
[0176] Taking the sixth component as arousal and the seventh component as attention intensity as an example, the more excited and attentive the target is, the faster it is necessary to approach the target.
[0177] For example, the target trajectory velocity v ref =v0 (1+w1 u A +w2 u ar ); Target trajectory acceleration a ref =a0 (1+w3 u ar ); where v0 is the preset trajectory velocity, u A To note the intensity, u ar The wake-up level is represented by a0, the preset trajectory acceleration is represented by w1, w2, and w3, which are adjustable coefficients that can be set according to actual needs.
[0178] Next, a first loss component is obtained based on the difference between the robot's current trajectory velocity and the target trajectory velocity; and a second loss component is obtained based on the difference between the robot's current trajectory acceleration and the target trajectory acceleration.
[0179] In other words, the difference between the robot's current trajectory velocity and the target trajectory velocity is used to obtain the first loss component representing the difference between the current trajectory velocity and the target trajectory velocity; the difference between the robot's current trajectory acceleration and the target trajectory acceleration is used to obtain the second loss component representing the difference between the current trajectory acceleration and the target trajectory acceleration.
[0180] Finally, the first loss is obtained by combining the first loss component and the second loss component.
[0181] In this way, the robot's actual trajectory velocity will get closer and closer to the target trajectory velocity that matches the style, and the robot's actual trajectory acceleration will get closer and closer to the target trajectory acceleration that matches the style.
[0182] For example, the first loss Among them, w v The weights corresponding to the first loss component. v is the current trajectory velocity. ref For the target trajectory velocity, For the first loss component, w a The weights corresponding to the second loss component. For the current trajectory acceleration, a ref Acceleration for the target trajectory, This is the second loss component.
[0183] In some alternative implementations, the second loss can be obtained in, but is not limited to, the following ways: A third loss component is obtained based on the difference between the robot's current angular velocity and the target angular velocity; wherein the target angular velocity is determined based on a preset angular velocity, the sixth component, and the seventh component; A fourth loss component is obtained based on the difference between the robot's current angular acceleration and the target angular acceleration; wherein the target angular acceleration is determined based on a preset angular acceleration and the sixth component; The second loss is obtained based on the third loss component and the fourth loss component.
[0184] First, the target angular velocity is determined based on the preset angular velocity, the sixth component, and the seventh component; the target angular acceleration is determined based on the preset angular acceleration and the sixth component.
[0185] Taking the sixth component as arousal and the seventh component as attention intensity as an example, the more excited and attentive the target is, the faster it is necessary to approach the target.
[0186] For example, the target angular velocity w ref =w0 (1+w4 u A +w4 u ar ); Target angular acceleration z ref =z0 (1+w6 u ar ); where w0 is the preset angular velocity, u A To note the intensity, u ar For wake-up level, z0 is the preset angular acceleration, and w4, w5, and w6 are adjustment coefficients that can be set according to actual needs.
[0187] Next, a third loss component is obtained based on the difference between the robot's current angular velocity and the target angular velocity; and a fourth loss component is obtained based on the difference between the robot's current angular acceleration and the target angular acceleration.
[0188] In other words, the difference between the robot's current angular velocity and the target angular velocity is used to obtain the third loss component, which represents the difference between the current angular velocity and the target angular velocity; the difference between the robot's current angular acceleration and the target angular acceleration is used to obtain the fourth loss component, which represents the difference between the current angular acceleration and the target angular acceleration.
[0189] Finally, the second loss is obtained by combining the third and fourth loss components.
[0190] In this way, the robot's actual angular velocity will get closer and closer to the target angular velocity that matches the style, and the robot's actual angular acceleration will get closer and closer to the target angular acceleration that matches the style.
[0191] For example, the second loss Among them, w ω The weights corresponding to the third loss component. w is the current angular velocity. ref For the target angular velocity, For the third loss component, w a The weights corresponding to the fourth loss component. Let z be the current angular acceleration. ref For the target angular acceleration, This is the fourth loss component.
[0192] In some optional implementations, based on any of the above embodiments, the following steps may also be performed: In response to the target action, the corresponding additional facial expression is determined, and the robot's facial expression is controlled based on the additional facial expression.
[0193] The above embodiments achieve control consistency, that is, for time t, a unified vector is obtained to perform coordinated control of facial expressions and actions.
[0194] Based on this, the embodiments of this application also trigger additional facial expressions for the target action to achieve rhythm consistency.
[0195] The following are some specific examples to illustrate this: When an action "enters," the facial expression "blinks" is triggered. When an action "enters a buffer," the "smiling" expression is triggered. When an action "pauses", the facial expression "shakes". When an action is "stopped," the facial expression "blinking" is triggered.
[0196] The above-mentioned correspondences between target actions and attached expressions are merely illustrative examples; in practice, other correspondences may be set.
[0197] In this embodiment, for a specific target action, by triggering a specific additional facial expression, the mechanical mechanism dominates the rhythm, and the facial expression follows and adjusts accordingly. In this way, the facial expression and the action have a "common beat" in terms of timing, achieving rhythmic consistency and further improving the cross-modal synchronization effect.
[0198] like Figure 5 As shown in the figure, this application embodiment provides a robot facial expression and motion coordinated control device 500, the device comprising: The vector acquisition module 501 is used to acquire an attention vector and an emotion vector; wherein, the attention vector includes a first initial component corresponding to each of multiple attention information for the target object, and the emotion vector includes a second initial component corresponding to each of multiple emotion information for the target object. The vector generation module 502 is used to generate a unified vector based on the attention vector and the emotion vector; wherein, the unified vector includes a first target component corresponding to each of the plurality of attention information and a second target component corresponding to each of the plurality of emotion information. The control module 503 is used to control the robot's facial expressions and movements based on the unified vector.
[0199] In some alternative implementations, the control module 503 is specifically used to control facial expressions in the following ways: For any expression control point, based on the preset vector corresponding to the expression control point and combined with the first component, the deformation offset of the expression control point is obtained; wherein, the first component is one or more of a plurality of second target components; Based on the second component, the translation offset is obtained; wherein the second component is one or more of a plurality of first target components; For any expression control point, a target offset is obtained based on the deformation offset and the translation offset of the expression control point, and the expression control point is adjusted based on the target offset.
[0200] In some alternative implementations, when a blink is triggered, the control module 503 is specifically used to: Based on the target offset and combined with the blinking model, the expression control points are dynamically adjusted. The blinking model is generated based on the target blink interval, the target blink period, and the target blink direction; the target blink interval and the target blink period are both determined based on a third component, which is one or more of the plurality of second target components.
[0201] In some optional implementations, before adjusting the expression control points based on the target offset, the control module 503 is further configured to: The target object is determined not to meet the out-of-focus condition; wherein the out-of-focus condition includes the target object being lost, and / or the target object undergoing a sudden change.
[0202] In some optional implementations, the control module 503 is further configured to: control facial expressions when the target object meets the defocus condition by means of the following: The target movement speed is obtained based on the current position of the target expression control point; wherein, the target expression control point is the center point of the pupil or the center point of the eye socket; Based on the target movement speed, the control points for each facial expression are dynamically adjusted.
[0203] In some alternative implementations, the control module 503 is specifically used for: Based on the target offset, the facial expression control points are adjusted in conjunction with the target compensation amount; The target compensation amount is determined based on some or all of the first compensation amount, the second compensation amount, and the third compensation amount; the first compensation amount is used to enhance the visual tension of the mechanical action; the second compensation amount is used to simulate changes in spatial depth; and the third compensation amount is used to compensate for the continuation of facial expression at the end of the action.
[0204] In some optional implementations, the control module 503 is specifically configured to obtain the first compensation amount in the following manner: The first compensation amount is obtained by adjusting the direction mapping vector based on the first coefficient. Wherein, the direction mapping vector is obtained based on the target offset; the first coefficient is positively correlated with the fourth component, the fifth component and the target angle variable respectively; the fourth component is one or more of the plurality of second target components; the fifth component is one or more of the plurality of first target components; the target angle variable is the difference between the robot's current target angle and the initial target angle.
[0205] In some alternative implementations, the control module 503 is specifically configured to obtain the second compensation amount in the following manner: Obtain the depth factor corresponding to the robot's current target angle; Based on the depth factor and the second coefficient, a target scaling factor is obtained; wherein the second coefficient is positively correlated with the fourth component and the fifth component respectively; the fourth component is one or more of the plurality of second target components; and the fifth component is one or more of the plurality of first target components. The second compensation amount is generated based on the target scaling factor.
[0206] In some optional implementations, the control module 503 is specifically configured to obtain the third compensation amount in the following manner: Obtain the terminal weight corresponding to the current angular velocity; wherein the terminal weight is negatively correlated with the current angular velocity; The third compensation amount is obtained based on the orientation mapping vector, the target angle variable, the end-effector weight, and the third coefficient; wherein the orientation mapping vector is obtained based on the target offset; the third coefficient is positively correlated with the fourth component and the fifth component respectively; the fourth component is one or more of the plurality of second target components; the fifth component is one or more of the plurality of first target components; and the target angle variable is the difference between the robot's current target angle and the initial target angle.
[0207] In some alternative implementations, the control module 503 is specifically used to control the action in the following ways: The robot's position trajectory is adjusted based on a first loss; and the robot's posture is adjusted based on a second loss. The first loss and the second loss are both determined based on the sixth component and the seventh component; the sixth component is one or more of a plurality of second target components; and the seventh component is one or more of a plurality of first target components.
[0208] In some alternative implementations, the control module 503 is specifically configured to obtain the first loss in the following manner: A first loss component is obtained based on the difference between the robot's current trajectory speed and the target trajectory speed; wherein the target trajectory speed is determined based on a preset trajectory speed, the sixth component, and the seventh component; A second loss component is obtained based on the difference between the robot's current trajectory acceleration and the target trajectory acceleration; wherein the target trajectory acceleration is determined based on a preset trajectory acceleration and the sixth component. The first loss is obtained based on the first loss component and the second loss component.
[0209] In some alternative implementations, the control module 503 is specifically configured to obtain the second loss in the following manner: A third loss component is obtained based on the difference between the robot's current angular velocity and the target angular velocity; wherein the target angular velocity is determined based on a preset angular velocity, the sixth component, and the seventh component; A fourth loss component is obtained based on the difference between the robot's current angular acceleration and the target angular acceleration; wherein the target angular acceleration is determined based on a preset angular acceleration and the sixth component; The second loss is obtained based on the third loss component and the fourth loss component.
[0210] In some alternative implementations, the control module 503 is also used for: In response to the target action, the corresponding additional facial expression is determined, and the robot's facial expression is controlled based on the additional facial expression.
[0211] Since this device is the same as the device in the method of this application embodiment, and the principle of the device in solving the problem is similar to that of the method, the implementation of the device can be referred to the implementation of the method, and the repeated parts will not be described again.
[0212] Based on the same technical concept, this application also provides a robot 600, such as... Figure 6 As shown, it includes at least one processor 601 and a memory 602 connected to at least one processor. In this embodiment, the specific connection medium between the processor 601 and the memory 602 is not limited. Figure 6 Taking the connection between processor 601 and memory 602 via bus 603 as an example. The bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0213] The processor 601 serves as the robot's facial expression and motion coordination control center. It connects to various parts of the robot via various interfaces and wiring, and performs data processing by running or executing instructions stored in the memory 602 and accessing data stored in the memory 602. Optionally, the processor 601 may include one or more processing units. The processor 601 may integrate an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and applications, while the modem processor primarily handles issuing instructions. It is understood that the modem processor may not be integrated into the processor 601. In some embodiments, the processor 601 and the memory 602 may be implemented on the same chip; in other embodiments, they may be implemented on separate chips.
[0214] The processor 601 can be a general-purpose processor, such as a CPU, digital signal processor, application-specific integrated circuit (ASIC), field-programmable gate array or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the robot expression and motion collaborative control method can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.
[0215] Memory 602, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 602 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. Memory 602 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In the embodiments of this application, memory 602 can also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.
[0216] In this embodiment, the memory 602 stores a computer program, which, when executed by the processor 601, causes the processor 601 to perform the following: Obtain an attention vector and an emotion vector; wherein, the attention vector includes a first initial component corresponding to each of multiple attention information for the target object, and the emotion vector includes a second initial component corresponding to each of multiple emotion information for the target object; A unified vector is generated based on the attention vector and the emotion vector; wherein the unified vector includes a first target component corresponding to each of the multiple attention information and a second target component corresponding to each of the multiple emotion information. The robot's facial expression and motion are controlled based on the unified vector.
[0217] In some alternative implementations, processor 601 specifically performs: For any expression control point, based on the preset vector corresponding to the expression control point and combined with the first component, the deformation offset of the expression control point is obtained; wherein, the first component is one or more of a plurality of second target components; Based on the second component, the translation offset is obtained; wherein the second component is one or more of a plurality of first target components; For any expression control point, a target offset is obtained based on the deformation offset and the translation offset of the expression control point, and the expression control point is adjusted based on the target offset.
[0218] In some alternative implementations, when a blink is triggered, processor 601 specifically executes: Based on the target offset and combined with the blinking model, the expression control points are dynamically adjusted. The blinking model is generated based on the target blink interval, the target blink period, and the target blink direction; the target blink interval and the target blink period are both determined based on a third component, which is one or more of the plurality of second target components.
[0219] In some alternative implementations, before adjusting the facial expression control points based on the target offset, the processor 601 further performs: The target object is determined not to meet the out-of-focus condition; wherein the out-of-focus condition includes the target object being lost, and / or the target object undergoing a sudden change.
[0220] In some alternative implementations, the processor 601 further performs the following: when the target object meets the defocus condition, expression control is performed in the following manner: The target movement speed is obtained based on the current position of the target expression control point; wherein, the target expression control point is the center point of the pupil or the center point of the eye socket; Based on the target movement speed, the control points for each facial expression are dynamically adjusted.
[0221] In some alternative implementations, processor 601 specifically performs: Based on the target offset, the facial expression control points are adjusted in conjunction with the target compensation amount; The target compensation amount is determined based on some or all of the first compensation amount, the second compensation amount, and the third compensation amount; the first compensation amount is used to enhance the visual tension of the mechanical action; the second compensation amount is used to simulate changes in spatial depth; and the third compensation amount is used to compensate for the continuation of facial expression at the end of the action.
[0222] In some alternative implementations, processor 601 specifically performs: The first compensation amount is obtained by adjusting the direction mapping vector based on the first coefficient. Wherein, the direction mapping vector is obtained based on the target offset; the first coefficient is positively correlated with the fourth component, the fifth component and the target angle variable respectively; the fourth component is one or more of the plurality of second target components; the fifth component is one or more of the plurality of first target components; the target angle variable is the difference between the robot's current target angle and the initial target angle.
[0223] In some alternative implementations, processor 601 specifically performs: Obtain the depth factor corresponding to the robot's current target angle; Based on the depth factor and the second coefficient, a target scaling factor is obtained; wherein the second coefficient is positively correlated with the fourth component and the fifth component respectively; the fourth component is one or more of the plurality of second target components; and the fifth component is one or more of the plurality of first target components. The second compensation amount is generated based on the target scaling factor.
[0224] In some alternative implementations, processor 601 specifically performs: Obtain the terminal weight corresponding to the current angular velocity; wherein the terminal weight is negatively correlated with the current angular velocity; The third compensation amount is obtained based on the orientation mapping vector, the target angle variable, the end-effector weight, and the third coefficient; wherein the orientation mapping vector is obtained based on the target offset; the third coefficient is positively correlated with the fourth component and the fifth component respectively; the fourth component is one or more of the plurality of second target components; the fifth component is one or more of the plurality of first target components; and the target angle variable is the difference between the robot's current target angle and the initial target angle.
[0225] In some alternative implementations, processor 601 specifically performs: The robot's position trajectory is adjusted based on a first loss; and the robot's posture is adjusted based on a second loss. The first loss and the second loss are both determined based on the sixth component and the seventh component; the sixth component is one or more of a plurality of second target components; and the seventh component is one or more of a plurality of first target components.
[0226] In some alternative implementations, processor 601 specifically performs: A first loss component is obtained based on the difference between the robot's current trajectory speed and the target trajectory speed; wherein the target trajectory speed is determined based on a preset trajectory speed, the sixth component, and the seventh component; A second loss component is obtained based on the difference between the robot's current trajectory acceleration and the target trajectory acceleration; wherein the target trajectory acceleration is determined based on a preset trajectory acceleration and the sixth component. The first loss is obtained based on the first loss component and the second loss component.
[0227] In some alternative implementations, processor 601 specifically performs: A third loss component is obtained based on the difference between the robot's current angular velocity and the target angular velocity; wherein the target angular velocity is determined based on a preset angular velocity, the sixth component, and the seventh component; A fourth loss component is obtained based on the difference between the robot's current angular acceleration and the target angular acceleration; wherein the target angular acceleration is determined based on a preset angular acceleration and the sixth component; The second loss is obtained based on the third loss component and the fourth loss component.
[0228] In some alternative implementations, processor 601 also performs: In response to the target action, the corresponding additional facial expression is determined, and the robot's facial expression is controlled based on the additional facial expression.
[0229] Based on the same technical concept, embodiments of this application also provide a computer-readable storage medium storing a computer program executable by a processor, which, when run on the processor, causes the processor to perform the steps of the above-described robot facial expression and motion coordinated control method.
[0230] In some alternative implementations, various aspects of the robot facial expression and motion coordinated control method provided in this application can also be implemented as a program product containing computer-executable instructions. When the program product is run on a computer device, the computer-executable instructions are used to cause the computer device to perform the steps of the robot facial expression and motion coordinated control method according to the various exemplary embodiments of this application described above.
[0231] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0232] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0233] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0234] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0235] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0236] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for coordinated control of robot facial expressions and movements, characterized in that, The method includes: Obtain an attention vector and an emotion vector; wherein, the attention vector includes a first initial component corresponding to each of multiple attention information for the target object, and the emotion vector includes a second initial component corresponding to each of multiple emotion information for the target object; A unified vector is generated based on the attention vector and the emotion vector; wherein the unified vector includes a first target component corresponding to each of the multiple attention information and a second target component corresponding to each of the multiple emotion information. The robot's facial expression and motion are controlled based on the unified vector.
2. The method as described in claim 1, characterized in that, Control facial expressions in the following ways: For any expression control point, based on the preset vector corresponding to the expression control point and combined with the first component, the deformation offset of the expression control point is obtained; wherein, the first component is one or more of a plurality of second target components; Based on the second component, the translation offset is obtained; wherein the second component is one or more of a plurality of first target components; For any expression control point, a target offset is obtained based on the deformation offset and the translation offset of the expression control point, and the expression control point is adjusted based on the target offset.
3. The method as described in claim 2, characterized in that, When a blink is triggered, the facial expression control point is adjusted based on the target offset, including: Based on the target offset and combined with the blinking model, the expression control points are dynamically adjusted. The blinking model is generated based on the target blink interval, the target blink period, and the target blink direction; the target blink interval and the target blink period are both determined based on a third component, which is one or more of the plurality of second target components.
4. The method as described in claim 2, characterized in that, Before adjusting the facial expression control points based on the target offset, the method further includes: The target object is determined not to meet the out-of-focus condition; wherein the out-of-focus condition includes the target object being lost, and / or the target object undergoing a sudden change.
5. The method as described in claim 4, characterized in that, Also includes: When the target object meets the defocus condition, facial expression control is performed in the following manner: The target movement speed is obtained based on the current position of the target expression control point; wherein, the target expression control point is the center point of the pupil or the center point of the eye socket; Based on the target movement speed, the control points for each facial expression are dynamically adjusted.
6. The method as described in claim 2, characterized in that, Adjusting the facial expression control points based on the target offset includes: Based on the target offset, the facial expression control points are adjusted in conjunction with the target compensation amount; The target compensation amount is determined based on some or all of the first compensation amount, the second compensation amount, and the third compensation amount; the first compensation amount is used to enhance the visual tension of the mechanical action; the second compensation amount is used to simulate changes in spatial depth; and the third compensation amount is used to compensate for the continuation of facial expression at the end of the action.
7. The method as described in claim 6, characterized in that, The first compensation amount is obtained in the following manner: The first compensation amount is obtained by adjusting the direction mapping vector based on the first coefficient. Wherein, the direction mapping vector is obtained based on the target offset; the first coefficient is positively correlated with the fourth component, the fifth component and the target angle variable respectively; the fourth component is one or more of the plurality of second target components; the fifth component is one or more of the plurality of first target components; the target angle variable is the difference between the robot's current target angle and the initial target angle.
8. The method as described in claim 6, characterized in that, The second compensation amount is obtained in the following manner: Obtain the depth factor corresponding to the robot's current target angle; Based on the depth factor and the second coefficient, a target scaling factor is obtained; wherein the second coefficient is positively correlated with the fourth component and the fifth component respectively; the fourth component is one or more of the plurality of second target components; and the fifth component is one or more of the plurality of first target components. The second compensation amount is generated based on the target scaling factor.
9. The method as described in claim 6, characterized in that, The third compensation amount is obtained in the following manner: Obtain the terminal weight corresponding to the current angular velocity; wherein the terminal weight is negatively correlated with the current angular velocity; The third compensation amount is obtained based on the orientation mapping vector, the target angle variable, the end-effector weight, and the third coefficient; wherein the orientation mapping vector is obtained based on the target offset; the third coefficient is positively correlated with the fourth component and the fifth component respectively; the fourth component is one or more of the plurality of second target components; the fifth component is one or more of the plurality of first target components; and the target angle variable is the difference between the robot's current target angle and the initial target angle.
10. The method as described in claim 1, characterized in that, Motion control can be achieved in the following ways: The robot's position trajectory is adjusted based on a first loss; and the robot's posture is adjusted based on a second loss. The first loss and the second loss are both determined based on the sixth component and the seventh component; the sixth component is one or more of a plurality of second target components; and the seventh component is one or more of a plurality of first target components.
11. The method as described in claim 10, characterized in that, The first loss is obtained in the following manner: A first loss component is obtained based on the difference between the robot's current trajectory speed and the target trajectory speed; wherein the target trajectory speed is determined based on a preset trajectory speed, the sixth component, and the seventh component; A second loss component is obtained based on the difference between the robot's current trajectory acceleration and the target trajectory acceleration; wherein the target trajectory acceleration is determined based on a preset trajectory acceleration and the sixth component. The first loss is obtained based on the first loss component and the second loss component.
12. The method as described in claim 10, characterized in that, The second loss is obtained in the following manner: A third loss component is obtained based on the difference between the robot's current angular velocity and the target angular velocity; wherein the target angular velocity is determined based on a preset angular velocity, the sixth component, and the seventh component; A fourth loss component is obtained based on the difference between the robot's current angular acceleration and the target angular acceleration; wherein the target angular acceleration is determined based on a preset angular acceleration and the sixth component; The second loss is obtained based on the third loss component and the fourth loss component.
13. The method according to any one of claims 1 to 12, characterized in that, Also includes: In response to the target action, the corresponding additional facial expression is determined, and the robot's facial expression is controlled based on the additional facial expression.
14. A robot facial expression and motion coordinated control device, characterized in that, The device includes: A vector acquisition module is used to acquire an attention vector and an emotion vector; wherein, the attention vector includes a first initial component corresponding to each of multiple attention information for the target object, and the emotion vector includes a second initial component corresponding to each of multiple emotion information for the target object. A vector generation module is used to generate a unified vector based on the attention vector and the emotion vector; wherein the unified vector includes a first target component corresponding to each of the plurality of attention information and a second target component corresponding to each of the plurality of emotion information. The control module is used to control the robot's facial expressions and movements based on the unified vector.
15. A robot, characterized in that, It includes at least one processor and at least one memory, wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform the robot facial expression and motion coordinated control method as described in any one of claims 1 to 13.