Intelligent control method and device for bionic animal toy
Patent Information
- Application Number
- CN202511160432.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2045-08-19
AI Technical Summary
[0003]然而,现有仿生交互系统多采用单一传感器(如触摸传感器、简单视觉模块)采集用户交互信息,通过预设规则或基础机器学习模型解析用户行为(如识别“触摸”、“拍打”等简单动作)并调用固定动作库生成响应,难以精准提取用户的行为语义与情感特征,例如:仅通过简单的触摸传感器识别用户触摸动作难以判断用户的行为意图是玩耍、安抚还是探索;面部表情识别精度不足,无法准确捕捉用户的细微情感变化;缺乏对用户情感与行为关联规律的捕捉;导致仿生动物玩具在交互时无法根据用户实时状态做出个性化、情感化的反馈,动作生硬且缺乏与用户情感的共鸣,交互体验单调、缺乏深度,因此,如何通过构建多模态交互线索解析机制提取用户的行为语义与情感特征以实现仿生动物玩具的个性情感反馈成为了业界面临的难题
本申请提供的仿生动物玩具的智能控制方法及装置中,首先采集目标用户与仿生动物玩具进行交互时的动作图像序列;在所述动作图像序列中提取目标用户手部关节点的三维空间轨迹,基于所述三维空间轨迹解析目标用户的行为语义,进而生成目标用户对仿生动物玩具的行为意图向量;同步对所述动作图像序列中的面部微表情进行识别,得到目标用户与仿生动物玩具进行交互时的情感特征张量,进而基于所述情感特征张量和所述行为意图向量构建目标用户在互动过程中情感和行为的时空关联图谱;根据目标用户与仿生动物玩具的历史交互数据确定仿生动物玩具人格化的状态特征,进而基于所述时空关联图谱和所述状态特征生成符合仿生动物玩具人格化状态的运动基元序列;将所述运动基元序列转化为舵机控制指令流控制仿生动物玩具执行带情感反馈的交互动作。
Smart Images

Figure CN121187439B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of human-computer interaction technology, and more specifically, to an intelligent control method and device for a biomimetic animal toy. Background Technology
[0002] With the improvement of people's living standards and the growth of demand for interactive experiences, biomimetic animal toys have received much attention in the fields of entertainment and education. People expect these toys to not only be highly realistic in appearance, but also to show natural, intelligent and emotional responses during interaction, just like interacting with real animals, thereby enhancing users' sense of participation and immersion and creating richer and more interesting interactive experiences.
[0003] However, existing biomimetic interaction systems mostly use a single sensor (such as a touch sensor or a simple vision module) to collect user interaction information. They analyze user behavior (such as recognizing simple actions like "touching" or "patting") through preset rules or basic machine learning models and generate responses by calling a fixed action library. This makes it difficult to accurately extract the semantic and emotional features of user behavior. For example, it is difficult to determine whether the user's intention is to play, soothe, or explore by simply recognizing the user's touch actions using a simple touch sensor; facial expression recognition is not accurate enough to accurately capture subtle changes in the user's emotions; and there is a lack of capture of the correlation between user emotions and behavior. As a result, biomimetic animal toys cannot provide personalized and emotional feedback based on the user's real-time state during interaction. The movements are stiff and lack resonance with the user's emotions, resulting in a monotonous and shallow interactive experience. Therefore, how to extract the semantic and emotional features of user behavior by constructing a multimodal interaction cue parsing mechanism to achieve personalized emotional feedback for biomimetic animal toys has become a challenge for the industry. Summary of the Invention
[0004] This application provides an intelligent control method and device for bionic animal toys, which can extract the user's behavioral semantics and emotional features by constructing a multimodal interactive cue parsing mechanism to achieve personalized emotional feedback of bionic animal toys.
[0005] Firstly, this application provides an intelligent control method for a biomimetic animal toy, which controls the biomimetic animal toy to provide emotional feedback interaction based on the target user's actions and facial micro-expressions in a user interaction scenario. The method includes: Collect motion image sequences of target users interacting with bionic animal toys; The three-dimensional spatial trajectory of the target user's hand joints is extracted from the motion image sequence. The behavioral semantics of the target user are analyzed based on the three-dimensional spatial trajectory, and then a behavioral intention vector of the target user towards the bionic animal toy is generated. Simultaneously, facial micro-expressions in the action image sequence are identified to obtain the emotional feature tensor when the target user interacts with the bionic animal toy. Then, based on the emotional feature tensor and the behavioral intention vector, a spatiotemporal correlation map of the target user's emotions and behaviors during the interaction is constructed. Based on the historical interaction data between the target user and the bionic animal toy, the state characteristics of the bionic animal toy's personification are determined, and then a sequence of motion primitives that conforms to the personification state of the bionic animal toy is generated based on the spatiotemporal correlation map and the state characteristics. The motion primitive sequence is converted into a servo control command stream to control the bionic animal toy to perform interactive actions with emotional feedback.
[0006] In some embodiments, extracting the three-dimensional spatial trajectory of the target user's hand joint points from the motion image sequence specifically includes: Hand instance segmentation is performed on each depth motion image in the motion image sequence to obtain the hand region mask for each depth motion image; By aligning all hand region masks to locate the pixel coordinates of hand joints, a three-dimensional coordinate set of the target user's hand joints is obtained; By performing occlusion point compensation on the three-dimensional coordinate set, an anti-disturbance joint point sequence of the target user's hand joint points is obtained; The three-dimensional spatial trajectory of the target user's hand joints is generated using the anti-disturbance joint sequence.
[0007] In some embodiments, the process of parsing the target user's behavioral semantics based on the three-dimensional spatial trajectory to generate a behavioral intent vector of the target user towards the bionic animal toy specifically includes: Extract the motion features of the target user's hand joints from the three-dimensional spatial trajectory; Based on the motion characteristics, the three-dimensional spatial trajectory is divided into multiple discrete spatiotemporal action primitives; Based on a pre-established biomimetic interactive dictionary, the semantic strength of each spatiotemporal action primitive is quantified, thereby generating a semantic encoding vector of the target user's hand actions. The target user's behavioral intent vector for the bionic animal toy is determined based on the semantic encoding vector.
[0008] In some embodiments, the simultaneous recognition of facial micro-expressions in the motion image sequence to obtain the emotional feature tensor of the target user interacting with the bionic animal toy specifically includes: Facial region localization is performed on the motion image sequence to obtain the facial image sequence of the target user; Dynamic features of the target user's facial micro-expressions are extracted from the facial image sequence; The three-dimensional emotion value of each instantaneous expression of the target user is determined based on the dynamic characteristics; Spatiotemporal expansion of all three-dimensional emotion values is performed to construct an emotion feature tensor for the target user when interacting with the bionic animal toy.
[0009] In some embodiments, constructing a spatiotemporal correlation graph of the target user's emotions and behaviors during the interaction process based on the emotional feature tensor and the behavioral intention vector specifically includes: The emotional feature tensor and the behavioral intention vector are time-aligned to obtain time-synchronized emotional-behavioral data pairs; A spatiotemporal correlation matrix is constructed using the emotional state and behavioral intention in the aforementioned emotion-behavior data pairs as nodes; Based on the spatiotemporal correlation matrix, a spatiotemporal correlation map of the target user's emotions and behaviors during the interaction process is determined.
[0010] In some embodiments, determining the personified state characteristics of the bionic animal toy based on historical interaction data between the target user and the bionic animal toy specifically includes: Acquire historical interaction data between target users and bionic animal toys; The historical interaction data is cleaned and segmented into data fragments to generate time-series interaction samples with emotion tags; The interaction features in the time-series interaction samples are extracted to construct a five-dimensional personality feature vector for the bionic toy; The state characteristics of the bionic animal toy are determined based on the five-dimensional personality feature vector.
[0011] In some embodiments, generating a sequence of motion primitives that conforms to the personified state of a biomimetic animal toy based on the spatiotemporal correlation map and the state features specifically includes: Extract the sentiment-behavior coupling matrix of the target user from the spatiotemporal correlation graph; Based on the state features and the emotion-behavior coupling matrix, the angular oscillation trajectory of each joint of the bionic animal toy is generated; Inverse kinematics solutions are performed on all angular oscillation trajectories to generate a sequence of motion primitives that conform to the personified state of the bionic animal toy.
[0012] Secondly, this application provides an intelligent control device for a biomimetic animal toy, the intelligent control device comprising: The acquisition module is used to acquire motion image sequences of target users interacting with bionic animal toys; The processing module is used to extract the three-dimensional spatial trajectory of the target user's hand joints from the motion image sequence, analyze the target user's behavioral semantics based on the three-dimensional spatial trajectory, and then generate a behavioral intention vector of the target user towards the bionic animal toy. The processing module is used to simultaneously recognize facial micro-expressions in the action image sequence, obtain the emotional feature tensor when the target user interacts with the bionic animal toy, and then construct a spatiotemporal correlation map of the target user's emotions and behaviors during the interaction based on the emotional feature tensor and the behavioral intention vector. The processing module is used to determine the personified state characteristics of the bionic animal toy based on the historical interaction data between the target user and the bionic animal toy, and then generate a motion element sequence that conforms to the personified state of the bionic animal toy based on the spatiotemporal correlation map and the state characteristics. The execution module is used to convert the motion primitive sequence into a servo control command stream to control the bionic animal toy to perform interactive actions with emotional feedback.
[0013] Thirdly, this application provides a computer device including a memory and a processor, the memory storing code, and the processor being configured to acquire the code and execute the above-described intelligent control method for bionic animal toys.
[0014] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described intelligent control method for biomimetic animal toys.
[0015] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: The intelligent control method and device for bionic animal toys provided in this application first collects a sequence of motion images of a target user interacting with the bionic animal toy; extracts the three-dimensional spatial trajectory of the target user's hand joints from the motion image sequence, analyzes the target user's behavioral semantics based on the three-dimensional spatial trajectory, and then generates a behavioral intention vector of the target user towards the bionic animal toy; simultaneously, facial micro-expressions in the motion image sequence are recognized to obtain an emotional feature tensor of the target user interacting with the bionic animal toy, and then a spatiotemporal correlation graph of the target user's emotions and behaviors during the interaction is constructed based on the emotional feature tensor and the behavioral intention vector; the personified state characteristics of the bionic animal toy are determined according to the historical interaction data of the target user and the bionic animal toy, and then a motion primitive sequence that conforms to the personified state of the bionic animal toy is generated based on the spatiotemporal correlation graph and the state characteristics; the motion primitive sequence is converted into a servo control command stream to control the bionic animal toy to perform interactive actions with emotional feedback.
[0016] Therefore, this application transforms the motion primitive sequence into a servo control command flow to control the bionic animal toy to perform interactive actions with emotional feedback. First, determining the behavioral intent vector yields a probability vector representing the interactive behavioral intent between the target user and the bionic animal toy. This determination is achieved by analyzing the behavioral semantics of the target user's hand joints' three-dimensional spatial trajectory, generating a probability vector containing specific interactive intents (such as playing, comforting, or giving commands). This enables accurate decoding of the user's complex actions, solving the problem of ambiguous intent judgment under a single sensor. It can be used to clarify the target user's interactive needs, providing a decision-making basis for the bionic animal toy to generate feedback actions, ensuring the matching degree between the bionic animal toy's response and the target user's intent, improving the naturalness of the interaction, and providing accurate behavioral dimension data for the subsequent construction of the "emotion-behavior spatiotemporal correlation graph," ensuring more accurate spatiotemporal correlation modeling of emotional features and behavioral intent. Then, determining the spatiotemporal correlation graph yields a graph structure model displaying the dynamic correlation strength between emotion and behavior when the target user interacts with the bionic animal toy. The determination of the spatiotemporal correlation graph captures the dynamic coupling law between the target user's emotion and behavior through the graph structure, intuitively displaying the real-time relationship between emotion and behavior. The correlation patterns (such as the strong correlation between "high pleasure" and "play intention" over time) overcome the limitations of existing technologies that separate the analysis of emotions and behaviors, and realize the spatiotemporal correlation modeling of multimodal interaction cues. This provides multi-dimensional correlation evidence for generating response actions in bionic toys that conform to the user's emotional state, improves the naturalness and intelligence of the interaction, and provides correlation evidence for the subsequent generation of personalized responses. Finally, the determination of the motion primitive sequence can obtain a coherent action sequence composed of executable basic action units (such as affectionate head rubbing and excited tail wagging) transformed from the joint angle trajectory of the bionic animal toy. The motion primitive sequence serves as a connection abstraction. As a core bridge between the association model and the execution of specific actions, it ensures that the response of the bionic animal toy not only conforms to the user's real-time emotional and behavioral logic, but also reflects stable personality traits. This compensates for the lack of flexibility and emotional adaptability of the fixed action library in the existing technology, making the interactive actions natural, smooth, and personalized. It significantly enhances the immersive experience and emotional resonance of the user experience, and provides key action generation support for the overall solution to achieve "personalized emotional feedback". In summary, based on the above solution, a multimodal interaction cue parsing mechanism can be constructed to extract the user's behavioral semantics and emotional features to achieve personalized emotional feedback for the bionic animal toy. Attached Figure Description
[0017] Figure 1 This is an exemplary flowchart of an intelligent control method for a biomimetic animal toy according to some embodiments of this application; Figure 2 This is a flowchart illustrating the operation of determining a three-dimensional spatial trajectory according to some embodiments of this application; Figure 3This is an exemplary flowchart illustrating the determination of an emotion feature tensor according to some embodiments of this application; Figure 4 This is a schematic diagram of the structure of an intelligent control device for a biomimetic animal toy according to some embodiments of this application; Figure 5 This is an internal structural diagram of a computer device for implementing an intelligent control method for biomimetic animal toys, according to some embodiments of this application. Detailed Implementation
[0018] To better understand the technical solution of this application, the technical solution of this application will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0019] refer to Figure 1 The figure is an exemplary flowchart of an intelligent control method for a bionic animal toy according to some embodiments of this application. The intelligent control method for the bionic animal toy mainly includes the following steps: In step 101, a sequence of motion images of the target user interacting with the bionic animal toy is acquired.
[0020] It should be noted that, in this application, the motion image sequence is a set of depth motion images recording the changes in limb movements and facial expressions of the target user during interaction with the bionic animal toy. This motion image sequence can provide continuous and complete raw data for subsequent extraction of hand joint trajectories, identification of emotional features, and analysis of interaction intentions. In specific implementation, the acquisition of the motion image sequence when the target user interacts with the bionic animal toy can be achieved in the following way: a depth binocular camera module integrated on the bionic animal toy can be used to acquire depth motion images of the target user interacting with the bionic animal toy, and the sequence of all acquired depth motion images arranged in chronological order can be used as the motion image sequence when the target user interacts with the bionic animal toy. The frame rate of the depth binocular camera module acquiring the depth motion images can be set to 30 frames per second to ensure the continuity of the movements. The above is only an example, and other visual sensors or image acquisition devices can also be used in practice. This is only an example and is not intended to limit the specific application.
[0021] In step 102, the three-dimensional spatial trajectory of the target user's hand joints is extracted from the motion image sequence. The behavioral semantics of the target user are analyzed based on the three-dimensional spatial trajectory, and then a behavioral intention vector of the target user towards the bionic animal toy is generated.
[0022] In some embodiments, reference Figure 2 The figure is a flowchart illustrating the operation of determining a three-dimensional spatial trajectory according to some embodiments of this application. The extraction of the three-dimensional spatial trajectory of the target user's hand joint points from the motion image sequence in this application can be achieved using the following steps: Hand instance segmentation is performed on each depth motion image in the motion image sequence to obtain the hand region mask for each depth motion image; By aligning all hand region masks to locate the pixel coordinates of hand joints, a three-dimensional coordinate set of the target user's hand joints is obtained; By performing occlusion point compensation on the three-dimensional coordinate set, an anti-disturbance joint point sequence of the target user's hand joint points is obtained; The three-dimensional spatial trajectory of the target user's hand joints is generated using the anti-disturbance joint sequence.
[0023] In specific implementation, hand instance segmentation is performed on each depth motion image in the motion image sequence to obtain the hand region mask for each depth motion image. This can be achieved in the following way: For each depth motion image in the motion image sequence, an existing instance segmentation algorithm (such as an improved mask region convolutional neural network algorithm) can be used to segment the hand instance of the depth motion image to obtain the hand region mask of the depth motion image. Through the above steps, the hand region mask of each depth motion image in the motion image sequence can be obtained. Specifically, the instance segmentation algorithm can first extract multi-scale features of the depth motion image through a 50-layer residual network and then input the feature to the region proposal network. A candidate bounding box for the hand is generated. Then, the motion region detected by optical flow is used as a priori filter for non-hand candidate bounding boxes by combining the dynamic foreground mask module. Finally, a pixel-level binary mask of the hand region is output through the mask branch (i.e., 1 represents hand pixels and 0 represents background), thereby achieving accurate extraction of the hand region in complex backgrounds and outputting the hand region mask of the depth motion image. The hand region mask is a pixel-level region containing the target user's hand extracted from the depth motion image. This hand region mask focuses on the key interactive parts of the target user and eliminates background interference, which can provide an accurate region range for subsequent joint point localization, reduce the interference of non-hand information on feature extraction, and improve the accuracy of joint point detection.
[0024] In practice, the pixel coordinates of hand joints are located by aligning all hand region masks. The resulting 3D coordinate set of the target user's hand joints can be achieved in the following way: It can be implemented using a hand keypoint detection and tracking model within a machine learning framework (such as MediaPipe). The Hand model aligns all hand region masks and locates the pixel coordinates of hand joints within the hand region masks to obtain a 3D coordinate set of the target user's hand joints. The hand keypoint detection and tracking model first extracts texture features by aligning all hand region masks using a convolutional neural network and then captures the topological relationships between hand joints using a self-attention mechanism module, outputting 2D pixel coordinates of 21 hand joints (such as finger joints). Then, using calibration parameters (such as intrinsic and extrinsic matrices) of a binocular vision system, the 2D coordinates are converted to 3D coordinates via triangulation. Finally, the 3D coordinates of consecutive frames of each hand joint are arranged in chronological order to construct a 3D coordinate set of the target user's hand joints. This 3D coordinate set displays the spatial position information of the target user's hand joints during the interaction process. This 3D coordinate set quantifies the spatial distribution of each hand joint, providing a numerical basis for analyzing the spatial characteristics of hand movements, extending subsequent trajectory analysis from a plane to three-dimensional space, and better reflecting real-world interaction scenarios.
[0025] In specific implementation, occlusion point compensation is performed on the three-dimensional coordinate set to obtain the disturbance-resistant joint point sequence of the target user's hand joints. This can be achieved in the following way: kinematic constraint interpolation can be used to interpolate and compensate for abnormal points in the three-dimensional coordinate set that cause coordinate jumps due to occlusion, thereby obtaining the joint point sequence of each hand joint. Then, the set of all hand joint point sequences is used as the disturbance-resistant joint point sequence of the target user's hand joints. Specifically, the kinematic constraint interpolation method first detects occlusion points by measuring the motion velocity of the hand joints in the adjacent 5 frames (i.e., the average difference between the coordinates of the current frame and the previous 4 frames). (For example, if the velocity change exceeds a threshold of 0.5...) Each frame (millimeters per frame represents occlusion) then predicts reasonable coordinates of occluded points based on the 3D coordinates of unoccluded hand joints. Finally, the 3D coordinates are smoothed and compensated using Kalman filtering to generate an occlusion-resistant joint sequence, ensuring the spatiotemporal continuity of hand joint coordinates. The occlusion-resistant joint sequence is a set of continuous joint coordinate sequences formed by compensating for abnormal points caused by occlusion in the 3D coordinate set. This occlusion-resistant joint sequence can eliminate information loss caused by occlusion during human-computer interaction, ensure the spatiotemporal continuity of joint data, provide reliable input for trajectory generation, and avoid trajectory breakage or distortion caused by occlusion.
[0026] It should be noted that in this application, the three-dimensional spatial trajectory is a continuous set of curves reflecting the movement path of the target user's hand joints. This three-dimensional spatial trajectory can completely record the spatial evolution process of the target user's hand movement, providing a core basis for parsing user behavior semantics and generating behavioral intent vectors. In specific implementation, the three-dimensional spatial trajectory of the target user's hand joints can be generated through the anti-disturbance joint sequence in the following way: For the joint sequence of each hand joint in the anti-disturbance joint sequence, the discrete joint coordinates in the joint sequence can be curve fitted by B-spline interpolation first, and then local jitter can be removed by trajectory smoothness constraints (such as the rate of curvature change of adjacent points per millimeter), and the continuous three-dimensional trajectory of the hand joints changing with time can be output. The continuous three-dimensional trajectory of each hand joint can be obtained through the above steps, and the set of all continuous three-dimensional trajectories is used as the three-dimensional spatial trajectory of the target user's hand joints.
[0027] In some embodiments, the generation of a behavioral intent vector of the target user towards the biomimetic animal toy by parsing the target user's behavioral semantics based on the three-dimensional spatial trajectory can be achieved through the following steps: Extract the motion features of the target user's hand joints from the three-dimensional spatial trajectory; Based on the motion characteristics, the three-dimensional spatial trajectory is divided into multiple discrete spatiotemporal action primitives; Based on a pre-established biomimetic interactive dictionary, the semantic strength of each spatiotemporal action primitive is quantified, thereby generating a semantic encoding vector of the target user's hand actions. The target user's behavioral intent vector for the bionic animal toy is determined based on the semantic encoding vector.
[0028] In specific implementation, the motion features of the target user's hand joints extracted from the three-dimensional spatial trajectory can be achieved in the following way: An improved spatiotemporal convolutional network model (such as a spatiotemporal convolutional neural network model optimized based on the C3D architecture) can be loaded. The three-dimensional spatial trajectory is input into the spatiotemporal convolutional network model. The spatial relationship between hand joints is captured by three layers of spatial convolution (such as a 3×3 kernel, a stride of 1, and 64, 128, and 256 channels respectively). The temporal changes between different frames of each hand joint sequence are captured by two layers of temporal convolution (such as a 3×1 kernel, a stride of 1, and 256 channels respectively). A 512-dimensional motion feature vector is output through global average pooling. The motion feature is a feature vector that quantifies the essential attributes of the target user's hand movement, including hand movement speed, acceleration, and joint coordination mode. This motion feature can provide core feature basis for subsequent action segmentation and semantic parsing, reduce redundant information interference, and improve the accuracy of behavior analysis.
[0029] In specific implementation, the segmentation of the three-dimensional spatial trajectory into multiple discrete spatiotemporal action primitives based on the motion features can be achieved in the following way: An improved dynamic time warping algorithm can be used to segment the three-dimensional spatial trajectory based on the extracted motion features to obtain multiple discrete spatiotemporal action primitives. Specifically, the cosine similarity of adjacent frame features in the motion features can be calculated first. When the similarity is lower than a preset similarity threshold (e.g., 0.6, which can be dynamically adjusted based on historical interaction data) for three consecutive frames, it is determined as an action boundary. Then, the trajectory segments between the action boundaries are divided into a spatiotemporal action primitive. Each spatiotemporal action primitive contains a start frame, an end frame, and a corresponding three-dimensional trajectory segment, thereby obtaining multiple discrete spatiotemporal action primitives (e.g., "raise hand - approach", "touch - leave", etc.). The spatiotemporal action primitive is a discrete action segment containing spatiotemporal information in the three-dimensional spatial trajectory. Determining the spatiotemporal action primitive can decompose complex motion into understandable basic action units, reduce the analysis complexity of continuous trajectories, focus semantic parsing on structured units, and improve the efficiency of action understanding.
[0030] It should be noted that in this application, the bionic interaction dictionary is a data structure used to store and manage information related to bionic interaction. It includes the mapping relationship between various user interaction behaviors (such as action commands, voice commands, etc.) and the corresponding feedback behaviors of bionic animal toys (such as bionic actions, interactive voice, etc.). It also includes various common user interaction actions (such as stroking, patting, grabbing, summoning, feeding, waving, pointing, no action) and the primitive feature templates labeled for each type of interaction action. It can provide bionic animal toys with behavioral output references corresponding to the target user's interaction behavior, enabling bionic animal toys to quickly find appropriate bionic behaviors to provide feedback according to different interaction scenarios and user commands, thereby achieving natural and intelligent human-computer interaction, improving user experience, and making the behavior of bionic animal toys closer to the interactive behavior of real organisms.
[0031] In specific implementation, the semantic strength of each spatiotemporal action primitive is quantified based on a pre-established biomimetic interaction dictionary, and then the semantic encoding vector of the target user's hand movements is generated. This can be achieved in the following way: For each spatiotemporal action primitive, the Euclidean distance between the spatiotemporal action primitive and the primitive feature templates corresponding to various interactive actions in the biomimetic interaction dictionary can be calculated. The primitive feature template category corresponding to the smallest Euclidean distance is taken as the semantic label of the spatiotemporal action primitive. The reciprocal of the Euclidean distance (i.e., normalized to 0-1) represents the semantic strength of the corresponding interactive action category. The semantic strength of the Euclidean distance between the spatiotemporal action primitive and various interactive actions is then set to a fixed value. Sequentially arrange the spatiotemporal action primitives to generate semantic vectors. The higher the value in the vector, the more significant the semantics of the corresponding action category. Through the above steps, the semantic vector of each spatiotemporal action primitive can be obtained. Then, the set of semantic vectors of all spatiotemporal action primitives is used as the semantic encoding vector of the target user's hand movements. The semantic encoding vector is a set of numerical vectors that quantify the semantic intensity of various spatiotemporal actions in the target user's hand movement trajectory. This semantic encoding vector can convert spatiotemporal action primitives into computable semantic symbols, realize the standardized expression of action semantics, provide a unified input format for intent reasoning, and enhance the comparability between different actions.
[0032] It should be noted that, in this application, the behavioral intent vector is a probability vector representing the interactive behavioral intent between the target user and the bionic animal toy. This behavioral intent vector can be used to clarify the target user's interaction needs, provide a decision basis for the bionic animal toy to generate feedback actions, ensure the matching degree between the bionic animal toy's response and the target user's intent, and improve the naturalness of the interaction. In specific implementation, the behavioral intent vector of the target user towards the bionic animal toy can be determined according to the semantic encoding vector in the following way: the semantic encoding vector can be input into a pre-trained intent classification network to generate the behavioral intent vector of the target user towards the bionic animal toy. Among them, the intent classification network can adopt a "fully connected + activation function" structure. After the semantic encoding vector is mapped by two fully connected layers (such as 32 and 7 hidden units respectively), a 7-dimensional probability vector (corresponding to 7 types of intent: play, soothe, command, feed, interact, avoid, no intent) is output through an activation function (such as the softmax function). The 3 types of intents with the highest probability and their probability values are taken to form the final behavioral intent vector.
[0033] In step 103, facial micro-expressions in the motion image sequence are simultaneously identified to obtain the emotional feature tensor when the target user interacts with the bionic animal toy. Then, based on the emotional feature tensor and the behavioral intention vector, a spatiotemporal correlation map of the target user's emotions and behaviors during the interaction is constructed.
[0034] In some embodiments, reference Figure 3The figure is an exemplary flowchart illustrating the determination of the emotion feature tensor according to some embodiments of this application. In this application, the simultaneous recognition of facial micro-expressions in the action image sequence to obtain the emotion feature tensor when the target user interacts with the bionic animal toy can be achieved through the following steps: In step 1031, facial region localization is performed on the motion image sequence to obtain the facial image sequence of the target user; In step 1032, dynamic features of the target user's facial micro-expressions are extracted from the facial image sequence; In step 1033, the three-dimensional emotion value of each instantaneous expression of the target user is determined based on the dynamic features; In step 1034, all three-dimensional emotion values are spatiotemporally expanded to construct an emotion feature tensor when the target user interacts with the bionic animal toy.
[0035] In specific implementation, facial region localization of the motion image sequence to obtain the target user's facial image sequence can be achieved in the following way: an improved multi-task convolutional neural network can be used to locate facial regions frame by frame in the motion image sequence to obtain the target user's facial image sequence; wherein, for each frame of depth motion image in the motion image sequence, the multi-task convolutional neural network can first use a candidate network to filter out facial candidate boxes in the depth motion image and filter out facial candidate boxes with a confidence score lower than 0.8, then use a refinement network to optimize the coordinates of the facial candidate boxes to remove non-facial regions; finally, the output network outputs the facial region. A precise facial bounding box (containing 68 key points such as the eyes, nose tip, and corners of the mouth) is generated and cropped into a 150×150 pixel facial image. This yields the facial image for each frame of the motion image sequence. All facial images are then arranged chronologically to obtain a facial image sequence. This facial image sequence displays the changes in facial regions when the target user interacts with the bionic animal toy. This sequence provides continuous and consistent visual input for micro-expression analysis, eliminating interference from image scale and angle differences on feature extraction and ensuring the stability of micro-expression dynamic analysis.
[0036] In specific implementation, extracting the dynamic features of the target user's facial micro-expressions from the facial image sequence can be achieved in the following way: First, 468 three-dimensional facial key points can be extracted from the facial image sequence using an existing facial mesh model (such as the FaceMesh model), and the displacement changes of the facial key points within 10 consecutive frames (such as the opening and closing of the corners of the eyes and the upward movement of the corners of the mouth) can be calculated; then, the pixel motion vectors of adjacent frames in the facial image sequence are calculated using the optical flow method to generate a dynamic optical flow map to capture the instantaneous movement of micro-expressions; finally, the displacement changes of all facial key points are calculated. The dynamic optical flow map is input into a spatiotemporal convolutional network, where spatiotemporal information is fused through multi-scale convolutional layers (e.g., 3×3×3 kernels) to output dynamic features of the target user's facial micro-expressions. These dynamic features are a set of quantified features reflecting the movement of facial key points and pixel displacements when the target user interacts with a bionic animal toy. They include the evolutionary patterns of subtle facial movements (such as blinking, frowning, and smiling). These dynamic features can capture the spatiotemporal changes of the target user's subtle expressions, accurately characterize the dynamic attributes of micro-expressions, provide a core basis for emotion value calculation, and improve the sensitivity of emotion analysis.
[0037] In specific implementation, determining the three-dimensional emotional value of each instantaneous expression of the target user based on the dynamic features can be achieved in the following way: a pre-trained emotion regression network can be used to calculate the three-dimensional emotional value of each instantaneous expression of the target user based on the dynamic features; wherein, after the dynamic features are mapped by two fully connected layers (e.g., with 256 or 128 hidden units) in the emotion regression network, the values of each instantaneous expression in the dynamic features in the three dimensions of pleasure, arousal, and dominance are output through an activation function (e.g., sigmoid activation function). The values range from 0 to 1, where 0 represents the lowest intensity and 1 represents the highest intensity, thereby quantifying the emotion of each instantaneous expression (e.g., "pleasure 0.8, arousal 0.6, dominance 0.5" can represent a state of positive excitement); the three-dimensional emotional value is a numerical combination representing the target user's facial micro-expressions in terms of pleasure, arousal, and dominance. This three-dimensional emotional value can transform instantaneous expressions into quantifiable emotional indicators, realize the objective numerical expression of emotion, and provide a unified measurement standard for the correlation analysis of emotion and behavioral intention.
[0038] It should be noted that, in this application, the emotion feature tensor is a structured data representing the dynamic changes in emotion over time and facial region when the target user interacts with the bionic animal toy. This emotion feature tensor can provide comprehensive data support for the fusion analysis of behavioral intention vector and emotion, improving the richness and accuracy of intention reasoning. In specific implementation, the spatiotemporal expansion of all three-dimensional emotion values and the construction of the emotion feature tensor when the target user interacts with the bionic animal toy can be achieved in the following way: For all three-dimensional emotion values, the three-dimensional emotion values of each facial micro-expression are retained in the time frame order in the time dimension, and the three-dimensional emotion values are associated with the facial key point coordinates of the corresponding frame in the spatial dimension to form a three-dimensional matrix of "time-space-emotion". Then, a four-dimensional tensor is constructed by adding the emotion change rate (i.e., the difference between the current frame and the previous 3 frames) channel dimension on the basis of the three-dimensional matrix through the tensor expansion algorithm, and this four-dimensional tensor is used as the emotion feature tensor when the target user interacts with the bionic animal toy.
[0039] In some embodiments, constructing a spatiotemporal correlation graph of the target user's emotions and behaviors during interaction based on the emotional feature tensor and the behavioral intention vector can be achieved through the following steps: The emotional feature tensor and the behavioral intention vector are time-aligned to obtain time-synchronized emotional-behavioral data pairs; A spatiotemporal correlation matrix is constructed using the emotional state and behavioral intention in the aforementioned emotion-behavior data pairs as nodes; Based on the spatiotemporal correlation matrix, a spatiotemporal correlation map of the target user's emotions and behaviors during the interaction process is determined.
[0040] In specific implementation, the temporal alignment of the emotional feature tensor and the behavioral intention vector to obtain temporally synchronized emotional-behavioral data pairs can be achieved in the following way: A dynamic time warping algorithm can be used to temporally align the emotional feature tensor and the behavioral intention vector to obtain temporally synchronized emotional-behavioral data pairs. Specifically, a distance matrix can be constructed by calculating the timestamp difference between the emotional feature tensor and the behavioral intention vector. The elements in the distance matrix represent the Euclidean distance between the emotional feature tensor and the behavioral intention vector at the corresponding time. The alignment path with the smallest cumulative distance is then found in this distance matrix to interpolate the frequency of the behavioral intention vector to the emotional feature tensor, outputting temporally synchronized emotional-behavioral data pairs. The emotional-behavioral data pairs are structured data units that correspond one-to-one between the target user's emotional features and behavioral intentions in the time dimension. They contain fragments of the emotional feature tensor and the behavioral intention vector at the corresponding time. This emotional-behavioral data pair can eliminate the temporal asynchrony deviation of data from different modalities, providing time-synchronized basic data for subsequent correlation analysis, ensuring the accuracy of the spatiotemporal correspondence between emotions and behaviors, and avoiding analysis errors caused by differences in sampling frequencies.
[0041] In specific implementation, the spatiotemporal correlation matrix constructed using the emotional state and behavioral intention in the emotion-behavior data pair as nodes can be achieved in the following way: First, emotional state nodes and behavioral intention nodes can be constructed based on the emotion-behavior data pair. The emotional state node is each facial micro-expression and its corresponding three-dimensional emotional value in the emotional feature tensor, and the behavioral intention node is the interaction intention of the target user and its probability value in the behavioral intention vector. Then, the number of co-occurrences of each emotional node and behavioral node within a sliding time window (e.g., a window size of 5 seconds) is counted, and the correlation strength value (i.e., co-occurrence probability = number of co-occurrences / total number of data pairs within the window) is calculated to construct the spatiotemporal correlation matrix. The correlation matrix consists of matrix elements representing the correlation strength values between corresponding emotional state nodes and behavioral intention nodes, ranging from 0 to 1. Higher correlation strength values indicate a stronger correlation between emotional state nodes and behavioral intention nodes within that time window. Specifically, the spatiotemporal correlation matrix quantifies the correlation strength between emotions and behaviors when a target user interacts with a bionic animal toy. This spatiotemporal correlation matrix reveals high-frequency emotion-behavior combinations during the interaction process through numerical correlation strength, providing calculable correlation weights for graph construction and helping the model quickly identify typical emotion-behavior patterns during user interaction.
[0042] In specific implementation, determining the spatiotemporal correlation graph of the target user's emotions and behaviors during the interaction process based on the spatiotemporal correlation matrix can be achieved in the following way: Existing graph structure transformation algorithms (such as graph neural network graph construction algorithms based on adjacency matrices) can be used to construct the spatiotemporal correlation graph of the target user's emotions and behaviors during the interaction process based on the spatiotemporal correlation matrix. Specifically, the emotional state nodes and behavioral intention nodes in the spatiotemporal correlation matrix can be used as vertices of the graph, and the correlation strength can be used as the weight of the edges between vertices. Edges with a strength higher than the weight threshold are retained according to a preset weight threshold (such as 0.3) to simplify the graph. Then, the correlation graphs of different windows are connected in chronological order through time axis expansion to form a dynamic graph containing the time dimension and output as a spatiotemporal correlation graph. The weight of the edges can be dynamically updated with the time window.
[0043] It should be noted that, in this application, the spatiotemporal correlation graph is a graph structure model that displays the dynamic correlation strength between emotions and behaviors when a target user interacts with a bionic animal toy. This spatiotemporal correlation graph captures the dynamic coupling pattern of the target user's emotions and behaviors through graph structure, and can intuitively show the real-time correlation pattern of emotions and behaviors (such as the change of the strong correlation between "high pleasure" and "play intention" over time), reflecting the spatiotemporal dependence of emotions and behaviors during the interaction between the target user and the bionic animal toy. This provides multi-dimensional correlation basis for the bionic toy to generate response actions that conform to the user's emotional state, improves the naturalness and intelligence of the interaction, and provides correlation basis for the subsequent generation of personalized responses.
[0044] In step 104, the state characteristics of the bionic animal toy are determined based on the historical interaction data between the target user and the bionic animal toy. Then, a sequence of motion primitives that conforms to the state of the bionic animal toy is generated based on the spatiotemporal correlation map and the state characteristics.
[0045] In some embodiments, determining the personified state characteristics of the bionic animal toy based on historical interaction data between the target user and the bionic animal toy can be achieved through the following steps: Acquire historical interaction data between target users and bionic animal toys; The historical interaction data is cleaned and segmented into data fragments to generate time-series interaction samples with emotion tags; The interaction features in the time-series interaction samples are extracted to construct a five-dimensional personality feature vector for the bionic toy; The state characteristics of the bionic animal toy are determined based on the five-dimensional personality feature vector.
[0046] In practice, obtaining historical interaction data between the target user and the bionic animal toy can be achieved in the following way: the historical interaction data between the target user and the bionic animal toy can be obtained from the bionic interaction log. This historical interaction data can be collected through the sensors (accelerometers, gyroscopes) and vision system (such as binocular cameras) built into the bionic animal toy. The historical interaction data is a multimodal historical record set containing the action type, interaction duration and emotional feedback information of the target user when interacting with the bionic animal toy. By accumulating long-term interaction information, this historical interaction data can capture the user's personalized interaction habits and provide rich historical evidence for constructing personality traits that conform to the user's preferences.
[0047] In specific implementation, cleaning and segmenting the historical interaction data to generate time-series interaction samples with emotion tags can be achieved in the following way: First, the historical interaction data can be cleaned by removing fragmented records with interaction durations of less than 1 second and normalizing the numerical features (such as interaction frequency) in the historical interaction data using existing standardization methods (such as Z-score standardization). After inclusion, the historical interaction data can be segmented using a sliding window method (the window size can be set to 20 interaction cycles and the step size to 10 interaction cycles) to obtain multiple data fragments. The system processes data segments within each window and filters user sentiment tags using a majority voting method to generate temporal interaction samples (e.g., [action sequence, response sequence, sentiment tag]) with dominant sentiment tags (e.g., positive, neutral, negative). These temporal interaction samples are interaction segments marked with dominant sentiment tags from the target user's historical interaction records with the bionic animal toy. This process transforms disordered data into analyzable structured units, preserving the temporal correlation between interactive actions and emotions, providing standardized input for extracting dynamic personality characteristics, and improving the spatiotemporal coherence of feature analysis.
[0048] In specific implementation, the extraction of interaction features from the temporal interaction samples to construct the five-dimensional personality feature vector of the bionic toy can be achieved in the following way: First, five core interaction features can be extracted from the temporal interaction samples using existing feature extraction methods (such as a temporal feature encoder based on an attention mechanism), including: Agility: the average switching speed of response actions per unit time, reflecting the frequency of the bionic animal toy's response to user actions; Affinity and Interaction: the proportion of interaction time under positive emotion tags, reflecting the emotional compatibility between the bionic animal toy and the user; Emotional Resonance: the matching degree between the user's emotional value and the emotional response of the bionic animal toy (which can be calculated by cosine similarity to determine the matching degree between the three-dimensional emotional value and the emotional parameters of the action); Exploration and Curiosity: the trigger probability of rare actions (i.e., interaction actions with a frequency of less than 5%). The system reflects the tendency of bionic animal toys to actively initiate new interactions; the "laziness and relaxation" is the average duration of the interval between single interactions, reflecting the "passive waiting" personality trait of bionic animal toys; then, through principal component analysis to reduce dimensionality and remove feature redundancy, a five-dimensional personality feature vector is constructed; wherein, the five-dimensional personality feature vector is a numerical vector that quantifies the personality traits of bionic animal toys in five dimensions: agility, affinity and interaction, emotional resonance, exploratory curiosity, and laziness and relaxation. The value range of each dimension is 0-1, and the higher the value, the more obvious the personality trait of the corresponding dimension. This five-dimensional personality feature vector transforms the user interaction pattern into a calculable personality dimension, and through multi-dimensional features, it portrays the "personality" differences of bionic animal toys, making the personalized state interpretable and adjustable, and providing a parameter basis for generating personalized responses.
[0049] It should be noted that, in this application, the state feature is a probability distribution vector reflecting the current dominant personality state (such as lively or gentle) of the bionic animal toy. This state feature can dynamically characterize the "behavioral personality" of the bionic animal toy, guiding the toy to adjust its response strategy according to the user's real-time interaction mode (such as increasing proactive interactive actions when the activity level is high), thereby improving the consistency and emotional fit of the interactive experience. Specifically, determining the personified state features of the bionic animal toy based on the five-dimensional personality feature vector can be achieved in the following way: the five-dimensional personality feature vector can be input into a pre-trained Hidden Markov Model to infer personality states, generating the personified state features of the bionic animal toy; wherein, the Hidden Markov Model includes... It contains three hidden states (e.g., "lively", "mild", and "lazy"), and the output layer corresponds to the probability distribution of 5-dimensional features. It can output the probability value of each hidden state (e.g., Pdynamic feature (lively) = 0.6, Pdynamic feature (mild) = 0.3, Pdynamic feature (lazy) = 0.1) and use the vector composed of all hidden states and their corresponding probability values as personalized state features. When training the Hidden Markov Model, the Baum-Welch algorithm can be used to optimize the state transition matrix and the observation probability matrix, and the input is the interaction feature sequence of the target user in the past 30 days for training to dynamically adjust the response strategy of the bionic animal toy (e.g., increase the trigger probability of rapid swinging action in the "lively" state).
[0050] In some embodiments, generating a sequence of motion primitives that conforms to the personified state of a biomimetic animal toy based on the spatiotemporal correlation map and the state features can be achieved by the following steps: Extract the sentiment-behavior coupling matrix of the target user from the spatiotemporal correlation graph; Based on the state features and the emotion-behavior coupling matrix, the angular oscillation trajectory of each joint of the bionic animal toy is generated; Inverse kinematics solutions are performed on all angular oscillation trajectories to generate a sequence of motion primitives that conform to the personified state of the bionic animal toy.
[0051] In specific implementation, the extraction of the target user's emotion-behavior coupling matrix from the spatiotemporal correlation graph can be achieved in the following way: the target user's emotion-behavior coupling weight matrix can be extracted from the spatiotemporal correlation graph through a graph attention mechanism. Specifically, a multi-head attention mechanism is used to calculate the attention scores (i.e., the degree of correlation between emotional state and behavioral intention) between each node in the spatiotemporal correlation graph, strengthen the weight of high-frequency co-occurrence pairs (such as "high pleasure - play intention"), and suppress low-frequency irrelevant correlations, thereby outputting a 128-dimensional emotion-behavior coupling matrix. The emotion-behavior weight matrix is a numerical matrix that quantifies the real-time correlation strength between the target user's emotional state and behavioral intention when interacting with the bionic animal toy. This emotion-behavior weight matrix can strengthen high-frequency co-occurrence emotion-behavior pairs and suppress irrelevant correlations, providing key driving parameters for the bionic animal toy's biological central pattern generator, ensuring that the toy's response focuses on the user's currently dominant interaction mode, and improving the coupling accuracy of emotion and behavior.
[0052] In specific implementation, the generation of angular oscillation trajectories of each joint of the bionic animal toy based on the state features and the emotion-behavior coupling matrix can be achieved in the following way: the state features and the emotion-behavior coupling weight matrix can be input into an improved Hopf oscillator network for dynamic parameter modulation and phase synchronization optimization to generate angular oscillation trajectories of each joint of the bionic animal toy; wherein, each joint of the bionic animal toy corresponds to a Hopf oscillator, and the improved Hopf oscillator network can optimize the Hopf parameters based on the state features and the emotion-behavior coupling weight matrix through the contained dynamic equations and adaptive gradient descent algorithm, and output the angular oscillation trajectory of each joint containing phase difference information and a trajectory time resolution of 10 milliseconds per frame (e.g., the tail joint leads the head joint by 30° to achieve the "excited tail wagging" action); the angular oscillation trajectory is a curve of the joint angle of the bionic animal toy changing with time. This angular oscillation trajectory can simulate the rhythmic oscillation characteristics of biological movement. By introducing personalized parameters (such as agility, laziness, etc.) to modulate the oscillation frequency and amplitude, the joint movement naturally presents personalized characteristics ("lively", "lazy", etc.), enhancing the bio-realism of the movement.
[0053] In practical implementation, the inverse kinematics solution for all angular oscillation trajectories to generate a sequence of motion primitives conforming to the personified state of the bionic animal toy can be achieved in the following way: An inverse kinematics algorithm based on geometry (such as the Piper algorithm) can be used to solve the inverse kinematics solution for the angular oscillation trajectories of all joints of the bionic animal toy to obtain the end effector coordinates, generating a sequence of motion primitives conforming to the personified state of the bionic animal toy. Specifically, for the limb joint chains of the bionic animal toy, the inverse kinematics algorithm can define the base coordinate system and the end effector coordinate system using the Danavitt-Hartenberg parameter table. The joint transformation matrix is used to inversely deduce the joint angle combination that satisfies the end position and posture. Cubic spline interpolation is introduced to smooth the joint angle sequence after inverse solution. All angle oscillation trajectories are mapped to various basic motion primitives (such as affectionate head rubbing, excited tail wagging, curious paw probing, etc.). Each basic motion primitive contains joint movement amplitude, speed, and timing synchronization parameters (such as "M2 tail wagging action" corresponds to tail joint angle ±30° oscillation, frequency 2Hz, forming a 0.5-second phase difference with the head joint). Finally, a coherent motion primitive sequence that conforms to the personification state of the bionic animal toy is generated.
[0054] It should be noted that in this application, the motion primitive sequence is a continuous sequence of basic action units (such as affectionate head rubbing and excited tail wagging) that are converted from the joint angle trajectory of the bionic animal toy. This motion primitive sequence generates action combinations that conform to the user's real-time emotions and intentions by integrating emotion-behavior association weights and personalized state features, thereby realizing the personalized and emotional expression of the bionic toy's response actions and improving the naturalness and smoothness of the interactive experience.
[0055] In step 105, the motion primitive sequence is converted into a servo control command stream to control the bionic animal toy to perform interactive actions with emotional feedback.
[0056] In some embodiments, converting the motion primitive sequence into a servo control command stream to control the bionic animal toy to perform interactive actions with emotional feedback can be achieved through the following steps: The motion primitive sequence is decomposed into joint-level motion parameters, and then transformed into a servo control command stream for the end effector. The end effector of the bionic animal toy is controlled by the servo control command stream to perform interactive actions with emotional feedback.
[0057] It should be noted that, in this application, the servo control command stream refers to a continuous set of commands including servo identifier, timestamp, and pulse width modulation parameters. This servo control command stream can precisely drive the joints of the bionic animal toy to move along a preset trajectory. This solution achieves multi-servo collaborative control through standardized command format and dynamically adjusts parameters in conjunction with emotional feedback to ensure that the toy's actions not only conform to the personified state but also match the user's interaction needs in real time, thereby improving the accuracy of action execution and emotional fit.
[0058] In specific implementation, the decomposition of the motion primitive sequence into joint-level motion parameters, and then the conversion into a servo control command stream for the end effector, can be achieved in the following way: the mapping relationship between the angular velocity of each joint in the motion primitive sequence and the velocity of the end effector can be calculated by using the Jacobian pseudo-inverse matrix, the end trajectory can be decomposed into 12 servo angle changes (where the angle change range is 0°-180°), and the angle parameters can be dynamically adjusted by combining the emotional intensity coefficient (e.g., the amplitude gain is 1.2 for happiness and 0.7 for sadness). Then, all angle changes are converted into a pulse width modulation command stream and the pulse width value of the intermediate angle is calculated by linear interpolation, thereby generating a servo control command stream with a timestamp.
[0059] In specific implementation, the end effector of the bionic animal toy is controlled to perform interactive actions with emotional feedback according to the servo control command stream. This can be achieved in the following way: a multi-threaded scheduler can be used to synchronously send the commands in the servo control command stream according to the timestamp, and the commands can be transmitted in real time through the controller local area network bus to drive the end effector (such as the head, tail, and limb joints) to perform actions.
[0060] In another aspect, in some embodiments, this application provides an intelligent control device for a biomimetic animal toy, see reference. Figure 4 The figure is a schematic diagram of the structure of an intelligent control device for a bionic animal toy according to some embodiments of this application. The intelligent control device 400 for the bionic animal toy includes: a data acquisition module 401, a processing module 402, and an execution module 403, which are described below: The acquisition module 401 in this application is mainly used to acquire the action image sequence of the target user when interacting with the bionic animal toy; Processing module 402, in this application, is mainly used to extract the three-dimensional spatial trajectory of the target user's hand joints in the motion image sequence, analyze the target user's behavioral semantics based on the three-dimensional spatial trajectory, and then generate the target user's behavioral intention vector for the bionic animal toy. It should be noted that the processing module 402 in this application is also used to simultaneously identify facial micro-expressions in the action image sequence, obtain the emotional feature tensor when the target user interacts with the bionic animal toy, and then construct a spatiotemporal correlation map of the target user's emotions and behaviors during the interaction based on the emotional feature tensor and the behavioral intention vector. Additionally, it should be noted that the processing module 402 in this application is also used to determine the personification state characteristics of the bionic animal toy based on the historical interaction data between the target user and the bionic animal toy, and then generate a motion element sequence that conforms to the personification state of the bionic animal toy based on the spatiotemporal correlation map and the state characteristics. The execution module 403 in this application is mainly used to convert the motion primitive sequence into a servo control command stream to control the bionic animal toy to perform interactive actions with emotional feedback.
[0061] The various modules in the intelligent control device of the aforementioned biomimetic animal toys can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0062] In another embodiment, this application provides a computer device, which may be a server, and its internal structure diagram may be as follows. Figure 5 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores data on intelligent control methods for bionic animal toys. The network interface communicates with external terminals via a network connection. When the processor executes the computer program, it implements an intelligent control method for a bionic animal toy.
[0063] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0064] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described embodiment of the intelligent control method for biomimetic animal toys.
[0065] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described embodiment of the intelligent control method for biomimetic animal toys.
[0066] In one embodiment, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps described in the embodiment of the intelligent control method for the biomimetic animal toy.
[0067] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0068] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0069] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for intelligent control of a biomimetic animal toy, wherein the method controls the biomimetic animal toy to provide emotional feedback interaction based on the target user's actions and facial micro-expressions in a user interaction scenario, characterized in that, The method includes the following steps: Collect motion image sequences of target users interacting with bionic animal toys; The three-dimensional spatial trajectory of the target user's hand joints is extracted from the motion image sequence. The behavioral semantics of the target user are analyzed based on the three-dimensional spatial trajectory, and then a behavioral intention vector of the target user towards the bionic animal toy is generated. Simultaneously, facial micro-expressions in the action image sequence are identified to obtain the emotional feature tensor when the target user interacts with the bionic animal toy. Then, based on the emotional feature tensor and the behavioral intention vector, a spatiotemporal correlation map of the target user's emotions and behaviors during the interaction is constructed. Based on the historical interaction data between the target user and the bionic animal toy, the state characteristics of the bionic animal toy's personification are determined, and then a sequence of motion primitives that conforms to the personification state of the bionic animal toy is generated based on the spatiotemporal correlation map and the state characteristics. The motion primitive sequence is converted into a servo control command stream to control the bionic animal toy to perform interactive actions with emotional feedback; Specifically, constructing a spatiotemporal correlation graph of the target user's emotions and behaviors during the interaction process based on the emotional feature tensor and the behavioral intention vector includes: The emotional feature tensor and the behavioral intention vector are time-aligned to obtain time-synchronized emotional-behavioral data pairs; A spatiotemporal correlation matrix is constructed using the emotional state and behavioral intention in the aforementioned emotion-behavior data pairs as nodes; Based on the spatiotemporal correlation matrix, a spatiotemporal correlation map of the target user's emotions and behaviors during the interaction process is determined; Specifically, determining the personification characteristics of bionic animal toys based on historical interaction data between target users and the toys includes: Acquire historical interaction data between target users and bionic animal toys; The historical interaction data is cleaned and segmented into data fragments to generate time-series interaction samples with emotion tags; The interaction features in the time-series interaction samples are extracted to construct a five-dimensional personality feature vector for the bionic toy; The state characteristics of the personified bionic animal toy are determined based on the five-dimensional personality feature vector. Specifically, generating a sequence of motion primitives that conforms to the personified state of a biomimetic animal toy based on the spatiotemporal correlation map and the state features includes: Extract the sentiment-behavior coupling matrix of the target user from the spatiotemporal correlation graph; Based on the state features and the emotion-behavior coupling matrix, the angular oscillation trajectory of each joint of the bionic animal toy is generated; Inverse kinematics solutions are performed on all angular oscillation trajectories to generate a sequence of motion primitives that conform to the personified state of the bionic animal toy.
2. The method as described in claim 1, characterized in that, Extracting the three-dimensional spatial trajectory of the target user's hand joints from the motion image sequence specifically includes: Hand instance segmentation is performed on each depth motion image in the motion image sequence to obtain the hand region mask for each depth motion image; By aligning all hand region masks to locate the pixel coordinates of hand joints, a three-dimensional coordinate set of the target user's hand joints is obtained; By performing occlusion point compensation on the three-dimensional coordinate set, an anti-disturbance joint point sequence of the target user's hand joint points is obtained; The three-dimensional spatial trajectory of the target user's hand joints is generated using the anti-disturbance joint sequence.
3. The method as described in claim 1, characterized in that, Based on the analysis of the target user's behavioral semantics using the three-dimensional spatial trajectory, a behavioral intent vector of the target user towards the bionic animal toy is generated, specifically including: Extract the motion features of the target user's hand joints from the three-dimensional spatial trajectory; Based on the motion characteristics, the three-dimensional spatial trajectory is divided into multiple discrete spatiotemporal action primitives; Based on a pre-established biomimetic interactive dictionary, the semantic strength of each spatiotemporal action primitive is quantified, thereby generating a semantic encoding vector of the target user's hand actions. The target user's behavioral intent vector for the bionic animal toy is determined based on the semantic encoding vector.
4. The method as described in claim 1, characterized in that, Simultaneously recognizing facial micro-expressions in the motion image sequence to obtain the emotional feature tensor of the target user interacting with the bionic animal toy specifically includes: Facial region localization is performed on the motion image sequence to obtain the facial image sequence of the target user; Dynamic features of the target user's facial micro-expressions are extracted from the facial image sequence; The three-dimensional emotion value of each instantaneous expression of the target user is determined based on the dynamic characteristics; Spatiotemporal expansion of all three-dimensional emotion values is performed to construct an emotion feature tensor for the target user when interacting with the bionic animal toy.
5. An intelligent control device for a biomimetic animal toy, comprising intelligent control of the biomimetic animal toy using the method described in any one of claims 1 to 4, characterized in that, The intelligent control device includes: The acquisition module is used to acquire motion image sequences of target users interacting with bionic animal toys; The processing module is used to extract the three-dimensional spatial trajectory of the target user's hand joints from the motion image sequence, analyze the target user's behavioral semantics based on the three-dimensional spatial trajectory, and then generate a behavioral intention vector of the target user towards the bionic animal toy. The processing module is used to simultaneously recognize facial micro-expressions in the action image sequence, obtain the emotional feature tensor when the target user interacts with the bionic animal toy, and then construct a spatiotemporal correlation map of the target user's emotions and behaviors during the interaction based on the emotional feature tensor and the behavioral intention vector. The processing module is used to determine the personified state characteristics of the bionic animal toy based on the historical interaction data between the target user and the bionic animal toy, and then generate a motion element sequence that conforms to the personified state of the bionic animal toy based on the spatiotemporal correlation map and the state characteristics. The execution module is used to convert the motion primitive sequence into a servo control command stream to control the bionic animal toy to perform interactive actions with emotional feedback.
6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the intelligent control method for the bionic animal toy as described in any one of claims 1 to 4.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the intelligent control method for the bionic animal toy as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Hybrid gesture tracking method, system, equipment and medium
CN119088204A
Toy interaction control method and device based on large language model, terminal and medium
CN119905092A