Intelligent pet robot system and control method thereof

Through multimodal perception and dynamic personality update technology, pet robots are able to accurately identify users' emotions and behaviors, solve the problems of single interaction methods and lack of adaptability, and provide personalized and emotional resonant interactive experiences, suitable for fields such as family companionship, children's education and elderly care.

CN120516705APending Publication Date: 2025-08-22PANOVASIC TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510870287.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

Existing pet robots lack multimodal perception ability, cannot accurately identify user emotions and behaviors, and lack adaptability, resulting in a single interaction mode, a lack of emotional resonance and personalized interaction.

Method used

Multimodal perception module, personality model and dynamic update module, data processing and emotion recognition module, interaction strategy generation and execution module, and data storage and communication module are adopted, combined with deep learning algorithms and reinforcement learning algorithms to realize multimodal perception and dynamic personality update.

Benefits of technology

It enhances the naturalness and personalization of interaction. The robot can understand user emotions and needs, provide personalized interactive experience, improve emotional resonance and companionship effects, and is suitable for areas such as family companionship, children's education and elderly care.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120516705A_ABST
    Figure CN120516705A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent pet robot system and a control method thereof, multi-modal sensing and dynamic character updating are realized, the system comprises a multi-modal sensing module, a character model and dynamic updating module, a data processing and emotion recognition module, an interaction strategy generation and execution module and a data storage and communication module which are connected in sequence; according to the method, the naturalness and individuation of interaction can be enhanced, the character dynamic updating is realized, the emotional resonance and accompanying effect is improved, the application scene is wide, and good social benefits and market prospects are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and robotics technology, and in particular to an intelligent pet robot system and a control method thereof. Background Art

[0002] With the development of society and the advancement of artificial intelligence technology, intelligent pet robots have gradually entered people's lives as new emotional companionship devices. However, existing pet robots have the following main shortcomings:

[0003] 1. Single interaction method: Most pet robots interact with users through preset programs and simple perception methods. They lack a deep understanding of users' emotions and behaviors, resulting in a mechanical and monotonous interaction process.

[0004] 2. Lack of emotional resonance: Unable to accurately identify the user's emotional state, unable to adjust its own behavior and expression according to the user's emotional changes, and difficult to form emotional resonance with the user.

[0005] 3. Lack of adaptability: Lack of learning ability, unable to dynamically adjust its own characteristics and interaction methods based on users' long-term usage behavior and preferences, making it difficult to meet the diverse and changing needs of users.

[0006] Therefore, how to give pet robots multimodal perception capabilities, accurately identify user emotions and behaviors, and dynamically update their own personality models to adapt to different users and provide personalized and humanized interactive experiences has become an urgent problem to be solved. Summary of the Invention

[0007] The purpose of the present invention is to provide an intelligent pet robot system and a control method thereof that realize multimodal perception and dynamic personality updating. It aims to use multimodal perception and dynamic personality updating technology to solve the technical problems that existing pet robots cannot accurately recognize user emotions, lack personalized interaction, and cannot adapt to user needs.

[0008] The present invention solves the above problems through the following technical solutions:

[0009] An intelligent pet robot system realizes multimodal perception and dynamic personality updating. The system includes: a multimodal perception module, a personality model and dynamic update module, a data processing and emotion recognition module, an interaction strategy generation and execution module, and a data storage and communication module, which are connected in sequence; wherein: the multimodal perception module collects the user's visual, audio, tactile information and environmental data in real time; the data processing and emotion recognition module preprocesses, extracts features and recognizes emotions for the information collected by the multimodal perception module; the personality model and dynamic update module dynamically updates the robot's personality model based on the user's emotional state and behavioral characteristics recognized by the data processing and emotion recognition module, as well as historical interaction data; the interaction strategy generation and execution module generates personalized interaction strategies according to the robot's personality model and controls the robot to perform corresponding behaviors; the data storage and communication module is used to manage system operation data, user information, and model parameters, and communicate with external devices.

[0010] As a further improvement, the multimodal perception module includes: a visual perception unit, an audio perception unit and a tactile perception unit; the visual perception unit is equipped with a high-definition camera and a depth sensor to capture the user's facial expressions, body movements and surrounding environment information; the audio perception unit has a built-in omnidirectional microphone array to collect the user's voice information for voice recognition and emotion analysis; the tactile perception unit arranges pressure sensors and touch sensors on the outside of the robot to sense the user's touch position, touch method, strength and frequency.

[0011] As a further improvement, the high-definition camera is used to capture the user's facial expressions, eyes, and posture information; the depth sensor is used to obtain the user's movements, gestures and the distance between the user and the robot.

[0012] As a further improvement, the omnidirectional microphone array is used to collect audio information of the user's voice, angle, pitch, speaking speed and volume.

[0013] As a further improvement, the data processing and emotion recognition module includes: a data preprocessing unit, a feature extraction unit, a multimodal feature fusion unit and an emotion recognition and behavior analysis unit; the data preprocessing unit: filters, reduces noise, and normalizes the collected multimodal data; the feature extraction unit: uses a deep learning algorithm to extract key features of visual, audio, and tactile data; the multimodal feature fusion unit: fuses visual, audio, and tactile features; the emotion recognition and behavior analysis unit: combines multimodal features to identify the user's emotional state and behavioral intentions.

[0014] As a further improvement, the feature extraction unit uses a deep learning model to implement: (1) visual feature extraction: using a convolutional neural network (CNN) to extract image features; (2) audio feature extraction: using a recurrent neural network (RNN) or a long short-term memory network (LSTM) to extract speech features; (3) tactile feature extraction: extracting touch force and frequency features.

[0015] As a further improvement, the personality model and dynamic update module include: a personality basic model library, a personality update engine and a personality memory unit; the personality basic model library: establishes a variety of pet personality models, covering different personality characteristics; the personality update engine: a machine learning algorithm based on the reinforcement learning algorithm generative adversarial network (GAN), adjusts the personality model parameters according to interaction data and user feedback; the personality memory unit: stores historical interaction data and personality change records, supporting long-term learning and personalized evolution.

[0016] As a further improvement, the personality update engine uses a deep reinforcement learning (DQN) reinforcement learning algorithm. The inputs are the current user's emotional state and behavioral intentions, as well as the user's historical interaction data and feedback; the output is the updated personality model parameters. The workflow is as follows: (1) Emotion recognition results and user feedback are input into the personality update engine; (2) The personality update engine calculates the optimal personality parameter adjustment strategy based on the current context, historical data, and the personality basic model library; (3) The updated personality model is stored in the personality memory unit for subsequent interaction strategy generation.

[0017] As a further improvement, the interaction strategy generation and execution module includes: an interaction strategy generation unit, an action control unit and a feedback collection unit; the interaction strategy generation unit is connected to the personality memory unit, and a user feedback unit is connected between the personality memory unit and the feedback collection unit. The personality memory unit generates relevant interaction strategies, and the interaction strategies generate relevant control unit commands.

[0018] At the same time, the present invention also solves the above problems through the following technical solutions:

[0019] A control method for an intelligent pet robot system, for controlling the intelligent pet robot system as described above, comprises the following steps:

[0020] Step A: Multimodal Information Acquisition: The robot uses a multimodal perception module to acquire the user's visual, audio, and tactile information and environmental data in real time. For the visual modality, a pre-trained image emotion recognition model and a CNN architecture with a ResNet50 + SE mechanism are used to extract the texture and posture features of facial expressions. For the audio modality, a BiLSTM-based emotional speech recognition model is used to model MFCC, pitch, and energy feature sequences to capture temporal emotional signals. For the tactile modality, the user's touch position, force, and frequency are convolved and extracted using a 1D-CNN to identify the user's interaction intent.

[0021] Step B: Data processing; To address the temporal and semantic differences between different modalities, the following fusion strategies are used: (1) Temporal alignment: A temporal alignment network based on an attention mechanism is used to align and weight the audio, visual, and tactile sequences to synchronize them in the semantic space; (2) Feature fusion structure: A multi-channel Transformer structure or MCB is used for deep fusion to ensure that each modality fully interacts in a unified feature space; (3) Modality weight adjustment: A modality attention gating mechanism is introduced to automatically adjust the weight distribution of different modalities according to the current interaction environment and historical behavior to enhance robustness;

[0022] Step C: Emotional state reasoning and adaptive learning: Fusion features are fed into a softmax classifier to output the current user's emotion label and confidence distribution. Emotional time series are constructed, using a sliding window mechanism to record historical emotion trends, providing input support for personality model evolution and interaction strategy generation.

[0023] Step D: Dynamic personality model update: The personality update engine adjusts the robot's personality model based on the user's identified emotional state, behavioral characteristics, and historical user interaction data. Personality updates consider both short-term interaction feedback and long-term evolution trends to ensure the stability and adaptability of the robot's personality.

[0024] Step E: Personalized interaction strategy generation: Based on the updated personality model, generate an interaction strategy that adapts to the current situation and user preferences;

[0025] Step F: Behavior execution and user feedback collection: The robot executes corresponding behaviors according to the interaction strategy and interacts with the user; monitors the user's reactions and comments, and records the interaction effects;

[0026] Step G: Continuous learning and optimization; accumulating interaction data through personality memory units;

[0027] Continuously use new data to optimize personality models and interaction strategies to enhance the interactive experience.

[0028] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0029] (1) Enhance the naturalness and personalization of interaction: Through multimodal perception and emotion recognition, robots can understand users' real emotions and needs and provide personalized interactive experiences.

[0030] (2) Dynamic personality update: The robot has the ability to self-learn and dynamically adjust its personality. It can continuously optimize its own behavior and enhance user stickiness based on the user's long-term usage habits and feedback.

[0031] (3) Improve emotional resonance and companionship: Robots can display rich emotions and personalities, establish deeper emotional connections with users, and meet users’ needs for emotional companionship.

[0032] (4) Wide range of application scenarios: Applicable to many fields such as family companionship, children's education, and elderly care, with good social benefits and market prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 This is the overall architecture diagram of the intelligent pet robot system of the present invention;

[0034] Figure 2 Schematic diagram of the composition structure of the multimodal perception module of the present invention;

[0035] Figure 3 This is a flow chart of the data processing and emotion recognition module of the present invention;

[0036] Figure 4 This is a working principle diagram of the personality model and dynamic update module of the present invention;

[0037] Figure 5 This is a functional diagram of the interaction strategy generation and execution module of the present invention. DETAILED DESCRIPTION

[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0039] Example 1:

[0040] Combined with attachment Figure 1-5As shown, an intelligent pet robot system that realizes multimodal perception and dynamic personality update includes: a multimodal perception module, a personality model and dynamic update module, a data processing and emotion recognition module, an interaction strategy generation and execution module, and a data storage and communication module connected in sequence; wherein:

[0041] Multimodal perception module: collects user's visual, audio, tactile information and environmental data in real time;

[0042] Data processing and emotion recognition module: preprocesses, extracts features and performs emotion recognition on the information collected by the multimodal perception module;

[0043] Personality model and dynamic update module: Dynamically updates the robot's personality model based on the user's emotional state and behavioral characteristics identified by the data processing and emotion recognition module, as well as historical interaction data;

[0044] Interaction strategy generation and execution module: Generates personalized interaction strategies based on the robot's personality model and controls the robot to perform corresponding behaviors;

[0045] Data storage and communication module: used to manage system operation data, user information, and model parameters, and communicate with external devices.

[0046] In an optional embodiment, the multimodal perception module includes: a visual perception unit, an audio perception unit, and a tactile perception unit;

[0047] Visual perception unit: equipped with high-definition cameras and depth sensors to capture the user's facial expressions, body movements and surrounding environment information.

[0048] In this embodiment, the high-definition camera is used to capture the user's facial expressions, eyes, posture and other information.

[0049] Depth sensor, used to obtain the user's movements, gestures, and distance to the robot.

[0050] Audio perception unit: Built-in omnidirectional microphone array to collect user voice information for speech recognition and emotion analysis.

[0051] In this embodiment, the omnidirectional microphone array is used to collect audio information such as the user's voice, angle, pitch, speaking speed and volume.

[0052] Tactile perception unit: Pressure sensors and touch sensors are placed on the outside of the robot to sense the user's touch position, touch method, strength and frequency.

[0053] In another optional embodiment, the data processing and emotion recognition module includes: a data preprocessing unit, a feature extraction unit, a multimodal feature fusion unit and an emotion recognition and behavior analysis unit;

[0054] Data preprocessing unit: Filter, reduce noise and normalize the collected multimodal data.

[0055] Specifically, the data preprocessing unit is used to: receive the raw data from the multimodal perception module; perform filtering, denoising, image enhancement and other processing on the visual data; perform denoising, echo removal, speech segmentation and other processing on the audio data; and perform signal filtering and normalization processing on the tactile data.

[0056] Feature extraction unit: Uses deep learning algorithms (such as CNN and RNN) to extract key features of visual, audio, and tactile data.

[0057] Specifically, the feature extraction unit uses a deep learning model:

[0058] (1) Visual feature extraction: Convolutional neural network (CNN) is used to extract image features, such as facial expression features and posture features.

[0059] (2) Audio feature extraction: Use recurrent neural networks (RNN) or long short-term memory networks (LSTM) to extract speech features such as pitch and emotional tone.

[0060] (3) Tactile feature extraction: extracting features such as touch force and frequency.

[0061] Multimodal feature fusion unit: fuses visual, audio, and tactile features, possibly using a multimodal deep learning model or feature fusion algorithm.

[0062] Emotion recognition and behavior analysis unit: Combining multimodal features and using emotional computing models, it identifies the user's emotional state (such as happiness, sadness, anger, etc.) and behavioral intentions.

[0063] Emotion and behavior recognition: Using the fused features, employing an emotional computing model or classification algorithm, we can identify the user’s emotional state (such as happiness, sadness, anger, and surprise) and behavioral intentions (such as interaction and comfort needs).

[0064] In another optional embodiment, the personality model and dynamic update module includes: a personality basic model library, a personality update engine, and a personality memory unit;

[0065] Basic personality model library: Build a variety of pet personality models, covering different personality traits (such as liveliness, docility, curiosity, etc.).

[0066] In this embodiment, the basic personality model library stores multiple preset personality models, such as lively, docile, curious, calm, etc. Each personality model contains different parameter settings to control the robot's behavior mode.

[0067] Personality Update Engine: Based on machine learning algorithms such as reinforcement learning and generative adversarial networks (GANs), personality model parameters are adjusted according to interaction data and user feedback.

[0068] In this embodiment, the personality update engine:

[0069] The core part uses reinforcement learning algorithms (such as deep reinforcement learning DQN) and generative adversarial networks (GAN).

[0070] Input: The current user's emotional state and behavioral intention (from the emotion recognition module) and the user's historical interaction data and feedback (from the personality memory unit).

[0071] Output: Updated personality model parameters.

[0072] Personality Memory Unit: Stores historical interaction data and personality change records to support long-term learning and personalized evolution.

[0073] In this embodiment, the personality memory unit stores the interaction history with the user, including the emotional state, behavior, and feedback results of each interaction, providing long-term data support for personality updates.

[0074] Workflow:

[0075] (1) Emotion recognition results and user feedback are input into the personality update engine.

[0076] (2) The personality update engine calculates the optimal personality parameter adjustment strategy based on the current situation, historical data and personality basic model library.

[0077] (3) The updated personality model is stored in the personality memory unit for use in subsequent interaction strategy generation.

[0078] In another optional embodiment, the interaction strategy generation and execution module includes: an interaction strategy generation unit, an action control unit, and a feedback collection unit; the interaction strategy generation unit is connected to the personality memory unit, and a user feedback unit is connected between the personality memory unit and the feedback collection unit. The personality memory unit generates relevant interaction strategies, and the interaction strategies generate relevant control unit commands;

[0079] Specifically, the interaction strategies include: action commands: such as approaching the user, jumping, wagging the tail, etc.

[0080] Voice commands: speaking content, tone, and speed.

[0081] Expression display: displayed eyes, expression lights, etc.

[0082] The relevant output units output according to the strategy

[0083] Then the feedback collection unit collects the user's feedback

[0084] Finally, the feedback is output to the personality memory unit to form a closed loop.

[0085] Figure 5 It reflects the cyclical process from strategy generation to execution, feedback collection and further updating.

[0086] Interaction strategy generation unit: Generates an interaction strategy that adapts to the current situation and user preferences based on the current user's emotional state, behavioral characteristics and updated personality model.

[0087] Action control unit: converts interaction strategies into specific action instructions to control the robot's movement, voice, expression and other outputs.

[0088] Feedback collection unit: monitors user reactions to the robot, collects evaluation information, and provides a basis for personality updates.

[0089] In another optional embodiment, the data storage and communication module includes: a data storage unit and a network communication unit;

[0090] Data storage unit: manages the storage of system operation data, user information and model parameters.

[0091] Network communication unit: supports communication between the robot and cloud servers or other devices to achieve data synchronization and remote updates.

[0092] Example 2:

[0093] A control method for an intelligent pet robot system that implements multimodal perception and dynamic personality updating, comprising the following steps:

[0094] Step 1: Multimodal information collection;

[0095] The robot uses a multimodal perception module to obtain the user's visual, audio, tactile information and environmental data in real time.

[0096] Visual modality: Use pre-trained image emotion recognition models, such as the ResNet50 + SE (Squeeze-and-Excitation) CNN architecture, to extract texture and posture features of facial expressions.

[0097] Audio modality: Use a BiLSTM (bidirectional long short-term memory)-based emotional speech recognition model to model feature sequences such as MFCC (Mel-frequency cepstral coefficients), pitch, and energy to capture temporal emotional signals.

[0098] Tactile modality: The user's touch position, force, frequency and other timing signals are convolutionally extracted through 1D-CNN to identify the user's interaction intention (such as soothing, patting, dragging, etc.).

[0099] Step 2: Data processing;

[0100] To address the temporal and semantic differences between different modalities, this system uses the following fusion strategy:

[0101] (1) Temporal alignment: A temporal alignment network (CTA) based on the attention mechanism is used to align and weight the audio, visual, and tactile sequences so that they are synchronized in the semantic space.

[0102] (2) Feature fusion structure: Use a multi-channel Transformer structure or MCB (Multimodal Compact Bilinear Pooling) for deep fusion to ensure that each modality fully interacts in a unified feature space.

[0103] (3) Modality weight adjustment: The Modality Attention Gating (MAG) mechanism is introduced to automatically adjust the weight distribution of different modalities according to the current interaction environment and historical behavior to enhance robustness.

[0104] Step 3: Affective state reasoning and adaptive learning;

[0105] The fused features are fed into the Softmax classifier, which outputs the current user emotion label (such as happy, sad, surprised, angry, etc.) and confidence distribution.

[0106] Construct an emotional time series and use a sliding window mechanism to record historical emotional change trends, providing input support for the evolution of personality models and the generation of interaction strategies.

[0107] To improve individual adaptability, the system supports the use of Lightweight Continual Learning strategy:

[0108] 1) Dynamically maintain small personalized model parameters for each user;

[0109] 2) Rapid fine-tuning based on new data;

[0110] 3) Supports incremental training on the client side to avoid overfitting and catastrophic forgetting.

[0111] Step 4: Dynamic update of personality model;

[0112] The robot's personality model is adjusted through the personality update engine based on the identified user's emotional state, behavioral characteristics, and historical user interaction data.

[0113] Personality updates take into account short-term interaction feedback and long-term evolution trends to ensure the stability and adaptability of the robot's personality.

[0114] Step 5: Generate personalized interaction strategies;

[0115] Based on the updated personality model, an interaction strategy that adapts to the current situation and user preferences is generated, including action selection, voice content, expression display, etc.

[0116] Step 6: Behavior execution and user feedback collection;

[0117] The robot performs corresponding actions according to the interaction strategy and interacts with the user.

[0118] Through the feedback collection unit, monitor user reactions and evaluations and record interaction effects.

[0119] Step 7: Continuous learning and optimization;

[0120] Accumulate interaction data through personality memory units.

[0121] Continuously use new data to optimize personality models and interaction strategies to enhance the interactive experience.

[0122] Example 3:

[0123] The emotional interaction process between the robot and the user:

[0124] Step 1: The robot captures the user's smiling expression through the visual perception unit, the auditory perception unit captures the user's brisk voice tone, and the tactile perception unit perceives the user's action of tapping the robot.

[0125] Step 2: The data preprocessing and emotion recognition modules perform a fusion analysis on the multimodal data to identify that the user is in a "happy" emotional state and intends to interact with the robot.

[0126] Step 3: The personality update engine adjusts the robot's personality model to a "lively" mode based on the user's emotional state and interaction history, enhancing the tendency to actively interact.

[0127] Step 4: The strategy generation unit formulates interaction strategies such as the robot actively approaching the user, shaking its body, and making pleasant sounds.

[0128] Step 5: The robot executes the above interaction strategy, and the user shows happier emotions and continues to interact with the robot.

[0129] Step 6: The personality memory unit records the positive feedback of this interaction, and the personality update engine further optimizes the personality model based on it.

[0130] Example 4:

[0131] Application of robots in children's education:

[0132] Step 1: The robot discovers that the user (child) is depressed, detects tears through the visual perception unit, and captures sobbing sounds through the auditory perception unit.

[0133] Step 2: The sentiment analysis result of the emotion recognition module shows that the user is in a "sad" state and may need comfort.

[0134] Step 3: Adjust the personality model to the "gentle and considerate" mode, and increase the behavioral tendencies of comforting and encouraging.

[0135] Step 4: The strategy generation unit developed an interactive strategy of telling interesting stories and playing cheerful music.

[0136] Step 5: The robot implements the strategy, gradually easing the child's emotions, and the child begins to smile.

[0137] Step 6: The system records the interaction effect and provides a reference for subsequent similar situations.

[0138] Although the present invention is described herein with reference to illustrative embodiments of the present invention, the above embodiments are merely preferred embodiments of the present invention, and the embodiments of the present invention are not limited to the above embodiments. It should be understood that those skilled in the art can design many other modifications and implementations, which will fall within the scope and spirit of the principles disclosed in this application.

Claims

1. An intelligent pet robot system that implements multimodal perception and dynamic personality update, characterized by: The system includes: a multimodal perception module, a personality model and dynamic update module, a data processing and emotion recognition module, an interaction strategy generation and execution module, and a data storage and communication module, which are connected in sequence; Multimodal perception module: collects user's visual, audio, tactile information and environmental data in real time; Data processing and emotion recognition module: preprocesses, extracts features and performs emotion recognition on the information collected by the multimodal perception module; Personality model and dynamic update module: Dynamically updates the robot's personality model based on the user's emotional state and behavioral characteristics identified by the data processing and emotion recognition module, as well as historical interaction data; Interaction strategy generation and execution module: Generates personalized interaction strategies based on the robot's personality model and controls the robot to perform corresponding behaviors; Data storage and communication module: used to manage system operation data, user information, and model parameters, and communicate with external devices.

2. The intelligent pet robot system according to claim 1, characterized in that: The multimodal perception module includes: a visual perception unit, an audio perception unit and a tactile perception unit; Visual perception unit: equipped with a high-definition camera and depth sensor to capture the user's facial expressions, body movements and surrounding environment information; Audio perception unit: built-in omnidirectional microphone array to collect user voice information for speech recognition and emotion analysis; Tactile perception unit: Pressure sensors and touch sensors are arranged on the outside of the robot to sense the user's touch position, touch method, strength and frequency.

3. The intelligent pet robot system according to claim 2, characterized in that: The high-definition camera is used to capture the user's facial expressions, eyes, and posture information; The depth sensor is used to obtain the user's movements, gestures, and the distance between the user and the robot.

4. The intelligent pet robot system according to claim 2, characterized in that: The omnidirectional microphone array is used to collect audio information of the user's voice, angle, pitch, speaking speed and volume.

5. The intelligent pet robot system according to claim 1, characterized in that: The data processing and emotion recognition module includes: a data preprocessing unit, a feature extraction unit, a multimodal feature fusion unit and an emotion recognition and behavior analysis unit; Data preprocessing unit: filtering, noise reduction, and normalization of the collected multimodal data; Feature extraction unit: uses deep learning algorithms to extract key features of visual, audio, and tactile data; Multimodal feature fusion unit: fuses visual, audio, and tactile features; Emotion recognition and behavior analysis unit: combines multimodal features to identify the user's emotional state and behavioral intentions.

6. The intelligent pet robot system according to claim 5, characterized in that: The feature extraction unit is implemented using a deep learning model: (1) Visual feature extraction: Convolutional neural network (CNN) is used to extract image features; (2) Audio feature extraction: Use recurrent neural network (RNN) or long short-term memory network (LSTM) to extract speech features; (3) Tactile feature extraction: extract touch force and frequency features.

7. The intelligent pet robot system according to claim 1, characterized in that: The personality model and dynamic update module includes: a personality basic model library, a personality update engine and a personality memory unit; Personality basic model library: establish a variety of pet personality models, covering different personality traits; Personality Update Engine: A machine learning algorithm based on the generative adversarial network (GAN) reinforcement learning algorithm adjusts personality model parameters based on interaction data and user feedback; Personality Memory Unit: Stores historical interaction data and personality change records to support long-term learning and personalized evolution.

8. The intelligent pet robot system according to claim 7, characterized in that: The character update engine uses a deep reinforcement learning (DQN) algorithm, where: The input is: the current user's emotional state and behavioral intention, as well as the user's historical interaction data and feedback; The output is: updated personality model parameters; The workflow is: (1) Emotion recognition results and user feedback are input into the personality update engine; (2) The personality update engine calculates the optimal personality parameter adjustment strategy based on the current situation, historical data, and the personality basic model library; (3) The updated personality model is stored in the personality memory unit for use in subsequent interaction strategy generation.

9. The intelligent pet robot system according to claim 7, characterized in that: The interaction strategy generation and execution module includes: an interaction strategy generation unit, an action control unit and a feedback collection unit; the interaction strategy generation unit is connected to the personality memory unit, and a user feedback unit is connected between the personality memory unit and the feedback collection unit. The personality memory unit generates relevant interaction strategies, and the interaction strategies generate relevant control unit commands.

10. A control method for an intelligent pet robot system, for controlling an intelligent pet robot system according to any one of claims 1 to 9, characterized in that: The following steps are involved: Step A: Multimodal information collection; The robot uses a multimodal perception module to acquire the user's visual, audio, and tactile information, as well as environmental data, in real time. The visual modality uses a pre-trained image emotion recognition model and a CNN architecture with a ResNet50 + SE mechanism to extract texture and posture features of facial expressions. The audio modality uses a BiLSTM-based emotional speech recognition model to model MFCC, pitch, and energy feature sequences to capture temporal emotional signals. The tactile modality uses a 1D-CNN to convolve the user's touch position, force, and frequency signals to identify the user's interaction intent. Step B: Data processing; To address the temporal and semantic differences between different modalities, the following fusion strategy is used: (1) Temporal alignment: A temporal alignment network based on an attention mechanism is used to align and weight the audio, visual, and tactile sequences so that they are synchronized in the semantic space; (2) Feature fusion structure: Use a multi-channel Transformer structure or MCB for deep fusion to ensure that each modality fully interacts in a unified feature space; (3) Modality weight adjustment: Introducing a modality attention gating mechanism to automatically adjust the weight distribution of different modalities according to the current interaction environment and historical behavior to enhance robustness; Step C: Affective state reasoning and adaptive learning; The fused features are fed into the Softmax classifier, which outputs the current user emotion label and confidence distribution; Construct an emotional time series and use a sliding window mechanism to record historical emotional trends, providing input support for the evolution of personality models and the generation of interaction strategies; Step D: Dynamic update of personality model; Adjust the robot's personality model through the personality update engine based on the user's identified emotional state, behavioral characteristics, and historical user interaction data; Personality updates take into account short-term interaction feedback and long-term evolution trends to ensure the stability and adaptability of the robot's personality; Step E: Personalized interaction strategy generation; Based on the updated personality model, generate an interaction strategy that adapts to the current situation and user preferences; Step F: Behavior execution and user feedback collection; The robot performs corresponding actions according to the interaction strategy and interacts with the user; monitors the user's reactions and comments, and records the interaction effects; Step G: Continuous learning and optimization; Accumulate interaction data through personality memory units; Continuously use new data to optimize personality models and interaction strategies to enhance the interactive experience.

Citation Information

Patent Citations

  • Emotion communication device of pet robot

    CN112536808A

  • Intelligent emotional interaction method for service robot

    CN119150099A

  • Scene-based emotional interactive accompanying doll system and method

    CN120179069A

  • System

    JP2025053706A

  • Robotic control system

    WO2024163990A2

Cited By

  • Pet accompanying robot state control system and method capable of dynamically updating character

    CN120921385A

  • Interaction control method and device, electronic equipment and computer readable storage medium

    CN121028654A

  • Multi-modal perception and bionic action coordinated pet interaction device

    CN121561319A