Cross-modal emotion interaction system based on dynamic self-evolution anthropomorphic calculation

By employing multimodal fusion perception and self-evolutionary learning mechanisms, the shortcomings of cross-modal emotion interaction systems in terms of emotion recognition accuracy, adaptability, and interaction response speed are addressed, achieving efficient and natural human-computer interaction.

CN121659205APending Publication Date: 2026-03-13刘力军
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing cross-modal emotion interaction systems are insufficient in terms of emotion recognition accuracy, adaptability, and interaction response speed, making it difficult to meet the high reliability and humanization requirements in complex scenarios. Furthermore, they lack self-evolution capabilities, leading to misjudgments and unnatural interaction experiences.

Method used

Employing a multimodal fusion perception system, dynamic modal fusion and personalized feedback are achieved through refined preprocessing of speech, image, and text data and cross-modal self-attention fusion, combined with a self-evolutionary learning mechanism and a low-latency emotional interaction execution system.

Benefits of technology

It improves the accuracy of emotion recognition, reduces the cost of model adaptation, enhances the naturalness and adaptability of interaction, adapts to the needs of different users and scenarios, and meets the needs of applications in all scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121659205A_ABST
    Figure CN121659205A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal emotion interaction system based on dynamic self-evolution anthropomorphic calculation. According to the system, the functions of cross-modal perception, self-evolution learning and full-scene emotional interaction are realized through collaborative design of an innovative algorithm architecture and hardware and by taking a robot as a carrier. According to the algorithm, a multi-modal fusion perception system is constructed, an emotion feature vector is generated through a self-attention module, and accurate emotion recognition is realized based on Transform. The self-evolution mechanism adopts federated learning and abnormal optimization dual drive, continuous iteration and rapid adaptation of the model are achieved, and data security is guaranteed. The low-time-delay interactive execution system realizes anthropomorphic actions through kinematics solution, trajectory planning and a control algorithm, and combines a dynamic interactive library and reinforcement learning optimization strategy selection. Hardware adopts a polylactide topological optimization skeleton, has 17 degrees of freedom, and integrates a microphone array, a visual module and the like. The robot can be widely applied to the fields of pension services and the like, and emotion recognition precision, self-evolution ability and interaction naturalness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

1. Technical Field

[0001] This invention belongs to the interdisciplinary field of artificial intelligence and affective computing. It deeply integrates theories and methods from multiple disciplines such as computer science, automation control, cognitive psychology, and multimodal information processing, aiming to construct a cross-modal emotional interaction system with dynamic self-evolution capabilities and human-like computational characteristics, using robots as carriers. Specifically, this invention focuses on overcoming the limitations of traditional emotional interaction technologies in terms of modal fusion depth, adaptive learning capabilities, and human-like interaction effects. Through an innovative dynamic self-evolutionary algorithm architecture and cross-modal perception collaborative design, the system can accurately collect and analyze human emotional signals in multiple modal dimensions such as speech, text, and vision (facial expressions, body movements). Simultaneously, relying on a human-like computational model to simulate human emotional cognition and feedback logic, it dynamically optimizes interaction strategies. This technology can be widely applied in fields such as smart elderly care companionship, mental health intervention, special education guidance, smart home services, and personalized customer service in commercial scenarios. It provides key technical support for improving the naturalness, accuracy, and scenario adaptability of human-computer emotional interaction, helping to solve problems such as insufficient emotional care resources and homogenized interactive experiences in social services, and has significant technical value and social application significance. 2 Background Technology

[0002] With the accelerating aging of the global population and the increasing emphasis on mental health, the demand for systems with emotional interaction capabilities is experiencing explosive growth. Cross-modal emotional interaction systems based on dynamic self-evolving anthropomorphic computing hold the promise of filling gaps in social service resources, providing continuous and personalized care and support for the elderly, those with mental illnesses, and special education groups. However, the field of cross-modal emotional interaction systems based on dynamic self-evolving anthropomorphic computing currently faces many severe challenges, significantly limiting their widespread application and effectiveness in real-world scenarios.

[0003] 2.1 Background Technology Aspects

[0004] In the field of affective computing, numerous achievements have emerged. In 2024, the Emo robot, developed by Chinese scholars at Columbia University, was able to predict and mimic smiles in real time, contributing to the development of emotion recognition technology. The Pepper 6.0, an improved version launched by SoftBank Group in 2021, uses tactile sensors to sense the strength of a handshake or adjusts pupil focus based on ambient light, making robot interaction more natural. The Figure 01, a general-purpose humanoid robot released in 2024, equipped with a multimodal large model from OpenAI, can understand fuzzy commands, such as automatically handing an apple after receiving the "I'm hungry" signal, and can adjust its movement rhythm based on human facial expressions.

[0005] Existing physics-based emotion recognition research primarily employs three methods: visual, text, and audio. Text-based modality performs emotion analysis and recognition based on implicit emotions within a conversation, defining a two-dimensional emotion space model that includes Valance (degree of pleasure or displeasure) and Arousal (degree of activation or deactivation). Audio emotion recognition detects embedded emotions by processing speech signals. Given the complexity of emotional states in long utterances, RNNs and their variants (such as LSTM and Bi-LSTM) can be used to address the uncertainty of emotion labels. Visual emotion recognition includes Facial Expression Recognition (FER) and Body Posture Recognition (EBGR). EBGR aims to reveal a person's hidden emotional state from full-body or upper-body visual information. In existing ML-based EBGR systems, the input is abstracted into human gestures or their dynamics using a whole-body part or kinematic model. Then, machine learning-based methods or statistical measures are used to distinguish the output emotion, mapping the input to an emotion feature space, which is then identified by a machine learning-based classifier. In addition, physiological emotional signals such as EEG or ECG are often used for emotion analysis and recognition. EEG can directly measure changes in brain activity and provide intrinsic characteristics of emotional states. EEG-based emotion recognition often uses ML-based classifiers such as SVM or its variants, while DL-based EEG emotion recognition learns features and classifies emotions based on EEG at the same time.

[0006] Many other systems, supported by specific technologies, have demonstrated certain emotional interaction capabilities. For example, the "Youleyou - Mental Companion Robot" released by Shanghai Guixu Robotics Technology Co., Ltd. utilizes the core technology EmotionSim. TM The emotion simulation engine, trained with 865 emotion models, achieves multimodal emotional responses through voice, facial expressions, and body language. CloudMinds has built a multimodal artificial intelligence technology system, and its CloudGinger robot leverages the multimodal fusion AI capabilities of its cloud-based brain to achieve edge-side multimodal intelligent interaction. The dynamic face technology developed by Gravity (Ningbo) Electronic Technology Co., Ltd. is applied to the "Lingxi X2" robot, enabling it to engage in emotional communication with humans.

[0007] Nevertheless, current emotional interaction technologies still face many challenges in terms of emotion recognition and empathy. Existing systems do not analyze emotional contexts deeply enough and struggle to understand subtext; feedback mechanisms are prone to becoming template-based and lack individual adaptability; in cross-cultural interactions, they are highly susceptible to triggering cognitive biases and cannot achieve emotional resonance at the value level.

[0008] 2.2 Summary of Existing Problems

[0009] 1. Inherent limitations of single-modal emotion recognition

[0010] Traditional emotion-based interaction systems mostly rely on a single data modality for emotion recognition, such as inferring a user's emotional state solely from voice tone or text semantics. However, human emotional expression is extremely complex, often involving the synergistic effect of multiple dimensions of information, including facial micro-expressions, body language, vocal rhythm, and deep text semantics. For example, in everyday communication, a person might use a cheerful tone to recount a sad event; relying solely on the voice modality cannot accurately capture their true emotion. According to authoritative industry research reports, current single-modality emotion recognition technologies generally have an accuracy rate of less than 65% in complex scenarios, failing to meet the demands for accurate emotion judgment in practical applications. This limitation of a single modality makes the system highly susceptible to misjudgment when faced with complex emotional expressions, leading to inappropriate or even erroneous interactive feedback and severely impacting the user experience.

[0011] 2. Lack of self-evolutionary learning ability

[0012] Existing emotion-based interaction systems mostly employ machine learning models with fixed parameters, which are then statically run in real-world applications after training. However, different users have vastly different language habits, emotional expression styles, and environments. Factors such as regional dialects, personalized language styles, and varying environmental noise interference significantly impact emotion recognition and interaction effectiveness. Faced with these complex and ever-changing realities, systems with fixed model parameters cannot dynamically optimize their algorithms based on real-time interaction data, exhibiting extremely poor adaptability. This necessitates significant manual intervention, requiring regular model retraining and parameter adjustments, which not only consumes enormous human, material, and time resources but also makes it difficult to guarantee consistently high performance in rapidly changing scenarios.

[0013] 3. Severe interaction response delays significantly impact the user experience.

[0014] In emotional interaction, the latency between receiving user commands and responding with corresponding actions is crucial. Most current emotional interaction systems suffer from excessively high end-to-end latency, typically exceeding 300ms. This prolonged latency leads to a disconnect between the system's voice responses and body language, severely compromising the smoothness of the human-like interaction. Users will experience a noticeable disconnect and unnaturalness when communicating with the system. For example, if a user asks a question and the system takes a long time to respond, with delayed actions, this significantly reduces user trust and positive feelings towards the system, hindering deeper emotional interaction.

[0015] In summary, existing emotional interaction system technologies have significant shortcomings in terms of emotion recognition accuracy, adaptability, interaction response speed, and ethical security. There is an urgent need for an innovative technical solution to meet the needs of emotional interaction in all scenarios, with high reliability and humanization. 3. Summary of the Invention

[0016] Purpose of the Invention: The purpose of this invention is to overcome the shortcomings of existing cross-modal emotional interaction technologies, such as fixed modal fusion, static emotional computing, insufficient anthropomorphism, and lack of self-evolution capabilities. It provides a cross-modal emotional interaction system based on dynamic, self-evolving anthropomorphic computing using a robot as the carrier. The aim is to achieve the following objectives: First, to construct a dynamic modal fusion mechanism that can dynamically adjust the weight allocation of multimodal emotional features such as speech, text, and vision according to the scenario, improving the accuracy of emotion recognition in complex environments; second, to design an anthropomorphic emotional computing model that simulates the continuous evolution of human emotions, achieving a fine depiction from discrete labels to dynamic changes in emotional intensity; third, to endow the system with self-evolution capabilities, autonomously optimizing emotional cognition logic through real-time learning of interaction data to adapt to the needs of different user groups and scenarios; fourth, to bridge the cross-modal semantic gap, uncovering implicit correlations between multimodal data and enhancing the consistency of emotional understanding; and fifth, to generate personalized feedback that conforms to human cognitive habits, improving the naturalness of human-computer interaction and the depth of emotional connection.

[0017] Technical Solution: (I) Cross-modal Perception and Fusion Architecture: This invention innovatively designs a multi-modal fusion perception system, breaking through the limitations of single-modal recognition: it simultaneously collects three types of data—speech, image, and text—and performs refined preprocessing based on the characteristics of different modalities. Speech data undergoes Webrtcvad endpoint detection to remove invalid signals, and is pre-emphasized, framed, and windowed using the SoX tool. Then, 1582-dimensional features are extracted using the OpenSMILE toolkit, and finally, principal component analysis reduces the dimensionality to 512 dimensions, retaining 98.5% of the variance information. Image data relies on the collaborative processing of YOLOv8 and CLIP dual models. YOLOv8 first achieves rapid object localization, outputting bounding box coordinates and class probabilities. Then, the CLIP model corrects misclassifications, simultaneously extracting a 256-dimensional facial feature vector, containing the dynamic change rate of 68 facial key points. Text data is processed by removing special symbols using regular expressions, then segmented by jieba and encoded using the BERT-base model. Jieba segmentation supplements the domain-defined dictionary, generating a 768-dimensional semantic feature vector. Simultaneously, an entity-relationship graph is constructed to achieve deep semantic analysis. A pioneering cross-modal self-attention fusion module unifies speech (512-dimensional), image (256-dimensional), and text (768-dimensional) features into the same dimensional space through linear mapping. A similarity matrix is ​​calculated among the three to obtain an association score, and weights are dynamically allocated based on the score: the weight increases to 0.6 when speech emotion features are significant, to 0.5 when image expression features are prominent, and to a maximum of 0.4 when text semantic tendency is clear. A 768-dimensional comprehensive emotion feature vector is generated through weighted summation, with the basic weight formula as shown in equation (1).

[0018] F 融合 =0.4×F 语音+0.35×F 图像 +0.25×F 文本 (1)

[0019] It also supports dynamic adjustment based on context, solving the problem of multimodal information alignment and enabling emotional features to more comprehensively reflect the user's state.

[0020] (II) Self-evolutionary learning mechanism

[0021] A self-evolving framework driven by federated learning and anomaly optimization is constructed to achieve continuous model iteration: The federated learning architecture is deployed on Alibaba Cloud ECS (8 cores, 16GB memory) using a dual-machine hot standby mode to ensure failover time is ≤30 seconds. The FedAvg algorithm is used for global model aggregation and parameter distribution. Each system acts as an edge node, retaining nearly 3 days of interaction data in its local storage module, including voice clips, image frames, text dialogues, and feedback information. The local model is updated using the SGD optimization algorithm, with a batch size of 32 and a learning rate of 1e-4. Training stops when the loss fluctuation of 5 consecutive batches is <0.001. A publish-subscribe message queue is built using the MQTT protocol, and model parameters are transmitted using TLS 1.3 encryption. A global iteration is completed every 24 hours, executed during low-load periods, with each node uploading ≤10MB of parameters. Data privacy is protected using a local differential privacy mechanism combined with model compression technology. Pruning retains 60% of important parameters, and INT8 quantization is used, reducing transmission volume by 75% while maintaining an accuracy loss of ≤1%. The system employs a closed-loop feedback mechanism to address anomalies. Anomaly triggering conditions include: emotion recognition accuracy <85% for 3 consecutive hours, user interaction interruption rate >30%, and action execution failures >5 times / day. The system automatically filters anomalous samples and pushes them to the LabelStudio platform. Psychology professionals label these samples with 7 emotion categories (anger, sadness, joy, etc.) and their intensity (0-100 points). Consistency is verified by checking if the Cohen's Kappa coefficient is ≥0.85 before incorporating them into the incremental training set. Model fine-tuning utilizes knowledge distillation: a federated global model serves as the teacher model, and a local model serves as the student model. Parameters are optimized using a soft-label loss (weight 0.7) with a temperature coefficient of 5 combined with a hard-label loss (weight 0.3). Only the classification layer and the last two Transformer layers are updated. Training is completed for 5 epochs (learning rate 5e-5), optimization is completed within 2 hours, and updates are applied when the test set accuracy improves by more than 3%.

[0022] (III) Low-latency emotional interaction execution system

[0023] An innovative motion control and interaction strategy generation scheme achieves human-like interaction: Based on a 17-DOF bionic joint design (3 DDOF for the head, 2 DDOF for the torso, 6 DDOF each for the arms, and 2 DDOF for the hand, with repeatability accuracy ±0.05°), a Denavit-Hartenberg parametric model is established to describe the joint relationships. Damped least squares method is used to solve the inverse kinematics, and joint angles are iteratively optimized to bring the end effector to the target pose, achieving a solution accuracy of ±0.1°. The calculation time for a single target pose is ≤5ms, meeting real-time control requirements. Trajectory planning employs a fifth-order polynomial interpolation method to ensure continuity of position, velocity, and acceleration, with jerk ≤5m / s². 3 To avoid mechanical impact, the "handshake" action is planned in three phases, taking it as an example: the initiation phase (0-0.2s, acceleration linearly increases from 0 to 0.5m / s²). 2 The system consists of a constant speed phase (0.2-0.5s, maintaining a speed of 0.1m / s) and a deceleration phase (0.5-0.7s, with acceleration decreasing to 0), ultimately achieving a position error ≤2mm. The control algorithm employs a PID + feedforward composite control: proportional coefficient Kp = 5.0 for rapid response to position deviation; integral coefficient Ki = 0.1 to eliminate steady-state error; and derivative coefficient Kd = 0.5 to suppress oscillation. The feedforward compensation term pre-calculates the control quantity based on the acceleration and velocity of the desired trajectory, achieving a position tracking error ≤0.5mm. When the hand contact force exceeds 1N, the driving torque is automatically reduced with a damping coefficient of 0.8 to prevent user discomfort. A dynamic interaction library containing over 1000 strategies is constructed. Each strategy includes a "script template + action combination + voice parameters." For example, the "sadness comfort" strategy consists of the script "I can sense you're not happy right now, would you like to talk?", the action being a light tap on the shoulder with the right hand (speed 5cm / s, force 0.5N) + a slight tilt of the head (15° pitch), and voice parameters including a 20% decrease in tone, a 15% decrease in speech rate, and a 5% increase in volume. Strategy selection is optimized using Q-Learning reinforcement learning, with user responsiveness (dialogue duration as a percentage of total time + positive facial expressions as a percentage of total expression) as the reward function, iteratively updating the Q-value to adapt to user preferences.

[0024] (iv) Hardware Collaborative Design for Full-Scenario Adaptation

[0025] Customized, highly integrated hardware and environment-adaptive solutions: The mechanical structure uses polylactide (PLA) material (density 1.24 g / cm³). 3The structure features a yield strength of 50MPa, a topology-optimized design, a 3.2mm wall thickness for key load-bearing components, and a 17-DOF bionic joint, including a metal gear servo motor. It boasts an accuracy of 0.3° and a repeatability of ±0.05°. The sensor system integrates: a 6-unit ReSpeaker6-MicArray ring microphone array supporting 360° pickup with a pickup radius of 0.5-5m, featuring beamforming directional gain ≥12dB and noise suppression ≤-40dB; an Intel RealSense D435i vision module with a resolution of 1920×1080@30fps, supporting low-light environments, detecting 68 facial key points at a minimum illumination of 0.01 lux with an error ≤1.5 pixels; and a multi-mode voice recognition module powered by DC5V. 2 The system supports C-band communication, 150 candidate recognition sentences, and a pickup distance of 1 meter in quiet environments and 30cm in noisy environments. The computing and control unit utilizes a Raspberry Pi 5 with a BCM2712 processor, integrating Bluetooth / Wi-Fi and a rich array of interfaces including USB 2.0 / 3.0 and Gigabit Ethernet. Powered by an 11.1V 2000mAh lithium battery, it provides up to 8 hours of static interaction or 4 hours of dynamic activity, enabling efficient local data processing. The output modules include a 0.96-inch SPIOLED abdominal display with a resolution of 128×64, a 128×64 resolution OLED head expression screen, and a highly integrated voice synthesis module for multi-dimensional emotional expression. The hardware design is adapted to operating temperatures of -50℃ to 80℃, relative humidity of 30% to 80% (non-condensing), light conditions of 50-10000 lux, and acoustic environments ≤70dB. It employs an over-threshold automatic noise reduction algorithm, improving the signal-to-noise ratio by 20dB. Combined with edge computing capabilities, it ensures stable operation in various scenarios such as elderly care, services, and healthcare.

[0026] Beneficial Effects: The beneficial effects of this invention are significant. (1) Improved accuracy of emotion recognition: The combination of multimodal fusion and Transformer classification model achieves an emotion recognition accuracy of 92.3%, accurately distinguishing 7 basic emotions and solving the problem of emotion misjudgment in complex scenarios. (2) Significant self-evolution capability: Federated learning enables the model to iterate once every 24 hours, and the abnormal optimization mechanism completes the adaptation within 2 hours, reducing the model adaptation cost by 85%. It can automatically adapt to different user habits and scene changes without frequent manual intervention. (3) Improved naturalness of interaction: The end-to-end latency of action control is as low as 80ms, and the trajectory smoothness reaches 90%. 91% of human behavior standards, combined with dynamic interaction strategies, make the user interaction experience more natural. In smart elderly care, it can reduce the loneliness index of the elderly by 42% and increase social interaction by 2.3 times in special education; (4) Strong applicability in all scenarios: The combination of wide temperature design and multimodal perception can work stably in various environments such as indoor and outdoor, noisy / quiet, adapting to multiple fields such as smart elderly care, mental health, and special education, providing intelligent support for social services; (5) Data security compliance: The combination of federated learning local differential privacy and transmission encryption (TLS1.3) ensures that user emotional data is not leaked and complies with ethical and legal requirements. 4. Attached Figure Descriptions

[0027] Figure 1 It is a flowchart of the overall system interaction process;

[0028] Figure 2 This is a diagram of the voice data preprocessing process;

[0029] Figure 3 This is a diagram of the endpoint detection language data processing procedure;

[0030] Figure 4 This is a diagram of the feature extraction process;

[0031] Figure 5 This is a diagram illustrating the entity recognition process using the YOLOv8+CLIP model combination.

[0032] Figure 6 This is a diagram of a sentiment classification model based on Transformer;

[0033] Figure 7 It is a standard emotional psychological model diagram;

[0034] Figure 8 It is a multi-view model diagram of the robot;

[0035] Figure 9 It is a full-body model of the robot;

[0036] Figure 10 This is a diagram showing the dimensions of the robot's head.

[0037] Figure 11This is a diagram showing the dimensions of the robot's abdomen;

[0038] Figure 12 This is a diagram showing the dimensions of the robot's limbs;

[0039] Figure 13 This is a diagram showing the dimensions of the robot's hand.

[0040] Figure 14 This is a diagram showing the dimensions of the robot's feet;

[0041] Figure 15 It is a programmable speech recognition module;

[0042] Figure 16 This is a size diagram of a USB camera;

[0043] Figure 17 This is a design diagram of the Raspberry Pi development block;

[0044] Figure 18 It is the servo motor design dimension drawing;

[0045] Figure 19 It is a lithium battery design dimension drawing;

[0046] Figure 20 This is a design drawing of the OLED expression screen on the robot's head;

[0047] Figure 21 This is a design diagram of a highly integrated speech synthesis module. 5. Detailed Implementation

[0048] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading the present invention, any modifications of the present invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.

[0049] 5.1 Algorithm Flow Technical Details

[0050] 5.1.2 Multimodal Data Acquisition and Preprocessing

[0051] 1. Voice data processing

[0052] Preprocessing: such as Figure 2 The diagram illustrates the speech data preprocessing procedure. Webrtc VAD is used for endpoint detection, with a silence threshold of -30 dBFS. By continuously checking whether the energy of audio frames falls below this threshold, speech and non-speech segments are accurately distinguished, and invalid signals such as environmental noise and breathing sounds are removed. Figure 3The diagram shows a language data processing procedure. The SoX tool is used for pre-emphasis, with a coefficient of 0.97, to enhance the energy of high-frequency signals and improve the intelligibility of speech details. The frame processing adopts an overlapping method with a frame length of 25ms and a step size of 10ms to ensure the continuity of the speech signal and reduce the loss of information between frames. The Hanning window is selected for windowing to reduce the discontinuity of the frame edge signal and avoid spectrum leakage from affecting subsequent analysis.

[0053] Feature extraction: 1582-dimensional features are extracted using the IS10_paraling configuration of the OpenSMILE toolkit, such as... Figure 4 As shown, feature extraction includes:

[0054] Prosodic features: The fundamental frequency is F0, ranging from 50 to 500 Hz, extracted using the YIN algorithm. This algorithm accurately captures the fundamental frequency period by calculating the signal autocorrelation function and is applicable to speech features of different genders and ages; the intensity is energy ranging from -60 to 0 dB, reflecting the changes in speech intensity through intra-frame signal energy calculation; the speech rate is measured in syllables per second, based on speech frame boundaries and phonetic feature statistics, reflecting the rhythm of speech.

[0055] Spectral characteristics: 128-dimensional Mel spectral coefficients, which map the linear spectrum to a Mel scale that conforms to the characteristics of human hearing through a Mel filter bank, enhancing the capture of emotionally relevant frequency bands in speech; spectral entropy, which reflects the complexity of the spectral distribution and is used to distinguish between stationary and abrupt states in speech; spectral centroid, which represents the frequency position where spectral energy is concentrated and is associated with the timbre characteristics of speech; and bandwidth, which describes the width of the spectral energy distribution and helps to judge the clarity of speech.

[0056] Dimensionality reduction: Principal component analysis (PCA) was used to reduce the 1582-dimensional features to 512-dimensional features. During the calculation process, the main variance information of the features was retained, with a cumulative retention of 98.5%. This significantly reduced the data dimensionality to reduce the computational load of subsequent models while ensuring that key sentiment features were not lost, thus improving the system's processing efficiency.

[0057] 2. Image Data Processing

[0058] (1) Building a basic image recognition model

[0059] A dual-model collaborative architecture combining the YOLOv8 object detection model and the CLIP model is adopted:

[0060] YOLOv8 model: Based on the CSPDarknet backbone network and PAN-FPN feature fusion structure, it achieves fast localization and preliminary classification of objects in images through anchor box clustering and dynamic label allocation mechanism. It supports the recognition of 80 common object classes and outputs the object bounding box coordinates (x,y,w,h) and preliminary class probability.

[0061] The CLIP model consists of a visual encoder (a variant of ResNet-50) and a text encoder (Transformer). It is pre-trained on large-scale image and text data through contrastive learning, enabling cross-modal semantic understanding. For candidate object regions output by YOLOv8, the CLIP model matches visual features with text descriptions, such as "phone" or "cup," correcting misclassifications, such as distinguishing between "cup" and "bowl," and improving the generalization of recognition for small samples or rare objects.

[0062] The two work together through feature cascading: YOLOv8 provides efficient object localization results, while the CLIP model optimizes classification accuracy based on semantic matching, together achieving complete functions of object recognition, scene classification, and extraction of visual features of user behavior in images.

[0063] (2) Object recognition implementation

[0064] Based on the above YOLOv8+CLIP model combination, such as Figure 5 As shown, the object recognition process is as follows:

[0065] Region extraction: YOLOv8 performs multi-scale feature detection on the input image with a resolution of 640×640, generates candidate object bounding boxes, outputs ≤300 candidate boxes per frame, and filters out low-quality results with confidence <0.5;

[0066] Category correction: The CLIP model performs secondary encoding on the candidate regions and calculates the cosine similarity between the visual features and the pre-set object category text library. The object category text library contains more than 1,000 kinds of everyday object names, and the category with the highest similarity is selected as the final result.

[0067] Accuracy Validation: Tested on the COCO2017 validation set, the combined model achieved an object recognition accuracy of 98.2% at mAP@0.5, i.e., an intersection-union ratio ≥0.5. Among them, the recognition accuracy for common objects such as furniture and electronic devices was ≥99%, and the recognition accuracy for small objects was ≥95%, meeting the needs for fast and accurate object recognition in real-world scenarios.

[0068] (3) Scene classification achieved

[0069] By leveraging model ensembles to jointly analyze global image features and local object distribution, scene classification is achieved.

[0070] Feature fusion: YOLOv8 extracts the category and location distribution features of all objects in the image, while the CLIP model extracts global scene semantic features. The two are then fused by weighting to generate a scene description vector.

[0071] Classification output: Input the fused vector into the SoftMax classifier to classify the preset scene categories;

[0072] (4) Extraction of visual features of user behavior

[0073] By combining models to perform fine-grained analysis of the user's visual information in the image, the following behavioral features are extracted:

[0074] Posture characteristics: After YOLOv8 locates the human body area, it extracts 17 human body key points through the built-in key point detection branch, calculates the relative position and angle changes between key points, and recognizes 8 common gestures such as waving, liking, and pointing. The action recognition accuracy is ≥96%.

[0075] Facial features: Combining the facial key point localization results mentioned above, the CLIP model helps to determine the facial expression category (6 basic expressions such as smiling, frowning, and surprised), and calculates the rate of change of expression intensity through the feature difference of consecutive frames, such as the transition of the degree of smiling from "slight" to "laughing".

[0076] Feature application: The extracted behavioral visual features are output as 128-dimensional vectors and will be fused with voice and text features for subsequent user intent recognition, such as "waving" corresponding to the intent of "greeting", or interactive response such as adjusting the robot's feedback tone based on the "smiling" expression to improve the naturalness of human-computer interaction.

[0077] 3. Text Data Processing

[0078] (1) Text semantic parsing and graph construction

[0079] Based on text data such as user comments and search history, we conduct text semantic analysis and entity-relationship graph construction. The specific process is as follows:

[0080] Text preprocessing: The original text was standardized by removing special characters such as punctuation, URLs, and garbled characters using regular expressions. A domain-specific dictionary was configured using the jieba word segmentation tool to supplement industry terms such as "elderly care services," and Chinese word segmentation was performed. Then, the English text was segmented using the NLTK tool, uniformly converting the text into word sequences. Simultaneously, the Harbin Institute of Technology's LTP tool was used for part-of-speech tagging and named entity recognition, identifying entities such as users, products, and points of interest, laying the foundation for subsequent semantic analysis.

[0081] Entity-Relationship Graph Construction: The BERT-base model from the BERT series is used to extract semantic features from the preprocessed text. Feature representation is optimized through the Masked Language Model and Next Sentence Prediction task. For entity pairs in the text, an attention mechanism is used to mine semantic relationships, combined with dependency parsing, to construct an entity-relationship graph. For example, parsing the text "User Xiaoming is interested in smartwatches and other technology products," the system can identify the entities "Xiaoming," "smartwatch," and "technology products," extracting the relationship "user-interest-technology product." This achieves entity disambiguation, distinguishing between users with the same name, products with multiple meanings, and extracting relationships. It supports multilingual text processing, including Chinese and English.

[0082] (2) Semantic understanding and knowledge application

[0083] Semantic understanding engine optimization: Based on the BERT pre-trained model, a semantic understanding engine suitable for elderly care, service, and medical scenarios is built through fine-tuning. To address the needs of scenario-based question answering, domain-specific question answering data (such as "elderly chronic disease care questions and answers" in the elderly care scenario) is collected to construct a training set. The data is organized in the form of triples, where each triple represents the question, intent classification, and answer. The Cross-Entropy Loss multi-classification loss function is used to optimize the model, enabling it to learn scenario-based semantic representations and recognize the intent of user questions. Testing shows that in question answering tasks in elderly care, education, and medical scenarios, the F1 score reaches 0.85±0.05, accurately understanding user semantics and matching answers.

[0084] Knowledge Graph Applications and User Profile Updates: The knowledge graph relies on the Neo4j graph database for storage and supports real-time reasoning. When text semantic parsing identifies new entities and relationships, the knowledge graph can use graph traversal algorithms to perform a depth-first search to connect with existing knowledge, while dynamically updating the user profile. The user profile integrates multimodal interaction data (text, voice, and images), and based on the knowledge graph reasoning results, supplements user interest tags and needs preferences to drive personalized recommendations, achieving a precise match between knowledge applications and user needs.

[0085] (3) Text data-driven interaction enhancement

[0086] The text data processing results, namely entity-relationship graphs and semantic understanding results, are deeply integrated with the multimodal system. Through multimodal data collaboration, the naturalness and intelligence of robot interaction are improved, making text data a core driving element for multimodal emotional interaction and personalized services, supporting the efficient application of robots in scenarios such as elderly care, education and tutoring, and medical consultation.

[0087] 5.1.2 Emotion Recognition and Decision-Making Mechanism

[0088] 1. Multimodal feature fusion

[0089] A cross-modal self-attention module is employed. First, the 512-dimensional speech features, 256-dimensional image features, and 768-dimensional text features are unified to the same dimensional space through linear mapping. Then, a similarity matrix is ​​calculated among the three, with dimensions equal to the unified feature dimensions multiplied by the unified feature dimensions. Association scores for each modality feature are obtained through matrix operations. Based on these scores, a dynamic weight allocation mechanism is used: when speech emotion features are significant, the speech feature weight is automatically increased to 0.6; when image expression features are more prominent, the image feature weight can be increased to 0.5; and when the text semantic sentiment is clear, the text feature weight can be adjusted up to 0.4, ensuring that the weight allocation matches the intensity of emotional expression in each modality.

[0090] Fusion strategy: A 768-dimensional comprehensive sentiment feature vector is generated by weighted summation, and its specific calculation formula is shown in equation (2):

[0091] F 融合 =0.4×F 语音 +0.35×F 图像 +0.25×F 文本 (2)

[0092] (The weight can be dynamically adjusted according to the scenario, such as increasing the text weight to 0.35 in the elderly care scenario.)

[0093] 2. Sentiment Classification Model

[0094] Model Structure: A Transformer-based sentiment classifier is employed, consisting of six stacked encoder layers. Each layer includes a multi-head self-attention mechanism and a feedforward neural network. The multi-head attention mechanism employs eight attention heads, each with a feature dimension of 96. The feedforward neural network has a dimension of 2048, and the activation function is ReLU. Figure 6 As shown, the input is a 768-dimensional comprehensive emotional feature vector. After the encoder extracts higher-order emotional features layer by layer, the output layer uses a softmax activation function to output the probability distribution of seven emotions, corresponding to anger, sadness, joy, fear, surprise, disgust, and neutrality. The standard emotional psychological model diagram is shown below. Figure 7 As shown, this enables fine-grained classification of user emotions.

[0095] Training Details: The Adam optimizer was used for training, with a learning rate of 0.001, a first-moment estimation decay rate β1 = 0.9, a second-moment estimation decay rate β2 = 0.99, and a weight decay coefficient of 0.0001 to prevent overfitting. The training dataset contained 300,000 multimodal labeled data points, each of which had been processed and included speech, image, text, and corresponding sentiment tags. Training was conducted over 100 epochs: the learning rate was linearly increased to 0.001 for the first 20 epochs, maintained at a stable learning rate for the middle 60 epochs, and cosine decayed to 0.0001 for the last 20 epochs. During training, the cross-entropy loss function gradually converged to 0.12. The sentiment classification accuracy on the independent test set (containing 50,000 data points not used in training) reached 92.3%, with an F1 score of 0.91. The accuracy for recognizing strong emotions such as joy and anger exceeded 95%, and the accuracy for recognizing neutral emotions reached 89%.

[0096] 3. Interaction Strategy Generation

[0097] Strategy library design: Contains 1000+ preset scenario strategies. Each strategy consists of "script template + action combination + voice parameters". For example, the "sadness comfort" strategy is defined as:

[0098] Script: "I can tell you're not very happy right now. Would you like to talk?"

[0099] Action: Gently tap the user's shoulder with your right hand (speed 5cm / s, force 0.5N), head slightly lowered (tilt angle 15°).

[0100] Voice: Pitch reduced by 20%, speech rate slowed by 15%, volume increased by 5%.

[0101] Dynamic Adjustment: Based on historical user interaction data, nearly 30 days of interaction records are stored, including user response duration, facial expressions, and verbal responses to different strategies. Q-Learning reinforcement learning is used to optimize strategy selection. The strategy selection process is modeled as a Markov decision process: the state is the current user's emotional category + user profile label, the action set is the candidate strategies in the strategy library, and the reward function is set to the user's responsiveness, with dialogue duration accounting for 60% and positive facial expressions accounting for 40%. By iteratively updating the Q-value (policy-state value), the robot gradually adapts to user preferences through repeated interactions—for example, for users who prefer quiet companionship, the frequency of verbal dialogue is automatically reduced, while the proportion of action-based strategies such as patting or offering water is increased, enhancing the personalization of the interaction. The overall robot interaction flow is as follows: Figure 1 As shown.

[0102] 5.1.3 Motion Control and Execution

[0103] 1. Kinematic solution

[0104] DH Parameter Table: Establish Denavit-Hartenberg parametric models for 17 joints, such as shoulder joint parameters:

[0105] Table 1. Parameters of the DH model for the shoulder joint

[0106]

[0107] This parametric model can transform the rotational and translational motions of joints into transformation matrices in a unified coordinate system, providing a foundation for kinematic analysis.

[0108] Inverse kinematics solution: Damped least squares method is used for inverse kinematics calculation. Joint angles are iteratively optimized to bring the end effector to the target pose. This method introduces a damping term in the solution process to avoid the divergence problem of numerical iteration. The solution accuracy can reach ±0.1°, and the calculation time for a single target pose is ≤5ms, which meets the requirements of real-time control.

[0109] 2. Trajectory Planning

[0110] A smooth trajectory is generated using fifth-order polynomial interpolation to ensure continuity of position, velocity, and acceleration, with a jerk value (jerk) ≤ 5 m / s². 3 To avoid mechanical impact. For example, the trajectory of a "handshake" motion:

[0111] Start-up phase (0-0.2s): Acceleration increases linearly from 0 to 0.5m / s². 2 The end effector is driven to accelerate from its initial position toward the target direction;

[0112] Uniform speed phase (0.2-0.5s): Maintain a uniform speed of 0.1m / s to ensure smooth movement;

[0113] Deceleration phase (0.5-0.7s): The acceleration decreases linearly to 0, and the end effector slowly stops at the target position, with a final position error ≤2mm.

[0114] 3. Control Algorithm

[0115] PID+feedforward control is adopted: PID parameters: proportional coefficient Kp = 5.0 (fast response to position deviation), integral coefficient Ki = 0.1 (eliminating steady-state error), derivative coefficient Kd = 0.5 (suppressing oscillation); feedforward compensation: based on the acceleration and velocity information of the desired trajectory, the control quantity is calculated in advance and superimposed on the PID output to compensate for system inertia and achieve a position tracking error ≤ 0.5mm.

[0116] Force control adjustment: When the contact force of the hand exceeds 1N, the force control adjustment logic is triggered: the contact force is monitored in real time by the torque sensor, and the driving torque is automatically reduced. The attenuation coefficient is 0.8, which reduces the pressure on the user and avoids discomfort, while maintaining the stability of the movement.

[0117] 5.2 Implementation Scheme of Self-Evolution Mechanism

[0118] 5.2.1 Deployment of the Federated Learning Framework

[0119] 1. Node Architecture

[0120] The federated server is deployed on Alibaba Cloud ECS, configured with 8 cores and 16GB of memory, running a Linux operating system and Python environment, and using the FedAvg algorithm as the core aggregation logic. The server primarily performs three functions: receiving local model parameters uploaded from each edge node, generating a global model through weighted averaging, and distributing the updated global model parameters to all edge nodes. To ensure high availability, the server adopts a dual-machine hot standby mode; when the primary node fails, the standby node can automatically switch over and take over the service within 30 seconds. Each robot acts as an edge node, i.e., a local training node, retaining the interaction data of the most recent 3 days in its local storage module, with an average daily data volume of approximately 500 records, covering voice clips, image frames, text dialogues, and corresponding interactive feedback information. Local training uses the SGD optimization algorithm, with a batch size of 32 and a learning rate of 1e-4, updating the local model parameters iteratively batch by batch. During training, the loss value is monitored in real time; when the loss fluctuation is less than 0.001 for 5 consecutive batches, the local training is stopped and preparations are made to upload the parameters. Communication utilizes the MQTT protocol to build a publish-subscribe message queue. Model parameter transmission is secured using TLS 1.3 encryption to prevent data leakage or tampering during transmission. The system is configured to perform a federated iteration every 24 hours. Each edge node completes parameter uploading and downloading during the low-load period from 3 AM to 5 AM, with the amount of parameters uploaded by a single node each time controlled within 10MB to ensure efficient and stable network transmission.

[0121] 2. Data privacy protection

[0122] To protect user data security, a local differential privacy mechanism is employed. Before uploading model parameters at edge nodes, Laplacian noise is added to the parameter values ​​with a noise coefficient of 1.0. This process masks the characteristics of the original interaction data while ensuring that the parameters can still effectively update the global model during aggregation, preventing any third party from inferring the user's original interaction information from the model parameters. Model compression technology reduces the amount of data transmitted, improving the efficiency of federated learning. Pruning operations retain 60% of the important parameters, and core weights are selected by calculating the L1 norm of parameters at each layer, removing redundant parameters with minimal impact on model performance. Quantization converts the parameter precision from FLOAT32 to INT8, reducing storage usage while accelerating transmission. Actual testing shows that the compressed model parameters reduce transmission volume by more than 75% while ensuring that the precision loss does not exceed 1%.

[0123] 5.2.2 Anomaly Feedback and Optimization Process

[0124] 1. Anomaly detection triggering conditions

[0125] Anomaly detection is triggered when the emotion recognition accuracy is below 85% for 3 consecutive hours. The accuracy is calculated by comparing with manually labeled samples. Anomaly detection is triggered when the user interaction interruption rate exceeds 30%. Interaction interruption refers to a conversation duration of less than 10 seconds. Anomaly detection is triggered when the number of failed action executions exceeds 5 per day. Failures include joint jamming, collisions, etc.

[0126] 2. Manual annotation process

[0127] The system automatically filters out abnormal samples, including speech segments with emotion recognition errors, images with blurry facial expression recognition, and text dialogues that interrupt interaction. These are then categorized by scenario and pushed to the LabelStudio annotation platform. The platform supports simultaneous annotation of multimodal data, with the annotation interface simultaneously displaying speech waveforms, image frames, and text content for easy comprehensive judgment by annotators. Annotators must have a psychology background and be familiar with the characteristics of seven basic emotions, performing dual annotation of emotion category and intensity on samples. Emotion categories are divided into seven types: anger, sadness, joy, fear, surprise, disgust, and neutral. Intensity scores range from 0 to 100, with higher scores indicating stronger emotion expression. After annotation, the system verifies annotation consistency using the Cohen's Kappa coefficient, requiring a coefficient of at least 0.85. Samples that do not meet this standard are reassigned to experienced annotators for re-annotation. Annotated data undergoes a three-level review before being included in the incremental training set. Each batch of training data contains 1000 annotated data points, divided into training, validation, and test sets in a 7:2:1 ratio. The review includes category accuracy, intensity reasonableness, and scenario matching to ensure the quality of the annotated data meets the model training requirements.

[0128] 3. Model fine-tuning strategy

[0129] A knowledge distillation approach is used for model optimization, with the global model generated by federated learning serving as the teacher model and the local robot model serving as the student model. The teacher model generates soft labels using a softmax function with a temperature coefficient of 5, while the student model learns both soft and ground truth labels simultaneously. Parameters are optimized using a weighted loss function, with a soft label loss weight of 0.7 and a hard label loss weight of 0.3. The training process consists of 5 epochs with a learning rate of 5e-5, using the Adam optimizer for parameter updates. An incremental update strategy is employed to reduce computational resource consumption, updating parameters only in the classification layer and the last two Transformer layers, while freezing the weights of other layers. This approach ensures improved model performance while keeping each fine-tuning session under 2 hours, avoiding disruption to the robot's normal interactive functions. After fine-tuning, the system automatically verifies model performance on the test set. If the accuracy improvement exceeds 3%, the updated model is saved and applied to actual interactions; otherwise, it reverts to the model state before fine-tuning.

[0130] 5.3 Detailed Hardware Configuration Scheme for Robot Carrier

[0131] 5.3.1 Core Parameters of the Robot

[0132] 1. Mechanical structure

[0133] Skeleton material: Polylactide (PLA) material (density 1.24 g / cm³) 3 (Yield strength 50MPa, tensile strength 34.56MPa), topology optimized design, key load-bearing components wall thickness 3.2mm. Specific body frame dimensions design drawings are shown below. Figure 8 , Figure 9 , Figure 10 , Figure 11 , Figure 12 , Figure 13 , Figure 14 As shown.

[0134] Degrees of freedom distribution: 3 degrees of freedom for the head (pitch ±45°, rotation ±180°, tilt ±30°), 2 degrees of freedom for the torso (bending ±30°, twisting ±60°), 6 degrees of freedom for each arm (shoulder 3-axis, elbow 1-axis, wrist 2-axis), and 2 degrees of freedom for the hand (opening ±180°, rotation ±30°), for a total of 17 degrees of freedom, with a repeatability accuracy of ±0.05°.

[0135] 2. Sensor System

[0136] Microphone array: The microphone model is ReSpeaker 6-Mic Array, as detailed below. Figure 15As shown, it features a 6-unit circular array with a sampling rate of 48kHz, a signal-to-noise ratio of 65dB, supports beamforming directional gain ≥12dB and noise suppression ≤-40dB, has a pickup radius of 0.5-5m, and can distinguish the location of sound sources within a 30° angle.

[0137] Vision module: The camera model is Intel RealSense D435i, as detailed below. Figure 16 As shown, the resolution is 1920×1080@30fps, the pixel count is 300,000 (640×480), the field of view is 87°×58°, the built-in IMU is an accelerometer ±16g, the gyroscope ±2000° / s, and it supports facial key point detection of 68 points in low light environment with a minimum illumination of 0.01 lux and an error ≤1.5 pixels.

[0138] Speech recognition module: Employs a programmable language control module, specifically as follows... Figure 15 As shown, the material is ABS, the power supply voltage is DC 5V, the communication method is IIC communication, and it is connected to the main controller with a 4-pin cable. The product size is 48mm×24mm. It adopts three modes: cyclic recognition mode, password mode, and button mode. The recognition range is 1 meter in a quiet environment and 30cm in a noisy environment. The recognition requirements are that a maximum of 150 candidate recognition sentences can be set for each recognition. The recognition sentences can be single characters, phrases, or short sentences, and the length cannot exceed 10 Chinese characters or 79 bytes of Pinyin string.

[0139] 3. Calculation and control unit

[0140] Raspberry Pi Development Block: Specific Design as follows Figure 17 As shown, the Raspberry Pi 5 is equipped with a BCM2712 processor, integrated Bluetooth / Wi-Fi functionality, a memory chip, a UART interface, an RP1 I / O controller, a fan header, USB 2.0 and USB 3.0 interfaces, an integrated network card, a Gigabit Ethernet port, a Type-C power port, a PoE power port, a Micro HDMI port, an RTC battery interface, a 2x4 channel MIPIDSI / CSI connector, a PCIe interface, a power switch button, a power chip, and heatsink mounting holes.

[0141] 4. Power System

[0142] Servo: Specific servo design drawings are as follows Figure 18As shown, the product weighs 57g, has dimensions of 40mm×20.14mm×51.1mm, operates at 9-12.6V, rotates at 0.20sec / 60° at 11.1V, has a stall torque of 17kg.cm at 11.1V, a rotation range of 0-240°, a no-load current of 100mA, a stall current of 1.7-2A, a servo accuracy of 0.3°, a control angle range of 0-1000° (corresponding to 0-240°), uses UART serial commands, has a communication baud rate of 115200, saves user settings even after power failure, supports angle readback, provides stall and over-temperature protection, provides parameter feedback on temperature, voltage, and position, operates in servo mode and geared motor mode, uses metal gears, and has a 20cm cable length.

[0143] Battery: Specifically, as follows Figure 19 As shown, the battery model is an 11.1V 2000mAh 10C lithium battery. The battery dimensions are 69×55.5×19mm, the battery weight is 142 grams, the standard voltage is 11.1V, the maximum voltage is 12.6V, the combination is three series and two parallel, the maximum discharge current is 20000mA, the discharge cutoff voltage is 8.1±0.24V, the rated charging current is 400mA, the maximum charging current is 1000mA, and it has overcharge protection, overcurrent protection, over-discharge protection, and short circuit protection. The battery life is 8 hours (static interaction) / 4 hours (dynamic movement), and the charging time is 3 hours.

[0144] 5. Output Module

[0145] Abdominal display: Uses an SPI OLED screen, 7-pin interface, 0.96-inch model, resolution up to 128×64, supports SPI and I... 2 The C interface can display two colors such as yellow-blue, white, and blue. The driver chip is SSD1306. The SPI interface includes power supply pins (VCC, GND), clock pin (SCLK), data pin (SDA), chip select pin (CS), reset pin (RST), data / command pin (D / C), and switch pins (BS1 / BS0) for setting the working mode, etc. The operating voltage is 3V to 5V and it is compatible with 3.3V and 5V logic levels.

[0146] Head-mounted display screen: Specifically as follows Figure 20 As shown, it uses an OLED emoji screen with a resolution of 128×64 and a brightness of 400cd / m2 (APL 25%, no glass). It can display a wide range of emojis with clear display effects, vibrant colors, and a wide viewing angle.

[0147] Speech synthesis module: specifically as follows Figure 21As shown, the highly integrated speech synthesis module can synthesize both Chinese and English speech, and integrates speech encoding and decoding functions. This module can achieve functions such as volume adjustment, intelligent speech rate adjustment, and tone adjustment, simulating the sound of a real human voice. The speech synthesis module is powered by DC 5V and uses I... 2 The device uses C-type communication, with a serial port baud rate of 9600 and a maximum communication rate of 50K. It features a storage temperature range of -55-125℃, an operating temperature range of -40-85℃, an operating current of 29-38mA, and a pin input voltage of -0.3 to 3.6V. It supports voice recognition and voice encoding, is compatible with multiple platforms such as Arduino and Raspberry Pi, and supports Chinese and English languages. The product dimensions are 48mm × 24mm.

[0148] 5.3.2 Deployment Environment Adaptation Parameters

[0149] Operating temperature: -50℃ to 80℃

[0150] Relative humidity: 30%-80% (non-condensing)

[0151] Lighting conditions: 50-10000 lux

[0152] Acoustic environment: Background noise ≤70dB (if it exceeds this, the noise reduction algorithm will be automatically activated, and the signal-to-noise ratio will be improved by 20dB).

Claims

1. A cross-modal emotional interaction system based on dynamic self-evolving anthropomorphic computing, characterized in that, Using a robot as a carrier, the system includes a multimodal fusion perception system, a self-evolution mechanism, a low-latency emotional interaction execution system, and all-scenario adaptable hardware. The multimodal fusion perception system simultaneously collects speech, image, and text data. Speech data is processed by Webrtcvad endpoint detection and SoX pre-emphasis, then OpenSMILE extracts 1582-dimensional features, which are reduced to 512 dimensions using PCA. Images are processed using a YOLOv8 and CLIP dual-model approach to extract 256-dimensional features. Text data is preprocessed and then encoded using BERT-base to generate 768-dimensional features. A cross-modal self-attention module unifies the feature dimensions and dynamically allocates weights to generate a 768-dimensional comprehensive emotional feature vector. The self-evolution mechanism... The system employs a dual-drive approach of federated learning and anomaly optimization. The federated server is deployed on an 8-core, 16GB Alibaba Cloud ECS instance with dual-machine hot standby. Edge nodes retain 3 days of data and update the model via SGD. Parameters are transmitted using MQTT+TLS1.3 encryption, with 24-hour global iteration. After the anomaly feedback closed-loop mechanism is triggered, model fine-tuning is completed within 2 hours. The low-latency emotional interaction execution system is based on 17-DOF joints and achieves motion control through DH parameter models, damped least squares, and PID+feedforward control. It constructs a dynamic interaction library containing over 1000 strategies and optimizes it using Q-Learning. The all-scenario adaptable hardware adopts a PLA topology-optimized skeleton, integrating a 6-unit microphone array, an Intel RealSense D435i vision module, and a Raspberry Pi 5 control unit, adapting to working environments from -50℃ to 80℃.

2. The system according to claim 1, characterized in that, The cross-modal self-attention module unifies the feature dimensions of each modality through linear mapping, calculates a similarity matrix to obtain an association score, and assigns a weight of 0.6 when speech features are significant, a weight of 0.5 when image features are prominent, and a weight of 0.4 when text semantics are clear. The basic weight formula is F. 融合 =0.4×F 语音 +0.35×F 图像 +0.25×F 文本 It also supports dynamic adjustments based on specific scenarios.

3. The system according to claim 1, characterized in that, In the federated learning, edge nodes are trained locally using the SGD optimization algorithm with a batch size of 32 and a learning rate of 1e-4. Training stops when the loss fluctuation of 5 consecutive batches is less than 0.

001. Data privacy is protected by local differential privacy and model compression. Pruning retains 60% of the parameters, INT8 quantization is used, and the accuracy loss is ≤1% while reducing the amount of data transmitted by 75%.

4. The system according to claim 1, characterized in that, The triggering conditions for the abnormal feedback closed-loop mechanism are: the emotion recognition accuracy rate is less than 85% for 3 consecutive hours, the user interaction interruption rate is greater than 30%, or the number of action execution failures is greater than 5 times / day. Abnormal samples are labeled with LabelStudio and included in the incremental training set after the Cohen's Kappa coefficient is ≥0.

85. The model is then fine-tuned through knowledge distillation.

5. The system according to claim 1, characterized in that, Of the 17 degrees of freedom joints, the head has 3 degrees of freedom, the torso has 2 degrees of freedom, the arms each have 6 degrees of freedom, and the hand has 2 degrees of freedom. The repeatability is ±0.05°. The inverse kinematics uses the damped least squares method with a solution accuracy of ±0.1° and a computation time of ≤5ms. The trajectory planning uses fifth-order polynomial interpolation with an acceleration ≤5m / s². 3 .

6. The system according to claim 1, characterized in that, In the PID+feedforward control, the proportional coefficient Kp = 5.0, the integral coefficient Ki = 0.1, the derivative coefficient Kd = 0.5, the position tracking error ≤ 0.5mm, and the driving torque is adjusted by a decay coefficient of 0.8 when the hand contact force exceeds 1N.

7. The system according to claim 1, characterized in that, The strategy of the dynamic interaction library includes dialogue templates, action combinations, and voice parameters. Q-Learning reinforcement learning uses the user's response positivity, which is composed of 60% of the dialogue duration and 40% of positive facial expressions, as the reward function.

8. The system according to claim 1, characterized in that, The full-scene adaptation hardware includes a ReSpeaker 6-Mic Array microphone array with a pickup radius of 0.5-5m and noise suppression ≤-40dB; a visual module with a resolution of 1920×1080@30fps and a detection error of ≤1.5 pixels for 68 facial key points in low light; and a Raspberry Pi 5 control unit equipped with an 11.1V 2000mAh lithium battery, providing 8 hours of static interaction or 4 hours of dynamic sports performance.