Robot-based interaction processing method and device, computer equipment and medium

By using multimodal perception arrays and cross-modal feature fusion technology, robots can accurately identify user emotions and generate personalized interaction strategies, solving the problems of low accuracy and insufficient intelligence in existing emotion recognition technologies and achieving more efficient user interaction.

CN122018680APending Publication Date: 2026-05-12PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-01-09
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing robot interaction technologies suffer from low accuracy and insufficient intelligence in emotion recognition, making it difficult to meet the personalized service needs of the financial insurance and medical fields, especially lacking effective solutions in multimodal data fusion and dynamic adaptability.

Method used

By using a multimodal sensing array to collect user data, and generating interaction strategies through preprocessing, cross-modal feature fusion, sentiment analysis, and scene recognition, the robot can accurately recognize user emotions and interact naturally.

Benefits of technology

It improves the accuracy of emotion recognition and the intelligence of robot interaction, enabling it to dynamically adapt to changes in user emotions and provide a personalized service experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122018680A_ABST
    Figure CN122018680A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and relates to a robot-based interaction processing method, which comprises the following steps: acquiring multi-modal original data of a user based on a multi-modal sensing array; preprocessing the multi-modal original data to obtain an emotion data set; performing cross-modal feature fusion on the emotion data set to obtain an emotion feature vector; performing sentiment analysis on the sentiment feature vector based on a preset sentiment basic model to generate sentiment prediction data; determining a current interaction scene based on a scene recognition module, and obtaining a scene constraint condition corresponding to the interaction scene; performing strategy generation on the interaction scene, the scene constraint condition and the emotion prediction data based on a basic interaction strategy library to obtain an interaction strategy; and controlling the robot to interact with the user based on the interaction strategy. The method can be applied to interactive processing business scenes in the financial science and technology field and the digital medical field, the emotion recognition precision is improved, and the robot interaction intelligence is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology and can be applied to fields such as fintech and digital healthcare, particularly to robot-based interactive processing methods, devices, computer equipment, and storage media. Background Technology

[0002] In the field of robot interaction, while embodied robot emotion modeling technology has been initially applied to real-world scenarios, it suffers from significant technical limitations. Existing technologies largely rely on single-modal data (such as voice tone, facial expressions, or text semantics) for emotion judgment, resulting in a limited dimension of emotion recognition and an inability to comprehensively capture the complex and ever-changing emotional states of users. For example, in intelligent customer service scenarios in the financial insurance sector, users may simultaneously exhibit anxiety (rapid speech) and dissatisfaction (frowning) due to cumbersome claims processes. However, traditional single-modal models can only identify one of these emotions, easily leading to misjudgments. Furthermore, the mapping relationship between existing emotion modeling results and robot interaction behavior is largely based on fixed rules (such as "switching to a soothing tone when anger is detected"), lacking dynamic adaptability and failing to address the rapid changes in user emotions and personalized needs in real-world scenarios.

[0003] This problem is equally prominent in the medical field. For example, in digital healthcare scenarios involving mental health counseling, patients may describe their emotions solely through text (such as "I haven't been feeling well lately") due to privacy concerns or expression difficulties. Traditional unimodal models cannot combine multi-dimensional information such as voice tremor and body language, resulting in an emotion recognition accuracy rate of less than 30%. At the same time, fixed-rule interaction strategies (such as forcibly pushing psychological assessment questionnaires) may exacerbate patient resistance and reduce service compliance.

[0004] The aforementioned technical limitations result in low accuracy and insufficient intelligence in embodied robots during interactions, making it difficult to meet the personalized service needs of the financial and insurance sectors and the highly sensitive interaction requirements of the medical field. Therefore, there is an urgent need for an intelligent robot interaction technology to improve the comprehensiveness of emotion recognition and the adaptability of interactive behavior, thereby optimizing user experience and service efficiency. Summary of the Invention

[0005] The purpose of this application is to propose a robot-based interaction processing method, device, computer equipment, and storage medium to solve the technical problems of low accuracy and insufficient intelligence in existing robots during interaction.

[0006] Firstly, a robot-based interaction processing method is provided, including: The system collects users' raw multimodal data based on a pre-set multimodal sensing array. The multimodal raw data is preprocessed to obtain the corresponding sentiment data set; The emotional data set is subjected to cross-modal feature fusion processing to obtain the corresponding emotional feature vector; Based on a preset sentiment model, the sentiment feature vectors are subjected to sentiment analysis processing to generate corresponding sentiment prediction data. The current interaction scenario is determined based on the preset scene recognition module, and the scene constraints corresponding to the interaction scenario are obtained. Based on a preset basic interaction strategy library, the interaction scenario, the scenario constraints, and the sentiment prediction data are processed to generate corresponding interaction strategies. Based on the interaction strategy, the robot is controlled to perform corresponding interactive processing on the user.

[0007] Secondly, a robot-based interactive processing device is provided, comprising: The acquisition module is used to acquire the user's raw multimodal data based on a preset multimodal sensing array; The preprocessing module is used to preprocess the multimodal raw data to obtain the corresponding sentiment data set; The processing module is used to perform cross-modal feature fusion processing on the emotional data set to obtain the corresponding emotional feature vector; The analysis module is used to perform sentiment analysis on the sentiment feature vector based on a preset sentiment model to generate corresponding sentiment prediction data. The acquisition module is used to determine the current interaction scenario based on the preset scene recognition module, and to acquire the scene constraint conditions corresponding to the interaction scenario. The generation module is used to perform strategy generation processing on the interaction scenario, the scenario constraints and the sentiment prediction data based on a preset basic interaction strategy library to obtain the corresponding interaction strategy. The execution module is used to control the robot to perform corresponding interactive processing on the user based on the interaction strategy.

[0008] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described robot-based interactive processing method.

[0009] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the robot-based interactive processing method described above.

[0010] In the above-mentioned robot-based interaction processing method, device, computer equipment, and storage medium, the following steps are taken: First, multimodal raw data of the user is collected based on a preset multimodal perception array; the multimodal raw data is preprocessed to obtain a corresponding emotional data set; then, cross-modal feature fusion processing is performed on the emotional data set to obtain a corresponding emotional feature vector; subsequently, emotional analysis processing is performed on the emotional feature vector based on a preset emotional basic model to generate corresponding emotional prediction data; next, the current interaction scenario is determined based on a preset scene recognition module, and the scene constraints corresponding to the interaction scenario are obtained; further, a strategy generation process is performed on the interaction scenario, the scene constraints, and the emotional prediction data based on a preset basic interaction strategy library to obtain a corresponding interaction strategy; finally, based on the interaction strategy, the robot is controlled to perform corresponding interaction processing on the user. Based on the above automated processing flow, this application constructs a full-process emotion processing method that integrates multimodal perception, data fusion, sentiment analysis, strategy generation, and interactive execution. With multimodal emotion perception and emotion model-based sentiment analysis as the core technologies, it can achieve accurate recognition of user emotions and natural interaction by the robot, effectively improving the accuracy of emotion recognition and enhancing the intelligence and adaptability of robot interaction. Attached Figure Description

[0011] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 This is a flowchart of an embodiment of the robot-based interaction processing method according to this application; Figure 3 This is a schematic diagram of the structure of one embodiment of the robot-based interactive processing device according to this application; Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation

[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0014] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0015] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0016] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0017] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0018] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.

[0019] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0020] It should be noted that the robot-based interaction processing method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the robot-based interaction processing device is generally set in the server / terminal device.

[0021] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0022] Continue to refer to Figure 2 The flowchart illustrates an embodiment of the robot-based interaction processing method according to this application. Depending on different needs, the order of the steps in the flowchart can be changed, and some steps can be omitted. The robot-based interaction processing method provided in this application embodiment can be applied to any scenario requiring robot interaction, and therefore can be applied to products in these scenarios, such as robot interaction products in the financial insurance field or the digital healthcare field. The robot-based interaction processing method includes the following steps: Step S201: Collect the user's original multimodal data based on the preset multimodal sensing array.

[0023] In this embodiment, the robot-based interaction processing method runs on an electronic device (e.g., Figure 1The server / terminal device shown can acquire the user's multimodal raw data via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wideband) connections, and other currently known or future-developed wireless connection methods. The implementing entity of this application is specifically a robot control system, which can be simply referred to as the system.

[0024] The process of configuring the aforementioned multimodal perception array includes: 1) Vision module: Selecting high-resolution, high-frame-rate cameras and installing them in appropriate locations on the embodied robot, such as the head or chest, to obtain clear and comprehensive visual information such as user facial expressions and body movements. Simultaneously, equipping it with depth sensors, such as LiDAR or structured light sensors, to acquire depth information of the scene and assist in understanding the user's interaction with the surrounding environment. 2) Auditory module: Installing a high-sensitivity microphone array capable of capturing sound signals from different directions and distances. Noise reduction technology is employed to reduce environmental noise interference, ensuring clear acquisition of user speech and voice characteristics, such as pitch, volume, and speech rate. 3) Tactile module: Installing pressure sensors and temperature sensors on areas where the robot may come into contact with the user, such as the arms and hands. Pressure sensors can sense the force of the user's touch, and temperature sensors can detect temperature changes during touch, thereby obtaining the user's tactile feedback information. 4) Physiological signal module: If conditions permit, the user can wear wearable devices, such as smart bracelets or heart rate monitors, to collect the user's physiological signals, such as heart rate and skin conductance. These physiological signals can reflect the user's emotional state, such as tension or excitement.

[0025] Furthermore, the aforementioned multimodal raw data includes collected visual data, audio data, tactile data, and physiological data. Data collection triggering mechanisms can be set according to actual needs, including: 1) Time-based triggering: Setting fixed time intervals, such as collecting data every few seconds or minutes, to obtain changes in the user's emotional state at different times. 2) Event-based triggering: Defining specific events, such as the user starting a conversation with the robot or the user making specific physical movements, triggering data collection when these events occur. For example, when a user smiles, the visual module detects this change in expression and immediately triggers data collection, recording the visual, auditory, and other multimodal data at that moment. 3) Emotional intensity-based triggering: Through preliminary emotion analysis algorithms, when the detected emotional intensity of the user reaches a certain threshold, data collection is triggered. For example, when the user's voice volume suddenly increases and the speaking speed accelerates, it may indicate that the user is in an excited state; at this time, data collection is triggered to obtain richer emotion-related information.

[0026] Furthermore, this application can be applied to interactive processing scenarios in the fintech and digital healthcare fields. For example, in the fintech field, it can be applied to intelligent customer service robots within insurance outlets (high-traffic scenarios). The scenario description includes: users entering insurance outlets to conduct business, exhibiting anxiety (e.g., frequently checking their phones, frowning) due to long waiting times or unfamiliarity with the procedures. The intelligent customer service robot needs to provide services in noisy environments (high traffic, multiple conversations causing interference). Alternatively, in the fintech field, it can also be applied to virtual claims assistants within insurance apps (family emergency scenarios). The scenario description includes: users submitting claims due to family accidents (e.g., fire, medical emergencies), experiencing low spirits (e.g., brief replies, somber tone), and needing to quickly complete the claims process to alleviate financial pressure.

[0027] In the field of digital healthcare, AI-powered guidance robots can be applied to hospital outpatient clinics (for elderly patients). The scenario involves elderly patients seeking medical attention alone who may exhibit confusion due to unfamiliarity with the hospital layout or hearing loss (e.g., repeatedly asking the same question, walking slowly). These robots need to provide clear directions in busy outpatient halls. Alternatively, AI-powered counselors within mental health apps can be used (for users with depressive tendencies). The scenario involves users conducting self-assessments through mental health apps, revealing mild depressive tendencies and low mood (e.g., delayed responses, negative language), requiring a low-stress, highly empathetic interactive experience.

[0028] Step S202: Preprocess the multimodal raw data to obtain the corresponding sentiment data set.

[0029] In this embodiment, the specific implementation process of preprocessing the multimodal raw data to obtain the corresponding emotional data set will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0030] Step S203: Perform cross-modal feature fusion processing on the emotional data set to obtain the corresponding emotional feature vector.

[0031] In this embodiment, the specific implementation process of performing cross-modal feature fusion processing on the emotional data set to obtain the corresponding emotional feature vector will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0032] Step S204: Perform sentiment analysis processing on the sentiment feature vector based on the preset sentiment model to generate corresponding sentiment prediction data.

[0033] In this embodiment, the specific implementation process of performing sentiment analysis on the sentiment feature vector based on the preset sentiment model to generate corresponding sentiment prediction data will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0034] Step S205: Determine the current interaction scenario based on the preset scene recognition module, and obtain the scene constraints corresponding to the interaction scenario.

[0035] In this embodiment, the scene recognition module uses the robot's multimodal perception array, combined with robot positioning information, environmental sensor data, and user behavior information, to identify the current interaction scene (such as a family living room, hospital ward, or classroom). For example, it determines whether the environment is at home, in an office, or in a public place, and the specific activity, such as chatting, playing games, or studying. Simultaneously, it identifies the scene constraints, such as time limits, spatial restrictions, and social rules. It can also load scene constraints corresponding to the current interaction scene from a preset scene rule library. For example, in an office environment, the interaction time should not be too long to avoid affecting work efficiency; in a public place, the interaction behavior needs to conform to social etiquette and public order; a hospital scene requires quiet; a classroom scene requires guidance, and so on.

[0036] Step S206: Based on a preset basic interaction strategy library, perform strategy generation processing on the interaction scenario, the scenario constraints, and the sentiment prediction data to obtain the corresponding interaction strategy.

[0037] In this embodiment, the above-mentioned specific implementation process of generating strategies based on the preset basic interaction strategy library for the interaction scenario, the scenario constraints and the sentiment prediction data to obtain the corresponding interaction strategy will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0038] Step S207: Based on the interaction strategy, control the robot to perform corresponding interaction processing on the user.

[0039] In this embodiment, the specific implementation process of controlling the robot to perform corresponding interactive processing on the user based on the interaction strategy will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0040] This application first collects multimodal raw data from users based on a preset multimodal perception array; then preprocesses the multimodal raw data to obtain a corresponding sentiment data set; next, it performs cross-modal feature fusion processing on the sentiment data set to obtain a corresponding sentiment feature vector; then, it performs sentiment analysis processing on the sentiment feature vector based on a preset sentiment base model to generate corresponding sentiment prediction data; subsequently, it determines the current interaction scenario based on a preset scene recognition module and obtains the scene constraints corresponding to the interaction scenario; further, it performs strategy generation processing on the interaction scenario, the scene constraints, and the sentiment prediction data based on a preset basic interaction strategy library to obtain a corresponding interaction strategy; finally, based on the interaction strategy, it controls the robot to perform corresponding interaction processing on the user. Based on the above automated processing flow, this application constructs a full-process sentiment processing method that integrates multimodal perception, data fusion, sentiment analysis, strategy generation, and interaction execution. With multimodal sentiment perception and sentiment analysis based on sentiment models as core technical support, it can achieve accurate recognition and natural interaction of user emotions by the robot, effectively improving the accuracy of sentiment recognition and enhancing the intelligence and adaptability of robot interaction.

[0041] In some optional implementations, the multimodal raw data includes visual data, audio data, tactile data, and physiological data; step S202 includes the following steps: The visual data is preprocessed based on a preset first preprocessing strategy to obtain the corresponding target visual data.

[0042] In this embodiment, the preprocessing of the aforementioned visual data includes denoising, enhancement, and normalization of the acquired image and video data. Denoising removes noise points from the image, enhancement improves the image's clarity and contrast, and normalization unifies the image's size and color range to standard values ​​for subsequent feature extraction and analysis.

[0043] The audio data is preprocessed based on a preset second preprocessing strategy to obtain the corresponding target audio data.

[0044] In this embodiment, the preprocessing of the audio data includes noise reduction, filtering, and framing. Noise reduction removes environmental noise, filtering extracts audio signals within a specific frequency range, and framing divides continuous audio signals into short frames for easier feature extraction.

[0045] The tactile data is preprocessed based on a preset third preprocessing strategy to obtain the corresponding target tactile data.

[0046] In this embodiment, the preprocessing of the aforementioned tactile data includes: calibrating and filtering the tactile data collected by the pressure sensor and temperature sensor. Calibration can eliminate sensor errors, and filtering can remove noise and interference from the data, making the data more accurate and reliable.

[0047] The physiological data is preprocessed based on a preset fourth preprocessing strategy to obtain the corresponding target physiological data.

[0048] In this embodiment, the preprocessing of the above-mentioned physiological data includes: filtering and smoothing the collected physiological signals such as heart rate and skin conductance to remove outliers and noise, and performing time alignment to ensure that the data of different modalities are synchronized in time.

[0049] Based on a preset integration strategy, the target visual data, target audio data, target tactile data, and target physiological data are integrated to obtain corresponding integrated data.

[0050] In this embodiment, the preprocessed modal data are integrated and stored according to a unified data format and standard. For example, the preprocessed visual data, audio data, tactile data, and physiological signal data are associated according to timestamps to form a data record containing multiple dimensions. Furthermore, the integrated data can be standardized to map data from different modalities to the same numerical range, facilitating subsequent cross-modal feature fusion and analysis. For example, visual features, auditory features, tactile features, and physiological signal features are all normalized to the [0, 1] interval.

[0051] The integrated data is used as the sentiment data set.

[0052] This application preprocesses visual data using a first preprocessing strategy to obtain target visual data; then preprocesses audio data using a second preprocessing strategy to obtain target audio data; next, preprocesses tactile data using a third preprocessing strategy to obtain target tactile data; subsequently, preprocesses physiological data using a fourth preprocessing strategy to obtain target physiological data; and further integrates the target visual data, target audio data, target tactile data, and target physiological data using a preprocessing strategy to obtain integrated data; finally, the integrated data is used as an emotional data set. Based on this processing flow, this application effectively improves data quality and usability, removes noise and interference, and makes the generated emotional data set more accurate and reliable by using multiple preprocessing and integration strategies to preprocess and integrate multimodal raw data. Furthermore, the final multi-dimensional emotional data set will provide high-quality input for subsequent cross-modal feature fusion and other processes.

[0053] In some optional implementations of this embodiment, step S203 includes the following steps: Invoke a preset target feature extractor; wherein, the target feature extractor includes multiple feature extractors corresponding to the modality types of the sentiment data set.

[0054] In this embodiment, dedicated feature extractors corresponding to various modalities are pre-built, including visual feature extractors, auditory feature extractors, tactile feature extractors, and physiological signal feature extractors.

[0055] Based on the target feature extractor, feature extraction is performed on the emotional data set to obtain the corresponding multimodal features.

[0056] In this embodiment, the process of feature extraction from the aforementioned emotional data set based on the target feature extractor includes: 1) Visual feature extractor: Deep learning models such as convolutional neural networks (CNNs) are used to extract features from the preprocessed visual data. CNNs can automatically learn hierarchical features in images, from low-level edge and texture features to high-level semantic features. For example, through multi-layer convolution and pooling operations, key feature points of the user's facial expressions, such as the shape of eyebrows, the degree of eye opening and closing, the outline of the mouth, etc., as well as features of body movements, such as gestures and postures, are extracted. 2) Auditory feature extractor: Traditional audio feature extraction methods such as Mel-frequency cepstral coefficients (MFCCs) are used in combination with deep learning models, such as recurrent neural networks (RNNs) or long short-term memory networks (LSTMs), to extract features from the preprocessed audio signals. MFCCs can extract the spectral features of audio, reflecting information such as pitch and timbre of speech, while RNNs or LSTMs can capture the temporal features of audio signals, such as changes in intonation and speech rate. 3) Tactile Feature Extractor: Based on the data characteristics of pressure and temperature sensors, corresponding feature extraction methods are designed. For example, for pressure sensor data, features such as pressure magnitude, trend of change, and duration can be extracted; for temperature sensor data, features such as the rate of temperature change and fluctuation range can be extracted. 4) Physiological Signal Feature Extractor: For physiological signals such as heart rate and skin conductance, features are extracted using a combination of time-domain and frequency-domain analysis. Time-domain analysis can extract statistical features such as the signal's mean, variance, and peak value, while frequency-domain analysis can convert the signal to the frequency domain using Fourier transform to extract features such as the energy distribution of different frequency components.

[0057] The multimodal features are fused using a preset cross-modal fusion engine to obtain the corresponding fused features.

[0058] In this embodiment, features extracted from each modality (i.e., multimodal features) are input into a cross-modal fusion engine. The cross-modal fusion engine can employ various fusion strategies, such as early fusion, mid-term fusion, and late-term fusion, to process the aforementioned multimodal features and obtain corresponding fused features. Early fusion concatenates the features from each modality at the input layer and then inputs them together into the subsequent model for processing; mid-term fusion performs feature fusion at the middle layer of the model; late-term fusion processes each modality feature separately and then fuses them at the output layer. An attention mechanism is introduced to assign weights to features from different modalities. The attention mechanism can automatically adjust the weights of each modality feature based on its importance to emotion recognition. For example, in a specific emotional scenario, visual features may be more important for emotion recognition; in this case, the attention mechanism will assign a higher weight to visual features, while the weights of other modal features will be relatively lower.

[0059] The processing features are optimized based on a preset optimization strategy to obtain the corresponding target features.

[0060] In this embodiment, the specific implementation process of optimizing the processing features based on the preset optimization strategy to obtain the corresponding target features will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0061] The target features are used as the emotion feature vector.

[0062] This application utilizes a pre-defined target feature extractor, which includes multiple feature extractors corresponding to the modality types of the sentiment data set. Features are then extracted from the sentiment data set based on the target feature extractor to obtain corresponding multimodal features. These multimodal features are then fused using a pre-defined cross-modal fusion engine to obtain fused features. Subsequently, the fused features are optimized using a pre-defined optimization strategy to obtain the target features. Finally, the target features are used as the sentiment feature vector. Based on this processing flow, the cross-modal feature fusion provided in this application aims to organically fuse sentiment features from different modalities to generate more accurate and comprehensive sentiment feature vectors, providing a core basis for subsequent sentiment modeling. Different modalities of data have different characteristics and advantages; by designing dedicated feature extractors, the potential information of each modality's data can be fully explored. Furthermore, the cross-modal fusion engine can effectively integrate features from various modalities and assign weights according to their importance, thereby improving the accuracy of sentiment recognition.

[0063] In some optional implementations, the process of optimizing the features based on a preset optimization strategy to obtain the corresponding target features includes the following steps: The fused features are adjusted based on a preset Bayesian inference model to obtain the corresponding first processed features.

[0064] In this embodiment, conflicting information may exist between different modalities of data. For example, visual data might show a user smiling, but auditory data might show a user speaking in a low voice and at a slow pace, potentially indicating that the user is not genuinely happy. To resolve this information conflict, a Bayesian inference model is employed. Specifically, the Bayesian inference model calculates the posterior probability of each modality feature based on prior knowledge and observational data. By comparing the posterior probabilities of different modal features, the degree of conflict between them is determined, and the data is fused and adjusted according to certain rules to obtain more accurate and consistent emotional features.

[0065] Obtain the user's basic profile.

[0066] In this embodiment, a basic profile can be obtained based on the user's name information. The user's basic profile includes information such as age, gender, cultural background, and personality traits. Different users may express the same emotions differently; for example, outgoing users may exaggerate their happiness, while introverted users may be more reserved.

[0067] Based on the basic archive, the first processing feature is personalized and calibrated to obtain the corresponding second processing feature.

[0068] In this embodiment, the personalized calibration process includes: 1. User profile analysis and feature mapping. Profile parsing: Extracting key information (age, gender, cultural background, personality traits, etc.) from the user's basic profile and converting it into quantifiable feature parameters. For example: Personality traits: Through questionnaires or behavioral analysis, personality is divided into dimensions such as "extroversion / introversion" and "rationality / emotionality," and assigned numerical values ​​(e.g., extroversion 0~1). Cultural background: Encoded as cultural group labels (e.g., "collectivist culture," "individualistic culture"), and associated with the emotional expression tendencies of that culture (e.g., collectivist culture may be more inclined to subtle expression). Feature weight allocation: Assigning weights to different profile information according to task requirements. For example, personality traits may have a higher weight in influencing emotional expression than age.

[0069] 2. Modeling the association between fused features and user profiles. Constructing a calibration rule base: Defining how different user attributes affect emotional features. For example: Extroverted users: Intensity features of happiness (e.g., amplitude, duration) increase by 20%. Introverted users: Intensity features of anger decrease by 15% to avoid overexpression. Users from high-context cultures (e.g., East Asian cultures): Expression of sadness relies more on indirect features (e.g., speech rate variations rather than direct tone). Rules can be trained based on psychological research, cultural difference theories, or user behavior data. Dynamic weight adjustment: Dynamically adjusting the weights of each dimension of the fused emotional features (e.g., multi-dimensional vectors: [pleasure, excitement, repression, ...]) according to user attributes. For example: If a user is extroverted, the weight of "pleasure" increases significantly; "excitement" increases slightly. If a user comes from a low-context culture (e.g., Western cultures), the weight of direct emotional features (e.g., tone intensity) is higher.

[0070] 3. Calibration Process. Feature Scaling and Shifting: Perform linear or nonlinear transformations on the original fused feature vector (e.g., F = [f1, f2, ..., fn]): Extroverted users: f1_new = f1 * (1 + extroversion * 0.2) (enhancing features). Introverted users: f2_new = f2 * (1 - introversion * 0.1) (suppressing features). Introduce Cultural Shifts: For example, the "anger" feature of collectivist users may be mapped to a lower threshold. Nonlinear Correction: Use piecewise functions or sigmoid functions to limit the feature range and avoid over-correction. For example, for the "excitement" feature, extroverted users may be limited to [original value, original value * 1.5]. Contextual Compensation: Further fine-tune the current features by incorporating user historical behavior data (e.g., the average intensity of past expressions of happiness). For example, if user historical data shows that their happiness expression intensity is consistently lower than the group mean, then increase it by an additional 10%.

[0071] 4. Validation and Feedback Loop. Calibration Result Validation: Input the calibrated features into the sentiment classifier and observe whether the output is consistent with the user's historical expression habits. For example, are extroverted users more likely to be identified as "intensely happy"? Adaptive Optimization: If there is a deviation between the calibrated features and the user's actual expression (e.g., the user's actual expression is more subtle than predicted), adjust the calibration rules (e.g., reduce the intensity correction coefficient for extroverted users). Multimodal Fusion Correction: If the sentiment features come from multimodal data (speech + text + facial expressions), personalize the weights of different modalities based on user preferences (e.g., "more trusting of textual expression").

[0072] 5. Output accurate sentiment feature vectors. The final calibrated feature vector F_calibrated is generated, with each dimension adjusted according to user attributes. For example: Original fused features: [Pleasure = 0.7, Excitement = 0.5, Suppression = 0.1] After calibration for extroverted users: [Pleasure = 0.84, Excitement = 0.55, Suppression = 0.1] After calibration for introverted users: [Pleasure = 0.63, Excitement = 0.45, Suppression = 0.1] The vector can be input into downstream tasks (such as sentiment classification, dialogue generation) to achieve personalized responses.

[0073] The second processed feature is used as the target feature.

[0074] This application adjusts the fused features based on a pre-defined Bayesian inference model to obtain corresponding first processed features; then, it obtains the user's basic profile; subsequently, it performs personalized calibration on the first processed features based on the basic profile to obtain corresponding second processed features; finally, it uses the second processed features as the target features. Based on the above processing flow, this application can resolve the information conflict problem between different modalities by using a Bayesian inference model, ensuring that the fused features are more consistent and reliable. Furthermore, personalized calibration considers the individual differences of different users, making the generated emotional features more consistent with the user's actual emotional expression, thereby improving the personalization and naturalness of emotional interaction.

[0075] In some alternative implementations, step S204 includes the following steps: Invoke a pre-built sentiment foundation model.

[0076] In this embodiment, a foundational emotion model can be pre-constructed using deep learning techniques, such as deep neural networks (DNNs) and deep belief networks (DBNs). This foundational emotion model takes a precise emotion feature vector obtained through multimodal fusion as input and is trained using a large amount of labeled emotion data to learn the mapping relationship between emotions and features. Furthermore, the trained foundational emotion model has the function of mapping emotions to a three-dimensional space of "pleasure-arousal-dominance." Pleasure represents the positive or negative degree of the emotion, arousal represents the intensity of the emotion, and dominance represents the sense of control or dominance. By quantifying emotions across these three dimensions, a more comprehensive and detailed description of the user's emotional state can be achieved.

[0077] Based on the aforementioned emotion foundation model, the emotion feature vector is processed by emotion mapping to obtain the corresponding emotion state data.

[0078] In this embodiment, the aforementioned emotional feature vector can be input into the aforementioned emotional foundation model, and then the input emotional feature vector can be quantified using this emotional feature vector to map to an emotional quantification value corresponding to the three-dimensional space of "pleasure-arousal-dominance," i.e., the aforementioned emotional state data. Here, the current emotional quantification value represents the specific numerical values ​​of the user's pleasure, arousal, and dominance at the current moment.

[0079] Based on a preset emotional state transition module, the emotional state data is analyzed for trends to obtain corresponding emotional trend prediction data.

[0080] In this embodiment, an emotional state transition module is pre-built. This module combines time series analysis algorithms to capture emotional change trends by comparing historical emotional features with current features. Specifically, the emotional state transition module calculates a probability matrix of emotional state transitions based on the user's historical emotional data and current emotional state data. By analyzing the probability matrix, the likelihood of a user's emotions shifting from one state to another can be understood, thereby predicting the future development trend of the user's emotions and obtaining corresponding emotional trend prediction data. The trend describes the direction of the user's emotions over a future period. For example, if the user is currently in a state of high pleasure and moderate arousal, the probability matrix can predict that the user may maintain this state or gradually transition to a more excited state in the coming period.

[0081] The emotional state data and the emotional trend prediction data are integrated to obtain the corresponding integrated emotional data.

[0082] In this embodiment, the emotional state data and emotional trend prediction data can be integrated and processed, and the resulting integrated emotional data can be used as the corresponding emotional prediction data.

[0083] The integrated sentiment data is used as the sentiment prediction data.

[0084] In this embodiment, the parameters of the basic emotion model can also be dynamically adjusted based on information in the user's emotion preference database. For example, if it is found that the user's emotional response to a certain interaction method is more positive, the weight of that interaction method in the model can be appropriately increased to make the model more consistent with the user's emotional preferences.

[0085] This application utilizes a pre-built emotional foundation model; it then performs emotional mapping processing on emotional feature vectors based on the emotional foundation model to obtain corresponding emotional state data; subsequently, it performs trend analysis on the emotional state data based on a preset emotional state transition module to obtain corresponding emotional trend prediction data; finally, it integrates the emotional state data and emotional trend prediction data to obtain corresponding integrated emotional data; and finally, it uses the integrated emotional data as emotional prediction data. Based on the above processing flow, this application quantifies complex emotional states into a three-dimensional space of "pleasure-arousal-dominance" by using the emotional foundation model, making emotional states more intuitive and operable. Furthermore, the emotional state transition module can capture the dynamic changes in user emotions and predict their future development trends, providing a basis for timely adjustments to interaction strategies. This enables automatic and intelligent completion of emotional analysis processing of user emotional feature vectors, ensuring the accuracy of the generated emotional prediction data.

[0086] In some optional implementations of this embodiment, step S206 includes the following steps: Invoke the preset strategy to generate the engine.

[0087] In this embodiment, the aforementioned strategy generation engine is a pre-built automated engine that assists in generating interactive strategies for robots.

[0088] Based on the strategy generation engine, the basic interaction strategy library is matched with the sentiment prediction data to obtain the corresponding basic interaction strategy.

[0089] In this embodiment, the strategy generation engine matches corresponding basic interaction strategies from the basic interaction strategy library based on the obtained emotion prediction data. This basic interaction strategy library is a pre-built database that defines the mapping relationship between emotional states and basic strategies. For example: high pleasure - high arousal → lively strategies (such as high-frequency interaction, exaggerated feedback, multimodal stimulation); low pleasure - low arousal → companionship strategies (such as gentle tone, slow pace, empathetic response); high anxiety → soothing strategies (such as guided deep breathing, recommendations of soothing music). Furthermore, the basic interaction strategy library needs to cover common emotion combinations and reserve logic for handling "mixed emotions" (such as prioritizing anxiety when pleasure + anxiety).

[0090] Furthermore, by inputting the aforementioned sentiment prediction data, including current sentiment quantification and sentiment trend prediction data, into the decision tree, the closest basic interaction strategy is matched from the aforementioned basic interaction strategy library. For example: if the pleasure level > 0.8 and the arousal level > 0.7, the "lively strategy" is selected. If the pleasure level < 0.3 and the arousal level < 0.3 and the trend is "continuously declining", the "companionship strategy" is selected and the "emotional support" sub-process is triggered.

[0091] Based on the interaction scenario and the scenario constraints, the basic interaction strategy is adjusted to obtain the corresponding specified interaction strategy.

[0092] In this embodiment, the basic interaction strategy obtained by matching can be adapted to the scenario based on the above-mentioned interaction scenario and scenario constraints to adjust the basic interaction strategy and use the obtained specified interaction strategy as the corresponding interaction strategy.

[0093] The scene adaptation processing includes: 1. Parameter coverage and modality conversion. Iterate through the basic interaction strategy parameters and cover conflicting items with scene constraints: for example, basic interaction strategy {speech rate: "slow", output modality: "voice"}, constraint rule {disable voice: true} → force the output modality to text. 2. Dynamic threshold adjustment. Adjust the range of strategy parameters according to scene characteristics: basic interaction strategy {volume: 50%}, constraint rule {maximum volume: 20%} → adjust to 20%. 3. Behavior supplementation and trimming. Add necessary scene behaviors or remove conflicting behaviors: constraint rule {preferred modality: "text + haptic"} → add haptic feedback parameters to the basic interaction strategy. If the scene is a driving environment, remove highly interfering behaviors such as complex gestures. 4. Conflict rollback mechanism. If the adjusted strategy violates the core constraints (e.g., disabling voice but the strategy depends on voice), trigger rollback: rollback scheme: enable the backup modality (e.g., text-to-speech requires user intervention).

[0094] Use the specified interaction strategy as the interaction strategy.

[0095] This application utilizes a pre-defined strategy generation engine. Based on this engine, it matches a basic interaction strategy library with sentiment prediction data to obtain corresponding basic interaction strategies. Then, it adjusts these basic interaction strategies based on the interaction scenario and its constraints to obtain a specified interaction strategy. Finally, this specified interaction strategy is used as the final interaction strategy. By employing this process, the application obtains basic interaction strategies by matching them with sentiment prediction data using a strategy generation engine. It then adjusts these strategies based on the interaction scenario and its constraints, using the resulting specified interaction strategy as the final interaction strategy. This ensures that the output interaction strategy comprehensively considers both scenario and sentiment factors, improving the quality and satisfaction of user-robot interaction and enhancing user trust and acceptance of the robot.

[0096] In some optional implementations of this embodiment, step S207 includes the following steps: Collect emotional feedback data corresponding to the user.

[0097] In this embodiment, user emotional feedback data in different scenarios is collected in advance, including users' preferences for different interactive content and their emotional reactions. A user emotional preference database is established through the analysis and mining of this data.

[0098] An emotional preference database corresponding to the user is constructed based on the emotional feedback data.

[0099] In this embodiment, the user's sentiment preference database can record the user's sentiment preferences for different types of topics, interaction methods, robot behaviors, etc. For example, some users may prefer humorous and witty interaction methods, while others may prefer serious and earnest communication; some users are interested in technology-related topics, while others are more concerned with culture and art-related topics.

[0100] The interaction strategy is personalized and optimized based on the sentiment preference library to obtain the optimized target interaction strategy.

[0101] In this embodiment, the interaction strategy can be personalized and optimized by combining user types or user emotional preferences in the emotion preference database to obtain an optimized target interaction strategy. Specifically, for elderly users, the voice volume and speaking speed can be slowed down, and for children, physical interaction actions can be added; or, if the user emotion preference database shows that the user prefers a humorous and witty communication style, humorous elements can be appropriately added to the basic strategy, such as telling jokes or using witty language; if the user has a strong interest in a specific topic, content related to that topic can be added during the interaction process.

[0102] Based on the target interaction strategy, the robot is controlled to perform corresponding interaction processing on the user.

[0103] In this embodiment, after obtaining the target interaction strategy, the robot can be controlled to perform matching interaction processing on the user according to the strategy content of the target interaction strategy.

[0104] This application collects emotional feedback data corresponding to users; then, based on this data, it constructs an emotional preference database corresponding to each user; subsequently, it personalizes and optimizes interaction strategies based on this database to obtain an optimized target interaction strategy; finally, based on this target interaction strategy, it controls the robot to perform corresponding interactive processing on the user. Based on this process, this application constructs an emotional preference database based on collected emotional feedback data corresponding to users, and then personalizes and optimizes interaction strategies using this database. This makes the generated target interaction strategy more in line with the user's actual needs and emotional characteristics. Furthermore, by controlling the robot to perform interactive processing on the user based on the target interaction strategy, this application can effectively improve the quality and satisfaction of the interaction between the user and the robot, and enhance the user's trust and acceptance of the robot.

[0105] In some optional implementations, this application also features interactive feedback collection and model iteration capabilities, the specific implementation process of which includes: 1. Collect user interaction feedback data. User interaction feedback data, including immediate and delayed feedback, is collected through the robot's multimodal perception array. Immediate feedback refers to the user's real-time reactions during the interaction, such as facial expressions, tone of voice, and body language. For example, when the robot asks a question, a user's immediate thoughtful expression or nod in agreement can be considered immediate feedback data. Delayed feedback refers to feedback given by the user some time after the interaction ends, such as user evaluations and suggestions regarding the interaction experience. Delayed feedback data can be collected through questionnaires, online reviews, etc. For example, after the interaction, the robot can pop up a small window inviting the user to rate and leave a comment about the interaction.

[0106] 2. Generate Evaluation Metrics. Input the collected feedback data into the performance evaluation module to generate evaluation metrics. Evaluation metrics may include emotion recognition accuracy, interaction satisfaction, and task completion rate. For emotion recognition accuracy, it can be calculated by comparing the user's emotional state recognized by the robot with the user's actual emotional state; interaction satisfaction can be quantitatively analyzed through user ratings and comments on the interaction experience; task completion rate can be evaluated based on whether the user successfully completes the preset task during the interaction.

[0107] 3. Targeted Adjustment of Preceding Parameters. The feedback optimization module adjusts the parameters of preceding stages based on the evaluation results. For example, if the emotion recognition accuracy is low, the feature extractor parameters and attention mechanism weights in the cross-modal feature fusion and calibration stage can be adjusted to improve the accuracy of feature extraction and fusion effect. If the interaction satisfaction is low, the issues mentioned in user feedback can be analyzed, and the strategy generation rules and personalized optimization methods in the interaction strategy generation stage can be adjusted. For cases with low task completion rates, the model parameters in the dynamic emotion modeling and state tracking stage can be checked to see if they are reasonable and can accurately predict user emotion changes and needs, and corresponding adjustments can be made.

[0108] 4. Output the optimized system parameter set to drive iterative system upgrades. The adjusted parameter set is organized and optimized, outputting an optimized system parameter set. This parameter set covers various stages, including multimodal emotion data acquisition, cross-modal feature fusion and calibration, dynamic emotion modeling and state tracking, and emotion-driven personalized interaction strategy generation. The optimized system parameter set is then used to drive iterative system upgrades, updating the system's model and algorithms, enabling the system to operate more accurately and efficiently in subsequent interactions, and continuously improving its emotion modeling and interaction capabilities.

[0109] Interactive feedback collection and model iteration are key steps in achieving continuous system optimization, aiming to ensure the continuous improvement of emotion modeling and interaction capabilities. Collecting user interaction feedback data allows us to understand the system's performance and existing problems during actual operation. Generating evaluation metrics enables quantitative assessment of system performance, providing a basis for subsequent optimization. Targeted adjustments to parameters in preceding stages can identify system weaknesses based on evaluation results, allowing for targeted improvements and optimizations. Outputting the optimized system parameter set and driving iterative system upgrades enables the system to continuously adapt.

[0110] In some alternative implementations, the user information obtained is subject to user consent and complies with relevant laws and policies.

[0111] Furthermore, any software tools or components not belonging to our company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.

[0112] Furthermore, the core innovations of this application can be summarized in the following three points: 1. Multimodal adaptive fusion perception technology: The innovative design of the fusion engine combines attention mechanism and Bayesian inference to achieve dynamic weight allocation and conflict resolution of visual, auditory, tactile and physiological signals. Combined with user personalized feature calibration, it solves the problem of low accuracy of single-modal perception and improves the comprehensiveness and accuracy of emotional feature extraction.

[0113] 2. Three-dimensional dynamic emotion modeling system: Breaking through the limitations of traditional fixed-category modeling, it constructs a three-dimensional emotion model of "quantitative description + trend prediction + personalized adaptation", introduces time series analysis to capture dynamic changes in emotions, and combines user preference library to realize personalized updates of the model, thus solving the problems of static and rigid emotion modeling and lack of adaptability.

[0114] 3. Closed-loop interaction mechanism of emotion-scenario-personality collaboration: Construct a full-link closed-loop system of "modeling-strategy-feedback-iteration", deeply integrate emotional models with scenario constraints and user personality to generate interaction strategies, and optimize the parameters of the entire process through feedback data to solve the problems of rigid interaction and inability to continuously optimize existing technologies, so as to achieve the naturalness and evolution of emotional interaction.

[0115] Furthermore, compared to existing technologies, this application achieves significant improvements in the accuracy of emotion perception, the naturalness of interaction, and scene adaptability. Specific benefits are as follows: 1. Improve the accuracy of emotion perception and modeling, and reduce interaction misjudgments. Multimodal fusion perception technology improves the comprehensiveness of emotion feature extraction by more than 70%, and the three-dimensional dynamic modeling system enables accurate quantification of continuous emotional states. The accuracy of emotion recognition has increased from below 65% in existing technologies to over 90%. Taking medical care scenarios as an example, the robot can accurately identify the patient's transitional emotional state of "suppressing sadness" through multimodal data such as "facial tear stains (visual) + choked voice (auditory) + increased heart rate (physiological)," rather than misjudging it as "calm" or "anger." The modeling error is reduced by 60%, providing a guarantee for accurate subsequent interactions.

[0116] 2. Optimize emotional interaction experience and enhance user emotional resonance. The interaction strategy of emotion-scenario-personality coordination improves the adaptability of robot interaction behavior by 80%, fundamentally improving the problem of stiff interaction. In the home service scenario, for teenagers' "disappointment after failing an exam", the robot can combine "low-wake-up companionship (scenario: quiet study environment) + personalized encouragement (using the example of the user's favorite basketball star) + gentle shoulder pat comfort (tactile interaction)" strategy. The user's emotional resonance is improved by 75% compared with existing technologies, and the willingness to actively interact is increased by more than 3 times, effectively solving the problem of "interaction without warmth" in emotional robots.

[0117] 3. Expand the scope of scenario adaptation and reduce industry application costs. Dynamic adaptive models and closed-loop optimization mechanisms enable the system to quickly adapt to different scenarios and user groups without requiring extensive customized development for specific scenarios. For example, when the same robot switches from a home care scenario to a school psychological counseling scenario, it only needs to iteratively optimize parameters based on feedback data, and the adaptation can be completed within 24 hours, reducing scenario switching costs by 80%. For users of different ages and personalities, the system can automatically adjust modeling and interaction strategies without the need for additional dedicated model development, significantly lowering the industry application threshold and promotion costs of emotional robots.

[0118] 4. Achieving continuous evolution capabilities to adapt to long-term interaction needs. The closed-loop iteration mechanism enables the system to continuously optimize based on changes in users' emotional expression habits, and the accuracy of emotional modeling under long-term interaction continues to improve. In long-term home care scenarios, as the interaction time with users increases, the system's ability to recognize users' "implicit emotional signals" (such as habitual frowning indicating thinking rather than displeasure) continuously improves, the personalization of interaction strategies continues to increase, and user satisfaction after 3 months of use is more than 50% higher than the initial state, perfectly adapting to the long-term and in-depth needs of emotional interaction.

[0119] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0120] It should be emphasized that, to further ensure the privacy and security of the above interaction strategy, the interaction strategy can also be stored in a node of a blockchain.

[0121] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.

[0122] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0123] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0124] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0125] Further reference Figure 3 As a response to the above Figure 2 The implementation of the method shown in this application provides an embodiment of a robot-based interactive processing device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0126] like Figure 3 As shown, the robot-based interactive processing device 300 described in this embodiment includes: a data acquisition module 301, a preprocessing module 302, a processing module 303, an analysis module 304, an acquisition module 305, a generation module 306, and an execution module 307. Wherein: The acquisition module 301 is used to acquire the user's multimodal raw data based on a preset multimodal sensing array; Preprocessing module 302 is used to preprocess the multimodal raw data to obtain the corresponding sentiment data set; Processing module 303 is used to perform cross-modal feature fusion processing on the emotional data set to obtain the corresponding emotional feature vector; Analysis module 304 is used to perform sentiment analysis processing on the sentiment feature vector based on a preset sentiment model to generate corresponding sentiment prediction data; The acquisition module 305 is used to determine the current interaction scene based on the preset scene recognition module, and to acquire the scene constraint conditions corresponding to the interaction scene; The generation module 306 is used to perform strategy generation processing on the interaction scenario, the scenario constraints and the sentiment prediction data based on a preset basic interaction strategy library to obtain the corresponding interaction strategy. The execution module 307 is used to control the robot to perform corresponding interactive processing on the user based on the interaction strategy.

[0127] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the robot-based interaction processing method in the aforementioned embodiments, and will not be repeated here.

[0128] In some optional implementations of this embodiment, the multimodal raw data includes visual data, audio data, tactile data, and physiological data; the preprocessing module 302 includes: The first preprocessing submodule is used to preprocess the visual data based on a preset first preprocessing strategy to obtain the corresponding target visual data. The second preprocessing submodule is used to preprocess the audio data based on a preset second preprocessing strategy to obtain the corresponding target audio data. The third preprocessing submodule is used to preprocess the tactile data based on a preset third preprocessing strategy to obtain the corresponding target tactile data. The fourth preprocessing submodule is used to preprocess the physiological data based on a preset fourth preprocessing strategy to obtain the corresponding target physiological data. The first integration submodule is used to integrate the target visual data, target audio data, target tactile data and target physiological data based on a preset integration strategy to obtain corresponding integrated data; The first determining submodule is used to treat the integrated data as the emotional data set.

[0129] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the robot-based interaction processing method in the aforementioned embodiments, and will not be repeated here.

[0130] In some optional implementations of this embodiment, the processing module 303 includes: The first calling submodule is used to call a preset target feature extractor; wherein, the target feature extractor includes multiple feature extractors corresponding to the modality types of the sentiment data set; The extraction submodule is used to extract features from the emotional data set based on the target feature extractor to obtain the corresponding multimodal features; The fusion submodule is used to perform fusion processing on the multimodal features based on a preset cross-modal fusion engine to obtain the corresponding fused features; The first optimization submodule is used to optimize the processing features based on a preset optimization strategy to obtain the corresponding target features; The second determining submodule is used to use the target feature as the emotion feature vector.

[0131] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the robot-based interaction processing method in the aforementioned embodiments, and will not be repeated here.

[0132] In some optional implementations of this embodiment, the first optimization submodule includes: The adjustment unit is used to adjust the fused features based on a preset Bayesian inference model to obtain the corresponding first processed features; The acquisition unit is used to acquire the user's basic profile; A calibration unit is used to perform personalized calibration processing on the first processing feature based on the basic file to obtain the corresponding second processing feature; A determining unit is used to take the second processed feature as the target feature.

[0133] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the robot-based interaction processing method in the aforementioned embodiments, and will not be repeated here.

[0134] In some optional implementations of this embodiment, the analysis module 304 includes: The second calling submodule is used to call the pre-built sentiment foundation model; The mapping submodule is used to perform emotion mapping processing on the emotion feature vector based on the emotion base model to obtain the corresponding emotion state data. The analysis submodule is used to perform trend analysis on the emotional state data based on the preset emotional state transition module to obtain corresponding emotional trend prediction data. The second integration submodule is used to integrate the emotional state data and the emotional trend prediction data to obtain the corresponding integrated emotional data. The third determining submodule is used to use the integrated emotion data as the emotion prediction data.

[0135] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the robot-based interaction processing method in the aforementioned embodiments, and will not be repeated here. In some optional implementations of this embodiment, the generation module 306 includes: The third submodule is used to invoke the preset strategy generation engine; The matching submodule is used to perform strategy matching on the basic interaction strategy library based on the strategy generation engine and the sentiment prediction data to obtain the corresponding basic interaction strategy. The adjustment submodule is used to adjust the basic interaction strategy based on the interaction scenario and the scenario constraints to obtain the corresponding specified interaction strategy. The fourth determining submodule is used to use the specified interaction strategy as the interaction strategy.

[0136] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the robot-based interaction processing method in the aforementioned embodiments, and will not be repeated here.

[0137] In some optional implementations of this embodiment, the execution module 307 includes: The collection submodule is used to collect emotional feedback data corresponding to the user; A submodule is constructed to build an emotional preference library corresponding to the user based on the emotional feedback data; The second optimization submodule is used to perform personalized optimization of the interaction strategy based on the sentiment preference library to obtain the optimized target interaction strategy. The execution submodule is used to control the robot to perform corresponding interaction processing on the user based on the target interaction strategy.

[0138] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the robot-based interaction processing method in the aforementioned embodiments, and will not be repeated here. To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.

[0139] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0140] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0141] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may also include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for robot-based interactive processing methods. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.

[0142] In some embodiments, the processor 42 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, for example, to execute computer-readable instructions of the robot-based interactive processing method.

[0143] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.

[0144] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the robot-based interactive processing method described above.

[0145] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0146] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

Claims

1. A robot-based interactive processing method, characterized in that, Includes the following steps: The system collects users' raw multimodal data based on a pre-set multimodal sensing array. The multimodal raw data is preprocessed to obtain the corresponding sentiment data set; The emotional data set is subjected to cross-modal feature fusion processing to obtain the corresponding emotional feature vector; Based on a preset sentiment model, the sentiment feature vectors are subjected to sentiment analysis processing to generate corresponding sentiment prediction data. The current interaction scenario is determined based on the preset scene recognition module, and the scene constraints corresponding to the interaction scenario are obtained. Based on a preset basic interaction strategy library, the interaction scenario, the scenario constraints, and the sentiment prediction data are processed to generate corresponding interaction strategies. Based on the interaction strategy, the robot is controlled to perform corresponding interactive processing on the user.

2. The robot-based interactive processing method according to claim 1, characterized in that, The multimodal raw data includes visual data, audio data, tactile data, and physiological data; the step of preprocessing the multimodal raw data to obtain the corresponding emotional data set specifically includes: The visual data is preprocessed based on a preset first preprocessing strategy to obtain the corresponding target visual data. The audio data is preprocessed based on a preset second preprocessing strategy to obtain the corresponding target audio data; The tactile data is preprocessed based on a preset third preprocessing strategy to obtain the corresponding target tactile data; The physiological data is preprocessed based on a preset fourth preprocessing strategy to obtain the corresponding target physiological data. Based on a preset integration strategy, the target visual data, target audio data, target tactile data, and target physiological data are integrated to obtain corresponding integrated data; The integrated data is used as the sentiment data set.

3. The robot-based interactive processing method according to claim 1, characterized in that, The step of performing cross-modal feature fusion processing on the sentiment data set to obtain the corresponding sentiment feature vector specifically includes: Invoke a preset target feature extractor; wherein, the target feature extractor includes multiple feature extractors corresponding to the modality types of the sentiment data set; Based on the target feature extractor, feature extraction is performed on the emotional data set to obtain the corresponding multimodal features; The multimodal features are fused based on a preset cross-modal fusion engine to obtain the corresponding fused features; The processing features are optimized based on a preset optimization strategy to obtain the corresponding target features; The target features are used as the emotion feature vector.

4. The robot-based interactive processing method according to claim 3, characterized in that, The step of optimizing the processed features based on a preset optimization strategy to obtain the corresponding target features specifically includes: The fused features are adjusted based on a preset Bayesian inference model to obtain the corresponding first processed features; Obtain the user's basic profile; Based on the basic archive, the first processing feature is personalized and calibrated to obtain the corresponding second processing feature; The second processed feature is used as the target feature.

5. The robot-based interactive processing method according to claim 1, characterized in that, The step of performing sentiment analysis on the sentiment feature vector based on a preset sentiment model to generate corresponding sentiment prediction data specifically includes: Invoke a pre-built sentiment model; Based on the aforementioned emotion foundation model, the emotion feature vector is processed by emotion mapping to obtain the corresponding emotion state data. Based on the preset emotional state transfer module, the emotional state data is analyzed for trends to obtain corresponding emotional trend prediction data. The emotional state data and the emotional trend prediction data are integrated to obtain corresponding integrated emotional data. The integrated sentiment data is used as the sentiment prediction data.

6. The robot-based interactive processing method according to claim 1, characterized in that, The step of generating corresponding interaction strategies by processing the interaction scenario, the scenario constraints, and the sentiment prediction data based on a preset basic interaction strategy library specifically includes: Invoke the preset strategy to generate the engine; Based on the strategy generation engine, the basic interaction strategy library is matched with the sentiment prediction data to obtain the corresponding basic interaction strategy. Based on the interaction scenario and the scenario constraints, the basic interaction strategy is adjusted to obtain the corresponding specified interaction strategy. Use the specified interaction strategy as the interaction strategy.

7. The robot-based interactive processing method according to claim 1, characterized in that, The step of controlling the robot to perform corresponding interactive processing on the user based on the interaction strategy further includes: Collect emotional feedback data corresponding to the user; Based on the emotional feedback data, an emotional preference database corresponding to the user is constructed; Based on the sentiment preference library, the interaction strategy is personalized and optimized to obtain the optimized target interaction strategy. Based on the target interaction strategy, the robot is controlled to perform corresponding interaction processing on the user.

8. A robot-based interactive processing device, characterized in that, include: The acquisition module is used to acquire the user's raw multimodal data based on a preset multimodal sensing array; The preprocessing module is used to preprocess the multimodal raw data to obtain the corresponding sentiment data set; The processing module is used to perform cross-modal feature fusion processing on the emotional data set to obtain the corresponding emotional feature vector; The analysis module is used to perform sentiment analysis on the sentiment feature vector based on a preset sentiment model to generate corresponding sentiment prediction data. The acquisition module is used to determine the current interaction scenario based on the preset scene recognition module, and to acquire the scene constraint conditions corresponding to the interaction scenario. The generation module is used to perform strategy generation processing on the interaction scenario, the scenario constraints and the sentiment prediction data based on a preset basic interaction strategy library to obtain the corresponding interaction strategy. The execution module is used to control the robot to perform corresponding interactive processing on the user based on the interaction strategy.

9. A computer device, characterized in that, It includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the robot-based interactive processing method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the robot-based interactive processing method as described in any one of claims 1 to 7.