Robot-based emotion perception collaborative processing method and system
By acquiring and preprocessing multimodal emotion feature data in the robot, dynamically adjusting the lightweight model cluster, and coordinating processing at the edge and cloud, the problems of low efficiency of multimodal data fusion, lack of dynamic adaptation of computing resources, and insufficient privacy protection in robot emotion perception are solved, thereby improving the real-time performance and accuracy of emotion perception.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SUPER LOVE (HANGZHOU) INTELLIGENT TECHNOLOGY CO LTD
- Filing Date
- 2026-04-07
- Publication Date
- 2026-07-03
AI Technical Summary
Existing robot emotion perception technologies suffer from problems such as low efficiency of multimodal data fusion, lack of dynamic adaptation of computing resources, insufficient privacy protection, and lack of edge-cloud collaboration logic, resulting in low accuracy and efficiency of emotion perception.
By acquiring and preprocessing multimodal emotion feature data, dynamically adjusting the lightweight model cluster, combining user intention decryption levels for privacy protection, and performing collaborative processing at the edge and cloud, refined analysis of emotion features and model optimization are achieved.
It improves the real-time performance and accuracy of robot emotion perception and the iterative optimization capability of the model, while taking into account both dynamic adaptation of computing power and protection of user privacy.
Smart Images

Figure CN122333352A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics technology, and more specifically, to a robot-based method and system for collaborative processing of emotion perception. Background Technology
[0002] With the rapid development of artificial intelligence and robotics, and the increasing prominence of emotional issues among people today, robots with emotion perception capabilities have a broad application market. Emotion perception is one of the core technologies for robots to achieve natural human-computer interaction. Currently, robot emotion perception technologies mostly rely on single-modal data recognition or independent processing on the edge or cloud using fixed models, which has many technical shortcomings and application pain points.
[0003] At the data processing level, existing technologies lack precise time synchronization and standardized preprocessing methods for collecting multimodal emotional feature data such as user voice, physiological and visual data. The fusion efficiency of multimodal data is low, and the accuracy of emotional feature extraction is easily reduced due to data misalignment and noise interference. At the same time, there is a lack of targeted privacy desensitization mechanisms in the transmission and processing of emotional data, and sensitive data such as user physiological and visual data are at risk of leakage, making it difficult to match the privacy protection needs of different users.
[0004] At the level of computing power and model scheduling, the computing power resources at the edge of the robot are limited, and in actual applications, the processor utilization and memory usage rate change dynamically with the task. Existing technologies have not achieved dynamic adaptation between the model and the computing power at the edge. Fixed models are prone to processing delays caused by insufficient computing power, or resource waste caused by computing power redundancy, and cannot take into account both the real-time performance of emotion perception and the efficiency of computing power utilization.
[0005] In addition, the existing robot emotion perception system lacks standardized collaborative logic in its edge-cloud interaction, and does not differentiate between basic emotions and negative emotions, resulting in the ineffective use of cloud computing resources and low overall processing efficiency.
[0006] To address the aforementioned issues, there is an urgent need for a collaborative processing method for robot emotion perception that can achieve standardized processing of multimodal emotion data, dynamic adaptation of computing power, graded protection of privacy, and graded collaborative analysis between the edge and cloud, in order to improve the overall performance of robot emotion perception and meet the practical application needs of human-computer interaction. Summary of the Invention
[0007] The purpose of this application is to provide a robot-based collaborative processing method and system for emotion perception. Through effective extraction of emotion features, dynamic adjustment of effective lightweight models, desensitization of multi-dimensional emotion feature vectors, analysis and execution of refined emotion analysis data, and optimization of models, the technology of collaborative processing of emotion perception between edge devices and the cloud is realized.
[0008] This application also provides a robot-based collaborative processing method for emotion perception, including the following steps: The system acquires users' emotional feature data and preprocesses it to obtain effective emotional feature data. Acquire real-time computing load data at the robot edge and calculate the comprehensive computing load value. Adjust the preset dynamic lightweight model cluster according to the comprehensive computing load value to obtain an effective lightweight model. Effective emotional feature data is processed through an effective lightweight model to obtain multi-dimensional emotional feature vectors, and de-identification is performed according to the obtained user intention de-identification level to obtain de-identified multi-dimensional emotional feature vectors. Preliminary emotion categories are obtained by processing the desensitized multi-dimensional emotion feature vectors through a pre-set lightweight classification model. Preliminary emotion categories and desensitized multi-dimensional emotion feature vectors that meet the preset conditions are uploaded to the cloud for analysis to obtain refined emotion analysis data; The refined sentiment analysis data is distributed to the edge devices, and user sentiment characteristic variable data is collected and uploaded to the cloud for model optimization.
[0009] Optionally, in the robot-based emotion perception collaborative processing method described in this application, the step of acquiring the user's emotion feature data and preprocessing the emotion feature data to obtain effective emotion feature data specifically includes: Acquire users' emotional characteristics data, including voice data, physiological data, and visual data; Timestamps are added to and aligned for voice, physiological and visual data using dual calibration technology of NTP network time protocol and local hardware clock; After preprocessing the emotional feature data, effective emotional feature data is obtained, including effective voice data, effective physiological data, and effective visual data. The preprocessing includes framing and windowing the speech data, filtering and denoising the physiological data, and normalizing the size of the visual data.
[0010] Optionally, in the robot-based emotion perception collaborative processing method described in this application, the step of acquiring real-time computing power load data at the robot's edge, calculating a comprehensive computing power load value, and adjusting a preset dynamic lightweight model cluster based on the comprehensive computing power load value to obtain an effective lightweight model specifically includes: Acquire real-time computing load data at the robot's edge, including processor utilization, memory usage, and task difficulty data; The real-time computing load data is normalized and then weighted by preset weighting coefficients to obtain the comprehensive computing load value. The comprehensive computing load value is compared with the preset model scheduling threshold to obtain the corresponding effective lightweight model.
[0011] Optionally, in the robot-based collaborative processing method for emotion perception described in this application, the step of processing effective emotion feature data through an effective lightweight model to obtain a multi-dimensional emotion feature vector, and performing desensitization processing according to the obtained user intention de-identification level to obtain a desensitized multi-dimensional emotion feature vector, specifically includes: A feature-level serial fusion strategy is adopted, firstly extracting the corresponding single-modal sentiment features through the single-modal lightweight model in the effective lightweight model; Then, the lightweight fusion model in the effective lightweight model is used to achieve deep fusion of multimodal features to generate multidimensional emotion feature vectors; Obtain the user's intended level of privacy, including Level 1, Level 2, or Level 3 privacy. Desensitization is performed based on the user's desired level of anonymization to obtain a desensitized multi-dimensional emotional feature vector.
[0012] Optionally, in the robot-based collaborative processing method for emotion perception described in this application, the step of obtaining a preliminary emotion category by processing the desensitized multi-dimensional emotion feature vector through a preset lightweight classification model specifically includes: The desensitized multi-dimensional emotion feature vector is input into a preset lightweight classification model for processing to obtain the emotion category and the corresponding probability value. Emotional categories include basic emotions, mild negative emotions, moderate negative emotions, or extreme emotions; The emotion categories are sorted in descending order of probability value, and the emotion category with the highest probability value is taken as the initial emotion category.
[0013] Optionally, in the robot-based collaborative processing method for emotion perception described in this application, the step of uploading preliminary emotion categories and desensitized multi-dimensional emotion feature vectors that meet preset conditions to the cloud for analysis to obtain refined emotion analysis data specifically includes: Determine the initial emotion category. If the initial emotion category is a basic emotion, then process it at the edge to obtain the corresponding basic response strategy. If the initial emotion category is any of mild negative, moderate negative, or extreme emotion, then the initial emotion category and the desensitized multi-dimensional emotion feature vector will be uploaded to the cloud. The desensitized multi-dimensional emotional feature vectors are analyzed using a pre-set deep healing model under a federated learning framework, combined with a pre-set user emotional profile knowledge graph, to obtain refined emotional analysis data. Refined emotion analysis data includes identifying emotion categories and emotion management intervention strategies; Emotional intervention strategies include mild reassurance strategies, moderate intervention strategies, or emergency healing strategies.
[0014] Optionally, in the robot-based collaborative processing method for emotion perception described in this application, the step of distributing refined emotion analysis data to the edge and collecting user emotion feature variable data and uploading it to the cloud for model optimization specifically includes: The identification of emotion categories and the corresponding intervention strategies for emotional management will be distributed to marginalized groups. At the edge, emotional guidance and intervention strategies are used to provide emotional guidance to users, and user emotional characteristic variable data are collected and uploaded to the cloud at preset time intervals. The cloud-based system uses small-sample incremental learning technology to iteratively optimize the deep healing model.
[0015] Secondly, this application provides a robot-based emotion perception collaborative processing system, which includes: a memory and a processor, wherein the memory stores a program for a robot-based emotion perception collaborative processing method, and when the program for the robot-based emotion perception collaborative processing method is executed by the processor, it performs the following steps: The system acquires users' emotional feature data and preprocesses it to obtain effective emotional feature data. Acquire real-time computing load data at the robot edge and calculate the comprehensive computing load value. Adjust the preset dynamic lightweight model cluster according to the comprehensive computing load value to obtain an effective lightweight model. Effective emotional feature data is processed through an effective lightweight model to obtain multi-dimensional emotional feature vectors, and de-identification is performed according to the obtained user intention de-identification level to obtain de-identified multi-dimensional emotional feature vectors. Preliminary emotion categories are obtained by processing the desensitized multi-dimensional emotion feature vectors through a pre-set lightweight classification model. Preliminary emotion categories and desensitized multi-dimensional emotion feature vectors that meet the preset conditions are uploaded to the cloud for analysis to obtain refined emotion analysis data; The refined sentiment analysis data is distributed to the edge devices, and user sentiment characteristic variable data is collected and uploaded to the cloud for model optimization.
[0016] Optionally, in the robot-based emotion perception collaborative processing system described in this application, the step of acquiring the user's emotion feature data and preprocessing the emotion feature data to obtain effective emotion feature data specifically includes: Acquire users' emotional characteristics data, including voice data, physiological data, and visual data; Timestamps are added to and aligned for voice, physiological and visual data using dual calibration technology of NTP network time protocol and local hardware clock; After preprocessing the emotional feature data, effective emotional feature data is obtained, including effective voice data, effective physiological data, and effective visual data. The preprocessing includes framing and windowing the speech data, filtering and denoising the physiological data, and normalizing the size of the visual data.
[0017] As described above, the robot-based collaborative emotion perception processing method and system provided in this application obtain effective emotion feature data by acquiring and preprocessing emotion feature data, calculating a comprehensive computing load value based on real-time computing load data, dynamically adjusting a lightweight model cluster accordingly to obtain an effective lightweight model, fusing the effective lightweight model to generate a multi-dimensional emotion feature vector, and then desensitizing it by combining the user's intention de-identification level to obtain a preliminary emotion category. The preliminary emotion category and desensitized feature vector that meet the conditions are uploaded to the cloud for analysis to obtain refined emotion analysis data, which is then distributed to the edge. The edge executes a guidance strategy and simultaneously collects user emotion feature variable data and uploads it to the cloud to complete model optimization. This achieves collaborative emotion perception processing between the edge and the cloud, balancing dynamic computing power adaptation and user privacy protection, and improving the real-time performance, accuracy, and iterative optimization capability of the robot's emotion perception.
[0018] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A flowchart illustrating the robot-based emotion perception and collaborative processing method provided in this application embodiment; Figure 2 A flowchart illustrating the process of obtaining effective emotional feature data using a robot-based emotion perception collaborative processing method provided in this application embodiment; Figure 3 A flowchart illustrating the robot-based collaborative emotion perception processing method provided in this application embodiment; Figure 4 This is a schematic diagram of a robot-based emotion perception and collaborative processing system provided in an embodiment of this application. Detailed Implementation
[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0022] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0023] Please refer to Figure 1 , Figure 1 This is a flowchart of a robot-based emotion perception collaborative processing method according to some embodiments of this application. This robot-based emotion perception collaborative processing method is used in terminal devices, such as computers and mobile phones. The robot-based emotion perception collaborative processing method includes the following steps: S11. Obtain the user's emotional feature data, and obtain effective emotional feature data after preprocessing the emotional feature data; S12. Obtain real-time computing power load data at the robot edge and calculate the comprehensive computing power load value. Adjust the preset dynamic lightweight model cluster according to the comprehensive computing power load value to obtain an effective lightweight model. S13. Process the effective emotional feature data through an effective lightweight model to obtain a multi-dimensional emotional feature vector, and perform de-identification processing according to the obtained user intention de-identification level to obtain a de-identified multi-dimensional emotional feature vector. S14. Based on the desensitized multi-dimensional emotion feature vector, a preliminary emotion category is obtained by processing it through a preset lightweight classification model. S15. Upload the preliminary emotion categories and desensitized multi-dimensional emotion feature vectors that meet the preset conditions to the cloud for analysis to obtain refined emotion analysis data; S16. Distribute refined sentiment analysis data to the edge device and collect user sentiment characteristic variable data and upload it to the cloud for model optimization.
[0024] Understandably, the robot's multimodal acquisition module acquires the user's raw emotional feature data. A standardized preprocessing procedure then removes noise and standardizes the data format, filtering out effective emotional feature data that truly reflects the user's emotional state. Simultaneously, a hardware monitoring module at the edge collects real-time computing load data. This comprehensive computing load value is then used as the basis for scheduling, dynamically selecting and adjusting a pre-set dynamic lightweight model cluster to determine the effective lightweight model suitable for the current computing power. Subsequently, the effective emotional feature data is input into this effective lightweight model, where feature extraction and fusion generate a multi-dimensional emotional feature vector that can represent the user's emotions from multiple dimensions. This is then combined with the user's autonomous... Based on the established privacy protection requirements, the feature vector is subjected to targeted privacy desensitization processing according to the user's desired desensitization level, resulting in a desensitized multi-dimensional emotion feature vector. Next, the desensitized feature vector is input into a preset lightweight classification model, which uses classification reasoning to determine the user's initial emotion category. Then, preset judgment conditions are set to filter the initial emotion categories, and the initial emotion categories that meet the conditions, along with the desensitized multi-dimensional emotion feature vector, are uploaded to a cloud server. The cloud performs in-depth, refined emotion analysis, outputting refined emotion analysis data. Finally, the refined emotion analysis data from the cloud is distributed to the robot's edge device, which executes corresponding emotion guidance and interaction strategies based on this data, and simultaneously... (The sentence is incomplete and ends abruptly). The system continuously collects user emotion characteristic variable data within a specified range and uploads it to the cloud. The cloud uses this real-time data to iteratively optimize relevant models, achieving end-to-end collaboration from data collection, emotion recognition, hierarchical processing to model optimization. The edge refers to the robot's local hardware computing unit, which has the ability to collect, process, and execute data locally. While its computing resources are relatively limited, it has low processing latency and can achieve real-time response. In this embodiment, the dynamic lightweight model cluster refers to a set of models composed of multiple lightweight machine learning models designed for different computing power scenarios and task requirements. Models can be dynamically selected and combined according to the real-time computing power status of the edge. Compared with traditional deep learning models, lightweight models have fewer parameters and lower computational cost. Adapting to limited computing power at the edge; Multi-dimensional emotional feature vectors refer to the digital and vectorized emotional representation formed by extracting and fusing user emotional features from different dimensions, which can comprehensively reflect the user's emotional state from multiple dimensions; Refined emotional analysis data refers to the more accurate and comprehensive emotional-related data obtained by processing data uploaded from the edge through deeper models and analysis methods in the cloud, compared with the initial emotional categories, including accurate emotional categories and targeted emotional processing strategies; Emotional feature variable data: refers to the feature data reflecting changes in the user's emotional state collected within a preset time period after the edge implements the emotional guidance strategy, which is the core basis for judging the guidance effect and used for model optimization.
[0025] Please refer to Figure 2 , Figure 2This is a flowchart illustrating the process of obtaining effective emotional feature data using a robot-based emotion perception collaborative processing method provided in this application embodiment. According to this embodiment, obtaining the user's emotional feature data and preprocessing the emotional feature data to obtain effective emotional feature data specifically includes: S21. Obtain the user's emotional characteristic data, including voice data, physiological data and visual data; S22. Timestamps are added to and aligned for voice data, physiological data, and visual data using dual calibration technology of NTP network time protocol and local hardware clock; S23. After preprocessing the emotional feature data, obtain effective emotional feature data, including effective voice data, effective physiological data and effective visual data; S24. The preprocessing includes framing and windowing the speech data, filtering and denoising the physiological data, and normalizing the size of the visual data.
[0026] Understandably, the robot acquires user voice data (including tone, speech rate, and speech semantics), physiological data (including heart rate, respiratory rate, and blood pressure), and visual data (including facial expressions, body posture, eye movements, and facial motion units) through its voice acquisition module, physiological sensors, and visual acquisition module, respectively. Then, it uses the NTP network time protocol to obtain standard time from the network and combines this with the robot's local hardware clock to achieve dual-source time calibration. This adds a unified format timestamp to the acquired voice, physiological, and visual multimodal data, achieving precise alignment of different modalities on the timeline and eliminating feature correlation bias caused by misaligned data acquisition times. Finally, it performs specific processing based on the different characteristics of the three types of data. Differentiated preprocessing operations are performed, namely, frame-by-frame windowing of the speech data, dividing the continuous speech signal into short time frames and adding window functions to reduce spectral leakage; filtering and denoising of the physiological data to remove interference signals such as environmental noise and equipment noise to retain effective physiological features; and size normalization of the visual data to adjust visual images or video frames of different resolutions and sizes to a uniform size to ensure the standardization of model input. After the above preprocessing, effective speech data, effective physiological data, and effective visual data are obtained, which are integrated into effective emotional feature data that can be used for subsequent feature extraction. In this embodiment, NTP Network Time Protocol refers to Network Time Protocol, a standardized protocol used to achieve time synchronization of various nodes in a computer network. It can obtain accurate standard time through the network and ensure the uniformity of timestamps.
[0027] According to an embodiment of the present invention, the step of acquiring real-time computing power load data at the robot edge, calculating a comprehensive computing power load value, and adjusting a preset dynamic lightweight model cluster based on the comprehensive computing power load value to obtain an effective lightweight model specifically includes: Acquire real-time computing load data at the robot's edge, including processor utilization, memory usage, and task difficulty data; The real-time computing load data is normalized and then weighted by preset weighting coefficients to obtain the comprehensive computing load value. The comprehensive computing load value is compared with the preset model scheduling threshold to obtain the corresponding effective lightweight model.
[0028] Understandably, the hardware monitoring program at the robot's edge collects core data reflecting the computing power status in real time, including processor (CPU / GPU / NPU) utilization, actual memory usage, and quantified data on the computational complexity of the robot's current task, i.e., task difficulty. The collected processor utilization, memory usage, and task difficulty data are then normalized, converting the raw data with different dimensions and value ranges into standardized data between 0 and 1, eliminating dimensional differences. Finally, based on the impact of each data point on the overall computing power at the edge, and combined with pre-set weighting coefficients, a weighted summation is performed on the three types of normalized data. The calculation yields a comprehensive computing load value that quantifies the current overall computing power occupancy status at the edge. Finally, the calculated comprehensive computing load value is compared with a pre-set model scheduling threshold. Based on the comparison results, a model combination matching the current computing power status is selected from a pre-set dynamic lightweight model cluster. Specifically, when the computing load is low, a lightweight model with higher accuracy and slightly higher computational load is selected; when the computing load is high, a lightweight model with lower computational load and faster response speed is selected. Ultimately, an effective lightweight model suitable for the current edge computing power is determined, achieving rational utilization of computing resources and dynamic model scheduling. In this embodiment, the pre-set model scheduling threshold is set to 0.75.
[0029] According to an embodiment of the present invention, the step of processing effective emotional feature data through an effective lightweight model to obtain a multi-dimensional emotional feature vector, and performing desensitization processing according to the obtained user intention de-identification level to obtain a de-identified multi-dimensional emotional feature vector, specifically includes: A feature-level serial fusion strategy is adopted, firstly extracting the corresponding single-modal sentiment features through the single-modal lightweight model in the effective lightweight model; Then, the lightweight fusion model in the effective lightweight model is used to achieve deep fusion of multimodal features to generate multidimensional emotion feature vectors; Obtain the user's intended level of privacy, including Level 1, Level 2, or Level 3 privacy. Desensitization is performed based on the user's desired level of anonymization to obtain a desensitized multi-dimensional emotional feature vector.
[0030] Understandably, a feature-level serial fusion strategy is employed for the extraction and fusion of multimodal emotional features. First, effective speech data, effective physiological data, and effective visual data are input into the corresponding speech, physiological, and visual single-modal lightweight models within the effective lightweight model. Each single-modal lightweight model then extracts the emotional features of its corresponding modality, resulting in independent single-modal emotional features for each modality. Next, all single-modal emotional features are input into the lightweight fusion model within the effective lightweight model. Through the model's feature fusion layer, deep fusion of the features across modalities is achieved, uncovering intermodal correlations and eliminating information redundancy. Ultimately, a multi-dimensional emotional feature vector is generated that comprehensively and accurately represents the user's emotional state from multiple dimensions, including speech, physiological, and visual aspects. Simultaneously, user-defined emotional features are obtained through robot interaction. The privacy protection level refers to the user's intended decryption level, which is divided into three gradients: Level 1 privacy, Level 2 privacy, and Level 3 privacy, corresponding to high, medium, and low privacy protection needs, respectively. Finally, based on the user's selected intention decryption level, targeted privacy desensitization processing is performed on the generated multi-dimensional emotion feature vector. Sensitive features at different levels are removed, blurred, or masked to maximize user privacy while retaining the core features of emotion recognition, ultimately resulting in a desensitized multi-dimensional emotion feature vector. In this embodiment, the feature-level serial fusion strategy refers to a strategy for multimodal feature fusion, belonging to the category of feature-level fusion. "Serial" means that features are first extracted from each single-modal data separately, and then the extracted single-modal features are input into the fusion model according to a certain logic for deep fusion.
[0031] According to an embodiment of the present invention, the step of obtaining a preliminary emotion category by processing the desensitized multi-dimensional emotion feature vector through a preset lightweight classification model specifically includes: The desensitized multi-dimensional emotion feature vector is input into a preset lightweight classification model for processing to obtain the emotion category and the corresponding probability value. Emotional categories include basic emotions, mild negative emotions, moderate negative emotions, or extreme emotions; The emotion categories are sorted in descending order of probability value, and the emotion category with the highest probability value is taken as the initial emotion category.
[0032] Understandably, the anonymized multi-dimensional emotion feature vector, after privacy desensitization, is first input into a pre-trained, lightweight classification model deployed at the robot's edge. The model then performs emotion category inference calculations, outputting the corresponding emotion category and its probability value. This probability value, between 0 and 1, represents the model's confidence in determining that the user belongs to that emotion category. The model pre-defines four main emotion categories: basic emotions without negative tendencies (calm, joy, relaxation, etc.), and mildly negative, moderately negative, and extremely negative emotions (irritability / loss, anxiety / grievance, anger / collapse, etc.), categorized by severity. All emotion categories output by the model are then sorted in descending order of their corresponding probability values. Finally, the emotion category with the highest probability value is determined as the user's initial emotion category. This determination method balances the computational limitations and processing efficiency of the edge, enabling rapid, real-time preliminary determination of the user's emotional state. In this embodiment, the probability values corresponding to basic, mildly negative, moderately negative, and extremely negative emotions are summed to 1.
[0033] According to an embodiment of the present invention, uploading preliminary emotion categories and desensitized multi-dimensional emotion feature vectors that meet preset conditions to the cloud for analysis to obtain refined emotion analysis data specifically includes: Determine the initial emotion category. If the initial emotion category is a basic emotion, then process it at the edge to obtain the corresponding basic response strategy. If the initial emotion category is any of mild negative, moderate negative, or extreme emotion, then the initial emotion category and the desensitized multi-dimensional emotion feature vector will be uploaded to the cloud. The desensitized multi-dimensional emotional feature vectors are analyzed using a pre-set deep healing model under a federated learning framework, combined with a pre-set user emotional profile knowledge graph, to obtain refined emotional analysis data. Refined emotion analysis data includes identifying emotion categories and emotion management intervention strategies; Emotional intervention strategies include mild reassurance strategies, moderate intervention strategies, or emergency healing strategies.
[0034] Understandably, using whether the initial emotion category is a basic emotion is a core preset criterion for edge-cloud graded processing, enabling differentiated allocation between local edge processing and deep cloud analysis. If the initial emotion category is a basic emotion, there's no need to upload to the cloud; the robot directly matches and generates the corresponding basic response strategy at the edge, completing rapid local processing. If the initial emotion category is any of mild negative, moderate negative, or extreme emotions, it indicates the user requires in-depth emotion analysis and intervention. In this case, the initial emotion category and desensitized multi-dimensional emotion feature vectors are transmitted to the cloud server via an encrypted network. After receiving the data, the cloud uses a privacy-preserving analysis environment built on a federated learning framework to call a preset deep healing model and perform joint analysis in conjunction with a pre-built user emotion profile knowledge graph. This user emotion profile knowledge graph integrates... The system uses multi-dimensional information, including users' historical emotional data, personality traits, emotional triggers, and past intervention effects, to achieve personalized emotional analysis. Finally, through in-depth cloud analysis, it outputs a more precise and specific emotional category than the initial emotional classification, along with matching emotional intervention strategies. These strategies are categorized into three types based on the severity of negative emotions: mild soothing strategies, moderate intervention strategies, and emergency healing strategies. This results in refined emotional analysis data containing both the identified emotional category and the intervention strategies. In this embodiment, the basic response strategy refers to a simple, real-time interactive response strategy developed by the robot's edge for the user's basic emotions, requiring no cloud intervention and adapting to the rapid interaction needs of basic emotions. The mild soothing strategy, moderate intervention strategy, and emergency healing strategy can be customized according to needs or user profiles.
[0035] According to an embodiment of the present invention, the step of distributing refined sentiment analysis data to the edge and collecting user sentiment characteristic variable data and uploading it to the cloud for model optimization specifically includes: The identification of emotion categories and the corresponding intervention strategies for emotional management will be distributed to marginalized groups. At the edge, emotional guidance and intervention strategies are used to provide emotional guidance to users, and user emotional characteristic variable data are collected and uploaded to the cloud at preset time intervals. The cloud-based system uses small-sample incremental learning technology to iteratively optimize the deep healing model.
[0036] Understandably, the determined emotion categories and corresponding emotional intervention strategies obtained from cloud analysis are distributed to the robot's edge terminal via an encrypted network protocol. Upon receiving the refined emotion analysis data, the edge terminal immediately invokes the corresponding interaction and execution modules to conduct targeted emotional guidance and interactive operations for the user according to the emotional intervention strategy. Simultaneously, the edge terminal periodically and continuously collects the user's emotional characteristic data during the emotional guidance process according to a pre-set time period. This data is dynamic data reflecting the changes in the user's emotional state as the guidance operation changes, i.e., emotional characteristic variable data, which can intuitively demonstrate the actual effect of emotional guidance. After collection, the emotional characteristic variable data is uploaded to the cloud server via a real-time transmission protocol. After receiving the user's real-time emotional characteristic variable data, the cloud uses small-sample incremental learning technology to iteratively update and optimize the deep healing model under the federated learning framework using a small amount of real-time scene data. This eliminates the need for large-scale batch sample retraining, avoiding "catastrophic forgetting" of the model while achieving continuous lightweight optimization of the model. This allows the model to continuously adapt to the user's emotional change patterns, improving the accuracy of subsequent emotion analysis and guidance strategy formulation.
[0037] It's worth mentioning that after obtaining and running the corresponding effective lightweight model, it also includes: Real-time comprehensive computing load values are collected according to a preset time period. The real-time comprehensive computing power load value is compared with the preset edge computing power status evaluation threshold to obtain the edge computing power status. The preset edge computing power status evaluation thresholds include a first threshold and a second threshold, and the first threshold is greater than the second threshold; If the real-time comprehensive computing load value is greater than the first threshold, the edge computing power status is a computing power shortage state; If the real-time comprehensive computing load value is less than the first threshold and greater than or equal to the second threshold, then the edge computing power status is a moderate computing power status. If the real-time comprehensive computing load value is less than the second threshold, the edge computing power status is a sufficient computing power status.
[0038] Understandably, after obtaining and running an effective lightweight model adapted to the current computing power at the robot's edge, the system will initiate a real-time dynamic monitoring and grading process for edge computing power. Specifically, the system continuously collects core computing load indicators such as edge processor utilization, memory usage, and task difficulty data according to a pre-set time period (which can be flexibly configured according to the robot's actual application scenario, such as milliseconds or seconds). It repeatedly executes data normalization and weighted summation operations to calculate and obtain a real-time comprehensive computing load value that reflects the overall state of the edge computing power. Simultaneously, the system pre-configures an edge computing power state evaluation threshold system containing a first threshold and a second threshold, with the first threshold value being greater than the second threshold. This serves as the core basis for grading the computing power state, and the real-time calculated comprehensive computing load value is compared with this dual-threshold system. The system compares values one by one to accurately classify the computing power status at the edge: if the real-time comprehensive computing power load value is greater than the first threshold, it indicates that the computing resources such as processors and memory at the edge are highly occupied, and there is no spare computing power to support the stable operation of the effective lightweight model, and the edge computing power status is determined to be a computing power shortage state; if the real-time comprehensive computing power load value is less than the first threshold but greater than or equal to the second threshold, it indicates that the usage of computing power resources at the edge is within a reasonable range and can support the normal operation of the effective lightweight model, and the edge computing power status is determined to be a computing power moderate state; if the real-time comprehensive computing power load value is less than the second threshold, it indicates that the usage rate of computing power resources at the edge is low, and there is sufficient spare computing power to support the operation of the model and even the processing of additional tasks, and the edge computing power status is determined to be a computing power sufficient state. Through this dynamic monitoring and hierarchical judgment, the real-time perception and accurate classification of the computing power status at the edge can be achieved.
[0039] It's worth mentioning that after obtaining the edge computing power status, it also includes: If the edge computing power status is in a state of computing power shortage, a computing power scheduling request will be triggered and sent to the cloud; The cloud allocates dedicated computing nodes to the edge through a preset global computing resource monitoring module; Dedicated computing nodes use the same efficient and lightweight model as the edge computing nodes to process effective sentiment feature data. When the edge computing power status shows that the computing power is sufficient, the cloud releases the corresponding dedicated computing node, and the edge device resumes normal working mode.
[0040] Understandably, after the system determines the edge computing power status and obtains the specific computing power status result, it will initiate a dynamic scheduling and elastic resource allocation process for edge-cloud collaboration, executing differentiated computing power scheduling strategies for different computing power statuses: If the edge computing power status is determined to be under pressure, it means that the local computing power at the edge can no longer support the normal processing of emotional feature data by the effective lightweight model. At this time, the edge will immediately and automatically trigger a computing power scheduling request. This request includes core content such as the edge device identifier, current computing power load data, effective lightweight model information, and emotional feature data processing requirements, and is sent to the cloud server in real time via an encrypted network communication protocol to initiate a cloud computing power support request. After receiving the computing power scheduling request, the cloud server scans the idle computing node resources in the cloud in real time through a pre-deployed preset global computing power resource monitoring module, and allocates a matching cloud computing node exclusively for the edge based on the processing needs of the edge and the computing power requirements of the effective lightweight model. Furthermore, this dedicated computing node will synchronously deploy an effective lightweight model that is completely consistent with the edge end, ensuring the uniformity of data processing logic. Subsequently, the edge end transmits the effective emotional feature data to be processed to this dedicated computing node, and the dedicated computing node in the cloud replaces the edge end to complete the feature extraction, fusion and other processing tasks of the effective emotional feature data, alleviating the computing power pressure on the edge end. If the real-time comprehensive computing power load value of the edge end subsequently decreases, it is determined that the edge computing power status has recovered to a sufficient computing power state, indicating that the edge end has regained the computing power conditions to independently support the operation of the effective lightweight model. At this time, the cloud will immediately release the dedicated computing node allocated to the edge end, return the node to the cloud computing power resource pool, and the edge end will stop transmitting data to the cloud, resuming the normal working mode of processing effective emotional feature data by the local effective lightweight model. Through the computing power scheduling mechanism of edge-cloud collaboration, the elastic allocation of cloud computing power resources and the dynamic supplementation of edge computing power are realized, ensuring the continuity and stability of emotional perception data processing.
[0041] The present invention also discloses a robot-based emotion perception collaborative processing system, including a memory 41 and a processor 42. The memory stores a robot-based emotion perception collaborative processing method program, which, when executed by the processor, performs the following steps: The system acquires users' emotional feature data and preprocesses it to obtain effective emotional feature data. Acquire real-time computing load data at the robot edge and calculate the comprehensive computing load value. Adjust the preset dynamic lightweight model cluster according to the comprehensive computing load value to obtain an effective lightweight model. Effective emotional feature data is processed through an effective lightweight model to obtain multi-dimensional emotional feature vectors, and de-identification is performed according to the obtained user intention de-identification level to obtain de-identified multi-dimensional emotional feature vectors. Preliminary emotion categories are obtained by processing the desensitized multi-dimensional emotion feature vectors through a pre-set lightweight classification model. Preliminary emotion categories and desensitized multi-dimensional emotion feature vectors that meet the preset conditions are uploaded to the cloud for analysis to obtain refined emotion analysis data; The refined sentiment analysis data is distributed to the edge devices, and user sentiment characteristic variable data is collected and uploaded to the cloud for model optimization.
[0042] Understandably, the robot's multimodal acquisition module acquires the user's raw emotional feature data. A standardized preprocessing procedure then removes noise and standardizes the data format, filtering out effective emotional feature data that truly reflects the user's emotional state. Simultaneously, a hardware monitoring module at the edge collects real-time computing load data. This comprehensive computing load value is then used as the basis for scheduling, dynamically selecting and adjusting a pre-set dynamic lightweight model cluster to determine the effective lightweight model suitable for the current computing power. Subsequently, the effective emotional feature data is input into this effective lightweight model, where feature extraction and fusion generate a multi-dimensional emotional feature vector that can represent the user's emotions from multiple dimensions. This is then combined with the user's autonomous... Based on the established privacy protection requirements, the feature vector is subjected to targeted privacy desensitization processing according to the user's desired desensitization level, resulting in a desensitized multi-dimensional emotion feature vector. Next, the desensitized feature vector is input into a preset lightweight classification model, which uses classification reasoning to determine the user's initial emotion category. Then, preset judgment conditions are set to filter the initial emotion categories, and the initial emotion categories that meet the conditions, along with the desensitized multi-dimensional emotion feature vector, are uploaded to a cloud server. The cloud performs in-depth, refined emotion analysis, outputting refined emotion analysis data. Finally, the refined emotion analysis data from the cloud is distributed to the robot's edge device, which executes corresponding emotion guidance and interaction strategies based on this data, and simultaneously... (The sentence is incomplete and ends abruptly). The system continuously collects user emotion characteristic variable data within a specified range and uploads it to the cloud. The cloud uses this real-time data to iteratively optimize relevant models, achieving end-to-end collaboration from data collection, emotion recognition, hierarchical processing to model optimization. The edge refers to the robot's local hardware computing unit, which has the ability to collect, process, and execute data locally. While its computing resources are relatively limited, it has low processing latency and can achieve real-time response. In this embodiment, the dynamic lightweight model cluster refers to a set of models composed of multiple lightweight machine learning models designed for different computing power scenarios and task requirements. Models can be dynamically selected and combined according to the real-time computing power status of the edge. Compared with traditional deep learning models, lightweight models have fewer parameters and lower computational cost. Adapting to limited computing power at the edge; Multi-dimensional emotional feature vectors refer to the digital and vectorized emotional representation formed by extracting and fusing user emotional features from different dimensions, which can comprehensively reflect the user's emotional state from multiple dimensions; Refined emotional analysis data refers to the more accurate and comprehensive emotional-related data obtained by processing data uploaded from the edge through deeper models and analysis methods in the cloud, compared with the initial emotional categories, including accurate emotional categories and targeted emotional processing strategies; Emotional feature variable data: refers to the feature data reflecting changes in the user's emotional state collected within a preset time period after the edge implements the emotional guidance strategy, which is the core basis for judging the guidance effect and used for model optimization.
[0043] According to an embodiment of the present invention, the step of obtaining user emotional feature data and preprocessing the emotional feature data to obtain effective emotional feature data specifically includes: Acquire users' emotional characteristics data, including voice data, physiological data, and visual data; Timestamps are added to and aligned for voice, physiological and visual data using dual calibration technology of NTP network time protocol and local hardware clock; After preprocessing the emotional feature data, effective emotional feature data is obtained, including effective voice data, effective physiological data, and effective visual data. The preprocessing includes framing and windowing the speech data, filtering and denoising the physiological data, and normalizing the size of the visual data.
[0044] Understandably, the robot acquires user voice data (including tone, speech rate, and speech semantics), physiological data (including heart rate, respiratory rate, and blood pressure), and visual data (including facial expressions, body posture, eye movements, and facial motion units) through its voice acquisition module, physiological sensors, and visual acquisition module, respectively. Then, it uses the NTP network time protocol to obtain standard time from the network and combines this with the robot's local hardware clock to achieve dual-source time calibration. This adds a unified format timestamp to the acquired voice, physiological, and visual multimodal data, achieving precise alignment of different modalities on the timeline and eliminating feature correlation bias caused by misaligned data acquisition times. Finally, it performs specific processing based on the different characteristics of the three types of data. Differentiated preprocessing operations are performed, namely, frame-by-frame windowing of the speech data, dividing the continuous speech signal into short time frames and adding window functions to reduce spectral leakage; filtering and denoising of the physiological data to remove interference signals such as environmental noise and equipment noise to retain effective physiological features; and size normalization of the visual data to adjust visual images or video frames of different resolutions and sizes to a uniform size to ensure the standardization of model input. After the above preprocessing, effective speech data, effective physiological data, and effective visual data are obtained, which are integrated into effective emotional feature data that can be used for subsequent feature extraction. In this embodiment, NTP Network Time Protocol refers to Network Time Protocol, a standardized protocol used to achieve time synchronization of various nodes in a computer network. It can obtain accurate standard time through the network and ensure the uniformity of timestamps.
[0045] According to an embodiment of the present invention, the step of acquiring real-time computing power load data at the robot edge, calculating a comprehensive computing power load value, and adjusting a preset dynamic lightweight model cluster based on the comprehensive computing power load value to obtain an effective lightweight model specifically includes: Acquire real-time computing load data at the robot's edge, including processor utilization, memory usage, and task difficulty data; The real-time computing load data is normalized and then weighted by preset weighting coefficients to obtain the comprehensive computing load value. The comprehensive computing load value is compared with the preset model scheduling threshold to obtain the corresponding effective lightweight model.
[0046] Understandably, the hardware monitoring program at the robot's edge collects core data reflecting the computing power status in real time, including processor (CPU / GPU / NPU) utilization, actual memory usage, and quantified data on the computational complexity of the robot's current task, i.e., task difficulty. The collected processor utilization, memory usage, and task difficulty data are then normalized, converting the raw data with different dimensions and value ranges into standardized data between 0 and 1, eliminating dimensional differences. Finally, based on the impact of each data point on the overall computing power at the edge, and combined with pre-set weighting coefficients, a weighted summation is performed on the three types of normalized data. The calculation yields a comprehensive computing load value that quantifies the current overall computing power occupancy status at the edge. Finally, the calculated comprehensive computing load value is compared with a pre-set model scheduling threshold. Based on the comparison results, a model combination matching the current computing power status is selected from a pre-set dynamic lightweight model cluster. Specifically, when the computing load is low, a lightweight model with higher accuracy and slightly higher computational load is selected; when the computing load is high, a lightweight model with lower computational load and faster response speed is selected. Ultimately, an effective lightweight model suitable for the current edge computing power is determined, achieving rational utilization of computing resources and dynamic model scheduling. In this embodiment, the pre-set model scheduling threshold is set to 0.75.
[0047] According to an embodiment of the present invention, the step of processing effective emotional feature data through an effective lightweight model to obtain a multi-dimensional emotional feature vector, and performing desensitization processing according to the obtained user intention de-identification level to obtain a de-identified multi-dimensional emotional feature vector, specifically includes: A feature-level serial fusion strategy is adopted, firstly extracting the corresponding single-modal sentiment features through the single-modal lightweight model in the effective lightweight model; Then, the lightweight fusion model in the effective lightweight model is used to achieve deep fusion of multimodal features to generate multidimensional emotion feature vectors; Obtain the user's intended level of privacy, including Level 1, Level 2, or Level 3 privacy. Desensitization is performed based on the user's desired level of anonymization to obtain a desensitized multi-dimensional emotional feature vector.
[0048] Understandably, a feature-level serial fusion strategy is employed for the extraction and fusion of multimodal emotional features. First, effective speech data, effective physiological data, and effective visual data are input into the corresponding speech, physiological, and visual single-modal lightweight models within the effective lightweight model. Each single-modal lightweight model then extracts the emotional features of its corresponding modality, resulting in independent single-modal emotional features for each modality. Next, all single-modal emotional features are input into the lightweight fusion model within the effective lightweight model. Through the model's feature fusion layer, deep fusion of the features across modalities is achieved, uncovering intermodal correlations and eliminating information redundancy. Ultimately, a multi-dimensional emotional feature vector is generated that comprehensively and accurately represents the user's emotional state from multiple dimensions, including speech, physiological, and visual aspects. Simultaneously, user-defined emotional features are obtained through robot interaction. The privacy protection level refers to the user's intended decryption level, which is divided into three gradients: Level 1 privacy, Level 2 privacy, and Level 3 privacy, corresponding to high, medium, and low privacy protection needs, respectively. Finally, based on the user's selected intention decryption level, targeted privacy desensitization processing is performed on the generated multi-dimensional emotion feature vector. Sensitive features at different levels are removed, blurred, or masked to maximize user privacy while retaining the core features of emotion recognition, ultimately resulting in a desensitized multi-dimensional emotion feature vector. In this embodiment, the feature-level serial fusion strategy refers to a strategy for multimodal feature fusion, belonging to the category of feature-level fusion. "Serial" means that features are first extracted from each single-modal data separately, and then the extracted single-modal features are input into the fusion model according to a certain logic for deep fusion.
[0049] According to an embodiment of the present invention, the step of obtaining a preliminary emotion category by processing the desensitized multi-dimensional emotion feature vector through a preset lightweight classification model specifically includes: The desensitized multi-dimensional emotion feature vector is input into a preset lightweight classification model for processing to obtain the emotion category and the corresponding probability value. Emotional categories include basic emotions, mild negative emotions, moderate negative emotions, or extreme emotions; The emotion categories are sorted in descending order of probability value, and the emotion category with the highest probability value is taken as the initial emotion category.
[0050] Understandably, the anonymized multi-dimensional emotion feature vector, after privacy desensitization, is first input into a pre-trained, lightweight classification model deployed at the robot's edge. The model then performs emotion category inference calculations, outputting the corresponding emotion category and its probability value. This probability value, between 0 and 1, represents the model's confidence in determining that the user belongs to that emotion category. The model pre-defines four main emotion categories: basic emotions without negative tendencies (calm, joy, relaxation, etc.), and mildly negative, moderately negative, and extremely negative emotions (irritability / loss, anxiety / grievance, anger / collapse, etc.), categorized by severity. All emotion categories output by the model are then sorted in descending order of their corresponding probability values. Finally, the emotion category with the highest probability value is determined as the user's initial emotion category. This determination method balances the computational limitations and processing efficiency of the edge, enabling rapid, real-time preliminary determination of the user's emotional state. In this embodiment, the probability values corresponding to basic, mildly negative, moderately negative, and extremely negative emotions are summed to 1.
[0051] According to an embodiment of the present invention, uploading preliminary emotion categories and desensitized multi-dimensional emotion feature vectors that meet preset conditions to the cloud for analysis to obtain refined emotion analysis data specifically includes: Determine the initial emotion category. If the initial emotion category is a basic emotion, then process it at the edge to obtain the corresponding basic response strategy. If the initial emotion category is any of mild negative, moderate negative, or extreme emotion, then the initial emotion category and the desensitized multi-dimensional emotion feature vector will be uploaded to the cloud. The desensitized multi-dimensional emotional feature vectors are analyzed using a pre-set deep healing model under a federated learning framework, combined with a pre-set user emotional profile knowledge graph, to obtain refined emotional analysis data. Refined emotion analysis data includes identifying emotion categories and emotion management intervention strategies; Emotional intervention strategies include mild reassurance strategies, moderate intervention strategies, or emergency healing strategies.
[0052] Understandably, using whether the initial emotion category is a basic emotion is a core preset criterion for edge-cloud graded processing, enabling differentiated allocation between local edge processing and deep cloud analysis. If the initial emotion category is a basic emotion, there's no need to upload to the cloud; the robot directly matches and generates the corresponding basic response strategy at the edge, completing rapid local processing. If the initial emotion category is any of mild negative, moderate negative, or extreme emotions, it indicates the user requires in-depth emotion analysis and intervention. In this case, the initial emotion category and desensitized multi-dimensional emotion feature vectors are transmitted to the cloud server via an encrypted network. After receiving the data, the cloud uses a privacy-preserving analysis environment built on a federated learning framework to call a preset deep healing model and perform joint analysis in conjunction with a pre-built user emotion profile knowledge graph. This user emotion profile knowledge graph integrates... The system uses multi-dimensional information, including users' historical emotional data, personality traits, emotional triggers, and past intervention effects, to achieve personalized emotional analysis. Finally, through in-depth cloud analysis, it outputs a more precise and specific emotional category than the initial emotional classification, along with matching emotional intervention strategies. These strategies are categorized into three types based on the severity of negative emotions: mild soothing strategies, moderate intervention strategies, and emergency healing strategies. This results in refined emotional analysis data containing both the identified emotional category and the intervention strategies. In this embodiment, the basic response strategy refers to a simple, real-time interactive response strategy developed by the robot's edge for the user's basic emotions, requiring no cloud intervention and adapting to the rapid interaction needs of basic emotions. The mild soothing strategy, moderate intervention strategy, and emergency healing strategy can be customized according to needs or user profiles.
[0053] According to an embodiment of the present invention, the step of distributing refined sentiment analysis data to the edge and collecting user sentiment characteristic variable data and uploading it to the cloud for model optimization specifically includes: The identification of emotion categories and the corresponding intervention strategies for emotional management will be distributed to marginalized groups. At the edge, emotional guidance and intervention strategies are used to provide emotional guidance to users, and user emotional characteristic variable data are collected and uploaded to the cloud at preset time intervals. The cloud-based system uses small-sample incremental learning technology to iteratively optimize the deep healing model.
[0054] Understandably, the determined emotion categories and corresponding emotional intervention strategies obtained from cloud analysis are distributed to the robot's edge terminal via an encrypted network protocol. Upon receiving the refined emotion analysis data, the edge terminal immediately invokes the corresponding interaction and execution modules to conduct targeted emotional guidance and interactive operations for the user according to the emotional intervention strategy. Simultaneously, the edge terminal periodically and continuously collects the user's emotional characteristic data during the emotional guidance process according to a pre-set time period. This data is dynamic data reflecting the changes in the user's emotional state as the guidance operation changes, i.e., emotional characteristic variable data, which can intuitively demonstrate the actual effect of emotional guidance. After collection, the emotional characteristic variable data is uploaded to the cloud server via a real-time transmission protocol. After receiving the user's real-time emotional characteristic variable data, the cloud uses small-sample incremental learning technology to iteratively update and optimize the deep healing model under the federated learning framework using a small amount of real-time scene data. This eliminates the need for large-scale batch sample retraining, avoiding "catastrophic forgetting" of the model while achieving continuous lightweight optimization of the model. This allows the model to continuously adapt to the user's emotional change patterns, improving the accuracy of subsequent emotion analysis and guidance strategy formulation.
[0055] It's worth mentioning that after obtaining and running the corresponding effective lightweight model, it also includes: Real-time comprehensive computing load values are collected according to a preset time period. The real-time comprehensive computing power load value is compared with the preset edge computing power status evaluation threshold to obtain the edge computing power status. The preset edge computing power status evaluation thresholds include a first threshold and a second threshold, and the first threshold is greater than the second threshold; If the real-time comprehensive computing load value is greater than the first threshold, the edge computing power status is a computing power shortage state; If the real-time comprehensive computing load value is less than the first threshold and greater than or equal to the second threshold, then the edge computing power status is a moderate computing power status. If the real-time comprehensive computing load value is less than the second threshold, the edge computing power status is a sufficient computing power status.
[0056] Understandably, after obtaining and running an effective lightweight model adapted to the current computing power at the robot's edge, the system will initiate a real-time dynamic monitoring and grading process for edge computing power. Specifically, the system continuously collects core computing load indicators such as edge processor utilization, memory usage, and task difficulty data according to a pre-set time period (which can be flexibly configured according to the robot's actual application scenario, such as milliseconds or seconds). It repeatedly executes data normalization and weighted summation operations to calculate and obtain a real-time comprehensive computing load value that reflects the overall state of the edge computing power. Simultaneously, the system pre-configures an edge computing power state evaluation threshold system containing a first threshold and a second threshold, with the first threshold value being greater than the second threshold. This serves as the core basis for grading the computing power state, and the real-time calculated comprehensive computing load value is compared with this dual-threshold system. The system compares values one by one to accurately classify the computing power status at the edge: if the real-time comprehensive computing power load value is greater than the first threshold, it indicates that the computing resources such as processors and memory at the edge are highly occupied, and there is no spare computing power to support the stable operation of the effective lightweight model, and the edge computing power status is determined to be a computing power shortage state; if the real-time comprehensive computing power load value is less than the first threshold but greater than or equal to the second threshold, it indicates that the usage of computing power resources at the edge is within a reasonable range and can support the normal operation of the effective lightweight model, and the edge computing power status is determined to be a computing power moderate state; if the real-time comprehensive computing power load value is less than the second threshold, it indicates that the usage rate of computing power resources at the edge is low, and there is sufficient spare computing power to support the operation of the model and even the processing of additional tasks, and the edge computing power status is determined to be a computing power sufficient state. Through this dynamic monitoring and hierarchical judgment, the real-time perception and accurate classification of the computing power status at the edge can be achieved.
[0057] It's worth mentioning that after obtaining the edge computing power status, it also includes: If the edge computing power status is in a state of computing power shortage, a computing power scheduling request will be triggered and sent to the cloud; The cloud allocates dedicated computing nodes to the edge through a preset global computing resource monitoring module; Dedicated computing nodes use the same efficient and lightweight model as the edge computing nodes to process effective sentiment feature data. When the edge computing power status shows that the computing power is sufficient, the cloud releases the corresponding dedicated computing node, and the edge device resumes normal working mode.
[0058] Understandably, after the system determines the edge computing power status and obtains the specific computing power status result, it will initiate a dynamic scheduling and elastic resource allocation process for edge-cloud collaboration, executing differentiated computing power scheduling strategies for different computing power statuses: If the edge computing power status is determined to be under pressure, it means that the local computing power at the edge can no longer support the normal processing of emotional feature data by the effective lightweight model. At this time, the edge will immediately and automatically trigger a computing power scheduling request. This request includes core content such as the edge device identifier, current computing power load data, effective lightweight model information, and emotional feature data processing requirements, and is sent to the cloud server in real time via an encrypted network communication protocol to initiate a cloud computing power support request. After receiving the computing power scheduling request, the cloud server scans the idle computing node resources in the cloud in real time through a pre-deployed preset global computing power resource monitoring module, and allocates a matching cloud computing node exclusively for the edge based on the processing needs of the edge and the computing power requirements of the effective lightweight model. Furthermore, this dedicated computing node will synchronously deploy an effective lightweight model that is completely consistent with the edge end, ensuring the uniformity of data processing logic. Subsequently, the edge end transmits the effective emotional feature data to be processed to this dedicated computing node, and the dedicated computing node in the cloud replaces the edge end to complete the feature extraction, fusion and other processing tasks of the effective emotional feature data, alleviating the computing power pressure on the edge end. If the real-time comprehensive computing power load value of the edge end subsequently decreases, it is determined that the edge computing power status has recovered to a sufficient computing power state, indicating that the edge end has regained the computing power conditions to independently support the operation of the effective lightweight model. At this time, the cloud will immediately release the dedicated computing node allocated to the edge end, return the node to the cloud computing power resource pool, and the edge end will stop transmitting data to the cloud, resuming the normal working mode of processing effective emotional feature data by the local effective lightweight model. Through the computing power scheduling mechanism of edge-cloud collaboration, the elastic allocation of cloud computing power resources and the dynamic supplementation of edge computing power are realized, ensuring the continuity and stability of emotional perception data processing.
[0059] This invention discloses a robot-based collaborative emotion perception processing method and system. It acquires and preprocesses emotion feature data to obtain effective emotion feature data, calculates a comprehensive computing load value based on real-time computing load data, dynamically adjusts a lightweight model cluster accordingly to obtain an effective lightweight model, and generates a multi-dimensional emotion feature vector through fusion of the effective lightweight model. After desensitization based on user intention de-identification levels, a preliminary emotion category is obtained. The preliminary emotion category and desensitized feature vector that meet the conditions are uploaded to the cloud for analysis to obtain refined emotion analysis data, which is then distributed to the edge. The edge executes a guidance strategy and simultaneously collects user emotion feature variable data and uploads it to the cloud to complete model optimization. This achieves collaborative emotion perception processing between the edge and cloud, balancing dynamic computing power adaptation with user privacy protection, and improving the real-time performance, accuracy, and iterative optimization capabilities of the robot's emotion perception.
[0060] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0061] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0062] In addition, in the various embodiments of the present invention, each functional unit can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0063] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0064] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.
Claims
1. A method for emotion-aware collaborative processing based on robots, characterized by, include: The system acquires users' emotional feature data and preprocesses it to obtain effective emotional feature data. Acquire real-time computing load data at the robot edge and calculate the comprehensive computing load value. Adjust the preset dynamic lightweight model cluster according to the comprehensive computing load value to obtain an effective lightweight model. Effective emotional feature data is processed through an effective lightweight model to obtain multi-dimensional emotional feature vectors, and de-identification is performed according to the obtained user intention de-identification level to obtain de-identified multi-dimensional emotional feature vectors. Preliminary emotion categories are obtained by processing the desensitized multi-dimensional emotion feature vectors through a pre-set lightweight classification model. Preliminary emotion categories and desensitized multi-dimensional emotion feature vectors that meet the preset conditions are uploaded to the cloud for analysis to obtain refined emotion analysis data; The refined sentiment analysis data is distributed to the edge devices, and user sentiment characteristic variable data is collected and uploaded to the cloud for model optimization.
2. The robot-based emotion perception collaborative processing method according to claim 1, characterized in that, The process of acquiring user emotional feature data and preprocessing the emotional feature data to obtain effective emotional feature data specifically includes: Acquire users' emotional characteristics data, including voice data, physiological data, and visual data; Timestamps are added to and aligned for voice, physiological and visual data using dual calibration technology of NTP network time protocol and local hardware clock; After preprocessing the emotional feature data, effective emotional feature data is obtained, including effective voice data, effective physiological data, and effective visual data. The preprocessing includes framing and windowing the speech data, filtering and denoising the physiological data, and normalizing the size of the visual data.
3. The robot-based emotion perception collaborative processing method according to claim 2, characterized in that, The process of acquiring real-time computing load data at the robot's edge and calculating a comprehensive computing load value, then adjusting a preset dynamic lightweight model cluster based on the comprehensive computing load value to obtain an effective lightweight model, specifically includes: Acquire real-time computing load data at the robot's edge, including processor utilization, memory usage, and task difficulty data; The real-time computing load data is normalized and then weighted by preset weighting coefficients to obtain the comprehensive computing load value. The comprehensive computing load value is compared with the preset model scheduling threshold to obtain the corresponding effective lightweight model.
4. The robot-based emotion perception and collaborative processing method according to claim 3, characterized in that, The process of processing effective emotional feature data through an effective lightweight model to obtain multi-dimensional emotional feature vectors, and then performing de-identification processing based on the obtained user intention de-identification level to obtain de-identified multi-dimensional emotional feature vectors, specifically includes: A feature-level serial fusion strategy is adopted, firstly extracting the corresponding single-modal sentiment features through the single-modal lightweight model in the effective lightweight model; Then, the lightweight fusion model in the effective lightweight model is used to achieve deep fusion of multimodal features to generate multidimensional emotion feature vectors; Obtain the user's intended level of privacy, including Level 1, Level 2, or Level 3 privacy. Desensitization is performed based on the user's desired level of anonymization to obtain a desensitized multi-dimensional emotional feature vector.
5. The robot-based emotion perception and collaborative processing method according to claim 4, characterized in that, The process of obtaining preliminary emotion categories by processing desensitized multi-dimensional emotion feature vectors through a preset lightweight classification model specifically includes: The desensitized multi-dimensional emotion feature vector is input into a preset lightweight classification model for processing to obtain the emotion category and the corresponding probability value. Emotional categories include basic emotions, mild negative emotions, moderate negative emotions, or extreme emotions; The emotion categories are sorted in descending order of probability value, and the emotion category with the highest probability value is taken as the initial emotion category.
6. The robot-based emotion perception collaborative processing method according to claim 5, characterized in that, The process of uploading preliminary emotion categories and desensitized multi-dimensional emotion feature vectors that meet preset conditions to the cloud for analysis to obtain refined emotion analysis data specifically includes: Determine the initial emotion category. If the initial emotion category is a basic emotion, then process it at the edge to obtain the corresponding basic response strategy. If the initial emotion category is any of mild negative, moderate negative, or extreme emotion, then the initial emotion category and the desensitized multi-dimensional emotion feature vector will be uploaded to the cloud. The desensitized multi-dimensional emotional feature vectors are analyzed using a pre-set deep healing model under a federated learning framework, combined with a pre-set user emotional profile knowledge graph, to obtain refined emotional analysis data. Refined emotion analysis data includes identifying emotion categories and emotion management intervention strategies; Emotional intervention strategies include mild reassurance strategies, moderate intervention strategies, or emergency healing strategies.
7. The robot-based collaborative processing method for emotion perception according to claim 6, characterized in that, The process of distributing refined sentiment analysis data to edge devices and collecting user sentiment characteristic variable data and uploading it to the cloud for model optimization specifically includes: The identification of emotion categories and the corresponding intervention strategies for emotional management will be distributed to marginalized groups. At the edge, emotional guidance and intervention strategies are used to provide emotional guidance to users, and user emotional characteristic variable data are collected and uploaded to the cloud at preset time intervals. The cloud-based system uses small-sample incremental learning technology to iteratively optimize the deep healing model.
8. A robot-based emotion perception and collaborative processing system, characterized in that, It includes a memory and a processor. The memory includes a robot-based emotion perception collaborative processing method program. When the robot-based emotion perception collaborative processing method program is executed by the processor, it performs the following steps: The system acquires users' emotional feature data and preprocesses it to obtain effective emotional feature data. Acquire real-time computing load data at the robot edge and calculate the comprehensive computing load value. Adjust the preset dynamic lightweight model cluster according to the comprehensive computing load value to obtain an effective lightweight model. Effective emotional feature data is processed through an effective lightweight model to obtain multi-dimensional emotional feature vectors, and de-identification is performed according to the obtained user intention de-identification level to obtain de-identified multi-dimensional emotional feature vectors. Preliminary emotion categories are obtained by processing the desensitized multi-dimensional emotion feature vectors through a pre-set lightweight classification model. Preliminary emotion categories and desensitized multi-dimensional emotion feature vectors that meet the preset conditions are uploaded to the cloud for analysis to obtain refined emotion analysis data; The refined sentiment analysis data is distributed to the edge devices, and user sentiment characteristic variable data is collected and uploaded to the cloud for model optimization.
9. The robot-based emotion perception and collaborative processing system according to claim 8, characterized in that, The process of acquiring user emotional feature data and preprocessing the emotional feature data to obtain effective emotional feature data specifically includes: Acquire users' emotional characteristics data, including voice data, physiological data, and visual data; Timestamps are added to and aligned for voice, physiological and visual data using dual calibration technology of NTP network time protocol and local hardware clock; After preprocessing the emotional feature data, effective emotional feature data is obtained, including effective voice data, effective physiological data, and effective visual data. The preprocessing includes framing and windowing the speech data, filtering and denoising the physiological data, and normalizing the size of the visual data.
10. The robot-based emotion perception and collaborative processing system according to claim 9, characterized in that, The process of acquiring real-time computing load data at the robot's edge and calculating a comprehensive computing load value, then adjusting a preset dynamic lightweight model cluster based on the comprehensive computing load value to obtain an effective lightweight model, specifically includes: Acquire real-time computing load data at the robot's edge, including processor utilization, memory usage, and task difficulty data; The real-time computing load data is normalized and then weighted by preset weighting coefficients to obtain the comprehensive computing load value. The comprehensive computing load value is compared with the preset model scheduling threshold to obtain the corresponding effective lightweight model.