Companion robot control method and system based on emotion perception classification
Through multi-dimensional data collection and voice interaction, combined with external personnel intervention, the accuracy of the escort robot in emotional recognition and intervention was solved, and the quality of escort service and user experience were improved.
Patent Information
- Application Number
- CN202510812667.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The existing accompanying robots have low recognition accuracy and high misjudgment rates in terms of emotional interaction and emotional intervention, which cannot relieve negative emotions in a timely manner, affecting the quality of accompanying services and user experience.
Through multi-dimensional data collection and analysis of abnormal emotions, combined with voice interaction, accurately classify emotions types and degrees, actively link external personnel to conduct hierarchical intervention, including emotional regulation personnel and safety personnel, breaking through the limitations of robot emotional processing.
Accurate identification and efficient processing of negative emotions have been achieved, significantly improving the quality of accompanying services and user emotional care experience.
Smart Images

Figure CN120307308B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robotics technology, and in particular to a control method and system for a companion robot based on emotion perception classification. Background Art
[0002] Against the backdrop of people's growing demand for a higher quality of life, companion robots, as intelligent assistive devices, have been widely used in areas such as home care, health monitoring, and emotional companionship. Currently, existing companion robots primarily focus on basic daily care functions, such as medication reminders, mobility assistance, and vital sign monitoring, but lack significant capabilities in emotional interaction and emotional intervention.
[0003] Existing technologies have limited capabilities for sensing and processing user emotions, relying on simple, pre-set dialogue strategies or fixed interaction patterns. When a target person experiences negative emotions such as anxiety, depression, or loneliness, the robot is unable to accurately identify the emotion, making it even more difficult to implement effective intervention measures. Some robots with emotion recognition capabilities can only judge emotions based on a single dimension, such as voice tone and facial expression. This results in low recognition accuracy and high misjudgment rates. Furthermore, after identifying negative emotions, they lack a mechanism to proactively contact external personnel for collaborative intervention, preventing the timely introduction of more effective emotion regulation resources.
[0004] It can be seen that the existing control scheme of accompanying robots cannot alleviate the negative emotions of the target objects in a timely manner, seriously affecting the quality of accompanying services and user experience, and it is difficult to meet the users' urgent needs for emotional care and mental health support, and needs to be improved. Summary of the Invention
[0005] In this regard, the present invention provides a companion robot control method, system, electronic device, computer storage medium and computer program product based on emotion perception classification to solve at least one of the above technical problems.
[0006] In a first aspect, the present invention provides a companion robot control method based on emotion perception classification, which is applied to the companion robot and includes the following method steps: obtaining multi-dimensional emotion-related data of the target object, and deriving the probability of emotion abnormality judgment based on emotion-related data analysis; when the probability of emotion abnormality judgment is higher than a probability threshold, performing voice interaction with the target object, and using a multimodal emotion classification model to process the voice interaction data and emotion-related data to perform emotion perception classification on the target object, and obtain the emotion type and emotion degree; when the emotion type is an abnormal emotion and the emotion degree exceeds the control ability of the companion robot, determining external personnel, including emotion regulation personnel and / or security personnel, based on the emotion type and the emotion degree, and dispatching the external personnel to perform emotional intervention on the target object.
[0007] According to a second aspect of the present invention, a companion robot control system based on emotion perception classification is provided, which is applied to the companion robot. The system includes a first emotion perception module, a second emotion perception module, and a scheduling module. The first emotion perception module is used to obtain multi-dimensional emotion-related data of the target object, and derive the probability of emotion abnormality judgment based on the emotion-related data analysis. The second emotion perception module performs voice interaction with the target object when the probability of emotion abnormality judgment is higher than the probability threshold, and uses a multimodal emotion classification model to process the voice interaction data and emotion-related data to perform emotion perception classification on the target object to obtain the emotion type and emotion degree. The scheduling module determines external personnel, including emotion regulation personnel and / or security personnel, based on the emotion type and the emotion degree when the emotion type is an abnormal emotion and the emotion degree exceeds the control ability of the companion robot, and dispatches the external personnel to perform emotional intervention on the target object.
[0008] In a third aspect of the present invention, an electronic device is provided, comprising: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the computer program implements any of the methods described above when executed by the processor.
[0009] According to a fourth aspect of the present invention, a computer storage medium is provided, wherein the computer storage medium stores a computer program executable by a processor to implement any of the methods described above.
[0010] According to a fifth aspect of the present invention, a computer program product is provided, which comprises a computer program executable by a processor to implement any of the methods described above.
[0011] To address the inadequate emotional intervention capabilities of existing companion robots, this invention uses multi-dimensional data collection and analysis to analyze the probability of abnormal emotions, combined with voice interaction to accurately classify the type and severity of emotions. When the companion robot's control capabilities are exceeded, it proactively engages external personnel for graded intervention. This effectively avoids single-dimensional misjudgments, enabling accurate identification and efficient processing of negative emotions. Furthermore, it overcomes the limitations of companion robots' emotional processing capabilities, significantly improving the quality of companion services and the user's emotional care experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0013] Figure 1 This is a flow chart of a method for controlling a companion robot based on emotion perception classification disclosed in an embodiment of the present invention.
[0014] Figure 2 It is a structural diagram of the joint classification model of the solution of the present invention.
[0015] Figure 3 This is a structural diagram of a companion robot control system based on emotion perception and classification disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0016] The following specific embodiments illustrate the implementation of this application. Those familiar with the art can easily understand the other advantages and functions of this application from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of this application, but not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0017] In addition, the technical features involved in the different embodiments of the present application described below do not constitute each other.
[0018] like Figure 1 As shown, an embodiment of the present invention discloses a companion robot control method based on emotion perception classification, which is applied to the companion robot and includes the following method steps: Step 10, obtaining multi-dimensional emotion-related data of the target object, and deriving the probability of emotion abnormality judgment based on emotion-related data analysis.
[0019] The solution of the present invention is applied to a companion robot, which is equipped with detection equipment such as cameras and infrared detectors, as well as a processor and memory (with built-in computer program codes for executing the various analysis and processing functions of the present invention). In addition, it should also be equipped with walking parts, power supply parts, communication parts, etc. The details will not be repeated here.
[0020] In this step, the companion robot tracks and identifies the target object in the scene, using its own detection equipment to detect multi-dimensional emotion-related data about the target object, mainly including behavioral characteristics, facial expression characteristics, and voice characteristics. Behavioral characteristics include but are not limited to movement frequency (such as pacing and duration of silence); facial expression characteristics are micro-expressions captured by the camera, such as frowning and averted gaze; and voice characteristics include intonation fluctuations, speech rate changes, and keyword frequency.
[0021] Single-dimensional emotion recognition has significant limitations. For example, judging emotions solely through facial expressions can lead to misjudgments due to changes in ambient lighting or the subject's deliberate concealment. Relying on voice intonation can be affected by dialect, accent, or physical illness. Therefore, the present invention utilizes multi-dimensional data to perform preliminary identification of whether the subject is experiencing abnormal emotions, thereby improving the accuracy and reliability of emotion recognition and avoiding misjudgments or missed detections.
[0022] It is understandable that the above-mentioned probability of abnormal emotion judgment can be calculated using a machine learning algorithm. The machine learning algorithm can adopt a traditional machine learning algorithm, such as logistic regression, support vector machine, random forest, etc., or a deep learning algorithm, such as convolutional neural network, Transformer, attention-based multimodal fusion algorithm, etc., without specific limitation.
[0023] Step 20: When the probability of emotion abnormality is higher than the probability threshold, voice interaction is performed with the target object, and the multimodal emotion classification model is used to process the voice interaction data and emotion-related data to classify the target object's emotion perception and obtain the emotion type and emotion degree.
[0024] If the probability of emotional anomaly determined in step 10 indicates an abnormal risk (but without specifying the specific type and degree of emotion), the companion robot approaches the target subject and actively engages in voice interaction (e.g., "You don't look very happy today. Can you tell me what happened?"), encouraging the target subject to express their emotions. Voice interaction data is collected simultaneously during the interaction: semantic content (using natural language processing (NLP) to extract keywords such as "stress" and "loneliness") and emotional features of the voice (e.g., fundamental frequency and energy values reflecting anxiety levels). Voice interaction is an important way to obtain information about the target subject's subjective emotions. Through guided conversations, the target subject can actively express their feelings, complementing the shortcomings of passive monitoring data.
[0025] Inputting real-time voice interaction data (semantic content and emotional features) and historical emotion-related data (multi-dimensional emotion-related data collected in step 10) into a multimodal emotion classification model can fully leverage the complementary nature of each modality. The multimodal emotion classification model outputs emotion types and corresponding emotion levels. For example, emotion types can be classified as anxiety, depression, loneliness, anger, and other specific categories; emotion levels can be classified as mild, moderate, or severe (which can also be quantified numerically, such as on a scale of 1-10).
[0026] Step 30, when the emotion type is abnormal and the emotion level exceeds the control ability of the accompanying robot, external personnel, including emotion regulators and / or security personnel, are determined based on the emotion type and the emotion level, and the external personnel are dispatched to perform emotional intervention on the target object.
[0027] The care robot's emotional regulation capabilities for various types of abnormal emotions are pre-set, such that a certain care robot can only handle mild emotional issues. For example, for mild abnormal emotions, the robot can intervene by playing music, guiding breathing exercises, etc. However, for moderate to severe abnormal emotions, such as deep depression or severe anxiety, the care robot's configured methods are difficult to effectively regulate the target person's abnormal emotions, and may even worsen the abnormal emotions due to incorrect regulation.
[0028] In this regard, the present invention introduces a mechanism for collaborative intervention by external personnel, that is, when the emotion type is an abnormal emotion type, the accompanying robot compares the emotional degree of the abnormal emotion type based on the preset upper limit of the emotion regulation ability for the abnormal emotion type. If it exceeds the upper limit of the emotion regulation ability, the accompanying robot will no longer actively regulate the emotion, but will dispatch appropriate external personnel to deal with it.
[0029] Different external personnel possess varying professional capabilities and emotional support strengths. Relatives and friends can provide affection and friendship, while psychological counselors possess professional psychological counseling skills. By identifying target external personnel based on the type of abnormal emotion and the corresponding degree of emotion, resources can be precisely matched to ensure the effectiveness of intervention measures. For example, for depression, prioritizing contact with a psychological counselor for professional intervention is recommended; for loneliness, contacting relatives or friends for companionship is more appropriate. This hierarchical intervention mechanism, enabled by human-robot collaboration, can overcome the limitations of the robot's own capabilities, effectively alleviate the target person's negative emotions, and improve the quality of care services and user experience.
[0030] External personnel include emotional management personnel and / or security personnel. Emotional management personnel utilize emotional connections and communication skills to manage the target's abnormal emotions to a reasonable level, while security personnel are responsible for controlling any potential aggressive behavior. It is understood that security personnel will be dispatched when the target's emotions are anger, rage, or other indicators.
[0031] To address the inadequate emotional intervention capabilities of existing companion robots, this invention uses multi-dimensional data collection and analysis to analyze the probability of abnormal emotions, combined with voice interaction to accurately classify the type and severity of emotions. When the companion robot's control capabilities are exceeded, it proactively engages external personnel for graded intervention. This effectively avoids single-dimensional misjudgments, enabling accurate identification and efficient processing of negative emotions. Furthermore, it overcomes the limitations of companion robots' emotional processing capabilities, significantly improving the quality of companion services and the user's emotional care experience.
[0032] Furthermore, the probability of determining an abnormal emotion based on the analysis of emotion-related data includes: obtaining behavioral characteristics of other people in the same scene, calculating the similarity between the behavioral characteristics of each other person and the target object, and if the proportion of similarities above the similarity threshold exceeds the proportion threshold, lowering the weight coefficient of the behavioral characteristics; making emotional abnormality judgments based on the behavioral characteristics, facial expression characteristics, and voice characteristics in the emotion-related data, and fusing the judgment probability based on the behavioral characteristics with the other judgment probabilities according to the lowered weight coefficients to obtain the probability of determining an abnormal emotion.
[0033] The present invention uses a comprehensive analysis of the target subject's behavioral, facial, and voice characteristics to determine abnormal emotions. The resulting probabilities are then combined to derive a probability of abnormal emotion. This significantly improves the accuracy of initial identification of abnormal emotions. However, in settings such as nursing homes, various group activities are often held, and the behavioral characteristics of subjects in these group activities may have a high similarity to abnormal emotions, which can easily lead to misjudgments.
[0034] To this end, the present invention first identifies whether there is a collective activity in the scene, that is, calculates the similarity between the behavioral characteristics of other people in the same scene and the target object, and counts the proportion of similarities above the similarity threshold. If the proportion exceeds the threshold, it is determined that there is a high probability of a collective activity in the scene. At this time, even if the behavioral characteristics of the target object have a certain degree of similarity with abnormal emotions, there is a high probability that there is no abnormal emotion. Therefore, the present invention sets a lower weight coefficient for the behavioral characteristics, and then merges the judgment probability based on the behavioral characteristics with other judgment probabilities according to the lowered weight coefficient to obtain the judgment probability of abnormal emotion.
[0035] It is understandable that algorithms such as dynamic time warping (DTW) and cosine similarity are used to calculate the similarity between the behavioral characteristics of other people and the target object.
[0036] Furthermore, the use of the multimodal emotion classification model includes a first classification model, a second classification model and a joint classification model; the use of the multimodal emotion classification model to process voice interaction data and emotion-related data to perform emotion perception classification on the target object to obtain the emotion type and emotion degree includes: using the first classification model to perform emotion perception classification on the voice interaction data to obtain a first classification result; using the second classification model to perform emotion perception classification on the emotion-related data to obtain a second classification result; calculating the semantic deviation degree of the first classification result and the second classification result, and using the semantic deviation degree as a guide to use the joint classification model to perform emotion perception classification on the voice interaction data and emotion-related data to obtain the emotion type and the emotion degree.
[0037] When it comes to identifying abnormal emotions, unimodal data has limitations, primarily manifesting in the following: Voice interaction data is susceptible to subjective expression. While it can directly reflect subjective emotions, it carries the risk of being insincere; historical emotion-related data, on the other hand, is susceptible to environmental interference or deliberate concealment. To address these issues, the present invention employs the following approach: First, since speech is a direct expression of subjective emotion, it is necessary to prioritize capturing explicit emotional cues actively expressed by the target subject to quickly identify their subjective emotional tendencies. A first classification model is pre-built using, for example, an NLP algorithm. This model is then used to analyze the voice interaction data, extracting semantic keywords (e.g., "stress") and speech emotional features (e.g., accelerated speech rate), and outputting a first classification result, such as "suspected anxiety."
[0038] A pre-built second-classification model, such as one using CNN+LSTM, is used to analyze historical emotion-related data (behavior frequency, facial micro-expressions) and output a second-classification result, such as "abnormal behavior, possible anxiety." This objective behavioral clues compensate for the potential deceptiveness of voice data, forming a two-dimensional preliminary judgment.
[0039] Through the independently modeled first and second classification models, we obtain preliminary judgments on subjective expression and objective behavior, respectively, providing basic data for cross-modal comparison. It is understandable that both the first and second classification results contain the corresponding emotion type and emotion level.
[0040] Next, the semantic deviation between the first and second classification results is calculated to identify modal consistency. Specifically, a deviation evaluation function (such as KL divergence) is constructed to compare the emotion type (e.g., "anxiety" vs. "anger") and emotion intensity (e.g., "mild" vs. "moderate") between the first and second classification results, quantifying the difference as semantic deviation. A higher value for semantic deviation indicates greater modal conflict, while a lower value indicates less modal conflict.
[0041] For example, if the voice interaction data indicates “mild stress” (the first classification result), but the behavioral data indicates “frequent pacing + frowning” (the second classification result indicates “moderate anxiety”), then the semantic deviation is high, suggesting that there may be “inconsistency between expression and behavior.”
[0042] Next, guided by the semantic deviation calculated above, a joint classification model is used to perform more reasonable emotion perception classification on the voice interaction data and emotion-related data, thereby obtaining more accurate emotion types and emotion levels.
[0043] The present invention can achieve more accurate emotion perception classification by constructing the above-mentioned multimodal emotion classification system of "single-modal independent classification-semantic deviation calculation-joint dynamic fusion".
[0044] Furthermore, guided by the semantic deviation degree, a joint classification model is used to perform emotion perception classification on the voice interaction data and emotion-related data to obtain the emotion type and the emotion degree, including: if the semantic deviation degree is lower than the deviation threshold, the joint classification model directly fuses the first classification result and the second classification result according to the preset weight, and calculates the emotion type and the emotion degree; if the semantic deviation degree is higher than the deviation threshold, the joint classification model enhances the weight of the data with higher credibility in the voice interaction data and emotion-related data through the attention mechanism, and fuses the first classification result and the second classification result based on the data weight to calculate the emotion type and the emotion degree.
[0045] In this embodiment, if the semantic deviation degree is low (for example, both the voice and the behavior indicate "anxiety"), the two types of data can verify each other, have high credibility, and can directly enter the fusion stage. If the semantic deviation degree is high (for example, the voice denies emotional problems but the behavior is abnormal), the deep fusion mechanism needs to be triggered to avoid relying on a single modality and causing misjudgment. Specifically: If the semantic deviation degree is lower than the deviation threshold (modality consistency), the joint classification model uses, for example, a weighted summation algorithm to directly fuse the two types of data, and calculates the final emotion type and degree, such as "moderate anxiety," according to preset weights (for example, 40% for voice interaction data and 60% for emotion-related data).
[0046] If the semantic deviation is higher than the deviation threshold (modal conflict), the joint classification model automatically increases the weight of data with higher credibility in the conflicting modality through the attention mechanism (for example, emotion-related data is more reliable when the environment is stable), or introduces additional features (such as historical emotion trends) to assist in judgment, and finally outputs a revised classification result, such as "potential depressive tendency").
[0047] Furthermore, the joint classification model enhances the weight of data with higher credibility in voice interaction data and emotion-related data through the attention mechanism, including: constructing a word vector sequence for the voice interaction data, calculating the context association matrix through the self-attention mechanism, calculating the sentence coherence score, and adjusting the credibility weight of the voice interaction data based on the sentence coherence score; mapping the voice interaction data and emotion-related data to multiple subspaces respectively, calculating the modal consistency score of each subspace, and adjusting the credibility weight of the emotion-related data based on each modal consistency score; deriving the first fusion weight of the voice interaction data based on the adjusted credibility weight, and then deriving the second fusion weight of the emotion-related data.
[0048] In this embodiment, the semantic coherence of the voice interaction data is first quantified to identify possible "insincere" or contradictory expressions. Specifically, the voice interaction data is converted into voice interaction text, and the voice interaction text is divided into token sequences. , mapped to a sequence of word vectors through a pre-trained language model (such as BERT) ,in , is the vector dimension.
[0049] The self-attention mechanism is used to calculate the association strength between each token and other tokens: ;in, is the linear transformation result ( is the learnable parameter matrix); It is a scaling factor to prevent the dot product result from being too large and causing the gradient to disappear.
[0050] Finally, we get the correlation matrix ,in Indicates the association strength between token i and j.
[0051] Calculate the overall coherence score based on the correlation matrix: .in, is the word vector similarity (such as cosine similarity), which reflects the semantic association between tokens.
[0052] Through a monotonically increasing function Map coherence scores to credibility weights: .in, is a learnable parameter, The sigmoid function normalizes the score to the range [0, 1]. If the sentence coherence is low (for example, inconsistency), the credibility weight of the voice interaction data is reduced.
[0053] Next, we evaluate the consistency of voice interaction data and emotion-related data from multiple feature dimensions to identify possible modal conflicts. Specifically, we map the voice interaction data S and emotion-related data B (including behavior, micro-expressions, and voice features) into h subspaces using linear transformations: .in, is the projection matrix of each subspace.
[0054] For each subspace , calculate the consistency score between voice interaction data and emotion-related data: .
[0055] in, is the importance weight of each subspace, which is learned through training data.
[0056] Through a monotonically increasing function Mapping modality ratings to credibility weights .in, is a learnable parameter. If the multimodal data is expressed consistently in multiple subspaces, the credibility weight of the behavioral data is increased.
[0057] Finally, based on the two-dimensional credibility evaluation results, the final fusion weight is generated. Specifically: the credibility weight of the voice interaction data is Convert to the first fusion weight ;in, is a smoothing factor to prevent the denominator from being zero. Similarly, the credibility weight of sentiment-related data is Convert to the second fusion weight .in, .
[0058] The final emotion classification result is: .
[0059] It should be noted that the present invention generally performs the following process of fusing the first classification result and the second classification result based on the data weight: converting the two into a unified vector space representation, for example: the first classification result (voice interaction data) is represented as a probability distribution vector ,in Indicates that the voice interaction data is judged to be The probability of each emotion type (e.g., anxiety probability 0.7, depression probability 0.2, etc.). The second classification result (emotion-related data) is represented as a probability distribution vector ,in, Indicates that the behavioral characteristics and / or facial expression characteristics and / or voice characteristics are judged as the first The probability of the emotion type. The emotion degree is expressed as scalar and Indicates the emotional level of the first and second classification results (e.g., a 1-10 scale).
[0060] Use the first fusion weight and the second fusion weight Perform weighted fusion of two probability distributions: Among them, the probability of the final emotion type is .
[0061] Take a weighted average of the sentiment levels: .
[0062] Finally, from the fused probability distribution The emotion type with the highest probability is selected as the final classification result.
[0063] Furthermore, the calculation of the context association matrix through the self-attention mechanism includes: calculating the semantic density uniformity of the word vector sequence, and adjusting the local density-based dynamic scaling factor of the self-attention mechanism in the dot product stage based on the semantic density uniformity; the semantic density uniformity is determined based on the Euclidean distance variance between adjacent word vectors; and calculating the context association matrix through the self-attention mechanism according to the dynamic scaling factor.
[0064] In this embodiment, the present invention analyzes the semantic density characteristics of the word vector sequence and dynamically adjusts the scaling factor of the self-attention mechanism accordingly to further improve the accuracy of the context association matrix extracted by the self-attention mechanism.
[0065] First, for each pair of adjacent word vectors in the word vector sequence, the Euclidean distance is calculated to form a distance sequence. The variance of this distance sequence is then calculated to quantify the degree of distance fluctuation. A large variance indicates uneven semantic density. For example, the keyword "pressure" causes a sudden change in local density, as seen in the sentence "I'm fine" → "But I've been very stressed lately."
[0066] Standard self-attention uses a fixed scaling factor , which can lead to: excessive dispersion of attention in semantically dense areas (such as near sentiment keywords); and excessive amplification of associations between unrelated words in semantically sparse areas. To address this, the present invention uses a dynamic scaling factor based on local density to adjust the self-attention mechanism's scaling strategy for areas with uneven semantic density, preventing the dilution of key information.
[0067] For each word vector , based on its local neighborhood Dynamic scaling factor for semantic density calculation: ;in, Representing word vectors The k-nearest neighbor set of (for example, 3 words before and after ); is a learnable parameter that controls the sensitivity of the scaling factor to local density.
[0068] Therefore, when the local density is high (the distance is small), setting Reduce to enhance the attention weight of the area; when the local density is low (the distance is large), set Increase to suppress irrelevant correlations.
[0069] Furthermore, whether the emotional level of the emotional type exceeds the control ability of the accompanying robot is judged in the following manner, specifically: a pre-established control ability threshold table for different abnormal emotional types is retrieved, a matching search is performed to obtain the upper limit of the control ability of the accompanying robot corresponding to the emotional type, the upper limit of the control ability is compared with the emotional level, and a judgment is made as to whether the emotional level of the emotional type exceeds the control ability of the accompanying robot based on the comparison result.
[0070] A table of control ability thresholds for different abnormal emotion types is pre-established and stored in the memory of the accompanying robot, or stored in a server and can be called up by the accompanying robot at any time. As shown in Table 1 below:
[0071]
[0072] The upper limit of the ability to regulate abnormal emotions can be set based on the companion robot's hardware performance (e.g., interactive functions, knowledge base size), preset strategies (e.g., processing only mild emotions), etc. Periodic updates and adjustments should also be made based on the actual control performance of the companion robot. For example, if the preset upper limit of the anxiety control ability is 5 points, but actual control performance shows that it can effectively regulate more severe anxiety, the upper limit of the anxiety control ability can be adjusted to 6 points.
[0073] like Figure 3 As shown, an embodiment of the present invention further provides a companion robot control system 10 based on emotion perception classification, which is applied to the companion robot. The system 10 includes a first emotion perception module 102, a second emotion perception module 104, and a scheduling module 106; the first emotion perception module 102 is used to obtain multi-dimensional emotion-related data of the target object, and obtain the probability of emotion abnormality judgment based on the emotion-related data analysis; the second emotion perception module 104, when the emotion abnormality judgment probability is higher than the probability threshold, performs voice interaction with the target object, uses a multimodal emotion classification model to process the voice interaction data and emotion-related data to perform emotion perception classification on the target object, and obtains the emotion type and emotion degree; the scheduling module 106, when the emotion type is an abnormal emotion and the emotion degree exceeds the control ability of the companion robot, determines external personnel, including emotion regulation personnel and / or security personnel, based on the emotion type and the emotion degree, and dispatches external personnel to perform emotional intervention on the target object.
[0074] Furthermore, the first emotion perception module 102 is used to: obtain the behavioral characteristics of other people in the same scene, calculate the similarity between the behavioral characteristics of each other person and the target object, and if the proportion of similarities above the similarity threshold exceeds the proportion threshold, the weight coefficient of the behavioral characteristics is lowered; based on the behavioral characteristics, facial expression characteristics, and voice characteristics in the emotion-related data, perform emotional abnormality judgments respectively, and fuse the judgment probability based on the behavioral characteristics with other judgment probabilities according to the lowered weight coefficient to obtain the probability of emotional abnormality judgment.
[0075] Furthermore, the multimodal emotion classification model includes a first classification model, a second classification model and a joint classification model; the second emotion perception module 104 is used to: use the first classification model to perform emotion perception classification on the voice interaction data to obtain a first classification result; use the second classification model to perform emotion perception classification on the emotion-related data to obtain a second classification result; calculate the semantic deviation degree of the first classification result and the second classification result, and use the joint classification model to perform emotion perception classification on the voice interaction data and emotion-related data based on the semantic deviation degree to obtain the emotion type and the emotion degree.
[0076] Furthermore, the second emotion perception module 104 is used to: if the semantic deviation degree is lower than the deviation threshold, the joint classification model directly fuses the first classification result and the second classification result according to the preset weight, and calculates the emotion type and the emotion degree; if the semantic deviation degree is higher than the deviation threshold, the joint classification model uses the attention mechanism to enhance the weight of the data with higher credibility in the voice interaction data and emotion-related data, and fuses the first classification result and the second classification result based on the data weight to calculate the emotion type and the emotion degree.
[0077] Furthermore, the second emotion perception module 104 is used to: construct a word vector sequence for the voice interaction data, calculate the context association matrix through the self-attention mechanism, calculate the sentence coherence score, and adjust the credibility weight of the voice interaction data based on the sentence coherence score; map the voice interaction data and emotion-related data to multiple subspaces respectively, calculate the modal consistency score of each subspace, and adjust the credibility weight of the emotion-related data based on each modal consistency score; derive the first fusion weight of the voice interaction data based on the adjusted credibility weight, and then derive the second fusion weight of the emotion-related data.
[0078] Furthermore, the second emotion perception module 104 is used to: calculate the semantic density uniformity of the word vector sequence, and adjust the local density-based dynamic scaling factor of the self-attention mechanism in the dot product stage based on the semantic density uniformity; the semantic density uniformity is determined based on the Euclidean distance variance between adjacent word vectors; and according to the dynamic scaling factor, calculate the context association matrix through the self-attention mechanism.
[0079] Furthermore, the scheduling module 106 is used to: retrieve a pre-established table of control ability thresholds for different abnormal emotion types, match and search to obtain the upper limit of the control ability of the accompanying robot corresponding to the emotion type, compare the upper limit of the control ability with the emotion level, and judge whether the emotion level of the emotion type exceeds the control ability of the accompanying robot based on the comparison result.
[0080] An embodiment of the present invention further provides an electronic device comprising: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the computer program implements any of the aforementioned methods when executed by the processor.
[0081] An embodiment of the present invention further provides a computer storage medium storing a computer program that can be executed by a processor to implement any of the methods described above.
[0082] An embodiment of the present invention further provides a computer program product, which includes a computer program that can be executed by a processor to implement any of the methods described above.
[0083] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A control method for a companion robot based on emotion perception and classification, applied to a companion robot, characterized by: The method comprises the following steps: obtaining multi-dimensional emotion-related data of a target object, and determining a probability of abnormal emotion determination based on analysis of the emotion-related data; when the probability of abnormal emotion determination is higher than a probability threshold, performing voice interaction with the target object, and using a multimodal emotion classification model to process the voice interaction data and emotion-related data to classify the target object's emotion perception, thereby obtaining an emotion type and an emotion degree; when the emotion type is abnormal and the emotion degree exceeds the control capability of the accompanying robot, determining an external person, including an emotion control person and / or a security person, based on the emotion type and the emotion degree, and dispatching the external person to perform emotional intervention on the target object; The multimodal emotion classification model includes a first classification model, a second classification model and a joint classification model; Using a multimodal emotion classification model to process voice interaction data and emotion-related data to perform emotion perception classification on a target object to obtain an emotion type and an emotion degree, including: using a first classification model to perform emotion perception classification on the voice interaction data to obtain a first classification result; using a second classification model to perform emotion perception classification on the emotion-related data to obtain a second classification result; calculating a semantic deviation degree between the first classification result and the second classification result, and using the semantic deviation degree as a guide, using a joint classification model to perform emotion perception classification on the voice interaction data and emotion-related data to obtain the emotion type and the emotion degree; Guided by the semantic deviation degree, a joint classification model is used to perform emotion perception classification on the voice interaction data and emotion-related data to obtain the emotion type and the emotion degree, including: if the semantic deviation degree is lower than the deviation threshold, the joint classification model directly fuses the first classification result and the second classification result according to the preset weight, and calculates the emotion type and the emotion degree; if the semantic deviation degree is higher than the deviation threshold, the joint classification model uses the attention mechanism to enhance the weight of data with higher credibility in the voice interaction data and emotion-related data, and fuses the first classification result and the second classification result based on the data weight to calculate the emotion type and the emotion degree.
2. The control method of a companion robot based on emotion perception and classification according to claim 1, characterized in that: The probability of emotion abnormality judgment is obtained based on the analysis of emotion-related data, including: obtaining the behavioral characteristics of other people in the same scene, calculating the similarity between the behavioral characteristics of each other person and the target object, and if the proportion of similarities above the similarity threshold exceeds the proportion threshold, the weight coefficient of the behavioral characteristic is lowered; based on the behavioral characteristics, facial expression characteristics, and voice characteristics in the emotion-related data, emotion abnormality judgment is performed separately, and the judgment probability based on the behavioral characteristics is integrated with the other judgment probabilities according to the lowered weight coefficient to obtain the probability of emotion abnormality judgment.
3. The control method of a companion robot based on emotion perception and classification according to claim 1, characterized in that: The joint classification model uses the attention mechanism to enhance the weight of data with higher credibility in voice interaction data and emotion-related data, including: constructing a word vector sequence for voice interaction data, calculating the context association matrix through the self-attention mechanism, calculating the sentence coherence score, and adjusting the credibility weight of the voice interaction data based on the sentence coherence score; mapping the voice interaction data and emotion-related data to multiple subspaces respectively, calculating the modal consistency score of each subspace, and adjusting the credibility weight of the emotion-related data based on the modal consistency score; deriving the first fusion weight of the voice interaction data based on the adjusted credibility weight, and then deriving the second fusion weight of the emotion-related data.
4. The method for controlling a companion robot based on emotion perception and classification according to claim 3, characterized in that: The method of calculating the context association matrix through the self-attention mechanism includes: calculating the semantic density uniformity of the word vector sequence, and adjusting the local density-based dynamic scaling factor of the self-attention mechanism in the dot product stage based on the semantic density uniformity; the semantic density uniformity is determined based on the Euclidean distance variance between adjacent word vectors; and calculating the context association matrix through the self-attention mechanism according to the dynamic scaling factor.
5. The control method of a companion robot based on emotion perception and classification according to claim 1, characterized in that: Whether the emotional level of the emotional type exceeds the control ability of the accompanying robot is judged in the following manner. Specifically, a pre-established control ability threshold table for different abnormal emotional types is retrieved, and a matching search is performed to obtain the upper limit of the control ability of the accompanying robot corresponding to the emotional type, and the upper limit of the control ability is compared with the emotional level. Based on the comparison result, a judgment is made as to whether the emotional level of the emotional type exceeds the control ability of the accompanying robot.
6. A companion robot control system based on emotion perception and classification, applied to a companion robot, characterized in that: The system includes a first emotion perception module, a second emotion perception module, and a scheduling module; the first emotion perception module is used to obtain multi-dimensional emotion-related data of the target object and obtain the probability of emotion abnormality based on the emotion-related data analysis; the second emotion perception module, when the probability of emotion abnormality is higher than the probability threshold, performs voice interaction with the target object and uses a multimodal emotion classification model to process the voice interaction data and emotion-related data to perform emotion perception classification on the target object and obtain the emotion type and emotion degree; The scheduling module determines external personnel, including emotion regulators and / or security personnel, based on the emotion type and the emotion level when the emotion type is abnormal and the emotion level exceeds the control capability of the accompanying robot, and dispatches the external personnel to perform emotional intervention on the target object; The multimodal emotion classification model includes a first classification model, a second classification model and a joint classification model; the second emotion perception module is used to: use the first classification model to perform emotion perception classification on the voice interaction data to obtain a first classification result; Use the second classification model to perform emotion perception classification on the emotion-related data to obtain the second classification result; Calculating the semantic deviation degree of the first classification result and the second classification result, and using the semantic deviation degree as a guide, using a joint classification model to perform emotion perception classification on the voice interaction data and the emotion-related data to obtain the emotion type and the emotion degree; The second emotion perception module is configured to: if the semantic deviation is lower than the deviation threshold, directly fuse the first classification result and the second classification result using a joint classification model according to a preset weight to calculate the emotion type and the emotion degree; If the semantic deviation is higher than the deviation threshold, the joint classification model uses the attention mechanism to increase the weight of the more credible data in the voice interaction data and emotion-related data, and fuses the first classification results and the second classification results based on the data weight to calculate the emotion type and the emotion degree.
7. An electronic device, characterized in that: The electronic device includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the computer program implements the method according to any one of claims 1 to 5 when executed by the processor.
8. A computer storage medium, characterized in that: The computer storage medium stores a computer program that can be executed by a processor to implement the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Emotion monitoring and early warning method and system
CN106361356A
Intelligent emotion recognition method based on multi-modal data fusion
CN118709146A