Accompanying robot control method and system based on emotion perception classification
The method and system for emotion recognition and classification in companion robots improve emotional support by using multi-dimensional data analysis and external interventions, addressing the limitations of single-dimensional analysis and enhancing user experience.
Patent Information
- Application Number
- CN202510812667.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The existing accompanying robots have low recognition accuracy and high misjudgment rates in terms of emotional interaction and emotional intervention, and cannot relieve negative emotions in a timely manner. They lack a mechanism to actively contact external personnel for collaborative intervention, which affects the quality of accompanying services and user experience.
Through multi-dimensional data collection, abnormal emotions probability is analyzed, combined with voice interaction, and accurately classify emotions types and degrees. When the regulatory capabilities of the accompanying robot are exceeded, external personnel will be actively linked to hierarchical intervention, including the dispatch of emotional regulation personnel and safety personnel.
It realizes accurate identification and efficient processing of negative emotions, breaks through the limitations of emotional processing of accompanying robots, and significantly improves the quality of accompanying services and user emotional care experience.
Smart Images

Figure CN120307308A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robots, and more particularly, to a control method and system for a companion robot based on emotion perception classification. Background Art
[0002] In the context of the growing demand for high-quality life, companion robots, as intelligent auxiliary devices, have been widely used in the fields of home care, health monitoring, and emotional companionship. At present, existing companion robots mainly focus on basic life care functions, such as reminding to take medicine, assisting with movement, monitoring vital signs, etc., and there are significant deficiencies in emotional interaction and emotion intervention.
[0003] In the prior art, the ability of companion robots to perceive and process users' emotions is limited, mostly relying on preset simple conversation strategies or fixed interaction modes. When the target object generates negative emotions such as anxiety, depression, and loneliness, the companion robot cannot accurately identify the emotion category, let alone take effective intervention measures. Some robots with emotion recognition functions can only judge emotions through a single dimension such as voice intonation and facial expressions, resulting in low recognition accuracy and high misjudgment rate. Moreover, after recognizing negative emotions, they lack a mechanism to actively contact external personnel for collaborative intervention and cannot introduce more effective emotion regulation resources in a timely manner.
[0004] It can be seen that the existing control schemes for companion robots cannot relieve the negative emotions of the target object in a timely manner, seriously affecting the quality of companion services and the user experience, and it is difficult to meet the urgent needs of users for emotional care and mental health support, thus improvement is needed. Summary of the Invention
[0005] In view of this, the present invention provides a control method, system, electronic device, computer storage medium, and computer program product for a companion robot based on emotion perception classification to solve at least one of the above technical problems.
[0006] In the first aspect of the present invention, there is provided a control method for a companion robot based on emotion perception classification, which is applied to the companion robot and includes the following method steps: obtaining multi-dimensional emotion-related data of the target object, and analyzing the emotion-related data to obtain the probability of emotion abnormality determination; when the probability of emotion abnormality determination is higher than the probability threshold, performing a voice interaction with the target object, and using a multi-modal emotion classification model to process the voice interaction data and emotion-related data to perform emotion perception classification on the target object, obtaining the emotion type and emotion degree; when the emotion type is an abnormal emotion and the emotion degree exceeds the regulation ability of the companion robot, determining external personnel based on the emotion type and the emotion degree, including emotion regulation personnel and / or security personnel, and dispatching the external personnel to perform emotion intervention on the target object.
[0007] In the second aspect of the present invention, a control system for a companion robot based on emotion perception classification is provided, which is applied to a companion robot. The system includes a first emotion perception module, a second emotion perception module, and a scheduling module. The first emotion perception module is configured to obtain multi-dimensional emotion-related data of a target object and analyze an emotion abnormality determination probability based on the emotion-related data. The second emotion perception module is configured to, when the emotion abnormality determination probability is higher than a probability threshold, perform a voice interaction with the target object, and process the voice interaction data and the emotion-related data using a multi-modal emotion classification model to perform emotion perception classification on the target object, so as to obtain an emotion type and an emotion degree. The scheduling module is configured to, when the emotion type is an abnormal emotion and the emotion degree exceeds the regulation ability of the companion robot, determine external personnel, including emotion regulation personnel and / or security personnel, based on the emotion type and the emotion degree, and schedule the external personnel to perform emotion intervention on the target object.
[0008] In the third aspect of the present invention, an electronic device is provided. The electronic device includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the computer program is executed by the processor, the method described in any one of the foregoing items is implemented.
[0009] In the fourth aspect of the present invention, a computer storage medium is provided. The computer storage medium stores a computer program that can be executed by a processor to implement the method described in any one of the foregoing items.
[0010] In the fifth aspect of the present invention, a computer program product is provided. The computer program product includes a computer program that can be executed by a processor to implement the method described in any one of the foregoing items.
[0011] Aiming at the problem of insufficient emotion intervention ability of existing companion robots, the present invention analyzes the probability of abnormal emotions through multi-dimensional data collection, combines voice interaction to accurately classify the emotion type and degree, and actively links external personnel for hierarchical intervention when it exceeds the regulation ability of the companion robot. On the one hand, it can effectively avoid misjudgment by single-dimensional recognition and achieve accurate recognition and efficient processing of negative emotions. On the other hand, it can break through the emotional processing limitations of companion robots and significantly improve the quality of companion services and the user's emotional care experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0013] Figure 1 It is a schematic flowchart of a control method for a companion robot based on emotion perception classification disclosed in an embodiment of the present invention.
[0014] Figure 2 It is a schematic structural diagram of the joint classification model of the solution of the present invention.
[0015] Figure 3 It is a schematic structural diagram of a control system for a companion robot based on emotion perception classification disclosed in an embodiment of the present invention. Specific embodiments
[0016] The following specific embodiments illustrate the implementation manners of the present application. Those skilled in the art can easily understand the other advantages and effects of the present application from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0017] In addition, the technical features involved in different implementation manners of the present application described below do not constitute each other.
[0018] Such as Figure 1 As shown, an embodiment of the present invention discloses a control method for a companion robot based on emotion perception classification, which is applied to a companion robot and includes the following method steps: Step 10, obtaining multi-dimensional emotion-related data of a target object, and analyzing an emotion abnormality determination probability based on the emotion-related data.
[0019] The solution of the present invention is applied to a companion robot. The companion robot is equipped with detection devices such as a camera and an infrared detector, as well as a processor and a memory (built-in with computer program codes for performing various analysis and processing functions of the present invention). In addition, it should also be equipped with walking components, power supply components, communication components, etc., which will not be elaborated here.
[0020] In this step, the companion robot tracks and identifies the target object in the scene it is in, and uses the detection devices it is equipped with to detect multi-dimensional emotion-related data of the target object, mainly including behavioral features, facial expression features, and voice features. Among them, the behavioral features include but are not limited to action frequencies (such as pacing, silence duration); the facial expression features are micro-expressions captured by the camera, such as frowning and eye avoidance; the voice features include intonation fluctuations, speech rate changes, keyword frequencies, etc.
[0021] Due to the significant limitations of single-dimensional emotion recognition. For example, judging emotions solely by facial expressions may lead to misjudgments due to changes in environmental light or deliberate concealment by the target object; relying on voice intonation may be affected by dialects, accents, or physical illnesses. Therefore, in the present invention, the companion robot uses the collected multi-dimensional data to preliminarily identify whether the target object has abnormal emotions, so as to improve the accuracy and reliability of abnormal emotion recognition and avoid misjudgments or missed judgments.
[0022] It can be understood that the above-mentioned probability of abnormal emotion determination can be calculated using machine learning algorithms. The machine learning algorithms can adopt traditional machine learning algorithms, such as logistic regression, support vector machines, random forests, etc., or can also be deep learning algorithms, such as convolutional neural networks, Transformers, attention-based multi-modal fusion algorithms, etc., and are not specifically limited.
[0023] Step 20, when the probability of abnormal emotion determination is higher than the probability threshold, conduct a voice interaction with the target object, and use the multi-modal emotion classification model to process the voice interaction data and emotion-related data to perform emotion perception classification on the target object, obtaining the emotion type and emotion degree.
[0024] When the probability of abnormal emotion determination obtained in the previous step 10 indicates an abnormal risk (but without clarifying the specific emotion type and degree), the companion robot approaches the target object and takes the initiative to conduct a voice interaction with it (such as "You don't seem very happy today. Can you tell me what happened?"), encouraging the target object to express emotions. During the interaction, voice interaction data is collected synchronously: semantic content (extracting keywords through natural language processing NLP, such as "stress", "loneliness"), and voice emotional features (such as fundamental frequency, energy value reflecting the degree of anxiety). Voice interaction is an important way to obtain the subjective emotion information of the target object. Through guided conversations, the target object can be encouraged to actively express their inner feelings and supplement the deficiencies of passive monitoring data.
[0025] Input the real-time voice interaction data (semantic content, voice emotional features) and historical emotion-related data (the multi-dimensional emotion-related data collected in step 10) into the multi-modal emotion classification model, which can make full use of the complementarity of each modal data. The multi-modal emotion classification model outputs the emotion type and the corresponding emotion degree. For example: Emotion types: specific categories such as anxiety, depression, loneliness, anger, etc.; Emotion degrees: mild, moderate, severe (can also be quantified numerically, such as 1-10 points).
[0026] Step 30, when the emotion type is an abnormal emotion and the emotion degree exceeds the regulation ability of the companion robot, determine external personnel based on the emotion type and the emotion degree, including emotion regulation personnel and / or security personnel, and dispatch external personnel to conduct emotion intervention on the target object.
[0027] Preset the upper limit of the emotional regulation ability of the companion robot for various abnormal emotion types. For example, a certain companion robot can only handle mild emotional problems. For example: for mild abnormal emotions, the robot can intervene by playing music, guiding breathing training, etc.; but in the face of moderate to severe abnormal emotions, such as deep depression or severe anxiety, the means configured by the companion robot are difficult to effectively regulate the abnormal emotions of the target object, and there is also a possibility of worsening the abnormal emotions due to incorrect regulation.
[0028] In response to this, the present invention sets up an external personnel collaborative intervention mechanism, that is, when the emotion type is an abnormal emotion type, the companion robot compares the emotional degree of the abnormal emotion type based on the preset upper limit of the emotional regulation ability for the abnormal emotion type. If it exceeds the upper limit of the emotional regulation ability, the companion robot no longer actively conducts emotional regulation, but schedules appropriate external personnel to come for handling.
[0029] Different external personnel have different professional capabilities and emotional support advantages. Relatives and friends can provide family and friendship care, while psychological counselors have professional psychological counseling skills. Determining the target external personnel according to the abnormal emotion type and the corresponding emotional degree can achieve precise matching of resources and ensure the effectiveness of the intervention measures. For example, for depressive emotions, it is preferable to contact a psychological counselor for professional intervention; for lonely emotions, it is more appropriate to contact relatives or friends for company. Through this human-machine collaborative hierarchical intervention mechanism, the limitations of the robot's own capabilities can be broken through, the negative emotions of the target object can be effectively relieved, and the quality of the companion service and the user experience can be improved.
[0030] The external personnel include emotion regulation personnel and / or security personnel. Among them, the emotion regulation personnel are used to regulate the abnormal emotions of the target object to a reasonable level by using emotional relationships, communication skills, etc., while the security personnel are used to control the excessive behaviors that the target object may implement. It can be understood that when the emotion type is anger, mania, etc., the scheduling of the security personnel is activated.
[0031] Aiming at the problem of insufficient emotional intervention ability of existing companion robots, the present invention analyzes the probability of abnormal emotions through multi-dimensional data collection, combines voice interaction to accurately classify the emotion type and degree, and actively links external personnel for hierarchical intervention when it exceeds the regulation ability of the companion robot. On the one hand, it can effectively avoid misjudgment by single-dimensional recognition and achieve accurate recognition and efficient processing of negative emotions; on the other hand, it can break through the emotional processing limitations of the companion robot and significantly improve the quality of the companion service and the user's emotional care experience.
[0032] Further, the determination probability of emotional abnormality obtained based on emotion-related data analysis includes: obtaining the behavioral characteristics of other persons in the same scenario, calculating the similarity between the behavioral characteristics of each other person and the target object, and if the proportion of the similarity higher than the similarity threshold exceeds the proportion threshold, then reducing the weight coefficient of the behavioral characteristics; respectively performing emotional abnormality determination based on the behavioral characteristics, facial expression characteristics, and voice characteristics in the emotion-related data, and fusing the determination probability based on the behavioral characteristics and other determination probabilities according to the reduced weight coefficient to obtain the determination probability of emotional abnormality.
[0033] The present invention comprehensively performs emotional abnormality determination on the behavioral characteristics, facial expression characteristics, and voice characteristics of the target object respectively, and then fuses the determination probabilities obtained based on the analysis of each characteristic to obtain the determination probability of emotional abnormality, which can significantly improve the accuracy of the initial recognition of abnormal emotions. However, in scenarios such as nursing homes, various group activities are often held, and the behavioral characteristics of the objects in the group activities may have a high similarity to abnormal emotions, which is likely to lead to misjudgment.
[0034] In this regard, the present invention first identifies whether there is a group activity in the current scenario, that is, calculates the similarity between the behavioral characteristics of other persons and the target object in the same scenario, and counts the proportion of the similarity higher than the similarity threshold. If this proportion exceeds the proportion threshold, it is determined that there is a high probability of a group activity in the current scenario. At this time, even if the behavioral characteristics of the target object have a certain degree of similarity to abnormal emotions, it is very likely that there is no emotional abnormality. Therefore, the present invention sets to reduce the weight coefficient of the behavioral characteristics, and then fuses the determination probability based on the behavioral characteristics and other determination probabilities according to the reduced weight coefficient to obtain the determination probability of emotional abnormality.
[0035] It can be understood that algorithms such as dynamic time warping (DTW) and cosine similarity are used to calculate the similarity between the behavioral characteristics of other persons and the target object.
[0036] Further, the multi-modal emotion classification model used includes a first classification model, a second classification model, and a joint classification model; the multi-modal emotion classification model is used to process voice interaction data and emotion-related data to perform emotion perception classification on the target object to obtain the emotion type and emotion degree, including: using the first classification model to perform emotion perception classification on the voice interaction data to obtain a first classification result; using the second classification model to perform emotion perception classification on the emotion-related data to obtain a second classification result; calculating the semantic deviation degree between the first classification result and the second classification result, and using the joint classification model to perform emotion perception classification on the voice interaction data and the emotion-related data under the guidance of the semantic deviation degree to obtain the emotion type and the emotion degree.
[0037] In abnormal emotion recognition, unimodal data has limitations, which are mainly reflected in: voice interaction data is vulnerable to subjective expression deception. Although voice interaction data can directly reflect subjective emotions, there is a risk of "saying one thing but meaning another"; while historical emotion-related data is vulnerable to environmental interference or deliberate concealment. To address the above problems, the present invention adopts the following solutions, specifically as follows: First, voice is a direct expression of subjective emotions, and it is necessary to first capture the obvious emotional clues actively expressed by the target object to quickly locate the subjective emotional tendency. For example, use NLP algorithms to pre-construct a first classification model, and use the first classification model to analyze voice interaction data, extract semantic keywords (such as "stress") and voice emotional features (such as increased speech rate), and output the first classification result, such as "suspected anxiety".
[0038] Adopt, for example, CNN+LSTM to pre-construct a second classification model, and use the second classification model to analyze historical emotion-related data (behavior frequency, facial micro-expressions), and output the second classification result, such as "abnormal behavior, possible anxiety". Make up for the potential deception of voice data through objective behavior clues to form a two-dimensional preliminary judgment.
[0039] Through the above first classification model and second classification model independently modeled, preliminary judgments of subjective expression and objective behavior are respectively obtained, providing basic data for cross-modal comparison. It can be understood that both the first classification result and the second classification result contain the corresponding emotion type and emotion degree.
[0040] Next, calculate the semantic deviation degree between the first classification result and the second classification result to identify modal consistency. Specifically: construct a deviation degree evaluation function (such as KL divergence), compare the differences in emotion types (such as "anxiety" vs "anger") and emotion degrees (such as "mild" vs "moderate") between the first classification result and the second classification result, and quantify them as the semantic deviation degree. Among them, the higher the value of the semantic deviation degree, the greater the modal conflict, and vice versa, the smaller the modal conflict.
[0041] For example: If the voice interaction data determines "mild stress" (the first classification result), but the behavior data shows "frequent pacing + frowning" (the second classification result determines "moderate anxiety"), then the semantic deviation degree is relatively high, indicating that there may be "inconsistency between expression and behavior".
[0042] Next, guided by the above calculated semantic deviation degree, use a joint classification model to perform a more reasonable emotion perception classification on voice interaction data and emotion-related data, so as to obtain a more accurate emotion type and emotion degree.
[0043] The present invention can achieve a more accurate emotion perception classification by constructing the above multi-modal emotion classification system of "unimodal independent classification - semantic deviation degree calculation - joint dynamic fusion".
[0044] Further, guided by the semantic deviation degree, a joint classification model is used to perform emotion perception classification on speech interaction data and emotion-related data to obtain the emotion type and the emotion degree, including: if the semantic deviation degree is lower than the deviation threshold, the joint classification model directly fuses the first classification result and the second classification result according to a preset weight, and calculates the emotion type and the emotion degree; if the semantic deviation degree is higher than the deviation threshold, the joint classification model enhances the weight of the data with higher credibility in the speech interaction data and the emotion-related data through an attention mechanism, and fuses the first classification result and the second classification result based on this data weight, and calculates the emotion type and the emotion degree.
[0045] In this embodiment, if the above semantic deviation degree belongs to a low deviation degree (for example, both speech and behavior point to "anxiety"), it means that the two types of data can verify each other, with high credibility, and can directly enter the fusion stage; while if the above semantic deviation degree belongs to a high deviation degree (for example, speech denies emotional problems but behavior is abnormal), at this time, a deep fusion mechanism needs to be triggered to avoid misjudgment caused by relying on a single modality. Specifically: if the semantic deviation degree is lower than the deviation threshold (modal consistency), the joint classification model directly fuses the two types of data using, for example, a weighted summation algorithm, and calculates the final emotion type and degree according to a preset weight (for example, speech interaction data accounts for 40% and emotion-related data accounts for 60%), such as "moderate anxiety".
[0046] If the semantic deviation degree is higher than the deviation threshold (modal conflict), the joint classification model automatically enhances the weight of the data with higher credibility in the conflicting modality through an attention mechanism (for example, emotion-related data is more reliable when the environment is stable), or introduces additional features (such as historical emotion trends) to assist in judgment, and finally outputs the corrected classification result, such as "potential depressive tendency").
[0047] Further, the joint classification model enhances the weight of the data with higher credibility in the speech interaction data and the emotion-related data through an attention mechanism, including: constructing a word vector sequence for the speech interaction data, calculating a context correlation matrix through a self-attention mechanism, calculating a sentence coherence score, and adjusting the credibility weight of the speech interaction data based on the sentence coherence score; mapping the speech interaction data and the emotion-related data into multiple subspaces respectively, calculating the modal consistency score of each subspace, and adjusting the credibility weight of the emotion-related data based on each modal consistency score; obtaining the first fusion weight of the speech interaction data based on the adjusted credibility weight, and further obtaining the second fusion weight of the emotion-related data.
[0048] In this embodiment, first, the semantic coherence of the voice interaction data is quantified to identify possible cases of "saying one thing but meaning another" or contradictory expressions. Specifically, the voice interaction data is converted into voice interaction text, and the voice interaction text is segmented into a token sequence , and mapped into a sequence of word vectors through a pre-trained language model (such as BERT) , where , is the vector dimension.
[0049] The association strength between each token and other tokens is calculated through the self-attention mechanism: ; where is the result of the linear transformation ( is a learnable parameter matrix); is the scaling factor to prevent the dot product result from being too large and causing the gradient to vanish.
[0050] Finally, the association matrix is obtained, where represents the association strength between token i and j.
[0051] Based on the association matrix, the overall coherence score is calculated: . Where is the word vector similarity (such as cosine similarity), reflecting the semantic association degree between tokens.
[0052] Through the monotonically increasing function the coherence score is mapped to a credibility weight: . Where is a learnable parameter, is the sigmoid function, which normalizes the score to the [0,1] interval. If the statement coherence is low (such as contradictory), the credibility weight of the voice interaction data is reduced.
[0053] Next, the consistency between the voice interaction data and the emotion-related data is evaluated from multiple feature dimensions to identify possible modality conflicts. Specifically, the voice interaction data S and the emotion-related data B (including behavior, micro-expression, voice features) are respectively mapped to h subspaces through linear transformation: . Where is the projection matrix of each subspace.
[0054] For each subspace , the consistency score between the voice interaction data and the emotion-related data is calculated: .
[0055] Where is the importance weight of each subspace, learned from the training data.
[0056] Map the modality score to a credibility weight through a monotonically increasing function where are learnable parameters. If the multi-modal data has consistent representations in multiple subspaces, the credibility weight of the behavior data is increased.
[0057] Finally, based on the two-dimensional credibility assessment results, generate the final fusion weight. Specifically: convert the credibility weight of the voice interaction data to the first fusion weight where is a smoothing factor to prevent the denominator from being zero. Similarly, convert the credibility weight of the emotion-related data to the second fusion weight where
[0058] The final emotion classification result is:
[0059] It should be noted that the general process of the present invention for fusing the first classification result and the second classification result based on the data weights is as follows: convert the two into a unified vector space representation. For example, the first classification result (voice interaction data) is represented as a probability distribution vector where represents the probability that the voice interaction data is judged to be the th emotion type (e.g., anxiety probability 0.7, depression probability 0.2, etc.). The second classification result (emotion-related data) is represented as a probability distribution vector where represents the probability that the behavioral features and / or facial expression features and / or voice features are judged to be the th emotion type. The emotion degrees are represented by scalars and respectively for the emotion degrees of the first classification result and the second classification result (e.g., on a 1-10 scale).
[0060] Use the first fusion weight and the second fusion weight to perform weighted fusion on the two probability distributions: where the probability of the final emotion type is
[0061] Perform a weighted average on the emotion degrees:
[0062] Finally, select the emotion type with the highest probability from the fused probability distribution as the final classification result.
[0063] Further, calculating the context correlation matrix through the self-attention mechanism includes: calculating the semantic density uniformity of the word vector sequence, and adjusting the dynamic scaling factor based on local density in the dot product stage of the self-attention mechanism based on the semantic density uniformity; the semantic density uniformity is determined based on the Euclidean distance variance between adjacent word vectors; calculating the context correlation matrix through the self-attention mechanism according to the dynamic scaling factor.
[0064] In this embodiment, the present invention analyzes the semantic density characteristics of the word vector sequence and dynamically adjusts the scaling factor of the self-attention mechanism accordingly to further improve the accuracy of the context correlation matrix obtained by the self-attention mechanism.
[0065] First, for each adjacent word vector pair in the word vector sequence, calculate its Euclidean distance to form a distance sequence; then calculate the variance of the distance sequence to quantify the degree of distance fluctuation. When the variance is large, it represents uneven semantic density. For example: "I'm fine" → "But recently I've been under a lot of pressure", the keyword "pressure" causes a local density mutation.
[0066] Standard self-attention uses a fixed scaling factor , which may lead to: in semantic dense regions (such as near sentiment keywords), the attention is too scattered; in semantic sparse regions, the associations between irrelevant words are over-amplified. In response, the present invention adopts a dynamic scaling factor based on local density to adjust the scaling strategy of the self-attention mechanism for uneven semantic density regions to avoid dilution of key information.
[0067] For each word vector , calculate the dynamic scaling factor based on the semantic density of its local neighborhood : ; where represents the k-nearest neighbor set of the word vector (for example, the first 3 and last 3 words); is a learnable parameter that controls the sensitivity of the scaling factor to local density.
[0068] Thus, when the local density is high (distance is small), set to decrease to enhance the attention weight in this region; when the local density is low (distance is large), set to increase to suppress irrelevant associations.
[0069] Further, the following method is used to determine whether the emotional intensity of the emotional type exceeds the regulation ability of the companion robot. Specifically: retrieve the pre-established regulation ability threshold table for different abnormal emotional types, match and find the upper limit of the regulation ability of the companion robot corresponding to the emotional type, compare the upper limit of the regulation ability with the emotional intensity, and make a judgment on whether the emotional intensity of the emotional type exceeds the regulation ability of the companion robot according to the comparison result.
[0070] The pre-established regulation ability threshold table for different abnormal emotional types is stored in the memory of the companion robot or stored in the server and can be called by the companion robot at any time. As shown in Table 1 below:
[0071] The upper limit of the regulation ability for abnormal emotions can be set according to the hardware performance of the companion robot (such as interaction function, knowledge base scale), preset strategies (such as only dealing with mild emotions), etc. At the same time, it is also necessary to focus on the actual regulation performance of the companion robot to update and adjust periodically. For example, the preset upper limit of the regulation ability for anxiety is 5 points, but the actual regulation performance reflects that it can effectively regulate deeper anxiety emotions, then the upper limit of the regulation ability for anxiety can be adjusted to 6 points.
[0072] As Figure 3 shown, the embodiment of the present invention also provides a companion robot control system 10 based on emotion perception classification, which is applied to the companion robot. The system 10 includes a first emotion perception module 102, a second emotion perception module 104, and a scheduling module 106; the first emotion perception module 102 is used to obtain multi-dimensional emotion-related data of the target object and analyze the emotion-related data to obtain the probability of emotion abnormality determination; the second emotion perception module 104, when the probability of emotion abnormality determination is higher than the probability threshold, conducts a voice interaction with the target object, and uses a multi-modal emotion classification model to process the voice interaction data and emotion-related data to perform emotion perception classification on the target object, and obtain the emotion type and emotion intensity; the scheduling module 106, when the emotion type is an abnormal emotion and the emotion intensity exceeds the regulation ability of the companion robot, determines external personnel based on the emotion type and the emotion intensity, including emotion regulation personnel and / or security personnel, and schedules the external personnel to perform emotion intervention on the target object.
[0073] Further, the first emotion perception module 102 is configured to: obtain the behavior characteristics of other persons in the same scene, calculate the similarity between the behavior characteristics of each other person and the target object, and if the proportion of the similarity higher than the similarity threshold exceeds the proportion threshold, lower the weight coefficient of the behavior characteristics; perform emotion abnormality determination respectively based on the behavior characteristics, facial expression characteristics, and voice characteristics in the emotion-related data, and fuse the determination probability based on the behavior characteristics and other determination probabilities according to the lowered weight coefficient to obtain the emotion abnormality determination probability.
[0074] Further, the multi-modal emotion classification model includes a first classification model, a second classification model, and a joint classification model; the second emotion perception module 104 is configured to: use the first classification model to perform emotion perception classification on the voice interaction data to obtain a first classification result; use the second classification model to perform emotion perception classification on the emotion-related data to obtain a second classification result; calculate the semantic deviation degree between the first classification result and the second classification result, and use the joint classification model to perform emotion perception classification on the voice interaction data and the emotion-related data under the guidance of the semantic deviation degree to obtain the emotion type and the emotion degree.
[0075] Further, the second emotion perception module 104 is configured to: if the semantic deviation degree is lower than the deviation threshold, the joint classification model directly fuses the first classification result and the second classification result according to a preset weight to calculate and obtain the emotion type and the emotion degree; if the semantic deviation degree is higher than the deviation threshold, the joint classification model enhances the weight of the data with higher credibility in the voice interaction data and the emotion-related data through an attention mechanism, and fuses the first classification result and the second classification result based on the data weight to calculate and obtain the emotion type and the emotion degree.
[0076] Further, the second emotion perception module 104 is configured to: construct a word vector sequence for the voice interaction data, calculate a context correlation matrix through a self-attention mechanism, calculate a sentence coherence score, and adjust the credibility weight of the voice interaction data based on the sentence coherence score; map the voice interaction data and the emotion-related data into multiple subspaces respectively, calculate the modal consistency score of each subspace, and adjust the credibility weight of the emotion-related data based on each modal consistency score; obtain a first fusion weight of the voice interaction data based on the adjusted credibility weight, and further obtain a second fusion weight of the emotion-related data.
[0077] Further, the second emotion perception module 104 is configured to: calculate the semantic density uniformity of the word vector sequence, and adjust the local density-based dynamic scaling factor of the self-attention mechanism in the dot product stage based on the semantic density uniformity; the semantic density uniformity is determined based on the Euclidean distance variance between adjacent word vectors; calculate the context correlation matrix through the self-attention mechanism according to the dynamic scaling factor.
[0078] Further, the scheduling module 106 is configured to: retrieve a pre-established regulation ability threshold table for different abnormal emotion types, match and find out the upper limit of the regulation ability of the escort robot corresponding to the emotion type, compare the upper limit of the regulation ability with the emotion degree, and make a judgment on whether the emotion degree of the emotion type exceeds the regulation ability of the escort robot according to the comparison result.
[0079] An embodiment of the present invention further provides an electronic device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, where the computer program, when executed by the processor, implements the method described in any one of the foregoing.
[0080] An embodiment of the present invention further provides a computer storage medium, which stores a computer program executable by a processor to implement the method described in any one of the foregoing.
[0081] An embodiment of the present invention further provides a computer program product, which includes a computer program executable by a processor to implement the method described in any one of the foregoing.
[0082] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by the present invention, and these modifications or substitutions should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A control method for an escort robot based on emotion perception classification, which is applied to the escort robot, and is characterized in that: The method comprises the following steps: acquiring multi-dimensional emotion-related data of a target object, and obtaining an abnormal emotion determination probability based on the emotion-related data analysis; when the abnormal emotion determination probability is higher than a probability threshold, performing voice interaction with the target object, and using a multimodal emotion classification model to process the voice interaction data and emotion-related data to perform emotion perception classification on the target object, and obtaining an emotion type and an emotion degree; when the emotion type is an abnormal emotion and the emotion degree exceeds the regulation capability of the accompanying robot, determining external personnel, including emotion regulation personnel and / or security personnel, based on the emotion type and the emotion degree, and dispatching the external personnel to perform emotional intervention on the target object.
2. The control method of a companion robot based on emotion perception classification according to claim 1, characterized in that: The probability of abnormal emotion determination is obtained based on the analysis of emotion-related data, including: obtaining the behavioral characteristics of other persons in the same scene, calculating the similarity between the behavioral characteristics of each other person and the target object, and if the proportion of similarities above the similarity threshold exceeds the proportion threshold, the weight coefficient of the behavioral characteristics is lowered; based on the behavioral characteristics, facial expression characteristics, and voice characteristics in the emotion-related data, emotional abnormality determination is performed separately, and the determination probability based on the behavioral characteristics is merged with the other determination probabilities according to the lowered weight coefficient to obtain the probability of abnormal emotion determination.
3. The control method of a companion robot based on emotion perception classification according to claim 1, wherein: The multimodal emotion classification model includes a first classification model, a second classification model and a joint classification model; A multimodal emotion classification model is used to process voice interaction data and emotion-related data to perform emotion perception classification on a target object to obtain an emotion type and an emotion degree, including: using a first classification model to perform emotion perception classification on voice interaction data to obtain a first classification result; using a second classification model to perform emotion perception classification on emotion-related data to obtain a second classification result; calculating the semantic deviation degree of the first classification result and the second classification result, and using a joint classification model to perform emotion perception classification on voice interaction data and emotion-related data based on the semantic deviation degree to obtain the emotion type and the emotion degree.
4. The control method of a companion robot based on emotion perception classification according to claim 3, characterized in that: Guided by the semantic deviation degree, a joint classification model is used to perform emotion perception classification on the voice interaction data and emotion-related data to obtain the emotion type and the emotion degree, including: if the semantic deviation degree is lower than the deviation threshold, the joint classification model directly fuses the first classification result and the second classification result according to the preset weight, and calculates the emotion type and the emotion degree; if the semantic deviation degree is higher than the deviation threshold, the joint classification model enhances the weight of the data with higher credibility in the voice interaction data and emotion-related data through the attention mechanism, and fuses the first classification result and the second classification result based on the data weight, and calculates the emotion type and the emotion degree.
5. The control method of an escort robot based on emotion perception classification according to claim 4, characterized in that: The joint classification model enhances the weights of more reliable data in speech interaction data and emotion-related data through an attention mechanism, including: constructing a word vector sequence for the speech interaction data, calculating a context correlation matrix through a self-attention mechanism, calculating a sentence coherence score, and adjusting the credibility weight of the speech interaction data based on the sentence coherence score; mapping the speech interaction data and emotion-related data into multiple subspaces respectively, calculating the modal consistency score of each subspace, and adjusting the credibility weight of the emotion-related data based on each modal consistency score; obtaining the first fusion weight of the speech interaction data based on the adjusted credibility weight, and further obtaining the second fusion weight of the emotion-related data.
6. A control method for an escort robot based on emotion perception classification according to claim 5, characterized in that: Calculating the context correlation matrix through the self-attention mechanism includes: calculating the semantic density uniformity of the word vector sequence, and adjusting the dynamic scaling factor based on local density in the dot product stage of the self-attention mechanism based on the semantic density uniformity; the semantic density uniformity is determined based on the variance of the Euclidean distance between adjacent word vectors; calculating the context correlation matrix through the self-attention mechanism according to the dynamic scaling factor.
7. A control method for an escort robot based on emotion perception classification according to claim 1, characterized in that: Judging whether the degree of the emotion of the emotion type exceeds the regulation ability of the escort robot through the following method. Specifically: retrieving a pre-established regulation ability threshold table for different abnormal emotion types, matching and finding out the upper limit of the regulation ability of the escort robot corresponding to the emotion type, comparing the upper limit of the regulation ability with the degree of the emotion, and making a judgment on whether the degree of the emotion of the emotion type exceeds the regulation ability of the escort robot according to the comparison result.
8. An escort robot control system based on emotion perception classification, applied to an escort robot, characterized in that, The system includes a first emotion perception module, a second emotion perception module, and a scheduling module; the first emotion perception module is used to obtain multi-dimensional emotion-related data of a target object, and analyze and obtain an emotion abnormality determination probability based on the emotion-related data; the second emotion perception module, when the emotion abnormality determination probability is higher than a probability threshold, conducts a speech interaction with the target object, and uses a multi-modal emotion classification model to process the speech interaction data and emotion-related data to perform emotion perception classification on the target object, and obtain an emotion type and an emotion degree. The scheduling module, when the emotion type is an abnormal emotion and the degree of the emotion exceeds the regulation ability of the escort robot, determines external personnel based on the emotion type and the degree of the emotion, including emotion regulation personnel and / or security personnel, and schedules the external personnel to conduct emotion intervention on the target object.
9. An electronic device, characterized in that: The electronic device includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, and the computer program, when executed by the processor, implements the method according to any one of claims 1-7.
10. A computer storage medium, characterized in that: The computer storage medium stores a computer program that can be executed by a processor to implement the method according to any one of claims 1-7.
Citation Information
Patent Citations
Emotion monitoring and early warning method and system
CN106361356A
Emotion recognition system and method and electronic equipment
CN112016367A
Child accompanying robot based on emotion intelligent recognition and control method
CN113246156A
Intelligent emotion recognition method based on multi-modal data fusion
CN118709146A
Intelligent accompanying method and emotion monitoring device
CN119770823A
Cited By
Emotion response method, robot system and storage medium
CN122463198A
Emotion response methods, robotic systems, and storage media
CN122463198B