User emotion recognition method based on text information, medium and equipment

By using text-based user intent recognition and emotion contagion models, the problem of inaccurate emotion recognition in digital customer service systems when voice and video input are lacking is solved, achieving high-accuracy emotion recognition and improved service quality in privacy-preserving scenarios.

CN121958554APending Publication Date: 2026-05-01MOBILE TECH COMPANY CHINA TRAVELSKY HLDG
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MOBILE TECH COMPANY CHINA TRAVELSKY HLDG
Filing Date
2026-01-21
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing digital customer service systems, lacking voice and video input, rely solely on text content for emotion analysis, resulting in low accuracy and an inability to accurately identify users' true emotional states. This leads to decreased service quality and a poor user experience, especially in high-density crowd locations where they cannot effectively model the contagion effect of group emotions.

Method used

By using a text-based user intent recognition and emotion contagion model, combined with a spatial distance weighting function and intent weight, a two-stage emotion assessment mechanism is constructed to identify user intent types and compensate for the impact of group emotions, accurately model the impact of emotion contagion, and differentiate contagion weights between common anxiety and non-common anxiety.

Benefits of technology

In the absence of multimodal input, it significantly improves the accuracy of emotion recognition and the quality of service response, accurately captures users' true anxiety levels in privacy-preserving scenarios, provides timely and effective reassurance and solutions, and enhances the emotion perception capabilities of digital customer service.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121958554A_ABST
    Figure CN121958554A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of user emotion recognition in digital customer service, in particular to a user emotion recognition method based on text information, a medium and equipment. The method comprises the following steps: determining an influence time period and range of emotional infection based on current time and position; and calculating a comprehensive negative emotion value in combination with the basic emotion score and the emotion scores of other users meeting the conditions. Emotion scores, spatial distance weights and intention weights of other users are considered in a calculation formula, wherein the weight of a high anxiety common intention is higher. The emotion infection effect is accurately modeled through the spatial distance weight function, infection weights of common anxiety and non-common anxiety are distinguished, the accuracy of emotion evaluation under a single text mode is improved, higher weights are distributed to common anxiety events, and the emotion infection effect is enhanced. Therefore, the emotion underestimation of the target user caused by the single text input can be compensated.
Need to check novelty before this filing date? Find Prior Art

Description

A method, medium, and device for user emotion recognition based on text information. Technical Field

[0001] This invention relates to the field of user emotion recognition technology in digital customer service, and in particular to a method, medium and device for user emotion recognition based on text information. Background Technology

[0002] In the field of digital customer service systems, accurately identifying users' emotional states is a crucial step in providing personalized services. With increasing awareness of user privacy, more and more users are choosing to disable their camera and microphone permissions when using digital customer service systems, interacting solely through text. This trend is particularly evident in public places such as airports and train stations, where users, out of concern for their privacy, are often unwilling to provide sensitive biometric data such as facial expressions and voice tone to the digital customer service system. However, the emotion recognition modules of existing digital customer service systems mainly rely on multimodal data fusion. When voice and video input are lacking, methods that rely solely on text content for emotion analysis generally suffer from low accuracy and superficial emotional understanding. This limitation of unimodal emotion recognition prevents the system from accurately grasping the user's true emotional state, resulting in customer service responses that are severely mismatched with the user's actual needs, thus reducing service quality and user experience.

[0003] Existing technologies have significant limitations in handling single-modal emotion recognition in text. Traditional methods typically rely solely on keyword matching or simple sentiment analysis models to determine the emotion of user input, failing to adequately consider the complex relationship between user intent type and emotion intensity. Especially in high-density crowds such as airport waiting halls and train station waiting areas, user emotions are significantly influenced by the surrounding environment and the emotions of other users, forming a complex emotion contagion network. Existing systems cannot effectively model this group emotion contagion effect, nor can they distinguish the different degrees of impact of shared anxiety (such as widespread flight delays) and non-shared anxiety (such as lost personal luggage) on user emotions. When the system can only acquire text input, this deficiency is further amplified, leading to a significant deviation from the actual situation in the judgment of the user's emotional state. This results in an inability to provide targeted reassurance and service responses, and may even exacerbate negative emotions due to erroneous emotion judgments, leading to decreased service quality and increased user complaints. Summary of the Invention

[0004] To address one of the aforementioned technical problems, the present invention provides the following technical solution: According to one aspect of the present invention, a user emotion recognition method based on text information is provided, characterized in that the method includes the following steps: identifying the user intent type based on the input text of the target user obtained at the current recognition time, and determining a basic emotion score F(e); determining the influence period and influence distance range that have an emotional contagion effect on the target user based on the current recognition time and the target user's current location information; and generating a comprehensive negative emotion value E corresponding to the target user at the current recognition time based on F(e) and the basic emotion scores of all other users that meet the influence period and influence distance range conditions. total The following conditions must be met: ; ; Among them, E individual E represents the individual emotion score of the target user at the current identification moment. k The base sentiment score of the Kth other user who meets the criteria of the time period and distance of influence; n is the total number of other users who meet the criteria of the time period and distance of influence; k = 1, 2, ..., n; d k w(d) represents the spatial distance between the current target user and the Kth other user; k P represents the distance weighting function for the Kth other user; k Let P1 be the intent weight based on the intent type corresponding to the Kth other user's input text; the intent types include high-anxiety common intents and high-anxiety non-common intents. The intent weight corresponding to the high-anxiety common intent is P1, and the intent weight corresponding to the high-anxiety non-common intent is P2, where P1>P2>1; E total It is positively correlated with the level of anxiety.

[0005] According to a second aspect of the present invention, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores a computer program, which, when executed by a processor, implements the above-described method for user emotion recognition based on text information.

[0006] According to a third aspect of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the above-described method for user emotion recognition based on text information.

[0007] This invention has at least one of the following beneficial effects: By classifying user intent and implementing an emotion compensation mechanism under a single text modality, this invention effectively solves the problem of inaccurate emotion recognition caused by the lack of multimodal data. In practical applications, when users only provide text input for privacy reasons, traditional systems can only judge user emotions based on limited text information, often underestimating the user's true anxiety level. This invention designs a two-stage emotion assessment mechanism: first, it identifies the user intent type based on the input text; only when it is determined to be a high-anxiety intent is a basic emotion score determined; then, by introducing a group emotion contagion model, it uses the anxiety of other users as a compensation factor to compensate for the shortcomings of emotion assessment under a single text modality. In particular, through the formula... This invention organically integrates individual emotion scores with the contagious influence of the group, where the intention weight P k The parameters (P1>P2>1) can amplify the contagion effect differentially based on the anxiety types (common or non-common) of other users. This design enables the system to accurately compensate for the target user's emotional assessment by the emotional state of surrounding users even in the absence of multimodal inputs such as voice and video, significantly improving the accuracy of emotion recognition and the quality of service response in privacy-preserving scenarios.

[0008] This invention further improves the accuracy of emotion assessment in a single text modality by accurately modeling the impact of spatial distance on emotion contagion. Under the constraint of only being able to obtain text input, traditional methods cannot verify the intensity of a user's emotions through voice tone or facial expressions, leading to overly conservative emotion judgments. This invention innovatively introduces a distance weighting function. This function dynamically adjusts the intensity of emotion contagion based on the actual spatial distance between users, accurately reflecting the natural law that the impact of emotions diminishes with increasing distance in the real world. When a user exhibits strong negative emotions during text interaction, the system can calculate the degree of emotional impact on other users within a certain spatiotemporal range and incorporate this contagion effect as a compensation factor into the target user's emotion score. This spatially accurate emotion contagion modeling enables digital customer service to enhance the reliability of emotion recognition through interaction data of adjacent users in physical space, even in the absence of multimodal input. This effectively overcomes the limitations of a single text modality and makes emotion assessment more closely reflect the user's actual psychological state.

[0009] This invention significantly improves the accuracy of emotion compensation in privacy-preserving scenarios by differentiating contagion weights between shared and non-shared anxiety. In real-world scenarios where users refuse to grant voice and video permissions, text-based emotion recognition often fails to capture subtle emotional changes, especially for group emotional fluctuations related to high-anxiety shared events (such as widespread flight delays). This invention innovatively assigns different weights (P1>P2>1) to the anxiety types of other users, amplifying the emotional contagion effect of shared anxiety events. When calculating the comprehensive negative emotion value E… total If multiple users in the surrounding area express common anxieties such as flight delays or widespread cancellations, the emotional impact will be significantly amplified, effectively compensating for the underestimation of the target user's emotional score due to a single text input. This differentiated weighting mechanism allows the system to accurately infer individual emotional states through the common characteristics of group emotions, while protecting user privacy. Especially in sudden events or crisis situations, it can quickly capture trends in group emotional fluctuations, adjust service strategies in advance, and provide timely and effective reassurance and solutions, significantly improving the emotional perception capabilities and service quality of digital customer service in privacy-sensitive environments. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 is a flowchart of a digital customer service information generation method based on multimodal information provided by an embodiment of the present invention; Figure 2 is a flowchart of a user emotion recognition method based on text information provided by another embodiment of the present invention. Detailed Implementation

[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0013] As a possible embodiment of the present invention, as shown in Figure 1, a method for generating digital customer service information based on multimodal information is provided. This method is used when all modal information (text, voice and facial video) of the user can be fully acquired. The method includes the following steps: S100: Based on the input text of the target user acquired at the current recognition time, identify the user intent type and determine the basic emotion score F(e).

[0014] Specifically, S100 includes: S101: Using a pre-trained natural language processing model to perform binary classification on the input text to determine whether it is a high-anxiety intent.

[0015] A pre-trained BERT-based model can be used here, which has been fine-tuned for domain adaptation on a corpus containing 100,000 airport customer service dialogues. The binary classification design considers the balance between computational efficiency and accuracy, because in practical applications, the primary task is to distinguish between ordinary inquiries and high-anxiety states requiring urgent attention, rather than refining the emotion classification. A model output probability exceeding a 0.4 threshold is considered a high-anxiety intent; a neutral intent does not require the processing described in this invention and can be directly responded to using the standard preset response method.

[0016] S102: When a high-anxiety intent is identified, the input text is classified into intent range categories to determine whether the input text represents a common intent affecting most users in the current area, or a non-common intent affecting only a single user or a small group of users. This intent range classification is crucial because in public scenarios, the emotional contagion effect caused by common events (such as widespread flight delays) is far greater than that of personal events (such as lost luggage), and different service strategies need to be adopted.

[0017] S102 includes: S112: Matching the input text with a preset keyword list, which includes a set of common intent keywords and a set of non-common intent keywords.

[0018] S122: When the input text contains any keyword from the set of common intent keywords, determine that the common intent type corresponding to the input text is the high anxiety common intent.

[0019] S132: When the input text does not contain common intent keywords but contains any keyword from the set of non-common intent keywords, the common intent type corresponding to the input text is determined to be high anxiety non-common intent.

[0020] The set of common intent keywords includes keywords related to all flight delays, all delays, large-scale cancellations, and extreme weather. The set of non-common intent keywords includes keywords related to flight changes, lost personal baggage, and ticket refunds / rescheduling. The keyword set is designed following a "precision-first" principle to avoid overgeneralization. For example, "delay" alone is not considered a common keyword; it must be combined with scope-limiting words such as "all," "complete," or "large-scale" to constitute a common intent trigger condition. A synonym expansion mechanism is also implemented; for example, "cancel" will match variations such as "cancel" and "revoke," improving recall.

[0021] In addition to determining the basic emotion score F(e), the system also determines the common anxiety identifier corresponding to the input text based on the matching results between the input text and the preset keyword list. This identifier is then added to the user's information vector so that the corresponding intent weight can be generated based on the label in subsequent use.

[0022] S142: If the common intent type corresponding to the input text is a high anxiety common intent, then the anxiety common identifier configured for the corresponding basic emotion score is a high anxiety common identifier. For example, the high anxiety common identifier is 1.

[0023] S152: If the common intent type corresponding to the input text is high anxiety non-common intent, then the anxiety commonality identifier configured for the corresponding basic emotion score is a high anxiety non-common identifier. For example, the high anxiety non-common identifier is 0.

[0024] This identifier design aims to preserve key emotional information while protecting user privacy. The information vector does not directly store the original text or detailed emotional intensity; only a binary identifier is retained, effectively preventing the leakage of personal information. Simultaneously, this identifier provides necessary parameters for subsequent emotion contagion calculations, achieving a balance between privacy protection and functional requirements.

[0025] S103: Determine the basic emotional score F(e) based on the classification results: When the common intention is judged to be high anxiety, F(e) = 0.7. Common intentions are given a higher score (0.7) because common anxiety events are often accompanied by stronger negative emotions and more urgent service needs.

[0026] When the intention is determined to be high anxiety non-common, F(e) = 0.6.

[0027] The threshold for determining high-anxiety intentions is 0.4, and the threshold for determining common intentions is 0.65.

[0028] Of course, there will be situations where keyword matching fails, namely when the input text contains neither common intent keywords nor non-common intent keywords. In this case, the intent consensus type corresponding to the input text is neither a high-anxiety consensus intent nor a high-anxiety non-consensus intent, but only a high-anxiety intent form. Therefore, in this case, a base score of F(e) = 0.5 can be assigned. At the same time, no anxiety consensus identifier is configured for it, and the intent weight corresponding to this intent is an empty set. This design handles edge cases, avoiding the system's over-interpretation of uncertain intents. The intermediate value of 0.5 reflects the state of "anxiety but uncertain scope of influence," leaving room for subsequent multimodal sentiment weight adjustments.

[0029] S200: Based on the target user's input voice and facial input video obtained at the current recognition time, obtain the emotion weight coefficient G(e). G(e) is used to adjust F(e).

[0030] When both the target user's input speech and facial input video are acquired simultaneously, the system calculates the speech emotion weight coefficient G_a(e) and the video emotion weight coefficient G_v(e) separately, and then takes their arithmetic mean as the final emotion weight coefficient G(e) = (G_a(e) + G_v(e)) / 2. The speech emotion weight coefficient G_a(e) is obtained using existing speech emotion recognition technologies: for example, the fundamental frequency, energy, MFCC (Mel-frequency cepstral coefficients), and prosodic features of the speech signal can be extracted using the open-source toolkit openSMILE. These features are then input into an SVM classifier pre-trained on the IEMOCAP dataset, outputting weight values ​​in the range of 0.8-1.6. The video emotion weight coefficient G_v(e) is obtained using existing facial expression recognition technologies: for example, the OpenFace toolkit is used to detect 68 facial keypoints, extract the intensity of facial action units (AMUs), and combine this with a convolutional neural network model pre-trained on the FER+ dataset to analyze the emotional intensity of the user's facial expressions, generating weight values ​​in the range of 0.8-1.6. Of course, in this embodiment, in addition to using existing open-source tools and some training data that can be collected in this field, existing model training methods are used to directly obtain the model that meets the needs of the application scenario of this invention, which is based on input speech and facial input video, and directly obtains G_a(e) and G_v(e).

[0031] The emotion weight coefficients generated based on input information from two different modalities are set in the range of 0.8 to 1.6, primarily based on the following considerations: Generally, the anxiety expressed in the user's text input is consistent with the emotional state reflected in their facial video and voice information. However, in certain special cases, the emotional state identified by voice information and facial video may differ significantly from the text information recognition results. Therefore, setting the minimum value of the weight coefficient range close to 1 allows for a moderate reduction in the adjustment of F(e) when the voice and facial information results are inconsistent; conversely, when the emotional state identified by voice and facial information is consistent, F(e) can be significantly increased, thereby improving overall recognition accuracy.

[0032] In addition, existing methods for predicting a user's current emotional state based on user-input voice information (such as the voice emotion recognition method disclosed in CN115881103A) are used to obtain the target user's current emotional state. The emotional state is then classified into the corresponding preset anxiety level and assigned a weight value in the range of 0.8-1.6 to obtain G_a(e).

[0033] Alternatively, a method can be used to predict the current emotional state of a user based on the obtained facial video (such as an emotional state assessment method disclosed in CN112472088A) to obtain the current emotional state of the target user, then classify the emotional state into the corresponding preset anxiety level, and assign it a corresponding weight value in the range of 0.8-1.6, thereby obtaining G_v(e).

[0034] When only voice input is obtained, the emotion weighting coefficient G(e) is directly equal to the voice emotion weighting coefficient G_a(e).

[0035] When only facial input video is obtained, the emotion weight coefficient G(e) is directly equal to the video emotion weight coefficient G_v(e).

[0036] When no voice or facial video input is available, the emotion weighting coefficient G(e) is set to 1.0 by default, and E... individual =F(e)×1.0=F(e), automatically downgraded to the W300 processing flow, but retaining the geographic grid queue mechanism to obtain group sentiment compensation.

[0037] In all multimodal input scenarios, the calculation of G(e) references standardized methods in the field of affective computing, such as the IEEE Affective Computing standard P2897 and the multimodal fusion strategies recommended by ACII (International Conference on Affective Computing). These existing technologies have been widely applied and verified in real-world scenarios such as customer service centers and smart terminals, and those skilled in the art can also use these existing technologies to obtain G(e) in this embodiment.

[0038] S300: Based on the current identification time and the target user's current location information, determine the period of time and distance range of influence that have an emotional contagion effect on the target user.

[0039] Specifically, the period of influence is within one minute of the current identification time. The spatial range of influence is within 10 meters of the target user. Emotional contagion exhibits significant spatiotemporal characteristics in public spaces. Typically, historical emotional contagion occurs primarily within a short period (<2 minutes) and at close range (<10 meters). This step, by limiting the spatiotemporal range of influence, avoids interference from irrelevant user emotions, improving computational efficiency and accuracy. Location information is obtained through indoor positioning systems (such as Wi-Fi fingerprinting or Bluetooth beacons), achieving centimeter-level accuracy to meet the needs of emotional contagion modeling.

[0040] S400: Based on F(e), G(e), and the basic sentiment scores of all other users who meet the conditions of the time period and distance of influence, generate the comprehensive negative sentiment value E corresponding to the target user at the current identification time. total The following conditions must be met: .

[0041] . In order to make E individual To better represent the level of anxiety in an individual's emotions, existing normalization techniques can be used, and then E... individual The value of F(e)×G(e) is limited to [0, 2.0] to facilitate the setting of relevant thresholds in subsequent calculations (such as the setting of Y1 and Y2) and the implementation of subsequent steps. Correspondingly... The value range can be set to [0, 1]. This ensures that the influence of other individuals' emotions on the target user's emotions is not only guaranteed, but also limits the emotions of other individuals by setting a corresponding value range. Therefore, the impact of other individuals' emotions on the final overall negative emotion value of the target user is relatively low compared to the proportion of individual emotions. From this, the final calculated E can also be obtained. totalThe value range is [0, 3]. By limiting this range, we can better distinguish the level of users' overall negative emotions corresponding to different range values, thereby expressing the user's anxiety level, so as to generate corresponding digital human customer service response prompts.

[0042] E total The design integrates individual emotions and group influences, reflecting the "individual-environment" interaction theory. The multiplicative relationship F(e)×G(e) reflects the non-linear superposition effect of multimodal emotion intensity: when the textual intent aligns with the voice / facial emotion (e.g., high-anxiety text + high-anxiety voice and facial expression), the emotional intensity is further amplified; when inconsistent, they mutually inhibit each other. This setting of mutual inhibition or reinforcement between multimodal results allows users' emotions to be expressed more accurately.

[0043] Among them, E individual E represents the individual emotion score of the target user at the current identification moment. k This represents the base sentiment score of the Kth other user who meets the criteria of the time period and distance of influence. n is the total number of other users who meet the criteria of the time period and distance of influence. k = 1, 2, ..., n. k Let w(d) be the spatial distance between the current target user and the Kth other user. k Let E be the distance weight function for the Kth other user. total It is positively correlated with the level of anxiety.

[0044] Individual emotions are often significantly influenced by the surrounding environment and the emotions of other users, but traditional emotion recognition systems typically ignore this group interaction effect. This invention constructs a spatial distance-based emotion contagion weighting mechanism, using a distance weighting function... This function precisely quantifies the intensity of emotional influence between users in different spatial locations. The 3.5-meter distance is based on the midpoint of social distance in interpersonal distance theory, and 0.8 controls the decay rate. The function employs a variation of the Sigmoid function to accurately characterize the non-linear characteristic of emotional contagion decreasing with increasing distance in reality. For example, this function in d... k The maximum value of approximately 0.94 is reached at 0 meters, indicating that the emotional contagion effect is strongest when users are closely adjacent; as the distance increases, the weight decreases smoothly, reaching a maximum at d. kAt a distance of 3.5 meters, the weight is 0.5, representing a half-range influence; when the distance exceeds 7 meters, the weight drops below 0.06, and the emotional contagion effect essentially disappears. This design conforms to the interpersonal space theory of "intimate distance (<0.5m), personal distance (0.5-1.2m), social distance (1.2-3.5m), and public distance (>3.5m)," and also smoothly transitions the intensity of emotional influence across different distance ranges through a mathematical function. The S-shaped decay characteristic of the function effectively avoids the abrupt change problem of piecewise functions, making the emotional contagion model more stable and consistent with the real laws of human emotional transmission in practical applications.

[0045] S500: According to E total The size of the prompts determines the content of the digital customer service avatar's response.

[0046] S500 includes: S501: If E total ≤0.7, generate standard response prompts. This level corresponds to typical consultation scenarios where the user is emotionally stable and a standardized, streamlined service can be provided. The prompt template includes accurate information and basic polite language, such as "According to the query, your flight is on time. Please wait at gate 15." S502: If 0.7 <E total ≤1.2, add emotional reassurance statements to the prompt. This level corresponds to moderate anxiety and requires emotional support. The system adds empathetic statements to the standard response, such as "We understand your concern about the flight delay. We are working hard to coordinate and expect to update the information in 30 minutes." S503: If 1.2 <E total For scores ≤1.6, add efficient solution strategies and priority processing indicators to the prompts. This level corresponds to high anxiety and requires quick problem-solving. The prompts include clear time commitments and priority information, such as "We have detected that you are experiencing significant anxiety. We have opened a priority channel for you and will arrange a dedicated customer service representative for you within 3 minutes." Simultaneously, mark the prompt as priority in the background to ensure service commitments are fulfilled.

[0047] S504: If 3 ≥ E total >1.6, Force the insertion of a live agent transfer option and emergency reassurance messages into the prompt. This level corresponds to users with extreme anxiety or potential risk, requiring human intervention. The system provides an immediate transfer option while using reassuring language trained in crisis intervention, such as, "I understand how you feel; this situation is indeed anxiety-inducing. To better assist you, we recommend transferring you to a live agent, who has more authority and resources to resolve your issues." Simultaneously, an alert is sent to the backend to ensure a rapid response. Furthermore, in this embodiment, the corresponding customer service avatar is a digital human. This allows us to adjust the digital human's speaking speed, tone, and facial expressions according to different user anxiety levels to help users resolve problems while alleviating their anxiety.

[0048] As another possible embodiment of the present invention, as shown in Figure 2, a user emotion recognition method based on text information is provided. This method is mainly applicable to use scenarios where only the user's text input can be obtained. Specifically, the method includes the following steps: W100: Identify the user's intent type based on the target user's input text obtained at the current recognition time, and determine the basic emotion score F(e).

[0049] W200: Based on the current identification time and the target user's current location information, determine the period of time and distance range of influence that have an emotional contagion effect on the target user.

[0050] The content corresponding to W100 and W200 is completely consistent with the content of S100 and S200 in the above embodiments, and will not be repeated here. You can directly refer to the content of S100 and S200 in the above embodiments for setting. The main difference lies in the specific settings of W300, which are as follows: W300: Based on F(e) and the basic emotion scores of all other users who meet the conditions of the time period and distance of influence, generate the comprehensive negative emotion value E corresponding to the target user at the current identification time. total The following conditions must be met: .

[0051] . .

[0052] Unlike the multimodal input information processing methods in the previous embodiments, due to user permission restrictions, this example can only obtain text information input by the user and cannot obtain voice information or facial video information. Therefore, this embodiment cannot generate adjustment weights for F(e) based on other modal information. In this case, E individual The value is F(e), not F(e) × G(e). To overcome this limitation, this embodiment introduces an intent weight P. k We apply weighted processing to the relevant sentiment values ​​in the herd contagion effect. By enhancing the weights of the herd sentiment model, the herd sentiment contagion effect can compensate for the uncertainty of individual sentiment assessments, thereby improving the accuracy of overall sentiment assessment.

[0053] Among them, E individual E represents the individual emotion score of the target user at the current identification moment. k This represents the base sentiment score of the Kth other user who meets the criteria of the time period and distance of influence. n is the total number of other users who meet the criteria of the time period and distance of influence. k = 1, 2, ..., n. k Let w(d) be the spatial distance between the current target user and the Kth other user. k Let P be the distance weight function for the Kth other user. kThis represents the intent weight based on the intent type corresponding to the Kth other user's input text. Intent types include high-anxiety common intents and high-anxiety non-common intents. The intent weight corresponding to a high-anxiety common intent is P1, and the intent weight corresponding to a high-anxiety non-common intent is P2, where P1 > P2 > 1. For example, P1 = 1.5, P2 = 1.1. E total It is positively correlated with the level of anxiety.

[0054] P k The design is based on the principle of emotional contagion: the emotional contagion intensity triggered by shared anxiety events (such as all flight delays) is greater than that triggered by non-shared events (such as lost personal luggage). Therefore, when other users express shared anxiety, the emotional contagion effect is significantly amplified, effectively compensating for the inability to confirm the target user's emotions through voice / facial expression.

[0055] By implementing user intent classification and emotion compensation mechanisms within a single text modality, this invention effectively addresses the problem of inaccurate emotion recognition caused by the lack of multimodal data. In practical applications, when users only provide text input for privacy reasons, traditional systems can only judge user emotions based on limited text information, often underestimating the user's true anxiety level. This invention designs a two-stage emotion assessment mechanism: first, it identifies the user's intent type based on the input text; only when it is determined to be a high-anxiety intent can a basic emotion score be further determined; then, by introducing a group emotion contagion model, the anxiety of other users is used as a compensation factor to compensate for the shortcomings of emotion assessment under a single text modality. Specifically, through the formula... This invention organically integrates individual emotion scores with the contagious influence of the group, where the intention weight P k The parameters (P1>P2>1) can amplify the contagion effect differentially based on the anxiety types (common or non-common) of other users. This design enables the system to accurately compensate for the target user's emotional assessment by the emotional state of surrounding users even in the absence of multimodal inputs such as voice and video, significantly improving the accuracy of emotion recognition and the quality of service response in privacy-preserving scenarios.

[0056] By accurately modeling the impact of spatial distance on emotion contagion, the accuracy of emotion assessment in a single text modality is further improved. Under the constraint of only obtaining text input, traditional methods cannot verify the intensity of a user's emotions through voice tone or facial expressions, leading to overly conservative emotion judgments. This invention introduces a distance weighting function. This function dynamically adjusts the intensity of emotion contagion based on the actual spatial distance between users, accurately reflecting the natural law that the impact of emotions diminishes with increasing distance in the real world. When a user exhibits strong negative emotions during text interaction, the system can calculate the degree of emotional impact on other users within a certain spatiotemporal range and incorporate this contagion effect as a compensation factor into the target user's emotion score. This spatially accurate emotion contagion modeling enables digital customer service to enhance the reliability of emotion recognition through interaction data of adjacent users in physical space, even in the absence of multimodal input. This effectively overcomes the limitations of a single text modality and makes emotion assessment more closely reflect the user's actual psychological state.

[0057] By differentiating contagion weights between shared and non-shared anxiety, the accuracy of emotion compensation in privacy-preserving scenarios is significantly improved. In real-world scenarios where users refuse to grant voice and video permissions, text-based emotion recognition often fails to capture subtle emotional changes, especially for group emotional fluctuations related to high-anxiety shared events (such as widespread flight delays). This invention innovatively assigns different weights (P1>P2>1) to the anxiety types of other users, amplifying the emotional contagion effect of shared anxiety events. When calculating the comprehensive negative emotion value E... total If multiple users in the surrounding area express common anxieties such as flight delays or widespread cancellations, the emotional impact will be significantly amplified, effectively compensating for the underestimation of the target user's emotional score due to a single text input. This differentiated weighting mechanism allows the system to accurately infer individual emotional states through the common characteristics of group emotions, while protecting user privacy. Especially in sudden events or crisis situations, it can quickly capture trends in group emotional fluctuations, adjust service strategies in advance, and provide timely and effective reassurance and solutions, significantly improving the emotional perception capabilities and service quality of digital customer service in privacy-sensitive environments.

[0058] As another possible embodiment of the present invention, after S300: determining the period of influence and the range of influence distance that have an emotional contagion effect on the target user based on the current identification time and the current location information of the target user, the method further includes: S310: obtaining the basic emotional scores of other users who meet the conditions from the queue of the two-dimensional geographic grid corresponding to the target area according to the filtering conditions defined by the period of influence and the range of influence distance.

[0059] Geographic gridding design solves the problem of efficient emotion data retrieval in high-density environments. The target area (e.g., a 50m x 50m waiting area) is divided into 1m x 1m grids, with each grid maintaining an independent queue. When calculating the contagion of a user's emotions, only 10 x 10 = 100 surrounding grids need to be queried (instead of all 10,000+ grids). The grid size can be adjusted according to scene density; high-density areas (e.g., security checkpoints) use 0.5m x 0.5m grids, while low-density areas (e.g., rest areas) use 2m x 2m grids.

[0060] S320: Each geographic raster maintains a first queue and a second queue. The first queue stores information vectors that have a short-term infectious effect on users. The second queue stores information vectors that have a long-term infectious effect on users. The validity period T1 of the information vectors that have a short-term infectious effect on users is shorter than the validity period T2 of the information vectors that have a long-term infectious effect on users. Specifically, T1 = 60 seconds, and T2 = 180 seconds.

[0061] The dual-queue design reflects the time-sensitive nature of emotion contagion. Short-term contagion (such as a sudden outburst of anger) has a strong but brief impact, while long-term contagion (such as persistent anxiety) has a lasting effect. Therefore, T1 = 60 seconds and T2 = 180 seconds are set, with short-term emotions weakening after 60 seconds and long-term emotions weakening after 180 seconds. Separate storage in the dual queues avoids data contamination and improves the accuracy of the emotion contagion model.

[0062] Specifically, the information vector includes the user's individual emotion score, the timestamp when the individual emotion score was generated, the user ID identifier, the user's current location information when the individual emotion score was generated, and common anxiety identifiers.

[0063] The information vector design follows the "minimum necessary data" principle, storing only the data necessary for emotion contagion calculations and not recording original dialogue content or personal identity details to protect user privacy. User IDs are used only for deduplication (avoiding multiple entries of the same user), and these IDs can be hashed information, not associated with personal identity. Anxiety commonality identifiers use binary encoding to further reduce the risk of information leakage. Meanwhile, the geographic grid precision automatically decreases to 5m×5m during low-traffic periods at night, reducing the risk of location tracking.

[0064] The information vectors in the first and second queues are generated according to the following steps: S321: If the target user E at a certain identification time... individual When ∈(Y1, Y2], the information vector of the target user corresponding to the identification time is stored in the first target queue maintained by the geographic raster to which the target user's location belongs at the identification time. individual It is positively correlated with the level of anxiety, with Y1 and Y2 being the first and second anxiety thresholds, respectively.

[0065] The settings for Y1 and Y2 are based on an analysis of emotion intensity distribution. In actual deployment, Y1=0.6, Y2=1.2. E individual Users with emotional states ∈(0.6,1.2] are capable of influencing those around them but have not reached a crisis level, making them suitable for short-term queue storage. The first queue adopts a FIFO (First-In, First-Out) strategy, where new data automatically replaces old data from 60 seconds ago, ensuring that the queue size is controllable.

[0066] E individual >Y2 (e.g., >1.2) represents an extreme emotional state, and these users can generate widespread and persistent emotional contagion. Because this state persists for a long time, it is necessary to predict the user's movement trajectory during this period to cover more potentially affected geographic grids based on the predicted movement trajectory. Movement trajectory prediction can employ an improved Kalman filtering algorithm combined with typical airport route maps (e.g., security checkpoint → boarding gate → waiting area), or use the method corresponding to S332 in this embodiment for prediction. The information vector is stored in a second queue of multiple grids to ensure that users along the route are affected by emotional contagion, simulating the real-world path of emotional spread.

[0067] S322: If the target user E at a certain identification time... individual When Y2, the predicted movement trajectory of the target user in T2 corresponding to the identification time is used as the target storage raster, and the geographical raster it passes through is used as the target storage raster.

[0068] S323: Store the information vector of the target user corresponding to the identification time into the second queue of each target storage grid.

[0069] If the target area is the waiting hall, then S322: the geographic grid that the target user's predicted movement trajectory generated in T2 corresponding to the identification time is used as the target storage grid, including: S332: determine the affected area according to the user's current location type: if the user is currently at the boarding gate, the affected area is the current grid and its adjacent grids.

[0070] If the user is currently in the security check area, the affected area is the current grid and the path grid leading to the nearest boarding gate.

[0071] If the user is currently in the dining area, the affected area is the current grid and the four grids adjacent to it in front, behind, left, and right.

[0072] The boarding gate area has high pedestrian traffic but frequent movement, so the impact is limited to adjacent grid cells. The security checkpoint is a high-risk area for emotional outbursts, and users tend to move along fixed paths, so the impact extends to the path grid cells. In the food and beverage area, users spend more time there, resulting in a longer-lasting emotional impact, but they move less, so a fixed adjacent grid cell strategy is used. Of course, these rules for determining the scope of the affected grid cells need to be adapted to the specific usage scenario to suit the characteristics of each scenario.

[0073] In addition, the individual sentiment scores of users in the information vector of the second queue satisfy the following condition: .

[0074] Among them, E t individual This represents the individual emotion score of the user at the current identification time. tk is the time corresponding to the timestamp in the user's information vector. E tk individual The individual sentiment score in the user information vector. t now T1 represents the current identification time. T2 represents the validity period of the information vector that has a long-term impact on users.

[0075] In embodiments of the present invention, a second queue is used to store information vectors of users with long-term emotional contagion effects. These users typically exhibit high-intensity negative emotions (such as extreme anxiety, anger, etc.), and their emotional state may have a lasting impact on those around them. However, over time, this emotional influence gradually weakens, and its decay process is not linear, but rather exhibits a non-linear characteristic of "slow decline in the early stage and rapid decay in the later stage."

[0076] To accurately simulate this psychological and behavioral pattern, this embodiment introduces the time decay formula for a quadratic function: This quadratic function can better fit the characteristics of this decay curve.

[0077] For example, in a specific scenario: Passenger A exhibits extremely high anxiety due to flight cancellation at time tk=0, and their initial individual emotional score is E. tk individual =1.6, and is stored in the second queue, valid for 180 seconds. The decay process of its emotion can be seen in Table 1 below: Table 1 As can be seen from the table above, in the first 30 seconds, the emotion score only dropped from 1.6 to 1.56, a very small change, reflecting the "inertia" of emotion transmission; while in the last 30 seconds (150→180 seconds), the emotion score dropped sharply from 0.49 to 0, reflecting the rapid fading of the influence of emotion.

[0078] By employing nonlinear decay, the system more realistically reflects the dynamic propagation of emotions within a population, avoiding the overestimation of long-term emotional impact by traditional linear models. Simultaneously, it ensures that only recent, intensely emotional events influence the current user's overall emotional assessment, preventing historically accumulated emotions from interfering with current service decisions. Furthermore, T2 can dynamically adjust based on actual scenarios (such as peak hours or nighttime periods), giving the system greater environmental adaptability.

[0079] In conclusion, this time decay formula not only scientifically simulates the natural decay law of human emotional influence, but also significantly improves the accuracy and practicality of the emotion contagion model, making it one of the key technical supports for achieving efficient and intelligent digital customer service.

[0080] Furthermore, although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.

[0081] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, mobile terminal, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0082] In an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described method is also provided.

[0083] Those skilled in the art will understand that various aspects of the present invention can be implemented as systems, methods, or program products. Therefore, various aspects of the present invention can be specifically implemented in the following forms: entirely in hardware, entirely in software (including firmware, microcode, etc.), or in a combination of hardware and software, collectively referred to herein as “circuit,” “module,” or “system.”

[0084] An electronic device according to this embodiment of the invention. The electronic device is merely an example and should not be construed as limiting the functionality or scope of the embodiments of the invention.

[0085] Electronic devices are manifested in the form of general-purpose computing devices. Components of an electronic device may include, but are not limited to: at least one processor, at least one memory, and buses connecting different system components (including memory and processor).

[0086] The memory stores program code that can be executed by a processor, causing the processor to perform the steps described in the "Exemplary Methods" section above, according to various exemplary embodiments of the present invention.

[0087] The storage may include readable media in the form of volatile storage, such as random access memory (RAM) and / or cache memory, and may further include read-only memory (ROM).

[0088] The storage may also include programs / utilities having a set (at least one) of program modules, including but not limited to: an operating system, one or more applications, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.

[0089] A bus can represent one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus that uses any of the various bus architectures.

[0090] The electronic device can also communicate with one or more external devices (e.g., keyboards, pointing devices, Bluetooth devices, etc.), one or more devices that enable a user to interact with the electronic device, and / or any device that enables the electronic device to communicate with one or more other computing devices (e.g., routers, modems, etc.). This communication can be performed via input / output (I / O) interfaces. Furthermore, the electronic device can communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter. The network adapter communicates with other modules of the electronic device via a bus. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the electronic device, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0091] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible embodiments, various aspects of the present invention may also be implemented as a program product comprising program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps of the various exemplary embodiments of the present invention described in the "Exemplary Methods" section above.

[0092] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0093] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0094] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0095] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0096] Furthermore, the accompanying drawings are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes shown in the above drawings do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0097] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0098] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for user emotion recognition based on text information, characterized in that, The method includes the following steps: identifying the user's intent type based on the input text of the target user obtained at the current identification time, and determining the basic emotion score F(e); determining the period of influence and the range of influence distance that have an emotional contagion effect on the target user based on the current identification time and the target user's current location information; and generating the comprehensive negative emotion value E corresponding to the target user at the current identification time based on F(e) and the basic emotion scores of all other users that meet the conditions of the period of influence and the range of influence distance. total Meets the following conditions: ; ; Among them, E individual E represents the individual emotion score of the target user at the current identification moment. k The base sentiment score of the Kth other user who meets the criteria of the time period and distance of influence; n is the total number of other users who meet the criteria of the time period and distance of influence; k = 1, 2, ..., n; d k w(d) represents the spatial distance between the current target user and the Kth other user; k P represents the distance weighting function for the Kth other user; k Let P1 be the intent weight based on the intent type corresponding to the Kth other user's input text; the intent types include high-anxiety common intents and high-anxiety non-common intents. The intent weight corresponding to the high-anxiety common intent is P1, and the intent weight corresponding to the high-anxiety non-common intent is P2, where P1>P2>1; E total It is positively correlated with the level of anxiety.

2. The method according to claim 1, characterized in that, The step of identifying the user's intent type based on the target user's input text obtained at the current identification time and determining the basic sentiment score F(e) includes: using a pre-trained natural language processing model to perform binary classification on the input text to determine whether it is a high-anxiety intent; when it is determined to be a high-anxiety intent, classifying the intent range of the input text to determine whether the input text is a common intent affecting most users in the current area, or a non-common intent affecting only a single user or a small group of users; determining the basic sentiment score F(e) based on the classification result: when it is determined to be a high-anxiety common intent, F(e) = 0.7; when it is determined to be a high-anxiety non-common intent, F(e) = 0.6; wherein, the threshold for determining the high-anxiety intent is 0.4, and the threshold for determining the common intent is 0.

65.

3. The method according to claim 2, characterized in that, The input text is categorized by intent scope to determine whether it represents a common intent affecting a majority of users in the current area or a non-common intent affecting only a single user or a small group of users. This includes: matching the input text with a preset keyword list, which includes a set of common intent keywords and a set of non-common intent keywords; when the input text contains any keyword from the set of common intent keywords, the intent type corresponding to the input text is determined to be a high-anxiety common intent; when the input text does not contain common intent keywords but contains any keyword from the set of non-common intent keywords, the intent type corresponding to the input text is determined to be a high-anxiety non-common intent; wherein, the set of common intent keywords includes keywords such as all flight delays, all delays, large-scale cancellations, and extreme weather; the set of non-common intent keywords includes keywords such as my flight change, lost personal baggage, and ticket refund / rescheduling.

4. The method according to claim 2, characterized in that, After determining whether the input text represents a common intent affecting most users in the current area or a non-common intent affecting only a single user or a small group of users, the method further includes: if the common intent type corresponding to the input text is a high-anxiety common intent, then the anxiety common identifier configured for the corresponding basic emotion score is a high-anxiety common identifier; if the common intent type corresponding to the input text is a high-anxiety non-common intent, then the anxiety common identifier configured for the corresponding basic emotion score is a high-anxiety non-common identifier.

5. The method according to claim 4, characterized in that, After determining the period and distance range of influence that have an emotional contagion effect on the target user based on the current identification time and the target user's current location information, the method further includes: obtaining the basic emotional scores of other users who meet the conditions from the queue of the two-dimensional geographic grid corresponding to the target area according to the filtering conditions defined by the period and distance range of influence; maintaining a first queue and a second queue in each geographic grid, wherein the first queue stores the information vectors of users with short-term contagion influence; the second queue stores the information vectors of users with long-term contagion influence; the validity period T1 of the information vectors of users with short-term contagion influence is less than the validity period T2 of the information vectors of users with long-term contagion influence; the information vectors include the user's individual emotional score, the timestamp when the individual emotional score was generated, the user ID identifier, the user's current location information when the individual emotional score was generated, and the anxiety commonality identifier.

6. The method according to claim 5, characterized in that, The information vectors in the first and second queues are generated according to the following steps: If the target user's E at a certain identification time... individual When ∈(Y1, Y2], the information vector of the target user corresponding to the identification time is stored in the first target queue maintained by the geographic raster to which the target user's location belongs at the identification time; E individual Positively correlated with anxiety level, Y1 and Y2 are the first and second anxiety level thresholds, respectively; if the target user's E at a certain identification time... individual When Y2, the predicted movement trajectory of the target user in T2 corresponding to the identification time is used as the target storage grid, and the geographical grid passed through is used as the target storage grid; the information vector of the target user corresponding to the identification time is stored in the second queue of each target storage grid.

7. The method according to claim 6, characterized in that, The target area is the waiting hall; the predicted movement trajectory of the target user in T2 corresponding to the identification time, and the geographical grids traversed by it are used as target storage grids, including: determining the affected area according to the user's current location type: if the user is currently at the boarding gate, the affected area is the current grid and its adjacent grids; if the user is currently at the security check area, the affected area is the current grid and the path grids leading to the nearest boarding gate; if the user is currently at the catering area, the affected area is the current grid and the four grids adjacent to it in front, behind, left, and right.

8. The method according to claim 5, characterized in that, The individual sentiment scores of users in the information vector of the second queue satisfy the following condition: Among them, E t individual tk is the individual emotion score of the user at the current identification time; tk is the time corresponding to the timestamp in the user's information vector; E tk individual The individual emotion score in the user information vector; t now T1 represents the current identification time; T2 represents the validity period of the information vector that the short-term infectious disease affects the user.

9. A non-transitory computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements a user emotion recognition method based on text information as described in any one of claims 1 to 8.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements a user emotion recognition method based on text information as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Emotional state evaluation method and device, intelligent terminal and storage medium

    CN112472088A

  • Voice emotion recognition model training method, voice emotion recognition method and device

    CN115881103A