Text-based emotion recognition method and device, terminal equipment and storage medium
By grading and coding analysis of text data, combining emotional and domain recognition methods, the problem of inaccurate response of smart products is solved, more accurate and fast response is achieved, and user experience is improved.
Patent Information
- Application Number
- CN202510693574.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-08-22
AI Technical Summary
Existing smart products cannot accurately identify users' emotions and space in scenarios such as smart homes and smart hospitals, resulting in inaccurate responses.
By performing word segmentation on the text data, after obtaining the text vector, feature hierarchy and coding analysis are performed, multi-grained features are extracted in combination with the local attention mechanism, a two-way gating cycle unit and a converter model, the domain adapter is used to determine the domain of the text data, and feature compression is performed through the multi-head attention mechanism, and finally emotion recognition and response are used to use the emotion classifier.
It improves the accuracy and speed of the response, and can provide more valuable response information based on user emotions and scenarios, enhancing the user experience.
Smart Images

Figure CN120523916A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine learning technology, and in particular to a text-based emotion recognition method, apparatus, terminal device, and storage medium. Background Art
[0002] Current smart products can respond based on the user's input text, but this is limited to semantic responses, and these responses often do not take into account the user's current emotions and the space they are in. Therefore, in some specific scenarios, such as smart homes or smart hospitals, the responses of smart products are not accurate and cannot provide better services. Summary of the Invention
[0003] In view of this, the embodiments of the present application provide a text-based emotion recognition method, apparatus, terminal device and storage medium to solve the problem of inaccurate responses.
[0004] In a first aspect, an embodiment of the present application provides a text-based emotion recognition method, comprising: When obtaining input text data, performing word segmentation processing on the text data to obtain a text vector; Performing feature classification on the text vector to obtain multiple granularity features, and fusing the features of each granularity to obtain a multi-granularity fused feature; Performing coding analysis on the text data, and determining the domain to which the text data belongs through a domain adapter; Performing feature compression on the multi-granularity fusion features to obtain compressed features, inputting the compressed features into an emotion classifier for recognition to obtain emotion recognition results; Outputting response information according to the emotion recognition result and the field to which the text data belongs.
[0005] In some embodiments, performing encoding analysis on the text data and determining the domain to which the text data belongs through a domain adapter includes: Encoding the text data to obtain encoded data; Calculating the cosine similarity between the encoded data and a pre-stored domain prototype through a domain adapter to determine the domain to which the text data belongs; The domain prototype includes one or more of a space type, a technology type, and an expression type of the text data.
[0006] In some embodiments, the feature classification of the text vector to obtain multiple granular features includes: Obtaining word-level sentiment features in the text vector through a local attention mechanism; Acquire sentence-level temporal features in the text vector through a bidirectional gated recurrent unit; The chapter-level global features in the text vector are obtained through the transformer model.
[0007] In some embodiments, the fusing of features at each of the granularities to obtain a multi-granularity fused feature includes: The word-level sentiment features, the sentence-level temporal features, and the passage-level global features are concatenated to obtain multi-granularity fusion features.
[0008] In some embodiments, compressing the multi-granularity fusion features to obtain compressed features includes: Group the multi-head attention into multiple attention head groups; Performing dimensionality reduction projection on the multi-granularity fusion features in each of the attention head groups according to their dimensions, and calculating the attention weight of each of the attention head groups; According to the attention weights of each of the attention head groups, the target feature head is screened out, and the data in the target feature head is fused to obtain a fused attention feature.
[0009] In some embodiments, outputting response information based on the emotion recognition result and the domain to which the text data belongs includes: In combination with the emotion recognition result and the domain to which the text data belongs, a response strategy corresponding to the emotion recognition result in the domain is determined, corresponding response information is generated according to the response strategy, and the response information is output.
[0010] In some embodiments, after outputting the response information, the method further includes: Acquire user feedback information on the response information, and perform self-supervised training based on the feedback information to update parameters of the adapter and the classifier.
[0011] In a second aspect, the present application further provides a text-based emotion recognition device, comprising: A text preprocessing module is used to obtain input text data, perform word segmentation on the text data, and obtain a text vector; A feature extraction module is used to perform feature classification on the text vector to obtain multiple granularity features, and fuse the features of each granularity to obtain a multi-granularity fusion feature; A domain identification module, configured to perform coding analysis on the text data and determine the domain to which the text data belongs through a domain adapter; An emotion recognition module is used to compress the multi-granularity fusion features to obtain compressed features, and input the compressed features into an emotion classifier for recognition to obtain emotion recognition results; The output module is used to output response information according to the emotion recognition result and the field to which the text data belongs.
[0012] In a third aspect, the present application also provides a terminal device, which includes a processor and a memory, wherein the memory stores a computer program, and the processor is used to execute the computer program to implement the text-based emotion recognition method.
[0013] In a fourth aspect, the present application also provides a readable storage medium storing a computer program, which, when executed on a processor, implements the text-based emotion recognition method.
[0014] The embodiments of the present application have the following beneficial effects: This application combines emotion and domain data for analysis to output responses and operations that are more appropriate for the current scenario, thereby improving the accuracy of responses. In the process of data processing, feature compression is used to reduce the amount of data processed, speed up data processing, and improve response speed, making responses faster and more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0016] Figure 1 A flowchart of a text-based emotion recognition method according to an embodiment of the present application is shown; Figure 2 A data flow diagram of a text-based emotion recognition method according to an embodiment of the present application is shown; Figure 3 A schematic structural diagram of a text-based emotion recognition device according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0017] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments.
[0018] The components of the embodiments of the present application generally described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but rather merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort are within the scope of protection of the present application.
[0019] Hereinafter, the terms "including", "having" and their cognates used in various embodiments of the present application are intended only to indicate specific features, numbers, steps, operations, elements, components or combinations of the aforementioned items, and should not be understood as excluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the aforementioned items or adding the possibility of one or more features, numbers, steps, operations, elements, components or combinations of the aforementioned items. In addition, the terms "first", "second", "third" and the like are only used to distinguish descriptions and should not be understood as indicating or implying relative importance.
[0020] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art to which the various embodiments of the present application belong. The terms (such as those defined in generally used dictionaries) will be interpreted as having the same meaning as in the context of the relevant technical field and will not be interpreted as having an idealized meaning or an overly formal meaning unless clearly defined in the various embodiments of the present application.
[0021] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments may be combined with each other.
[0022] Current smart products can only respond based on the semantics of the input text. This application provides a text-based emotion recognition method, which performs feature classification processing on text data for emotion recognition, and outputs response information based on the field to which the text data belongs. The output response information can take into account the user's current emotional state. In some special application scenarios, more valuable response information can be output, thereby improving the user experience.
[0023] The text-based emotion recognition method is described below with reference to some specific embodiments.
[0024] Figure 1 A flow chart of a text-based emotion recognition method according to an embodiment of the present application is shown. Exemplarily, the text-based emotion recognition method includes the following steps: Step S100 , when obtaining input text data, performing word segmentation processing on the text data to obtain a text vector.
[0025] Text data refers to text input by the user through natural language. The text data can be in the form of text or voice. That is, this solution can be used to recognize text in text form or through voice recognition.
[0026] The text-based emotion recognition method of this embodiment can be used in a variety of user-interactive fields, such as the smart home field, where voice interaction with users is common. It can also be used in the medical field, where, for example, when a user inquires about a symptom, in addition to conventional semantic understanding, it is also necessary to recognize the user's emotion in order to select appropriate feedback.
[0027] When performing word segmentation on text data, the segmented text can be converted into a vector representation through a pre-trained word embedding model (such as BERT) to obtain a text vector.
[0028] For example, for text data received in the form of text, word segmentation can be performed to obtain words from long sentences, and then these words can be converted, such as ACSII code conversion, so that each segmented part becomes data in digital format, thereby converting the text data into a text vector.
[0029] For example, for text data received as voice data, the voice can be segmented first, such as into segments every 10 milliseconds, and then feature extraction can be performed based on the data used in voice emotion analysis, such as the energy, frequency, Mel-frequency cepstrum, etc. of each voice segment, so that each voice segment can obtain multi-dimensional data, thereby processing an entire segment of voice data into a text vector in the form of multiple vector segments.
[0030] It can be understood that after word segmentation of a sentence, multiple text vectors can be obtained. In order to facilitate the processing of these text vectors, they can be combined into a feature matrix consisting of multiple text vectors for the sentence.
[0031] Step S200 , performing feature classification on the text vector to obtain multiple granularity features, and fusing the features of each granularity to obtain a multi-granularity fused feature.
[0032] After obtaining the text vector, this embodiment further performs feature classification on the text vector to obtain multiple granular features.
[0033] Among them, the specific method of feature classification in this embodiment includes: obtaining word-level emotional features in the text vector through a local attention mechanism, obtaining sentence-level temporal features in the text vector through a bidirectional gated recurrent unit, and obtaining paragraph-level global features in the text vector through a converter model.
[0034] As can be seen from the names of the various hierarchical features, the hierarchical principles are to classify words, sentences and entire articles. For example, for short sentences, there are word-level emotional features and sentence-level temporal features, and the entire sentence is equivalent to the global feature at the chapter level.
[0035] After obtaining the above-mentioned multiple granular features, these features are fused to obtain multi-granularity fusion features. The fusion method is to splice these features. It can be understood that for word-level emotional features, sentence-level temporal features, and paragraph-level global features, they are actually feature vectors. However, due to different extraction and analysis principles, the specific values will be different, but the dimensions of each feature vector are the same, so they can still be spliced together to obtain a feature matrix, which is the multi-granularity fusion feature.
[0036] For example, the multi-granularity fusion feature can be recorded as F∈R LxD , L is the column length, and D is the vector dimension. The multi-granularity fusion feature F can be represented as a matrix with L rows and D columns. Therefore, it can also be called a multi-granularity fusion feature matrix.
[0037] Step S300: performing coding analysis on the text data and determining the domain to which the text data belongs through a domain adapter.
[0038] This embodiment also analyzes the field to which the text data belongs. The field mentioned in this embodiment refers to one or more of the spatial type, technical type, and expression type of the text data.
[0039] For example, some sentences are only used in certain situations. For example, if a sentence contains highly specialized medical terms, then the entire sentence may belong to the medical field, which is a technical field. Some sentences need to be analyzed in conjunction with their spatial location. For example, in smart home applications, the command "turn off the lights" is related to the user's location. If the user is in the living room, the living room lights must be turned off, and if the user is in the bedroom, the bedroom lights must be turned off. Similarly, whether a sentence is command-based or interactive-based is also different. For example, the command "turn off the lights" and the sentence "What is xxx?" are actually two completely different sentence structures. Command-based sentences require execution, while interactive sentences require corresponding text feedback.
[0040] It is understandable that there are more sub-fields under the above three fields. For example, in the field of space types, different spaces are specific sub-fields. In the field of technology types, different technology categories are also different sub-fields, such as computers, medicine, physics, etc.
[0041] After identifying the domain to which a sentence belongs, combining semantics and sentiment allows for more accurate and effective analysis. It's understandable that the aforementioned domains aren't mutually exclusive. For example, the spatial domain and the technological domain aren't mutually exclusive. A single sentence can possess attributes from multiple domains, but in some monotonous scenarios, it might only be applicable to one domain.
[0042] To this end, this embodiment performs encoding analysis on the text data to obtain encoded text, and then calculates the cosine similarity between the encoded data and a pre-stored domain prototype through a domain adapter to determine the domain to which the text data belongs.
[0043] The encoding method can be an existing encoding method such as ASCII, Unicode, GB2312, or a custom encoding method, such as constructing an alphabet or Chinese character table, and then encoding based on the sequence number of the constructed alphabet or Chinese character table.
[0044] The domain prototype is a domain that is manually set based on prior conditions. The encoding method can be custom encoding or existing encoding methods. After encoding, the text data will also become a feature vector, and then the cosine similarity between the encoded data and each domain prototype can be calculated.
[0045] It is understandable that through the above calculation, a text data can calculate multiple cosine similarities. For example, from 5 domain prototypes, the similarity of domain A is 30%, domain B is 15%, domain C is 10%, domain D is 20%, and domain E is 15%. It can be seen that a text data can have different similarities with multiple domains. In order to determine which domain the text data belongs to, or which domain should be used for subsequent recognition operations, the first few domains with higher similarity can be selected as the domain to which it belongs, such as the top two domains with the highest similarity. In this example, domain A and domain D are the domains corresponding to the text data.
[0046] Step S400 , performing feature compression on the multi-granularity fusion feature to obtain compressed features, and inputting the compressed features into a sentiment classifier for recognition to obtain sentiment recognition results.
[0047] Combined with the aforementioned acquisition of multi-granularity fusion features, it can be seen that each feature vector has a certain dimension, which is determined by the type of features extracted. The more types of features extracted, the higher the dimension. The higher the dimension, the greater the computational burden during calculation, and the longer the time to recognize emotions. Therefore, dimensionality reduction and compression operations are required.
[0048] In this embodiment, dimensionality reduction can be performed by using PCA (principal component analysis), multi-head attention grouping, or local linear embedding.
[0049] For example, taking multi-head attention grouping as an example of dimensionality reduction, the transformer model has a multi-head attention mechanism. To perform dimensionality reduction, the multi-head attention can be grouped first, that is, several attention heads are grouped together to obtain multiple attention head groups. For example, if there are 100 attention heads, 10 heads can be grouped together to obtain 10 attention head groups.
[0050] Then, based on the multi-granularity fusion features obtained in the previous step, they are split based on dimension. For example, if the dimension of the multi-granularity fusion feature matrix is 780, then for 10 attention head groups, the dimension can be divided equally into 10 feature matrices with 78 dimensions. These 10 feature matrices are then input into the 10 attention head groups for dimensionality reduction projection and calculation of the attention weights for each attention head group.
[0051] Each attention head only needs to process a subset of features, significantly reducing the amount of processing required for each attention head and improving processing speed. Furthermore, after dimensionality reduction projection, each attention head group simplifies the original features. After processing, the features can be reassembled to restore the decomposed features to their original dimensions, ensuring no information is lost. In addition, this embodiment can also perform denoising by calculating the attention weights of each attention head group, where the weight can be considered as the importance of the features assigned to each attention head group. The more important the feature, the less likely it is to be noise. Therefore, when performing the final fusion, only the features of the attention head groups with the highest weights can be fused, or the features of the attention head groups with weights greater than a preset value can be fused. This can achieve the purpose of denoising.
[0052] Emotion recognition is performed using an emotion classifier, which is a pre-trained classifier. By inputting the aforementioned multi-granularity fusion features, it can perform emotion recognition and output the matching degree of the multi-granularity fusion features corresponding to various emotions. For example, if the classifier can recognize 7 emotions, it will output the matching degree of the multi-granularity fusion features corresponding to these 7 emotions. These 7 matching degrees are the emotion recognition results.
[0053] Therefore, after obtaining the emotion recognition result, we can further determine what emotion the text data belongs to. For example, we can select a preset number of emotions before the matching degree as the emotion of the text data, or select an emotion with a matching degree greater than a preset value as the emotion of the text data.
[0054] Understandably, everyday conversations don't necessarily contain just one emotion; they can be a mix of multiple. Therefore, we use emotion recognition results to determine the sentiment of the current text input. If several emotions have similar matching scores, around 20%, these can be used as a reference for subsequent feedback, indicating that the text contains complex emotions. If a clear, highly matched emotion, such as 80%, is identified, the remaining emotions can be disregarded and the emotion with an 80% match is directly determined as the sentiment of the current text data.
[0055] Step S500: outputting response information according to the emotion recognition result and the domain to which the text data belongs.
[0056] After determining the emotion type and the field to which the text data belongs, the corresponding response information can be output based on these two results. In this embodiment, the output can rely on a large language model, or the response information can be output through a decision-making AI model.
[0057] The response information of this embodiment is not limited to language and text interaction, but also includes automatic control of some devices.
[0058] For example, in a smart home scenario, by acquiring the noise from the range hood, the current location can be determined to be in the kitchen. Then, the conversations acquired in this background sound can be determined to be related to kitchen-related areas. For example, if the user says "it's so choking", it can be inferred that the smoke is too thick and the power of the range hood is insufficient. Therefore, the output response information is to control the working power of the range hood to increase and enhance the smoke extraction effect.
[0059] In addition, many non-command sentences are collected and identified at home. These sentences often do not require a response, such as some daily sentences such as "I'm so tired after get off work", "You are awesome", "Boring", etc. These sentences do not require a response, but through the above analysis, the user's emotions can be identified. These sentences account for most of the user's words at home, so these sentences can be used to calculate the user's daily mood fluctuation curve. For example, the stress index is continuously high and the mood is bad from 9:00 to 11:00 for several consecutive days, and the user's mood improves every time at 3 pm. In the field of smart homes, soothing music can be played as a response information between 9:00 and 11:00 to adjust the user's mood, while no special adjustment is made after 3 pm, thus reflecting the intelligence of the smart home.
[0060] For example, if a continuous high-decibel conversation is detected in the living room (the voiceprint is recognized as a family member), and text analysis finds the keyword frequencies: "always" (appears 5 times) and "unfair" (3 times), then the method of this embodiment can identify that the family members are emotionally agitated, and that anger, complaints and other negative emotions are predominant, which can be identified as a family conflict. In this case, the response information output can be automatically dimming the lights to a soothing tone, playing ASMR natural sounds, and sending an alert to the mobile phones of other family members.
[0061] It can be seen that in the above two scenarios, the combination of spatial domain and language emotion is applied, so that in the smart home scenario, the smart home system can output more appropriate output, thereby giving users a better experience.
[0062] In addition, user feedback can be obtained for some interactive scenarios. This feedback can be re-input into the model used in this embodiment to optimize parameters. For example, it can be input into the Buddhist, Confucian, and Taoist emotion recognition classifier to adjust the accuracy of emotion recognition, and into the domain adapter to adjust the adapter parameters and optimize the domain recognition results. This achieves the technical effect of self-supervised learning.
[0063] like Figure 2 As shown, this is a data flow diagram of this embodiment. When text data is obtained, the text data will be processed into text vectors to obtain multi-granularity fusion features, and then emotion recognition will be performed. In addition, the text data will be used for encoding to perform field identification to obtain the corresponding field. Finally, the field and emotion recognition results are combined to obtain the corresponding response information.
[0064] The method of this embodiment uses multi-granularity fusion features to identify emotions and combines them with domain recognition to output responses. Furthermore, during data processing, multi-head attention splitting is used to compress and project the multi-granularity fusion features, and then weighting is used to perform noise removal. This reduces computational complexity while also reducing interference data and enhancing recognition results. This results in the method of this embodiment achieving rapid and high-speed response and improving recognition accuracy.
[0065] Figure 3 The text-based emotion recognition device according to an embodiment of the present application is shown, and the device includes: The text preprocessing module 10 is used to obtain input text data, perform word segmentation processing on the text data, and obtain a text vector; A feature extraction module 20 is used to perform feature classification on the text vector to obtain multiple granularity features, and fuse the features of each granularity to obtain a multi-granularity fused feature; A domain identification module 30 is used to perform coding analysis on the text data and determine the domain to which the text data belongs through a domain adapter; The emotion recognition module 40 is used to compress the multi-granularity fusion features to obtain compressed features, and input the compressed features into the emotion classifier for recognition to obtain emotion recognition results; The output module 50 is configured to output response information based on the emotion recognition result and the domain to which the text data belongs.
[0066] It can be understood that the device of this embodiment corresponds to the text-based emotion recognition method of the above embodiment, and the optional options in the above embodiment are also applicable to this embodiment, so they will not be repeated here.
[0067] The present application also provides a terminal device. Exemplarily, the terminal device includes a processor and a memory, wherein the memory stores a computer program, and the processor runs the computer program to enable the terminal device to execute the functions of each module in the above-mentioned multi-granularity feature emotion recognition method or the above-mentioned multi-granularity feature emotion recognition device.
[0068] The processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including at least one of a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor, etc., and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application.
[0069] The memory may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and electrically erasable programmable read-only memory (EEPROM). The memory is used to store computer programs, and the processor can execute the computer programs accordingly after receiving an execution instruction.
[0070] The present application also provides a readable storage medium for storing the computer program used in the above-mentioned terminal device.
[0071] For example, the readable storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0072] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in an alternative implementation, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the structure diagram and / or flowchart, and the combination of boxes in the structure diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0073] In addition, the functional modules or units in the various embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0074] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a smart phone, personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.
[0075] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. A text-based emotion recognition method, characterized in that: include: Obtain input text data, perform word segmentation on the text data, and obtain a text vector; Performing feature classification on the text vector to obtain multiple granularity features, and fusing the features of each granularity to obtain a multi-granularity fused feature; Performing coding analysis on the text data, and determining the domain to which the text data belongs through a domain adapter; Performing feature compression on the multi-granularity fusion feature to obtain compressed features, inputting the compressed features into an emotion classifier for recognition, and obtaining an emotion recognition result; Outputting response information according to the emotion recognition result and the field to which the text data belongs.
2. The text-based emotion recognition method according to claim 1, characterized in that The performing coding analysis on the text data and determining the domain to which the text data belongs through a domain adapter includes: Encoding the text data to obtain encoded data; Calculating the cosine similarity between the encoded data and a pre-stored domain prototype through a domain adapter to determine the domain to which the text data belongs; The domain prototype includes one or more of the spatial type, technical type and expression type of the text data.
3. The text-based emotion recognition method according to claim 1, characterized in that The feature classification of the text vector is performed to obtain multiple granular features, including: Obtaining word-level sentiment features in the text vector through a local attention mechanism; Acquire sentence-level temporal features in the text vector through a bidirectional gated recurrent unit; The chapter-level global features in the text vector are obtained through the transformer model.
4. The text-based emotion recognition method according to claim 3, characterized in that The fusion of the features of each granularity to obtain a multi-granularity fusion feature includes: The word-level sentiment features, the sentence-level temporal features, and the passage-level global features are concatenated to obtain multi-granularity fusion features.
5. The text-based emotion recognition method according to claim 1, characterized in that The step of compressing the multi-granularity fusion features to obtain compressed features includes: Group the multi-head attention into multiple attention head groups; Performing dimensionality reduction projection on the multi-granularity fusion features in each of the attention head groups according to their dimensions, and calculating the attention weight of each of the attention head groups; According to the attention weights of each of the attention head groups, the target feature head is screened out, and the data in the target feature head is fused to obtain a fused attention feature.
6. The text-based emotion recognition method according to claim 2, characterized in that Outputting response information according to the emotion recognition result and the domain to which the text data belongs includes: In combination with the emotion recognition result and the domain to which the text data belongs, a response strategy corresponding to the emotion recognition result in the domain is determined, and corresponding response information is output according to the response strategy.
7. The text-based emotion recognition method according to claim 1, characterized in that After outputting the response information, the method further includes: Acquire user feedback information on the response information, and perform self-supervised training based on the feedback information to update parameters of the adapter and the classifier.
8. A text-based emotion recognition device, characterized in that: include: A text preprocessing module is used to obtain input text data, perform word segmentation on the text data, and obtain a text vector; A feature extraction module is used to perform feature classification on the text vector to obtain multiple granularity features, and fuse the features of each granularity to obtain a multi-granularity fusion feature; A domain identification module, configured to perform coding analysis on the text data and determine the domain to which the text data belongs through a domain adapter; An emotion recognition module is used to compress the multi-granularity fusion features to obtain compressed features, and input the compressed features into an emotion classifier for recognition to obtain emotion recognition results; The output module is used to output response information according to the emotion recognition result and the field to which the text data belongs.
9. A terminal device, characterized in that: The terminal device includes a processor and a memory, the memory stores a computer program, and the processor is used to execute the computer program to implement the text-based emotion recognition method according to any one of claims 1 to 7.
10. A readable storage medium, characterized in that: The device stores a computer program, which, when executed on a processor, implements the text-based emotion recognition method according to any one of claims 1 to 7.