Text and voice emotion recognition method and device and electronic equipment
By combining continuous text fragments with the same emotional tendencies and identifying their emotional values, the problem of not being able to accurately identify complete emotional expressions in text emotion recognition is solved, and accurate recognition of persistent emotions is achieved.
Patent Information
- Application Number
- CN202510359354.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-25
AI Technical Summary
Existing text emotion recognition methods are difficult to accurately identify the emotional value of the identified person's complete emotional expression, mainly because the fragment division method cannot contain complete emotional expression.
By obtaining the number of positive and negative affective words in the text fragments, the complete text is constructed using timestamps, and the continuous fragments with the same emotional tendencies are merged to form a second text fragment containing the complete emotional expression, and the emotional values of these fragments are then identified.
Accurate recognition of the complete emotional expression of the identified object is achieved, allowing for better identification of persistent emotional states, such as the overall degree of anger during anger.
Smart Images

Figure CN120373309A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of emotion recognition, and in particular to an emotion recognition method, device and electronic device for text and speech. Background Art
[0002] In speech emotion recognition, usually the speech is first recognized and converted into text, then the text is split into multiple segments, and finally the emotion words in each segment are respectively recognized, and the emotion value of each segment is determined based on the emotion words included in each segment. Among them, since the segments are usually divided at punctuation positions, that is, a sentence forms a segment, and the emotional expression of a person is usually continuous, such as continuous happiness or continuous sadness, it is difficult for a single segment obtained by the above division method to contain a complete emotional expression, and a complete emotional expression refers to the complete expression of a single emotion.
[0003] Therefore, the current text emotion recognition method is difficult to accurately recognize the emotion value of the complete emotional expression of the person to be recognized. Summary of the Invention
[0004] In the present invention, an emotion recognition method, device and electronic device for text and speech are provided to solve the problem that the current text emotion recognition method is difficult to accurately recognize the emotion value of the complete emotional expression of the person to be recognized.
[0005] In a first aspect, the present invention provides an emotion recognition method for text, including:
[0006] Obtain a plurality of first text segments to be recognized, respectively recognize and determine the number of positive emotion words and negative emotion words in the plurality of first text segments, and the plurality of first text segments form the complete text to be recognized according to time stamps;
[0007] Determine the first text segments with more positive emotion words than negative emotion words as positive segments, and determine the first text segments with more negative emotion words than positive emotion words as negative segments;
[0008] Merge the continuous positive segments and the continuous negative segments in the plurality of first text segments to obtain a plurality of second text segments;
[0009] Determine the emotion values of the plurality of second text segments.
[0010] In a second aspect, the present invention provides an emotion recognition method for speech, including:
[0011] Recognize a plurality of speech segments of the object to be recognized to obtain a plurality of first text segments to be recognized;
[0012] Performing sentiment recognition on the several first text segments according to the text sentiment recognition method described in the first aspect to obtain the sentiment values of several second text segments;
[0013] Outputting the sentiment recognition results of several said speech segments, where the sentiment recognition results include several said second text segments and their sentiment values.
[0014] In a third aspect, a speech sentiment recognition system is provided in the present invention, including:
[0015] At least one speech acquisition device for acquiring the speech of the object to be recognized;
[0016] A sentiment recognition device for obtaining several speech segments of the object to be recognized from the speech acquisition device and performing sentiment recognition on the several speech segments of the object to be recognized by the speech sentiment recognition method described in the second aspect to obtain the sentiment recognition result of the object to be recognized corresponding to each speech acquisition device
[0017] In a fourth aspect, a text sentiment recognition device is provided in the present invention, including:
[0018] A text recognition module for obtaining several first text segments to be recognized, respectively identifying and determining the number of positive sentiment words and negative sentiment words in the several first text segments, and the several first text segments form the complete text to be recognized according to the time stamp;
[0019] A segment recognition module for determining the first text segments with more positive sentiment words than negative sentiment words as positive segments and the first text segments with more negative sentiment words than positive sentiment words as negative segments;
[0020] A segment merging module for merging consecutive positive segments and consecutive negative segments among the several first text segments to obtain several second text segments;
[0021] A sentiment determination module for determining the sentiment values of the several second text segments.
[0022] In a fifth aspect, a speech sentiment recognition device is provided in the present invention, including:
[0023] A speech recognition module for recognizing and segmenting the speech of the object to be recognized to obtain several first text segments to be recognized;
[0024] A sentiment recognition module for performing sentiment recognition on the several first text segments according to the text sentiment recognition method described in the first aspect to obtain the sentiment values of several second text segments;
[0025] A result output module, configured to output the emotion recognition result of the speech, where the emotion recognition result includes a plurality of the second text segments and their emotion values
[0026] In a sixth aspect, the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the text emotion recognition method according to any one of the first aspect or the speech emotion recognition method according to the second aspect.
[0027] Compared with the related art, the text and speech emotion recognition methods provided in the present invention include two stages: rough recognition and fine recognition. In the rough recognition stage, it is determined that consecutive text segments with the same emotion tendency jointly constitute a complete emotion expression. Therefore, consecutive first text segments with the same emotion flag bits are merged to obtain each second text segment, and each second text segment contains a complete emotion expression. By recognizing the emotion value of each second text segment, the emotion values of each complete emotion expression of the object to be recognized can be recognized more accurately, solving the problem that the current text emotion recognition method is difficult to accurately recognize the emotion value of the complete emotion expression of the person to be recognized.
[0028] Details of one or more embodiments of the present application are set forth in the following drawings and description, so that other features, objects, and advantages of the present application will become more concise and understandable. Description of the Drawings
[0029] Figure 1 is a flowchart of the speech emotion recognition method provided in this embodiment;
[0030] Figure 2 is a flowchart of the text emotion recognition method provided in this embodiment. Detailed Embodiments
[0031] To understand the purpose, technical solution, and advantages of the present application more clearly, the present application will be described and illustrated below with reference to the drawings and embodiments.
[0032] Unless otherwise defined, technical or scientific terms used in this application shall have the ordinary meanings understood by those with ordinary skills in the technical field to which this application belongs. In this application, words such as "a", "an", "one", "the", "these" and the like do not indicate a limitation in quantity, and they can be singular or plural. The terms "comprising", "including", "having" and any variants thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connected", "coupled" and the like used in this application do not limit to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The term "plurality" used in this application means two or more. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third" and the like used in this application only distinguish similar objects and do not represent a specific order for the objects.
[0033] In this embodiment, a method for emotion recognition of speech is provided. Figure 1 It is a flowchart of the method for emotion recognition of speech provided in this embodiment, as Figure 1 shown. This process includes step S110, step S120 and step S130.
[0034] Step S110: Recognize a plurality of speech segments of the object to be recognized to obtain a plurality of first text segments to be recognized.
[0035] In this embodiment, the speech of the object to be recognized is collected by a speech acquisition device. Among them, based on the automatic speech recognition technology (ASR) in deep learning, the speech can be converted into text T, and the i-th first text segment is represented as T i . The corresponding i-th speech segment is represented as N i . In addition, multiple speech acquisition devices can be used to collect the speech of multiple objects to be recognized simultaneously. After multiple speeches are collected, the method for emotion recognition of speech in this embodiment can be used to recognize multiple speeches simultaneously. The method for emotion recognition of speech in this embodiment is applied to a speech processing system. As follows, an example is used to illustrate the operation process of the speech acquisition device and the speech processing system.
[0036] 1. The voice processing system starts, loads relevant communication configuration files, and waits to receive data (the voice of the object to be recognized).
[0037] 2. Connect the voice acquisition device to the voice processing system or log in to the voice processing system to ensure successful two-way data communication.
[0038] 3. The voice acquisition device sets audio information such as the sampling rate, number of channels, and encoding method required for voice recording, and temporarily stores the data locally to facilitate retransmission in case of transmission failure. At the same time, save information such as the device address, device number, connection time segment to the system, communication link status, communication serial number, and number of accesses.
[0039] 4. The voice acquisition device uploads the recorded voice to the voice processing system after segmenting or chunking it (obtaining multiple voice segments) according to different methods such as time slicing or capacity chunking. And at the same time upload relevant information such as segmenting or chunking information, device address, and device number. For the segment or chunk of voice with transmission failure, retransmit it after timeout to ensure complete transmission. Among them, segmenting or chunking is a voice segment.
[0040] 5. The voice processing system simultaneously receives data transmissions from multiple voice acquisition devices through methods such as multi-GPU or multi-threading, and the processor preprocesses the multi-device voice segment or chunk set in parallel, including noise filtering, normalization processing, and outlier processing, etc.
[0041] 6. Assume that the number of a certain voice acquisition device is N, then after data cleaning, the m segments or chunks of audio are stored in the form of Map mapping [device number: audio queue], and the following results can be obtained:
[0042] N: [N1 N2 … N m
[0043] Among them, N m represents the m-th voice segment in the voice acquisition device N.
[0044] 7. The voice processing system will process the mapped data obtained from multiple voice acquisition devices in parallel, and apply the voice emotion recognition method in this embodiment to return the emotion recognition results to each voice acquisition device according to the device number N for storage and display on the system.
[0045] Step S120, perform emotion recognition on a number of first text segments according to the text emotion recognition method in this embodiment to obtain the emotion values of a number of second text segments.
[0046] After obtaining a number of first text segments, the emotion recognition method provided in this embodiment can be used to recognize them. Specifically, refer to Figure 2 , the method for emotion recognition of text provided in this embodiment includes step S210, step S220, step S230, and step S240.
[0047] Step S210: Obtain a number of first text segments to be recognized, and respectively identify and determine the number of positive emotion words and negative emotion words in the number of first text segments. The number of first text segments constitutes the complete text to be recognized according to the time stamp.
[0048] Specifically, the emotion value of each emotion word in each first text segment can be determined through an emotion dictionary. The emotion dictionary records the emotion values of a number of preset emotion words, and the emotion value is used to represent the degree of positive emotion or negative emotion; then, in each first text segment, the emotion words with positive emotion are determined as positive emotion words, and the emotion words with negative emotion are determined as negative emotion words; finally, the number of positive emotion words and negative emotion words in each first text segment is counted.
[0049] Among them, the emotion dictionary is a pre-customized dictionary for recording different emotion values.
[0050] Step S220: Determine the first text segment with more positive emotion words than negative emotion words as a positive segment, and determine the first text segment with more negative emotion words than positive emotion words as a negative segment.
[0051] In this embodiment, the positive emotion value can be represented by a positive number, and the negative emotion value can be represented by a negative number. The absolute value of the positive and negative numbers reflects different emotion degrees. For example, the value of "like" is set to 4, and the value of "hate" is set to "-4". The higher the score of the word with stronger emotional color.
[0052] According to the word segmentation result of the i-th first text segment T i , all word segments in the text segment are matched and searched for the document content, and the number of positive and negative values is counted (the number of positive emotion values and negative emotion words in the text segment), and the difference between the two is used to measure the emotional tendency of the entire text. The result greater than 0 indicates a positive emotional tendency, and the corresponding first text segment is determined as a positive segment; otherwise, it is negative, and the corresponding first text segment is determined as a negative segment. The positive segment and the negative segment can be represented by 1 and 0 respectively, and stored in the flag bit F to achieve the first step of rough classification of emotion analysis. Among them, Fi represents the flag bit of the i-th first text segment.
[0053] By repeatedly executing the above steps, all voice segments (first text segments) of the current voice acquisition device can be traversed, and finally the emotion flag bit matrix of all voice segments (first text segments) can be obtained: F = [F1 F2 …F m .
[0054] Step S230: Merge consecutive positive segments and consecutive negative segments in several first text segments to obtain several second text segments.
[0055] After obtaining the above sentiment flag matrix, that is, after determining whether each first text segment is a positive segment or a negative segment, segment integration can be performed. Specifically, the platform value that appears each time can be calculated according to the sentiment flag matrix F, that is, consecutive 1s or 0s. If the flag bits for a period of time are all 1, it indicates that there is a relatively obvious emotional bias during this period. Combining the time stamps, merge consecutive segments in the same direction (merge the first text segments with consecutive sentiment flag bits of 1, and merge the first text segments with consecutive sentiment flag bits of 0). Finally, obtain k (k < m) second text segments and their corresponding k voice segments N = [N1 N2 … N k .
[0056] Step S240: Determine the sentiment values of several second text segments.
[0057] Specifically, the sentiment values of each sentiment word in each second text segment can be determined through a sentiment dictionary. The sentiment dictionary records the sentiment values of several preset sentiment words, and the sentiment values are used to represent the degree of positive sentiment or negative sentiment; determine the sentiment values of each second text segment according to the sentiment values of each sentiment word in each second text segment respectively.
[0058] Among them, in this step, different sentiment values are marked for each sentiment word in combination with its sentiment characteristics. For any sentiment word t in the sentiment dictionary, its specific sentiment value is represented by DT(t). The number of negation words before the sentiment word t is represented by x(t), and the degree adverb that restricts the sentiment word t is represented by L(t), which represents the strength of the sentiment. The higher the value defined for a word with more obvious sentiment color, calculate the specific emotion type of each second text segment (voice segment), and delimit different emotion categories according to different numerical ranges to obtain the emotion recognition label E(N i ) until the sentiment recognition of k second text segments (voice segments) is completed, realizing refined sentiment recognition.
[0059] Among them, the formula for determining the sentiment value of each second text segment is:
[0060]
[0061] Among them, E(N i ) represents the sentiment value of the i-th second text segment, N iLet \(n_i\) denote the number of sentiment words in the \(i\)-th second text segment, \(L(t)\) denote the degree value of the adverb of the \(t\)-th sentiment word in the \(i\)-th second text segment, \(DT(t)\) denote the sentiment value of the \(t\)-th sentiment word in the \(i\)-th second text segment, and \(x(t)\) denote the number of negation words of the \(t\)-th sentiment word in the \(i\)-th second text segment.
[0062] Through the above formula, the overall sentiment value of each second text segment can be calculated, that is, the overall sentiment value of the corresponding voice segments is obtained. The description of the above formula is as follows:
[0063] Entropy is used to calculate the energy value of the \(i\)-th sentiment word, that is, \(p(DT(t))\times\log p(DT(t))\). The level of information contained in the entropy value here indicates the intensity of the sentiment, and the positive or negative value indicates the goodness or badness of the sentiment. However, due to the existence of negation words and degree adverbs in the language, the influence of these two on the sentiment entropy value of the word segmentation also needs to be considered. When there is a negation word before a sentiment word, such as "not" before "like", the sentiment entropy value should be multiplied by -1. When there are \(x\) negation words, it is \((-1)^{x}\) x ; Similarly, for degree adverbs, there are also set values in the sentiment dictionary. For example, "a little" is set to 0.4, and "very" is set to 3. Such words also indicate the intensity of the emotion and are used as the weight of the sentiment entropy value. Therefore, the final calculation result of the sentiment value of the \(i\)-th sentiment word is: \((-1)^{x}\) x \(\times L(t)\times p(DT(t))\times\log p(DT(t))\). After the sentiment values of all sentiment words in a second text segment are calculated and accumulated, the sentiment value of this second text segment is obtained. Finally, it is classified according to the score. For example, \((0.5\sim - 0.5)\) is defined as a neutral sentiment, and \((-3\sim -4)\) is defined as a sad sentiment, and finally the sentiment label of this second text segment is obtained.
[0064] Step S130, output the sentiment recognition result of the voice, and the sentiment recognition result includes several second text segments and their sentiment values.
[0065] The sentiment recognition result of the voice can also include several first text segments and the number of positive sentiment words and negative sentiment words in them (or directly give the difference between the number of positive sentiment words and negative sentiment words and the sentiment flag bit), and the timestamps of several first text segments and several second text segments. Then the sentiment recognition result corresponding to the voice acquisition device N can be expressed as: N: [T D F E]. Wherein, T is the speech recognition result, including the first text segment and the second text segment, D is the timestamp information, including the timestamps of each first text segment and the second text segment, F includes the flag bits of each first text segment, and E includes the sentiment values of each second text segment.
[0066] As can be seen from the above description, the method for emotion recognition of text and speech provided in this embodiment includes two stages: rough recognition and fine recognition. In the rough recognition stage, it is determined that consecutive text segments with the same emotional tendency jointly constitute a complete emotional expression. Therefore, consecutive first text segments with the same emotional flag bits are merged to obtain respective second text segments, and each second text segment contains a complete emotional expression. By recognizing the emotional values of each second text segment, the emotional values of each complete emotional expression of the object to be recognized can be recognized more accurately, solving the problem that the current text emotion recognition method is difficult to accurately recognize the emotional values of the complete emotional expressions of the person to be recognized. Exemplarily, if the object to be recognized is relatively angry during a certain period of time, the method for emotion recognition of text and speech in this embodiment can determine the overall anger level of the object to be recognized during the entire angry period.
[0067] In this embodiment, a speech emotion recognition system is also provided, which includes at least one speech acquisition device and an emotion recognition device.
[0068] At least one speech acquisition device is used to acquire the speech of the object to be recognized.
[0069] The emotion recognition device is used to obtain a plurality of speech segments of the object to be recognized from the speech acquisition device and perform emotion recognition on the plurality of speech segments of the object to be recognized by the speech emotion recognition method in this embodiment, so as to obtain the emotion recognition result of the object to be recognized corresponding to each speech acquisition device.
[0070] In this embodiment, a text emotion recognition device is also provided. This device is used to implement the text emotion recognition provided in this embodiment, and the parts that have been described will not be elaborated again. The following terms such as "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0071] The text emotion recognition device provided in this embodiment includes: a text recognition module, a segment recognition module, a segment merging module, and an emotion determination module.
[0072] The text recognition module is used to obtain a plurality of first text segments to be recognized, and respectively recognize and determine the number of positive emotion words and negative emotion words in the plurality of first text segments. The plurality of first text segments form the complete text to be recognized according to the time stamps.
[0073] The segment recognition module is used to determine the first text segment with more positive emotion words than negative emotion words as a positive segment, and the first text segment with more negative emotion words than positive emotion words as a negative segment.
[0074] The segment merging module is used to merge consecutive positive segments and consecutive negative segments in a number of first text segments to obtain a number of second text segments.
[0075] The emotion determination module is used to determine the emotion values of a number of second text segments.
[0076] In this embodiment, an emotion recognition device for speech is further provided. This device is used to implement the emotion recognition method for speech provided in this embodiment, and those that have been described will not be elaborated here. The following terms such as "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0077] The emotion recognition device for speech provided in this embodiment includes: a speech recognition module, an emotion recognition module, and a result output module.
[0078] The speech recognition module is used to recognize and segment the speech of the recognized object to obtain a number of first text segments to be recognized;
[0079] The emotion recognition module is used to perform emotion recognition on a number of first text segments according to the emotion recognition method for text in this embodiment to obtain the emotion values of a number of second text segments;
[0080] The result output module is used to output the emotion recognition result of the speech, and the emotion recognition result includes a number of second text segments and their emotion values.
[0081] It should be noted that the above-mentioned each module can be a functional module or a program module, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned each module can be located in the same processor; or the above-mentioned each module can also be located in different processors in any combined form.
[0082] In this embodiment, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to perform the emotion recognition method for text or the emotion recognition method for speech in this embodiment.
[0083] It should be understood that the specific embodiments described here are only used to explain this application, rather than to limit it. According to the embodiments provided in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of this application.
[0084] Obviously, the accompanying drawings are only some examples or embodiments of the present application. For those of ordinary skill in the art, the present application can also be applied to other similar situations based on these drawings without creative efforts. Additionally, it can be understood that although the work done during this development process may be complex and time-consuming, for those of ordinary skill in the art, certain design, manufacturing, or production changes based on the technical content disclosed in the present application are only routine technical means and should not be regarded as insufficient disclosure of the present application.
Claims
1. A method for sentiment recognition of text, characterized in that, Including: Obtain a plurality of first text segments to be recognized, respectively recognize and determine the number of positive sentiment words and negative sentiment words in the plurality of first text segments, and the plurality of first text segments form a complete text to be recognized according to timestamps; Determine the first text segments with more positive sentiment words than negative sentiment words as positive segments, and determine the first text segments with more negative sentiment words than positive sentiment words as negative segments; Merge the continuous positive segments and the continuous negative segments in the plurality of first text segments to obtain a plurality of second text segments; Determine the sentiment values of the plurality of second text segments.
2. The method for text emotion recognition according to claim 1, characterized in that, Respectively recognizing and determining the number of positive sentiment words and negative sentiment words in the plurality of first text segments includes: Determine the sentiment value of each sentiment word in each first text segment through a sentiment dictionary, where the sentiment dictionary records the sentiment values of a plurality of preset sentiment words, and the sentiment value is used to represent the degree of positive sentiment or negative sentiment; In each of the first text segments, determine the sentiment words with positive sentiment as positive sentiment words, and determine the sentiment words with negative sentiment as negative sentiment words; Count the number of positive sentiment words and negative sentiment words in each first text segment.
3. The emotional recognition method of the text according to claim 1, characterized in that Determining the sentiment values of the plurality of second text segments includes: Determine the sentiment value of each sentiment word in each second text segment through a sentiment dictionary, where the sentiment dictionary records the sentiment values of a plurality of preset sentiment words, and the sentiment value is used to represent the degree of positive sentiment or negative sentiment; Respectively determine the sentiment values of the second text segments according to the sentiment values of the sentiment words in each second text segment.
4. The method for sentiment recognition of the text according to claim 3, wherein The determination formula for the sentiment value of each second text segment is: Among them, E(N i ) represents the sentiment value of the i-th second text segment, N i represents the number of sentiment words in the i-th second text segment, L(t) represents the degree value of the adverb of the t-th sentiment word in the i-th second text segment, DT(t) represents the sentiment value of the t-th sentiment word in the i-th second text segment, and x(t) represents the number of negation words of the t-th sentiment word in the i-th second text segment.
5. A method for emotional recognition of speech, characterized in that, Including: Recognize a plurality of speech segments of the object to be recognized to obtain a plurality of first text segments to be recognized; Perform sentiment recognition on the plurality of first text segments according to the text sentiment recognition method according to any one of claims 1-4 to obtain the sentiment values of a plurality of second text segments; Output the sentiment recognition results of the plurality of speech segments, where the sentiment recognition results include the plurality of second text segments and their sentiment values.
6. The method for emotion recognition of speech according to claim 5, characterized in that, The sentiment recognition results further include the plurality of first text segments and the number of their positive sentiment words and negative sentiment words, and the timestamps of the plurality of first text segments and the plurality of second text segments.
7. A voice emotion recognition system, characterized in that, Including: At least one speech acquisition device for acquiring the speech of the object to be recognized; A sentiment recognition device for obtaining a plurality of speech segments of the object to be recognized from the speech acquisition device and performing sentiment recognition on the plurality of speech segments of the object to be recognized through the speech sentiment recognition method according to claim 5 or 6 to obtain the sentiment recognition results of the object to be recognized corresponding to each speech acquisition device.
8. An emotional recognition device for text, characterized in that, Including: A text recognition module for obtaining a plurality of first text segments to be recognized, respectively recognizing and determining the number of positive sentiment words and negative sentiment words in the plurality of first text segments, and the plurality of first text segments form a complete text to be recognized according to timestamps; A segment recognition module, configured to determine the first text segment with more positive sentiment words than negative sentiment words as a positive segment, and determine the first text segment with more negative sentiment words than positive sentiment words as a negative segment; A segment merging module, configured to merge consecutive positive segments and consecutive negative segments among several of the first text segments to obtain several second text segments; An emotion determination module, configured to determine the emotion values of several of the second text segments.
9. An emotional recognition device for speech, characterized in that, It includes: A speech recognition module, configured to recognize and segment the speech of the recognized object to obtain several first text segments to be recognized; An emotion recognition module, configured to perform emotion recognition on the several first text segments according to the text emotion recognition method described in any one of claims 1-4 to obtain the emotion values of several second text segments; A result output module, configured to output the emotion recognition result of the speech, where the emotion recognition result includes several of the second text segments and their emotion values.
10. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to run the computer program to execute the text emotion recognition method described in any one of claims 1 to 4 or the speech emotion recognition method described in claim 5 or 6.