Emotion recognition method and device, computer device and storage medium

CN116884442BActive Publication Date: 2026-09-22CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310952605.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-31
Publication Date
2026-09-22
Estimated Expiration
2043-07-31

AI Technical Summary

Technical Problem

[0004]但是,在客户服务过程中,需要关注的不只是客户的问题,还需要及时关注客户情绪;在人工服务过程中,可以由人工客服根据自身经验来判断,及时安抚客户,但在自助服务过程中,是无法关注客户情绪,若情绪未能及时安抚,则可能会导致客户升级投诉

Benefits of technology

[0043]上述情绪识别方法、装置、计算机设备和存储介质,在与目标用户进行语音交互的过程中,从目标用户的语音记录中,提取目标用户的当前声纹信息和语音内容信息,从当前声纹信息和语音内容信息两个维度对用户情绪进行识别,保证了对情绪分析的全面性,提高了对用户情绪分析的准确度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116884442B_ABST
    Figure CN116884442B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, in particular to an emotion recognition method and device, computer equipment and a storage medium. The method comprises the following steps: extracting current voiceprint information and voice content information of a target user from a voice record of the target user; analyzing the current voiceprint information to obtain a sound emotion recognition result; analyzing the voice content information to obtain a keyword emotion recognition result; and determining a target emotion recognition result according to the sound emotion recognition result and the keyword emotion recognition result. The application can accurately recognize user emotions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to an emotion recognition method, apparatus, computer device, and storage medium. Background Technology

[0002] Customer service is a primary way for businesses to obtain user feedback and resolve user issues.

[0003] Currently, with the development of artificial intelligence technology, enterprises are using intelligent technologies to improve the efficiency of self-service problem solving in order to better reduce service costs.

[0004] However, in customer service, it is not only necessary to pay attention to the customer's problems, but also to pay attention to the customer's emotions. In the process of human service, the customer service representative can judge and soothe the customer in a timely manner based on their own experience. However, in the process of self-service, it is impossible to pay attention to the customer's emotions. If the emotions are not soothed in time, it may lead to the customer escalating to a complaint. Summary of the Invention

[0005] Therefore, it is necessary to provide an emotion recognition method, device, computer equipment, and storage medium that can accurately identify user emotions in order to address the aforementioned technical problems.

[0006] Firstly, this application provides an emotion recognition method, which includes:

[0007] Extract the target user's current voiceprint information and voice content information from the target user's voice recording;

[0008] The current voiceprint information is analyzed to obtain the voice emotion recognition result;

[0009] The speech content information is analyzed to obtain the keyword emotion recognition results;

[0010] Based on the voice emotion recognition results and keyword emotion recognition results, the target emotion recognition result is determined.

[0011] In one embodiment, the current voiceprint information is analyzed to obtain a voice emotion recognition result, including:

[0012] Verify the continuity of the current voiceprint information;

[0013] If the verification is successful, the current voiceprint information is compared with the basic voiceprint information corresponding to the target user to obtain the voice emotion recognition result.

[0014] In one embodiment, the continuity verification of the current voiceprint information includes:

[0015] The current voiceprint information is compared with the target user's historical voiceprint information to obtain the continuity verification result.

[0016] In one embodiment, the method further includes:

[0017] Among multiple candidate voiceprint information, the candidate voiceprint information that is located before and adjacent to the current voiceprint information is identified as the target user's historical voiceprint information; wherein, the multiple candidate voiceprint information is extracted from the target user's voice records according to a preset sampling interval.

[0018] In one embodiment, determining the target emotion recognition result based on the voice emotion recognition result and the keyword emotion recognition result includes:

[0019] If the voice emotion recognition result or the keyword emotion recognition result indicates an abnormal emotion, then the target emotion recognition result is determined to be an abnormal emotion.

[0020] If the voice emotion recognition result is inconsistent with the keyword emotion recognition result, then the voice emotion recognition result will be determined as the target emotion recognition result.

[0021] In one embodiment, the method further includes:

[0022] If the target emotion recognition result is abnormal, then emotion regulation information is output to the target user.

[0023] Secondly, this application also provides an emotion recognition device, which includes:

[0024] The extraction module is used to extract the target user's current voiceprint information and voice content information from the target user's voice recording;

[0025] The voice comparison module is used to analyze the current voiceprint information and obtain the voice emotion recognition result;

[0026] The keyword comparison module is used to analyze the speech content information and obtain the keyword emotion recognition results;

[0027] The result determination module is used to determine the target emotion recognition result based on the voice emotion recognition result and the keyword emotion recognition result.

[0028] Thirdly, this application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0029] Extract the target user's current voiceprint information and voice content information from the target user's voice recording;

[0030] The current voiceprint information is analyzed to obtain the voice emotion recognition result;

[0031] The speech content information is analyzed to obtain the keyword emotion recognition results;

[0032] Based on the voice emotion recognition results and keyword emotion recognition results, the target emotion recognition result is determined.

[0033] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:

[0034] Extract the target user's current voiceprint information and voice content information from the target user's voice recording;

[0035] The current voiceprint information is analyzed to obtain the voice emotion recognition result;

[0036] The speech content information is analyzed to obtain the keyword emotion recognition results;

[0037] Based on the voice emotion recognition results and keyword emotion recognition results, the target emotion recognition result is determined.

[0038] Fifthly, this application also provides a computer program product comprising a computer program that, when executed by a processor, performs the following steps:

[0039] Extract the target user's current voiceprint information and voice content information from the target user's voice recording;

[0040] The current voiceprint information is analyzed to obtain the voice emotion recognition result;

[0041] The speech content information is analyzed to obtain the keyword emotion recognition results;

[0042] Based on the voice emotion recognition results and keyword emotion recognition results, the target emotion recognition result is determined.

[0043] The aforementioned emotion recognition method, device, computer equipment, and storage medium, during voice interaction with the target user, extract the target user's current voiceprint information and voice content information from the target user's voice recording, and identify the user's emotions from two dimensions: current voiceprint information and voice content information. This ensures the comprehensiveness of emotion analysis and improves the accuracy of user emotion analysis. Attached Figure Description

[0044] Figure 1 This is a diagram illustrating the application environment of an emotion recognition method in one embodiment;

[0045] Figure 2 This is a flowchart illustrating an emotion recognition method in one embodiment;

[0046] Figure 3 This is a flowchart illustrating the voice emotion recognition results in one embodiment;

[0047] Figure 4 This is a flowchart illustrating the process of determining the target emotion recognition result in one embodiment;

[0048] Figure 5 This is a flowchart illustrating the emotion recognition method in another embodiment;

[0049] Figure 6 This is a structural block diagram of an emotion recognition device in one embodiment;

[0050] Figure 7 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0052] The emotion recognition method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or placed in the cloud or on other network servers. Server 104 receives voice recordings from the target user's terminal 102. From the target user's voice recordings, server 104 extracts the target user's current voiceprint information and voice content information; analyzes the current voiceprint information to obtain voice emotion recognition results; analyzes the voice content information to obtain keyword emotion recognition results; and determines the target emotion recognition result based on the voice emotion recognition results and keyword emotion recognition results. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. Server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0053] In one embodiment, such as Figure 2 As shown, an emotion recognition method is provided, which can be applied to... Figure 1 Taking server 104 as an example, the following steps are included:

[0054] S201, extract the target user's current voiceprint information and voice content information from the target user's voice recording.

[0055] Among them, the target user's voice record can refer to the target user's current call record data.

[0056] Specifically, a voiceprint extraction model is used to analyze the speech recordings to extract the current voiceprint information, and a keyword recognition model is used to extract the keyword information in the speech recordings.

[0057] S202, Analyze the current voiceprint information to obtain the voice emotion recognition result.

[0058] Voiceprint information can include information such as sound frequency and sound wavelength.

[0059] Specifically, a voice emotion recognition model can be used to analyze information such as voice frequency and wavelength to predict the emotional changes of the target user. The output of this voice emotion recognition model is the voice emotion recognition result, which can include abnormal emotions and normal emotions. Furthermore, abnormal emotions can also include obvious abnormal emotions and suspected abnormal emotions.

[0060] S203, analyze the speech content information to obtain the keyword emotion recognition results.

[0061] Understandably, keyword extraction refers to extracting keywords from text content data. For example, in speech content information, the statistical information, part-of-speech, and positional information of words are weighted and comprehensively calculated to extract the most semantically relevant core words from the speech content information.

[0062] Specifically, a keyword extraction model is used to segment the speech content information to obtain a word set. By generating word vectors, text vectors are generated. Based on the word vectors and text vectors, keywords are determined from the word set, thereby achieving the goal of effectively extracting keywords from the speech content information.

[0063] Furthermore, the keyword emotion recognition model is used to verify the keywords extracted by the keyword extraction model, and the keyword emotion recognition result is obtained. The keyword emotion recognition result indicates whether there are preset keywords corresponding to abnormal emotions among the extracted keywords.

[0064] Specifically, if there are preset keywords that represent obvious emotional abnormalities, the keyword emotion recognition result is determined to be obviously abnormal; if there are preset keywords that represent suspected emotional abnormalities, the keyword emotion recognition result is determined to be suspected emotional abnormal; if there are no preset keywords that represent either obviously abnormal or suspected emotional abnormalities, the keyword emotion recognition result is determined to be normal.

[0065] S204. Based on the voice emotion recognition results and the keyword emotion recognition results, determine the target emotion recognition result.

[0066] Specifically, the target emotion recognition result is determined by combining the voice emotion recognition result and the keyword emotion recognition result. The target emotion recognition result can include obvious abnormal emotion, suspected abnormal emotion, normal emotion, etc.

[0067] In the aforementioned emotion recognition method, during the voice interaction with the target user, the current voiceprint information and voice content information of the target user are extracted from the target user's voice recording. The user's emotions are identified from two dimensions: the current voiceprint information and the voice content information. This ensures the comprehensiveness of the emotion analysis and improves the accuracy of the user's emotion analysis.

[0068] like Figure 3 As shown, this embodiment provides an optional method for analyzing current voiceprint information to obtain voice emotion recognition results, that is, a method for refining S202. The specific implementation process may include:

[0069] S301, perform continuity verification on the current voiceprint information.

[0070] The process involves comparing the current voiceprint information with the target user's historical voiceprint information to obtain a continuity verification result.

[0071] Specifically, if the similarity between the current voiceprint information and the target user's historical voiceprint information is greater than the similarity threshold, then it is determined that the current voiceprint information and the historical voiceprint information are continuous.

[0072] Optionally, the candidate voiceprint information that is located before and adjacent to the current voiceprint information among multiple candidate voiceprint information can be identified as the target user's historical voiceprint information.

[0073] In this embodiment, multiple candidate voiceprint information are extracted from the target user's voice recording according to a preset sampling interval. In this embodiment, voiceprint sampling can be performed at 5-second intervals to obtain multiple candidate voiceprint information; and in this embodiment, the number of current voiceprint information can be multiple.

[0074] S302. If the verification is successful, the current voiceprint information is compared with the basic voiceprint information corresponding to the target user to obtain the voice emotion recognition result.

[0075] Specifically, basic voiceprint information is extracted from the target user's historical voice recordings and stored in the target user's basic information database. This basic voiceprint information represents the user's voiceprint information when their emotions are stable. Then, both the current voiceprint information and the target user's basic voiceprint information are input into the voice emotion recognition model. In the voice emotion recognition model, multiple current voiceprint information and basic voiceprint information are superimposed to generate a voiceprint composite image. Further, the voice emotion recognition model calculates the deviation between the current voiceprint information and the basic voiceprint information in the voiceprint composite image to obtain the deviation value. The larger the deviation value, the greater the user's emotional fluctuation; the smaller the deviation value, the smaller the user's emotional fluctuation.

[0076] Furthermore, the voice emotion recognition result is determined based on the deviation value. For example, if the deviation value falls within the first numerical range, the voice emotion recognition result is determined to be normal; if the deviation value falls within the second numerical range, the voice emotion recognition result is determined to be suspected abnormal; if the deviation value falls within the third numerical range, the voice emotion recognition result is determined to be obviously abnormal; wherein, the first numerical range is smaller than the second numerical range, and the second numerical range is smaller than the third numerical range.

[0077] For example, if the deviation value is greater than 1.2 but less than 1.5 (equivalent to the second numerical range), the voice emotion recognition result is determined to be a suspected abnormal voice emotion, and the voice recording needs to be marked in time; if the deviation value is greater than 1.5 (equivalent to the third numerical range), the voice emotion recognition result is determined to be a significant abnormal voice emotion, and then the system will interact with the emotion instance library in time to perform emotion soothing.

[0078] Optionally, the aforementioned basic voiceprint information is obtained by the GPT (Generative Pre-trained Transformer) module through machine self-learning based on the extracted basic voiceprint. The GPT module is a machine learning module deployed based on the Transform model and is also used for training voice emotion recognition models and keyword recognition models.

[0079] In this embodiment, by comparing the current voiceprint information with the basic voiceprint information corresponding to the target user, the emotional fluctuations of the target user can be analyzed in a targeted manner, thereby improving the accuracy of the analysis.

[0080] like Figure 4As shown, this embodiment provides an optional method for determining the target emotion recognition result based on the voice emotion recognition result and the keyword emotion recognition result, that is, a method for refining S204. The specific implementation process may include:

[0081] S401, if the voice emotion recognition result or the keyword emotion recognition result indicates an abnormal emotion, then the target emotion recognition result is determined to be an abnormal emotion.

[0082] S402, if the voice emotion recognition result is inconsistent with the keyword emotion recognition result, then the voice emotion recognition result shall be determined as the target emotion recognition result.

[0083] Specifically, if the voice emotion recognition result indicates that the voice emotion is suspected to be abnormal, and the keyword emotion recognition result indicates that the emotion is normal, then the target emotion recognition result is determined to be that the voice emotion is suspected to be abnormal.

[0084] In this embodiment, keyword information and voiceprint recognition emotion changes are combined for calculation; if either voiceprint recognition or keyword recognition determines that the customer's emotion changes significantly, an emotion abnormality alarm is executed; if either voiceprint recognition or keyword recognition shows a slight change, the voiceprint recognition emotion change is the main alarm output.

[0085] Furthermore, the method also includes: if the target emotion recognition result is an abnormal emotion, then outputting emotion regulation information to the target user.

[0086] Specifically, the system retrieves keywords corresponding to abnormal emotions and then uses these keywords to retrieve relevant emotion regulation information from an emotion instance database. This emotion instance database stores instances used to identify and process abnormal customer emotions, and is used to soothe and manage customer emotional fluctuations.

[0087] In this embodiment, emotion regulation information is output to the target user in order to automatically soothe the target user.

[0088] For example, based on the above embodiments, this embodiment provides an optional example of an emotion recognition method. For instance... Figure 5 As shown, the specific implementation process includes:

[0089] S501 extracts the target user's current voiceprint information and voice content information from the target user's voice recording.

[0090] S502, perform continuity verification on the current voiceprint information.

[0091] S503 compares the current voiceprint information with the target user's historical voiceprint information to obtain the continuity verification result.

[0092] S504 If the verification is successful, the current voiceprint information is compared with the basic voiceprint information corresponding to the target user to obtain the voice emotion recognition result.

[0093] S505 analyzes the speech content information to obtain keyword emotion recognition results.

[0094] S506, if the voice emotion recognition result or the keyword emotion recognition result indicates an abnormal emotion, then the target emotion recognition result is determined to be an abnormal emotion.

[0095] S507 If the voice emotion recognition result is inconsistent with the keyword emotion recognition result, then the voice emotion recognition result shall be determined as the target emotion recognition result.

[0096] S508, if the target emotion recognition result is abnormal, then output emotion regulation information to the target user.

[0097] The specific processes of S501-S508 described above can be found in the description of the above method embodiments. Their implementation principles and technical effects are similar, and will not be repeated here.

[0098] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0099] Based on the same inventive concept, this application also provides an emotion recognition device for implementing the emotion recognition method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more emotion recognition device embodiments provided below can be found in the limitations of the emotion recognition method described above, and will not be repeated here.

[0100] In one embodiment, such as Figure 6 As shown, an emotion recognition device 1 is provided, including: an extraction module 11, a voice comparison module 12, a keyword comparison module 13, and a result determination module 14, wherein:

[0101] Extraction module 11 is used to extract the current voiceprint information and voice content information of the target user from the target user's voice recording;

[0102] The voice comparison module 12 is used to analyze the current voiceprint information and obtain the voice emotion recognition result;

[0103] Keyword comparison module 13 is used to analyze the speech content information and obtain the keyword emotion recognition result;

[0104] The result determination module 14 is used to determine the target emotion recognition result based on the voice emotion recognition result and the keyword emotion recognition result.

[0105] In one embodiment, the sound comparison module 12 includes:

[0106] The verification submodule is used to verify the continuity of the current voiceprint information;

[0107] The comparison submodule is used to compare the current voiceprint information with the basic voiceprint information corresponding to the target user if the verification is successful, so as to obtain the voice emotion recognition result.

[0108] In one embodiment, the verification submodule is further configured to: compare the current voiceprint information with the target user's historical voiceprint information to obtain a continuity verification result.

[0109] In one embodiment, the emotion recognition device further includes a storage module, which is used to: determine the candidate voiceprint information that is located before and adjacent to the current voiceprint information from among a plurality of candidate voiceprint information as the historical voiceprint information of the target user; wherein the plurality of candidate voiceprint information is extracted from the target user's voice recordings according to a preset sampling interval.

[0110] In one embodiment, the result determination module 14 is further configured to: if the voice emotion recognition result or the keyword emotion recognition result indicates an emotional abnormality, then determine that the target emotion recognition result is an emotional abnormality.

[0111] If the voice emotion recognition result is inconsistent with the keyword emotion recognition result, then the voice emotion recognition result will be determined as the target emotion recognition result.

[0112] In one embodiment, the emotion recognition device further includes a feedback module, which is used to output emotion regulation information to the target user if the target emotion recognition result is an abnormal emotion.

[0113] Each module in the aforementioned emotion recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0114] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores data for emotion recognition methods. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements an emotion recognition method.

[0115] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0116] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0117] Extract the target user's current voiceprint information and voice content information from the target user's voice recording;

[0118] The current voiceprint information is analyzed to obtain the voice emotion recognition result;

[0119] The speech content information is analyzed to obtain the keyword emotion recognition results;

[0120] Based on the voice emotion recognition results and keyword emotion recognition results, the target emotion recognition result is determined.

[0121] In one embodiment, when the processor executes a computer program to analyze the current voiceprint information and obtain the voice emotion recognition result, the following steps are specifically implemented: verifying the continuity of the current voiceprint information; if the verification is successful, comparing the current voiceprint information with the basic voiceprint information corresponding to the target user to obtain the voice emotion recognition result.

[0122] In one embodiment, when the processor executes the logic of the computer program to perform continuity verification of the current voiceprint information, the following steps are specifically implemented: comparing the current voiceprint information with the historical voiceprint information of the target user to obtain the continuity verification result.

[0123] In one embodiment, when the processor executes the computer program, it further implements the following steps: determining the candidate voiceprint information that is located before and adjacent to the current voiceprint information among a plurality of candidate voiceprint information as the historical voiceprint information of the target user; wherein the plurality of candidate voiceprint information is extracted from the voice records of the target user according to a preset sampling interval.

[0124] In one embodiment, when the processor executes the logic of a computer program to determine the target emotion recognition result based on the voice emotion recognition result and the keyword emotion recognition result, the following steps are specifically implemented: if the voice emotion recognition result or the keyword emotion recognition result indicates an abnormal emotion, then the target emotion recognition result is determined to be an abnormal emotion; if the voice emotion recognition result and the keyword emotion recognition result are inconsistent, then the voice emotion recognition result is determined to be the target emotion recognition result.

[0125] In one embodiment, when the processor executes the computer program, it also performs the following steps: if the target emotion recognition result is an abnormal emotion, it outputs emotion regulation information to the target user.

[0126] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0127] Extract the target user's current voiceprint information and voice content information from the target user's voice recording;

[0128] The current voiceprint information is analyzed to obtain the voice emotion recognition result;

[0129] The speech content information is analyzed to obtain the keyword emotion recognition results;

[0130] Based on the voice emotion recognition results and keyword emotion recognition results, the target emotion recognition result is determined.

[0131] In one embodiment, when the logic of the computer program analyzing the current voiceprint information to obtain the voice emotion recognition result is executed by the processor, the following steps are specifically implemented: the continuity of the current voiceprint information is verified; if the verification is successful, the current voiceprint information is compared with the basic voiceprint information corresponding to the target user to obtain the voice emotion recognition result.

[0132] In one embodiment, when the logic for the computer program to perform continuity verification of the current voiceprint information is executed by the processor, the following steps are specifically implemented: comparing the current voiceprint information with the historical voiceprint information of the target user to obtain the continuity verification result.

[0133] In one embodiment, when the computer program is executed by the processor, it further implements the following steps: determining the candidate voiceprint information that is located before and adjacent to the current voiceprint information among a plurality of candidate voiceprint information as the historical voiceprint information of the target user; wherein the plurality of candidate voiceprint information is extracted from the voice records of the target user according to a preset sampling interval.

[0134] In one embodiment, when the logic of determining the target emotion recognition result based on the voice emotion recognition result and the keyword emotion recognition result is executed by the processor, the following steps are specifically implemented: if the voice emotion recognition result or the keyword emotion recognition result indicates an abnormal emotion, then the target emotion recognition result is determined to be an abnormal emotion; if the voice emotion recognition result and the keyword emotion recognition result are inconsistent, then the voice emotion recognition result is determined to be the target emotion recognition result.

[0135] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: if the target emotion recognition result is an abnormal emotion, then output emotion regulation information to the target user.

[0136] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:

[0137] Extract the target user's current voiceprint information and voice content information from the target user's voice recording;

[0138] The current voiceprint information is analyzed to obtain the voice emotion recognition result;

[0139] The speech content information is analyzed to obtain the keyword emotion recognition results;

[0140] Based on the voice emotion recognition results and keyword emotion recognition results, the target emotion recognition result is determined.

[0141] In one embodiment, when the logic of the computer program analyzing the current voiceprint information to obtain the voice emotion recognition result is executed by the processor, the following steps are specifically implemented: the continuity of the current voiceprint information is verified; if the verification is successful, the current voiceprint information is compared with the basic voiceprint information corresponding to the target user to obtain the voice emotion recognition result.

[0142] In one embodiment, when the logic for the computer program to perform continuity verification of the current voiceprint information is executed by the processor, the following steps are specifically implemented: comparing the current voiceprint information with the historical voiceprint information of the target user to obtain the continuity verification result.

[0143] In one embodiment, when the computer program is executed by the processor, it further implements the following steps: determining the candidate voiceprint information that is located before and adjacent to the current voiceprint information among a plurality of candidate voiceprint information as the historical voiceprint information of the target user; wherein the plurality of candidate voiceprint information is extracted from the voice records of the target user according to a preset sampling interval.

[0144] In one embodiment, when the logic of determining the target emotion recognition result based on the voice emotion recognition result and the keyword emotion recognition result is executed by the processor, the following steps are specifically implemented: if the voice emotion recognition result or the keyword emotion recognition result indicates an abnormal emotion, then the target emotion recognition result is determined to be an abnormal emotion; if the voice emotion recognition result and the keyword emotion recognition result are inconsistent, then the voice emotion recognition result is determined to be the target emotion recognition result.

[0145] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: if the target emotion recognition result is an abnormal emotion, then output emotion regulation information to the target user.

[0146] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0147] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0148] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0149] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An emotion recognition method, characterized in that, The method includes: Extract the target user's current voiceprint information and voice content information from the target user's voice recording; If the similarity between the current voiceprint information and the target user's historical voiceprint information is greater than the similarity threshold, then it is determined that the current voiceprint information and the historical voiceprint information have continuity. The current voiceprint information and the basic voiceprint information are input into the voice emotion recognition model, which then superimposes the current voiceprint information and the basic voiceprint information to obtain a voiceprint composite image. The deviation between the current voiceprint information and the basic voiceprint information in the composite image is calculated to obtain a deviation value. The historical voiceprint information refers to candidate voiceprint information that precedes and is adjacent to the current voiceprint information from multiple candidate voiceprint information, which are extracted from the target user's voice recordings according to a preset sampling interval. The basic voiceprint information represents the user's voiceprint information when their emotions are stable. The voice emotion recognition result is determined based on the deviation value; The voice content information is analyzed to obtain keyword emotion recognition results; Based on the voice emotion recognition results and the keyword emotion recognition results, the target emotion recognition result is determined.

2. The method according to claim 1, characterized in that, The step of determining the voice emotion recognition result based on the deviation value includes: If the deviation value falls within the first numerical range, the voice emotion recognition result is determined to be normal voice emotion; If the deviation value falls into the second numerical range, the voice emotion recognition result is determined to be a suspected abnormal voice emotion. If the deviation value falls into the third numerical range, the voice emotion recognition result is determined to be that the voice emotion is obviously abnormal; where the first numerical range is smaller than the second numerical range, and the second numerical range is smaller than the third numerical range.

3. The method according to claim 1, characterized in that, The number of current voiceprint information is one or more.

4. The method according to claim 1, characterized in that, The step of determining the target emotion recognition result based on the voice emotion recognition result and the keyword emotion recognition result includes: If the voice emotion recognition result or the keyword emotion recognition result indicates an abnormal emotion, then the target emotion recognition result is determined to be an abnormal emotion. If the voice emotion recognition result is inconsistent with the keyword emotion recognition result, then the voice emotion recognition result is determined as the target emotion recognition result.

5. The method according to claim 4, characterized in that, The method further includes: If the target emotion recognition result is an abnormal emotion, then emotion regulation information is output to the target user.

6. An emotion recognition device, characterized in that, The device includes: The extraction module is used to extract the current voiceprint information and voice content information of the target user from the target user's voice recording; A voice comparison module is used to determine that there is continuity between the current voiceprint information and the historical voiceprint information of the target user if the similarity between the current voiceprint information and the historical voiceprint information of the target user is greater than a similarity threshold; input the current voiceprint information and the basic voiceprint information into a voice emotion recognition model, so that the voice emotion recognition model superimposes the current voiceprint information and the basic voiceprint information to obtain a voiceprint composite image, and calculates the deviation between the current voiceprint information and the basic voiceprint information in the voiceprint composite image to obtain a deviation value; wherein, the historical voiceprint information is the candidate voiceprint information that is located before and adjacent to the current voiceprint information among multiple candidate voiceprint information, and the multiple candidate voiceprint information is extracted from the voice recording of the target user according to a preset sampling interval; the basic voiceprint information is the voiceprint information representing the user's emotional stability; and determine the voice emotion recognition result based on the deviation value; The keyword comparison module is used to analyze the voice content information and obtain keyword emotion recognition results; The result determination module is used to determine the target emotion recognition result based on the voice emotion recognition result and the keyword emotion recognition result.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and system for monitoring voice communication of call center

    CN103701999A

  • Customer service control method and device and computer readable storage medium

    CN113099043A