Emotion Recognition Method, Device, Storage Medium, System and Program Product Using Adaptive Two-Stage Task
By constructing an adaptive two-stage task, using emotional image tasks and semantic fluent tasks to obtain user physiological data, and combining the first and second emotion recognition models for emotion recognition, the problem of limited long-term emotion recognition ability in the prior art is solved, and more accurate emotion recognition is achieved.
Patent Information
- Application Number
- CN202411833030.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-12-12
AI Technical Summary
Existing emotional recognition methods are difficult to accurately identify users' long-term emotions, especially their ability to recognize emotions such as happiness, happiness, pride, nostalgia, sadness, guilt, resentment, anxiety, depression, etc. are limited.
By constructing an adaptive two-stage task, firstly, the user's first physiological data is obtained using the first task (such as emotional image task and semantic fluency task), and the first emotion recognition model is used to perform rough screening emotional information processing; then the second task is determined based on the rough screening emotional information, the second physiological data is further obtained, and the second emotion recognition model is used for emotion recognition.
Through this method, users' long-term emotions can be more accurately identified, and the accuracy and reliability of emotions recognition are improved.
Smart Images

Figure CN119293738B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of data processing, and more particularly, to an emotion recognition method, apparatus, storage medium, system, and program product. Background Art
[0002] Existing emotion recognition methods usually focus on capturing users' instantaneous emotions, and lack relatively accurate and reliable recognition means for users' long-term emotions that relatively stably exist in a recent period of time, such as happiness, joy, pride, nostalgia, sadness, guilt, resentment, anxiety, and depression.
[0003] Specifically, although there are solutions in the related art to detect the near-infrared brain imaging data of users in a task state to understand the blood perfusion state of the prefrontal cortex (PFC) of users, and compare it with the standard data of blood perfusion states under several different types of long-term emotions, so as to identify the long-term emotions of users, such solutions can only roughly distinguish the types of long-term emotions, and the recognition ability is limited. Summary of the Invention
[0004] The present disclosure provides an emotion recognition method, apparatus, storage medium, system, and program product for solving at least one of the above problems.
[0005] According to one aspect of the present disclosure, an emotion recognition method is provided. The emotion recognition method includes: outputting first task information to prompt a target user to perform a first task corresponding to the first task information; during the process of the target user performing the first task, acquiring first physiological data of the target user; using a first emotion recognition model to process the first physiological data to obtain rough screening emotion information; determining second task information from multiple candidate task information according to the rough screening emotion information; outputting the second task information to prompt the target user to perform a second task corresponding to the second task information; during the process of the target user performing the second task, acquiring second physiological data of the target user; using a second emotion recognition model to process the second physiological data to obtain emotion recognition information.
[0006] Optionally, the first task includes an emotion image task lasting for a first duration and a semantic fluency task lasting for a second duration, and there is an interval lasting for a preset interval duration between the emotion image task and the semantic fluency task.
[0007] Optionally, the value range of the first duration is 65 seconds to 105 seconds; the value range of the second duration is 190 seconds to 230 seconds; the value range of the preset interval duration is 20 seconds to 30 seconds.
[0008] Optionally, the rough-screened emotion information includes emotion indices of multiple preset emotions, each emotion index being used to represent the likelihood that the user has the corresponding preset emotion. The multiple candidate task information includes general task information and emotion task information corresponding one-to-one to the multiple preset emotions. Among them, determining the second task information from the multiple candidate task information according to the rough-screened emotion information includes: determining, from the emotion indices of the multiple preset emotions, the emotion indices greater than the index threshold, and taking the preset emotions corresponding to the determined emotion indices as target emotions; in the case where the number of target emotions is less than the preset number, determining, from the multiple candidate task information, the emotion task information corresponding to the target emotions as the second task information; in the case where the number of target emotions is greater than or equal to the preset number, determining the general task information as the second task information.
[0009] Optionally, the emotion task corresponding to the emotion task information includes an interview task, and the general task corresponding to the general task information includes a picture-description task.
[0010] Optionally, the second task information lasts for a third duration, and the value range of the third duration is from 100 seconds to 140 seconds.
[0011] Optionally, the second physiological data includes voice data and near-infrared brain imaging data. The voice data is data obtained by recording the content dictated by the target user. The second emotion recognition model includes a natural language processing model, a near-infrared emotion recognition model, and an information fusion model. Among them, using the second emotion recognition model to process the second physiological data to obtain emotion recognition information includes: using the natural language processing model to perform natural language processing on the voice data to identify the emotion of the target user and obtain semantic emotion information; using the near-infrared emotion recognition model to process the near-infrared brain imaging data to identify the emotion of the target user and obtain near-infrared emotion information; using the information fusion model to perform fusion processing on the semantic emotion information and the near-infrared emotion information to obtain the emotion recognition information.
[0012] Optionally, the emotion recognition method further includes: performing a consistency comparison according to the rough-screened emotion information, the semantic emotion information, and the near-infrared emotion information to obtain consistency information, where the consistency information is used to reflect the degree of consistency between the identified different emotion information; in the case where the consistency information indicates that the degree of consistency between the identified different emotion information is lower than the preset degree, outputting a prompt message to prompt that the recognition result is inaccurate.
[0013] Optionally, the emotion recognition method further includes: receiving feedback information of the target user regarding the emotion recognition information; and updating the second emotion recognition model according to the feedback information.
[0014] According to another aspect of the present disclosure, there is provided an emotion recognition device, including: a first output unit configured to output first task information to prompt the target user to perform a first task corresponding to the first task information; a first acquisition unit configured to acquire first physiological data of the target user during the process of the target user performing the first task; a rough screening unit configured to process the first physiological data by using a first emotion recognition model to obtain rough screening emotion information; a determination unit configured to determine second task information from a plurality of candidate task information according to the rough screening emotion information; a second output unit configured to output the second task information to prompt the target user to perform a second task corresponding to the second task information; a second acquisition unit configured to acquire second physiological data of the target user during the process of the target user performing the second task; and an identification unit configured to process the second physiological data by using a second emotion recognition model to obtain emotion recognition information.
[0015] Optionally, the first task includes an emotion image task lasting for a first duration and a semantic fluency task lasting for a second duration, and there is an interval lasting for a preset interval duration between the emotion image task and the semantic fluency task.
[0016] Optionally, the value range of the first duration is from 65 seconds to 105 seconds; the value range of the second duration is from 190 seconds to 230 seconds; and the value range of the preset interval duration is from 20 seconds to 30 seconds.
[0017] Optionally, the rough screening emotion information includes emotion indexes of a plurality of preset emotions, each emotion index is used to represent the possibility that the user has the corresponding preset emotion, the plurality of candidate task information includes general task information and emotion task information corresponding to the plurality of preset emotions one by one, and the determination unit is further configured to: determine, from the emotion indexes of the plurality of preset emotions, the emotion indexes greater than an index threshold, and use the preset emotions corresponding to the determined emotion indexes as target emotions; in the case where the number of the target emotions is less than a preset number, determine, from the plurality of candidate task information, the emotion task information corresponding to the target emotions as the second task information; and in the case where the number of the target emotions is greater than or equal to the preset number, determine the general task information as the second task information.
[0018] Optionally, the emotion task corresponding to the emotion task information includes an interview task, and the general task corresponding to the general task information includes a task of describing a picture in words.
[0019] Optionally, the second task information lasts for a third duration, and the value range of the third duration is from 100 seconds to 140 seconds.
[0020] Optionally, the second physiological data includes voice data and near-infrared brain imaging data. The voice data is data obtained by recording the content dictated by the target user. The second emotion recognition model includes a natural language processing model, a near-infrared emotion recognition model, and an information fusion model. The recognition unit is further configured to: use the natural language processing model to perform natural language processing on the voice data to identify the emotion of the target user and obtain semantic emotion information; use the near-infrared emotion recognition model to process the near-infrared brain imaging data to identify the emotion of the target user and obtain near-infrared emotion information; use the information fusion model to perform fusion processing on the semantic emotion information and the near-infrared emotion information to obtain the emotion recognition information.
[0021] Optionally, the emotion recognition device further includes: a comparison unit configured to perform a consistency comparison based on the rough screening emotion information, the semantic emotion information, and the near-infrared emotion information to obtain consistency information, where the consistency information is used to reflect the degree of consistency between the recognized different emotion information; a prompt unit configured to output a prompt message to prompt that the recognition result is inaccurate when the consistency information indicates that the degree of consistency between the recognized different emotion information is lower than a preset degree.
[0022] Optionally, the emotion recognition device further includes: a receiving unit configured to receive feedback information of the target user regarding the emotion recognition information; an updating unit configured to update the second emotion recognition model according to the feedback information.
[0023] According to another aspect of the present disclosure, there is provided a computer-readable storage medium storing instructions, where when the instructions are run by at least one computing device, the at least one computing device is caused to execute the emotion recognition method as described above.
[0024] According to another aspect of the present disclosure, there is provided a system including at least one computing device and at least one storage device storing instructions, where when the instructions are run by the at least one computing device, the at least one computing device is caused to execute the emotion recognition method as described above.
[0025] According to another aspect of the present disclosure, there is provided a computer program product including instructions, where when the instructions are run by at least one computing device, the at least one computing device is caused to execute the emotion recognition method as described above.
[0026] The emotion recognition method, device, storage medium, system and program product according to the exemplary embodiments of the present disclosure break the conventional idea of improving the ability to mine data patterns. By constructing an adaptive two-stage task, the first task is used to roughly identify the user's emotion to obtain rough screening emotion information, and then the corresponding second task is issued according to the rough screening emotion information to stimulate the user to generate richer physiological data, i.e., the second physiological data, in the corresponding emotion, which can specifically increase the effective basic data required for emotion recognition, contribute to realizing more detailed emotion recognition from the perspective of data-driven, and improve the ability to recognize the long-term emotion of the target user. Specifically, since the long-term emotion is the emotion that relatively stably exists in the user in a recent period of time, it will affect the user's thinking tendency and the instantaneous emotion reaction when facing specific things. For example, when facing the same emotion image, users in two different long-term emotions of anxiety and happiness are very likely to have completely different instantaneous emotions. Based on this, by inviting the target user to perform a specific task, the user's instantaneous emotion can be stimulated, and then the long-term emotion can be recognized by identifying the stable emotion tendency in the instantaneous emotion. By adopting the adaptive task arrangement scheme of the present disclosure, the stable emotion tendency of the target user can be deeply stimulated and explored, and the accurate recognition of the long-term emotion of the target user can be realized.
[0027] Some other aspects and / or advantages of the general concept of the present disclosure will be partially described in the following description, and some will be clear from the description, or can be learned through the implementation of the general concept of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] By combining the drawings, these and / or other aspects and advantages of the present disclosure will become clear and more easily understood from the following description of the embodiments.
[0029] Figure 1 is a flowchart showing the emotion recognition method according to the exemplary embodiments of the present disclosure.
[0030] Figure 2 is a schematic diagram showing the task release logic according to the specific embodiments of the present disclosure.
[0031] Figure 3 is a block diagram showing the emotion recognition device according to the exemplary embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] The following description with reference to the accompanying drawings is provided to assist in a comprehensive understanding of embodiments of the present invention defined by the claims and their equivalents. Various specific details are included to aid understanding, but these are considered to be merely exemplary. Thus, those of ordinary skill in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. In addition, descriptions of well-known functions and structures are omitted for clarity and conciseness.
[0033] It should be noted here that "at least one of several items" as used in this disclosure all represents the inclusion of three parallel cases: "any one of the several items", "a combination of any multiple of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel cases: (1) including A; (2) including B; (3) including A and B. Another example, "performing at least one of step one and step two" means the following three parallel cases: (1) performing step one; (2) performing step two; (3) performing step one and step two.
[0034] The following refers to Figures 1 to 3 A method, apparatus, storage medium, system, and program product for emotion recognition according to an exemplary embodiment of the present disclosure will be described in detail.
[0035] Figure 1 is a flowchart showing a method for emotion recognition according to an exemplary embodiment of the present disclosure. The method for emotion recognition according to an exemplary embodiment of the present disclosure can be implemented in a computing device with sufficient computing power.
[0036] Referring to Figure 1 , in step S101, first task information is output to prompt a target user to perform a first task corresponding to the first task information.
[0037] The target user is the object of emotion recognition, that is, the user whose emotion needs to be recognized. When outputting the first task information, as an example, the first task information can be output in the form of text, video, or voice through devices such as a display screen and a loudspeaker to implement the release of the first task, so that the target user can perform the first task according to the prompt.
[0038] Optionally, the first task includes an emotion image task lasting for a first duration and a semantic fluency task lasting for a second duration, with an intermittent period of a preset intermittent duration between the emotion image task and the semantic fluency task. By sequentially releasing different types of tasks, on the one hand, different tasks can be used to provide different stimuli to the target user, facilitating the acquisition of more information. On the other hand, it can enable the target user to gradually enter the state, which helps to obtain effective physiological data reflecting the target user's emotions in the subsequent step S102. Specifically, through research and analysis, it is found that the effect of activating emotions in the emotion image task is relatively weaker compared to the semantic fluency task. Based on this, by releasing the emotion image task first, it can help the target user establish initial adaptability, enable the target user to enter a focused state, and establish a reference baseline for physiological data. After the emotion image task ends, after a short break, by releasing the semantic fluency task, the target user can be fully activated to obtain physiological data that effectively reflects the target user's emotions. In addition, the emotion image task can also capture the immediate emotional changes of the target user through visual stimuli, providing basic data for subsequent emotional state analysis. The semantic fluency task can also test the target user's semantic generation ability, language expression fluency, and emotional response when facing task pressure.
[0039] Specifically, the goal of the Emotion Picture Task (EPT) is to trigger the target user's emotional response by having them observe emotion-related images. Before the task officially starts, a brief task explanation can be given first. After the explanation ends, a number of emotion images randomly selected from the image library are sequentially shown to the target user. The content of these images is designed to evoke different emotional responses, and the display order is also random. In other words, the emotion image task includes two stages: task explanation and emotion image display, each with a specific duration. In the emotion image display stage, a number of randomly selected emotion images are sequentially shown. Correspondingly, the information in the first task information corresponding to the emotion image task includes emotion task explanation information, a number of randomly selected emotion images, task explanation duration, and single-image display duration. When outputting this part of the information, the emotion task explanation information is output within the task explanation duration first, and then the number of randomly selected emotion images is output one by one, and the duration of each emotion image output is the single-image display duration.
[0040] The goal of the Semantic Fluency Task (SFT) is to evaluate the language fluency of target users by asking them to name as many words as possible related to specific categories (e.g., quadrupeds, vegetables, stationery) within a limited time. In addition to the semantic fluency task, the Verbal Fluency Task (VFT) also includes the Phonetic Fluency Task (FFT), whose stimulus pattern focuses on naming words based on phonetic cues, such as "Please list words starting with the letter 'T'" or "Please list words starting with 'white'". Through research and analysis, it is found that the phonetic fluency task sometimes cannot effectively distinguish different emotions. Based on this, choosing a more effective semantic fluency task can improve the reliability of the initial emotion screening. Before the task officially starts, a brief task explanation can also be given first. After the explanation, several subtasks randomly selected from the task library are sequentially issued to the target user. Each subtask is completed within the same limited time duration, and there is a certain rest time after each subtask, so as to repeatedly activate the emotional response of the target user. In other words, the semantic fluency task includes two stages: task explanation and category display, each with a specific duration. The category display stage further includes several sub-stages, and each sub-stage includes the display of the category name and a break between groups. In particular, the last sub-stage may not include a break between groups. Correspondingly, the information in the first task information corresponding to the semantic fluency task includes semantic task explanation information, several randomly selected category names, task explanation duration, single category name display duration, and break between groups duration. When outputting this part of the information, first output the semantic task explanation information within the task explanation duration, then output a randomly selected category name, stop after the single category name display duration, and after the break between groups duration, continue to output the next randomly selected category name, and so on until all randomly selected category names are output.
[0041] Optionally, the value range of the first duration is 65 seconds to 105 seconds; the value range of the second duration is 190 seconds to 230 seconds. By limiting the value range of the durations of different stages in the first task, the task duration can be reasonably arranged. The lower limit value of the task duration can ensure that the task lasts long enough to obtain sufficient information, and the upper limit value of the task duration can prevent the target user from developing obvious negative emotions due to the overly long task and interfering with the emotional response of the target user, which helps to ensure the accuracy of the emotion recognition result.
[0042] Optionally, the value range of the preset intermittent duration is from 20 seconds to 30 seconds. The lower limit value can ensure that the target user gets sufficient rest and reduces the fatigue of the target user, and the upper limit value can reduce the risk of distracting the user's attention due to too long a rest time. By reserving a reasonable rest duration, the effectiveness of the physiological data collected during the task can be improved, and the emotion recognition effect can be enhanced.
[0043] In step S102, during the process of the target user performing the first task, the first physiological data of the target user is obtained.
[0044] Specifically, the first physiological data can be detected by a biosignal sensor. Optionally, the first physiological data can be the physiological data collected during the execution of the entire first task, or the physiological data collected during the execution of a part of the first task (such as a semantic fluency task). The present disclosure does not limit this. Through software control, the data collection of the biosignal sensor is synchronized with the output of the corresponding task information, and a timestamp can be added to the collected first biological data, so as to correspond the first physiological data to the corresponding task stage, which helps to achieve more accurate emotion recognition. Of course, it is also possible to start collecting physiological data before the corresponding task starts, stop collecting physiological data for a period of time after the corresponding task is completed, and then intercept the physiological data during the execution of the corresponding task as the first physiological data. In other words, here it is only limited to obtaining the physiological data of the target user as the first physiological data during the process of the target user performing the first task, but the method of collecting and obtaining the physiological data is not limited.
[0045] As an example, the first physiological data includes near-infrared brain imaging data. The near-infrared brain imaging data is obtained through functional near-infrared spectroscopy (fNIRS), which is a spectral imaging technique. By continuously irradiating the human brain scalp with near-infrared light with a wavelength of 650nm - 850nm, the near-infrared light can penetrate the scalp and skull and enter the cerebral cortex to a depth of 2cm - 3cm, forming a propagation arc and being detected by another detector on the scalp. During the propagation of near-infrared light in human tissues, the absorption of photons by tissues follows certain rules, especially for oxygenated hemoglobin and deoxygenated hemoglobin in the capillaries in the cerebral cortex. Due to this particularity, the change in the concentration of oxygenated hemoglobin and deoxygenated hemoglobin in the irradiated area can be inferred from the absorption rate obtained from the detector. When the human brain is working, nerve cells continuously discharge and consume a large amount of energy. And this energy is provided by the oxygen carried by the capillaries around the nerve cells. Therefore, when the brain is activated, the supply of oxygenated hemoglobin increases, accelerating oxygen supply and reducing the concentration of deoxygenated hemoglobin. After activation, the energy is converted, the concentration of oxygenated hemoglobin decreases, and the concentration of deoxygenated hemoglobin increases. The cerebral cortex is again in a state of oxygen demand, waiting for the next oxygen supply. This process repeats, which is the neurovascular coupling mechanism in neuroscience.
[0046] The prefrontal lobe of the human brain is responsible for various high-level emotional activities and regulation of a person. Current research shows that emotions originate from the limbic system deep in the brain, and the regulation of emotions is a very complex process, which is organically regulated by various parts of the prefrontal lobe. Therefore, studying the blood oxygen dynamics feedback mechanism of the prefrontal lobe can reveal an individual's response and regulation ability to different emotions. By using near-infrared brain imaging technology on the prefrontal lobe, the cerebral cortex can be activated through task states (including emotional stimulation tasks, language fluency tasks, resting state tasks, etc., where the resting state task refers to the state where the subject remains still and relaxed without performing specific cognitive tasks). By observing the changes in oxygenated hemoglobin and deoxygenated hemoglobin, an individual's emotional characteristics can be obtained, thereby establishing an individual emotional regulation model.
[0047] It should be understood that the near-infrared brain imaging data needs to be collected in conjunction with corresponding equipment, such as including a prefrontal lobe headgear and a detector for collecting data. It should also be understood that to ensure sufficient acquisition of near-infrared brain imaging data during the process of the target user performing the first task, data collection can start before the target user begins to perform the first task and stop for a period of time after the target user completes the first task, and then the data corresponding to the execution period of the first task is intercepted. In other words, here it is only limited to obtaining the first physiological data during the process of the target user performing the first task, but there is no restriction on how the data is collected and obtained.
[0048] In step S103, the first emotion recognition model is used to process the first physiological data to obtain the roughly screened emotion information.
[0049] As an example, the first emotion recognition model can be trained by machine learning methods. The training samples used for training may include the sample physiological data of the sample users and the sample emotion labels. Each sample emotion label is used to mark a long-term emotion state, which can be a specific emotion type or a set of multiple emotion types. By configuring appropriate sample emotion labels, the first emotion recognition model can be made to have different emotion rough screening capabilities. Correspondingly, the roughly screened emotion information can be a long-term emotion state in which the target user is most likely to be, or the possibility that the target user is in each long-term emotion state. In addition, before using the first emotion recognition model to process the first physiological data, preprocessing can be performed on the first physiological data, such as signal baseline correction, sliding window low-pass filtering, and artifact signal processing, to reduce noise interference and improve data quality.
[0050] In step S104, according to the roughly screened emotion information, the second task information is determined from multiple pieces of candidate task information.
[0051] This step is used to determine, from multiple pieces of candidate task information, the one that best matches the roughly screened emotion information identified in step S103 as the second task information, so as to achieve adaptive task release and targeted emotion stimulation.
[0052] Optionally, the rough screening emotion information includes emotion indices of multiple preset emotions. The preset emotions can be long-term emotion types, such as including multiple long-term emotion types like happiness, joy, pride, nostalgia, sadness, guilt, resentment, anxiety, depression, etc., to achieve rich emotion recognition. For another example, it can only include negative long-term emotion types like guilt, resentment, anxiety, depression, etc., to achieve the monitoring and early warning of negative long-term emotions. Each emotion index is used to represent the possibility that the user has the corresponding preset emotion. The multiple candidate task information includes general task information and emotion task information corresponding one by one to the multiple preset emotions. Step S104 includes: determining, from the emotion indices of the multiple preset emotions, the emotion indices greater than the index threshold, and taking the preset emotions corresponding to the determined emotion indices as target emotions; in the case where the number of target emotions is less than the preset number, determining, from the multiple candidate task information, the emotion task information corresponding to the target emotions as the second task information; in the case where the number of target emotions is greater than or equal to the preset number, determining the general task information as the second task information. By configuring the rough screening emotion information as the emotion indices of multiple specific preset emotions, and configuring the multiple candidate tasks as general tasks and emotion tasks corresponding one by one to these preset emotions, and then selecting the matching task as the second task according to the identified emotion indices of each preset emotion, the adaptive arrangement of the second task can be achieved. Specifically, by configuring the index threshold as a reference standard, the intensity of each preset emotion identified during rough screening can be reliably judged, so as to find out the target emotions with higher intensity. By configuring the preset number, the data processing volume during the second-stage task arrangement and emotion precise recognition can be controlled, that is, when the number of target emotions with higher intensity is small, corresponding second tasks are arranged for each target emotion to achieve in-depth excitation and exploration of each target emotion, and when the number of target emotions with higher intensity is large, the general task is directly arranged to do in-depth excitation and exploration of the comprehensive emotion on the basis of the first task, so as to appropriately improve the recognition accuracy of long-term emotions. As an example, the preset number is 1 or 2, and the present disclosure does not limit this. The value of the index threshold can be reasonably set according to actual needs, and the present disclosure also does not limit this.
[0053] As an example, in the case where the preset emotions simultaneously include long-term emotion types of both positive and negative, general tasks can be configured for positive emotions and negative emotions respectively. If the number of target emotions is greater than or equal to the preset number, first determine the proportion of positive emotions and / or negative emotions in the target emotions. Then, when the proportion of positive emotions is greater than or equal to the preset proportion (such as 50%, or 80%), select the general task corresponding to the positive emotions. When the proportion of negative emotions is greater than or equal to the preset proportion, select the general task corresponding to the negative emotions. Further, if the preset proportion is set relatively high, there may also be a situation where the proportions of both positive emotions and negative emotions in the target emotions are less than the preset proportion. For this reason, a neutral emotion can be further included in the preset emotions, and a general task applicable to the neutral emotion can be configured accordingly. And when the number of target emotions is greater than or equal to the preset number, and the proportions of both positive emotions and negative emotions in the target emotions are less than the preset proportion, select the general task applicable to the neutral emotion.
[0054] Optionally, the emotion task corresponding to the emotion task information includes an interview task (Interview Task, abbreviated as INT). Corresponding interview tasks can be arranged for different preset emotions, which can effectively stimulate the corresponding preset emotions by arranging interview content as needed, contribute to the accurate quantification of the preset emotions, and improve the emotion recognition accuracy. The general task corresponding to the general task information includes a storytelling picture task (Storytelling Picture Task, abbreviated as SPT), which can be applicable to multiple preset emotions and further improve the emotion recognition accuracy on the basis of the first task.
[0055] Optionally, the second task information lasts for a third duration, and the value range of the third duration is from 100 seconds to 140 seconds. The lower limit value can ensure that the task duration is long enough to obtain sufficient information, and the upper limit value can prevent the target user from generating obvious negative emotions due to the overly long task and interfering with the emotional response of the target user, which helps to ensure the accuracy of the emotion recognition result.
[0056] In step S105, output the second task information to prompt the target user to perform the second task corresponding to the second task information.
[0057] The output method of the second task information can refer to the output method of the first task information. For example, for the interview task, the interview questions can be displayed on the display screen, or the interview question audio can be played through the loudspeaker. Additionally, a sound pickup device such as a microphone can be further configured to record the audio of the target user's answers to provide more reference data for emotion recognition. Of course, the answer audio may not be recorded, and emotion recognition can be achieved only relying on the second physiological data of the target user obtained in the subsequent step S106. The present disclosure does not limit this.
[0058] In step S106, during the process of the target user performing the second task, second physiological data of the target user is acquired.
[0059] For the content and acquisition method of the second physiological data, reference can be made to the content and acquisition method of the first physiological data, which will not be repeated here.
[0060] In step S107, the second emotion recognition model is used to process the second physiological data to obtain emotion recognition information.
[0061] Optionally, the second physiological data includes voice data and near-infrared brain imaging data. The voice data is data obtained by recording the content spoken by the target user. For example, it is the recorded data of the content stated by the target user when performing the interview task and the picture description task described in the previous example, or the recorded data of the content stated by the target user when performing other types of second tasks. The present disclosure places no restrictions on this. Correspondingly, the second emotion recognition model includes a natural language processing model, a near-infrared emotion recognition model, and an information fusion model. Step S106 includes: using the natural language processing model to perform natural language processing on the voice data to identify the emotion of the target user and obtain semantic emotion information; using the near-infrared emotion recognition model to process the near-infrared brain imaging data to identify the emotion of the target user and obtain near-infrared emotion information; using the information fusion model to perform fusion processing on the semantic emotion information and the near-infrared emotion information to obtain emotion recognition information. By first using the natural language processing model and the near-infrared emotion recognition model respectively to process the voice data and the near-infrared brain imaging data collected during the execution of the second task to obtain semantic emotion information and near-infrared emotion information, and then using the information fusion model to fuse the two emotion recognition results, accurate emotion recognition is achieved.
[0062] As an example, when the natural language processing model processes the voice data, it can first perform speech recognition on the voice data to convert the voice data into text data, and then perform word segmentation and part-of-speech tagging on the text data, and determine the semantic emotion information based on the part-of-speech tagging results. When determining the semantic emotion information, for example, the method of using an emotion dictionary can be used. A series of words and their corresponding emotional polarities (positive or negative) are pre-recorded in the emotion dictionary. By calculating the ratio of positive words and negative words in the text, the emotional tendency of the text is judged as the semantic emotion information. Another example is that machine learning algorithms can be used, such as including but not limited to Support Vector Machine (SVM), Naive Bayes, decision trees, etc. These algorithms learn emotional features by training a large amount of labeled data and are used for the emotional classification of new texts. The classification results obtained are used as the semantic emotion information.
[0063] It should be understood that the first emotion recognition model is used to implement rough emotion screening, and the data types of the first physiological data it processes can be relatively few. For example, it only processes near-infrared brain imaging data. In order to improve the recognition accuracy, the second physiological data processed by the second emotion recognition model can include more data types, such as the speech data and near-infrared brain imaging data introduced here. In order to implement the processing of near-infrared brain imaging data, the type of the near-infrared recognition model in the second emotion recognition model can be the same as or different from the type of the first emotion recognition model. In the case of the same type, the parameters of the two can also be the same (that is, the near-infrared recognition models in the first emotion recognition model and the second emotion recognition model are the same model) or different. Of course, in the case of different types, the parameters of the two must also be different. For the case where the near-infrared recognition model in the second emotion recognition model and the first emotion recognition model are different models, different emotion recognition models can be further configured according to different candidate task information, so as to use the emotion recognition model corresponding to the second task information as the second emotion recognition model.
[0064] Regarding the recognizable emotion labels, that is, the content of the emotion labels that the recognized emotion information can represent. As an example, the content of the emotion labels that the semantic emotion information and the near-infrared emotion information can each represent can be the same to facilitate the information fusion model to fuse the semantic emotion information and the near-infrared emotion information. Of course, they can also be different, and the present disclosure does not limit this. In addition, the content of the emotion labels that the rough-screened emotion information, the semantic emotion information, and the near-infrared emotion information can represent can also be the same or different, and the present disclosure also does not limit this. For the specific emotion labels, reference can be made to the introduction of steps S103 and S104 above, and details will not be elaborated here.
[0065] Optionally, according to different detection requirements, different label information can be configured for the training samples when training the second emotion recognition model, such as emotion type labels, or emotion intensity labels. The emotion intensity labels can correspond to several levels, such as strong, general, weak, or can also be continuous quantities, such as reflected in the form of scores. The present disclosure does not limit this.
[0066] Optionally, the content of the emotion tags that can be represented by the rough-screened emotion information, semantic emotion information, and near-infrared emotion information may be the same. The emotion recognition method according to an exemplary embodiment of the present disclosure may further include: performing a consistency comparison based on the rough-screened emotion information, semantic emotion information, and near-infrared emotion information to obtain consistency information, where the consistency information is used to reflect the degree of consistency between the recognized different emotion information; in the case where the consistency information indicates that the degree of consistency between the recognized different emotion information is lower than a preset degree, output a prompt message to prompt that the recognition result is inaccurate. By further adding the processing step of consistency comparison, it helps to automatically identify abnormal situations where the target user has unstable emotions, such as including but not limited to deliberately concealing emotions, answering off-topic, having an unstable mental state, and resisting answering, and output corresponding prompt messages, providing an additional layer of guarantee for the recognition result. It should be understood that the consistency information is used to reflect the degree of consistency between the recognized different emotion information, which can be achieved by the value of an index, and the value of this index may be positively correlated with the degree of consistency or negatively correlated with the degree of consistency, which are all implementation manners of the present disclosure and fall within the protection scope of the present disclosure.
[0067] As an example, in combination with the rough-screened emotion information, semantic emotion information, and near-infrared emotion information, the cooperation degree (Cooperation Index, abbreviated as CI) can be calculated through the following exclusive OR weighted formula as the consistency information:
[0068] .
[0069] In this formula, N is the total number of recognizable emotion tags, represents an emotion tag. , , respectively represent the recognition results of the first emotion recognition model, the second emotion recognition model, and the near-infrared recognition model in the second emotion recognition model for this emotion tag. All three are boolean values, 0 indicating that the recognition result is not , 1 indicating that the recognition result is . is the operator representing the exclusive OR operation (Exclusive OR, abbreviated as XOR), which is a binary operation. Its basic rule is: if the two quantities being compared are the same, the operation result is 0; if the two quantities being compared are different, the operation result is 1. Therefore, for each emotion tag, the value of the consistency calculation result may be 0, 1, 2, 3, so as to be able to reflect whether the recognition results of the three models for each emotion tag are consistent, and the smaller the value of the calculation result, the higher the degree of consistency. Then, add up the consistency calculation results for each emotion tag to obtain the cooperation degree CI, which can represent the consistency of the overall recognition result. It can be understood that the cooperation degree CI ranges from [0, 3N], and CI the larger it is, the more unstable the emotion of the target user is, and at the same time, it proves that there is a problem with the detection credibility. For example, in N the case of r = 3, if CI > 7, it represents an untrustworthy emotion recognition result or an extremely unstable result; if CI < 3, it represents a stable emotion and a trustworthy recognition result. And if necessary, in CI the case of a very large r, manual intervention can be carried out to judge the specific reason, which belongs to the specific application of the prompt information of the output, and the present disclosure does not limit this. It should be understood that the exclusive-or weighting formula given here is only an example, and other specific forms of calculation formulas can also be adopted based on the exclusive-or operation, as long as the meaning of the consistency information is satisfied, and they are not listed one by one here.
[0070] In addition to Figure 1 the steps shown, optionally, the emotion recognition method according to the exemplary embodiment of the present disclosure may further include: receiving feedback information of the target user for the emotion recognition information; updating the second emotion recognition model according to the feedback information. The feedback information of the target user for the emotion recognition information can represent the evaluation of the target user on the accuracy of the emotion recognition information. By using the feedback information to update the second emotion recognition model, the model can be continuously optimized as the model is used, which helps to continuously improve the accuracy of emotion recognition and the emotion recognition ability during the use process.
[0071] As an example, the feedback information may include simple evaluations of accurate and inaccurate, and may further include the emotion information self-evaluated by the target user. The content items included in the self-evaluated emotion information may be the same as the content items of the emotion recognition information. At this time, the second physiological data and the corresponding self-evaluated emotion information can be used to form an update sample for updating the second emotion recognition model.
[0072] Generally speaking, to achieve emotion recognition, an emotion recognition system can be constructed, which includes a task information output device (such as including a display screen), a biosignal sensor (such as including a prefrontal headgear and a data acquisition instrument), and a computing device that executes the emotion recognition method according to the exemplary embodiment of the present disclosure.
[0073] Next, in combination with Figure 2 , the emotion recognition method of a specific embodiment of the present disclosure will be introduced.
[0074] In this specific embodiment, a brain activation induction task paradigm for near-infrared brain imaging is designed, with reference to Figure 2, the first stage of this task paradigm (i.e., the first task) includes: a 75-second (first duration) emotional image task, a 25-second (preset interval duration) rest, and a 215-second (second duration) semantic fluency task. The emotional image task is the initial adaptation task with the lowest activation of emotional effects, which can enable the target user to enter a focused state and establish a reference baseline for blood oxygen signals. The role of the semantic fluency task is to repeatedly activate the prefrontal lobe through four subtasks so as to obtain corresponding emotional indices subsequently. After the semantic fluency task, the system will analyze the blood oxygen dynamics characteristics of the semantic fluency task through a neural network algorithm (i.e., the first emotion recognition model) to achieve preliminary emotion recognition, obtain the emotion indices of multiple preset emotions, including anxiety emotion index, depression emotion index, and obsessive-compulsive emotion index, as rough screening emotion information, and return it to the system.
[0075] After another 25-second rest, a 120-second (third duration) second task determined based on the emotion indices will be given according to the result of emotion index recognition as an extended task in the second stage. The second task is a task determined from interview task INT-A, interview task INT-B, interview task INT-C, and picture description task, which can induce the activation of the prefrontal cortex of the target user's brain through a continuous task. If the anxiety emotion index is greater than the index threshold, interview task INT-A for anxiety emotion will be pushed; if the depression emotion index is greater than the index threshold, interview task INT-B for depression emotion will be pushed; if the obsessive-compulsive emotion index is greater than the index threshold, interview task INT-C for obsessive-compulsive emotion will be pushed; if all three types of emotion indices exceed the index threshold, the picture description task will be pushed.
[0076] Specifically, the goal of the emotional image task is to trigger the target user's emotional reaction by having the target user observe emotion-related images. Before the task starts, there is a 15-second task explanation. After the explanation ends, the target user will see three groups of a total of 6 emotional images (randomly displayed from 24 image libraries divided into three groups with similar emotions in each group, and each image is displayed for 10 seconds). The content of these images is designed to stimulate different emotional reactions, and the display order is random. The goal of this part is to capture the immediate emotional changes of the target user through visual stimuli and provide basic data for subsequent emotional state analysis.
[0077] The goal of the semantic fluency task is to evaluate the language fluency of the target users by asking them to say as many words as possible related to specific categories (e.g., quadrupeds, vegetables, stationery) within a limited time. There is a 15-second task introduction at the beginning. There are a total of 4 sub-tasks (randomly assigned from 16 tasks), each sub-task is 35 seconds, and there is a 15-second break after each sub-task. The goal of this part is to test the semantic generation ability and language expression fluency of the target users, as well as their emotional responses when facing task pressure, and to provide data for the extraction of semantic features.
[0078] The picture description task requires the target users to verbally describe the content of the images. The task introduction is 15 seconds. The discussion of each picture (negative, neutral, positive) includes a 5-second preparation time and a 20-second description time. This task combines visual and language information and aims to analyze the emotional and language response patterns of the target users under specific visual stimuli.
[0079] During the execution of the task, the region of interest (ROI) for near-infrared brain imaging data acquisition focuses on the prefrontal cortex. The functions of the prefrontal cortex are mainly related to cognition, emotion, and behavior management, etc., especially in processing various recently acquired information, and its function is called working memory. Basic research shows that under normal circumstances, the neuronal activity in the brain region will increase the oxygen consumption in this region, and at the same time, the cerebral blood volume and blood flow will also increase, carrying a large amount of oxyhemoglobin. This leads to a significant increase in the local oxyhemoglobin concentration, while the concentration of deoxyhemoglobin fluctuates accordingly. The total hemoglobin concentration and blood oxygen saturation also increase, making this brain region in a state of high oxygen concentration blood perfusion. The absorption and scattering rates of oxyhemoglobin and deoxyhemoglobin in human tissues for infrared light in the wavelength range of 700 - 900 nm are significantly different and have an intersection. Therefore, the relative concentration changes of oxyhemoglobin and deoxyhemoglobin in the capillaries of the prefrontal cortex can be indirectly calculated through near-infrared light imaging technology and mathematical methods, so as to obtain relevant data on the blood perfusion state of the prefrontal cortex.
[0080] In order to obtain near-infrared brain imaging data at the same time series as that induced by the brain activation task, the near-infrared data acquisition module is controlled and started by the task paradigm module of the software. The time stamp contained in the acquired near-infrared brain imaging data is equal to the time stamp of the task paradigm operation. At the same time, under the action of the time stamp, the near-infrared data in the brain activation state is divided into three segments: the emotional image task for 75 seconds, the semantic fluency task for 215 seconds, and the second task for 75 seconds.
[0081] To retain the metrics highly correlated with the prefrontal working memory network, it is necessary to preprocess the near-infrared brain imaging data collected from the prefrontal region. Among them, there are three main steps: signal baseline correction, sliding window low-pass filtering, and artifact signal processing.
[0082] When performing signal baseline correction, the cut-off frequency is usually set relatively low (e.g., 0.01 Hz) to ensure that all frequency variations related to cognitive tasks are retained while filtering out slower physiological changes and other drifts.
[0083] Suppose there is a time series signal x(t), and a high-pass filter H1(f) can be used to process this signal, where f represents frequency. A simple implementation of the high-pass filter is the first-order Butterworth filter, and its transfer function can be expressed as:
[0084] 。
[0085] where f c is the cut-off frequency. After applying this high-pass filter to the signal x(t), its output y(t) can be obtained through Fourier transform , filtering, and inverse Fourier transform :
[0086] 。
[0087] Sliding window low-pass filtering is an effective signal processing technique that can be used to smooth the noise in near-infrared brain imaging data, thereby restoring the true fluctuations of the blood oxygen signal. This method uses the Butterworth filter of the time series and processes the signal through a progressive algorithm to achieve the low-pass filtering effect. The Butterworth filter is named for its flat amplitude response and can effectively remove high-frequency noise without introducing too much signal distortion. For the low-pass filter, its transfer function H2(f) is usually defined as:
[0088] 。
[0089] where n represents the length of the sliding window.
[0090] Regarding artifact signal processing, artifacts cannot be deleted when processing near-infrared brain imaging data in the task state. The ASR (Artifact Subspace Reconstruction) algorithm is a subspace separation-based technique that can extract artifact components from complex signals. ASR learns the statistical characteristics of clean calibration data and compares these statistics with new data that may contain artifacts. The ASR algorithm consists of two parts: calibration and processing. During the calibration process, the data is filtered, and then the covariance matrix is estimated using the geometric median of the subsequent samples of the input data segment.
[0091] Figure 3 is a block diagram showing an emotion recognition device according to an exemplary embodiment of the present disclosure.
[0092] Referring to Figure 3 , the emotion recognition device 300 includes a first output unit 301, a first acquisition unit 302, a rough screening unit 303, a determination unit 304, a second output unit 305, a second acquisition unit 306, and an identification unit 307.
[0093] The first output unit 301 is configured to output first task information to prompt the target user to perform a first task corresponding to the first task information.
[0094] Optionally, the first task includes an emotion image task lasting for a first duration and a semantic fluency task lasting for a second duration, and there is an interval lasting for a preset interval duration between the emotion image task and the semantic fluency task.
[0095] Optionally, the value range of the first duration is 65 seconds to 105 seconds; the value range of the second duration is 190 seconds to 230 seconds; the value range of the preset interval duration is 20 seconds to 30 seconds.
[0096] The first acquisition unit 302 is configured to acquire first physiological data of the target user during the execution of the first task by the target user.
[0097] The rough screening unit 303 is configured to process the first physiological data using a first emotion recognition model to obtain rough screening emotion information.
[0098] The determination unit 304 is configured to determine second task information from multiple candidate task information according to the rough screening emotion information.
[0099] Optionally, the rough-screened emotion information includes emotion indices of multiple preset emotions, where each emotion index is used to represent the likelihood that the user has the corresponding preset emotion. The multiple candidate task information includes general task information and emotion task information corresponding one-to-one to the multiple preset emotions. The determination unit is further configured to: determine, from the emotion indices of the multiple preset emotions, the emotion indices greater than the index threshold, and use the preset emotions corresponding to the determined emotion indices as target emotions; in a case where the number of target emotions is less than the preset number, determine, from the multiple candidate task information, the emotion task information corresponding to the target emotions as the second task information; and in a case where the number of target emotions is greater than or equal to the preset number, determine the general task information as the second task information.
[0100] Optionally, the emotion task corresponding to the emotion task information includes an interview task, and the general task corresponding to the general task information includes a picture description task.
[0101] Optionally, the second task information lasts for a third duration, and the value range of the third duration is 100 seconds to 140 seconds.
[0102] The second output unit 305 is configured to output the second task information to prompt the target user to perform the second task corresponding to the second task information.
[0103] The second acquisition unit 306 is configured to acquire second physiological data of the target user during the process of the target user performing the second task.
[0104] The recognition unit 307 is configured to process the second physiological data using a second emotion recognition model to obtain emotion recognition information.
[0105] Optionally, the second physiological data includes voice data and near-infrared brain imaging data. The voice data is data obtained by recording the content spoken by the target user. The second emotion recognition model includes a natural language processing model, a near-infrared emotion recognition model, and an information fusion model. The recognition unit 307 is further configured to: use the natural language processing model to perform natural language processing on the voice data to recognize the emotion of the target user and obtain semantic emotion information; use the near-infrared emotion recognition model to process the near-infrared brain imaging data to recognize the emotion of the target user and obtain near-infrared emotion information; and use the information fusion model to perform fusion processing on the semantic emotion information and the near-infrared emotion information to obtain emotion recognition information.
[0106] Optionally, the emotion recognition device 300 further includes a comparison unit (not shown in the figure) and a prompt unit (not shown in the figure). The comparison unit is configured to perform a consistency comparison based on the roughly screened emotion information, semantic emotion information, and near-infrared emotion information to obtain consistency information, where the consistency information is used to reflect the degree of consistency between the recognized different emotion information; the prompt unit is configured to output a prompt message to prompt that the recognition result is inaccurate when the consistency information indicates that the degree of consistency between the recognized different emotion information is lower than a preset degree.
[0107] Optionally, the emotion recognition device 300 further includes a receiving unit (not shown in the figure) and an updating unit (not shown in the figure). The receiving unit is configured to receive feedback information of a target user regarding the emotion recognition information; the updating unit is configured to update the second emotion recognition model according to the feedback information.
[0108] The above has been described with reference to Figures 1 to 3 the emotion recognition method and device according to the exemplary embodiments of the present disclosure.
[0109] Figure 3 Each unit in the emotion recognition device shown can be configured as software, hardware, firmware, or any combination of the above items that perform specific functions. For example, each unit can correspond to an application-specific integrated circuit, or can correspond to pure software code, or can also correspond to a module combining software and hardware. In addition, one or more functions implemented by each unit can also be uniformly executed by components in a physical entity device (such as a processor, client, or server, etc.).
[0110] In addition, with reference to Figure 1 the described emotion recognition method can be implemented by a program (or instruction) recorded on a computer-readable storage medium. For example, according to the exemplary embodiments of the present disclosure, a computer-readable storage medium storing instructions can be provided, where when the instructions are run by at least one computing device, the at least one computing device is prompted to execute the emotion recognition method according to the present disclosure.
[0111] The computer program in the above computer-readable storage medium can run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. It should be noted that the computer program can also be used to execute additional steps other than the above steps or perform more specific processing when executing the above steps. The content of these additional steps and further processing has been mentioned during the description of the related method with reference to Figure 1 and will not be repeated here to avoid redundancy.
[0112] It should be noted that each unit in the emotion recognition device according to the exemplary embodiments of the present disclosure can completely rely on the operation of a computer program to implement corresponding functions, that is, each unit corresponds to each step in the functional architecture of the computer program, so that the entire system can be called through a dedicated software package (for example, a lib library) to implement corresponding functions.
[0113] On the other hand, Figure 3 Each of the units shown can also be implemented by hardware, software, firmware, middleware, microcode, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segment for performing the corresponding operations can be stored in a computer-readable medium such as a storage medium, so that the processor can perform the corresponding operations by reading and running the corresponding program code or code segment.
[0114] For example, the exemplary embodiments of the present disclosure can also be implemented as a computing device, which includes a storage component and a processor. A set of computer-executable instructions is stored in the storage component. When the set of computer-executable instructions is executed by the processor, an emotion recognition method according to the exemplary embodiments of the present disclosure is executed.
[0115] Specifically, the computing device can be deployed in a server or a client, or can be deployed on a node device in a distributed network environment. In addition, the computing device can be a PC computer, a tablet device, a personal digital assistant, a smart phone, a web emotion recognition device, or other devices capable of executing the above instruction set.
[0116] Here, the computing device does not have to be a single computing device, but can also be any aggregate of devices or circuits that can individually or jointly execute the above instructions (or instruction sets). The computing device can also be a part of an integrated control system or a system manager, or can be configured to be interconnected with a local or remote (for example, via wireless transmission) interface of a portable electronic device.
[0117] In the computing device, the processor can include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor can also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0118] Some of the operations described in the emotion recognition method according to the exemplary embodiments of the present disclosure can be implemented in software, some can be implemented in hardware, and in addition, these operations can also be implemented in a combination of software and hardware.
[0119] The processor can execute instructions or code stored in one of the storage components, where the storage components can also store data. The instructions and data can also be sent and received via a network interface device over a network, where the network interface device can employ any known transmission protocol.
[0120] The storage components can be integrated with the processor, for example, by arranging RAM or flash memory within an integrated circuit microprocessor, etc. Additionally, the storage components can include independent devices such as external disk drives, storage arrays, or other storage devices that can be used by any database system. The storage components and the processor can be operatively coupled or can communicate with each other, for example, via I / O ports, network connections, etc., such that the processor can read files stored in the storage components.
[0121] Furthermore, the computing device can also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, mouse, touch input device, etc.). All components of the computing device can be connected to each other via a bus and / or a network.
[0122] The emotion recognition method according to an exemplary embodiment of the present disclosure can be described as various interconnected or coupled functional blocks or functional diagrams. However, these functional blocks or functional diagrams can be equally integrated into a single logical device or operate with non-exact boundaries.
[0123] Therefore, referring to Figure 1 the described emotion recognition method can be implemented by a system including at least one computing device and at least one storage device storing instructions.
[0124] According to an exemplary embodiment of the present disclosure, at least one computing device is a computing device for executing the emotion recognition method according to an exemplary embodiment of the present disclosure, and a set of computer-executable instructions is stored in the storage device. When the set of computer-executable instructions is executed by at least one computing device, the emotion recognition method described referring to Figure 1 is executed.
[0125] According to an exemplary embodiment of the present disclosure, a computer program product can also be provided, including instructions that, when run by at least one computing device, cause the at least one computing device to execute the emotion recognition method according to the present disclosure.
[0126] The above describes various exemplary embodiments of the present disclosure. It should be understood that the above description is merely exemplary and not exhaustive. The present disclosure is not limited to the disclosed exemplary embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the present disclosure. Therefore, the protection scope of the present disclosure should be determined by the scope of the claims.
Claims
1. An emotion recognition method using an adaptive two-stage task, characterized in that: The emotion recognition method is used to recognize long-term emotions, and the emotion recognition method includes: In the first stage, first task information is output to prompt the target user to perform a first task corresponding to the first task information, wherein the first task includes an emotional image task lasting a first duration and a semantic fluency task lasting a second duration; Acquiring first physiological data of the target user during the process of the target user performing the first task; Using a first emotion recognition model, processing the first physiological data to obtain coarse-screened emotion information; Determine second task information from a plurality of candidate task information according to the roughly screened emotional information, wherein the tasks corresponding to the plurality of candidate task information include interview tasks; After a preset interval, the second stage is entered, and the second task information is output to prompt the target user to perform the second task corresponding to the second task information; Acquiring second physiological data of the target user during the process of the target user performing the second task; The second physiological data is processed using a second emotion recognition model to obtain emotion recognition information.
2. The emotion recognition method according to claim 1, characterized in that: There is an interval of a preset duration between the emotional image task and the semantic fluency task; The first duration ranges from 65 seconds to 105 seconds; The second duration ranges from 190 seconds to 230 seconds; The preset interval duration ranges from 20 seconds to 30 seconds.
3. The emotion recognition method according to claim 1, characterized in that: The roughly screened emotional information includes emotional indexes of a plurality of preset emotions, each emotional index is used to indicate the possibility that the user has a corresponding preset emotion, the plurality of candidate task information includes general task information and emotional task information corresponding to the plurality of preset emotions one by one, wherein determining the second task information from the plurality of candidate task information according to the roughly screened emotional information includes: Determine an emotion index greater than an index threshold from the emotion indexes of the plurality of preset emotions, and use the preset emotion corresponding to the determined emotion index as the target emotion; When the number of the target emotions is less than a preset number, determining the emotion task information corresponding to the target emotion from the plurality of candidate task information as the second task information; When the number of the target emotions is greater than or equal to the preset number, the general task information is determined as the second task information.
4. The emotion recognition method according to claim 3, characterized in that: The emotional task corresponding to the emotional task information includes an interview task, and the general task corresponding to the general task information includes a picture description task; and / or The second task lasts for a third duration, and a value range of the third duration is 100 seconds to 140 seconds.
5. The emotion recognition method according to any one of claims 1 to 4, characterized in that: The second physiological data includes voice data and near-infrared brain imaging data, the voice data is data obtained by recording the oral content of the target user, the second emotion recognition model includes a natural language processing model, a near-infrared emotion recognition model, and an information fusion model, wherein the second physiological data is processed using the second emotion recognition model to obtain emotion recognition information, including: Using the natural language processing model, performing natural language processing on the voice data to identify the emotion of the target user and obtain semantic emotion information; Using the near-infrared emotion recognition model, processing the near-infrared brain imaging data to identify the target user's emotions, and obtaining near-infrared emotion information; The information fusion model is used to fuse the semantic emotion information and the near-infrared emotion information to obtain the emotion recognition information.
6. The emotion recognition method according to claim 5, characterized in that: The emotion recognition method further comprises: Performing consistency comparison on the coarse-screened emotion information, the semantic emotion information and the near-infrared emotion information to obtain consistency information, wherein the consistency information is used to reflect the consistency degree between the identified different emotion information; When the consistency information indicates that the consistency between the identified different emotion information is lower than a preset degree, a prompt message is output to indicate that the recognition result is inaccurate.
7. An emotion recognition device using an adaptive two-stage task, characterized in that: The emotion recognition device is used to recognize long-term emotions, and the emotion recognition device includes: A first output unit is configured to output first task information in a first stage to prompt a target user to perform a first task corresponding to the first task information, wherein the first task includes an emotional image task lasting a first duration and a semantic fluency task lasting a second duration; A first acquiring unit is configured to acquire first physiological data of the target user during the process in which the target user performs the first task; a coarse screening unit, configured to process the first physiological data using a first emotion recognition model to obtain coarse screening emotion information; A determination unit is configured to determine second task information from a plurality of candidate task information according to the roughly screened emotion information, wherein the tasks corresponding to the plurality of candidate task information include interview tasks; A second output unit is configured to enter a second stage after a preset interval time and output the second task information to prompt the target user to perform a second task corresponding to the second task information; A second acquiring unit is configured to acquire second physiological data of the target user during the process of the target user performing the second task; The recognition unit is configured to use a second emotion recognition model to process the second physiological data to obtain emotion recognition information.
8. A computer-readable storage medium storing instructions, characterized in that: When the instructions are executed by at least one computing device, the at least one computing device is prompted to execute the emotion recognition method according to any one of claims 1 to 6.
9. A system comprising at least one computing device and at least one storage device storing instructions, characterized in that: When the instructions are executed by the at least one computing device, the at least one computing device is prompted to perform the emotion recognition method according to any one of claims 1 to 6.
10. A computer program product comprising instructions, characterized in that When the instructions are executed by at least one computing device, the at least one computing device is prompted to perform the emotion recognition method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Content recommendation method and device based on user emotion and readable storage medium
CN117688246A
Emotion recognition method and device, storage medium, system and program product
CN118861785A
Multi-mode depressive emotion feature recognition method based on brain imaging
CN118969241A