Electronic apparatus, psychological state fluctuation evaluation method and program
The electronic device estimates and visualizes psychological state changes in conversations by analyzing user utterances, addressing the lack of dialogue impact visualization in existing systems.
Patent Information
- Application Number
- JP2024046039
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-22
- Publication Date
- 2025-10-03
AI Technical Summary
Existing systems fail to visualize the impact of dialogue content on a user's psychological state before and after a speaker change in a conversation.
An electronic device that estimates psychological states based on user utterances using video, audio, and natural language data, derives psychological state change information, and displays it with user identification and utterance content for visualization.
Enables visualization of the influence of conversation content on a user's psychological state, facilitating effective dialogue analysis.
Smart Images

Figure 2025145716000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an electronic device, a psychological state fluctuation evaluation method, and a program. [Background technology]
[0002] BACKGROUND ART Conventionally, a speech content output system has been disclosed that detects speech made by a user, associates the speech content with the speaker, and displays the speech content on a screen (see Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2011-192048 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the system disclosed in Patent Document 1 cannot display changes in the user's psychological state before and after a speaker (speaker) changes in a dialogue between users in association with the dialogue, making it impossible to visualize the impact of the content of the dialogue on the user's psychological state.
[0005] The present invention has been made in view of such problems, and aims to visualize the influence that the content of a conversation has on the psychological state of a user. [Means for solving the problem]
[0006] In order to solve the above problem, the electronic device of the present invention is characterized by comprising: a first estimation means for estimating a first psychological state, which is the psychological state of the first user, based on a first utterance uttered by the first user; a second estimation means for estimating a second psychological state, which is the psychological state of the first user, based on a second utterance uttered by a second user immediately before the first utterance; a derivation means for deriving psychological state change information indicating a change in the psychological state of the first user from the second psychological state estimated by the second estimation means to the first psychological state estimated by the first estimation means; and a control means for displaying the psychological state change information derived by the derivation means on a display unit together with at least one of identification information identifying the second user, the content of the second utterance, and video data showing the first user during the first utterance and the second utterance. [Effects of the Invention]
[0007] According to the present invention, it is possible to visualize the influence of the content of a conversation on the psychological state of a user. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 2 is a block diagram showing the functional configuration of a PC. [Figure 2] FIG. 10 is a diagram showing a control procedure for psychological state change evaluation processing. [Figure 3] FIG. 1 is a diagram showing the flow of a one-on-one meeting between a first speaker and a second speaker. [Figure 4] FIG. 10 is a diagram illustrating the relationship between fluctuations in emotion estimation results and psychological state fluctuation information. [Figure 5] FIG. 10 is a diagram showing an example of the contents of an utterance information database. [Figure 6] FIG. 10 is a diagram showing an example of a psychological state change evaluation screen. [Figure 7] FIG. 10 is a diagram showing an example of a psychological state change evaluation screen. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, an example of an embodiment in which an electronic device according to the present invention is applied to a PC (Personal Computer) will be described with reference to the drawings. The PC is assumed to be a desktop PC used by a user (e.g., a human resources officer of a company), but may also be a notebook PC, tablet PC, etc.
[0010] 1, the PC 10 includes a CPU 11, a RAM 12, a storage unit 13, an operation unit 14, a display unit 15, a communication unit 16, and a bus 17. The components of the PC 10 are connected to each other via the bus 17.
[0011] The CPU (first estimating means, second estimating means, derivation means, control means) 11 is a processor that reads and executes a program 131 stored in the storage unit 13 and performs various arithmetic processing, thereby controlling the operation of each unit of the PC 10. The RAM 12 provides a working memory space for the CPU 11 and stores temporary data. The storage unit 13 is a non-transitory recording medium readable by the CPU 11 as a computer, and stores the program 131 and various data. The program 131 is stored in the storage unit 13 in the form of computer-readable program code. The various data stored in the storage unit 13 include, for example, a video database 132, a first emotion estimating model 133, a second emotion estimating model 134, a third emotion estimating model 135, and an utterance information database 136.
[0012] The video database 132 is a database for storing conference recording data of one-on-one meetings between employees of the company (e.g., a superior and a subordinate). Here, a one-on-one meeting is an online meeting held over a network (e.g., the Internet) using a terminal device such as a PC by each participating employee. The terminal device used in this one-on-one meeting is assumed to be capable of recording at least the audio and video (conference recording data) when the one-on-one meeting is being held. In other words, the audio and video (conference recording data) recorded by this terminal device are stored in the video database 132. The video is displayed in a manner that allows the movements and facial expressions (gestures) of each employee participating in the one-on-one meeting to be seen.
[0013] The first emotion estimation model 133 is a learning model (trained model) machine-learned based on learning data that takes video data as input data and outputs positive (Pos), neutral (Neu), and negative (Neg) information (information indicating emotions) as correct answer information. By inputting video data (estimation information) acquired from conference recording data stored in the video database 132 into this first emotion estimation model 133, it is possible to estimate the emotions of employees who participated in a one-on-one meeting. The second emotion estimation model 134 is a learning model (trained model) machine-learned based on learning data that takes audio data as input data and outputs positive (Pos), neutral (Neu), and negative (Neg) information (information indicating emotions) as correct answer information. By inputting audio data (estimation information) acquired from conference recording data stored in the video database 132 into this second emotion estimation model 134, it is possible to estimate the emotions of employees who participated in a one-on-one meeting. The third emotion estimation model 135 is a learning model (trained model) that has been machine-learned based on learning data in which natural language data (utterance content) is input data and positive (Pos), neutral (Neu), and negative (Neg) information (information indicating emotions) is output data (correct answer information). By inputting natural language data (estimation information) acquired from conference recording data stored in the video database 132 into the third emotion estimation model 135, the emotions of employees participating in a one-on-one meeting can be estimated. Here, the natural language data (utterance content) is acquired by performing speech recognition on audio data acquired from the conference recording data. Emotion estimation using the first emotion estimation model 133, emotion estimation using the second emotion estimation model 134, and emotion estimation using the third emotion estimation model 135 are each performed for each speech section (see FIG. 3) during a one-on-one meeting. A speech section refers to a section from when a speaker starts speaking to when the speaker finishes speaking.
[0014] The speech information database 136 is a database for storing speech information from one-on-one meetings in association with changes in the psychological state of each employee who participated in the one-on-one meeting. The speech information database 136 will be described in detail later.
[0015] The operation unit 14 has a key input unit such as a keyboard and a pointing device such as a mouse, and accepts key operation inputs and position operation inputs from the user, outputting the operation information to the CPU 11. The CPU 11 accepts the user's input operations based on the information transmitted from the operation unit 14. The display unit 15 is composed of an LCD (Liquid Crystal Display), an EL (Electro Luminescence) display, etc., and displays various information according to the display information instructed by the CPU 11. The communication unit 16 performs communication operations in accordance with a predetermined communication standard. Through this communication operation, the communication unit 16 transmits and receives information to and from external devices via a communication network.
[0016] Next, we will explain the operation of the PC 10. Specifically, we will explain the psychological state variation evaluation process executed by the PC 10. The psychological state variation evaluation process is started, for example, when new conference recording data is stored in the above-mentioned video database 132, but this is merely an example.
[0017] As shown in Fig. 2, first, the CPU 11 of the PC 10 estimates the emotions of a first speaker (e.g., a superior) and a second speaker (e.g., a subordinate) for each speech section (speech section a, utterance section b, utterance section c, utterance section d, ...) based on the conference recording data newly stored in the video database 132 as shown in Fig. 3 (step S1). Specifically, the CPU 11 estimates the emotions of the first speaker and the second speaker by inputting video data acquired from the conference recording data for each speech section into a first emotion estimation model 133. Furthermore, the CPU 11 estimates the emotions of the first speaker and the second speaker by inputting audio data acquired from the conference recording data for each speech section into a second emotion estimation model 134. Furthermore, the CPU 11 estimates the emotions of the first and second speakers by inputting the natural language data (utterance content) acquired from the conference recording data for each utterance section into the third emotion estimation model 135. Here, utterance sections a and c indicate sections in which the first speaker spoke. On the other hand, utterance sections b and d indicate sections in which the second speaker spoke. Note that the second speaker did not speak in utterance sections a and c. Therefore, the emotion estimation result of the second speaker using the audio data in these sections and the emotion estimation result of the second speaker using the natural language are both treated as neutral (Neu). Similarly, the first speaker did not speak in utterance sections b and d. Therefore, the emotion estimation result of the first speaker using the audio data in these sections and the emotion estimation result of the first speaker using the natural language are both treated as neutral (Neu).
[0018] Next, CPU 11 derives psychological state variation information for each of the first speaker and the second speaker (step S2). Specifically, CPU 11 derives the psychological state variation information by defining an utterance section in which the other speaker is speaking as an utterance section before the psychological state changes and defining an utterance section in which the speaker himself is speaking as an utterance section after the psychological state changes. For example, when deriving psychological state variation information for the second speaker, as shown in FIG. 3, CPU 11 defines an utterance section a in which the first speaker (other speaker) is speaking as an utterance section before the psychological state changes and defines an utterance section b in which the second speaker (speaker himself) is speaking as an utterance section after the psychological state changes, and derives psychological state variation information from emotion estimation results for each utterance section (emotion estimation result based on video data, emotion estimation result based on audio data, emotion estimation result based on natural language data). Similarly, the speech section c in which the first speaker (the other speaker) is speaking is defined as the speech section before the psychological state changes, and the speech section d in which the second speaker (the speaker himself) is speaking is defined as the speech section after the psychological state changes, and psychological state change information is derived from the emotion estimation results for each speech section. On the other hand, when deriving the psychological state change information for the first speaker, as shown in Figure 3, the speech section b in which the second speaker (the other speaker) is speaking is defined as the speech section before the psychological state changes, and the speech section c in which the first speaker (the speaker himself) is speaking is defined as the speech section after the psychological state changes, and psychological state change information is derived from the emotion estimation results for each speech section.
[0019] Here, as shown in Figure 4, psychological state change information includes "Pos," which indicates a change in a positive direction, such as when the emotion estimation result (psychological state) changes from "Neu" to "Pos," "Neg" to "Neu," or "Neg" to "Pos" before and after the psychological state change. It also includes "Neg," which indicates a change in a negative direction, such as when the emotion estimation result changes from "Neu" to "Neg," "Pos" to "Neu," or "Pos" to "Neg" before and after the psychological state change. It also includes "Non," which indicates no change, such as when the emotion estimation result changes from "Neu" to "Neu," "Neg" to "Neg," or "Pos" to "Pos" before and after the psychological state change.
[0020] Next, a method for deriving psychological state variation information will be described in detail. First, the CPU 11 derives psychological state variation information based on video data using emotion estimation results based on the video data (emotion estimation results before and after a psychological state variation). Similarly, the CPU 11 derives psychological state variation information based on audio data using emotion estimation results based on the audio data (emotion estimation results before and after a psychological state variation). Furthermore, the CPU 11 derives psychological state variation information based on natural language data using emotion estimation results based on the natural language data. Then, the CPU 11 counts the number of “Pos” and the number of “Neg” in the derived three pieces of psychological state variation information. If the number of “Pos” is greater than the number of “Neg”, the psychological state variation information is derived as “Pos”. If the number of “Pos” is less than the number of “Neg”, the psychological state variation information is derived as “Neg”. Furthermore, when the number of "Pos" and the number of "Neg" are the same, the psychological state fluctuation information is derived as "Neg".
[0021] Next, CPU 11 associates the psychological state variation information derived in step S2 with the corresponding utterance information and stores it in utterance information database 136 (step S3). Specifically, as shown in Fig. 3, CPU 11 associates the psychological state variation information (e.g., "Non") of the second speaker in utterance sections a and b with information (utterance information) such as the utterance content Aa (e.g., "This term's evaluation is..."), utterance format (e.g., "Other"), utterance date (e.g., "2024 / 3 / 1"), and speaker (e.g., "Mr. A") uttered by the first speaker in utterance section a, and stores the information in utterance information database 136 (see Fig. 5). Furthermore, the psychological state fluctuation information of the first speaker in speech sections b and c (e.g., "Non") is linked to the information (speech information) of the utterance content Bb (e.g., "It was good"), speech format (e.g., "Other"), utterance date (e.g., "2024 / 3 / 1"), and speaker (e.g., "Mr. B") uttered by the second speaker in speech section b, and stored in the speech information database 136. Similarly, the psychological state fluctuation information of the first speaker or the second speaker for subsequent speech sections is linked to the information (speech information) of the corresponding utterance content, speech format, utterance date, and speaker, and stored in the speech information database 136. Here, the utterance content is extracted using natural language data verbalized by speech recognition from the speech data in the corresponding utterance section. Furthermore, when the utterance content is stored in the utterance information database 136, it is stored in the form of a sentence summarizing the utterance content, but it may also be stored in the form of words such as nouns, verbs, and adjectives that appeared in the utterance. The utterance format consists of three formats: "question," "answer," and "other," and is extracted using the above-mentioned voice data and natural language data. Speaker information is identification information that identifies the speaker.
[0022] Next, CPU 11 refers to utterance information database 136 (see FIG. 5) and determines whether the same utterance content by the same speaker has existed in the past, for the utterance information stored in step S3 (step S4). Here, as a method for determining whether the utterance content is the same, the utterance content (summary sentence) is converted into an embedded representation (vector), COS similarity is calculated, and whether the utterance content is the same is determined based on the value of COS similarity. Note that, if the utterance content is in the form of the above-mentioned words, whether the utterance content is the same is determined based on the number of matching words. If it is determined in step S4 that the same utterance content by the same speaker has not existed in the past (step S4; NO), CPU 11 ends the psychological state variation evaluation process. Also, in step S4, if it is determined that the same utterance content has been made by the same speaker in the past (step S4; YES), CPU 11 refers to the utterance information database 136 (see Figure 5) and determines whether the psychological state change information linked to the corresponding utterance content is "Pos" or "Neg" (step S5).
[0023] In step S5, if it is determined that the psychological state variation information linked to the corresponding utterance content is neither "Pos" nor "Neg" (step S5; NO), that is, if it is determined that the psychological state variation information linked to the corresponding utterance content is "Non", the CPU 11 ends the psychological state variation evaluation process. Also, in step S5, if it is determined that the psychological state variation information linked to the corresponding utterance content is "Pos" or "Neg" (step S5; YES), the CPU 11 displays the occurrence state (occurrence frequency) of the corresponding utterance content on the psychological state variation evaluation screen G (step S6). Then, the CPU 11 ends the psychological state variation evaluation process.
[0024] As shown in FIG. 6, the psychological state fluctuation evaluation screen G displays the occurrence status (e.g., “continuous”) of the corresponding utterance content, along with the speaker who uttered the utterance content (e.g., “Mr. A”), the utterance content (e.g., “You did a great job.”), and the psychological state fluctuation information linked to the utterance content (e.g., “Pos”). Here, the occurrence status (occurrence frequency) displayed on the psychological state fluctuation evaluation screen G includes three occurrence statuses: “single,” “periodic,” and “continuous.” “Single” indicates that the corresponding utterance content occurs irregularly. “Periodic” indicates that the corresponding utterance content occurs periodically. “Continuous” indicates that the corresponding utterance content occurs continuously. Whether the occurrence status is “single,” “periodic,” or “continuous” is determined using information in the “utterance date” field of the utterance information database 136 (see FIG. 5).
[0025] As described above, the CPU 11 of the PC 10 estimates a first psychological state, which is the psychological state of a first user (e.g., a second speaker; see FIG. 3), based on a first utterance (e.g., speech section b; see FIG. 3) uttered by the first user. The CPU 11 estimates a second psychological state, which is the psychological state of the first user, based on a second utterance (e.g., speech section a; see FIG. 3) uttered by a second user (e.g., the first speaker; see FIG. 3) immediately before the first utterance. The CPU 11 derives psychological state change information indicating a change in the psychological state of the first user from the estimated second psychological state to the first psychological state. The CPU 11 displays the derived psychological state change information on the display unit 15 together with at least one of identification information identifying the second user, the content of the second utterance, and video data showing the first user during the first utterance and the second utterance (see FIG. 6). Therefore, the PC 10 can visualize the effect that the content of the speech of the second user (e.g., the first speaker; see FIG. 3) has on the psychological state of the first user (e.g., the second speaker; see FIG. 3).
[0026] Furthermore, the CPU 11 estimates a first psychological state based on each of the video, audio, and speech content (natural language) related to a first utterance (e.g., speech section b; see FIG. 3), and estimates a second psychological state based on each of the video, audio, and speech content (natural language) related to a second utterance (e.g., speech section a; see FIG. 3). The CPU 11 derives psychological state change information based on each of the change in the psychological state of the first user (e.g., second speaker; see FIG. 3) from the second psychological state based on the estimated video related to the second utterance to the first psychological state based on the video related to the first utterance, the change in the psychological state of the first user from the second psychological state based on the estimated audio related to the second utterance to the first psychological state based on the audio related to the first utterance, and the change in the psychological state of the first user from the second psychological state based on the estimated audio related to the second utterance to the first psychological state based on the audio related to the first utterance. Therefore, according to the PC 10, psychological state variation information is derived based on each data of moving images, sounds, and speech content (natural language), and therefore the psychological state variation information can be appropriately derived.
[0027] Furthermore, the CPU 11 expresses the change in the psychological state of the first user (e.g., the second speaker; see FIG. 3) from the second psychological state corresponding to each of the video, audio, and speech content (natural language) to the first psychological state with one of three indices ("Pos," "Non," "Neg") corresponding to a positive change, no change, or negative change. If the positive change is the most prevalent among the three indices, the CPU 11 derives the positive change (Pos) as the psychological state change information. If the negative change is the most prevalent among the three indices, the CPU 11 derives the negative change (Neg) as the psychological state change information. If the positive and negative changes are equal in number, the CPU 11 derives no change (Non) as the psychological state change information. Therefore, according to the PC 10, by using the three indices ("Pos," "Non," "Neg") corresponding to a positive change, no change, or negative change, the psychological state change information can be appropriately and simply derived.
[0028] Furthermore, the CPU 11 associates the derived psychological state fluctuation information of the first user (e.g., the second speaker; see FIG. 3) with at least one of identification information for identifying the second user (e.g., the first speaker; see FIG. 3), the content of the second utterance (e.g., utterance section a; see FIG. 3), and video data showing the first user during the first utterance and the second utterance, and stores the information in the utterance information database 136 of the storage unit 13. Therefore, by referring to the utterance information database 136, the PC 10 can grasp the occurrence state of the psychological state fluctuation information derived due to the utterance of the second user (e.g., the first speaker; see FIG. 3). As a result, this information can be used to handle future one-on-one meetings.
[0029] In addition, when the derived psychological state fluctuation information is a predetermined index ("Pos" or "Neg") out of three indexes ("Pos", "Non", "Neg"), the CPU 11 determines whether the same utterance content as the utterance content relating to the second utterance (for example, utterance section a; see Figure 3) used to derive the psychological state fluctuation information is stored in the utterance information database 136, and displays the result of the determination on the display unit 15. Specifically, when the derived psychological state variation information is a predetermined index ("Pos" or "Neg") among three indices ("Pos", "Non", "Neg"), the CPU 11 determines, based on the content of the second utterance stored in the utterance information database 136, the utterance frequency (occurrence state) of the utterance content that is the same as the utterance content of the second utterance (e.g., utterance section a; see FIG. 3) used to derive the psychological state variation information and that is made by the same second user (e.g., first speaker; see FIG. 3), and causes the display unit 15 to display the result of the determination. Therefore, the PC 10 can grasp the occurrence state of the psychological state variation information derived due to the utterance of the second user, and therefore it is possible to verify the utterance of the second user and evaluate whether the one-on-one meeting is being conducted appropriately.
[0030] Although the present invention has been specifically described above based on the embodiments, the present invention is not limited to the above embodiments and can be modified within the scope of the invention. For example, in the above embodiment, the psychological state change evaluation screen G (see FIG. 6) displays the occurrence state (e.g., "continuous") of the corresponding utterance content together with the speaker who uttered the utterance content (e.g., "Mr. A"), the utterance content (e.g., "You did a great job."), and psychological state change information (e.g., "Pos") linked to the utterance content. However, as shown in FIG. 7, this information may be displayed superimposed on a moving image of the corresponding utterance section. Furthermore, although not shown, the psychological state before the psychological state change (e.g., the psychological state (emotion estimation result) of the first user in utterance section a) and the psychological state after the change (e.g., the psychological state (emotion estimation result) of the first user in utterance section b) may be displayed together with this information.
[0031] In addition, in the above embodiment, the CPU 11 of the PC 10 may be configured to display on the display unit 15 the psychological state fluctuation information ("Pos", "Non" or "Neg") regarding the first user (e.g., the second speaker; see Figure 3) derived in step S2 of the psychological state fluctuation evaluation process (see Figure 2), together with at least one of identification information identifying the second user (e.g., the first speaker; see Figure 3), the speech content of the second utterance (e.g., speech section a; see Figure 3), and video data showing the first user during the first utterance (e.g., speech section b; see Figure 3) and the second utterance. In addition, when the derived psychological state fluctuation information regarding the first user (e.g., the second speaker; see Figure 3) is a predetermined index ("Pos" or "Neg") out of three indexes ("Pos", "Non", "Neg"), the CPU 11 may cause the display unit 15 to display the psychological state fluctuation information represented by the predetermined index together with at least one of identification information identifying the second user (e.g., the first speaker; see Figure 3), the speech content of the second utterance (e.g., speech section a; see Figure 3), and video data showing the first user during the first utterance (e.g., speech section b; see Figure 3) and the second utterance.
[0032] Furthermore, in the above embodiment, for example, the CPU 11 of the PC 10 may create a graph showing psychological state variation information ("Pos", "Non", "Neg") along the vertical axis and the elapsed time of the one-on-one meeting along the horizontal axis, and display the graph on the display unit 15. By displaying this graph, it becomes possible to grasp the change in the psychological state variation information over time. Furthermore, instead of the psychological state variation information ("Pos", "Neu", "Neg"), the psychological state (emotion estimation result ("Pos", "Neu", "Neg")) in each utterance section may be displayed along the vertical axis.
[0033] Furthermore, in the above embodiment, the number of pieces of information output by the first emotion deduction model 133, the second emotion deduction model 134, and the third emotion deduction model 135 does not have to be three: "Pos", "Neu", and "Neg". For example, learning data may be output as five pieces of information. In this case, the psychological state fluctuation information may use the three indices "Pos", "Neu", and "Neg" as in the above embodiment, or different indices may be used depending on the magnitude of the change even if the psychological state fluctuation is in the same direction.
[0034] Furthermore, in the above embodiment, each piece of information ("Pos", "Neu", or "Neg") of the first emotion deduction model 133, the second emotion deduction model 134, and the third emotion deduction model 135 may be displayed on the display unit together with video data corresponding to each piece of information and utterance content corresponding to each piece of information, or psychological state variation information ("Pos", "Non", or "Neg") may be displayed on the display unit together with video data corresponding to the psychological state variation information and utterance content corresponding to each piece of psychological state variation information. [Explanation of symbols]
[0035] 10 PC, 11 CPU, 13 memory unit, 131 program, 132 video database, 133 first emotion estimation model, 134 second emotion estimation model, 135 third emotion estimation model, 136 speech information database, 15 display unit
Claims
1. a first estimation means for estimating a first psychological state, which is a psychological state of the first user, based on a first utterance uttered by the first user; a second estimation means for estimating a second psychological state of the first user based on a second utterance uttered by a second user immediately before the first utterance; a derivation means for deriving psychological state change information indicating a change in the psychological state of the first user from the second psychological state estimated by the second estimation means to the first psychological state estimated by the first estimation means; a control means for displaying the psychological state fluctuation information derived by the derivation means on a display unit together with at least one of identification information for identifying the second user, the content of the second utterance, and video data showing the first user during the first utterance and the second utterance; and An electronic device comprising:
2. the first estimation means estimates the first psychological state based on a moving image, a sound, and a speech content related to the first utterance; the second estimation means estimates the second psychological state based on a moving image, a sound, and a speech content related to the second utterance; the derivation means derives the psychological state change information based on a change in the psychological state of the first user from the second psychological state based on a moving image related to the second utterance estimated by the second estimation means to the first psychological state based on the moving image related to the first utterance estimated by the first estimation means, a change in the psychological state of the first user from the second psychological state based on a voice related to the second utterance estimated by the second estimation means to the first psychological state based on the voice related to the first utterance estimated by the first estimation means, and a change in the psychological state of the first user from the second psychological state based on a speech content related to the second utterance estimated by the second estimation means to the first psychological state based on a speech content related to the first utterance estimated by the first estimation means.
2. The electronic device according to claim 1, wherein the electronic device is a semiconductor device.
3. the derivation means represents a change in the psychological state of the first user from the second psychological state to the first psychological state corresponding to each of the video, the audio, and the speech content by one of three indices corresponding to a positive change, no change, or a negative change, and when the positive change is the most common among the three indices, derives the positive change as the psychological state change information, when the negative change is the most common, derives the negative change as the psychological state change information, and when the positive change and the negative change are the same in number, derives no change as the psychological state change information.
3. The electronic device according to claim 2.
4. When the psychological state variation information derived by the derivation means is a predetermined index among the three indexes, the control means causes a display unit to display the psychological state variation information represented by the predetermined index together with at least one of identification information for identifying the second user, the utterance content of the second utterance, and video data showing the first user during the first utterance and the second utterance.
4. The electronic device according to claim 3.
5. the control means stores the psychological state fluctuation information derived by the derivation means in a storage unit in association with at least one of identification information for identifying the second user, the utterance content of the second utterance, and video data showing the first user during the first utterance and the second utterance.
4. The electronic device according to claim 3.
6. the control means, when the psychological state fluctuation information derived by the derivation means is a predetermined index among the three indexes, determines whether or not the same utterance content as the utterance content related to the second utterance used to derive the psychological state fluctuation information is stored in the storage unit, and causes the display unit to display the result of the determination.
6. The electronic device according to claim 5,
7. When the psychological state fluctuation information derived by the derivation means is a predetermined index among the three indexes, the control means determines, based on the content of the second utterance stored in the storage unit, the frequency of utterance content that is the same as the content of the second utterance used to derive the psychological state fluctuation information and that is the same as the content of the second utterance used to derive the psychological state fluctuation information, and that is the same as the content of the second user, and causes the display unit to display the result of the determination.
6. The electronic device according to claim 5,
8. A psychological state fluctuation evaluation method executed by a computer of an electronic device, comprising: a first estimation step of estimating a first psychological state, which is a psychological state of the first user, based on a first utterance uttered by the first user; a second estimation step of estimating a second psychological state of the first user based on a second utterance uttered by a second user immediately before the first utterance; a derivation step of deriving psychological state change information indicating a change in the psychological state of the first user from the second psychological state estimated by the second estimation step to the first psychological state estimated by the first estimation step; a control step of displaying the psychological state fluctuation information derived by the derivation step on a display unit together with at least one of identification information for identifying the second user, the content of the second utterance, and video data showing the first user during the first utterance and the second utterance; A method for evaluating psychological state fluctuations, comprising:
9. Electronic equipment computers, a first estimation means for estimating a first psychological state, which is a psychological state of the first user, based on a first utterance uttered by the first user; a second estimation means for estimating a second psychological state, which is a psychological state of the first user, based on a second utterance uttered by a second user immediately before the first utterance; a derivation means for deriving psychological state change information indicating a change in the psychological state of the first user from the second psychological state estimated by the second estimation means to the first psychological state estimated by the first estimation means; a control means for displaying the psychological state fluctuation information derived by the derivation means on a display unit together with at least one of identification information for identifying the second user, the content of the second utterance, and video data showing the first user during the first utterance and the second utterance; A program characterized by functioning as
Citation Information
Patent Citations
Speech content output system, speech content output device, and speech content output method
JP2011192048A