Information processing device, information processing method, and information processing program

The information processing system addresses the issue of inadequate reflection by reading user answers aloud in different voices, enhancing introspection and feedback quality.

JP2026060263APending Publication Date: 2026-04-08유겐가이샤티아이에스
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2026-04-08

AI Technical Summary

Technical Problem

Users may not adequately reflect on their experiences, emotions, and thoughts when answering reflection questions without sufficient introspection, leading to inappropriate feedback.

Method used

An information processing system that includes a terminal device and a server device, utilizing a reception unit, voice generation unit, and voice control unit to receive answers, generate voice data for reading aloud, and output the answers in various voices, encouraging users to reflect on their experiences.

Benefits of technology

Encourages users to reflect on their experiences by reading their answers aloud in different voices, allowing them to recall memories and verify the appropriateness of their responses, thereby providing appropriate feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026060263000001_ABST
    Figure 2026060263000001_ABST
Patent Text Reader

Abstract

Encouraging users to reflect on their experiences. [Solution] The server device 100 includes a reception unit 131, a voice generation unit 133, and a voice control unit 134. The reception unit 131 receives answers from user U to questions regarding reflection. The voice generation unit 133 generates voice data that reads aloud the answers received by the reception unit 131. The voice control unit 134 controls the output of the voice data generated by the voice generation unit 133 to user U.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, and an information processing program.

Background Art

[0002] Conventionally, various techniques for supporting the growth of human resources using a computer have been known. For example, there is a technique of prompting reflection (introspection) to a user by posing a "question" from an application and having the user answer it.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in order to perform reflection appropriately, it is necessary for the user to answer questions regarding reflection while looking back on the user's own experiences, emotions, and thoughts at that time. However, if the user simply answers the questions, there may be cases where the user answers without sufficiently looking back. In this case, it may not be possible to appropriately provide feedback on reflection to the user.

Means for Solving the Problems

[0005] The information processing apparatus according to the present invention includes a reception unit, a voice generation unit, and a voice control unit. The reception unit receives an answer to a question regarding the reflection of a user. The voice generation unit generates voice data for reading out the answer received by the reception unit in voice. The voice control unit outputs the voice of the voice data generated by the voice generation unit to the user.

Effects of the Invention

[0006] According to the present invention, it is possible to encourage users to reflect on their experiences. [Brief explanation of the drawing]

[0007] [Figure 1] Figure 1 is a diagram illustrating the overview of the service according to the embodiment. [Figure 2] Figure 2 shows an example of an input screen for an application relating to a terminal device according to the embodiment. [Figure 3] Figure 3 shows an example of the configuration of a terminal device according to this embodiment. [Figure 4] Figure 4 shows an example of the configuration of a server device according to this embodiment. [Figure 5] Figure 5 shows an example of a screen displayed on a terminal device according to this embodiment. [Figure 6] Figure 6 shows an example of a prompt that instructs the estimation of emotion according to the embodiment. [Figure 7] Figure 7 shows an example of the output result of the generation AI according to the embodiment. [Figure 8] Figure 8 shows an example of corresponding data according to the embodiment. [Figure 9] Figure 9 shows an example of a prompt that instructs the generation of audio data according to the embodiment. [Figure 10] Figure 10 shows an example of a screen displayed on a terminal device according to this embodiment. [Figure 11] Figure 11 is a flowchart showing an example of information processing according to the embodiment. [Figure 12] Figure 12 is a hardware configuration diagram showing an example of a computer that implements the functions of a server device. [Modes for carrying out the invention]

[0008] The following describes in detail, with reference to the drawings, embodiments for implementing the information processing device, information processing method, and information processing program according to the present application (hereinafter referred to as "embodiments"). Note that these embodiments do not limit the information processing device, information processing method, and information processing program according to the present application. Furthermore, the same parts are denoted by the same reference numerals in each of the following embodiments, and redundant descriptions are omitted.

[0009] (Embodiment) (1. Introduction) Figure 1 is a diagram illustrating the overview of a service according to an embodiment. The service according to the embodiment is implemented by an information processing system 1. The information processing system 1 includes a terminal device 10 and a server device 100. The terminal device 10 and the server device 100 are connected to each other via a network N, whether wired or wireless, so as to be able to communicate. Network N is, for example, the Internet, a WAN (Wide Area Network), a LAN (Local Area Network), or the like.

[0010] Terminal device 10 is a smart device such as a smartphone or tablet used by user U, and is a mobile terminal device capable of communicating via a wireless communication network such as 4G (Generation) or LTE (Long Term Evolution). Terminal device 10 has a screen such as an LCD display that has touch panel functionality, and accepts various operations on displayed data such as content from user U using a finger or stylus, such as tapping, sliding, and scrolling. Operations performed on the area of ​​the screen where content is displayed may also be considered as operations on the content. Furthermore, terminal device 10 may be an information processing device such as a desktop PC (Personal Computer) or notebook PC, not just a smart device. In the example in Figure 1, terminal device 10 is shown as a smartphone.

[0011] The terminal device 10 can install a dedicated application (hereinafter referred to as an application) for using the service according to the embodiment. When the terminal device 10 uses a web application on the terminal device 10 that has the same function as the application, the application may not be installed. Hereinafter, the case where the application is installed on the terminal device 10 of the user U will be described.

[0012] In the embodiment, the server device 100 corresponds to the information processing device according to the present application. The server device 100 is an information processing device that cooperates with the terminal device 10 and provides services such as services by applications and various data to the terminal device 10, and is realized by a server device, a cloud system, or the like.

[0013] The server device 100 provides the service according to the embodiment. The service according to the embodiment is a human resource development support service for each person to grow into an autonomous talent and create an organization where they learn from each other. Here, an autonomous talent refers to a person who can think for themselves and perform tasks proactively and actively, rather than waiting for instructions. The user U is a user who uses the service according to the embodiment. The user U uses the service according to the embodiment provided by the server device 100 using the terminal device 10. The terminal device 10 poses a "question" to the user U through the application and prompts the user U to reflect on their experience (hereinafter also referred to as introspection).

[0014] Here, reflection (introspection) in the field of human resource development is "to step away from one's own work and mindset once and objectively review the work process, way of thinking, actions, etc.". Reflection is a future-oriented methodology in which the user U reexamines all experiences, including failed and successful experiences, gains insights, and connects them to new actions.

[0015] Specifically, the reflection framework involves looking back on oneself along the items of "opinions", "experiences", "emotions", and "values" to understand why one took such actions, what feelings one had at that time, and why those feelings occurred.

[0016] More specifically, for performing reflections in various situations, the application displays a plurality of questions and accepts selection of questions for which reflections are to be carried out. When any question is selected, the application sequentially displays each item along the reflection framework for the selected question and sequentially prompts for input of answers for each item.

[0017] FIG. 2 is a diagram showing an example of an input screen of an application related to the terminal device 10 according to an embodiment. For example, for a question, the application sequentially displays, as shown in (A) to (D) of FIG. 2, an "opinion" item asking for the opinion of the user U, an "experience" item asking for the experience behind the user U's opinion, an "emotion" item asking for the emotion felt by the user U in the experience, and a "value" item asking for the value underlying the emotion felt by the user U, and sequentially prompts for input to the answer fields for each item. The items of "opinion", "experience", and "value" are items for which answers are to be given in the answer fields by text input. The "emotion" item is an item for which a plurality of emotions such as anger are displayed as selection candidates, and answers are to be given in the answer field by selection from the selection candidates.

[0018] The user U organizes his or her thoughts about each item of opinion, experience, emotion, and value along the reflection framework and inputs them in order in response to the inquiries from the application. Thereby, the application can teach the user U the reflection framework that there is experience behind an opinion, the experience is stored together with emotions, and emotions arise from values.

[0019] Figure 1 schematically shows the flow of reflection by the service according to the embodiment. The server device 100 receives answers from the terminal device 10 via an application to questions about reflection from user U (step S1). The server device 100 performs an analysis of the received answers regarding reflection and generates analysis results (step S2). For example, the server device 100 uses an analysis model to analyze the answers regarding opinions, experiences, feelings, and values, and generates analysis results. Note that an existing so-called generative AI (Artificial Intelligence) may be used as the analysis model. The server device 100 then transmits the analyzed results to the terminal device 10, which displays them to user U (step S3).

[0020] This allows the service according to the embodiment to encourage introspection about user U's experience. As a result, the service according to the embodiment can appropriately support user U's growth.

[0021] (2. Configuration of terminal device 10) Next, the configuration of the terminal device 10 according to the embodiment will be described using Figure 3. Figure 3 is a diagram showing an example of the configuration of the terminal device 10 according to the embodiment. As shown in Figure 3, the terminal device 10 has a communication unit 20, an input unit 30, a display unit 40, an audio output unit 50, and a control unit 60.

[0022] The communication unit 20 is connected to the network N by wire or wireless connection and transmits and receives information to and from the server device 100 via the network N. For example, the communication unit 20 transmits reflection information related to reflection to the server device 100 and receives analysis results related to reflection and voice data from the server device 100. The communication unit 20 is implemented, for example, by a NIC (Network Interface Card).

[0023] The input unit 30 accepts various types of information input through various operations from the user U. For example, the input unit 30 accepts answers from the user U to questions about reflection. The input unit 30 may also accept various operations from the user U via a touch panel display. Furthermore, the input unit 30 may accept various operations from buttons provided on the terminal device 10, or from a keyboard or mouse connected to the terminal device 10.

[0024] The display unit 40 is a display device for displaying various types of information. For example, the display unit 40 displays various screens related to reflection. The display unit 40 also displays analysis results related to reflection. The display unit 40 is a display screen for a tablet terminal or the like, implemented using a liquid crystal display or an organic EL (Electro-Luminescence) display.

[0025] The audio output unit 50 is, for example, a speaker. The audio output unit 50 outputs sound related to reflection. The audio output unit 50 may also be wired or wirelessly connected earphones, headphones, or a headset.

[0026] The control unit 60 is, for example, a controller, and is implemented by a processor such as a CPU (Central Processing Unit) or MPU (Micro Processing Unit) executing various programs stored in the memory device inside the terminal device 10 using RAM (Random Access Memory) as the working area. For example, these various programs include application programs installed on the terminal device 10. For example, these various programs include application programs that display information transmitted from the server device 100. The control unit 60 is also implemented by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0027] As shown in Figure 3, the control unit 60 has a transmission unit 61 and an output control unit 62, and realizes or executes the information processing operations described below.

[0028] The transmission unit 61 transmits operation information performed by user U to the server device 100. The transmission unit 61 also transmits various types of reflection information to the server device 100. For example, the transmission unit 61 transmits a question about reflection and the answer to the question received from user U by the input unit 30 as reflection information to the server device 100. Specifically, the transmission unit 61 transmits the question content and the answers from user U to the server device 100 as reflection information, including opinions, experiences, feelings, and values.

[0029] The output control unit 62 controls the output of various types of information. For example, the output control unit 62 controls the display on the display unit 40 and the output of sound from the sound output unit 50. For instance, when the output control unit 62 receives analysis results related to reflection from the server device 100, it controls the display of the analysis results on the display unit 40. Also, when the output control unit 62 receives sound data from the server device 100, it controls the output of the sound data from the sound output unit 50.

[0030] (3. Configuration of Server Device 100) Next, the configuration of the server device 100 according to the embodiment will be described using Figure 4. Figure 4 is a diagram showing an example of the configuration of the server device 100 according to the embodiment. As shown in Figure 4, the server device 100 has a communication unit 110, a storage unit 120, and a control unit 130. The server device 100 may also have an input unit (for example, keywords or mouse input) that accepts various operations from the user U of the server device 100, a display unit (for example, a liquid crystal display) for displaying various information, and an audio output unit.

[0031] The communication unit 110 is connected to the network N by wire or wireless connection and transmits and receives information to and from terminal devices 10, etc., via the network N. For example, the communication unit 110 receives answers to questions about reflection from the terminal device 10 and transmits the analysis results and voice data related to the reflection to the terminal device 10. The communication unit 110 is implemented by, for example, a NIC.

[0032] The memory unit 120 stores the OS (Operating System) and various programs executed by the control unit 130. For example, the memory unit 120 stores a program that executes the information processing of the information processing method described later. Furthermore, the memory unit 120 stores various data used by the program executed by the control unit 130. For example, the memory unit 120 stores the analysis model 121, the emotion estimation model 122, and the voice generation model 123. The memory unit 120 is implemented by, for example, a semiconductor memory element such as RAM or flash memory, or a storage device such as a hard disk or optical disc. Programs and data can be used in a state where they are stored on a computer-readable recording medium, or they can be transmitted from other devices as needed, for example via a dedicated line, and used online. Examples of computer recording media include hard disks, optical discs such as DVDs, flexible disks, and semiconductor memory.

[0033] Analysis model 121 is a model trained to take answers to reflection-related questions as input and output analysis results. For example, analysis model 121 is generated by learning from the results of experts such as career consultants, coaches, and teachers analyzing answers to reflection-related questions and identifying what should be reflected upon. Note that existing AI models, such as so-called generative AI, may also be used as the analysis model.

[0034] The emotion estimation model 122 is a model that estimates emotions from text. For example, the emotion estimation model 122 performs linguistic analysis on the text, estimates emotions from the words and phrases contained in the text, and outputs the estimated emotion. The speech generation model 123, when given text and emotions as input, estimates and outputs which part of the input text corresponds to the input emotion. It should be noted that the emotion estimation model 122 may also use existing AI models such as so-called generative AI.

[0035] The speech generation model 123 is a model that generates audio data from a text document, reading the text aloud. The speech generation model 123 generates audio data according to the specified emotion. Note that the speech generation model 123 may also utilize existing AI models, such as so-called generative AI.

[0036] The control unit 130 is a controller that, for example, executes various programs stored in the storage device inside the server device 100 using RAM as a working area, using a processor such as a CPU or MPU. The control unit 130 is also implemented by an integrated circuit such as an ASIC or FPGA.

[0037] As shown in Figure 4, the control unit 130 includes a reception unit 131, an estimation unit 132, a voice generation unit 133, a voice control unit 134, a notification unit 135, an analysis unit 136, and a display control unit 137, and realizes or executes the information processing operations described below. Note that the internal configuration of the control unit 130 is not limited to the configuration shown in Figure 4, and other configurations are also acceptable as long as they perform the information processing described later.

[0038] The reception unit 131 receives various types of information from the terminal device 10. For example, the reception unit 131 receives operation information transmitted from the terminal device 10. The reception unit 131 also receives answers to reflection-related questions transmitted from the terminal device 10. Specifically, the reception unit 131 receives reflection information including the question content and answers from user U regarding their opinions, experiences, feelings, and values ​​in response to the question.

[0039] By the way, in order to conduct reflection properly, user U needs to answer questions about reflection while reflecting on their own experiences, feelings, and thoughts related to the question. However, if user U simply answers the question without sufficient reflection, they may not be able to provide appropriate feedback on their reflection.

[0040] Here, people sometimes experience different emotions or recall memories when they read aloud what they have typed. Similarly, listening to their written text read aloud can evoke different emotions or bring back forgotten memories. "Writing" is the process of putting images in the brain into words. The same images are generated when reading what has been written. On the other hand, "listening" and "reading" generate different "images."

[0041] Therefore, the server device 100 according to the embodiment provides a function to read aloud the received response in the service according to the embodiment. For example, the server device 100 provides a function to estimate the emotion of user U from the response and to read the response aloud in the voice of the estimated emotion.

[0042] The estimation unit 132 estimates user U's emotions from the reflection information. For example, the estimation unit 132 estimates user U's emotions from the answers to the reflection questions received by the reception unit 131. For example, the estimation unit 132 uses the emotion estimation model 122 to analyze the opinions, experiences, emotions, and values ​​answers and estimates which part of the experience-related answers corresponds to the emotions in the emotion-related answers. For example, the estimation unit 132 estimates the corresponding emotion from the emotions in the emotion-related answers for each sentence of the experience-related answers. Note that the emotion estimation model 122 may use an existing so-called generative AI. For example, the estimation unit 132 may input a prompt to the generative AI instructing it to estimate emotions, thereby causing the generative AI to estimate user U's emotions.

[0043] The voice generation unit 133 generates audio data that reads aloud the answers to the reflection questions received by the reception unit 131. The voice generation unit 133 uses the voice generation model 123 to generate audio data that reads aloud the answers regarding opinions, experiences, feelings, and values. The voice generation unit 133 generates audio data that reads the answers in the voice of a male, female, child, adolescent, middle-aged, or elderly person. For example, the voice generation unit 133 generates audio data that reads the answers in a voice that has the characteristics of at least one of the following: male, female, child, adolescent, middle-aged, or elderly person. For example, if the answer is to be in the voice of an "old man," the voice generation unit 133 generates audio data that reads the answer in a voice that has the characteristics of both "male" and "elderly."

[0044] Furthermore, if the estimation unit 132 has estimated an emotion, the voice generation unit 133 generates audio data that reads the answer in the voice of the estimated emotion. For example, if an emotion is estimated in any part of the answer regarding the experience, the voice generation unit 133 generates audio data that reads the part of the answer regarding the experience that is associated with an emotion in the voice of the associated emotion. For example, the voice generation unit 133 generates audio data that reads each sentence of the answer regarding the experience in the voice of the associated emotion. Note that the voice generation model 123 may use an existing so-called generative AI. For example, the voice generation unit 133 may input a prompt to the generative AI instructing it to generate audio data that reads the answer aloud, thereby causing the generative AI to generate audio data that reads the answer aloud.

[0045] The voice control unit 134 controls the output of the voice data generated by the voice generation unit 133 to the user U.

[0046] Here, let's explain a specific example. On terminal device 10, the application displays several questions related to reflection and accepts the user's selection of a question to reflect on. If any question is selected, the application displays the items of opinion, experience, feeling, and values ​​in order, according to the framework of reflection, as shown in Figure 2 (A) to (D), and prompts the user to input answers for each item in order. When the user selects that they have finished inputting their values, the application displays the selected reflection question and the answers for opinion, experience, feeling, and values.

[0047] Figure 5 shows an example of a screen 200 displayed on the terminal device 10 according to the embodiment. The screen 200 shown in Figure 5 displays the name of the responding user U, a question about reflection, and an example of the answers regarding opinions, experiences, feelings, and values. The question is, "What has recently made you angry or upset?" The opinion answer to the question is, "A comment from an old acquaintance on something I posted on Facebook." The experience answer is, "An acquaintance commented on someone's post after I added a comment and shared it, saying something like, 'No, you don't get it.' This person has been posting negative comments recently, but I never thought they would post a negative comment on my post as well. I thought about posting an aggressive comment, but after reading it carefully, it was off-topic from my original post, and what they wrote wasn't wrong, so I just liked it." The feelings answer is, "Anger," and "Acceptance." The values ​​answer is, "Various opinions," "Well, something must have bothered them," and "A big heart."

[0048] Screen 200 has a menu button 201, which consists of three dots arranged vertically, to the right of the header. When menu button 201 is selected, menu 210 is displayed. Menu 210 includes a read-aloud button 211. User U selects the read-aloud button 211 to instruct the system to read out questions and answers.

[0049] When the read-aloud button 211 is selected in the terminal device 10, the transmission unit 61 transmits the question content and the user U's answers to the question, including opinions, experiences, feelings, and values, as reflection information to the server device 100.

[0050] In the server device 100, the reception unit 131 receives the question content and the user U's responses to the question, including opinions, experiences, feelings, and values, as reflection information.

[0051] The estimation unit 132 estimates user U's emotions from the reflection information. For example, the estimation unit 132 analyzes the sentences of the experience responses among the opinions, experiences, emotions, and values ​​responses using the emotion estimation model 122, and estimates which part of the experience response corresponds to the emotion in the emotion response. For example, the estimation unit 132 estimates the corresponding emotion from the emotion in the emotion response for each sentence of the experience response. For example, the estimation unit 132 uses a generative AI as the emotion estimation model 122, and by inputting a prompt to the generative AI instructing it to estimate emotions, it causes the generative AI to estimate user U's emotions.

[0052] Figure 6 shows an example of a prompt that instructs the estimation of emotions according to the embodiment. Figure 6 shows an example of a prompt that instructs the user to estimate whether the emotion in the emotion response corresponds to the experience response among the opinion, experience, emotion, and value responses shown in Figure 5. The prompt in Figure 6 instructs the user to analyze the sentences of question, opinion, experience, emotion, and value, divide the experience into emotions, and rearrange them into the format of question, opinion, experience (emotion), and value.

[0053] Figure 7 shows an example of the output result of the generating AI according to the embodiment. Figure 7 shows an example of the output result when the prompt shown in Figure 6 is input to the generating AI. In Figure 7, the sentences of question, opinion, experience, emotion, and values ​​shown in Figure 6 are organized into the format of question, opinion, experience (emotion), and value. In addition, in the experience (emotion) section, the corresponding emotion is indicated as (emotion) for each sentence of the experience sentence.

[0054] The voice generation unit 133 generates audio data that reads aloud the reflection-related questions and answers received by the reception unit 131. The voice generation unit 133 uses the voice generation model 123 to generate audio data that reads aloud the answers to the questions, opinions, experiences, emotions, and values. The voice generation unit 133 generates audio data that reads the answers in the voice of a male, female, child, young adult, middle-aged, or elderly person. If an emotion has been estimated by the estimation unit 132, the voice generation unit 133 generates audio data that reads the answers in the voice of the estimated emotion. For example, if an emotion has been estimated for any part of the answer regarding experience, the voice generation unit 133 generates audio data that reads aloud the part of the answer regarding experience that has been associated with an emotion, in the voice of the associated emotion. For example, as shown in Figure 7, if an emotion has been associated with each sentence of the answer regarding experience, the voice generation unit 133 generates audio data that reads aloud each sentence of the answer regarding experience in the voice of the associated emotion.

[0055] Here, the emotion estimated by the estimation unit 132 and the emotion that can be specified by the speech generation model 123 may not correspond. In this case, correspondence data that associates the emotion estimated by the estimation unit 132 with the emotion that can be specified by the speech generation model 123 is prepared in advance and stored in the storage unit 120. Figure 8 is a diagram showing an example of correspondence data according to the embodiment. The speech generation model 123 is capable of specifying the emotion of the voice by a combination of emotion parameters for four emotions: happiness, enjoyment, anger, and sadness. In the example in Figure 8, the values ​​of the four emotion parameters for happiness, enjoyment, anger, and sadness are stored for each emotion. The speech generation unit 133 uses the speech generation model 123 to generate speech data that reads aloud the answers to opinions, experiences, emotions, and values. Furthermore, for sentences to which emotions are associated, such as the experience sentences in Figure 7, the speech generation unit 133 specifies the emotion parameters corresponding to the associated emotion for each sentence and generates speech data that reads aloud in the voice of the associated emotion.

[0056] The speech generation model 123 may use an existing so-called generative AI. For example, the speech generation unit 133 may input a prompt to the generative AI instructing it to generate audio data that reads the answer aloud, thereby causing the generative AI to generate audio data that reads the answer aloud.

[0057] Figure 9 shows an example of a prompt that instructs the generation of audio data according to the embodiment. In the example in Figure 9, the question about reflection is "What did you learn this week?", the opinion answer to the question is "Dialogue is important", the experience answer is "I had trouble with my partner. I got a response that was completely different from what I asked for. When I told them that it was different from what I asked, they showed me the email I sent and explained how they understood it. I realized that my writing was poor and it was being misunderstood.", the emotion answers are "Anger", "Surprise", and "Sadness", and the values ​​answers are "Work", "Sharing Challenges", and "Communication". In the prompt in Figure 9, fixed reading sections that are to be read aloud are indicated in parentheses (), and the questions, opinions, experiences, emotions, and values ​​answers to be read aloud are indicated in quotation marks ("". In the prompt in Figure 9, the person (voice) to read aloud is instructed to be randomly selected from male, female, child, young adult, middle-aged, and elderly. Note that in the prompt in Figure 9, the specification for the voice to read out emotions is omitted. When the response sentence is to be read aloud in an emotional voice, for example, the prompt will include a description that specifies the emotion to be read aloud, or a description that specifies the emotion parameter corresponding to the emotion to be read aloud, in relation to the sentence to be read aloud in an emotional voice.

[0058] The generating AI generates audio data that aligns with the prompt once a prompt is entered. For example, in the example in Figure 9, the generating AI generates audio data that reads out the question, opinion, experience, emotion, values, and fixed reading points, in accordance with the prompt, such as "The question is, what did you learn this week?..."

[0059] The voice control unit 134 controls the output of the voice data generated by the voice generation unit 133 to the user U.

[0060] Figure 10 shows an example of a screen 200 displayed on the terminal device 10 according to the embodiment. When the read-aloud button 211 shown in Figure 5 is selected and audio data is generated, a play button 202 that instructs playback of the audio data is displayed on the screen 200. When user U wants to play the audio data, they select the play button 202. When the play button 202 is selected in the terminal device 10, the transmission unit 61 sends the operation information that the play button 202 has been selected to the server device 100.

[0061] In the server device 100, the reception unit 131 receives operation information indicating that the play button 202 has been selected. When the audio control unit 134 receives operation information indicating that the play button 202 has been selected from the reception unit 131, it outputs the audio data generated by the audio generation unit 133 to the terminal device 10.

[0062] In the terminal device 10, the output control unit 62 controls the output of the audio data received from the server device 100 from the audio output unit 50.

[0063] This allows user U to hear their own opinions, experiences, feelings, and values ​​expressed in audio form in response to reflection-related questions. Even when user U has typed their own answers in text, listening to them aloud can evoke different emotions or recall memories compared to when they were writing, allowing them to reflect on their own experiences, feelings, and thoughts in response to the questions. For example, when user U listens to sentences associated with emotions in audio form for answers about experiences, it becomes easier to recall memories and reflect on experiences, feelings, and thoughts. Furthermore, even when user U has typed their own answers in text, listening to them aloud allows them to verify whether the answers appropriately represent their own experiences, feelings, and thoughts at the time. For example, when user U listens to sentences associated with emotions in audio form for answers about experiences, they can verify whether the answer corresponds to the experience and emotions they remember at the time.

[0064] In this way, the server device 100 generates audio data that reads aloud the answers to the reflection questions and controls the output of the audio data to user U, thereby encouraging user U to reflect. Furthermore, by generating audio data that reads aloud the answers to the reflection questions and controlling the output of the audio data to user U, the server device 100 allows user U to confirm whether the answers are appropriate.

[0065] The terminal device 10 allows users to re-enter their answers for each item—opinions, experiences, feelings, and values—by returning to the screen shown in Figure 2(D). When returning to the screen shown in Figure 2(D), the previously entered answers are displayed in the answer fields for each item on the screen.

[0066] If user U listens to the audio recording of the answers to the reflection questions and finds the answers insufficient, they can operate the terminal device 10 to return to, for example, the screen shown in Figure 2(D), and re-enter the answers for each item, such as opinions, experiences, feelings, and values, by reviewing or adding to them. When input completion is selected on the screen shown in Figure 2(D), the terminal device 10 displays the questions and the re-entered answers for opinions, experiences, feelings, and values, as shown in Figure 5. User U can listen to the re-entered answers by selecting the read-aloud button 211 and then the play button 202 that appears.

[0067] Incidentally, the estimation unit 132 estimates the corresponding emotion from the emotion in the emotion response for each sentence in the experience response. However, if the experience response contains sentences expressing emotions not specified in the emotion response, or if the content of the sentences in the experience response is unclear, the emotion estimated for the experience response may differ from the actual emotion of user U. In this case, it may not be possible to provide appropriate feedback on the reflection to user U.

[0068] Therefore, if the emotion estimated from the experience-related response by the estimation unit 132 differs from the emotion received by the reception unit 131, the notification unit 135 notifies the user U to revise their response. For example, the estimation unit 132 estimates an emotion for each sentence in the experience-related response. If the emotion estimated by the estimation unit 132 from a sentence in the experience-related response is not among the emotions received by the reception unit 131, the notification unit 135 controls the terminal device 10 to display a message indicating that there is a sentence in the experience-related response that does not correspond to the emotion in the emotion-related response.

[0069] When a message is displayed on the terminal device 10, user U operates the terminal device 10 to return to, for example, the screen shown in Figure 2(D), and revise either or both of the answers to the experience item and the emotion item so that the sentences in the experience answer and the emotions in the emotion answer correspond.

[0070] In this way, the server device 100 can assist user U in properly performing reflection by prompting user U to revise their response so that the sentences in the experience response correspond to the emotions in the emotion response.

[0071] Once user U confirms that the answer is appropriate, they perform a predetermined operation to instruct terminal device 10 to perform the analysis.

[0072] When a predetermined operation to instruct analysis is performed on the terminal device 10, the transmission unit 61 transmits the question content and the responses from user U to each item of opinion, experience, emotion, and values ​​to the question as reflection information to the server device 100.

[0073] In the server device 100, the reception unit 131 receives the question content and the user U's responses to the question, including opinions, experiences, feelings, and values, as reflection information.

[0074] The analysis unit 136 analyzes the reflection information and generates the analysis results. For example, with respect to the reflection information received by the reception unit 131, the analysis unit 136 generates analysis results that analyze at least one of user U's opinions, experiences, feelings, and values ​​from the perspective of a professional such as a career consultant, coach, or teacher, or an equivalent perspective, depending on the purpose of the analysis. For example, the analysis unit 136 uses the analysis model 121 to analyze the opinions, experiences, feelings, and values ​​and generates the analysis results. To give one example, the analysis unit 136 generates the analysis results using the analysis model 121 stored in the memory unit 120. For example, the analysis unit 136 inputs the reflection information into the analysis model 121 and obtains the analysis results output from the analysis model 121. Note that the analysis model 121 may use an existing so-called generative AI. For example, the analysis unit 136 inputs a prompt indicating the analysis results to be output and the reflection information into the analysis model 121, causing the analysis model 121 to generate the analysis results.

[0075] The display control unit 137 controls the display of the analysis results to the user U. For example, the display control unit 137 controls the display of the analysis results on the terminal device 10 as a bulleted list or as a reflection card. For example, the display control unit 137 displays the user U's name, opinions, experiences, feelings, and values ​​in a table as a reflection card.

[0076] In this way, the server device 100 can provide appropriate feedback on reflection to user U by analyzing the user U's responses to questions about reflection, after the user U has reflected on the matter and confirmed that the content is appropriate.

[0077] (4. Flowchart of information processing) Next, using Figure 11, the information processing procedure of the information processing method according to the embodiment performed by the server device 100 will be described. Figure 11 is a flowchart of an example of information processing according to the embodiment. The information processing shown in Figure 11 is executed when the read-aloud button 211 is selected and reflection information is received by the server device 100.

[0078] The reception unit 131 receives responses to user U's reflection questions. For example, the reception unit 131 receives the question content and responses from user U regarding their opinions, experiences, feelings, and values ​​in relation to the question as reflection information (step S10).

[0079] The estimation unit 132 estimates the user U's emotions from the reflection information (step S11). For example, the estimation unit 132 analyzes the sentences of the experience responses among the opinions, experiences, emotions, and values ​​responses using the emotion estimation model 122, and estimates which part of the experience response corresponds to the emotion in the emotion response. For example, the estimation unit 132 estimates the corresponding emotion from the emotion in the emotion response for each sentence of the experience response.

[0080] The voice generation unit 133 generates audio data that reads aloud the reflection-related questions and answers received by the reception unit 131 (step S12). The voice generation unit 133 uses the voice generation model 123 to generate audio data that reads aloud the answers to the questions, opinions, experiences, emotions, and values. The voice generation unit 133 generates audio data that reads the answers in the voice of a male, female, child, young adult, middle-aged, or elderly person. If an emotion has been estimated by the estimation unit 132, the voice generation unit 133 generates audio data that reads the answers in the voice of the estimated emotion.

[0081] The audio control unit 134 determines whether the play button 202 has been selected (step S13). If the play button 202 has not been selected (step S13: No), the system proceeds back to step S13.

[0082] On the other hand, when the play button 202 is selected (step S13: Yes), the audio control unit 134 controls the output of the audio data generated by the audio generation unit 133 to the user U (step S14), and then terminates the process.

[0083] (5. Effects) The server device 100 according to this embodiment includes a reception unit 131, a voice generation unit 133, and a voice control unit 134. The reception unit 131 receives the user U's answer to a question about reflection. The voice generation unit 133 generates voice data that reads aloud the answer received by the reception unit 131. The voice control unit 134 controls the output of the voice data generated by the voice generation unit 133 to the user U. This allows the server device 100 to encourage the user U to reflect. The server device 100 also allows the user U to confirm whether the answer is appropriate.

[0084] Furthermore, the voice generation unit 133 generates audio data in which the answer is read aloud in the voice of a male, female, child, young adult, middle-aged, or elderly person. This allows the server device 100 to let user U hear the answer in the voice of another person, such as a male, female, child, young adult, middle-aged, or elderly person. By hearing the answer in someone else's voice rather than their own, user U can calmly reflect on their answer.

[0085] Furthermore, the server device 100 according to this embodiment further includes an estimation unit 132. The estimation unit 132 estimates the emotions of user U from the answers received by the reception unit 131. The voice generation unit 133 generates voice data that reads out the answers in the voice of the emotions estimated by the estimation unit 132. As a result, the server device 100 can make user U listen to the answers in the voice of the emotions estimated from the answers. By listening to the answers in the voice of the emotions estimated from the answers, user U can more easily recall memories of that time and reflect on experiences, emotions, and thoughts.

[0086] Furthermore, the reception unit 131 accepts input of answers to questions, including answers to multiple items, such as experiences and emotions. The estimation unit 132 estimates which parts of the experience-related answers received by the reception unit 131 correspond to the emotions received by the reception unit 131. The voice generation unit 133 generates voice data that reads aloud the parts of the experience-related answers that have been associated with emotions, using the voice of the associated emotion. As a result, the server device 100 can allow user U to hear the parts of the answers that have been associated with emotions, using the voice of the associated emotion. By listening to the parts of the answers that have been associated with emotions, user U can more easily recall memories from that time and reflect on their experiences, emotions, and thoughts. In addition, user U can confirm whether the answers represent the experiences, emotions, and thoughts from that time.

[0087] Furthermore, the reception unit 131 accepts input of answers to questions, including multiple items related to experience and emotion. The estimation unit 132 estimates the corresponding emotion for each sentence of the experience-related answer received by the reception unit 131, based on the emotion received by the reception unit 131. The voice generation unit 133 generates voice data that reads aloud in the voice of the corresponding emotion for each sentence of the experience-related answer that has been associated with an emotion. As a result, the server device 100 can make user U listen to each sentence of the experience-related answer in the voice of the corresponding emotion. User U can reflect on the experience, emotion, and thoughts at the time for each sentence of the answer. User U can also confirm whether the answer represents user U's own experience, emotion, and thoughts at the time. For example, user U can confirm whether the experience-related answer corresponds to the experience and emotion in their memory by listening to the sentence associated with the emotion in the voice of the corresponding emotion.

[0088] Furthermore, the server device 100 according to this embodiment further includes a notification unit 135. If the emotion estimated from the experience-related answers as a result of the estimation unit 132 differs from the emotion received by the reception unit 131, the notification unit 135 notifies the user U to revise their answer. Upon notification, the user U revises either or both of their answers to the experience-related items and / or the emotion-related items. In this way, the server device 100 can assist the user U in appropriately conducting reflection.

[0089] In the above embodiment, we have described a case in which the voice generation unit 133 generates voice data in which the answer is read aloud in the voice of a male, female, child, young adult, middle-aged, or elderly person. However, the processing of the disclosed server device 100 is not limited to this. The voice generation unit 133 may determine the type of voice to read the answer according to the attributes of user U, the content of the answer, and the emotion estimated from the answer. For example, the voice generation unit 133 may generate voice data in which the answer is read aloud in the voice of a person older than user U. When user U hears the answer read aloud in the voice of a person older than user U, they may feel as if they are being advised by an older person, and be able to reflect on the situation calmly. Also, for example, the voice generation unit 133 may generate voice data in which an answer that is estimated to evoke a negative emotion, such as sadness, is read aloud in the voice of a child. When an answer that is estimated to evoke a negative emotion is read aloud in the voice of a child, user U may feel empathy for the answer, and be able to reflect on the situation of the answer more deeply. Also, for example, the voice generation unit 133 may change the voice to read aloud depending on the estimated emotion. Furthermore, for example, the voice generation unit 133 may change the speed, intonation, and pitch of the voice it reads aloud for each estimated emotion. Also, for example, the voice generation unit 133 may generate audio data in which background music (BGM) corresponding to the emotion estimated from the answer is played along with the voice of the answer.

[0090] Furthermore, in the above embodiment, when the read-aloud button 211 on the screen 200 is selected, the server device 100 generates audio data that reads the answer aloud, and when the play button 202 on the screen 200 is selected, the server device 100 controls the output of the audio data to the user U. However, this disclosure is not limited thereto. For example, when the read-aloud button 211 on the screen 200 is selected, the server device 100 generates audio data that reads the answer aloud, and then controls the output of the generated audio data to the user U.

[0091] Furthermore, in the above embodiment, we described a case in which the items of opinion, experience, emotion, and values ​​are displayed in order, as shown in Figures 2(A) to (D), and the user is prompted to input answers for each item. However, this disclosure is not limited to this. For example, the screen shown in Figure 2(D) may be displayed, allowing the user to input answers for each of the items of opinion, experience, emotion, and values.

[0092] Furthermore, the above embodiment described a case in which the server device 100 estimates the emotions of user U from the response and generates voice data. However, this disclosure is not limited thereto. For example, the control unit 60 of the terminal device 10 may have the functions of a reception unit 131, an estimation unit 132, a voice generation unit 133, a voice control unit 134, a notification unit 135, an analysis unit 136, and a display control unit 137, and the terminal device 10 may perform the estimation of the emotions of user U from the response and generate voice data. In this case, the terminal device 10 corresponds to the information processing apparatus according to the present application.

[0093] (6. Other) [6.1. Hardware Configuration] Furthermore, the server device 100 according to the embodiment described above is realized by a computer 1000 having a configuration such as that shown in Figure 12. Figure 12 is a hardware configuration diagram showing an example of a computer that realizes the functions of the server device 100. The server device 100 according to the embodiment will be described below as an example. The computer 1000 is connected to an output device 1010 and an input device 1020, and an arithmetic unit 1030, a cache 1040, a memory 1050, an output interface 1060, an input interface 1070, and a network interface 1080 are connected by a bus 1090.

[0094] The arithmetic unit 1030 operates based on programs stored in the cache 1040 and memory 1050, as well as programs read from the input device 1020, and executes various processes. The cache 1040 is a cache that temporarily stores data used by the arithmetic unit 1030 for various calculations, such as RAM. The memory 1050 is a storage device in which data used by the arithmetic unit 1030 for various calculations and various databases are registered, and is implemented as ROM (Read Only Memory), HDD (Hard Disk Drive), flash memory, etc.

[0095] The output IF1060 is an interface for transmitting information to be output to output devices 1010 that output various types of information, such as monitors and printers. This interface may be implemented using connectors of standards such as USB (Universal Serial Bus), DVI (Digital Visual Interface), or HDMI (High Definition Multimedia Interface). On the other hand, the input IF1070 is an interface for receiving information from various input devices 1020, such as mice, keyboards, and scanners. This interface may be implemented using USB, for example.

[0096] For example, the input device 1020 may be implemented by a device that reads information from optical recording media such as CDs (Compact Discs), DVDs (Digital Versatile Discs), PDs (Phase Change Rewritable Disks), magneto-optical recording media such as MOs (Magneto-Optical disks), tape media, magnetic recording media, or semiconductor memory. Alternatively, the input device 1020 may be implemented by an external storage medium such as a USB memory stick.

[0097] The network IF1080 has the function of receiving data from other devices via network N and sending it to the arithmetic unit 1030, and also transmitting data generated by the arithmetic unit 1030 to other devices via network N.

[0098] Here, the arithmetic unit 1030 controls the output device 1010 and the input device 1020 via the output IF 1060 and the input IF 1070. For example, the arithmetic unit 1030 loads a program from the input device 1020 or memory 1050 onto the cache 1040 and executes the loaded program. For example, if computer 1000 functions as server device 100, the arithmetic unit 1030 of computer 1000 will realize the functions of the control unit 130 by executing the program loaded onto the cache 1040.

[0099] The embodiments of the present application have been described in detail above with reference to the drawings. However, these are illustrative examples, and the embodiments of the present application can be implemented in various forms with modifications and improvements based on the knowledge of those skilled in the art, starting with the embodiments described in the disclosure section of the invention.

[0100] [6.2. Other] Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically by known methods. In addition, the processing procedures, specific names, and various data and parameters shown in the above document and drawings can be changed at will unless otherwise specified. For example, the various information shown in each figure is not limited to the information shown.

[0101] Furthermore, each component of the illustrated device is a functional concept and does not necessarily have to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions. In addition, the embodiments described above can be combined as appropriate, as long as the processing content is not contradictory. [Explanation of Symbols]

[0102] 1. Information Processing System 10 Terminal devices 20 Communications Department 30 Input section 40 Display section 50 Audio output section 60 Control Unit 61 Transmitter 62 Output Control Unit 100 Server Devices 110 Communications Department 120 Storage section 121 Analysis Models 122 Emotion Estimation Models 123 Speech Generation Models 130 Control Unit 131 Reception Department 132 Estimation Department 133 Voice generation unit 134 Audio Control Unit 135 Notification Department 136 Analysis Department 137 Display Control Unit N Network U User

Claims

1. A reception desk that accepts answers to user reflection questions, A voice generation unit generates voice data that reads aloud the response received by the reception unit, A voice control unit that controls the output of the voice data generated by the voice generation unit to the user, An information processing device characterized by having the following features.

2. The voice generation unit generates audio data that reads out the answer in one of the following voices: male, female, child, young adult, middle-aged, or elderly. The information processing apparatus according to claim 1.

3. The system further includes an estimation unit that estimates the user's emotions from the responses received by the reception unit, The voice generation unit generates voice data that reads out the answer using the emotion estimated by the estimation unit. The information processing apparatus according to claim 1.

4. The aforementioned reception unit accepts input of answers to questions, including multiple items such as experiences and emotions. The estimation unit estimates which part of the response regarding the experience received by the reception unit corresponds to the emotion received by the reception unit. The voice generation unit generates audio data in which it reads aloud the portion of the response regarding the experience that is associated with an emotion, using the voice corresponding to that emotion. The information processing apparatus according to claim 3.

5. The aforementioned reception unit accepts input of answers to questions, including multiple items such as experiences and emotions. The estimation unit estimates the corresponding emotion from the emotions received by the reception unit for each sentence of the response regarding the experience received by the reception unit. The voice generation unit generates audio data that reads aloud each sentence in the response text about the experience that is associated with an emotion, using the voice corresponding to the emotion. The information processing apparatus according to claim 4.

6. If the estimation result of the estimation unit differs from the emotion estimated from the response regarding the experience, the system further includes a notification unit that notifies the user to revise the response. The information processing apparatus according to claim 4.

7. On the computer, It generates audio data that reads aloud the answers to the user's reflection-related questions. Control is performed to output the generated audio data to the user. An information processing method characterized by executing a process.

8. It generates audio data that reads aloud the answers to the user's reflection-related questions. Control is performed to output the generated audio data to the user. An information processing program characterized by having a computer perform the processing.

Citation Information

Patent Citations

  • Information processing device, information processing method and information processing program

    JP2023110739A