Information processing system and information processing method
The information processing system addresses the lack of emotional rhythm in phoneme sequences by separating frequency bands for biometric information analysis, enabling emotional expression in output information, thus improving communication accuracy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2024-11-25
- Publication Date
- 2026-06-04
AI Technical Summary
Existing technologies fail to accurately convey emotional rhythm in phoneme sequences estimated from biometric information, leading to potential miscommunication due to slight intonation differences.
An information processing system that acquires biometric information from body movements, separates high-frequency components for phoneme estimation and low-frequency components for emotion estimation, and generates output information with emotional expression using trained models based on neural networks.
Enables the output of audio or textual information with emotional expression, accurately reflecting user emotions, thereby enhancing communication clarity.
Smart Images

Figure 2026091491000001_ABST
Abstract
Description
Technical Field
[0005]
[0001] The present invention relates to an information processing system that converts biometric information related to the movement of a user's body part, such as a speech operation, into character information or voice information and outputs it.
Background Art
[0002] In recent years, it has been attempted to estimate the speech content from the movement of a user's mouth. For example, biometric information indicating the movement of a user's mouth is acquired, and a phoneme estimation result based on the biometric information is output (for example, Non-Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In Non-Patent Document 1, phoneme sequences are estimated from biometric information based on a user's non-vocal speech, and the speech content of the user is recognized. However, although rhythm based on the user's emotion can be conveyed in the case of voice, in the case of phonemes estimated from biometric information other than voice, such as mouth movement, the rhythm based on emotion is missing, and it becomes a mere phoneme sequence. Therefore, in daily life, there is a risk that a slight difference in intonation may greatly change the impression or cause an obstacle to communication.
[0005] An object of the present invention is to provide an information processing device capable of outputting voice information or character information with emotional expression based on biometric information related to the movement of a user's body part. However, the problems that the embodiments disclosed in this specification and drawings aim to solve are not limited to those described above. Problems corresponding to the effects of each configuration shown in the embodiments described later can also be positioned as other problems. [Means for solving the problem]
[0006] To achieve the above objective, the information processing system according to the present invention is An acquisition unit that acquires biological information regarding the movement of the user's body parts, A language estimation unit estimates language information, including textual information or audio information, based on the biometric information acquired by the acquisition unit, Based on the biometric information acquired by the acquisition unit, an emotion estimation unit estimates the user's emotions, A generation unit generates output information based on the language information estimated by the language estimation unit and the emotion estimated by the emotion estimation unit. The device is characterized by having an output unit that outputs the output information generated by the generation unit. [Effects of the Invention]
[0007] According to the present invention, it is possible to output audio or textual information accompanied by emotional expression based on biological information relating to the movement of the user's body parts. [Brief explanation of the drawing]
[0008] [Figure 1] A diagram showing an example of the functional configuration of an information processing system according to the first embodiment. [Figure 2] An example of user behavior detected by the information processing system according to the first embodiment. [Figure 3] An example of emotion information detected by the information processing system according to the first embodiment and the resulting voice processing. [Figure 4] An example of emotion information detected by the information processing system according to the first embodiment and its text processing. [Figure 5] A diagram showing the configuration of the information processing system according to Example 1. [Figure 6]A flowchart illustrating the operation according to Example 1. [Modes for carrying out the invention]
[0009] The embodiments will be described in detail below. The present invention is not limited to the embodiments shown below, unless it exceeds the essence of the present invention.
[0010] [First Embodiment] Figure 1 shows the functional configuration of the information processing system according to the first embodiment.
[0011] The information processing system 1000 mainly consists of a detection device 100 and an information processing device 101. The detection device 100 detects biometric information based on muscle movements caused by the user's speech. The information processing device 101 acquires the biometric information detected by the detection device 100 and converts it into text information or voice information. The detection device 100 has a biometric information detection unit 102 that detects biometric information from one or more locations on the user, and a transmission unit 104 that transmits the biometric information detected by the biometric information detection unit 102 to the information processing device 101. Note that the detection device 100 and the information processing device 101 may each be composed of separate devices connected by a network, or they may be composed of an integrated device that has the functionality of a detection device as an information processing device.
[0012] Figure 1 shows a configuration in which the detection device 100 is a single device, but it may consist of multiple devices, and multiple biometric information detection units 102 may be configured with each detection device. The detection device 100 can also be described as an acquisition unit that acquires biometric information related to the movement of the user's body parts.
[0013] The biometric information detection unit 102 consists of sensors that detect biometric information related to the user's muscle movements, skin movements, tongue movements, etc. The biometric information detection unit 102 is at least one of the following: an electromyography sensor, an acceleration sensor, an ultrasonic sensor, a tactile sensor, an optical sensor, a pressure sensor, etc. However, the biometric information detection unit 102 may be composed of sensors other than those listed above, as long as they can detect biometric information. The biometric information detection unit 102 is preferably installed in areas such as the neck, lower jaw, around the mouth, or temples in order to detect biometric information related to the user's mouth and tongue movements. The biometric information detection unit 102 may be installed in areas other than those listed above, as long as it is a location where biometric information related to mouth and tongue movements can be detected.
[0014] The information processing device 101 includes a receiving unit 104, a first extraction unit 105, and a second extraction unit. The receiving unit 104 receives biological information transmitted from the detection device 100. The first extraction unit 105 extracts high-frequency components of a first frequency band from the biological information received by the receiving unit 104. The second extraction unit extracts low-frequency components of a second frequency band, which is a lower frequency band than the first frequency band, from the biological information received by the receiving unit 104.
[0015] The information processing device 101 also includes a phoneme learning unit 107, an emotion learning unit 108, a phoneme estimation unit 109, and an emotion estimation unit 110. The phoneme learning unit 107 learns the correspondence between the high-frequency components of the biometric information extracted by the first extraction unit 105 and the phoneme sequence (linguistic information). The emotion learning unit 108 learns the correspondence between the low-frequency components of the biometric information extracted by the second extraction unit 106 and emotions. The phoneme estimation unit 109 estimates the linguistic information, which is the phoneme sequence, from the high-frequency components of the biometric information. The emotion estimation unit 110 estimates emotions from the low-frequency components of the biometric information.
[0016] Further, the information processing apparatus 101 has a generation unit 111 that generates output information including voice information and character information with emotion based on the estimated language information and emotion. The generation unit 111 includes a voice generation unit 112 that generates voice information and a character generation unit 113 that generates character information. Also, the information processing system 1000 has a voice information output unit 114 that outputs the voice information generated by the generation unit 111 and a display unit 115 that outputs character information.
[0017] The information processing apparatus 101 receives biometric information from the detection device 100 by the reception unit 104. After reception, the first extraction unit 105 and the second extraction unit 106 extract high-frequency components and low-frequency components. The high-frequency components are in the first frequency band, and the low-frequency components are in the second frequency band that is a frequency band lower than the first frequency band. As a method for extracting frequency components, it may be extracted by passing the biometric information through a frequency filter, or may be extracted by dividing it into frequency bands after converting it into a spectrogram or the like by fast Fourier transform.
[0018] Since the user's emotion appears mainly in the low-frequency components in the user's body movement, it is preferable to use the low-frequency components 106 for emotion estimation. The separation between high frequency and low frequency is preferably set generally between 10 Hz and 100 Hz. Also, in FIG. 1, the first extraction unit 105 and the second extraction unit 106 are described separately, but it may be configured to extract the high-frequency components and the low-frequency components separately by one extraction unit having both functions of the first extraction unit 105 and the second extraction unit 106. The high-frequency components 105 are used for estimating the arrangement of phonemes, but it is sufficient if at least the high-frequency components 105 are included, and the low-frequency components 106 may be further included.
[0019] As a method for estimating the correspondence between high-frequency and low-frequency components extracted from biological information and the corresponding phoneme sequences and emotions, for example, a trained model based on an architecture composed of a neural network is used. The information processing device 101 has a storage unit (not shown) for storing the trained model. The phoneme estimation unit 109 and the emotion estimation unit 110 have the function of performing inference using the trained model.
[0020] The trained models are those generated using deep learning techniques such as CNNs (Convolutional Neural Networks) and RNNs (Recurrent Neural Networks). In addition to models derived from CNNs and RNNs, other machine learning techniques such as support vector machines, logistic regression, and random forests may also be used, as may rule-based methods.
[0021] To create a trained model for use in the phoneme estimation unit 109, the high-frequency components extracted by the first extraction unit 105 are sent to the phoneme learning unit 107 to learn the correspondence between phonemes and biometric information. Similarly, to create a trained model for use in the emotion estimation unit 108, the low-frequency components extracted by the second extraction unit 106 are sent to the emotion learning unit 108 to learn the correspondence between emotions and biometric information. In this embodiment, an example is described in which the information processing device 101 has the phoneme learning unit 107 and the emotion learning unit 108, but the phoneme learning unit 107 and the emotion learning unit 108 may be configured to be located outside the information processing device 101. Furthermore, the phoneme learning unit 107 and the emotion learning unit 108 located outside the information processing device 101 may perform their learning on the cloud.
[0022] After a series of learning processes have generated a trained model, the high-frequency and low-frequency components contained in the biometric information are sent to the phoneme estimation unit 109 and the emotion estimation unit 110, respectively. Then, phoneme estimation and emotion estimation are performed, and phoneme information and emotion information can be obtained.
[0023] The phoneme estimation unit 109 uses a trained model learned by the phoneme learning unit 107 to estimate linguistic information consisting of a sequence of phonemes based on biometric information. More specifically, the high-frequency components of the biometric information extracted by the first extraction unit are input to the trained model, and the linguistic information estimated from the trained model is output. The estimated linguistic information includes character information and speech information. In other words, the phoneme estimation unit 109 can be described as a language estimation unit that estimates linguistic information including character information or speech information based on biometric information.
[0024] The emotion estimation unit 110 estimates the user's emotion based on biometric information using a trained model learned by the emotion learning unit 108. More specifically, the low-frequency components of the biometric information extracted by the second extraction unit are input to the trained model, and the estimated user emotion is obtained as the output of the trained model.
[0025] The relationship between user emotions and body movements resulting from user actions is explained using Figure 2. In Figure 2, the types of emotions are classified into groups 1 to 3. Note that the number of groups is illustrative, and more detailed classifications are possible.
[0026] Group 1 includes positive emotions such as joy, pleasure, greeting, understanding, and agreement. When a user experiences joy or pleasure, their body movements may manifest as short vibrations throughout the body or vibrations associated with suppressing laughter. Similarly, when a user feels the urge to greet, their body movements may include actions such as lowering their head forward. When a user experiences understanding or agreement, their body movements may manifest as short, up-and-down vibrations, such as those associated with nodding or giving a thumbs-up.
[0027] Group 2 includes negative emotions such as frustration, anger, sadness, and pity. For example, frustration and anger may manifest as physical movements such as banging on a table, stomping feet, or shaking knees. Sadness and pity may manifest as physical movements such as looking up at the sky or bowing one's head.
[0028] Group 3 includes emotions related to having doubts. For example, when a doubt arises, it manifests as physical movements such as tilting the head to one side or the other.
[0029] Body movements associated with the emotions described above are detected by the sensor of the biometric information detection unit 102 in the information processing system configuration diagram of Figure 1, and acquired as biometric information. The acquired biometric information is sent to the receiving unit 104 of the information processing device 101. Low-frequency components are extracted from the biometric information received by the receiving unit 104 by the second extraction unit 106. The low-frequency components of the extracted biometric information are then sent to the emotion learning unit 108, where the correspondence between the emotions of groups 1 to 3 and body movements described above is learned.
[0030] The emotion estimation unit 110 uses the trained model, which has undergone the above-described learning process, to estimate emotions. The low-frequency components of the biometric information extracted by the second extraction unit 106 are input to the trained model, and it is estimated which of the groups 1 to 3 the emotion corresponds to. The estimated emotion is sent to the generation unit 111 as emotion information.
[0031] While there may be variations in behavior depending on the individual and the time, the accuracy of estimation can be improved by learning to account for these variations.
[0032] Next, we will explain the relationship between emotional information and the processing of corresponding auditory information using Figure 3. Similar to Figure 2, Figure 3 classifies emotions into groups 1 through 3. Specifically, joy, pleasure, greetings, understanding, and agreement are classified as positive emotions in Group 1. Irritation, anger, sadness, and pity are classified as negative emotions in Group 2. Questions are classified in a separate group, Group 3.
[0033] The generation unit 111 uses the estimation results of linguistic information (arrangement of phonemes) from the phoneme estimation unit 109, which acts as a language estimation unit, and the emotion estimation results (emotion information) estimated by the emotion estimation unit 110 to generate speech and / or text with emotional expression. Specifically, it uses the emotion estimation results to process linguistic information including speech (speech information) and / or text (text information) to generate linguistic information with emotional expression.
[0034] The processing of the audio information is carried out according to the grouping described above, as follows: For positive emotions in Group 1, the audio is processed with a raised tone; for negative emotions in Group 2, the audio is processed with a lower tone; and for questions in Group 3, the audio is processed with a rising intonation at the end of the word. This processing of audio information is performed by the audio generation unit 112. The tone is adjusted by adjusting the pitch, length of the sound, and the intensity of the sound. All three can be adjusted, or one or two of them can be selected and adjusted.
[0035] For example, the toned-up meter adjustment in Group 1, as mentioned earlier, is expressed by raising the pitch overall, shortening the length of the sounds, and adding dynamics to the sounds. Similarly, the toned-down meter adjustment in Group 2, as mentioned earlier, is expressed by lowering the pitch overall, lengthening the length of the sounds, and reducing the dynamics to the sounds. In the case of questions in Group 3, the adjustment is made by raising the pitch only at the end of the word.
[0036] Next, the relationship between emotional information and the processing of textual information will be explained using Figure 4. In Figure 4, emotions are classified into groups 1 to 3, similar to Figures 2 and 3. The number of groups is illustrative, and more detailed classifications are possible. In the example in Figure 4, joy, pleasure, greetings, understanding, and agreement are classified as positive emotions in Group 1. Irritation, anger, sadness, and pity are classified as negative emotions in Group 2. Doubt is classified in a different group, Group 3.
[0037] In the character generation unit 113 of the generation unit 111, the estimation results of linguistic information (arrangement of phonemes) from the phoneme estimation unit 109, which acts as a language estimation unit, and the emotion estimation results estimated by the emotion estimation unit 110 are used to generate characters with emotional expressions. Specifically, the character generation unit 113 uses the emotion estimation results to process the linguistic information including characters, thereby generating linguistic information with emotional expressions.
[0038] For example, in the case of positive emotions in Group 1, the text color can be changed to a warm color, the font can be changed to a Gothic / sans-serif typeface, or cheerful text or smiling emojis can be added, and these can be combined as appropriate to convey positive emotions. Similarly, in the case of negative emotions in Group 2, the text color can be changed to a cool color, the font can be changed to a Mincho / serif typeface, or text or emojis representing sadness or anger can be added, and these can be combined as appropriate to convey negative emotions. In the case of questions in Group 3, question marks or emojis representing questions can be added. Additionally, the background color of the area where text is displayed can be changed, symbols can be added, and icons can be added.
[0039] The information processing device 101 can be a smartphone, personal computer (PC), tablet PC, etc., but is not limited to these. If the information processing device 101 is, for example, a personal computer, the audio information processed by the generation unit 111 is transmitted to the audio information output unit 114. The audio information output unit 114 is a speaker or the like, and can play back the audio information.
[0040] The character information processed by the generation unit 111 is transmitted to the display unit 115. The display unit 115 is a display or the like and can display the character information. The audio information output unit 114 and the display unit 115 only need to be able to receive the audio information and / or character information transmitted from the generation unit 111 and may be installed separately from the information processing device. Such a configuration is necessary, for example, when sending audio information and / or character information to a person in a remote location.
[0041] The above describes a method for estimating phonemes and emotions from biometric information and generating and outputting speech and / or text using emotional information. The timing for acquiring biometric information to estimate the user's emotions can be either when the user is speaking or when the communication partner (dialogue partner) is speaking and the user is listening. The former is a means of capturing the emotional shifts of the user when they are speaking. The latter is a means of capturing the emotional shifts of the user when they are listening to the other person's speech. Depending on the purpose, it may be possible to switch between estimating emotions using biometric information at the time the user is speaking or when the other person is speaking.
[0042] The timing of communication between a user and another party can be measured using information correlated with each utterance. Specific methods are illustrated below.
[0043] First, when one party is speaking, speech can be detected by the presence or absence of a signal capturing vibrations from the voice using a microphone or bone conduction microphone, allowing for the determination of the start and end of speech. When silent speech occurs, speech can be detected by the presence or absence of output from a sensor that acquires biometric information. Alternatively, the high-frequency portion can be extracted and used for speech detection. Furthermore, the results of phoneme estimation can also be used for speech detection.
[0044] [Example 1] The information processing system according to Example 1 will now be described. The configuration of the information processing system in Example 1 is shown in Figure 5. A flowchart of the processing of the information processing system in Example 1 is shown in Figure 6.
[0045] The detection device 100 is a headset-type device having detection units in a part of the headset. The detection device 100 consists of a transmitter 1601, a biometric information detection unit 1602, a sound information detection unit 1603, and earphones 1604 and 1605. The information processing device 101 consists of a smartphone.
[0046] The biometric information detection unit 1602 is a 3-axis accelerometer. Using the accelerometer, the biometric information detection unit 1602 can detect the acceleration when the user moves their mouth. The biometric information detection unit 1602 can also use the accelerometer to detect body movements associated with the user's emotional changes. The sound information detection unit 1603 is a microphone that detects sounds emitted by the user and is used for voice-based conversations, etc.
[0047] The transmitting unit 1601 is connected to a biometric information detection unit 1602 and a sound information detection unit 1603. The transmitting unit 1601 can transmit to the outside the biometric information detected by the biometric information detection unit 1602 and the sound information detected by the sound information detection unit 1603.
[0048] By powering on the detection device 100, the detection device 100 is connected to the information processing device 101. Subsequently, the biological information detected by the biological information detection unit 1602 and the sound information detected by the sound information detection unit 1603 are transferred to the information processing device 101 by the transmission unit 1601.
[0049] The receiving unit 104 within the application of the information processing device 101 receives the biometric information detected by the biometric information detection unit 1602. The high-frequency components extracted from the detected biometric signal are sent to the phoneme estimation unit 109, where linguistic information, which is a sequence of phonemes, is estimated. The low-frequency components extracted from the detected biometric signal are sent to the emotion estimation unit 110, where emotions are estimated. The estimated linguistic information and emotion information, which are sequences of phonemes, are sent to the generation unit 111, where the linguistic information and emotion information are combined to generate speech with appropriate prosody, or characters suitable for appropriate display. The generated speech is output from the speech information output unit 114, for example, the speaker of the other party's information processing device, such as a smartphone. The processed characters are displayed from the display unit 115, for example, on the other party's information processing device, such as a smartphone.
[0050] Figure 6 illustrates the processing flow in the information processing system.
[0051] In step S100, the biological information detection unit 102 is used to detect biological information related to the movement of the user's body parts. The detected information is then transmitted to the information processing device 101 by the transmission unit 103 (step S101).
[0052] In step S102, the first extraction unit 105 and the second extraction unit 106 extract high-frequency and low-frequency components from the biological information.
[0053] In step S103, the phoneme estimation unit 109, acting as a language estimation unit, estimates language information, which is a sequence of phonemes, based on the extracted high-frequency components.
[0054] In step S104, the emotion estimation unit 108 estimates the user's emotions based on the extracted low-frequency components.
[0055] In step S105, the language information estimated in step S103 and the emotion information estimated in step S104 are used to process the speech and / or text to generate output information with emotion. Then, in step S106, the output information, which is the speech information and / or text information with emotion, is output.
[0056] The above-described embodiment can also be implemented by supplying a program that implements one or more functions to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. Furthermore, it can also be implemented by a circuit (e.g., an ASIC) that implements one or more functions. This program and the computer-readable storage medium storing the program are included in the embodiment.
[0057] The embodiments described above are merely examples of how the present invention can be implemented, and the technical scope of the present invention should not be interpreted as being limited by them. In other words, the present invention can be implemented in various forms without departing from its technical concept or its main features.
[0058] This embodiment includes the following configurations, methods, and programs.
[0059] (Composition 1) An acquisition unit that acquires biological information regarding the movement of the user's body parts, A language estimation unit estimates language information, including textual information or audio information, based on the biometric information acquired by the acquisition unit, Based on the biometric information acquired by the acquisition unit, an emotion estimation unit estimates the user's emotions, A generation unit generates output information based on the language information estimated by the language estimation unit and the emotion estimated by the emotion estimation unit. An information processing system characterized by having an output unit that outputs the output information generated by the generation unit.
[0060] (Configuration 2) The language estimation unit estimates the language information based on the signals of a first frequency band included in the biological information. The information processing system according to configuration 1, characterized in that the emotion estimation unit estimates the emotion based on a signal in a second frequency band which is a lower frequency band than the first frequency band included in the biological information.
[0061] (Composition 3) The information processing system according to Configuration 1, characterized in that the generation unit generates output information by processing the voice information based on the emotion estimated by the emotion estimation unit.
[0062] (Composition 4) The information processing system according to Configuration 1, characterized in that the generation unit generates output information by processing the character information based on the emotion estimated by the emotion estimation unit.
[0063] (Composition 5) The information processing system according to configuration 3, characterized in that the generation unit generates the output information by processing the voice information with at least one of pitch, voice length, and voice intensity based on the emotion estimated by the emotion estimation unit.
[0064] (Composition 6) The information processing system according to configuration 4, characterized in that the generation unit generates the output information by performing at least one of the following on the character information: processing the color of the characters, processing the font, processing the background color, adding symbols, or adding icons, based on the emotion estimated by the emotion estimation unit.
[0065] (Composition 7) It is equipped with a detection unit that detects what the other party says, The information processing system according to Configuration 1, characterized in that the emotion estimation unit estimates the user's emotions based on the biometric information acquired when the conversation partner is speaking.
[0066] (Composition 8) The information processing system according to Configuration 1, characterized in that the emotion estimation unit estimates the user's emotions based on the biometric information acquired when the user is performing a speech action.
[0067] (Composition 9) Steps include acquiring biometric information regarding the movement of the user's body parts, A step of estimating linguistic information including textual information or audio information based on the acquired biometric information, A step of estimating the user's emotions based on the acquired biometric information, A step of generating output information based on the estimated language information and the estimated emotion, An information processing method characterized by comprising the step of outputting the generated output information.
[0068] (Composition 10) A program for causing a computer to execute the information processing method described in Configuration 9. [Explanation of symbols]
[0069] 1000 Information Processing Systems 100 detection devices 101 Information Processing Device 110 Emotion estimation part 111 Generation part
Claims
1. An acquisition unit that acquires biological information regarding the movement of the user's body parts, A language estimation unit estimates language information, including textual information or audio information, based on the biometric information acquired by the acquisition unit, Based on the biometric information acquired by the acquisition unit, an emotion estimation unit estimates the user's emotions, A generation unit generates output information based on the language information estimated by the language estimation unit and the emotion estimated by the emotion estimation unit. An information processing system characterized by having an output unit that outputs the output information generated by the generation unit.
2. The language estimation unit estimates the language information based on the signals of the first frequency band included in the biological information. The information processing system according to claim 1, characterized in that the emotion estimation unit estimates the emotion based on a signal in a second frequency band which is a lower frequency band than the first frequency band included in the biological information.
3. The information processing system according to claim 1, characterized in that the generation unit generates output information by processing the voice information based on the emotion estimated by the emotion estimation unit.
4. The information processing system according to claim 1, characterized in that the generation unit generates output information by processing the character information based on the emotion estimated by the emotion estimation unit.
5. The information processing system according to claim 3, characterized in that the generation unit generates output information by processing the voice information with at least one of pitch, voice length, and voice intensity based on the emotion estimated by the emotion estimation unit.
6. The information processing system according to claim 4, characterized in that the generation unit generates the output information by performing at least one of the following on the character information: processing the color of the characters, processing the font, processing the background color, adding symbols, or adding icons, based on the emotion estimated by the emotion estimation unit.
7. It is equipped with a detection unit that detects what the other party says, The information processing system according to claim 1, characterized in that the emotion estimation unit estimates the user's emotions based on the biometric information acquired when the person the other party is speaking.
8. The information processing system according to claim 1, characterized in that the emotion estimation unit estimates the user's emotions based on the biometric information acquired when the user is performing a speech action.
9. Steps include acquiring biometric information regarding the movement of the user's body parts, A step of estimating linguistic information including textual information or audio information based on the acquired biometric information, A step of estimating the user's emotions based on the acquired biometric information, A step of generating output information based on the estimated language information and the estimated emotion, An information processing method characterized by comprising the step of outputting the generated output information.
10. A program for causing a computer to execute the information processing method described in claim 9.