Information processing method, information processing device, and program
The information processing system addresses the challenge of accurately capturing speaker intent in telemedicine by creating a conversation log through speech and image recognition, ensuring clear interpretation of dialogues.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- TOSHIBA TEC KK
- Filing Date
- 2023-08-24
- Publication Date
- 2026-06-02
AI Technical Summary
Existing telemedicine systems struggle to accurately capture the speaker's intent in conversation logs, leading to potential misinterpretation of multiple intentions in spoken sentences.
An information processing system that includes a server and communication devices to acquire call information, create a conversation log, and determine the user's intent based on speech and image recognition, associating intent with the log for accurate interpretation.
The system effectively generates a conversation log that accurately reflects the speaker's intent, enhancing the clarity and understanding of dialogues in telemedicine consultations.
Smart Images

Figure 0007868789000001 
Figure 0007868789000002 
Figure 0007868789000003
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to an information processing method, an information processing apparatus, and a program.
Background Art
[0002] In recent years, as a medical treatment method in hospitals, not only "face-to-face treatment" in which patients visit the hospital and have a face-to-face consultation with a doctor, but also "telemedicine" in which doctors and patients use a communication device so that patients can receive a diagnosis without directly visiting a medical institution has come to be carried out. In telemedicine, medical equipment is installed at the patient's home or a predetermined location, and vital data obtained by the medical equipment is sent to a doctor in a remote location via communication, and the doctor makes a diagnosis based on the data.
[0003] As one type of telemedicine, there is an information processing system called a telemedicine booth, a telemedicine box, a telemedicine kiosk, etc., which are equipped with a device capable of communicating with medical equipment at a predetermined location and enabling medical treatment by means of a real-time video call or the like between a patient and a doctor. As a result, not only medical treatment by dialogue using a communication device but also medical treatment using medical equipment such as a sphygmomanometer and a stethoscope can be realized in telemedicine, and symptoms can be examined in more detail.
[0004] In telemedicine, the dialogue via a communication device may be converted into text data, and a conversation log that can be reviewed by a doctor, a patient, the patient's family, etc. after the dialogue is completed may be used.
[0005] However, in a conversation log, when a plurality of intentions are included in a spoken sentence, the intention of the speaker may not be correctly understood.
[0006] Therefore, there is a need for a call processing system that can create a conversation log including the intention of the speaker.
Prior Art Documents
Patent Documents
[0007]
Patent Document 1
[0008] The problem that the embodiments of the present invention aim to solve is to provide an information processing method, an information processing device, and a program that can create a conversation log that includes the speaker's intent. [Means for solving the problem]
[0009] In one embodiment, an information processing device that processes calls between multiple terminals performs an information processing method which includes acquiring user call information from a terminal, creating a conversation log based on the call information, determining the user's intent to speak based on the call information, and saving the intent to speak in association with the conversation log. [Brief explanation of the drawing]
[0010] [Figure 1] Figure 1 is an external view illustrating a remote medical consultation booth included in the information processing system according to the embodiment. [Figure 2] Figure 2 is a block diagram illustrating an information processing system according to an embodiment. [Figure 3] Figure 3 illustrates an example of the data structure of conversation log information according to the embodiment. [Figure 4] Figure 4 illustrates another example of the data structure of conversation log information according to the embodiment. [Figure 5] Figure 5 is a flowchart showing an example of the information processing procedure performed by the server according to this embodiment. [Figure 6] Figure 6 is a flowchart showing an example of the information processing procedure performed by the server according to this embodiment. [Modes for carrying out the invention]
[0011] (Embodiment) (Example configuration) The embodiments will be described below with reference to the drawings.
[0012] In each drawing, the same reference numeral is used for identical components whenever possible, and redundant explanations are omitted. Figure 1 is an external view illustrating a remote medical consultation booth included in the information processing system 100 according to the embodiment. A telemedicine booth is a facility enclosed on all four sides by walls, making it difficult for conversations to be overheard and ensuring patient privacy during consultations. A telemedicine booth is also called a consultation booth or simply a booth. Inside the booth, there are chairs, desks, medical equipment for measurements, and communication devices for communicating with a doctor outside the booth. For example, multiple consultation booths may be set up in each area, such as a city or town. Patients make a reservation in advance to receive a telemedicine consultation at a consultation booth and enter the reserved booth at the scheduled time. The patient communicates with the doctor via video call or other means using the communication device inside the consultation booth, and uses medical equipment to measure biometric data based on the doctor's instructions. The measurement data from the medical equipment can be viewed in real time on medical equipment used by the doctor. The doctor can then conduct a consultation based on the measurement data. Each consultation booth may be set up in a hospital, or it may be used by multiple doctors from multiple hospitals.
[0013] Figure 2 is a block diagram illustrating an information processing system 100 according to an embodiment. The information processing system 100 includes a server 1, a second communication device 4, and at least one remote medical consultation booth. The remote medical consultation booth includes a remote medical consultation support device 2 and a first communication device 3. The server 1, the remote medical consultation support device 2, the first communication device 3, and the second communication device 4 are connected to each other via a network. For example, the network consists of one or more networks from among various networks such as the Internet, a mobile communication network, and a LAN (Local Area Network). The one or more networks may include a wireless network or a wired network. The remote medical consultation support device 2 and the first communication device 3 are connected to each other via the network so as to be able to communicate with each other. The network is a LAN, etc. The LAN may be a wireless LAN or a wired LAN. The remote medical consultation support device 2 is connected to at least one medical device so as to be able to communicate with each other. The remote medical consultation support device 2 and at least one medical device 251-1 to 251-m are directly connected to each other so as to be able to communicate with each other via wired or wireless. The telemedicine support device 2 and at least one medical device 251-1 to 251-m may be connected by, for example, LAN, Bluetooth®, Wi-Fi®, etc. The information processing system 100 may also refer to a system including server 1, telemedicine support device 2, first communication device 3, second communication device 4, and at least two of the medical devices 251-1 to 251-m.
[0014] Server 1 is an electronic device that collects and processes data. The electronic device includes a computer. Server 1 is freely connected to the telemedicine support device 2, the first communication device 3, and the second communication device 4 via a network. Server 1 receives various data from the telemedicine support device 2, the first communication device 3, and the second communication device 4, and outputs various data to the telemedicine support device 2, the first communication device 3, and the second communication device 4. Server 1 may also be a server used in a cloud service.
[0015] Server 1 can implement a call service for performing video communication or the like between the first communication device 3 and the second communication device 4. Note that the call service may be a call service based on voice communication. The call service may not involve communication using video images. Server 1 can implement a telemedicine service between the telemedicine support device 2, the first communication device 3, and the second communication device 4. A configuration example of server 1 will be described later.
[0016] The telemedicine support device 2 is an electronic device that can communicate with other electronic devices. The telemedicine support device 2 is a device installed in a consultation booth. For example, the telemedicine support device 2 is a PC (Personal Computer), a smartphone, a tablet terminal, or the like. A participant may be read as a user or a person. A configuration example of the telemedicine support device 2 will be described later.
[0017] The first communication device 3 is an electronic device that can communicate with other electronic devices. The first communication device 3 is a device installed in a consultation booth. The first communication device 3 is, for example, a device used by a patient receiving telemedicine. For example, the first communication device 3 is a PC, a smartphone, a tablet terminal, or the like. A patient may be read as a user or a person. A configuration example of the first communication device 3 will be described later. The first communication device 3 is an example of a terminal.
[0018] The second communication device 4 is an electronic device that can communicate with other electronic devices. The second communication device 4 is, for example, a device used by a doctor performing telemedicine. For example, the second communication device 4 is a PC, a smartphone, a tablet terminal, or the like. A doctor may be read as a medical worker, a user, or a person. The second communication device 4 is, for example, installed at a location different from the consultation booth. The location where the consultation booth is installed is an example of the first base. The location where the second communication device 4 is installed is an example of the second base. The second base is, for example, a medical institution such as a hospital. A configuration example of the second communication device 4 will be described later. The second communication device 4 is an example of a terminal.
[0019] A configuration example of server 1 will be described. Server 1 is an electronic device including a processor 11, a main memory 12, an auxiliary storage device 13, and a communication interface 14. Each part constituting Server 1 is connected so as to be able to input and output signals to and from each other. In FIG. 1, the interface is described as "I / F".
[0020] The processor 11 corresponds to the central part of Server 1. The processor 11 is an element constituting the computer of Server 1. For example, the processor 11 is a CPU (Central Processing Unit), but is not limited thereto. The processor 11 may be composed of various circuits. The processor 11 expands a program stored in advance in the main memory 12 or the auxiliary storage device 13 into the main memory 12. The program is a program that causes the processor 11 of Server 1 to realize or execute each part described later. The processor 11 executes various operations by executing the program expanded in the main memory 12.
[0021] The main memory 12 corresponds to the main storage part of Server 1. The main memory 12 is an element constituting the computer of Server 1. The main memory 12 includes a non-volatile memory area and a volatile memory area. The main memory 12 stores an operating system or a program in the non-volatile memory area. The main memory 12 uses the volatile memory area as a work area where data is appropriately rewritten by the processor 11. For example, the main memory 12 includes ROM (Read Only Memory) as the non-volatile memory area. For example, the main memory 12 includes RAM (Random Access Memory) as the volatile memory area. The main memory 12 stores a program.
[0022] The auxiliary storage device 13 corresponds to the auxiliary storage portion of server 1. The auxiliary storage device 13 is an element that constitutes the computer of server 1. The auxiliary storage device 13 is an EEPROM (registered trademark) (Electric Erasable Programmable Read-Only Memory), HDD (Hard Disk Drive), or SSD (Solid State Drive), etc. The auxiliary storage device 13 stores the above-mentioned program, data used by the processor 11 in performing various processes, and data generated by the processing of the processor 11. The auxiliary storage device 13 stores the above-mentioned program. The auxiliary storage device 13 is an example of a storage unit.
[0023] The auxiliary storage device 13 stores conversation log information. The conversation log information is information recording a conversation between the user of the first communication device and the user of the second communication device. The conversation log information includes at least speaker information and conversation information. The speaker information is, for example, speaker identification information that can identify the speaker. The speaker includes the user of the first communication device and the user of the second communication device. The speaker identification information includes speaker ID, speaker name, etc. The speaker name may be the speaker's full name or user name, etc. The speaker identification information may also be information indicating the speaker's attributes, such as "doctor" or "patient". The conversation information is associated with the speaker identification information. The conversation information indicates the content of the utterance spoken by the speaker. The conversation information is information that has undergone speech recognition processing based on the speech information related to the utterance. Speech recognition includes, for example, converting the speech information related to the utterance into text data and segmenting it using known techniques. The conversation information includes at least one utterance. Speech segments are, for example, segments based on the boundaries of a user's utterances. Each speech segment may contain identifying information. Each speech segment contains at least one word. A word is the smallest meaningful linguistic unit. Words include parts of speech such as nouns, verbs, and interjections. A conversation log shows data of speech segments arranged chronologically. Speech segments are also called conversational sentences.
[0024] The conversation log information includes utterance time information. This utterance time information may include time information indicating the time the utterance was made, elapsed time information indicating the time elapsed since the conversation began, etc. The time the utterance was made corresponds to the time when Server 1 acquired the call information from the first communication device 3 or the second communication device 4.
[0025] Conversation log information includes recognition results regarding the user's state. The user's state indicates the user's emotions, gestures, etc. The recognition results indicate the result of recognizing the user's emotions or gestures based on at least one of the audio information and / or image information about the user. Image information may include moving images. Image information may include still images. Recognition of the user's emotions can be achieved by known emotion recognition technology. Recognition of the user's gestures can be achieved by known gesture recognition technology. The recognition results may include, for example, recognition results of the user's state based on the user's biometric data. The recognition results may include results of recognizing the user's state based on multiple types of data. The user's emotions may include positive, neutral, negative, etc. Positive may indicate affirmation, admiration, relief, etc. Neutral may indicate normalcy. Negative may indicate denial, doubt, anxiety, etc. The user's emotions may include joy, anger, sadness, normalcy, etc. The user's gestures may include body language, hand movements, gestures, etc. User gestures may include head gestures such as nodding, shaking, and tilting. The recognition result includes time information on the user's state. This time information may be a specific time, or it may be elapsed time information indicating the time elapsed since the conversation began.
[0026] Conversation log information includes information about the user's utterance intent based on the recognition results regarding the user's state. The information about utterance intent indicates the determination result of the user's utterance intent based on the recognition results regarding the user's state. The user's utterance intent may also indicate the type of utterance. The type of utterance includes affirmative, negative, interrogative, neutral, etc. Affirmative indicates an affirmative sentence or affirmative form. Negative indicates a negative sentence or negative form. Interrogative indicates an interrogative sentence or interrogative form. Neutral indicates a declarative sentence or declarative form. Neutral may include affirmative. The type of utterance may also include exclamations, commands, etc. The type of utterance may distinguish utterances that can be interpreted in multiple ways. The type of utterance may also be a classification to distinguish between affirmative and negative utterances that have the same string of characters and can be interpreted in both an affirmative and a negative way. For example, the utterance "Hmm" can be interpreted in multiple ways, such as indicating affirmation ("Hmm (That's right)") or questioning ("Hmm (I suppose so)"). The type of utterance should be one that can classify the intent behind utterances with the same string of characters. The type of utterance for utterances with the same string of characters may differ depending on the recognition result of the user's state. Information regarding the user's intent may include time information about the time the user's state was acquired. The correspondence between recognition results and intent may be set in advance. The correspondence between recognition results and intent may differ depending on the scene in which the conversation takes place, the speaker, etc.
[0027] Conversation log information is information that links at least speaker information, conversation information, perception results regarding the user's state, and information regarding the user's intent.
[0028] The auxiliary memory device 13 can store information about the medical booth. The information about the medical booth includes information that can identify the medical booth. The information about the medical booth may also include information about the medical equipment installed in the medical booth.
[0029] The auxiliary storage device 13 can store user information. User information is information about the user (patient) using the examination booth. User information includes user identification information. User identification information is unique identification information assigned to each user to identify them individually. User information may include information such as the user's location information and medical history information. User information may include measurement data from medical devices.
[0030] The communication interface 14 includes various interfaces that enable communication between the server 1 and other electronic devices via a network, according to a predetermined communication protocol.
[0031] Note that the hardware configuration of Server 1 is not limited to the configuration described above. Server 1 may, as appropriate, omit or change the above-described components and add new components.
[0032] The various components implemented in the aforementioned processor 11 will now be described. The processor 11 implements a call information acquisition unit 110, a log creation unit 111, a memory control unit 112, a recognition unit 113, a determination unit 114, and an output unit 115. Each part implemented in the processor 11 can also be called a function. Each part implemented in the processor 11 can also be said to be implemented in the control unit which includes the processor 11 and the main memory 12. The processor 11 is an example of a processing circuit.
[0033] The call information acquisition unit 110 acquires call information from the first communication device 3 and the second communication device 4 via the communication interface 14. The call information includes voice information. The voice information includes voice information relating to the user's speech. The voice information may also include information indicating the characteristics of the voice. These characteristics may include pitch, timbre, etc. The call information also includes image information. The image information includes information captured by camera 351-2 or camera 451-2. The image information may include moving images and still images. The call information includes speaker information and speech time information. The call information includes time information relating to the time the image information was acquired. The time information relating to the time the image information was acquired corresponds to the time information relating to the time the user's state was acquired. The call information acquisition unit 110 may also acquire the user's biometric data from the first communication device 3 and the second communication device 4. The biometric data may be measurement data from medical devices 251-1 to 251-m. The biometric data may also be data measured by sensors, etc., connected to the first communication device 3 and the second communication device 4.
[0034] The log creation unit 111 creates a conversation log based on the call information. The log creation unit 111 creates a conversation log in chronological order based on the voice information related to the user's utterances, the speaker's information, and the utterance time information.
[0035] The log creation unit 111 updates the conversation log based on information about the user's intent. The log creation unit 111 updates the conversation log based on time information included in the information about the user's intent. The log creation unit 111 associates the information about the user's intent with each utterance included in the conversation log. The information about the conversation log created by the log creation unit 111 is also called conversation log information.
[0036] The memory control unit 112 stores the information acquired by the call information acquisition unit 110 in the auxiliary storage device 13. The memory control unit 112 stores the conversation log information. The memory control unit 112 updates the conversation log information. The memory control unit 112 stores the recognition result from the recognition unit 113 (described later) in the auxiliary storage device 13. The memory control unit 112 stores the determination result from the determination unit 114 (described later) in the auxiliary storage device 13.
[0037] The recognition unit 113 recognizes the user's state based on at least one of audio information and image information. The recognition unit 113 may recognize the user's emotions based on the audio information. The recognition unit 113 may recognize the user's emotions based on the image information. The recognition unit 113 may recognize the user's gestures based on the image information. The recognition unit 113 may recognize the user's emotions based on biometric data. The recognition unit 113 may recognize the user's emotions based on known technologies. The recognition unit 113 may recognize the user's gestures based on known technologies.
[0038] The determination unit 114 determines the user's utterance intent based on the recognition result by the recognition unit 113. The determination unit 114 classifies the recognition result of the user's state into utterance intent. The determination unit 114 may also determine the utterance intent based on the user's emotions. For example, if the user's emotions are "positive", the determination unit 114 determines the utterance intent as "affirmative". If the user's emotions are "neutral", the determination unit 114 determines the utterance intent as "normal". If the user's emotions are "negative", the determination unit 114 determines the utterance intent as "questionable". The determination unit 114 may also determine the utterance intent as "negative" if the user's emotions are "negative".
[0039] The determination unit 114 may determine the speaker's intent based on the user's gestures. For example, if the user's gesture is a nod, the determination unit 114 determines the speaker's intent to be "affirmative." If the user's gesture is a tilt of the head, the determination unit 114 determines the speaker's intent to be "questionable." If the user's gesture is a shake of the head, the determination unit 114 determines the speaker's intent to be "negative." The determination unit 114 may also determine the speaker's intent based on the recognition result and the content of the utterance.
[0040] The output unit 115 outputs conversation log information to the first communication device 3 and the second communication device 4 via the communication interface 14. The output unit 115 may also output conversation log information based on conversation log display requests from the first communication device 3 and the second communication device 4.
[0041] This explains the conversation log information. Figure 3 illustrates an example of the data structure of conversation log information according to the embodiment. Figure 3 shows conversation log information when a conversation takes place between a doctor and a patient. The conversation log information includes at least speaker identification information, conversation information, recognition results, and information regarding the speaker's intent. Speaker identification information is, for example, information indicating "doctor" or "patient". Speaker identification information is associated with each utterance. Speaker identification information may also be information based on information that can identify the communication device from which the audio information was output. Speaker identification information may also be information based on information that can identify the user of the communication device from which the audio information was output. Conversation information is, for example, text information indicating the content of the utterance. Conversation information is associated with each utterance. Recognition results indicate the recognition results regarding the user's state. Recognition results indicate the results of recognizing the user's emotions or gestures based on at least one of the audio information and image information about the user. Recognition results indicate the results of recognizing the speaker's emotions or gestures associated with each utterance. Audio information and image information are information output from the first communication device 3 and the second communication device 4. The recognition results are linked to each utterance. Information regarding the utterance intent indicates the utterance intent based on the recognition results. Information regarding the utterance intent is linked to each utterance.
[0042] In the example in Figure 3, the perceived state of the speaker for the conversation content with utterance IDs "1" and "2" is "neutral." In this case, the utterance intention can be judged as "normal." The conversation content "Hmm (I guess)" with utterance ID "3" can contain multiple utterance intentions. For example, the speaker may utter "Hmm (I guess)" with an affirmative intention, or with a questioning intention. Note that the content in parentheses is not uttered. In the example in Figure 3, the perceived state of the doctor when the conversation content "Hmm" is uttered indicates that the utterance intention is "negative." The doctor's utterance intention when the conversation content "Hmm" is uttered indicates the type of utterance based on the perceived result of "negative." In this example, if the speaker's emotion is "negative," the determination unit 114 determines the speaker's intent to speak as "question," and if the speaker's emotion is "positive," the determination unit 114 determines the speaker's intent to speak as "affirmative." In this case, the determination unit 114 determines that the doctor's intent to speak when he utters "Hmm" is "question."
[0043] In this example, when multiple utterance intentions can be determined from the content of an utterance, the appropriate utterance intention can be determined based on the recognition result regarding the speaker's state. Therefore, users viewing the conversation log can clearly recognize the utterance intention of each utterance. As a result, the information processing system 100 can create a conversation log in which the speaker's intention has been estimated.
[0044] Figure 4 illustrates another example of the data structure of conversation log information according to the embodiment. Figure 4 shows conversation log information when a conversation takes place between a doctor and a patient. The data structure of the conversation log information is the same as in the example in Figure 3.
[0045] In the example in Figure 4, the recognition result of the speaker's state for the conversation content with utterance IDs "4" and "5" is "neutral". In this case, the utterance intention can be judged as "normal". For the conversation content "Hmm (That's right)" with utterance ID "6", multiple utterance intentions can be judged. In the example in Figure 4, the recognition result of the doctor's state when the conversation content "Hmm" was uttered indicates "positive" and "nodding". The doctor's utterance intention when the conversation content "Hmm" was uttered indicates the type of utterance based on the recognition result "positive" and "nodding". In this example, since the speaker's emotion is "positive", the judgment unit 114 can determine that the utterance intention is "affirmative". Also, since the speaker's gesture is "nodding", the judgment unit 114 can determine that the utterance intention is "affirmative". The determination unit 114 may determine the speaker's intent based on at least one of the speaker's emotions and / or gestures. In this example, the determination unit 114 determines that the doctor's intent when he utters the word "Hmm" is "affirmative."
[0046] In this example, when multiple utterance intentions can be determined from the content of an utterance, the appropriate utterance intention can be determined based on the recognition result regarding the speaker's state. Therefore, users viewing the conversation log can clearly recognize the utterance intention of each utterance. As a result, the information processing system 100 can create a conversation log in which the speaker's intention has been estimated.
[0047] This section describes an example configuration of the telemedicine support device 2. The telemedicine support device 2 is an electronic device that includes a processor 21, main memory 22, auxiliary storage device 23, communication interface 24, input / output interface 25, display device 26, speaker 27, and input device 28. Each component of the telemedicine support device 2 is connected to each other so that signals can be input and output.
[0048] The processor 21 is the central part of the telemedicine support device 2. The processor 21 is an element that makes up the computer of the telemedicine support device 2. For example, the processor 21 has the same hardware configuration as the processor 11 described above. The processor 21 performs various operations by executing programs loaded into the main memory 22. The processor 21 is an example of a processing circuit.
[0049] Main memory 22 corresponds to the main memory portion of the telemedicine support device 2. Main memory 22 is an element that constitutes the computer of the telemedicine support device 2. Main memory 22 has the same hardware configuration as main memory 12 described above. Main memory 22 stores programs.
[0050] The auxiliary storage device 23 corresponds to the auxiliary storage portion of the telemedicine support device 2. The auxiliary storage device 23 is an element that constitutes the computer of the telemedicine support device 2. The auxiliary storage device 23 has the same hardware configuration as the auxiliary storage device 13 described above. The auxiliary storage device 23 stores the program described above.
[0051] The communication interface 24 includes various interfaces that enable the remote medical care support device 2 to communicate with other devices via a network, in accordance with a predetermined communication protocol.
[0052] The input / output interface 25 is an interface for connecting the telemedicine support device 2 with external devices. The external devices include at least one medical device 251-1 to 251-m (where m is an integer greater than or equal to 1). Medical devices 251-1 to 251-m include, for example, electrocardiographs, blood pressure monitors, digital stethoscopes, pulse oximeters, slit lamps, otoscopes, and other medical devices. Medical devices 251-1 to 251-m have communication functions and output measurement data to the telemedicine support device 2.
[0053] The display device 26 is a device capable of displaying various screens under the control of the processor 21. For example, the display device 26 may be a liquid crystal display or an EL display.
[0054] Speaker 27 is a device capable of outputting sound under the control of processor 21. Speaker 27 is an example of an output device.
[0055] The input device 28 is a device capable of inputting data or instructions to the telemedicine support device 2. For example, the input device 28 includes a built-in microphone capable of inputting voice, and a built-in camera capable of acquiring image data of the shooting range. The input device 28 may also include a keyboard or a touch panel.
[0056] The hardware configuration of the telemedicine support device 2 is not limited to the configuration described above. The telemedicine support device 2 allows for the omission and modification of the above-mentioned components, as well as the addition of new components, as appropriate.
[0057] An example configuration of the first communication device 3 will be described. The first communication device 3 is an electronic device that includes a processor 31, main memory 32, auxiliary storage device 33, communication interface 34, input / output interface 35, and input device 38. Each component of the first communication device 3 is connected to each other so that signals can be input and output.
[0058] The processor 31 corresponds to the central part of the first communication device 3. The processor 31 is an element that constitutes the computer of the first communication device 3. The processor 31 has the same hardware configuration as the processor 11 described above. The processor 31 performs various operations by executing programs that are pre-stored in the main memory 32 or auxiliary storage device 33. The processor 31 is an example of a processing circuit.
[0059] Main memory 32 corresponds to the main memory portion of the first communication device 3. Main memory 32 is an element that constitutes the computer of the first communication device 3. Main memory 32 has the same hardware configuration as the main memory 12 described above. Main memory 32 stores programs.
[0060] The auxiliary storage device 33 corresponds to the auxiliary storage portion of the first communication device 3. The auxiliary storage device 33 is an element that constitutes the computer of the first communication device 3. The auxiliary storage device 33 has the same hardware configuration as the auxiliary storage device 13 described above. The auxiliary storage device 33 stores the program described above.
[0061] The communication interface 34 includes various interfaces that enable the first communication device 3 to communicate with other devices via a network, in accordance with a predetermined communication protocol.
[0062] The input / output interface 35 is an interface for connecting the first communication device 3 to external devices. The external devices include a display device 351-1, a camera 351-2, a microphone 351-3, and a speaker 351-4. The display device 351-1 is a device that can display various screens under the control of the processor 31. For example, the display device 351-1 is a liquid crystal display or an EL display. The camera 351-2 is a device that can acquire shooting data of the shooting range under the control of the processor 31. The microphone 351-3 is a device that can input sound under the control of the processor 31. The speaker 351-4 is a device that can output sound under the control of the processor 31.
[0063] The input device 38 is a device capable of inputting data or instructions to the first communication device 3. For example, the input device 38 includes a keyboard or a touch panel.
[0064] The hardware configuration of the first communication device 3 is not limited to the configuration described above. The first communication device 3 may be modified or have its components omitted or changed as appropriate, and new components added as needed.
[0065] A configuration example of the second communication device 4 will be described. The second communication device 4 is an electronic device that includes a processor 41, main memory 42, auxiliary storage device 43, communication interface 44, input / output interface 45, and input device 48. Each component of the second communication device 4 is connected to each other so that signals can be input and output.
[0066] The processor 41 corresponds to the central part of the second communication device 4. The processor 41 is an element that constitutes the computer of the second communication device 4. The processor 41 has the same hardware configuration as the processor 11 described above. The processor 41 performs various operations by executing programs that are pre-stored in the main memory 42 or auxiliary storage device 43. The processor 41 is an example of a processing circuit.
[0067] Main memory 42 corresponds to the main memory portion of the second communication device 4. Main memory 42 is an element that constitutes the computer of the second communication device 4. Main memory 42 has the same hardware configuration as the main memory 12 described above. Main memory 42 stores programs.
[0068] The auxiliary storage device 43 corresponds to the auxiliary storage portion of the second communication device 4. The auxiliary storage device 43 is an element that constitutes the computer of the second communication device 4. The auxiliary storage device 43 has the same hardware configuration as the auxiliary storage device 13 described above. The auxiliary storage device 43 stores the program described above.
[0069] The communication interface 44 includes various interfaces that enable the second communication device 4 to communicate with other devices via a network, in accordance with a predetermined communication protocol.
[0070] The input / output interface 45 is an interface for connecting the second communication device 4 to external devices. The external devices include a display device 451-1, a camera 451-2, a microphone 451-3, and a speaker 451-4. The display device 451-1 is a device capable of displaying various screens under the control of the processor 41. For example, the display device 451-1 is a liquid crystal display or an EL display. The camera 451-2 is a device capable of acquiring shooting data of the shooting range under the control of the processor 41. The microphone 451-3 is a device capable of inputting sound under the control of the processor 41. The speaker 451-4 is a device capable of outputting sound under the control of the processor 41.
[0071] The input device 48 is a device capable of inputting data or instructions to the second communication device 4. For example, the input device 48 includes a keyboard or a touch panel.
[0072] The hardware configuration of the second communication device 4 is not limited to the configuration described above. The second communication device 4 may be modified or have its components omitted or changed as appropriate, and new components added as needed.
[0073] (Example of operation) The procedure for processing by the information processing system 100 will be explained. In the following explanation focusing on Server 1, you may substitute Server 1 for Processor 11. In the explanation focusing on the first communication device 3, you may substitute the first communication device 3 for Processor 31. In the explanation focusing on the second communication device 4, you may substitute the second communication device 4 for Processor 41. The processing procedure described below is merely an example, and each process may be modified as much as possible. Furthermore, depending on the embodiment, steps in the processing procedure described below may be omitted, replaced, or added as appropriate.
[0074] Figure 5 is a flowchart showing an example of the information processing procedure performed by Server 1 according to this embodiment.
[0075] The following process shall commence based on the commencement of a call between the first communication device 3 and the second communication device 4.
[0076] The call information acquisition unit 110 acquires call information from the first communication device 3 and the second communication device 4 (ACT1). In ACT1, for example, the call information acquisition unit 110 acquires voice information from the first communication device 3 and the second communication device 4. The voice information may include speaker identification information for identifying the user of the first communication device 3 or the second communication device 4. The call information acquisition unit 110 acquires image information from the first communication device 3 and the second communication device 4. The image information may include speaker identification information for identifying the user of the first communication device 3 or the second communication device 4. The memory control unit 112 may store the acquired call information in the auxiliary storage device 13.
[0077] The log creation unit 111 creates conversation log information based on the call information (ACT2). In ACT2, for example, the log creation unit 111 performs speech recognition processing based on the voice information related to the user's utterances. The log creation unit 111 segments the text information based on the voice information related to the user's utterances into utterance units. The log creation unit 111 links the segmented text information to speaker identification information. The log creation unit 111 creates conversation log information that includes the utterances linked to the speaker identification information. The log creation unit 111 may also create conversation log information that includes speech time information for each utterance.
[0078] The memory control unit 112 saves the conversation log information to the auxiliary storage device 13 (ACT3).
[0079] The recognition unit 113 recognizes the user's state based on the call information (ACT4). In ACT4, for example, the recognition unit 113 recognizes the user's state based on at least one of the audio information and the image information. The recognition unit 113 recognizes the user's emotions based on at least one of the audio information and the image information. For example, the recognition unit 113 may perform emotion recognition based on the audio information. The recognition unit 113 may perform emotion recognition based on the image information. The recognition unit 113 may perform emotion recognition based on the audio information and the image information. The recognition unit 113 recognizes the user's gestures based on at least one of the audio information and the image information. For example, the recognition unit 113 may perform gesture recognition based on the image information. The recognition unit 113 may associate the recognition results of the user's state with each utterance in chronological order. The recognition unit 113 may associate the recognition results of the user's state with each utterance based on the time information included in the recognition results of the user's state and the speech time information related to the utterance.
[0080] The determination unit 114 determines the user's intent to speak based on the call information (ACT5). In ACT5, for example, the determination unit 114 determines the user's intent to speak based on at least one of the audio information and the image information. The determination unit 114 determines the user's intent to speak based on the recognition result by the recognition unit 113. For example, the determination unit 114 may determine the intent to speak based on the recognition result of recognizing the user's emotions. The determination unit 114 may determine the intent to speak based on the recognition result of recognizing the user's gestures.
[0081] For example, the determination unit 114 describes the case where the recognition result of the user's emotion obtained when uttering utterance ID "01" is "positive". In this case, the determination unit 114 may determine that the intention of utterance ID "01" is "affirmative". The determination unit 114 describes the case where the recognition result of the user's emotion is "negative". In this case, the determination unit 114 may determine that the intention of utterance ID "01" is "negative" or "questionable". The determination unit 114 describes the case where the recognition result of the user's emotion is "neutral". In this case, the determination unit 114 may determine that the intention of utterance ID "01" is "normal". The recognition result of the user's emotion is not limited to "positive", "negative", and "neutral". The recognition result of the user's emotion may differ depending on the scene in which the conversation takes place, the speaker, etc.
[0082] For example, the determination unit 114 will explain the case where the recognition result of the user's gesture obtained when the utterance with utterance ID "02" is "nodding". In this case, the determination unit 114 may determine that the intent of the utterance with utterance ID "02" is "affirmative". The determination unit 114 will explain the case where the recognition result of the user's emotion is "tilting the head". In this case, the determination unit 114 may determine that the intent of the utterance with utterance ID "02" is "questioning". The determination unit 114 will explain the case where the recognition result of the user's emotion is "shaking the head". In this case, the determination unit 114 may determine that the intent of the utterance with utterance ID "02" is "negative". The recognition results of the user's gesture are not limited to "nodding", "tilting the head", and "shaking the head". The recognition results of the user's gesture may differ depending on the scene in which the conversation takes place, the speaker, etc.
[0083] The determination unit 114 may determine the intent of the utterance based on the recognition result of the user's emotion and the recognition result of the user's gesture. For example, the determination unit 114 will describe a case where the recognition result of the user's emotion obtained when utterance ID "03" was uttered is "negative" and the recognition result of the user's gesture is "tilting the head". In this case, the determination unit 114 may determine that the intent of utterance ID "03" is "questioning".
[0084] The determination unit 114 may determine the intent of the utterance based on the recognition result and the content of the utterance. For example, the case in which the content of the utterance with utterance ID "04" may include either a "question" or an "affirmation" intent will be described. The case in which the recognition result of the user's emotion is "positive" will be described. In this case, the determination unit 114 may determine that the intent of the utterance with utterance ID "04" is "affirmation". The case in which the recognition result of the user's emotion is "negative" will be described. In this case, the determination unit 114 may determine that the intent of the utterance with utterance ID "04" is "question".
[0085] The following describes the case where the user's gesture is recognized as "nodding." In this case, the determination unit 114 may determine that the intent of the utterance with utterance ID "04" is "affirmative." The following describes the case where the user's gesture is recognized as "tilting the head." In this case, the determination unit 114 may determine that the intent of the utterance with utterance ID "04" is "questioning."
[0086] The memory control unit 112 stores the user's utterance intent determined by the determination unit 114, associating it with each utterance (ACT6). The memory control unit 112 may also update the conversation log information based on the user's utterance intent.
[0087] Processor 11 determines whether the call between the first communication device 3 and the second communication device 4 has ended (ACT7). If Processor 11 determines that the call has ended (ACT7:YES), the process ends. If Processor 11 determines that the call has not ended (ACT7:NO), the process transitions from ACT7 to ACT1. Processor 11 repeats the processes of ACT1 to ACT6 until the call between the first communication device 3 and the second communication device 4 has ended.
[0088] The processor 11 may perform the processing of ACT2 to ACT6 after the call between the first communication device 3 and the second communication device 4 has ended. For example, the call information acquisition unit 110 may acquire call information while the call between the first communication device 3 and the second communication device 4 is in progress. The log creation unit 111 may create conversation log information after the call between the first communication device 3 and the second communication device 4 has ended. The recognition unit 113 may recognize the user's emotions after the call between the first communication device 3 and the second communication device 4 has ended. The determination unit 114 may determine the user's intent to speak after the call between the first communication device 3 and the second communication device 4 has ended.
[0089] In this example, Server 1 can create a conversation log that includes the speaker's intent, which may be difficult to understand from text information alone. For example, if multiple intents can be determined from the content of an utterance, Server 1 can determine the appropriate intent based on the recognition result regarding the speaker's state. Therefore, users viewing the conversation log can clearly understand the intent of each utterance. As a result, the information processing system 100 can create a conversation log that estimates the speaker's intent. Furthermore, Server 1 can determine the speaker's intent for each utterance and create a conversation log that includes the speaker's intent for each utterance. Therefore, users viewing the conversation log can understand the speaker's intent on an utterance-by-utterance basis. For example, when a third party other than the speaker reviews the conversation log, they can appropriately understand its contents.
[0090] Figure 6 is a flowchart showing an example of the information processing procedure performed by Server 1 according to this embodiment.
[0091] The following process shall begin after the call between the first communication device 3 and the second communication device 4 has ended.
[0092] The processor 11 receives a conversation log display request from at least one of the first communication device 3 and the second communication device 4 (ACT 11). In ACT 11, for example, the processor 11 receives the conversation log display request based on a user operation of the first communication device 3 or the second communication device 4. The user operation may include, for example, clicking or touching a display button to display the conversation log displayed on a display device.
[0093] Based on the conversation log display request, the output unit 115 outputs conversation log information to at least one of the first communication device 3 and the second communication device 4, which are the sources of the conversation log display request (ACT12). In ACT12, for example, the output unit 115 outputs information for displaying the conversation log and the speaker's intent. The first communication device 3 or the second communication device 4 acquires the conversation log information. Based on the conversation log information, the first communication device 3 or the second communication device 4 displays the conversation log and the speaker's intent on a display device.
[0094] In this example, Server 1 can output a conversation log, including the speaker's intent, to at least one of the first communication device 3 and the second communication device 4. Therefore, users of the first communication device 3 and the second communication device 4 can view the conversation log, which includes the intent of each utterance. Users can easily understand the intent of each utterance.
[0095] Furthermore, the above-described process may be performed while a call is taking place between the first communication device 3 and the second communication device 4.
[0096] The conversation log display request may be output from an electronic device other than the first communication device 3 and the second communication device 4. For example, the conversation log display request may be output from a user terminal (not shown). In this case, the output unit 115 may output the conversation log information to the user terminal that is the source of the conversation log display request.
[0097] (Other embodiments) In the embodiments described above, an example was described in which the medical examination booth includes a telemedicine support device 2 and a first communication device 3, but it is not limited to this. The medical examination booth may include either the telemedicine support device 2 or the first communication device 3. If the medical examination booth does not include the first communication device 3 but includes the telemedicine support device 2, the telemedicine support device 2 may implement the functions of the first communication device 3. If the medical examination booth does not include the telemedicine support device 2 but includes the first communication device 3, the first communication device 3 may implement the functions of the telemedicine support device 2.
[0098] The above-described embodiment uses a call between a doctor and a patient in telemedicine as an example, but is not limited to this. The above-described embodiment is applicable when a call is made by users of multiple communication devices, such as web conferencing, video conferencing, and distance education.
[0099] The information processing device may be implemented as a single device, such as Server 1, or as multiple devices with distributed functions.
[0100] The embodiments described above may apply not only to the apparatus but also to the methods performed by the apparatus. The embodiments described above may apply to a program that can cause the computer of the apparatus to perform each function. The embodiments described above may apply to a recording medium that stores the program. The embodiments described above may apply not only to the system but also to the methods performed by multiple elements included in the system.
[0101] A processing circuit includes one or more circuits that perform multiple processes through multiple functions. For example, the circuit may be, but is not limited to, a processor, an ASIC (Application Specific Integrated Circuit), or an FPGA (Field-Programmable Gate Array).
[0102] Each of the one or more circuits that make up a processing circuit performs one or more of the multiple processes. If the processing circuit consists of a single circuit, the single circuit performs all of the multiple processes. If the processing circuit consists of multiple circuits, each of the multiple circuits performs some of the multiple processes. Some of the multiple processes may be one of the multiple processes, or two or more of the multiple processes. If the processing circuit consists of multiple circuits, the multiple circuits may be contained in a single device, or they may be distributed across multiple devices.
[0103] The program may be transferred while stored in the device, or it may be transferred without being stored in the device. In the latter case, the program may be transferred via a network, or it may be transferred while recorded on a recording medium. The recording medium is a non-temporary tangible medium. The recording medium is a computer-readable medium. The recording medium can be any medium that is capable of storing a program and is readable by a computer, such as a CD-ROM or memory card, and its form is not limited.
[0104] Although embodiments of the present invention have been described in detail above, the above description is merely illustrative in all respects of the present invention. Needless to say, various improvements and modifications can be made without departing from the scope of the present invention. In other words, when implementing the present invention, specific configurations may be adopted as appropriate depending on the embodiment.
[0105] In short, this invention is not limited to the embodiments described above, and in the implementation stage, the components can be modified and materialized without departing from the gist of the invention. Furthermore, various inventions can be formed by appropriately combining the multiple components disclosed in the embodiments. For example, some components may be deleted from all the components shown in the embodiments. Moreover, components from different embodiments may be appropriately combined.
[0106] (Note) The above-described embodiments may be represented as follows: (1) An information processing method performed by an information processing device that processes calls between multiple terminals, Obtaining user call information from the device, Based on the aforementioned call information, a conversation log will be created, Determining the user's intent to speak based on the aforementioned call information, The purpose of the utterance is saved in association with the aforementioned conversation log, An information processing method comprising the following: (2) Further comprising recognizing the user's status based on the call information, The aforementioned determination includes determining the utterance intent based on the recognition result. (1) The information processing method described above. (3) The call information includes voice information and image information, The determination includes determining the intent of the utterance based on at least one of the audio information and the image information. (1) The information processing method described above. (4) Further comprising outputting the conversation log and the utterance intent, (1) The information processing method described above. (5) An information processing device for processing calls between multiple terminals, A call information acquisition unit that acquires user call information from the terminal, A log creation unit creates a conversation log based on the aforementioned call information, A determination unit that determines the user's intent to speak based on the aforementioned call information, A memory unit that stores the utterance intent linked to the conversation log, An information processing device equipped with the following features. (6) The computer of the information processing device that processes calls between multiple terminals, A function to obtain user call information from the device, A function to create a conversation log based on the aforementioned call information, A function to determine the user's intent to speak based on the aforementioned call information, A function to save the utterance intent linked to the aforementioned conversation log, A program that can execute [this action]. [Explanation of Symbols]
[0107] 1…Server, 2…Remote medical support device, 3…First communication device, 4…Second communication device, 11…Processor, 12…Main memory, 13…Auxiliary storage device, 14…Communication interface, 21…Processor, 22…Main memory, 23…Auxiliary storage device, 24…Communication interface, 25…Input / output interface, 26…Display device, 27…Speaker, 28…Input device, 31…Processor, 32…Main memory, 33…Auxiliary storage device, 34…Communication interface, 35…Input / output interface, 38…Input device, 41… Processor, 42...Main memory, 43...Auxiliary storage device, 44...Communication interface, 45...Input / output interface, 48...Input device, 100...Information processing system, 110...Call information acquisition unit, 111...Log creation unit, 112...Memory control unit, 113...Recognition unit, 114...Determination unit, 115...Output unit, 251-1~251-m...Medical equipment, 351-1...Display device, 351-2...Camera, 351-3...Microphone, 351-4...Speaker, 451-1...Display device, 451-2...Camera, 451-3...Microphone, 451-4...Speaker.
Claims
1. An information processing method performed by an information processing device that processes calls between multiple terminals, To obtain call information, including user voice information and image information, from the terminal, Based on the aforementioned audio information, a conversation log containing conversation information is created, Recognizing the user's gestures based on the aforementioned image information, Based on the recognition results of the gesture, the purpose of the utterance is determined, which is a classification for distinguishing the intent of the conversational information. The conversation log is saved, in which the recognition result and the utterance intent are linked to the conversation information. An information processing method comprising the following:
2. The system further comprises outputting the aforementioned conversation log and the aforementioned utterance intent. The information processing method according to claim 1.
3. An information processing device that processes calls between multiple terminals, A call information acquisition unit that acquires call information including user voice information and image information from the terminal, A log creation unit creates a conversation log including conversation information based on the aforementioned audio information, A recognition unit that recognizes the user's gestures based on the aforementioned image information, A determination unit that determines the intent of an utterance, which is a classification for distinguishing the intent of the conversational information, based on the recognition result of the gesture, A storage unit that stores the conversation log in which the recognition result and the utterance intent are linked to the conversation information, An information processing device equipped with the following features.
4. In the computer of the information processing device that handles calls between multiple terminals, A function to acquire call information, including user voice information and image information, from the terminal, A function to create a conversation log including conversation information based on the aforementioned audio information, A function to recognize the user's gestures based on the aforementioned image information, Based on the recognition results of the aforementioned gestures, a function is provided to determine the utterance intent, which is a classification for distinguishing the aforementioned conversational information. A function to save the conversation log in which the recognition result and the utterance intent are linked to the conversation information, A program that can execute [this action].