Information processing system and program
The information processing system addresses the challenge of coordinating speech timing by providing visual notifications based on situation information about surrounding individuals, enhancing communication efficiency and reducing interference.
Patent Information
- Application Number
- JP2023194489
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-11-15
AI Technical Summary
Speakers often struggle to coordinate their speech timing with third parties, leading to interference or missed opportunities for communication.
An information processing system that acquires situation information about individuals around a target person and generates control information to provide notifications, such as display images, indicating the possibility of speech by surrounding persons, thereby helping the target person to better coordinate their speech timing.
The system facilitates easier timing coordination for speakers by providing visual notifications about potential speech from surrounding individuals, reducing interference and improving communication effectiveness.
Smart Images

Figure 2025081017000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing system and a program.
Background Art
[0002] Patent Document 1 discloses a process of determining, when there is a user who has obtained the line-of-sight of the majority of users among users excluding the user who is looking at the current speaker in a conversation situation, this user as the next speaker. Patent Document 2 discloses an apparatus provided with notification voice storage means for storing, for each conference participant, a notification voice for notifying the next speaker to the conference participants. Patent Document 3 discloses a configuration including chat text input means for receiving input of chat text and voice synthesis means for synthesizing the chat text into chat voice data.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Patent Document 3
Summary of the Invention
Problems to be Solved by the Invention
[0004] When a speaker attempts to communicate with a third party other than themselves, they usually speak at a time that misses the timing of the third party's speech. If a speaker speaks while a third party is speaking, it may interfere with the third party's speech or make it difficult for the third party to recognize the speaker's speech. An object of the present invention is to make it easier for a speaker to specify the timing of a speech that the speaker is about to make, as compared with a configuration in which a notification regarding the speech is not given to the speaker.
Means for Solving the Problems
[0005] The invention according to claim 1 includes a processor, and the processor acquires situation information which is information about the situation of a person around the target person, and when the situation specified by the acquired situation information is a specific situation, generates control information which is control information used for controlling the device of the target person and which causes the device to give a notification indicating that there is a possibility that the person around makes a speech. The invention according to claim 2 is the information processing system according to claim 1, wherein the processor acquires the situation information of the person around based on an image in which the person around is reflected. The invention according to claim 3 is the information processing system according to claim 1, wherein the processor generates control information which causes the device to give a notification indicating that there is no speech when there is a situation where the person around does not make a speech. The invention according to claim 4 is the information processing system according to claim 1, wherein the processor generates, as the control information which causes the device to give the notification indicating that there is a possibility that the speech is made, control information which causes the device to give a notification appealing to the vision of the target person. The invention according to claim 5 is the information processing system according to claim 1, wherein the processor generates, as the control information which causes the device to give the notification indicating that there is a possibility that the speech is made, control information which causes a display image which is an image displayed on the device and indicates that there is a possibility that the person around makes a speech to be displayed on a display unit of the device. The invention according to claim 6 is the information processing system according to claim 5, wherein, after the processor generates control information for causing the display image to be displayed on the display unit of the device, when there is no actual speech by the surrounding person, the processor generates control information for causing the display image being displayed on the display unit to be erased. The invention according to claim 7 is the information processing system according to claim 5, wherein the processor generates, as control information for causing the display image to be displayed on the display unit of the device, control information for causing the display image to be displayed in a form associated with the surrounding person. The invention according to claim 8 is the information processing system according to claim 7, wherein, when a plurality of the surrounding persons are in the specific situation, the processor generates, as control information for causing the display image to be displayed in a form associated with the surrounding person, control information for causing the display image to be displayed in a form associated with each of the plurality of surrounding persons. The invention according to claim 9 is the information processing system according to claim 5, wherein, when the surrounding person who may speak actually speaks, speech content that is the content of the speech is displayed on the display unit of the device, and the processor generates, as control information for causing the display image to be displayed on the display unit of the device, control information for causing the display image to be displayed in a form associated with a display location on the display unit of the device where the speech content is displayed. The invention according to claim 10 is the information processing system according to claim 9, wherein the processor generates, as control information for causing the display image to be displayed in a form associated with the display location, control information for causing the display image having a shape surrounding the display location to be displayed on the display unit. The invention according to claim 11 is the information processing system according to claim 1, wherein, when a surrounding person outside the range of the field of view of the subject who visually recognizes the front of the subject through the device is in the specific situation, the processor generates control information for causing a display image, which is an image indicating that there is a possibility of speech by the surrounding person outside the range of the field of view, to be displayed on the display unit of the device. The invention according to claim 12 is the information processing system according to claim 11, wherein the processor generates, as the control information for causing the display image of the peripheral person outside the range of the field of view to be displayed on the display unit of the device, control information for causing the display image to be displayed on a line connecting the position of the peripheral person when the peripheral person is projected onto a virtual plane along the display unit and the central portion of the display unit. The invention according to claim 13 is the information processing system according to claim 1, wherein the processor further acquires, in addition to the pre-speech situation information, which is the situation information acquired before the actual speech by the peripheral person, in-speech situation information, which is information about the situation of the peripheral person who has actually started speaking, and when the situation specified by the acquired in-speech situation information is in a specific situation, generates control information for causing an end suggestion notification, which is a notification indicating that there is a possibility that the speech of the peripheral person will end and which is control information used for controlling the device, to be performed by the device, or generates control information for causing an end notification, which is a notification indicating that the speech of the peripheral person has ended, to be performed by the device. The invention according to claim 14 is the information processing system according to claim 13, wherein the processor generates control information for causing a predetermined display image to be displayed on the display unit provided in the device when the situation specified by the pre-speech situation information is in the specific situation, and generates, as the control information for causing the end suggestion notification to be performed by the device, control information for changing the display form of the display image being displayed on the display unit. The invention according to claim 15 is the information processing system according to claim 14, wherein the processor generates control information for further changing the display form of the display image whose display form has been changed by the control information for changing the display form, as the control information for causing the end notification to be performed by the device. The invention according to claim 16 is the information processing system according to claim 13, wherein the processor acquires the pre - speech situation information based on the video in which the peripheral person appears, and acquires the during - speech situation information based on the voice information which is information about the voice of the peripheral person. The invention according to claim 17 is an information processing system including a processor, wherein the processor acquires situation information which is information about the situation of a peripheral person who is a person located around a target person and is speaking, and when the situation specified by the acquired situation information is in a specific situation, generates control information which is control information used for controlling a device possessed by the target person and for which a notification indicating that there is a possibility that the speech of the peripheral person will end is to be given by the device, and / or generates control information for which a notification indicating that the speech of the peripheral person has ended is to be given by the device. The invention according to claim 18 is the information processing system according to claim 17, wherein the processor acquires the situation information of the peripheral person based on voice information which is information about the voice of the peripheral person. The invention according to claim 19 is the information processing system according to claim 17, wherein the processor generates, as the control information for which a possibility - suggesting notification which is the notification indicating that there is a possibility that the speech will end is to be given by the device, control information for changing the display image being displayed on the display unit of the device. The invention according to claim 20 is the information processing system according to claim 19, wherein the processor generates, as the control information for which the possibility - suggesting notification is to be given by the device, control information for changing the display image being displayed on the display unit and associated with the display location of the speech content of the peripheral person. The invention according to claim 21 is the information processing system according to claim 17, wherein the processor generates, as the control information for which an end - notification which is the notification indicating that the speech has ended is to be given by the device, control information for changing the display image being displayed on the display unit of the device. According to the invention described in claim 22, the information processing system according to claim 21, wherein the processor generates control information for changing a display image that is displayed on the display unit and is associated with a display location of the speech content of the surrounding person so that the end notification is made by the device. According to the invention described in claim 23, the information processing system according to claim 20 or 22, wherein the processor generates control information for changing at least one of the shape, thickness, and color of the corresponding display image that is displayed in a form surrounding the display location as the control information for changing the corresponding display image that is displayed in association with the display location of the speech content. A program for causing a computer to realize an acquisition function of acquiring situation information that is information about the situation of a surrounding person who is a person located around a target person, and a generation function of generating control information that is used for controlling a device possessed by the target person so that a notification indicating that there is a possibility of speech of the surrounding person is made by the device when the situation specified by the situation information acquired by the acquisition function is in a specific situation. A program for causing a computer to realize an acquisition function of acquiring situation information that is information about the situation of a surrounding person who is a surrounding person located around a target person and is speaking, and a generation function of generating control information that is used for controlling a device possessed by the target person so that a notification indicating that there is a possibility that the speech of the surrounding person will end is made by the device when the situation specified by the situation information acquired by the acquisition function is in a specific situation, and / or generating control information that causes the device to make a notification indicating that the speech of the surrounding person has ended. [Effect of the Invention]
[0006] According to the invention of claim 1, compared with a configuration in which a notification regarding speech is not made to the speaker, it is possible to make it easier for the speaker to specify the timing of the speech that the speaker is about to make. According to the invention of claim 2, it is possible to acquire situation information based on the movements of surrounding people. According to the invention of claim 3, compared with a configuration in which a device does not give a notice indicating the absence of speech, it is easier for the target person to determine whether a surrounding person is speaking. According to the invention of claim 4, compared with a configuration in which a device gives a notice appealing to the target person's sense of hearing, it is easier for a person with a hearing impairment to determine whether there is a possibility that a surrounding person is speaking. According to the invention of claim 5, the target person can obtain information about the possibility of speech by a surrounding person by referring to the display unit of the device. According to the invention of claim 6, it is possible to notify the target person that the possibility of speech by a surrounding person has decreased. According to the invention of claim 7, compared with a configuration in which a display image is not associated with a surrounding person, it is easier for the target person to identify a surrounding person who may be speaking. According to the invention of claim 8, compared with a configuration in which display images are not displayed in a form associated with each of a plurality of surrounding people, it is easier for the target person to identify a surrounding person who may be speaking. According to the invention of claim 9, compared with a configuration in which a display image is displayed without being associated with a display location where the speech content is displayed, it is easier for the target person to recognize the location where the speech content is displayed. According to the invention of claim 10, compared with a configuration in which the display image does not have a shape surrounding the display location, it is easier for the target person to more clearly recognize the location where the speech content is displayed. According to the invention of claim 11, it is also possible to display a display image corresponding to a surrounding person on the display unit even for a surrounding person located at a position outside the visual field range of the target person who visually recognizes the front of himself / herself via the device. According to the invention of claim 12, compared with a configuration in which a display image is displayed at a location deviated from a line connecting the position of this surrounding person and the central part of the display unit when the surrounding person is projected onto a virtual plane along the display unit, it is easier for the target person to identify in which direction the surrounding person is located. According to the invention of claim 13, compared with a configuration in which no notification regarding the end of a speech is made, it is possible to make it easier for the speaker to specify the timing of the speech that the speaker is about to make. According to the invention of claim 14, the target person can determine whether there is a possibility that the speech of the surrounding person will end by referring to the display image displayed on the display unit of the device. According to the invention of claim 15, the target person can determine whether the speech of the surrounding person has ended by referring to the display image displayed on the display unit of the device. According to the invention of claim 16, compared with a configuration in which pre-speech situation information is acquired based on the voice information of the surrounding person, the accuracy of the determination regarding the possibility of the surrounding person's speech can be improved, and compared with a configuration in which in-speech situation information is acquired based on the video in which the surrounding person appears, the accuracy of the determination regarding the end of the surrounding person's speech can be improved. According to the invention of claim 17, compared with a configuration in which no notification regarding the speech is made to the speaker, it is possible to make it easier for the speaker to specify the timing of the speech that the speaker is about to make. According to the invention of claim 18, compared with a configuration in which situation information about the surrounding person who is speaking is acquired based on the video in which the surrounding person appears, the accuracy of the determination regarding the end of the surrounding person's speech can be improved. According to the invention of claim 19, the target person can determine whether there is a possibility that the speech of the surrounding person will end by referring to the display image displayed on the display unit of the device. According to the invention of claim 20, compared with a configuration in which the display image that is not displayed in association with the display location of the speech content of the surrounding person is changed, it is possible to make it easier to identify the target person who may end the speech. According to the invention of claim 21, the target person can determine whether the speech of the surrounding person has ended by referring to the display image displayed on the display unit of the device. According to the invention of claim 22, compared with a configuration in which the display image that is not displayed in association with the display location of the speech content of the surrounding person is changed, it is possible to make it easier to identify the target person who has ended the speech. According to the invention of claim 23, the person concerned can determine whether there is a possibility that the conversation of the surrounding person will end or whether the conversation of the surrounding person has ended by referring to the display image displayed on the display unit of the device. According to the invention of claim 24, compared with a configuration in which no notification regarding the speech is given to the speaker, it is possible to make it easier to specify the timing of the speech that the speaker is about to make. According to the invention of claim 25, compared with a configuration in which no notification regarding the speech is given to the speaker, it is possible to make it easier to specify the timing of the speech that the speaker is about to make.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Mode for Carrying Out the Invention
[0008] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. FIG. 1 is a diagram showing the overall configuration of the information processing system 1 of the present embodiment. The information processing system 1 is provided with a management server 300 as an example of an information processing device. Further, the information processing system 1 is provided with devices 200 to be worn by each of the target persons described later. In FIG. 1, only one device 200 is shown, but a plurality of devices 200 are provided according to the number of target persons. The device 200 is a glasses-type device 200 and is worn on the head of the target person. The target person visually recognizes the surroundings through the device 200.
[0009] Furthermore, in the present embodiment, an overall camera 500 which is a camera for photographing the target person wearing the device 200 and the surrounding persons (described later) located around this target person is provided. The overall camera 500 is provided for each place where the target person is located, and when there are a plurality of target persons, a plurality of overall cameras 500 are also provided. Furthermore, in the present embodiment, individual microphones 600 to be worn by each of the surrounding persons described later are provided. This individual microphone 600 acquires the voices of the surrounding persons and generates voice information. The individual microphones 600 are provided for each surrounding person, and when there are a plurality of surrounding persons, a plurality of individual microphones 600 are provided. Each of the device 200, the overall camera 500, and the individual microphone 600 is connected to the management server 300 through a communication line 400 such as the Internet.
[0010] 〔Configuration of Management Server〕 Figure 2 is a diagram showing the configuration of the management server 300. The management server 300 is implemented by a computer. The management server 300 includes an arithmetic processing unit 111 that executes digital arithmetic processing according to a program, and an information storage unit 19 that stores information. The information storage unit 19 is realized by an existing information storage device such as an HDD (Hard Disk Drive), a semiconductor memory, or a magnetic tape.
[0011] The arithmetic processing unit 111 is provided with a CPU 11a as an example of a processor. The arithmetic processing unit 111 is also provided with a RAM 11b used as a working memory of the CPU 11a and a ROM 11c that stores programs and the like executed by the CPU 11a. The arithmetic processing unit 111 is also provided with a non-volatile memory 11d that is configured to be rewritable and can hold data even when the power supply is interrupted, and an interface unit 11e that controls each unit such as a communication unit connected to the arithmetic processing unit 111.
[0012] The non-volatile memory 11d is composed of, for example, an SRAM backed up by a battery or a flash memory. The information storage unit 19 stores various types of information such as programs executed by the arithmetic processing unit 111. In this embodiment, the CPU 11a provided in the arithmetic processing unit 111 reads the programs stored in the ROM 11c and the information storage unit 19, thereby executing various processes performed by the management server 300.
[0013] The program executed by the CPU 11a can be provided to the management server 300 in a state stored in a computer-readable recording medium such as a magnetic recording medium (magnetic tape, magnetic disk, etc.), an optical recording medium (optical disk, etc.), a magneto-optical recording medium, or a semiconductor memory. Also, the program executed by the CPU 11a may be provided to the management server 300 using communication means such as the Internet.
[0014] 〔Configuration of the Device〕 FIG. 3 is a diagram showing the hardware configuration of device 200. Device 200 includes an arithmetic processing unit 211, an information storage unit 212, a sensor 213, a device camera 214, a device microphone 215, a speaker 216, and a display unit 217. The arithmetic processing unit 211 is provided with a CPU 21a as an example of a processor. In addition, the arithmetic processing unit 211 is provided with a RAM 21c used as a working memory of the CPU 21a and a ROM 21b in which programs executed by the CPU 21a and the like are stored.
[0015] The information storage unit 212 is realized by an existing information storage device such as a semiconductor memory. Examples of the sensor 213 include a GPS sensor and a direction sensor. By referring to the output from this sensor 213, the current position of device 200 and the orientation of device 200 can be specified. The device camera 214 is a camera that photographs the surroundings of device 200. The device camera 214 faces the front direction of the subject and photographs this front direction when device 200 is worn by the subject. In other words, the device camera 214 faces the direction the subject is facing and photographs the front of this subject.
[0016] The device microphone 215 acquires the voice of the subject and generates voice information. The speaker 216 outputs sounds and voices and performs a notification process to the subject on whom device 200 is worn. The display unit 217 is a so-called display and displays various types of information. The display unit 217 is arranged in front of the eyes of the subject when device 200 is worn by the subject. In this embodiment, the video obtained by the device camera 214 is displayed on the display unit 217. When device 200 is worn by the subject, the video showing the state in front of the subject is displayed on the display unit 217. In this embodiment, the subject visually recognizes the area in front of himself / herself by referring to the video reflected on the display unit 217.
[0017] In addition, there is also a transmissive device 200. In this case, as the display unit 217, a transparent display unit 217 is installed so that the subject can visually recognize the area behind the display unit 217. The subject visually recognizes the area behind the display unit 217 through the display unit 217. In other words, the subject visually recognizes the area in front of himself / herself through the display unit 217. In the transmissive device 200, when an image is displayed on the display unit 217, the user visually recognizes both the real space located behind the display unit 217 and the image displayed on the display unit 217.
[0018] The program executed by the CPU 21a can be provided to the device 200 while being stored in a computer-readable recording medium such as a magnetic recording medium (magnetic tape, magnetic disk, etc.), an optical recording medium (optical disk, etc.), a magneto-optical recording medium, or a semiconductor memory. Also, the program executed by the CPU 21a may be provided to the device 200 using a communication means such as the Internet.
[0019] Note that the device 200 is not limited to the glasses-type device 200, and other examples include smartphones, tablet terminals, etc. The glasses-type device 200, smartphones, tablet terminals, etc. are all devices that can be carried by the subject. In smartphones and tablet terminals, a display unit and a device camera are also provided. The subject can visually recognize the area in front of himself / herself by referring to the video captured by the device camera and reflected on the display unit. In other words, in this case, the subject can visually recognize the space located in front of himself / herself, that is, the space located behind the smartphone or tablet terminal, by referring to the display unit provided in the smartphone or tablet terminal placed in front of his / her eyes. In this embodiment, notification processing described below is performed via the device 200. However, the device 200 is not limited to a glasses-type device, and this notification processing can be performed even when using a smartphone or a tablet terminal.
[0020] In this specification, the processor refers to a processor in a broad sense and includes a general-purpose processor (e.g., CPU: Central Processing Unit, etc.) and a dedicated processor (e.g., GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, programmable logic device, etc.). Also, the operation of the processor may be achieved not only by one processor but also by a plurality of physically separated processors cooperating with each other. Also, the order of each operation of the processor is not limited to the order described in this embodiment and may be changed.
[0021] [Explanation of Processing Executed in the Information Processing System] FIGS. 4(A) to 4(D) are diagrams for explaining the processing executed in the information processing system 1 of this embodiment. In FIG. 4(A), a person 41 to whom notification processing is to be performed and a person 42 in the vicinity of the person 41 are shown. In this embodiment, as shown in FIG. 4(A), the device 200 is worn by the person 41 to whom notification processing is to be performed. In this embodiment, as shown in FIG. 4(A), the display unit 217 provided in the device 200 shows a situation in which the person 42 in the vicinity is reflected. Furthermore, in this embodiment, as shown in FIG. 4(A), the individual microphone 600 is worn by the person 42 in the vicinity.
[0022] As described above, the device 200 of this embodiment is a glasses-type device. This glasses-type device 200 is worn on the head of the subject 41. The subject 41 visually recognizes the surrounding person 42 located around him / herself through this device 200. In other words, the subject 41 visually recognizes the surrounding person 42 located in front of him / herself through this device 200. As described above, the device 200 is provided with a device camera 214 (see FIG. 3) and a display unit 217 capable of displaying the video obtained by this device camera 214. The subject 41 visually recognizes the surrounding person 42 by referring to the surrounding person 42 photographed by the device camera 214 and shown on the display unit 217. When the device 200 is the above-described transmissive device, the subject 41 visually recognizes the surrounding person 42 located behind this display unit 217 through the transparent display unit 217.
[0023] In this embodiment, in the state shown in FIG. 4(A), the CPU 11a, which is an example of a processor provided in the management server 300 (see FIG. 2), acquires situation information, which is information about the situation of this surrounding person 42 who is a person located around the subject 41. Specifically, the CPU 11a acquires the situation information of the surrounding person 42 based on the video in which the surrounding person 42 is shown and voice information, which is information about the voice of the surrounding person 42.
[0024] When acquiring the situation information of the surrounding person 42 based on the video in which the surrounding person 42 is shown, the CPU 11a acquires the situation information of the surrounding person 42 based on the video obtained by the device camera 214 provided in the device 200 (see FIG. 3). In this embodiment, the video obtained by the device camera 214 is transmitted to the management server 300 through the communication line 400 (see FIG. 1). The CPU 11a of the management server 300 analyzes this video and acquires the situation information of the surrounding person 42. The CPU 11a acquires the situation information of the surrounding person 42 based on the video in which the surrounding person 42 is shown.
[0025] When acquiring the situation information of the peripheral person 42 based on the voice information of the peripheral person 42, the CPU 11a acquires the situation information of the peripheral person 42 based on the voice information obtained by the individual microphones 600 each worn by the peripheral persons 42. In the present embodiment, the voice information obtained by the individual microphones 600 is transmitted to the management server 300 through the communication line 400. The CPU 11a of the management server 300 analyzes this voice information to acquire the situation information of the peripheral person 42.
[0026] 〔Explanation of Database〕 FIG. 5 is a diagram showing the database stored in the information storage unit 19 (see FIG. 2) of the management server 300. In the present embodiment, as shown in FIG. 5, information about the peripheral person 42 is registered in the database stored in the information storage unit 19 for each peripheral person 42. In the present embodiment, in advance, in the database, for each peripheral person 42, an identification ID which is information used for identifying each of the peripheral persons 42, microphone identification information which is identification information of the individual microphones 600 each possessed by the peripheral persons 42, face information of the peripheral person 42, etc. are registered. In the present embodiment, the face of the peripheral person 42 is photographed in advance. Then, face information which is information of the face of the peripheral person 42 is registered in the database. As the face information, an image of the face of the peripheral person 42 and information about the feature amount of the face of the peripheral person 42 obtained by analyzing this image are registered.
[0027] In the present embodiment, as described above, the video acquired by the device camera 214 and the voice information obtained by the individual microphones 600 are transmitted to the management server 300. The CPU 11a of the management server 300 acquires this video and voice information, and acquires the situation information of the peripheral person 42 based on this video and voice information. Specifically, when the CPU 11a of the management server 300 acquires the video acquired by the device camera 214, based on the image of the face of the peripheral person 42 shown in this video and the face information stored in the database, the peripheral person 42 shown in this video is specified. Furthermore, the CPU 11a of the management server 300 analyzes this video and acquires the situation information of this identified person in the vicinity 42.
[0028] Also, when the CPU 11a of the management server 300 acquires the voice information acquired by the individual microphone 600, based on the microphone identification information transmitted to the management server 300 together with this voice information and the microphone identification information stored in advance in the database, it identifies the person in the vicinity 42 from whom the voice information by this individual microphone 600 was acquired. Furthermore, the CPU 11a of the management server 300 analyzes this voice information and acquires the situation information of this identified person in the vicinity 42.
[0029] Each of the individual microphones 600 stores microphone identification information for identifying each of the individual microphones 600. From the individual microphone 600 to the management server 300, the voice information acquired by the individual microphone 600 and this microphone identification information are transmitted. The CPU 11a of the management server 300 identifies the person in the vicinity 42 from whom the voice information by the individual microphone 600 was acquired based on this microphone identification information and the microphone identification information registered in the database.
[0030] In the present embodiment, an individual microphone 600 (see FIG. 4(A)) is prepared for each person in the vicinity 42. In the present embodiment, the voice information of each person in the vicinity 42 is acquired by this individual microphone 600 prepared for each person in the vicinity 42. In the present embodiment, the voice information obtained by the individual microphone 600 is transmitted to the management server 300 together with the microphone identification information as described above. The CPU 11a of the management server 300 identifies the person in the vicinity 42 based on the microphone identification information, and further acquires the situation information based on the voice information for this identified person in the vicinity 42.
[0031] More specifically, in this embodiment, when transmitting voice information from the individual microphone 600 to the management server 300, the individual microphone 600 selects voice information whose sound pressure exceeds a predetermined threshold from among the voice information obtained by the individual microphone 600. Then, this selected voice information is transmitted from the individual microphone 600 to the management server 300 together with the microphone identification information. Thereby, in this embodiment, it is suppressed that voice information of other peripheral persons 42 different from the peripheral person 42 wearing the individual microphone 600 is transmitted to the management server 300 through this individual microphone 600.
[0032] The other peripheral persons 42 are away from the microphone-wearing peripheral person 42 who is the peripheral person 42 wearing the individual microphone 600, and usually, the sound pressure of the voice of these other peripheral persons 42 acquired by this individual microphone 600 becomes small. In the case of a configuration in which voice information whose sound pressure exceeds a predetermined threshold is selected and the voice information is transmitted to the management server 300, it is suppressed that voice information of other peripheral persons 42 is transmitted to the management server 300 through the individual microphone 600 of the microphone-wearing peripheral person 42.
[0033] Note that the selection of voice information whose sound pressure exceeds a predetermined threshold may be performed by the management server 300. In this case, the management server 300 selects voice information whose sound pressure exceeds a predetermined threshold from among the voice information transmitted from the individual microphone 600. Then, the management server 300 acquires this selected voice information as the voice information of the microphone-wearing peripheral person 42.
[0034] In addition, the CPU 11a of the management server 300 may also acquire the voice information of each of the peripheral persons 42 based on the voice information obtained by a microphone provided in a terminal device owned by each of the peripheral persons 42, such as a smartphone or a tablet terminal owned by each of the peripheral persons 42. In addition, alternatively, a common microphone may be provided, and the CPU 11a of the management server 300 may acquire the voice information of each of the surrounding persons 42 based on the voice information acquired by this common microphone.
[0035] When using a common microphone, in advance, feature information, which is information about the features of the voice of each of the surrounding persons 42, is registered in the database. The CPU 11a of the management server 300 identifies the voice information of each of the surrounding persons 42 based on this feature information registered in the database, and acquires the situation information of each of the surrounding persons 42 based on this voice information.
[0036] 〔Explanation of Specific Processing〕 In the present embodiment, as described above, the CPU 11a of the management server 300 acquires the situation information of the surrounding persons 42 based on the video acquired by the device camera 214 and the voice information obtained by the individual microphone 600. Then, when the situation specified by the acquired situation information is in a specific situation, the CPU 11a generates control information used for controlling the device 200 possessed by the target person 41 (see FIG. 4(A)). More specifically, the CPU 11a generates control information such that a predetermined notification is made to the target person 41 via this device 200 as this control information.
[0037] More specifically, the CPU 11a generates control information such that a notification is made to the target person 41 when the situation specified by the acquired situation information is a situation where there is a possibility of the surrounding person 42 speaking. More specifically, the CPU 11a generates control information such that a notification indicating that there is a possibility of the surrounding person 42 speaking is made by the device 200.
[0038] The CPU 11a determines that there is a possibility of the surrounding person 42 speaking when the situation specified by the acquired situation information is any of the following situations, for example. · When there is an exhalation sound of the surrounding person 42 · When a bystander 42 makes specific sounds such as "hmm", "um", "ah", "eh", etc. · When the expression of the bystander 42 becomes a specific expression, such as when the opening of the mouth of the bystander 42 becomes large · When the bystander 42 visually recognizes the direction in which the target person 41 is located for a time exceeding a predetermined time · When the bystander 42 faces the direction in which the target person 41 is located · When the bystander 42 performs a predetermined specific action, such as nodding, moving their own hand close to their face, or stretching their posture
[0039] In addition, the situation information of the bystander 42 may be obtained based on the biometric information of the bystander 42. Specifically, the CPU 11a may obtain the situation information of the bystander 42 based on biometric information such as pulse, heart rate, and blood pressure obtained by a sensor (not shown) attached to the bystander 42. When obtaining the situation information of the bystander 42 based on biometric information, the biometric information obtained by the sensor and the sensor identification information for each sensor are transmitted from the sensor to the management server 300 through a communication line (not shown).
[0040] The CPU 11a of the management server 300 identifies the bystander 42 based on the sensor identification information, and when the situation specified by the transmitted biometric information is a specific situation, it determines that there is a possibility that the identified bystander 42 will speak. Specifically, for example, when numerical values such as pulse, heart rate, and blood pressure increase, the CPU 11a of the management server 300 determines that there is a possibility that the bystander 42 identified based on the sensor identification information will speak.
[0041] In the processing example shown in FIG. 4, the CPU 11a generates control information for causing a display image indicating the possibility that the bystander 42 (see FIG. 4(A)) may speak to be displayed on the display unit 217 of the device 200 as control information for causing a notification indicating the possibility that the bystander 42 may speak to be made by the device 200. This generated control information is transmitted to device 200, and device 200 performs display control of the display unit 217 based on this control information. Thus, in this embodiment, as shown in FIG. 4(B), a display image 45 indicating the possibility of the peripheral person 42 speaking is displayed on the display unit 217 of device 200.
[0042] This display image 45 is an image representing a so-called "speech bubble". In this embodiment, when it is determined that there is a possibility that the peripheral person 42 will speak, before the display of the speech content described later is performed, this display image 45 composed of an image representing a speech bubble is displayed on the display unit 217 of device 200 as shown in FIG. 4(B). Thereby, the target person 41 recognizes that there is a possibility that the peripheral person 42 will speak. The display image 45 is not limited to a 2D (2 Dimension) image, and may be a 3D (3 Dimension) image. When the display image 45 is a 3D (3 Dimension) image, images with different angles for each eye are displayed on the display unit 217. In other words, when the display image 45 is a 3D (3 Dimension) image, a plurality of images with different viewing angles are displayed on the display unit 217 as the display image 45.
[0043] Note that the generation of the control information may be performed by other devices than the management server 300. In this embodiment, the management server 300 generates the control information, but not limited thereto, the generation of the control information may be performed by a device other than the management server 300. The generation of the control information may be performed by, for example, device 200. When device 200 generates the control information, device 200 determines whether the peripheral person 42 is in a specific situation based on, for example, video of the peripheral person 42 obtained by the device camera 214 (see FIG. 3) that device 200 has and audio information obtained by the individual microphone 600.
[0044] Then, when the peripheral person 42 is in a specific situation, the device 200 generates control information that enables a notification indicating the possibility of the peripheral person 42 speaking to be given by this device 200. Specifically, the device 200 generates control information that causes the display image 45 to be displayed on its own display unit 217. As a result, the display image 45 is displayed on the display unit 217 of the device 200, in the same manner as when the CPU 11a of the management server 300 generates control information.
[0045] The CPU 11a of the management server 300 generates control information such that the display image 45 is displayed in a form associated with the peripheral person 42 reflected in this display unit 217, as control information for causing the display image 45 (see FIG. 4(B)) to be displayed on the display unit 217. As a result, as shown in FIG. 4(B), the display image 45 is displayed in a form associated with the peripheral person 42 reflected in the display unit 217. More specifically, in the present embodiment, the display image 45 is displayed in a form associated with the periphery of the head of the peripheral person 42.
[0046] The CPU 11a of the management server 300 identifies each of the peripheral persons 42 reflected in the display unit 217 based on the video acquired by the device camera 214 (see FIG. 3). Specifically, the CPU 11a of the management server 300 identifies each of the peripheral persons 42 reflected in the display unit 217 of the device 200 based on the video of the peripheral person 42 reflected in the video acquired by the device camera 214 and the face information registered in the database.
[0047] Then, the CPU 11a of the management server 300 generates control information such that the display image 45 is associated with the peripheral person 42 (hereinafter sometimes referred to as the "peripheral person 42 with speaking possibility") among the identified peripheral persons 42 who are determined to have the possibility of speaking. Specifically, when generating this control information, the CPU 11a generates control information including position information, which is information about the display position of the display image 45.
[0048] The CPU 11a determines the position of the potential speaker 42 shown in the video acquired by the device camera 214 as the display position of the display image 45, and generates control information including position information, which is information about this display position. More specifically, the CPU 11a determines the position around the head of the potential speaker 42 on the video acquired by the device camera 214 as the display position of the display image 45. Then, the CPU 11a generates control information including position information, which is information about this determined display position.
[0049] Then, in this embodiment, the control information including this position information is transmitted to the device 200. Then, the device 200 performs display control so that the display image 45 is displayed at the position specified by this position information. As a result, as shown in FIG. 4(B), on the display unit 217 of the device 200, the display image 45 is displayed in a form associated with the potential speaker 42 who may be speaking.
[0050] Note that when the device 200 is the above-described transparent device 200, the CPU 11a of the management server 300 generates control information for causing the display image 45 to be displayed on the portion of the display unit 217 of this device 200 that is located on the straight line connecting the eyes of the target person 41 and the potential speaker 42. In this case, the CPU 11a of the management server 300 first acquires the angle formed by the front direction of the device 200 and the direction from the device 200 toward the potential speaker 42. Specifically, the CPU 11a of the management server 300 analyzes the video acquired by the device camera 214 to acquire the angle formed by the front direction and the direction toward the potential speaker 42.
[0051] Then, based on this angle, the CPU 11a determines the display position of the display image 45 on the display unit 217, and generates control information including information about this determined display position. As a result, even in the transmissive device 200, the display image 45 is displayed in a form associated with the potential speaker 42 who may speak. In this case, the subject 41 will visually recognize the potential speaker 42 existing in the real space and the display image 45 reflected on the display unit 217 and located between the subject's own eyes and the potential speaker 42.
[0052] After that, in this processing example, as shown by the reference numeral 4C in FIG. 4(C), the actual speech by the person 42 starts. In other words, the actual speech by the potential speaker 42 starts. When the actual speech by the person 42 starts, as shown in FIG. 4(C), an image 46 indicating that the speech by the person 42 has started is displayed inside the display image 45 associated with this person 42. When there is an actual speech by the person 42 in the state where the display image 45 is associated, an image 46 indicating that the speech by this person 42 has started is displayed inside this display image 45.
[0053] In other words, in this embodiment, when there is an actual speech by the person 42 in the state where the display image 45 is associated, an image indicating that the acquisition of audio information by the individual microphone 600 has started is displayed inside the display image 45. Whether there is an actual speech by the person 42 in the state where the display image 45 is associated is determined based on, for example, the output from the individual microphone 600 worn by this person 42.
[0054] In this embodiment, an image 46 indicating that the speech by the person 42 has started is displayed within the area surrounded by the display image 45. This image 46 indicating that it has started is displayed on the display unit 217 of the device 200 until the speech content of the peripheral person 42 is acquired by the CPU 11a of the management server 300.
[0055] It takes time to acquire the speech content by the CPU 11a. In this embodiment, until the speech content is acquired by the CPU 11a, an image 46 indicating that the speech by the peripheral person 42 has started is displayed on the display unit 217. Note that the display of this image 46 indicating that it has started is not essential, and the display shown in FIG. 4(D) to be described next may be performed without going through the display shown in FIG. 4(C) from the display shown in FIG. 4(B).
[0056] When there is an actual speech by the peripheral person 42, the CPU 11a of the management server 300 acquires the speech content that is the content of the speech by the peripheral person 42. The CPU 11a of the management server 300 analyzes the voice information transmitted from the individual microphone 600 worn by the peripheral person 42 in a state where the display image 45 is associated, and acquires the speech content of this peripheral person 42. Note that the acquisition of the speech content based on the voice information may be performed using a known method.
[0057] Next, the CPU 11a of the management server 300 generates control information so that the acquired speech content is displayed on the display unit 217 in a form associated with the peripheral person 42 who made the speech with this speech content. Then, the CPU 11a of the management server 300 transmits the generated control information to the device 200.
[0058] Thereby, in this embodiment, as shown in FIG. 4(D), the speech content 48 of the peripheral person 42 is displayed at a predetermined display location 47 in the display unit 217 of the device 200. In this embodiment, when the above-mentioned peripheral person 42 who may have made a speech actually makes a speech, the speech content 48 of this peripheral person 42 is displayed on the display unit 217 of the device 200. In the present embodiment, when the utterance content 48 is to be displayed, the utterance content 48 is displayed inside a display image 45 composed of an image representing a speech balloon. In the present embodiment, the utterance content 48 of the surrounding person 42 is displayed inside the display image 45 that is associated with and displayed for the surrounding person 42 who is determined to have a possibility of speaking.
[0059] In the processing example described above, when there is a possibility that the surrounding person 42 is speaking, the CPU 11a first generates control information to cause the display image 45 to be displayed on the display unit 217 of the device 200 as described above. As a result, as shown in FIG. 4(B), the display image 45 is displayed on the display unit 217. As this control information for causing the display image 45 to be displayed, the CPU 11a generates control information such that the display image 45 is displayed in a form associated with a display location 47 where the utterance content 48 (not shown in FIG. 4(B)) is to be displayed in the display unit 217.
[0060] In the present embodiment, the inside of the display image 45 (see FIG. 4(B)) is the display location 47 where the utterance content 48 is to be displayed. As control information for causing the display image 45 to be displayed, the CPU 11a generates control information such that the display image 45 is associated with this display location 47. More specifically, as control information for causing the display image 45 to be associated with the display location 47, as shown in FIG. 4(B), the CPU 11a generates control information such that an image representing a speech balloon having a shape surrounding the display location 47 is displayed on the display unit 217.
[0061] Then, in the present embodiment, after this control information is generated, when there is an actual utterance of the surrounding person 42, as shown in FIG. 4(D), the utterance content 48 of this surrounding person 42 is displayed in the area surrounded by the display image 45 composed of an image representing a speech balloon. In other words, when there is an actual utterance of the surrounding person 42, the utterance content 48 is displayed at the display location 47 located inside the display image 45.
[0062] The CPU 11a of the management server 300 generates control information that causes the device 200 to issue a notification appealing to the vision of the target person 41, such as causing the display image 45 to be displayed, as control information indicating that there is a possibility of speech. Thus, in this embodiment, the device 200 issues a notification that the target person 41 can visually confirm.
[0063] 〔Form of notification processing〕 Here, the display image 45 is not limited to an image having a shape surrounding the above-described display location 47. The shape of the display image 45 is not particularly limited, and any shape may be used as long as the target person 41 can visually confirm it. As another example of the display image 45, for example, a dot-shaped image can be cited. When the dot-shaped image is to be displayed on the display unit 217, similar to the image having the surrounding shape described above, the dot-shaped image is displayed in a form associated with the surrounding person 42. Also, when the dot-shaped image is to be displayed on the display unit 217, the speech content 48 of the surrounding person 42 is displayed around the dot-shaped image.
[0064] In addition, as the display image 45 displayed on the display unit 217 of the device 200, for example, an image of characters indicating that there is a possibility of speech, such as "There is a possibility of speech", may be displayed. Also, the notification appealing to the vision of the target person 41 is not limited to being by an image, and may be performed by lighting a light source (not shown) provided in the device 200. Also, the notification appealing to the vision of the target person 41 may be performed by changing the color of the entire display screen displayed on the display unit 217 or changing the color of a part of the display screen, such as the edge of the display screen displayed on the display unit 217.
[0065] In addition, as control information for causing the device 200 to issue a notification indicating the possibility of the peripheral person 42 speaking, for example, control information for vibrating a vibration source (not shown) provided in the device 200 may be generated. In this case, the target person 41 recognizes the possibility of the peripheral person 42 speaking based on the vibration of the device 200. In addition, for example, control information for causing a sound or voice to be emitted from a speaker 216 (see FIG. 3) provided in the device 200 may be generated. When the target person 41 has a hearing impairment, it is difficult to notify by sound. However, when the target person 41 does not have a hearing impairment, the target person 41 recognizes the possibility of the peripheral person 42 speaking by the sound.
[0066] When the target person 41 has a hearing impairment, as shown in FIG. 4(D), when the speech content 48 is displayed, the target person 41 can recognize that the peripheral person 42 is speaking. Here, the display of the speech content 48 is performed after the actual speech by the peripheral person 42. If the display of the speech content 48 is delayed, a situation may occur where the speech content 48 has not yet been displayed even though the actual speech by the peripheral person 42 has already started. In this case, a situation may occur where the target person 41 mistakenly recognizes that the peripheral person 42 is not speaking and the target person 41 starts speaking while the peripheral person 42 is speaking.
[0067] In contrast, in the present embodiment, as described above, the target person 41 is notified of the possibility of speaking before the actual speech by the peripheral person 42 starts. In this case, when the target person 41 is notified of the possibility of speaking, the target person 41 becomes cautious about his or her own speech. In this case, a situation where the speech by the target person 41 starts after the start of the speech by the peripheral person 42 is less likely to occur. Note that the information processing system 1 of the present embodiment functions effectively even when the target person 41 is not a person with a hearing impairment. If the subject 41 is notified that there is a possibility that the surrounding person 42 may speak, even if the subject 41 is not hearing impaired, it becomes less likely for the subject 41 to speak after the surrounding person 42 starts speaking.
[0068] [Other display examples] FIG. 6 is a diagram showing other display examples in the display unit 217 of the device 200. FIG. 6 shows a display example in a situation where there is no speech from the surrounding person 42 and there is also no possibility of speech from the surrounding person 42. When there is no speech from the surrounding person 42, the CPU 11a of the management server 300 may generate control information so that a notification indicating the absence of speech is made by the device 200. Accordingly, in this case, as shown in FIG. 6, an image 51 indicating the absence of speech is displayed on the display unit 217 of the device 200.
[0069] This image 51 indicating the absence of speech shown in FIG. 6 is an image composed of the characters "No speech". By displaying the image 51 indicating the absence of speech, the subject 41 recognizes that there is no speech from the surrounding person 42. Note that the image 51 indicating the absence of speech is not limited to an image composed of characters, and may be an image other than an image composed of characters, such as a symbol or a figure.
[0070] Here, assume a case where the possibility of speech from the surrounding person 42 occurs from the situation shown in FIG. 6. In this case, the CPU 11a of the management server 300 generates control information so that the display on the device 200 switches to the display shown in FIG. 4(B). More specifically, the CPU 11a generates control information such that the image 51 indicating the absence of speech is erased and the display image 45 shown in FIG. 4(B) is displayed on the display unit 217. When the display image 45 shown in FIG. 4(B) is displayed, the subject 41 recognizes that there is a possibility that the surrounding person 42 may speak.
[0071] 〔Other processing examples〕 Figures 7(A) to (C) are diagrams showing other processing examples. The processing when there is no actual speech by the surrounding person 42 will be described. In this processing example, as in the above, first, as shown in FIGS. 7(A) and (B), the possibility of speech by the surrounding person 42 occurs, and in response thereto, the display image 45 is displayed on the display unit 217 of the device 200. Thereafter, in this processing example, a situation where there is no actual speech by the surrounding person 42 occurs. In this case, in this processing example, as shown in FIG. 7(C), the display image 45 corresponding to this surrounding person 42 is erased.
[0072] After the CPU 11a generates the above control information for causing the display image 45 to be displayed on the display unit 217, when there is no actual speech by the surrounding person 42, the CPU 11a generates control information for erasing the display image 45 that is being displayed on the display unit 217. More specifically, the CPU 11a generates control information for erasing the display image 45 when there is no actual speech by the surrounding person 42 corresponding to the displayed display image 45 during a predetermined time period after generating the above control information for causing the display image 45 to be displayed, or during a predetermined time period after the display image 45 is displayed on the display unit 217. In this case, accordingly, the device 200 erases the display image 45. As a result, as shown in FIGS. 7(B) and (C), the display image 45 that was being displayed on the display unit 217 is erased.
[0073] 〔Other processing examples〕 Figures 8(A) to (C) are diagrams showing other processing examples. When a plurality of surrounding persons 42 reflected in the display unit 217 of the device 200 are in a specific situation, the CPU 11a of the management server 300 generates, as control information, control information for causing the display image 45 to be displayed in a form in which the display image 45 is associated with each of the plurality of surrounding persons 42. As a result, in this case, as shown in FIG. 8(B), on the display unit 217 of the device 200, the display image 45 is displayed in a form in which the display image 45 is associated with each of the plurality of surrounding persons 42. In this case, the person 41 who refers to the display unit 217 of the device 200 recognizes that there is a possibility of speaking to the plurality of surrounding persons 42.
[0074] After that, when a surrounding person 42 included in the plurality of surrounding persons 42 actually speaks, as indicated by reference numeral 8D in FIG. 8(C), the speech content 48 is displayed within the area surrounded by the display image 45 that is associated with and displayed for the surrounding person 42 who actually spoke. Also, in this processing example shown in FIG. 8, for the other surrounding persons 42 indicated by reference numeral 8E in FIG. 8(C) who did not actually speak, as shown in FIGS. 8(B) and 8(C), the display images 45 that were associated with and displayed for these other surrounding persons 42 are erased.
[0075] Although not shown, in the state shown in FIG. 8(C), if there is a possibility of speaking to the other surrounding person 42 indicated by reference numeral 8E, the display image 45 corresponding to this other surrounding person 42 will be displayed again. In the present embodiment, while a surrounding person 42 indicated by reference numeral 8F, who is one of the surrounding persons 42, is speaking, if there is a possibility of speaking to another surrounding person 42, the display image 45 corresponding to this one of the surrounding persons 42 and the speech content 48 are being displayed, and a new display image 45 corresponding to this other surrounding person 42 is displayed.
[0076] 〔Other Processing Examples〕 FIGS. 9(A) to 9(C) and FIG. 10 are diagrams showing other processing examples. FIG. 10 shows the state when viewing the device 200, the surrounding persons 42, and the person 41 from above in the vertical direction. In the processing examples shown in FIGS. 9 and 10, as shown in FIG. 10, some of the surrounding persons 42 indicated by reference numeral 10B are outside the imaging range 10A by the device camera 214 (not shown in FIG. 10) provided in the device 200. The shooting range 10A can also be regarded as the range of the field of view of the subject 41 who visually recognizes the front of himself / herself via the device 200. In the processing examples shown in FIGS. 9 and 10, in this range of the field of view, some of the surrounding persons 42 indicated by the reference numeral 10B are excluded. Hereinafter, this part of the surrounding person 42 will be referred to as the "non-display surrounding person 42B". As shown in FIG. 9(A), on the display unit 217 of the device 200, two surrounding persons 42 indicated by the reference numeral 9D other than the non-display surrounding person 42B are displayed, and the non-display surrounding person 42B is not shown on the display unit 217 of the device 200.
[0077] Furthermore, in this processing example shown in FIGS. 9 and 10, this non-display surrounding person 42B who is not shown on the display unit 217 of the device 200 is in a specific situation, and there is a possibility that the non-display surrounding person 42B may speak. In this case, the CPU 11a of the management server 300 generates control information so that a display image 45 indicating the possibility that the non-display surrounding person 42B may speak is displayed on the display unit 217. In the present embodiment, even when there is a possibility that the non-display surrounding person 42B may speak, control information is generated so that the display image 45 is displayed on the display unit 217.
[0078] As a result, in this processing example, as shown in FIG. 9(B), a display image 45 corresponding to the non-display surrounding person 42B is displayed on the display unit 217 of the device 200. When the subject 41 refers to the display unit 217 in the state of FIG. 9(B), the subject 41 recognizes that there is a possibility that the surrounding person 42 located at a position outside his / her field of view may speak. When the non-display surrounding person 42B actually speaks, as shown in FIG. 9(C), the speech content 48 of the non-display surrounding person 42B is displayed in a form associated with the display image 45 corresponding to the non-display surrounding person 42B. Also in this processing example, the speech content 48 of the non-display surrounding person 42B is displayed within the area surrounded by the display image 45 corresponding to the non-display surrounding person 42B. In the processing example shown in FIGS. 9 and 10, the speech content 48 of the non-display peripheral person 42B not reflected on the display unit 217 of the device 200 is also displayed on the display unit 217 of the device 200.
[0079] In this processing example shown in FIGS. 9 and 10, when causing a display image 45 corresponding to the non-display peripheral person 42B to be displayed on the display unit 217 of the device 200, as shown in FIG. 9(B), a display is performed so that the direction in which the non-display peripheral person 42B is located can be understood by the target person 41. In FIG. 9(B), the non-display peripheral person 42B is located on the left side of the front of the device 200, and the display image 45 corresponding to the non-display peripheral person 42B is also located on the left side in the figure from the central portion 217C of the display unit 217. In the present embodiment, the display position of the display image 45 corresponding to the non-display peripheral person 42B changes according to the position of the non-display peripheral person 42B.
[0080] FIG. 11 is a view when the display unit 217 and the peripheral person 42 are viewed from the direction indicated by the arrow XI in FIG. 10. As shown in FIG. 11, the CPU 11a generates control information for causing a display image 45 corresponding to the non-display peripheral person 42B to be displayed on a portion of the display unit 217 located on a straight line 11L connecting the non-display peripheral person 42B and the central portion 217C of the display unit 217. This will be described in detail with reference to FIG. 10. Here, an imaginary plane 10K along the display unit 217 of the device 200 is assumed. Further, a line 10H connecting the non-display peripheral person 42B from the center of the target person 41 is assumed.
[0081] Furthermore, here, a case where the non-display peripheral person 42B is projected onto the imaginary plane 10K is assumed. More specifically, a case where the non-display peripheral person 42B is projected onto the imaginary plane 10K in the direction in which the above-described line 10H extends from the location where the non-display peripheral person 42B is located is assumed. In this case, on the imaginary plane 10K, the non-display peripheral person 42B is located at a position indicated by reference numeral 10M. Hereinafter, when the non-display peripheral person 42B is projected onto this virtual plane 10K along the display unit 217, the position of this non-display peripheral person 42B is referred to as "plane position 10M".
[0082] The CPU 11a generates control information for causing the display image 45 (see FIG. 11) of the non-display peripheral person 42B to be displayed on the display unit 217, in a direction located on a straight line 11L connecting the plane position 10M and the central portion 217C of the display unit 217 in the display unit 217. Thereby, the display image 45 is displayed at the location indicated by reference numeral 11X in FIG. 11 on the display unit 217 of the device 200. By referring to the display unit 217 shown in FIG. 11, the subject 41 can identify in which direction the non-display peripheral person 42B who may speak is located.
[0083] When acquiring the situation information about the non-display peripheral person 42B who is not reflected in the display unit 217 of the device 200 based on the video in which this non-display peripheral person 42B is reflected, the situation information about this non-display peripheral person 42B is acquired based on the video obtained by the overall camera 500 (see FIG. 10). Specifically, in this case, based on the video obtained by the overall camera 500 and the face information registered in the database (see FIG. 5), the non-display peripheral person 42B is identified, and based on this video, the situation information of the identified non-display peripheral person 42B is acquired.
[0084] Also, when displaying the display image 45 on the display unit 217 of the device 200, it is necessary to identify the position of the identified non-display peripheral person 42B. In this case, the CPU 11a of the management server 300 analyzes, for example, the video obtained by the overall camera 500 to identify the position of the non-display peripheral person 42B. Furthermore, in this case, the CPU 11a of the management server 300 analyzes the video obtained by the overall camera 500 to identify the position of the central portion 217C of the display unit 217 of the device 200 worn by the subject 41 and the orientation of the device 200.
[0085] Then, based on the position of the non-display peripheral person 42B, the position of the central portion 217C of the display unit 217 of the device 200, and the orientation of the device 200, the CPU 11a of the management server 300 identifies the above-described planar position 10M. Next, based on the identified planar position 10M and the position of the central portion 217C of the display unit 217, the CPU 11a of the management server 300 determines the display position of the display image 45 on the display unit 217.
[0086] Then, the CPU 11a of the management server 300 generates control information including information about the determined display position. The device 200 performs display control on the display unit 217 according to this control information. As a result, on the display unit 217 of the device 200, as shown in FIG. 11, the display image 45 is displayed on the straight line 11L connecting the planar position 10M and the central portion 217C of the display unit 217.
[0087] In addition, when the situation specified by the acquired situation information is a situation where one peripheral person 42 is looking at another peripheral person 42, the CPU 11a may generate control information so that a notification indicating that there is a possibility of speech by the other peripheral person 42 is given by the device 200. In the above, based on the situation information of the peripheral person 42, it is determined whether there is a possibility of speech of this peripheral person 42. However, the present invention is not limited to this, and based on the situation information of one peripheral person 42, it may be determined whether there is a possibility of speech of another peripheral person 42.
[0088] When the CPU 11a determines that there is a possibility of speech of this other peripheral person 42, for example, the CPU 11a generates control information so that the display image 45 is associated with this other peripheral person 42. More specifically, for example, the CPU 11a generates control information so that the display image 45 is associated with this other peripheral person 42 reflected on the display unit 217.
[0089] Specifically, for example, when one peripheral person 42 shown in the video obtained by the device camera 214 continuously visually recognizes other peripheral persons 42 shown in this video for a time exceeding a predetermined time, the CPU 11a determines that the situation of these other peripheral persons 42 is a situation where there is a possibility of speech. And in this case, the CPU 11a generates control information so that the display image 45 is displayed in association with this other peripheral person 42.
[0090] 〔Explanation of the processing flow〕 FIG. 12 is a flowchart showing the processing flow executed when the above notification process is performed. A series of the processing flows described above will be explained. In this embodiment, first, the CPU 11a of the management server 300 determines for each of the peripheral persons 42 located around the target person 41 whether the situation of the peripheral person 42 is in the above specific situation (step S101). And when the CPU 11a determines that the situation of the peripheral person 42 is in a specific situation, the CPU 11a identifies the peripheral person 42 in this specific situation (step S102).
[0091] Thereafter, the CPU 11a generates control information so that the display image 45 corresponding to this peripheral person 42 in the specific situation is displayed on the display unit 217 (step S103). Thereby, a display image 45 indicating that there is a possibility of speech is displayed on the display unit 217 of the device 200. Thereafter, the CPU 11a determines whether the identified peripheral person 42 actually made a speech based on the voice information of the peripheral person 42 identified as being in a specific situation (step S104).
[0092] And when the CPU 11a does not determine that the identified peripheral person 42 actually made a speech, the CPU 11a generates control information so that the display image 45 corresponding to this peripheral person 42 is erased (step S105). Thereby, the display image 45 displayed on the display unit 217 of the device 200 is erased. On the one hand, when the CPU 11a determines that the identified peripheral person 42 has actually spoken, it analyzes the voice information to obtain the speech content 48 corresponding to this peripheral person 42 (step S106).
[0093] Next, the CPU 11a generates control information so that this speech content 48 is displayed within the display image 45 (step S107). In this case, the CPU 11a generates control information so that this speech content 48 is displayed within the display image 45 that is displayed in association with the identified peripheral person 42. As a result, the speech content 48 is displayed within the display image 45.
[0094] 〔Notification process for the end of speech〕 Next, the notification process for the end of speech will be described. Above, the notification process for the possibility of speech has been described. In addition, a notification indicating the possibility of the end of speech or a notification indicating that the speech has ended may be made to the target person 41 through the device 200. In the present embodiment, as described above, the CPU 11a of the management server 300 acquires situation information, which is information about the situation of the peripheral person 42 before actually speaking. And when there is a possibility that the peripheral person 42 will speak, as described above, a notification indicating the possibility of speech is made. Hereinafter, in this specification, this situation information, which is information about the situation of the peripheral person 42 before actually speaking, is referred to as "pre-speech situation information".
[0095] Furthermore, in the process described below, after the peripheral person 42 starts actual speech, the CPU 11a of the management server 300 acquires situation information, which is information about the situation of this peripheral person 42 who is speaking. Hereinafter, in this specification, this situation information about the situation of the peripheral person 42 who is speaking is referred to as "during-speech situation information". When the situation specified by the acquired in-conversation situation information is a specific situation, the CPU 11a also generates control information to be used for controlling the device 200.
[0096] Specifically, as this control information, the CPU 11a generates control information to cause the device 200 to issue a notification indicating that there is a possibility that the speech of the surrounding person 42 will end (hereinafter referred to as an "end possibility suggestion notification"). In other words, as this control information, the CPU 11a generates control information to cause the device 200 to issue an end possibility suggestion notification indicating that there is a possibility that the speech of the surrounding person 42, who is the target of display of the display image 45, will end.
[0097] In addition, when the situation specified by the acquired in-conversation situation information is a specific situation, the CPU 11a generates control information to cause the device 200 to issue a notification indicating that the speech of the surrounding person 42 has ended (hereinafter referred to as an "end notification"). In other words, as this control information, the CPU 11a generates control information to cause the device 200 to issue an end notification indicating that the speech of the surrounding person 42, who is the target of display of the display image 45, has ended.
[0098] 〔End possibility suggestion notification〕 The end possibility suggestion notification will be described. When the situation specified by the acquired in-conversation situation information is a situation where there is a possibility that the speech of the surrounding person 42 will end, the CPU 11a generates control information to cause the device 200 to issue an end possibility suggestion notification. When the situation specified by the acquired in-conversation situation information is, for example, the following situation, the CPU 11a generates control information to cause the device 200 to issue an end possibility suggestion notification. · When the situation specified by the voice information, such as the tone of the voice of the surrounding person 42 dropping or the pitch of the speech decreasing, is a specific situation · When the surrounding person 42 performs a specific predetermined action, such as lowering the hand that the surrounding person 42 had raised
[0099] 〔End Notification〕 Next, the end notification will be explained. When the situation specified by the acquired speaking situation information indicates that the speech of the surrounding person 42 has ended, the CPU 11a generates control information for causing the end notification to be performed by the device 200. Specifically, when the situation specified by the acquired speaking situation information is, for example, in the following situations, the CPU 11a generates control information for causing the end notification to be performed by the device 200. · When voice information is no longer acquired · When the expression of the surrounding person 42 is in a specific state, such as when the mouth of the surrounding person 42 is closed
[0100] The CPU 11a acquires the speaking situation information of the surrounding person 42 based on the voice information of the surrounding person 42 and the video in which the surrounding person 42 is reflected. More specifically, the CPU 11a acquires the speaking situation information of the surrounding person 42 based on the voice information obtained by the individual microphone 600 and the video obtained by the device camera 214 or the overall camera 500. Then, when the situation specified by this speaking situation information is in a predetermined situation, the CPU 11a generates control information for causing a possible end suggestion notification or an end notification to be performed by the device 200.
[0101] Note that the information used for acquiring the pre-speaking situation information and the information used for acquiring the speaking situation information may be different. Specifically, for example, for the pre-speaking situation information, the pre-speaking situation information may be acquired based on the video in which the surrounding person 42 is reflected, and for the speaking situation information, the speaking situation information may be acquired based on the voice information of the surrounding person 42.
[0102] When acquiring the pre-speaking situation information, since the surrounding person 42 often does not make a clear speech, it is easier to improve the accuracy of judgment about the possibility of making a speech by acquiring the pre-speaking situation information based on the video in which the surrounding person 42 is reflected. On the other hand, when acquiring the in-conversation situation information, since the bystander 42 is actually speaking, acquiring the in-conversation situation information based on voice information makes it easier to increase the accuracy of judging the possibility of ending the conversation and the end of the conversation compared to acquiring the in-conversation situation information based on video.
[0103] The above control information for causing the end possibility suggestion notification and the end notification to be performed by the device 200 is transmitted to the device 200 in the same manner as above. Accordingly, in the present embodiment, the device 200 performs control based on this control information to perform an end possibility suggestion notification or an end notification to the target person 41. Thereby, the target person 41 recognizes that the conversation of the bystander 42 is likely to end or that the conversation of the bystander 42 has ended.
[0104] 〔Specific example of processing〕 FIGS. 13(A) to 13(D) are diagrams showing specific examples of processing. FIG. 13(A) shows a situation where the bystander 42 is speaking. In the present embodiment, in the situation shown in FIG. 13(A), the conversation content 48 is displayed within the area surrounded by the display image 45. When the bystander 42 is speaking as shown in FIG. 13(A), the CPU 11a of the management server 300 acquires the in-conversation situation information. More specifically, the CPU 11a acquires the in-conversation situation information of the bystander 42 in a state where the display image 45 is associated based on the video obtained by the overall camera 500 or the device camera 214 and the voice information obtained by the individual microphone 600.
[0105] Then, when the situation specified by the acquired in-conversation situation information is a situation where there is a possibility of ending the conversation of the bystander 42, the CPU 11a generates control information for causing the end possibility suggestion notification to be performed by the device 200. Also, when the situation specified by the acquired in-conversation situation information indicates a situation where the speech of the surrounding person 42 has ended, the CPU 11a generates control information to cause an end notification to be issued by this device 200.
[0106] FIG. 13(B) shows the state of the display unit 217 of the device 200 when the situation specified by the in-conversation situation information is a situation where there is a possibility that the speech of the surrounding person 42 will end. When the situation is such that there is a possibility that the speech of the surrounding person 42 will end, as described above, the CPU 11a generates control information to cause a notification suggesting the possibility of end to be issued by the device 200. In this processing example, as this control information for causing a notification suggesting the possibility of end to be issued by the device 200, the CPU 11a generates control information for changing the display image 45 displayed on the display unit 217 of the device 200. Hereinafter, in this specification, this control information for causing a notification suggesting the possibility of end to be issued by the device 200 is referred to as "first control information".
[0107] In this processing example, as the first control information, the CPU 11a generates control information for changing the display image 45 that is displayed in association with the display location 47 of the speech content 48 of the surrounding person 42 (see FIG. 13(A)). Here, the display image 45 can be regarded as a corresponding display image that is displayed in association with the display location 47 of the speech content 48 of the surrounding person 42. In this embodiment, as this corresponding display image, a display image 45 representing a speech bubble is displayed. As the first control information for causing a notification suggesting the possibility of end to be issued by the device 200, the CPU 11a generates control information for changing the display image 45 representing this speech bubble, which is an example of the corresponding display image.
[0108] More specifically, as this first control information for changing the corresponding display image, the CPU 11a generates control information for changing the shape of the display image 45 representing the speech bubble, which is displayed in a form surrounding the display location 47. Specifically, as the first control information for changing the corresponding display image, the CPU 11a generates control information to cause the protruding portion 45G (see FIG. 13(A)), which is provided as a part of the display image 45 representing the speech bubble, to be erased.
[0109] As a result, in the present embodiment, as shown in FIGS. 13(A) and (B), the protruding portion 45G is erased. In the present embodiment, the protruding portion 45G provided in the display image 45 that is displayed in association with the surrounding person 42 in a situation where there is a possibility that the speech has ended is erased. The target person 41 recognizes that there is a possibility that the speech of the surrounding person 42 has ended by recognizing that the protruding portion 45G has been erased.
[0110] In the present embodiment, as described above, when the situation specified by the pre-speech situation information is a specific situation, the CPU 11a generates control information to cause the display image 45 representing the speech bubble to be displayed on the display unit 217 provided in the device 200. As a result, in the present embodiment, first, as shown in FIG. 13(A), this display image 45 representing the speech bubble is displayed on the display unit 217 of the device 200 in a form associated with the surrounding person 42. This display image 45 is provided with a protruding portion 45G that protrudes toward the surrounding person 42. In the present embodiment, in front of the protruding portion 45G in the protruding direction, the surrounding person 42 who may speak or the surrounding person 42 who is speaking is in a positioned state.
[0111] The CPU 11a generates control information to change the display form of the display image 45 displayed on the display unit 217 as the first control information for causing the end suggestion notification to be performed by the device 200. Specifically, as this first control information, the CPU 11a generates control information to cause the protruding portion 45G of the display image 45 to be non-displayed. As a result, in the present embodiment, as described above, the protruding portion 45G of the display image 45 becomes non-displayed.
[0112] Also, as described above, when the speech of the peripheral person 42 actually ends, the CPU 11a generates control information for causing the device 200 to issue an end notification, which is a notification indicating that the speech has ended. Hereinafter, this control information for causing the device 200 to issue an end notification is referred to as "second control information". Also in this case, the CPU 11a generates, as this second control information, control information for changing the display image 45 displayed on the display unit 217 of the device 200. More specifically, the CPU 11a generates, as the second control information, control information for changing the above-described display image 45 that is displayed in association with the display position 47 of the speech content 48 of the peripheral person 42.
[0113] More specifically, the CPU 11a generates, as the second control information, control information for further changing the display form of the display image 45 after the display form of the display image 45 has been changed by the above-described first control information so that the display form of the display image 45 is changed. More specifically, the CPU 11a generates, as the second control information, control information for changing the thickness of the lines constituting the display image 45.
[0114] As a result, in the present embodiment, as shown in FIGS. 13(B) and (C), the lines constituting the display image 45 displayed on the display unit 217 of the device 200 become thinner. In other words, in the present embodiment, the lines constituting the display image 45 that are displayed in association with the peripheral person 42 who is speaking become thinner. As a result, the target person 41 recognizes that the speech of the peripheral person 42 has ended.
[0115] In the present embodiment, there is a time difference between the timing when the speech of the peripheral person 42 ends and the timing when the speech content 48 at the end of the speech of the peripheral person 42 is displayed on the display unit 217. In this case, in the present embodiment, as shown in FIGS. 13(C) and (D), after the timing of the end of the speech of the peripheral person 42, the display process of the speech content 48 is continuously performed until all of the speech content 48 is displayed. In other words, in the present embodiment, even if the lines constituting the display image 45 become thinner, the display process of the speech content 48 does not end, and this display process is continuously performed until all of the speech content 48 is displayed.
[0116] The target person 41 can also recognize the end of the speech of the peripheral person 42 by recognizing the end of this display process of the speech content 48. By the way, in the present embodiment, even though the speech of the peripheral person 42 has already ended, the display process is continuously performed until all of the speech content 48 is displayed. In this case, the target person 41 is likely to misrecognize that the peripheral person 42 is still speaking even though the speech of the peripheral person 42 has already ended.
[0117] In this case, it is easy for a blank time without speech to occur between the end of the speech of the peripheral person 42 and the start of the speech of the target person 41. On the other hand, when a termination possibility suggestion notification or a termination notification is performed as in the present embodiment, the target person 41 can recognize the end of the speech of the peripheral person 42 at an earlier stage. In this case, the target person 41 can speak his or her own speech at a stage shortly after the speech of the peripheral person 42 ends.
[0118] FIGS. 14(A) to (I) are diagrams showing a series of flows of the display process. In FIG. 14(A), there is no speech of the peripheral person 42 and there is also no possibility of the speech of the peripheral person 42. In this case, no sound is detected. Also, in this case, the CPU 11a does not generate control information for causing the display image 45 to be displayed, and the display image 45 is not displayed on the display unit 217 of the device 200. In FIG. 14(B), a situation where there is a possibility of speech is shown. In this case, the display image 45 is displayed on the display unit 217 of the device 200.
[0119] Figures 14(C) to (F) show the situation while the bystander 42 is speaking. In this case, the CPU 11a acquires the speech content 48, and further generates control information for causing the speech content 48 to be displayed. As a result, as indicated by reference numeral 13X, the speech content 48 is sequentially displayed on the display unit 217 of the device 200. More specifically, the speech content 48 is sequentially displayed inside the display image 45 displayed on the display unit 217 of the device 200. In the present embodiment, there is a time difference between the timing at which the CPU 11a acquires the audio information and the timing at which the speech content 48 is displayed on the display unit 217 of the device 200. Therefore, in the present embodiment, as indicated by the arrow 14Y in FIG. 14, the display of the speech content 48 is performed with a delay relative to the acquisition of the audio information by the CPU 11a.
[0120] FIG. 14(F) shows a situation where there is a possibility that the speech of the bystander 42 has ended. In this case, in the present embodiment, the protruding portion 45G (see FIG. 14(E)) that was displayed as a part of the display image 45 is erased. From FIG. 14(G) onward, the situation where the speech of the bystander 42 has ended is shown. In this case, as shown in FIGS. 14(G) and (H), the lines constituting the display image 45 become thinner.
[0121] In the present embodiment, when there is a possibility that the speech has ended, at least one of the shape, thickness, and color of the display image 45 is changed, and when the speech has actually ended, at least one of the shape, thickness, and color of the display image 45 is further changed. In the above description, the case where the shape of the display image 45 is changed when there is a possibility that the speech has ended and the thickness of the line constituting the display image 45 is changed when the speech has actually ended has been described as an example. In other words, in the above description, the case where the shape of the display image 45 is first changed and then the thickness of the line constituting the display image 45 is changed has been described as an example.
[0122] As another example, for instance, first, the thickness of the lines constituting the display image 45 may be changed, and then, the shape of the display image 45 may be changed. In addition, alternatively, first, the shape of the display image 45 may be changed, and then, the shape of this display image 45 may be further changed. In addition, alternatively, first, the lines constituting the display image 45 may be made thinner, and then, the lines constituting the display image 45 may be made even thinner. Or, first, the lines constituting the display image 45 may be made thicker, and then, the lines constituting the display image 45 may be made even thicker. In addition, alternatively, first, the color of the display image 45 may be changed, and then, the color of the display image 45 may be further changed.
[0123] 〔Form of Notification Processing〕 In the present embodiment, even at the end of a speech, a notification appealing to the vision of the target person 41 is performed in the same manner as in the case of the start of a speech. Specifically, in the present embodiment, as the notification appealing to the vision of the target person 41, as described above, the first change of the display image 45 is performed, and then, the second change of the display image 45 is performed. At the end of a speech, the notification appealing to the vision of the target person 41 may, alternatively, be performed, for example, such that an image of characters such as "There may be an end of speech" or "Speech ended" is displayed on the display unit 217 of the device 200.
[0124] In addition, at the end of a speech, the notification appealing to the vision of the target person 41 may, alternatively, be performed, for example, by turning on or off a light source (not shown) provided in the device 200. In addition, at the end of a speech, the notification appealing to the vision of the target person 41 may, alternatively, be performed by changing the color of the entire display screen displayed on the display unit 217 or by changing the color of a part of the display screen such as the edge of the display screen displayed on the display unit 217.
[0125] In addition, the notification at the end of a speech may, for example, be performed by vibrating a vibration source (not shown) provided in the device 200. In addition, the notification at the end of the conversation may be made, for example, by causing sound or voice to be emitted from the speaker 216 (see FIG. 3) provided in the device 200. When the target person 41 has a hearing impairment, it is difficult to give a notification by sound. However, when the target person 41 does not have a hearing impairment, an end possibility notification or an end notification can be given to the target person 41 by sound.
[0126] 〔Explanation of the processing flow〕 FIG. 15 is a flowchart showing the processing flow when notification processing is also performed at the end of the conversation. Note that the processing in steps S201 to S207 in FIG. 15 is the same as the processing in steps S101 to S107 shown in FIG. 12. In the present embodiment, first, in the same manner as described above, the CPU 11a determines whether the situation specified by the pre-conversation situation information for each of the surrounding persons 42 located around the target person 41 is the above-specified situation (step S201).
[0127] Then, when the CPU 11a determines that the situation specified by the pre-conversation situation information is the specified situation, the CPU 11a specifies the surrounding person 42 in this specified situation (step S202). In other words, when the CPU 11a determines that the situation specified by the pre-conversation situation information is a situation where there is a possibility of conversation, the CPU 11a specifies the surrounding person 42 who has a possibility of conversation.
[0128] Next, the CPU 11a generates control information for causing the display image 45 corresponding to this surrounding person 42 specified as being in the specified situation to be displayed on the display unit 217 (step S203). Thereby, the display image 45 shown in FIG. 14(B) is displayed on the display unit 217 of the device 200 in a form corresponding to the surrounding person 42 who has a possibility of conversation. The displayed display image 45 is provided with a protruding portion 45G.
[0129] After that, based on the voice information of the surrounding person 42 targeted for display of the display image 45, the CPU 11a determines whether this surrounding person 42 has actually spoken (step S204). And when the CPU 11a determines that the surrounding person 42 has not actually spoken, it generates control information to cause the display image 45 to be erased (step S205). As a result, the display image 45 displayed on the display unit 217 of the device 200 is erased. On the other hand, when the CPU 11a determines that the surrounding person 42 has actually spoken, it analyzes the voice information to obtain the speech content 48 (step S206).
[0130] Next, the CPU 11a generates control information to cause the speech content 48 to be displayed within the display image 45 (step S207). As a result, as shown by reference numeral 13X in FIG. 14, the speech content 48 is displayed within the display image 45. In the processing examples shown in FIGS. 13 and 14 above, the case where the display image 45 formed by a thick line is displayed at a stage where there is a possibility of the surrounding person 42 speaking has been described, but the expression form of the display image 45 is not limited to this. At a stage where there is a possibility of the surrounding person 42 speaking, a display image 45 formed by a thin line may be displayed. And when the actual speech by the surrounding person 42 starts, the line constituting the display image 45 may be thickened.
[0131] Next, in step S208, the CPU 11a determines whether there is a possibility that the speech of the surrounding person 42 will end. And when the CPU 11a determines that there is a possibility that the speech of the surrounding person 42 will end, it generates control information to cause the protruding portion 45G to be erased (step S209). As a result, as described above, the protruding portion 45G is erased. Next, the CPU 11a determines whether the speech of the surrounding person 42 has ended (step S210). And when the CPU 11a determines that the speech of the surrounding person 42 has ended, it generates control information to cause the line constituting the display image 45 to become thinner (step S211).
[0132] [Others] In the above, the case where three notification processes are performed has been described, namely, the notification process regarding the possibility of speech before the speech, the notification process regarding the possibility of the end of the speech after the speech has started, and the notification process regarding the end of the speech after the speech has started. It is not essential that all of these three processes are performed, and only one of the processes may be performed, or two of the processes may be performed. In the above, the case where two notification processes are performed at the end of the speech has been described, namely, the notification process regarding the possibility of the end of the speech and the notification process when the speech actually ends. However, at the end of the speech, only one of these two notification processes may be performed.
[0133] (Appendix) (((1))) Comprising a processor, The processor, Obtains situation information, which is information about the situation of a person around the target person, and Generates control information used for controlling the device of the target person, which is control information for causing the device to issue a notification indicating that there is a possibility that the person around the target person will speak when the situation specified by the obtained situation information is a specific situation. An information processing system. (((2))) The processor, Obtains the situation information of the person around the target person based on the video in which the person around the target person appears, The information processing system according to ((1)). (((3))) The processor, Generates control information for causing the device to issue a notification indicating that there is no speech when the person around the target person is in a situation where there is no speech, The information processing system according to ((1)) or ((2)). (((4))) The processor generates control information for causing the device to perform a notification that appeals to the vision of the target person, as the control information for causing the device to perform the notification indicating the possibility of the speech The information processing system according to any one of ((1)) to ((3)). ((5)) The processor generates control information for causing a display image, which is an image displayed on the device and indicates the possibility of speech of the surrounding person, to be displayed on the display unit of the device, as the control information for causing the device to perform the notification indicating the possibility of the speech The information processing system according to any one of ((1)) to ((4)). ((6)) The processor generates control information for causing the display image displayed on the display unit to be erased when there is no actual speech by the surrounding person after generating the control information for causing the display image to be displayed on the display unit of the device The information processing system according to ((5)). ((7)) The processor generates control information for causing the display image to be displayed in a form associated with the surrounding person, as the control information for causing the display image to be displayed on the display unit of the device The information processing system according to ((5)) or ((6)). ((8)) The processor When a plurality of the surrounding persons are in the specific situation, generates control information for causing the display image to be displayed in a form associated with each of the plurality of surrounding persons, as the control information for causing the display image to be displayed in a form associated with the surrounding person The information processing system according to ((7)). ((9)) When the person in the vicinity who may speak actually speaks, the content of the speech, which is the content of the speech, is displayed on the display unit of the device. The processor As control information for causing the display image to be displayed on the display unit of the device, control information is generated so that the display image is displayed in a form associated with the display location on the display unit of the device where the content of the speech is displayed. (((5))) to the information processing system according to any one of (((8))). (((10))) The processor As control information for causing the display image to be displayed in a form associated with the display location, control information is generated so that the display image having a shape surrounding the display location is displayed on the display unit. (((9))) The information processing system according to. (((11))) The processor When a person in the vicinity outside the visual field range of the person viewing the front through the device is in the specific situation, control information is generated so that a display image, which is an image indicating the possibility of speech of the person in the vicinity outside the visual field range, is displayed on the display unit of the device. (((1))) to the information processing system according to any one of (((10))). (((12))) The processor As control information for causing the display image about the person in the vicinity outside the visual field range to be displayed on the display unit of the device, control information is generated so that the display image is displayed on the line connecting the position of the person in the vicinity when projected onto a virtual plane along the display unit and the central portion of the display unit. (((11))) The information processing system according to. (((13))) The processor In addition to the pre-speech situation information, which is the situation information obtained before the actual speech by the surrounding person, speech-in progress situation information, which is information about the situation of the surrounding person who has actually started speaking, is further obtained. When the situation specified by the obtained speech-in progress situation information is a specific situation, control information is generated so that an end suggestion notification, which is control information used for controlling the device and is a notification indicating that the speech of the surrounding person may end, is issued by the device, or control information is generated so that an end notification, which is a notification indicating that the speech of the surrounding person has ended, is issued by the device. (((1))) to (((12))) The information processing system according to any one of the above. (((14))) The processor When the situation specified by the pre-speech situation information is the specific situation, control information is generated so that a predetermined display image is displayed on a display unit provided in the device. As the control information for causing the device to issue the end suggestion notification, control information for changing the display form of the display image displayed on the display unit is generated. (((13))) The information processing system according to the above. (((15))) The processor As the control information for causing the device to issue the end notification, control information for further changing the display form of the display image whose display form has been changed by the control information for changing the display form of the display image is generated. (((14))) The information processing system according to the above. (((16))) The processor Based on the video in which the surrounding person appears, the pre-speech situation information is obtained. Based on voice information, which is information about the voice of the surrounding person, the speech-in progress situation information is obtained. (((13))) to (((15))) The information processing system according to any one of the above. (((17))) Comprising a processor, The processor Obtains situation information, which is information about the situation of a peripheral person who is located around the target person and is speaking, When the situation specified by the obtained situation information is in a specific situation, generates control information used for controlling the device possessed by the target person, the control information being such that a notification indicating that the speech of the peripheral person may end is to be made by the device, and / or generates control information such that a notification indicating that the speech of the peripheral person has ended is to be made by the device, An information processing system. (((18))) The processor Obtains the situation information of the peripheral person based on voice information, which is information about the voice of the peripheral person. The information processing system according to ((17)). (((19))) The processor As the control information for causing the device to issue the possibility suggestion notification, which is the notification indicating that the speech may end, generates control information for changing the display image displayed on the display unit of the device. The information processing system according to ((17)) or ((18)). (((20))) The processor As the control information for causing the device to issue the possibility suggestion notification, generates control information for changing the display image displayed on the display unit and associated with the display location of the speech content of the peripheral person. The information processing system according to ((19)). (((21))) The processor As the control information for causing the device to issue the end notification, which is the notification indicating that the speech has ended, generates control information for changing the display image displayed on the display unit of the device. The information processing system according to any one of ((17)) to ((20)). (((22))) The processor As the control information for causing the end notification to be performed by the device, generate control information for changing a display image that is displayed on the display unit and associated with a display location of the speech content of the surrounding person (((21))) The information processing system according to the above. (((23))) The processor As the control information for changing the corresponding display image that is the display image associated with the display location of the speech content, generate control information for changing at least one of the shape, thickness, and color of the corresponding display image that is displayed in a form surrounding the display location (((20))) or (((22))) The information processing system according to the above. (((24))) An acquisition function for acquiring situation information, which is information about the situation of a person located around the target person; When the situation specified by the situation information acquired by the acquisition function is in a specific situation, generate control information for use in controlling the device of the target person, and the control information is for causing a notification indicating that there is a possibility of speech by the surrounding person to be performed by the device A program for causing a computer to realize the above. (((25))) An acquisition function for acquiring situation information, which is information about the situation of a person located around the target person and who is speaking; When the situation specified by the situation information acquired by the acquisition function is in a specific situation, generate control information for use in controlling the device of the target person, and the control information is for causing a notification indicating that there is a possibility that the speech of the surrounding person will end to be performed by the device, and / or generate control information for causing a notification indicating that the speech of the surrounding person has ended to be performed by the device A program for causing a computer to realize the above.
[0134] According to the information processing system according to ((1)), compared with a configuration in which a notification regarding speech is not given to the speaker, it is possible to make it easier for the speaker to specify the timing of the speech that the speaker is about to make next. According to the information processing system according to ((2)), it is possible to acquire situation information based on the movements of the surrounding people. According to the information processing system according to ((3)), compared with a configuration in which a device does not give a notification indicating the absence of speech, it is easier for the target person to determine whether or not the surrounding people are speaking. According to the information processing system according to ((4)), compared with a configuration in which a device gives a notification that appeals to the hearing of the target person, it is easier for a hearing-impaired person to determine whether or not there is a possibility that the surrounding people are speaking. According to the information processing system according to ((5)), the target person can acquire information about the possibility of speech by the surrounding people by referring to the display unit of the device. According to the information processing system according to ((6)), it is possible to notify the target person that the possibility of speech by the surrounding people has decreased. According to the information processing system according to ((7)), compared with a configuration in which a display image is not associated with the surrounding people, it is easier for the target person to identify the surrounding people who may speak. According to the information processing system according to ((8)), compared with a configuration in which display images are not displayed in a form associated with each of a plurality of surrounding people, it is easier for the target person to identify the surrounding people who may speak. According to the information processing system according to ((9)), compared with a configuration in which the display image is displayed without being associated with the display location where the speech content is displayed, it is easier for the target person to recognize the location where the speech content is displayed. According to the information processing system according to ((10)), compared with a configuration in which the display image does not have a shape surrounding the display location, it is easier for the target person to more clearly recognize the location where the speech content is displayed. According to the information processing system according to ((11)), even for a bystander located at a position outside the field of view of the subject who visually recognizes the front of himself / herself via the device, it is possible to display a display image corresponding to this bystander on the display unit. According to the information processing system according to ((12)), compared with a configuration in which a display image is displayed at a position deviated from a line connecting the position of this bystander and the central portion of the display unit when the bystander is projected onto a virtual plane along the display unit, it becomes easier for the subject to identify in which direction the bystander is located. According to the information processing system according to ((13)), compared with a configuration in which no notification regarding the end of a speech is performed, it is possible to make it easier for the speaker to identify the timing of the speech that the speaker is about to make. According to the information processing system according to ((14)), the subject can determine whether there is a possibility that the speech of the bystander will end by referring to the display image displayed on the display unit of the device. According to the information processing system according to ((15)), the subject can determine whether the speech of the bystander has ended by referring to the display image displayed on the display unit of the device. According to the information processing system according to ((16)), compared with a configuration in which pre-speech situation information is acquired based on the voice information of the bystander, the accuracy of the determination regarding the possibility of the bystander's speech can be improved, and compared with a configuration in which in-speech situation information is acquired based on the video in which the bystander is shown, the accuracy of the determination regarding the end of the bystander's speech can be improved. According to the information processing system according to ((17)), compared with a configuration in which no notification regarding the speech is performed to the speaker, it is possible to make it easier for the speaker to identify the timing of the speech that the speaker is about to make. According to the information processing system according to ((18)), compared with a configuration in which situation information regarding the bystander who is speaking is acquired based on the video in which the bystander is shown, the accuracy of the determination regarding the end of the bystander's speech can be improved. According to the information processing system according to ((19)), the subject can determine whether there is a possibility that the conversation of the surrounding person has ended by referring to the display image displayed on the display unit of the device. According to the information processing system according to ((20)), compared with a configuration in which the display image that is not displayed in association with the display location of the conversation content of the surrounding person is changed, it can be made easier to identify the subject who may end the conversation. According to the information processing system according to ((21)), the subject can determine whether the conversation of the surrounding person has ended by referring to the display image displayed on the display unit of the device. According to the information processing system according to ((22)), compared with a configuration in which the display image that is not displayed in association with the display location of the conversation content of the surrounding person is changed, it can be made easier to identify the subject who has ended the conversation. According to the information processing system according to ((23)), the subject can determine whether there is a possibility that the conversation of the surrounding person has ended or whether the conversation of the surrounding person has ended by referring to the display image displayed on the display unit of the device. According to the program according to ((24)), compared with a configuration in which a notification regarding the conversation is not given to the speaker, it can be made easier to specify the timing of the conversation that the speaker is about to make. According to the program according to ((25)), compared with a configuration in which a notification regarding the conversation is not given to the speaker, it can be made easier to specify the timing of the conversation that the speaker is about to make.
Explanation of Signs
[0135] 1... Information processing system, 10K... Virtual plane, 11a... CPU, 11L... Straight line, 41... Subject, 42... Surrounding person, 45... Display image, 47... Display location, 48... Conversation content, 200... Device, 217... Display unit, 217C... Central part
Claims
1. An information processing system comprising a processor, wherein the processor acquires situation information which is information about the situation of a person in the vicinity of the target person, and generates control information which is control information used for controlling a device possessed by the target person, and which causes a notification indicating that there is a possibility of the person in the vicinity speaking to be made by the device when the situation specified by the acquired situation information is a specific situation.
2. The processor acquires the situation information of the person in the vicinity based on an image in which the person in the vicinity is reflected, The information processing system according to Claim 1.
3. The processor generates control information which causes a notification indicating that there is no speech by the person in the vicinity to be made by the device when the person in the vicinity is in a situation where there is no speech. The information processing system according to Claim 1.
4. The processor generates control information which causes a notification appealing to the vision of the target person to be made by the device as the control information which causes the notification indicating that there is a possibility of the speech to be made by the device. The information processing system according to Claim 1.
5. The processor generates control information which causes a display image, which is an image displayed on the device and indicates that there is a possibility of the person in the vicinity speaking, to be displayed on a display unit of the device as the control information which causes the notification indicating that there is a possibility of the speech to be made by the device. The information processing system according to Claim 1.
6. The processor generates control information which causes the display image displayed on the display unit to be erased when there is no actual speech by the person in the vicinity after generating the control information which causes the display image to be displayed on the display unit of the device. The information processing system according to Claim 5.
7. The processor generates control information which causes the display image to be displayed in a form associated with the person in the vicinity as the control information which causes the display image to be displayed on the display unit of the device. The information processing system according to Claim 5.
8. The processor generates control information which causes the display image to be displayed in a form associated with each of the plurality of persons in the vicinity when the plurality of persons in the vicinity are in the specific situation, as the control information which causes the display image to be displayed in a form associated with the person in the vicinity. The information processing system according to claim 7.
9. When the person around who may speak actually speaks, the content of the speech, which is the content of the speech, is displayed on the display unit of the device. The processor is As the control information for causing the display image to be displayed on the display unit of the device, control information is generated so that the display image is displayed in a form associated with the display location on the display unit of the device where the content of the speech is displayed. The information processing system according to claim 5.
10. The processor is As the control information for causing the display image to be displayed in a form associated with the display location, control information is generated so that the display image having a shape surrounding the display location is displayed on the display unit. The information processing system according to claim 9.
11. The processor is When a person around who is outside the range of the field of view of the person viewing in front of himself / herself through the device is in the specific situation, control information is generated so that a display image, which is an image indicating that there is a possibility that the person around who is outside the range of the field of view may speak, is displayed on the display unit of the device. The information processing system according to claim 1.
12. The processor is As the control information for causing the display image about the person around who is outside the range of the field of view to be displayed on the display unit of the device, control information is generated so that the display image is displayed on the line connecting the position of the person around when projected onto a virtual plane along the display unit and the central part of the display unit. The information processing system according to claim 11.
13. The processor is In addition to the pre-speech situation information, which is the situation information acquired before the actual speech by the person around, information about the situation of the person around who has actually started speaking, which is in-speech situation information, is further acquired. When the situation specified by the acquired in-speech situation information is in a specific situation, control information is generated so that an end suggestion notification, which is a notification indicating that there is a possibility that the speech of the person around will end, which is control information used for controlling the device, is performed by the device, or control information is generated so that an end notification, which is a notification indicating that the speech of the person around has ended, is performed by the device. The information processing system according to claim 1.
14. The processor is When the situation specified by the pre-speech situation information is the specified situation, control information is generated so that a predetermined display image is displayed on the display unit provided in the device. As the control information for causing the end suggestion notification to be performed by the device, control information for changing the display form of the display image displayed on the display unit is generated. The information processing system according to claim 13.
15. The processor is As the control information for causing the end notification to be performed by the device, control information is generated so that the display form of the display image whose display form has been changed by the control information for changing the display form is further changed. The information processing system according to claim 14.
16. The processor is Based on the video in which the surrounding person appears, the pre-speech situation information is acquired. Based on the voice information which is information about the voice of the surrounding person, the in-speech situation information is acquired. The information processing system according to claim 13.
17. Comprising a processor, The processor is Situation information which is information about the situation of a surrounding person who is a person located around the target person and is speaking is acquired. When the situation specified by the acquired situation information is in a specific situation, control information which is control information used for controlling the device possessed by the target person and for which a notification indicating that there is a possibility that the speech of the surrounding person will end is performed by the device is generated, and / or control information for causing a notification indicating that the speech of the surrounding person has ended to be performed by the device is generated. Information processing system.
18. The processor is Based on the voice information which is information about the voice of the surrounding person, the situation information of the surrounding person is acquired. The information processing system according to claim 17.
19. The processor is As the control information for causing the possibility suggestion notification which is the notification indicating that there is a possibility that the speech will end to be performed by the device, control information for changing the display image displayed on the display unit of the device is generated. The information processing system according to claim 17.
20. The processor is As the control information for causing the possibility suggestion notification to be performed by the device, control information for changing the display image displayed on the display unit and associated with the display location of the speech content of the surrounding person is generated. The information processing system according to claim 19.
21. The processor generates control information for changing a display image displayed on a display unit of the device, as the control information for causing the device to issue an end notification indicating that the utterance has ended. The information processing system according to claim 17.
22. The processor generates control information for changing a display image that is displayed on the display unit and is associated with a display location of the utterance content of the surrounding person, as the control information for causing the end notification to be issued by the device. The information processing system according to claim 21.
23. The processor generates control information for changing at least one of the shape, thickness, and color of the corresponding display image that is displayed in a form surrounding the display location, as the control information for changing the corresponding display image that is displayed in association with the display location of the utterance content. The information processing system according to claim 20 or 22.
24. An acquisition function for acquiring situation information, which is information about the situation of a person located around the target person; A generation function for generating control information that is used to control the device possessed by the target person and that causes the device to issue a notification indicating that there is a possibility that the surrounding person will speak, when the situation specified by the situation information acquired by the acquisition function is in a specific situation. A program for causing a computer to implement the above.
25. An acquisition function for acquiring situation information, which is information about the situation of a person located around the target person and who is speaking; A generation function for generating control information that is used to control the device possessed by the target person and that causes the device to issue a notification indicating that there is a possibility that the utterance of the surrounding person will end, and / or control information that causes the device to issue a notification indicating that the utterance of the surrounding person has ended, when the situation specified by the situation information acquired by the acquisition function is in a specific situation. A program for causing a computer to implement the above.
Citation Information
Patent Citations
Method, device, and program for detecting next speaker
JP2006338493A
Next speaker guidance system, next speaker guidance method and next speaker guidance program
JP2012146072A
Speech controller and electronic apparatus
JP2016224393A
Audio processing apparatus, audio processing method, and audio processing program
JP2019101385A
Voice interaction device
JP5405381B2
Cited By
Gear machining apparatus
US12350749B2