Information processing system and program

The information processing system addresses the challenge of coordinating speech timing by displaying and managing speech content on a device worn by the target person, enhancing timing coordination and reducing interference.

JP2025081287AActive Publication Date: 2025-05-27FUJIFILM BUSINESS INNOVATION CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024223938
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-05-27
Estimated Expiration
2043-11-15

AI Technical Summary

Technical Problem

Speakers often struggle to coordinate their speech timing with third parties, leading to potential interference or missed recognition of their speech.

Method used

An information processing system that acquires voice information from individuals around a target person and displays the speech content on a device worn by the target person, changing the display image when the surrounding speech ends or is about to end.

Benefits of technology

Facilitates easier timing coordination for speakers by visually indicating when surrounding speech has concluded or is nearing its end, reducing the likelihood of interference or missed recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025081287000001_ABST
    Figure 2025081287000001_ABST
Patent Text Reader

Abstract

To facilitate determination of a starting timing of a speech for a speaker compared to a configuration in which a notification related to a speech is not issued to the speaker.SOLUTION: Fig (B) shows a state of a display unit 217 of a device 200 when a status specified by speech status information indicates that a speech of a nearby person 42 is likely to end. When the speech of the nearby person 42 is likely to end, a CPU generates control information so that the device 200 displays an end possibility sign notification. In this processing example, the CPU generates the control information to change a display image 45 displayed on the display unit 217 of the device 200, the control information controlling the device 200 to display the end possibility sign notification.SELECTED DRAWING: Figure 13
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing system and a program.

Background Art

[0002] Patent Document 1 discloses a process of determining the next speaker as a user who has obtained the line of sight of the majority of users among the users excluding the user who is looking at the current speaker in a conversation situation. Patent Document 2 discloses an apparatus including notification voice storage means for storing a notification voice for notifying the next speaker to each conference participant. Patent Document 3 discloses a configuration including chat text input means for receiving input of chat text and voice synthesis means for synthesizing the chat text into chat voice data.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0004] When a speaker attempts to communicate with a third party other than themselves, they usually speak at a time that misses the timing of the third party's speech. If a speaker speaks while a third party is speaking, it may interfere with the third party's speech or make it difficult for the speaker's speech to be recognized by the third party. An object of the present invention is to make it easier for a speaker to specify the timing of a speech that the speaker is about to make, as compared with a configuration in which a notification regarding the speech is not given to the speaker.

Means for Solving the Problem

[0005] The invention according to claim 1 includes a processor that acquires voice information of a person around a target person, and based on the voice information, displays on a display unit of a device that the target person has, the speech content that is the content of the speech of the person around who is speaking. When the speech of the person around ends or there is a possibility that it will end, it is an information processing system that changes a display image that is displayed in association with a display location of the speech content displayed on the display unit. The invention according to claim 2 is the information processing system according to claim 1, wherein the processor changes the shape of the display image when the speech of the person around ends or there is a possibility that it will end. The invention according to claim 3 is the information processing system according to claim 2, wherein the processor changes the shape of the display image having a protrusion when the speech of the person around ends or there is a possibility that it will end, and makes the protrusion not present in the display image. The invention according to claim 4 is the information processing system according to claim 1, wherein the processor changes the color of the display image when the speech of the person around ends or there is a possibility that it will end. The invention according to claim 5 is the information processing system according to claim 1, wherein the processor changes the thickness of a line constituting the display image when the speech of the person around ends or there is a possibility that it will end. The invention according to claim 6 is the information processing system according to claim 5, wherein the processor makes the line constituting the display image thinner when the speech of the person around ends or there is a possibility that it will end. The invention according to claim 7 is a program for causing a computer to realize an acquisition function of acquiring voice information of a person around a target person, a display function of displaying, on a display unit of a device possessed by the target person, the content of the speech of the person around who is speaking, based on the voice information, and a change function of changing a display image that is displayed in association with a display location of the speech content displayed on the display unit when the speech of the person around ends or may end.

Advantages of the Invention

[0006] According to the inventions of claims 1 to 7, compared with a configuration in which a notification regarding speech is not given to the speaker, it is possible to make it easier to specify the timing of the speech that the speaker is about to make.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Embodiments for Carrying Out the Invention

[0008] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. FIG. 1 is a diagram showing the overall configuration of the information processing system 1 of the present embodiment. The information processing system 1 is provided with a management server 300 as an example of an information processing apparatus. Further, the information processing system 1 is provided with devices 200 to be worn by each of the target persons described later. In FIG. 1, only one device 200 is shown, but a plurality of devices 200 are provided according to the number of target persons. The device 200 is a glasses-type device 200 and is worn on the head of the target person. The target person visually recognizes the surroundings through the device 200.

[0009] Furthermore, in the present embodiment, an overall camera 500 which is a camera for photographing the target person wearing the device 200 and the surrounding persons (described later) located around this target person is provided. The overall camera 500 is provided for each place where the target person is present, and when there are a plurality of target persons, a plurality of overall cameras 500 are also provided. Furthermore, in the present embodiment, individual microphones 600 to be worn by each of the surrounding persons described later are provided. This individual microphone 600 acquires the voices of the surrounding persons and generates voice information. The individual microphones 600 are provided for each surrounding person, and when there are a plurality of surrounding persons, a plurality of individual microphones 600 are provided. Each of the machine 200, the overall camera 500, and the individual microphone 600 is connected to the management server 300 through a communication line 400 such as the Internet.

[0010] 〔Configuration of Management Server〕 FIG. 2 is a diagram showing the configuration of the management server 300. The management server 300 is realized by a computer. The management server 300 includes an arithmetic processing unit 111 that executes digital arithmetic processing according to a program, and an information storage unit 19 that stores information. The information storage unit 19 is realized by an existing information storage device such as an HDD (Hard Disk Drive), a semiconductor memory, or a magnetic tape.

[0011] The arithmetic processing unit 111 is provided with a CPU 11a as an example of a processor. In addition, the arithmetic processing unit 111 is provided with a RAM 11b used as a working memory of the CPU 11a and a ROM 11c that stores programs executed by the CPU 11a. In addition, the arithmetic processing unit 111 is provided with a non-volatile memory 11d configured to be rewritable and capable of holding data even when the power supply is interrupted, and an interface unit 11e that controls each unit such as a communication unit connected to the arithmetic processing unit 111.

[0012] The non-volatile memory 11d is composed of, for example, an SRAM backed up by a battery or a flash memory. The information storage unit 19 stores various types of information such as programs executed by the arithmetic processing unit 111. In this embodiment, the CPU 11a provided in the arithmetic processing unit 111 reads the programs stored in the ROM 11c and the information storage unit 19, thereby executing various processes performed by the management server 300.

[0013] The program executed by the CPU 11a can be provided to the management server 300 while being stored in a computer-readable recording medium such as a magnetic recording medium (magnetic tape, magnetic disk, etc.), an optical recording medium (optical disk, etc.), a magneto-optical recording medium, or a semiconductor memory. Also, the program executed by the CPU 11a may be provided to the management server 300 using communication means such as the Internet.

[0014] 〔Configuration of the Device〕 FIG. 3 is a diagram showing the hardware configuration of the device 200. The device 200 includes an arithmetic processing unit 211, an information storage unit 212, a sensor 213, a device camera 214, a device microphone 215, a speaker 216, and a display unit 217. The arithmetic processing unit 211 is provided with a CPU 21a as an example of a processor. Also, the arithmetic processing unit 211 is provided with a RAM 21c used as a working memory etc. of the CPU 21a, and a ROM 21b in which programs etc. executed by the CPU 21a are stored.

[0015] The information storage unit 212 is realized by an existing information storage device such as a semiconductor memory. Examples of the sensor 213 include a GPS sensor and a direction sensor. By referring to the output from this sensor 213, the current position of the device 200 and the orientation of the device 200 can be specified. The device camera 214 is a camera that photographs the surroundings of the device 200. The device camera 214 faces the front direction of the subject when the device 200 is worn by the subject, and photographs this front direction. In other words, the device camera 214 faces the direction the subject is facing and photographs the front of this subject.

[0016] The device microphone 215 acquires the voice of the subject and generates voice information. The speaker 216 outputs sounds and voices and performs notification processing to the subject on whom the device 200 is worn. The display unit 217 is a so-called display and displays various types of information. The display unit 217 is arranged in front of the eyes of the subject when the device 200 is worn by the subject. In the present embodiment, the video obtained by the device camera 214 is displayed on the display unit 217. When the device 200 is worn by the subject, a video showing the state in front of the subject is displayed on the display unit 217. In the present embodiment, the subject visually recognizes the front of himself / herself by referring to the video reflected on the display unit 217.

[0017] In addition, there is also a transmissive device 200. In this case, as the display unit 217, a transparent display unit 217 that allows the subject to visually recognize the back of the display unit 217 is installed. The subject visually recognizes the area behind the display unit 217 through this display unit 217. In other words, the subject visually recognizes the front of himself / herself through this display unit 217. In the transmissive device 200, when an image is displayed on the display unit 217, the user will visually recognize both the real space located behind the display unit 217 and the image displayed on this display unit 217.

[0018] The program executed by the CPU 21a can be provided to the device 200 in a state of being stored in a computer-readable recording medium such as a magnetic recording medium (magnetic tape, magnetic disk, etc.), an optical recording medium (optical disk, etc.), a magneto-optical recording medium, or a semiconductor memory. Also, the program executed by the CPU 21a may be provided to the device 200 using communication means such as the Internet.

[0019] Note that the device 200 is not limited to the glasses-type device 200, and other examples include smartphones, tablet terminals, etc. The glasses-type device 200, smartphones, tablet terminals, etc. are all devices that can be carried by the subject. In smartphones and tablet terminals, a display unit and a device camera are also provided. The subject can visually recognize the front of himself / herself by referring to the video captured by the device camera and reflected on the display unit. In other words, in this case, the subject can visually recognize the space in front of himself / herself, which is located behind the smartphone or tablet terminal, by referring to the display unit provided on the smartphone or tablet terminal placed in front of his / her eyes. In this embodiment, the notification process described later is performed via the device 200. However, this notification process can be performed not only for the glasses-type device 200 but also when using a smartphone or a tablet terminal.

[0020] In this specification, the processor refers to a processor in a broad sense, including a general-purpose processor (e.g., CPU: Central Processing Unit, etc.) and a dedicated processor (e.g., GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, programmable logic device, etc.). Moreover, the operation of the processor may be achieved not only by one processor but also by a plurality of physically separated processors cooperating with each other. Also, the order of each operation of the processor is not limited to the order described in this embodiment and may be changed.

[0021] 〔Explanation of the process executed in the information processing system〕 Figs. 4(A) to (D) are diagrams for explaining the process executed in the information processing system 1 of this embodiment. Fig. 4(A) shows the subject 41 to whom the notification process is to be performed and the peripheral person 42 who is located around the subject 41. In this embodiment, as shown in Fig. 4(A), the device 200 is worn by the subject 41 to whom the notification process is to be performed. In this embodiment, as shown in Fig. 4(A), the situation where the peripheral person 42 is reflected on the display unit 217 provided on this device 200 is shown. Furthermore, in this embodiment, as shown in Fig. 4(A), the individual microphone 600 is worn by the peripheral person 42.

[0022] As described above, the device 200 of this embodiment is a glasses-type device. This glasses-type device 200 is worn on the head of the subject 41. The subject 41 visually recognizes the surrounding person 42 located around himself / herself through this device 200. In other words, the subject 41 visually recognizes the surrounding person 42 located in front of himself / herself through this device 200. As described above, the device 200 is provided with a device camera 214 (see FIG. 3) and a display unit 217 capable of displaying the video obtained by this device camera 214. The subject 41 visually recognizes the surrounding person 42 by referring to the surrounding person 42 photographed by the device camera 214 and shown on the display unit 217. When the device 200 is the above-described transmissive device, the subject 41 visually recognizes the surrounding person 42 located behind this display unit 217 through the transparent display unit 217.

[0023] In this embodiment, in the state shown in FIG. 4(A), the CPU 11a, which is an example of the processor provided in the management server 300 (see FIG. 2), acquires situation information, which is information about the situation of this surrounding person 42 who is a person located around the subject 41. Specifically, the CPU 11a acquires the situation information of the surrounding person 42 based on the video in which the surrounding person 42 is shown and the audio information, which is information about the voice of the surrounding person 42.

[0024] When acquiring the situation information of the surrounding person 42 based on the video in which the surrounding person 42 is shown, the CPU 11a acquires the situation information of the surrounding person 42 based on the video obtained by the device camera 214 (see FIG. 3) provided in the device 200. In this embodiment, the video obtained by the device camera 214 is transmitted to the management server 300 through the communication line 400 (see FIG. 1). The CPU 11a of the management server 300 analyzes this video and acquires the situation information of the surrounding person 42. The CPU 11a acquires the situation information of the surrounding person 42 based on the video in which the surrounding person 42 is shown.

[0025] When acquiring the situation information of the peripheral person 42 based on the voice information of the peripheral person 42, the CPU 11a acquires the situation information of the peripheral person 42 based on the voice information obtained by the individual microphones 600 each worn by the peripheral person 42. In the present embodiment, the voice information obtained by the individual microphones 600 is transmitted to the management server 300 through the communication line 400. The CPU 11a of the management server 300 analyzes this voice information to acquire the situation information of the peripheral person 42.

[0026] 〔Explanation of Database〕 FIG. 5 is a diagram showing the database stored in the information storage unit 19 (see FIG. 2) of the management server 300. In the present embodiment, as shown in FIG. 5, information about the peripheral person 42 is registered in the database stored in the information storage unit 19 for each peripheral person 42. In the present embodiment, in advance, in the database, for each peripheral person 42, an identification ID which is information used for identifying each of the peripheral persons 42, microphone identification information which is identification information of the individual microphones 600 each owned by the peripheral person 42, face information of the peripheral person 42, etc. are registered. In the present embodiment, the face of the peripheral person 42 is photographed in advance. And face information which is information of the face of the peripheral person 42 is registered in the database. As the face information, an image of the face of the peripheral person 42 and information about the feature amount of the face of the peripheral person 42 obtained by analyzing this image are registered.

[0027] In the present embodiment, as described above, the video acquired by the device camera 214 and the voice information obtained by the individual microphones 600 are transmitted to the management server 300. The CPU 11a of the management server 300 acquires this video and voice information, and acquires the situation information of the peripheral person 42 based on this video and voice information. Specifically, when the CPU 11a of the management server 300 acquires the video acquired by the device camera 214, based on the image of the face of the peripheral person 42 shown in this video and the face information stored in the database, the peripheral person 42 shown in this video is specified. Furthermore, the CPU 11a of the management server 300 analyzes this video and obtains the situation information of this identified person 42 in the vicinity.

[0028] Also, when the CPU 11a of the management server 300 obtains the voice information acquired by the individual microphone 600, based on the microphone identification information transmitted to the management server 300 together with this voice information and the microphone identification information pre-stored in the database, the CPU 11a identifies the person 42 in the vicinity from whom the voice information by this individual microphone 600 was obtained. Furthermore, the CPU 11a of the management server 300 analyzes this voice information and obtains the situation information of this identified person 42 in the vicinity.

[0029] Each of the individual microphones 600 stores microphone identification information for identifying each of the individual microphones 600. From the individual microphone 600 to the management server 300, the voice information acquired by the individual microphone 600 and this microphone identification information are transmitted. The CPU 11a of the management server 300 identifies the person 42 in the vicinity from whom the voice information by the individual microphone 600 was obtained based on this microphone identification information and the microphone identification information registered in the database.

[0030] In this embodiment, an individual microphone 600 (see FIG. 4(A)) is prepared for each person 42 in the vicinity. In this embodiment, the voice information of each person 42 in the vicinity is acquired by this individual microphone 600 prepared for each person 42 in the vicinity. In this embodiment, the voice information obtained by the individual microphone 600 is transmitted to the management server 300 together with the microphone identification information as described above. The CPU 11a of the management server 300 identifies the person 42 in the vicinity based on the microphone identification information, and further obtains the situation information based on the voice information for this identified person 42 in the vicinity.

[0031] More specifically, in the present embodiment, when transmitting voice information from the individual microphone 600 to the management server 300, the individual microphone 600 selects voice information whose sound pressure exceeds a predetermined threshold from the voice information obtained by the individual microphone 600. Then, this selected voice information is transmitted from the individual microphone 600 to the management server 300 together with the microphone identification information. Thereby, in the present embodiment, it is suppressed that voice information of other peripheral persons 42 different from the peripheral person 42 wearing the individual microphone 600 is transmitted to the management server 300 through this individual microphone 600.

[0032] The other peripheral persons 42 are away from the microphone-wearing peripheral person 42 who is the peripheral person 42 wearing the individual microphone 600, and usually, the sound pressure of the voice of this other peripheral person 42 acquired by this individual microphone 600 becomes small. In the case of a configuration in which voice information whose sound pressure exceeds a predetermined threshold is selected and the voice information is transmitted to the management server 300, it is suppressed that voice information of other peripheral persons 42 is transmitted to the management server 300 through the individual microphone 600 of the microphone-wearing peripheral person 42.

[0033] Note that the selection of voice information whose sound pressure exceeds a predetermined threshold may be performed by the management server 300. In this case, the management server 300 selects voice information whose sound pressure exceeds a predetermined threshold from the voice information transmitted from the individual microphone 600. Then, the management server 300 acquires this selected voice information as the voice information of the microphone-wearing peripheral person 42.

[0034] In addition, the CPU 11a of the management server 300 may also acquire the voice information of each peripheral person 42 based on the voice information obtained by the microphones provided in the terminal devices of each peripheral person 42, such as the smartphones or tablet terminals of each peripheral person 42. Alternatively, in addition, a common microphone may be provided, and the CPU 11a of the management server 300 may acquire the voice information of each of the surrounding persons 42 based on the voice information acquired by this common microphone.

[0035] When using a common microphone, in advance, characteristic information, which is information about the characteristics of the voice of each of the surrounding persons 42, is registered in the database. The CPU 11a of the management server 300 identifies the voice information of each of the surrounding persons 42 based on this characteristic information registered in the database, and acquires the situation information of each of the surrounding persons 42 based on this voice information.

[0036] 〔Explanation of specific processing〕 In the present embodiment, as described above, the CPU 11a of the management server 300 acquires the situation information of the surrounding persons 42 based on the video acquired by the device camera 214 and the voice information obtained by the individual microphone 600. And when the situation specified by the acquired situation information is in a specific situation, the CPU 11a generates control information used for controlling the device 200 possessed by the target person 41 (see FIG. 4(A)). More specifically, the CPU 11a generates control information so that a predetermined notification is made to the target person 41 via this device 200 as this control information.

[0037] More specifically, when the situation specified by the acquired situation information is a situation where there is a possibility that the surrounding person 42 will speak, the CPU 11a generates control information so that a notification is made to the target person 41. More specifically, the CPU 11a generates control information so that a notification indicating that there is a possibility that the surrounding person 42 will speak is made by the device 200.

[0038] The CPU 11a determines that there is a possibility that the surrounding person 42 will speak when the situation specified by the acquired situation information is any of the following situations, for example. · When there is an exhalation sound of the surrounding person 42 · When a bystander 42 makes specific voices such as "hmm", "um", "ah", "eh", etc. · When the expression of the bystander 42 becomes a specific expression, such as the opening of the mouth of the bystander 42 becoming larger · When the bystander 42 visually recognizes the direction where the target person 41 is located for a time exceeding a predetermined time · When the bystander 42 faces the direction where the target person 41 is located · When the bystander 42 makes a predetermined specific movement, such as nodding, moving their own hand close to the face, or stretching their posture

[0039] In addition, the situation information of the bystander 42 may be obtained based on the biological information of the bystander 42. Specifically, the CPU 11a may obtain the situation information of the bystander 42 based on biological information such as pulse, heart rate, and blood pressure obtained by a sensor (not shown) worn by the bystander 42. When obtaining the situation information of the bystander 42 based on biological information, the biological information obtained by the sensor and the sensor identification information for each sensor are transmitted from the sensor to the management server 300 through a communication line (not shown).

[0040] The CPU 11a of the management server 300 identifies the bystander 42 based on the sensor identification information, and when the situation identified by the transmitted biological information is in a specific situation, it determines that there is a possibility that the identified bystander 42 will speak. Specifically, the CPU 11a of the management server 300 determines that there is a possibility that the bystander 42 identified based on the sensor identification information will speak when, for example, numerical values such as pulse, heart rate, and blood pressure increase.

[0041] In the processing example shown in FIG. 4, the CPU 11a generates control information for causing a display image indicating the possibility that the bystander 42 (see FIG. 4(A)) may speak to be displayed on the display unit 217 of the device 200 as control information for causing a notification indicating the possibility that the bystander 42 may speak to be performed by the device 200. This generated control information is transmitted to device 200, and device 200 performs display control of the display unit 217 based on this control information. Thus, in the present embodiment, as shown in FIG. 4(B), a display image 45 indicating the possibility of the peripheral person 42 speaking is displayed on the display unit 217 of device 200.

[0042] This display image 45 is an image representing a so-called "speech bubble". In the present embodiment, when it is determined that there is a possibility that the peripheral person 42 will speak, before the display of the speech content described later is performed, this display image 45 composed of an image representing a speech bubble is displayed on the display unit 217 of device 200 as shown in FIG. 4(B). Thereby, the target person 41 recognizes that there is a possibility of the peripheral person 42 speaking. The display image 45 is not limited to a 2D (2 Dimension) image, and may be a 3D (3 Dimension) image. When the display image 45 is a 3D (3 Dimension) image, images with different angles for each eye are displayed on the display unit 217. In other words, when the display image 45 is a 3D (3 Dimension) image, a plurality of images with different viewing angles are displayed on the display unit 217 as the display image 45.

[0043] Note that the generation of the control information may be performed by other than the management server 300. In the present embodiment, the management server 300 generates the control information, but not limited thereto, the generation of the control information may be performed by a device other than the management server 300. The generation of the control information may be performed by, for example, device 200. When device 200 generates the control information, device 200 determines whether the peripheral person 42 is in a specific situation based on, for example, video of the peripheral person 42 obtained by the device camera 214 (see FIG. 3) that device 200 has and voice information obtained by the individual microphone 600.

[0044] Then, when the peripheral person 42 is in a specific situation, the device 200 generates control information that causes the device 200 to issue a notification indicating the possibility of the peripheral person 42 speaking. Specifically, the device 200 generates control information that causes the display image 45 to be displayed on its own display unit 217. As a result, the display image 45 is displayed on the display unit 217 of the device 200 in the same manner as when the CPU 11a of the management server 300 generates control information.

[0045] The CPU 11a of the management server 300 generates control information that causes the display image 45 to be displayed in a form associated with the peripheral person 42 reflected in the display unit 217, as control information for causing the display image 45 (see FIG. 4(B)) to be displayed on the display unit 217. As a result, as shown in FIG. 4(B), the display image 45 is displayed in a form associated with the peripheral person 42 reflected in the display unit 217. More specifically, in the present embodiment, the display image 45 is displayed in a form associated with the periphery of the head of the peripheral person 42.

[0046] The CPU 11a of the management server 300 identifies each of the peripheral persons 42 reflected in the display unit 217 based on the video acquired by the device camera 214 (see FIG. 3). Specifically, the CPU 11a of the management server 300 identifies each of the peripheral persons 42 reflected in the display unit 217 of the device 200 based on the video of the peripheral person 42 reflected in the video acquired by the device camera 214 and the face information registered in the database.

[0047] Then, the CPU 11a of the management server 300 generates control information that causes the display image 45 to be associated with the peripheral person 42 (hereinafter sometimes referred to as the "peripheral person 42 with speaking possibility") among the identified peripheral persons 42 who are determined to have a possibility of speaking. Specifically, when generating this control information, the CPU 11a generates control information including position information, which is information about the display position of the display image 45.

[0048] The CPU 11a determines the position of the potential speaker 42 shown in the video acquired by the device camera 214 as the display position of the display image 45, and generates control information including position information, which is information about this display position. More specifically, the CPU 11a determines the position around the head of the potential speaker 42 on the video acquired by the device camera 214 as the display position of the display image 45. Then, the CPU 11a generates control information including position information, which is information about this determined display position.

[0049] Then, in this embodiment, the control information including this position information is transmitted to the device 200. Then, the device 200 performs display control so that the display image 45 is displayed at the position specified by this position information. As a result, as shown in FIG. 4(B), on the display unit 217 of the device 200, the display image 45 is displayed in a form associated with the potential speaker 42 who may be speaking.

[0050] Note that when the device 200 is the above-described transparent device 200, the CPU 11a of the management server 300 generates control information for causing the display image 45 to be displayed at a portion of the display unit 217 of this device 200 that is located on the straight line connecting the eyes of the target person 41 and the potential speaker 42. In this case, the CPU 11a of the management server 300 first acquires the angle formed by the front direction of the device 200 and the direction from the device 200 toward the potential speaker 42. Specifically, the CPU 11a of the management server 300 analyzes the video acquired by the device camera 214 to acquire the angle formed by the front direction and the direction toward the potential speaker 42.

[0051] Then, based on this angle, the CPU 11a determines the display position of the display image 45 on the display unit 217, and generates control information including information about the determined display position. As a result, also in the transmissive device 200, the display image 45 is displayed in a form associated with the potential speaker 42 who may speak. In this case, the subject 41 will visually recognize the potential speaker 42 existing in the real space and the display image 45 shown on the display unit 217 and located between his own eyes and the potential speaker 42.

[0052] After that, in this processing example, as shown by the reference numeral 4C in FIG. 4(C), the actual speech by the peripheral person 42 starts. In other words, the actual speech by the potential speaker 42 starts. When the actual speech by the peripheral person 42 starts, as shown in FIG. 4(C), an image 46 indicating that the speech by the peripheral person 42 has started is displayed inside the display image 45 associated with this peripheral person 42. When there is an actual speech by the peripheral person 42 in a state where the display image 45 is associated, an image 46 indicating that the speech by this peripheral person 42 has started is displayed inside this display image 45.

[0053] In other words, in this embodiment, when there is an actual speech by the peripheral person 42 in a state where the display image 45 is associated, an image indicating that the acquisition of audio information by the individual microphone 600 has started is displayed inside the display image 45. Whether there is an actual speech by the peripheral person 42 in a state where the display image 45 is associated is determined based on, for example, the output from the individual microphone 600 worn by this peripheral person 42.

[0054] In this embodiment, an image 46 indicating that the speech by the peripheral person 42 has started is displayed within the area surrounded by the display image 45. This image 46 indicating the start is displayed on the display unit 217 of the device 200 until the speech content of the surrounding person 42 is acquired by the CPU 11a of the management server 300.

[0055] It takes time to acquire the speech content by the CPU 11a. In the present embodiment, an image 46 indicating that the speech by the surrounding person 42 has started is displayed on the display unit 217 until the speech content is acquired by the CPU 11a. Note that the display of this image 46 indicating the start is not essential, and the display shown in FIG. 4(D) described below may be performed without going through the display shown in FIG. 4(C) from the display shown in FIG. 4(B).

[0056] When there is an actual speech by the surrounding person 42, the CPU 11a of the management server 300 acquires the speech content that is the content of the speech by the surrounding person 42. The CPU 11a of the management server 300 analyzes the voice information transmitted from the individual microphone 600 worn by the surrounding person 42 in a state where the display image 45 is associated, and acquires the speech content of this surrounding person 42. Note that the acquisition of the speech content based on the voice information may be performed using a known method.

[0057] Next, the CPU 11a of the management server 300 generates control information so that the acquired speech content is displayed on the display unit 217 in a form associated with the surrounding person 42 who made the speech with this speech content. Then, the CPU 11a of the management server 300 transmits the generated control information to the device 200.

[0058] Thereby, in the present embodiment, as shown in FIG. 4(D), the speech content 48 of the surrounding person 42 is displayed at a predetermined display location 47 in the display unit 217 of the device 200. In the present embodiment, when the above-mentioned surrounding person 42 who may have spoken actually speaks, the speech content 48 of this surrounding person 42 is displayed on the display unit 217 of the device 200. In the present embodiment, when the utterance content 48 is displayed, it is displayed inside a display image 45 composed of an image representing a speech bubble. In the present embodiment, the utterance content 48 of the surrounding person 42 is displayed inside the display image 45 that is displayed in association with the surrounding person 42 who is determined to have a possibility of speaking.

[0059] In the processing example described above, when there is a possibility that the surrounding person 42 will speak, the CPU 11a first generates control information so that the display image 45 is displayed on the display unit 217 of the device 200 as described above. As a result, as shown in FIG. 4(B), the display image 45 is displayed on the display unit 217. As this control information for causing the display image 45 to be displayed, the CPU 11a generates control information so that the display image 45 is displayed in a form associated with a display location 47 where the utterance content 48 (not shown in FIG. 4(B)) is to be displayed in the display unit 217.

[0060] In the present embodiment, the inside of the display image 45 (see FIG. 4(B)) is the display location 47 where the utterance content 48 is to be displayed. As control information for causing the display image 45 to be displayed, the CPU 11a generates control information so that the display image 45 is associated with this display location 47. More specifically, as control information for causing the display image 45 to be associated with the display location 47, as shown in FIG. 4(B), the CPU 11a generates control information so that an image representing a speech bubble having a shape surrounding the display location 47 is displayed on the display unit 217.

[0061] Then, in the present embodiment, after this control information is generated, when there is an actual utterance of the surrounding person 42, as shown in FIG. 4(D), the utterance content 48 of this surrounding person 42 is displayed within a region surrounded by the display image 45 composed of an image representing a speech bubble. In other words, when there is an actual utterance of the surrounding person 42, the utterance content 48 is displayed at the display location 47 located inside the display image 45.

[0062] The CPU 11a of the management server 300 generates control information that causes the device 200 to issue a notification appealing to the vision of the target person 41, such as causing the display image 45 to be displayed, as control information for causing the device 200 to issue a notification indicating that there is a possibility of speech. Accordingly, in the present embodiment, a notification that can be visually confirmed by the target person 41 is issued by the device 200.

[0063] 〔Form of Notification Processing〕 Here, the display image 45 is not limited to an image having a shape surrounding the above-described display location 47. The shape of the display image 45 is not particularly limited, and any shape may be used as long as it is an image that can be visually confirmed by the target person 41. As another example of the display image 45, for example, a dot-shaped image can be cited. When the dot-shaped image is to be displayed on the display unit 217, similar to the image having the surrounding shape described above, the dot-shaped image is to be displayed in a form associated with the bystander 42. Also, when the dot-shaped image is to be displayed on the display unit 217, the speech content 48 of the bystander 42 is to be displayed around the dot-shaped image.

[0064] In addition, as the display image 45 displayed on the display unit 217 of the device 200, for example, an image of characters indicating that there is a possibility of speech, such as "There is a possibility of speech", may be displayed. Also, the notification appealing to the vision of the target person 41 is not limited to being by an image, and may be performed by lighting a light source (not shown) provided in the device 200. Also, the notification appealing to the vision of the target person 41 may be performed by changing the color of the entire display screen displayed on the display unit 217 or by changing the color of a part of the display screen, such as the edge of the display screen displayed on the display unit 217.

[0065] In addition, as control information for causing the device 200 to issue a notification indicating the possibility of the peripheral person 42 speaking, for example, control information for vibrating a vibration source (not shown) provided in the device 200 may be generated. In this case, the target person 41 recognizes the possibility of the peripheral person 42 speaking based on the vibration of the device 200. In addition, for example, control information for causing a sound or voice to be emitted from a speaker 216 (see FIG. 3) provided in the device 200 may be generated. When the target person 41 has a hearing impairment, it is difficult to notify by sound. However, when the target person 41 does not have a hearing impairment, the target person 41 recognizes the possibility of the peripheral person 42 speaking by the sound.

[0066] When the target person 41 has a hearing impairment, as shown in FIG. 4(D), when the speech content 48 is displayed, the target person 41 can recognize that the peripheral person 42 is speaking. Here, the display of the speech content 48 is performed after the actual speech by the peripheral person 42. If the display of the speech content 48 is delayed, a situation may occur where the speech content 48 has not yet been displayed even though the actual speech by the peripheral person 42 has already started. In this case, a situation may occur where the target person 41 misrecognizes that the peripheral person 42 is not speaking and the target person 41 starts speaking while the peripheral person 42 is speaking.

[0067] In contrast, in the present embodiment, as described above, the target person 41 is notified of the possibility of speaking before the actual speech by the peripheral person 42 starts. In this case, when the target person 41 is notified of the possibility of speaking, the target person 41 becomes cautious about his or her own speech. In this case, situations such as the target person 41 starting to speak after the peripheral person 42 starts speaking are less likely to occur. Note that the information processing system 1 of the present embodiment also functions effectively when the target person 41 is other than a person with a hearing impairment. If the subject 41 is notified that there is a possibility that the surrounding person 42 may speak even if the subject 41 is not hearing impaired, it becomes less likely for the subject 41 to speak after the surrounding person 42 starts speaking.

[0068] 〔Other display examples〕 FIG. 6 is a diagram showing other display examples in the display unit 217 of the device 200. FIG. 6 shows a display example in a situation where there is no speech by the surrounding person 42 and there is also no possibility of speech by the surrounding person 42. When there is no speech by the surrounding person 42, the CPU 11a of the management server 300 may generate control information to cause the device 200 to issue a notification indicating that there is no speech. As a result, in this case, as shown in FIG. 6, an image 51 indicating that there is no speech is displayed on the display unit 217 of the device 200.

[0069] This image 51 indicating that there is no speech shown in FIG. 6 is an image composed of the characters "No speech". By displaying the image 51 indicating that there is no speech, the subject 41 recognizes that there is no speech by the surrounding person 42. Note that the image 51 indicating that there is no speech is not limited to an image composed of characters, and may be an image other than an image composed of characters, such as a symbol or a figure.

[0070] Here, assume a case where the possibility of speech by the surrounding person 42 occurs from the situation shown in FIG. 6. In this case, the CPU 11a of the management server 300 generates control information to cause the display on the device 200 to switch to the display shown in FIG. 4(B). More specifically, the CPU 11a generates control information to erase the image 51 indicating that there is no speech and to display the display image 45 shown in FIG. 4(B) on the display unit 217. When the display image 45 shown in FIG. 4(B) is displayed, the subject 41 recognizes that there is a possibility that the surrounding person 42 may speak.

[0071] 〔Other processing example〕 Figs. 7(A) to (C) are diagrams showing other processing examples. The processing when there is no actual speech by the surrounding person 42 will be described. In this processing example, as in the above, first, as shown in Figs. 7(A) and (B), the possibility of speech by the surrounding person 42 occurs, and in response thereto, the display image 45 is displayed on the display unit 217 of the device 200. After that, in this processing example, a situation where there is no actual speech by the surrounding person 42 occurs. In this case, in this processing example, as shown in Fig. 7(C), the display image 45 corresponding to this surrounding person 42 is erased.

[0072] After the CPU 11a generates the above control information for causing the display image 45 to be displayed on the display unit 217, when there is no actual speech by the surrounding person 42, the CPU 11a generates control information for erasing the display image 45 being displayed on the display unit 217. More specifically, the CPU 11a generates control information for erasing the display image 45 when there is no actual speech by the surrounding person 42 corresponding to the display image 45 during a predetermined time period after generating the above control information for causing the display image 45 to be displayed, or during a predetermined time period after the display image 45 is displayed on the display unit 217. In this case, accordingly, the device 200 erases the display image 45. As a result, as shown in Figs. 7(B) and (C), the display image 45 being displayed on the display unit 217 is erased.

[0073] 〔Other processing example〕 Figs. 8(A) to (C) are diagrams showing other processing examples. When a plurality of surrounding persons 42 reflected in the display unit 217 of the device 200 are in a specific situation, the CPU 11a of the management server 300 generates, as control information, control information for causing the display image 45 to be displayed in a form in which the display image 45 is associated with each of the plurality of surrounding persons 42. As a result, in this case, as shown in FIG. 8(B), on the display unit 217 of the device 200, the display image 45 is displayed in a form in which the display image 45 is associated with each of the plurality of surrounding persons 42. In this case, the person 41 who refers to the display unit 217 of the device 200 recognizes that there is a possibility of speaking to the plurality of surrounding persons 42.

[0074] After that, when a surrounding person 42 included in the plurality of surrounding persons 42 actually speaks, as indicated by reference numeral 8D in FIG. 8(C), the speech content 48 is displayed within a region surrounded by the display image 45 that is associated with and displayed for the surrounding person 42 who actually spoke. Also, in this processing example shown in FIG. 8, for other surrounding persons 42 indicated by reference numeral 8E in FIG. 8(C) who did not actually speak, as shown in FIGS. 8(B) and 8(C), the display image 45 that was associated with and displayed for this other surrounding person 42 is erased.

[0075] Although not shown, in the state shown in FIG. 8(C), when there is a possibility of speaking to another surrounding person 42 indicated by reference numeral 8E, the display image 45 corresponding to this other surrounding person 42 will be displayed again. In the present embodiment, while a surrounding person 42 indicated by reference numeral 8F, who is one of the surrounding persons 42, is speaking, if there is a possibility of speaking to another surrounding person 42, the display image 45 corresponding to this one of the surrounding persons 42 and the speech content 48 are being displayed, and a new display image 45 corresponding to this other surrounding person 42 is also displayed.

[0076] 〔Other Processing Examples〕 FIGS. 9(A) to 9(C) and FIG. 10 are diagrams showing other processing examples. FIG. 10 shows a state when the device 200, the surrounding persons 42, and the person 41 are viewed from above in the vertical direction. In the processing examples shown in FIGS. 9 and 10, as shown in FIG. 10, a part of the surrounding persons 42 indicated by reference numeral 10B is outside the shooting range 10A by the device camera 214 (not shown in FIG. 10) provided in the device 200. The imaging range 10A can also be regarded as the range of the field of view of the subject 41 who visually recognizes the front of himself / herself via the device 200. In the processing examples shown in FIGS. 9 and 10, in this range of the field of view, some of the surrounding persons 42 indicated by the reference numeral 10B are excluded. Hereinafter, this part of the surrounding persons 42 will be referred to as "non-display surrounding persons 42B". As shown in FIG. 9(A), on the display unit 217 of the device 200, two surrounding persons 42 indicated by the reference numeral 9D other than the non-display surrounding persons 42B are displayed, and the non-display surrounding persons 42B are not shown on the display unit 217 of the device 200.

[0077] Furthermore, in this processing example shown in FIGS. 9 and 10, this non-display surrounding person 42B who is not shown on the display unit 217 of the device 200 is in a specific situation, and there is a possibility that the non-display surrounding person 42B may speak. In this case, the CPU 11a of the management server 300 generates control information so that a display image 45 indicating the possibility that the non-display surrounding person 42B may speak is displayed on the display unit 217. In the present embodiment, even when there is a possibility that the non-display surrounding person 42B may speak, control information is generated so that the display image 45 is displayed on the display unit 217.

[0078] As a result, in this processing example, as shown in FIG. 9(B), a display image 45 corresponding to the non-display surrounding person 42B is displayed on the display unit 217 of the device 200. When the subject 41 refers to the display unit 217 in the state of FIG. 9(B), the subject 41 recognizes that there is a possibility that the surrounding person 42 located at a position outside his / her field of view may speak. When the non-display surrounding person 42B actually speaks, as shown in FIG. 9(C), the speech content 48 of the non-display surrounding person 42B is displayed in a form associated with the display image 45 corresponding to the non-display surrounding person 42B. Also in this processing example, the speech content 48 of the non-display surrounding person 42B is displayed within the area surrounded by the display image 45 corresponding to the non-display surrounding person 42B. In the processing example shown in FIGS. 9 and 10, the speech content 48 of the non-display peripheral person 42B not shown on the display unit 217 of the device 200 is also displayed on the display unit 217 of the device 200.

[0079] In this processing example shown in FIGS. 9 and 10, when a display image 45 corresponding to the non-display peripheral person 42B is to be displayed on the display unit 217 of the device 200, as shown in FIG. 9(B), a display is performed so that the direction in which the non-display peripheral person 42B is located can be understood by the target person 41. In FIG. 9(B), the non-display peripheral person 42B is located on the left side of the front of the device 200, and the display image 45 corresponding to the non-display peripheral person 42B is also located on the left side in the drawing from the central portion 217C of the display unit 217. In the present embodiment, the display position of the display image 45 corresponding to the non-display peripheral person 42B changes according to the position of the non-display peripheral person 42B.

[0080] FIG. 11 is a view when the display unit 217 and the peripheral person 42 are viewed from the direction indicated by the arrow XI in FIG. 10. As shown in FIG. 11, the CPU 11a generates control information for displaying the display image 45 corresponding to the non-display peripheral person 42B on the portion of the display unit 217 located on the straight line 11L connecting the non-display peripheral person 42B and the central portion 217C of the display unit 217. This will be described in detail with reference to FIG. 10. Here, an imaginary plane 10K along the display unit 217 of the device 200 is assumed. Further, a line 10H connecting the non-display peripheral person 42B from the center of the target person 41 is assumed.

[0081] Furthermore, here, a case where the non-display peripheral person 42B is projected onto the imaginary plane 10K is assumed. More specifically, a case where the non-display peripheral person 42B is projected onto the imaginary plane 10K in the direction in which the line 10H extends from the location where the non-display peripheral person 42B is located is assumed. In this case, on the imaginary plane 10K, the non-display peripheral person 42B is located at the position indicated by the reference numeral 10M. Hereinafter, when the non-display peripheral person 42B is projected onto this virtual plane 10K along the display unit 217, the position of this non-display peripheral person 42B is referred to as "plane position 10M".

[0082] The CPU 11a generates control information for causing the display image 45 (see FIG. 11) of the non-display peripheral person 42B to be displayed on the display unit 217 in a direction located on a straight line 11L connecting the plane position 10M and the central portion 217C of the display unit 217 in the display unit 217. As a result, the display image 45 is displayed at the location indicated by reference numeral 11X in FIG. 11 on the display unit 217 of the device 200. By referring to the display unit 217 shown in FIG. 11, the target person 41 can identify in which direction the non-display peripheral person 42B who may be speaking is located.

[0083] When acquiring the situation information about the non-display peripheral person 42B who is not reflected in the display unit 217 of the device 200 based on the video in which this non-display peripheral person 42B is reflected, the situation information about this non-display peripheral person 42B is acquired based on the video obtained by the overall camera 500 (see FIG. 10). Specifically, in this case, based on the video obtained by the overall camera 500 and the face information registered in the database (see FIG. 5), the non-display peripheral person 42B is identified, and based on this video, the situation information of the identified non-display peripheral person 42B is acquired.

[0084] Also, when displaying the display image 45 on the display unit 217 of the device 200, it is necessary to identify the position of the identified non-display peripheral person 42B. In this case, the CPU 11a of the management server 300 analyzes, for example, the video obtained by the overall camera 500 to identify the position of the non-display peripheral person 42B. Furthermore, in this case, the CPU 11a of the management server 300 analyzes the video obtained by the overall camera 500 to identify the position of the central portion 217C of the display unit 217 of the device 200 worn by the target person 41 and the orientation of the device 200.

[0085] Then, the CPU 11a of the management server 300 specifies the above-described planar position 10M based on the position of the non-display peripheral person 42B, the position of the central portion 217C of the display unit 217 of the device 200, and the orientation of the device 200. Next, the CPU 11a of the management server 300 determines the display position of the display image 45 on the display unit 217 based on the specified planar position 10M and the position of the central portion 217C of the display unit 217.

[0086] Then, the CPU 11a of the management server 300 generates control information including information about the determined display position. The device 200 performs display control on the display unit 217 according to this control information. As a result, on the display unit 217 of the device 200, as shown in FIG. 11, the display image 45 is displayed on the straight line 11L connecting the planar position 10M and the central portion 217C of the display unit 217.

[0087] In addition, when the situation specified by the acquired situation information is a situation where one peripheral person 42 is looking at another peripheral person 42, the CPU 11a may generate control information so that a notification indicating that there is a possibility of the other peripheral person 42 speaking is given by the device 200. In the above, based on the situation information of the peripheral person 42, it is determined whether there is a possibility of this peripheral person 42 speaking. However, the present invention is not limited to this, and based on the situation information of one peripheral person 42, it may be determined whether there is a possibility of the other peripheral person 42 speaking.

[0088] When the CPU 11a determines that there is a possibility of the other peripheral person 42 speaking, for example, the CPU 11a generates control information so that the display image 45 is associated with this other peripheral person 42. More specifically, for example, the CPU 11a generates control information so that the display image 45 is associated with this other peripheral person 42 reflected on the display unit 217.

[0089] Specifically, for example, when one peripheral person 42 shown in the video obtained by the device camera 214 continuously visually recognizes other peripheral persons 42 shown in this video for a time exceeding a predetermined time, the CPU 11a determines that the situation of these other peripheral persons 42 is a situation where there is a possibility of speech. And in this case, the CPU 11a generates control information so that the display image 45 is displayed in association with this other peripheral person 42.

[0090] 〔Explanation of the processing flow〕 FIG. 12 is a flowchart showing the processing flow executed when the above notification process is performed. A series of the processing flows described above will be explained. In the present embodiment, first, the CPU 11a of the management server 300 determines for each of the peripheral persons 42 located around the target person 41 whether the situation of the peripheral person 42 is in the above specific situation (step S101). And when the CPU 11a determines that the situation of the peripheral person 42 is in a specific situation, it identifies the peripheral person 42 in this specific situation (step S102).

[0091] Thereafter, the CPU 11a generates control information so that the display image 45 corresponding to this peripheral person 42 in the specific situation is displayed on the display unit 217 (step S103). As a result, a display image 45 indicating the possibility of speech is displayed on the display unit 217 of the device 200. Thereafter, the CPU 11a determines whether the identified peripheral person 42 actually made a speech based on the voice information of the peripheral person 42 identified as being in a specific situation (step S104).

[0092] And when the CPU 11a does not determine that the identified peripheral person 42 actually made a speech, it generates control information so that the display image 45 corresponding to this peripheral person 42 is erased (step S105). As a result, the display image 45 displayed on the display unit 217 of the device 200 is erased. On the one hand, when the CPU 11a determines that the identified peripheral person 42 has made an actual speech, it analyzes the voice information to obtain the speech content 48 corresponding to this peripheral person 42 (step S106).

[0093] Next, the CPU 11a generates control information so that this speech content 48 is displayed within the display image 45 (step S107). In this case, the CPU 11a generates control information so that this speech content 48 is displayed within the display image 45 that is displayed in association with the identified peripheral person 42. As a result, the speech content 48 is displayed within the display image 45.

[0094] 〔Notification processing regarding the end of speech〕 Next, the notification processing regarding the end of speech will be described. In the above, the notification processing regarding the possibility of speech has been described. In addition, a notification indicating the possibility of the end of speech or a notification indicating that the speech has ended may be given to the target person 41 through the device 200. In this embodiment, as described above, the CPU 11a of the management server 300 acquires situation information, which is information about the situation of the peripheral person 42 before actually making a speech. And when there is a possibility that the peripheral person 42 makes a speech, as described above, a notification indicating the possibility of speech is given. Hereinafter, in this specification, this situation information, which is information about the situation of the peripheral person 42 before actually making a speech, is referred to as "pre-speech situation information".

[0095] Furthermore, in the process described below, after the peripheral person 42 starts an actual speech, the CPU 11a of the management server 300 acquires situation information, which is information about the situation of this peripheral person 42 who is making a speech. Hereinafter, in this specification, this situation information, which is information about the situation of the peripheral person 42 who is making a speech, is referred to as "during-speech situation information". Then, even when the situation specified by the acquired in-conversation situation information is a specific situation, the CPU 11a generates control information to be used for controlling the device 200.

[0096] Specifically, as this control information, the CPU 11a generates control information to cause the device 200 to issue a notification indicating that the speech of the surrounding person 42 may end (hereinafter referred to as the "ending possibility suggestion notification"). In other words, as this control information, the CPU 11a generates control information to cause the device 200 to issue an ending possibility suggestion notification indicating that the speech of the surrounding person 42, who is the target of display of the display image 45, may end.

[0097] Also, when the situation specified by the acquired in-conversation situation information is a specific situation, the CPU 11a generates control information to cause the device 200 to issue a notification indicating that the speech of the surrounding person 42 has ended (hereinafter referred to as the "ending notification"). In other words, as this control information, the CPU 11a generates control information to cause the device 200 to issue an ending notification indicating that the speech of the surrounding person 42, who is the target of display of the display image 45, has ended.

[0098] 〔Ending possibility suggestion notification〕 The ending possibility suggestion notification will be described. When the situation specified by the acquired in-conversation situation information is a situation where the speech of the surrounding person 42 may end, the CPU 11a generates control information to cause the device 200 to issue an ending possibility suggestion notification. When the situation specified by the acquired in-conversation situation information is, for example, the following situation, the CPU 11a generates control information to cause the device 200 to issue an ending possibility suggestion notification. · When the situation specified by the voice information, such as the tone of the voice of the surrounding person 42 dropping or the pitch of the speech decreasing, is a specific situation · When the surrounding person 42 performs a specific predetermined action, such as lowering the hand that was raised

[0099] End Notification Next, the end notification will be explained. When the situation specified by the acquired in-conversation situation information indicates that the conversation of the surrounding person 42 has ended, the CPU 11a generates control information to cause the end notification to be issued by the device 200. Specifically, when the situation specified by the acquired in-conversation situation information is, for example, in the following situations, the CPU 11a generates control information to cause the end notification to be issued by the device 200. · When voice information is no longer acquired · When the expression of the surrounding person 42 is in a specific state, such as when the mouth of the surrounding person 42 is closed

[0100] The CPU 11a acquires the in-conversation situation information of the surrounding person 42 based on the voice information of the surrounding person 42 and the video in which the surrounding person 42 is reflected. More specifically, the CPU 11a acquires the in-conversation situation information of the surrounding person 42 based on the voice information obtained by the individual microphone 600 and the video obtained by the device camera 214 or the overall camera 500. Then, when the situation specified by this in-conversation situation information is in a predetermined situation, the CPU 11a generates control information to cause a possible end notification or an end notification to be issued by the device 200.

[0101] Note that the information used to acquire the pre-conversation situation information and the information used to acquire the in-conversation situation information may be different. Specifically, for example, for the pre-conversation situation information, the pre-conversation situation information may be acquired based on the video in which the surrounding person 42 is reflected, and for the in-conversation situation information, the in-conversation situation information may be acquired based on the voice information of the surrounding person 42.

[0102] When acquiring the pre-conversation situation information, since the surrounding person 42 often does not make a clear speech, it is easier to improve the accuracy of judgment about the possibility of making a speech by acquiring the pre-conversation situation information based on the video in which the surrounding person 42 is reflected. On the other hand, when acquiring the in-conversation situation information, since the surrounding person 42 is actually speaking, acquiring the in-conversation situation information based on voice information makes it easier to increase the possibility of the end of the conversation and the accuracy of judgment regarding the end of the conversation compared to the case of acquiring the in-conversation situation information based on video.

[0103] The above control information for causing the end possibility suggestion notification and the end notification to be performed by the device 200 is transmitted to the device 200 in the same manner as above. Accordingly, in the present embodiment, the device 200 performs control based on this control information to give the target person 41 an end possibility suggestion notification or an end notification. Thereby, the target person 41 recognizes that the speech of the surrounding person 42 is likely to end or that the speech of the surrounding person 42 has ended.

[0104] 〔Specific Example of Processing〕 FIGS. 13(A) to (D) are diagrams showing specific examples of the processing. FIG. 13(A) shows a situation where the surrounding person 42 is speaking. In the present embodiment, in the situation shown in FIG. 13(A), the speech content 48 is displayed within the area surrounded by the display image 45. When the surrounding person 42 is speaking as shown in FIG. 13(A), the CPU 11a of the management server 300 acquires the in-conversation situation information. More specifically, the CPU 11a acquires the in-conversation situation information of the surrounding person 42 in a state where the display image 45 is associated based on the video obtained by the overall camera 500 or the device camera 214 and the voice information obtained by the individual microphone 600.

[0105] Then, when the situation specified by the acquired in-conversation situation information is a situation where there is a possibility of the end of the speech of the surrounding person 42, the CPU 11a generates control information for causing the end possibility suggestion notification to be performed by the device 200. Also, when the situation specified by the acquired speaking situation information indicates a situation where the speech of the surrounding person 42 has ended, the CPU 11a generates control information to cause an end notification to be made by this device 200.

[0106] FIG. 13(B) shows the state of the display unit 217 of the device 200 when the situation specified by the speaking situation information is a situation where there is a possibility that the speech of the surrounding person 42 will end. When the situation is such that there is a possibility that the speech of the surrounding person 42 will end, as described above, the CPU 11a generates control information to cause a notification suggesting the possibility of end to be made by this device 200. In this processing example, as this control information for causing a notification suggesting the possibility of end to be made by the device 200, the CPU 11a generates control information for changing the display image 45 displayed on the display unit 217 of the device 200. Hereinafter, in this specification, this control information for causing a notification suggesting the possibility of end to be made by the device 200 is referred to as "first control information".

[0107] In this processing example, as the first control information, the CPU 11a generates control information for changing the display image 45 that is displayed in association with the display location 47 of the speech content 48 (see FIG. 13(A)) of the surrounding person 42. Here, the display image 45 can be regarded as a corresponding display image that is displayed in association with the display location 47 of the speech content 48 of the surrounding person 42. In this embodiment, as this corresponding display image, a display image 45 representing a speech bubble is displayed. As the first control information for causing a notification suggesting the possibility of end to be made by the device 200, the CPU 11a generates control information for changing the display image 45 representing this speech bubble, which is an example of the corresponding display image.

[0108] More specifically, as this first control information for changing the corresponding display image, the CPU 11a generates control information for changing the shape of the display image 45 representing the speech bubble, which is displayed in a form surrounding the display location 47. Specifically, the CPU 11a generates control information as the first control information for changing the corresponding display image such that a protruding portion 45G (see FIG. 13(A)) provided as a part of the display image 45 representing a speech bubble is erased.

[0109] As a result, in the present embodiment, as shown in FIGS. 13(A) and (B), the protruding portion 45G is erased. In the present embodiment, the protruding portion 45G provided in the display image 45 that is displayed in association with a peripheral person 42 in a situation where there is a possibility that the speech has ended is erased. The target person 41 recognizes that there is a possibility that the speech of the peripheral person 42 has ended by recognizing that the protruding portion 45G has been erased.

[0110] In the present embodiment, as described above, when the situation specified by the pre-speech situation information is a specific situation, the CPU 11a generates control information for causing the display unit 217 provided in the device 200 to display the display image 45 representing a speech bubble. As a result, in the present embodiment, first, as shown in FIG. 13(A), the display image 45 representing a speech bubble is displayed on the display unit 217 of the device 200 in a form associated with the peripheral person 42. The display image 45 is provided with a protruding portion 45G that protrudes toward the peripheral person 42. In the present embodiment, a peripheral person 42 who may speak or a peripheral person 42 who is in the middle of speaking is located at the tip in the protruding direction of the protruding portion 45G.

[0111] The CPU 11a generates control information as the first control information for causing the end suggestion notification to be performed by the device 200 so that the display form of the display image 45 displayed on the display unit 217 is changed. Specifically, the CPU 11a generates control information as this first control information so that the protruding portion 45G of the display image 45 is not displayed. As a result, in the present embodiment, as described above, the protruding portion 45G of the display image 45 is not displayed.

[0112] Also, as described above, when the speech of the peripheral person 42 actually ends, the CPU 11a generates control information to cause the device 200 to issue an end notification, which is a notification indicating that the speech has ended. Hereinafter, this control information for causing the device 200 to issue an end notification is referred to as "second control information". Also in this case, the CPU 11a generates, as this second control information, control information for changing the display image 45 displayed on the display unit 217 of the device 200. More specifically, the CPU 11a generates, as the second control information, control information for changing the above-described display image 45 that is displayed in association with the display position 47 of the speech content 48 of the peripheral person 42.

[0113] More specifically, the CPU 11a generates, as the second control information, control information for further changing the display form of the display image 45 after the display form has been changed by the above-described first control information so that the display form of the display image 45 is changed. More specifically, the CPU 11a generates, as the second control information, control information for changing the thickness of the lines constituting the display image 45.

[0114] As a result, in the present embodiment, as shown in FIGS. 13(B) and (C), the lines constituting the display image 45 displayed on the display unit 217 of the device 200 become thinner. In other words, in the present embodiment, the lines constituting the display image 45 that are displayed in association with the peripheral person 42 who is speaking become thinner. As a result, the target person 41 recognizes that the speech of the peripheral person 42 has ended.

[0115] In the present embodiment, there is a time difference between the timing when the speech of the peripheral person 42 ends and the timing when the speech content 48 at the end of the speech of the peripheral person 42 is displayed on the display unit 217. In this case, in the present embodiment, as shown in FIGS. 13(C) and (D), after the timing of the end of the speech of the peripheral person 42, the display process of the speech content 48 continues until all of the speech content 48 is displayed. In other words, in the present embodiment, even if the lines constituting the display image 45 become thinner, the display process of the speech content 48 does not end, and this display process continues until all of the speech content 48 is displayed.

[0116] The target person 41 can also recognize the end of the speech of the peripheral person 42 by recognizing the end of this display process of the speech content 48. By the way, in the present embodiment, even though the speech of the peripheral person 42 has already ended, the display process continues until all of the speech content 48 is displayed. In this case, the target person 41 is likely to misrecognize that the peripheral person 42 is still speaking even though the speech of the peripheral person 42 has already ended.

[0117] In this case, it is easy for a blank time without speech to occur between the end of the speech of the peripheral person 42 and the start of the speech of the target person 41. On the other hand, when a termination possibility suggestion notification or a termination notification is performed as in the present embodiment, the target person 41 can recognize the end of the speech of the peripheral person 42 at an earlier stage. In this case, the target person 41 can speak his or her own speech at a stage shortly after the speech of the peripheral person 42 ends.

[0118] FIGS. 14(A) to (I) are diagrams showing a series of flows of the display process. In FIG. 14(A), there is no speech of the peripheral person 42 and there is also no possibility of the peripheral person 42 speaking. In this case, no sound is detected. Also, in this case, the CPU 11a does not generate control information for causing the display image 45 to be displayed, and the display image 45 is not displayed on the display unit 217 of the device 200. In FIG. 14(B), a situation where there is a possibility of speech is shown. In this case, the display image 45 is displayed on the display unit 217 of the device 200.

[0119] Figures 14(C) to (F) show the situation while the peripheral person 42 is speaking. In this case, the CPU 11a acquires the speech content 48 and further generates control information to cause the speech content 48 to be displayed. As a result, as indicated by reference numeral 13X, the speech content 48 is sequentially displayed on the display unit 217 of the device 200. More specifically, the speech content 48 is sequentially displayed inside the display image 45 displayed on the display unit 217 of the device 200. In the present embodiment, there is a time difference between the timing at which the CPU 11a acquires the audio information and the timing at which the speech content 48 is displayed on the display unit 217 of the device 200. Therefore, in the present embodiment, as indicated by the arrow 14Y in FIG. 14, the display of the speech content 48 is performed with a delay relative to the acquisition of the audio information by the CPU 11a.

[0120] FIG. 14(F) shows a situation where there is a possibility that the speech of the peripheral person 42 has ended. In this case, in the present embodiment, the protruding portion 45G (see FIG. 14(E)) that was displayed as a part of the display image 45 is erased. From FIG. 14(G) onward, a situation where the speech of the peripheral person 42 has ended is shown. In this case, as shown in FIGS. 14(G) and (H), the lines constituting the display image 45 become thinner.

[0121] In the present embodiment, when there is a possibility that the speech has ended, at least one of the shape, thickness, and color of the display image 45 is changed, and when the speech has actually ended, at least one of the shape, thickness, and color of the display image 45 is further changed. In the above, the case where the shape of the display image 45 is changed when there is a possibility that the speech has ended and the thickness of the line constituting the display image 45 is changed when the speech has actually ended has been described as an example. In other words, in the above, the case where the shape of the display image 45 is first changed and then the thickness of the line constituting the display image 45 is changed has been described as an example.

[0122] As another example, for instance, first, the thickness of the lines forming the display image 45 may be changed, and then, the shape of the display image 45 may be changed. Alternatively, first, the shape of the display image 45 may be changed, and then, the shape of this display image 45 may be further changed. Alternatively, first, the lines forming the display image 45 may be made thinner, and then, the lines forming the display image 45 may be made even thinner. Or, first, the lines forming the display image 45 may be made thicker, and then, the lines forming the display image 45 may be made even thicker. Alternatively, first, the color of the display image 45 may be changed, and then, the color of the display image 45 may be further changed.

[0123] 〔Form of notification processing〕 In this embodiment, at the end of the speech, similar to the case at the start of the speech, a notification appealing to the vision of the target person 41 is performed. Specifically, in this embodiment, as the notification appealing to the vision of the target person 41, as described above, the first change of the display image 45 is made, and then, the second change of the display image 45 is made. At the end of the speech, the notification appealing to the vision of the target person 41 may, alternatively, be performed, for example, such that an image of text such as "There may be an end of the speech" or "Speech ended" is displayed on the display unit 217 of the device 200.

[0124] Also, at the end of the speech, the notification appealing to the vision of the target person 41 may, alternatively, be performed by, for example, turning on or off a light source (not shown) provided in the device 200. Also, at the end of the speech, the notification appealing to the vision of the target person 41 may be performed by changing the overall color of the display screen displayed on the display unit 217 or by changing the color of a part of the display screen such as the edge of the display screen displayed on the display unit 217.

[0125] Also, the notification at the end of the speech may be performed, for example, by vibrating a vibration source (not shown) provided in the device 200. In addition, the notification at the end of the speech may be made, for example, by emitting sound or voice from the speaker 216 (see FIG. 3) provided in the device 200. When the target person 41 has a hearing impairment, it is difficult to give a notification by sound. However, when the target person 41 does not have a hearing impairment, an end possibility notification or an end notification can be given to the target person 41 by sound.

[0126] 〔Explanation of the processing flow〕 FIG. 15 is a flowchart showing the processing flow when notification processing is also performed at the end of the speech. Note that the processing in steps S201 to S207 in FIG. 15 is the same as the processing in steps S101 to S107 shown in FIG. 12. In the present embodiment, first, in the same manner as above, the CPU 11a determines, for each of the peripheral persons 42 located around the target person 41, whether the situation specified by the pre-speech situation information is the above-specified situation (step S201).

[0127] Then, when the CPU 11a determines that the situation specified by the pre-speech situation information is the specified situation, the CPU 11a specifies the peripheral person 42 in this specified situation (step S202). In other words, when the CPU 11a determines that the situation specified by the pre-speech situation information is a situation where there is a possibility of speech, the CPU 11a specifies the peripheral person 42 who has a possibility of speech.

[0128] Next, the CPU 11a generates control information for causing the display image 45 corresponding to this peripheral person 42 specified in the specified situation to be displayed on the display unit 217 (step S203). Thereby, the display image 45 shown in FIG. 14(B) is displayed on the display unit 217 of the device 200 in a form corresponding to the peripheral person 42 who has a possibility of speech. A protruding portion 45G is provided in the displayed display image 45.

[0129] After that, based on the voice information of the surrounding person 42 who is the target of the display of the display image 45, the CPU 11a determines whether this surrounding person 42 has actually spoken (step S204). And when the CPU 11a determines that the surrounding person 42 has not actually spoken, it generates control information to cause the display image 45 to be erased (step S205). As a result, the display image 45 displayed on the display unit 217 of the device 200 is erased. On the other hand, when the CPU 11a determines that the surrounding person 42 has actually spoken, it analyzes the voice information to obtain the speech content 48 (step S206).

[0130] Next, the CPU 11a generates control information to cause the speech content 48 to be displayed within the display image 45 (step S207). As a result, as shown by reference numeral 13X in FIG. 14, the speech content 48 is displayed within the display image 45. In the processing examples shown in FIGS. 13 and 14 above, the case where the display image 45 composed of thick lines is displayed at the stage where there is a possibility of the surrounding person 42 speaking has been described, but the expression form of the display image 45 is not limited to this. At the stage where there is a possibility of the surrounding person 42 speaking, a display image 45 composed of thin lines may be displayed. And when the actual speech by the surrounding person 42 is started, the lines constituting the display image 45 may be thickened.

[0131] Next, in step S208, the CPU 11a determines whether there is a possibility that the speech of the surrounding person 42 will end. And when the CPU 11a determines that there is a possibility that the speech of the surrounding person 42 will end, it generates control information to cause the above-mentioned protruding portion 45G to be erased (step S209). As a result, as described above, the protruding portion 45G is erased. Next, the CPU 11a determines whether the speech of the surrounding person 42 has ended (step S210). And when the CPU 11a determines that the speech of the surrounding person 42 has ended, it generates control information to cause the lines constituting the display image 45 to become thinner (step S211).

[0132] [Others] In the above, the case where three notification processes are performed has been described: the notification process regarding the possibility of speech before the speech, the notification process regarding the possibility of the end of the speech after the speech has started, and the notification process regarding the end of the speech after the speech has started. It is not essential that all of these three processes are performed, and only one of the processes may be performed, or two of the processes may be performed. In the above, the case where two notification processes are performed at the end of the speech has been described: the notification process regarding the possibility of the end of the speech and the notification process when the speech actually ends. However, at the end of the speech, only one of these two notification processes may be performed.

[0133] (Supplementary Note) (((1))) Comprising a processor, The processor, Acquires voice information of a person around the target person, Based on the voice information, displays the speech content, which is the content of the speech of the person around who is speaking, on the display unit of the device possessed by the target person, When the speech of the person around ends or there is a possibility of ending, changes the display image that is displayed in association with the display location of the speech content displayed on the display unit An information processing system. (((2))) The processor, When the speech of the person around ends or there is a possibility of ending, changes the shape of the display image The information processing system according to ((1)). (((3))) The processor, When the speech of the person around ends or there is a possibility of ending, changes the shape of the display image having a protruding portion so that the protruding portion is not present in the display image The information processing system described in ((2)). (((4))) The processor When the speech of the surrounding person has ended or may end, change the color of the display image The information processing system according to any one of ((1)) to ((3)). (((5))) The processor When the speech of the surrounding person has ended or may end, change the thickness of the lines constituting the display image The information processing system according to any one of ((1)) to ((4)). (((6))) The processor When the speech of the surrounding person has ended or may end, make the lines constituting the display image thinner The information processing system described in ((5)). (((7))) An acquisition function for acquiring voice information of a surrounding person located around the target person, A display function for displaying, on a display unit of a device possessed by the target person, the speech content which is the content of the speech of the surrounding person who is speaking, based on the voice information, A change function for changing a display image that is displayed in association with a display location of the speech content displayed on the display unit when the speech of the surrounding person has ended or may end, A program for causing a computer to realize the above.

[0134] According to the information processing system according to ((1)) to ((6)), compared with a configuration in which a notification regarding speech is not given to the speaker, it is possible to make it easier for the speaker to specify the timing of the speech that the speaker is about to make next. According to the program according to ((7)), compared with a configuration in which a notification regarding speech is not given to the speaker, it is possible to make it easier for the speaker to specify the timing of the speech that the speaker is about to make next.

Explanation of Signs

[0135] 1... Information processing system, 10K... Virtual plane, 11a... CPU, 11L... Straight line, 41... Target person, 42... Surrounding person, 45... Display image, 47... Display location, 48... Spoken content, 200... Device, 217... Display unit, 217C... Central part

Claims

1. A processor is provided. The processor, Acquire voice information of people around the target person, Based on the voice information, a content of the speech of the surrounding person who is speaking is displayed on a display unit of a device held by the target person; When the speech of the surrounding person has ended or is likely to end, a display image displayed in association with a display portion of the speech content displayed on the display unit is changed. Information processing system.

2. The processor, When the speech of the surrounding person has ended or is likely to end, the shape of the display image is changed. The information processing system according to claim 1 .

3. The processor, When the speech of the surrounding person has ended or is likely to end, the shape of the display image having a protruding portion is changed so that the display image has no protruding portion. The information processing system according to claim 2 .

4. The processor, When the speech of the nearby person has ended or is likely to end, the color of the displayed image is changed. The information processing system according to claim 1 .

5. The processor, When the speech of the surrounding person has ended or is likely to end, the thickness of the lines constituting the display image is changed. The information processing system according to claim 1 .

6. The processor, When the speech of the surrounding person has ended or is likely to end, the lines constituting the display image are made thinner.

6. The information processing system according to claim 5.

7. An acquisition function for acquiring voice information of people around the subject; A display function for displaying the contents of the speech of the surrounding person who is speaking on a display unit of a device held by the target person based on the voice information; A change function of changing a display image displayed in association with a display portion of the speech content displayed on the display unit when the speech of the surrounding person has ended or is likely to end; A program to make the above happen on a computer.

Citation Information

Patent Citations

  • Video call device, control apparatus to be used for the same, and control method

    JP2022112784A

  • Method, device, and program for detecting next speaker

    JP2006338493A

  • Next speaker guidance system, next speaker guidance method and next speaker guidance program

    JP2012146072A

  • Program assisting user who cannot make utterance during online conference and terminal and method

    JP2023112602A