Information processing systems and programs
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- FUJIFILM BUSINESS INNOVATION CORP
- Filing Date
- 2024-12-19
- Publication Date
- 2026-05-15
Smart Images

Figure 0007859479000001 
Figure 0007859479000002 
Figure 0007859479000003
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing system and a program.
Background Art
[0002] Patent Document 1 discloses a process of determining the next speaker as a user who has obtained the line-of-sight of the majority of users among users excluding the user who is looking at the current speaker in a conversation situation. Patent Document 2 discloses a device including notification voice storage means for storing a notification voice for notifying the next speaker to each conference participant. Patent Document 3 discloses a configuration including chat text input means for receiving input of chat text and voice synthesis means for synthesizing the chat text into chat voice data.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Patent Document 3
Summary of the Invention
Problems to be Solved by the Invention
[0004] When a speaker attempts to communicate with a third party other than themselves, they usually make their own speech at a timing that misses the timing of the third party's speech. If a speaker makes a speech while a third party is making a speech, it may interfere with the third party's speech or make it difficult for the speaker's speech to be recognized by the third party. The objective of the present invention is to make it easier to identify the timing of an utterance that the speaker is about to make, compared to a configuration in which no notification regarding utterances is given to the speaker. [Means for solving the problem]
[0005] The invention described in claim 1 comprises a processor, the processor acquires voice information of a person in the vicinity of the subject, and based on the voice information, displays the content of the speech of the person in the vicinity who is speaking on the display unit of the device owned by the subject, and when the person in the vicinity finishes speaking, or When the voice of the person speaking in the vicinity becomes a specific state or when the person speaking in the vicinity performs a specific action, This is an information processing system that, before the entire content of the utterance is displayed on the display unit at the end of the utterance, changes the display image that is associated with the display location of the utterance content displayed on the display unit. The invention described in claim 2 is such that when the peripheral person finishes speaking, the processor... If the voice of the person speaking in the vicinity becomes as described above, or if the person speaking in the vicinity performs the described above, The information processing system according to claim 1, which changes the shape of the displayed image. The invention described in claim 3 is such that when the peripheral person finishes speaking, the processor... If the voice of the person speaking in the vicinity becomes as described above, or if the person speaking in the vicinity performs the described above, The information processing system according to claim 2, wherein the shape of the display image having a protrusion is changed to a state in which the display image does not have the protrusion. The invention described in claim 4 is such that when the peripheral person finishes speaking, the processor... If the voice of the person speaking in the vicinity becomes as described above, or if the person speaking in the vicinity performs the described above, The information processing system according to claim 1, wherein the color of the displayed image is changed. The invention described in claim 5 is such that when the peripheral person finishes speaking, the processor... If the voice of the person speaking in the vicinity becomes as described above, or if the person speaking in the vicinity performs the described above, The information processing system according to claim 1, wherein the thickness of the lines constituting the displayed image is changed. The invention described in claim 6 is such that when the peripheral person finishes speaking, the processor... If the voice of the person speaking in the vicinity becomes as described above, or if the person speaking in the vicinity performs the described above, The information processing system according to claim 5, wherein the lines constituting the displayed image are made thinner. The invention described in claim 7 includes an acquisition function for acquiring voice information of surrounding persons located around the subject, a display function for displaying the content of the speech of the surrounding person who is speaking, based on the voice information, on the display unit of a device owned by the subject, and when the surrounding person finishes speaking, or When the voice of the person speaking in the vicinity becomes a specific state or when the person speaking in the vicinity performs a specific action, This is a program for a computer to implement a change function that changes the display image displayed on the display unit, which is associated with the display location of the utterance content, before the entirety of the utterance content is displayed on the display unit at the end of the utterance. [Effects of the Invention]
[0006] According to the inventions of claims 1 to 7, compared to a configuration in which no notification regarding speech is given to the speaker, it becomes easier to identify the timing of the speech the speaker is about to make. [Brief explanation of the drawing]
[0007] [Figure 1] This is a diagram showing the overall configuration of the information processing system. [Figure 2] This is a diagram showing the configuration of the management server. [Figure 3] This is a diagram showing the hardware configuration of the device. [Figure 4] Figures (A) through (D) illustrate the processes performed in the information processing system. [Figure 5] This diagram shows the database stored in the information storage section of the management server. [Figure 6] This figure shows other display examples on the device's display unit. [Figure 7] (A) to (C) are diagrams illustrating other processing examples. [Figure 8] (A) to (C) are diagrams illustrating other processing examples. [Figure 9] (A) to (C) are diagrams illustrating other processing examples. [Figure 10] This figure shows another example of processing. [Figure 11] It is a view when looking at the display unit and the surrounding people from the direction indicated by arrow XI in FIG. 10. [Figure 12] It is a flowchart showing the flow of processing executed when notification processing is performed. [Figure 13] (A) to (D) are diagrams showing specific examples of processing. [Figure 14] (A) to (I) are diagrams showing a series of flows of display processing. [Figure 15] It is a flowchart showing the flow of processing when notification processing is also performed at the end of a speech.
Mode for Carrying Out the Invention
[0008] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. FIG. 1 is a diagram showing the overall configuration of the information processing system 1 of the present embodiment. The information processing system 1 is provided with a management server 300 as an example of an information processing device. Further, the information processing system 1 is provided with devices 200 to be worn by each of the target persons described later. In FIG. 1, only one device 200 is shown, but a plurality of devices 200 are provided according to the number of target persons. The device 200 is a glasses-type device 200 and is worn on the head of the target person. The target person visually recognizes the surroundings through the device 200.
[0009] Furthermore, in the present embodiment, an overall camera 500 which is a camera for photographing the target person wearing the device 200 and the surrounding persons (described later) located around this target person is provided. The overall camera 500 is provided for each place where the target person is located, and when there are a plurality of target persons, a plurality of overall cameras 500 are also provided. Furthermore, in the present embodiment, individual microphones 600 to be worn by each of the surrounding persons described later are provided. This individual microphone 600 acquires the voices of the surrounding persons and generates voice information. The individual microphones 600 are provided for each surrounding person, and when there are a plurality of surrounding persons, a plurality of individual microphones 600 are provided. Each of the devices 200, the overall camera 500, and the individual microphones 600 are connected to the management server 300 via a communication line 400 such as the internet.
[0010] [Management Server Configuration] Figure 2 shows the configuration of the management server 300. The management server 300 is implemented by a computer. The management server 300 includes an arithmetic processing unit 111 that performs digital arithmetic processing according to a program, and an information storage unit 19 that stores information. The information storage unit 19 is implemented using existing information storage devices such as an HDD (Hard Disk Drive), semiconductor memory, or magnetic tape.
[0011] The arithmetic processing unit 111 is equipped with a CPU 11a, which is an example of a processor. Furthermore, the arithmetic processing unit 111 is equipped with RAM 11b, which is used as working memory for the CPU 11a, and ROM 11c, which stores programs executed by the CPU 11a. Furthermore, the arithmetic processing unit 111 is provided with a non-volatile memory 11d that is rewritable and can retain data even if the power supply is interrupted, and an interface unit 11e that controls various parts such as the communication unit connected to the arithmetic processing unit 111.
[0012] The non-volatile memory 11d consists of, for example, SRAM or flash memory backed up by a battery. The information storage unit 19 stores various types of information, such as programs executed by the arithmetic processing unit 111. In this embodiment, the CPU 11a provided in the arithmetic processing unit 111 reads the programs stored in the ROM 11c and the information storage unit 19, thereby executing various processes performed by the management server 300.
[0013] The program executed by the CPU 11a can be provided to the management server 300 while stored on a computer-readable recording medium such as a magnetic recording medium (magnetic tape, magnetic disk, etc.), an optical recording medium (optical disk, etc.), a magneto-optical recording medium, or semiconductor memory. Alternatively, the program executed by the CPU 11a may be provided to the management server 300 using communication means such as the internet.
[0014] [Equipment Configuration] Figure 3 shows the hardware configuration of device 200. The device 200 includes a calculation processing unit 211, an information storage unit 212, a sensor 213, a device camera 214, a device microphone 215, a speaker 216, and a display unit 217. The arithmetic processing unit 211 is equipped with a CPU 21a, which is an example of a processor. Furthermore, the arithmetic processing unit 211 is equipped with RAM 21c, which is used as working memory for the CPU 21a, and ROM 21b, which stores programs executed by the CPU 21a.
[0015] The information storage unit 212 is implemented using existing information storage devices such as semiconductor memory. Examples of sensors 213 include GPS sensors and compass sensors. By referring to the output from these sensors 213, the current location and orientation of the device 200 can be determined. The equipment camera 214 is a camera that takes pictures of the area around the equipment 200. The device camera 214 faces forward of the subject when the device 200 is attached to the subject, and captures the area in front of the subject. In other words, the device camera 214 faces the direction the subject is facing and captures the area in front of the subject.
[0016] The device microphone 215 acquires the subject's voice and generates audio information. Speaker 216 outputs sound and voice and performs notification processing to the person wearing device 200. The display unit 217 is a so-called display that shows various kinds of information. The display unit 217 is positioned in front of the subject's eyes when the device 200 is worn by the subject. In this embodiment, the display unit 217 displays the image obtained by the equipment camera 214. When the equipment 200 is attached to the subject, the display unit 217 displays an image of the area in front of the subject. In this embodiment, the subject observes what is in front of them by referring to the image displayed on the display unit 217.
[0017] In addition, there is also a transparent device 200, in which case a transparent display unit 217 is installed as the display unit 217, allowing the user to see behind the display unit 217. The subject uses the display unit 217 to see what is behind it. In other words, the subject uses the display unit 217 to see what is in front of them. In the transparent device 200, when an image is displayed on the display unit 217, the user will perceive both the real space located behind the display unit 217 and the image displayed on the display unit 217.
[0018] The program executed by the CPU 21a can be provided to the device 200 in a computer-readable recording medium such as a magnetic recording medium (magnetic tape, magnetic disk, etc.), an optical recording medium (optical disk, etc.), a magneto-optical recording medium, or semiconductor memory. Alternatively, the program executed by the CPU 21a may be provided to the device 200 using communication means such as the Internet.
[0019] It should be noted that device 200 is not limited to eyeglasses-type devices 200; other examples include smartphones, tablet devices, etc. All of these devices, including eyeglasses-type devices 200, smartphones, and tablet devices, are portable devices for the target user. Smartphones and tablet devices are also equipped with a display unit and a camera. Users can see what is in front of them by referring to the image captured by the camera and displayed on the screen. In other words, in this case, the subject can see the space in front of them, but behind the smartphone or tablet, by referring to the display on the smartphone or tablet placed in front of them. In this embodiment, the notification process described later is performed via the device 200, but this notification process can be performed not only with the glasses-type device 200, but also with smartphones and tablet terminals.
[0020] In this specification, "processor" refers to a processor in a broad sense, including general-purpose processors (e.g., CPU: Central Processing Unit, etc.) and specialized processors (e.g., GPU: Graphics Processing Unit, ASIC: Application Specific Integrated Circuit, FPGA: Field Programmable Gate Array, programmable logic device, etc.). Furthermore, the operation of the processor may not be performed solely by a single processor, but may also be performed collaboratively by multiple processors located in physically separate locations. Additionally, the order of the processor operations is not limited to the order described in this embodiment and may be changed.
[0021] [Description of processes performed by the information processing system] Figures 4(A) to 4(D) illustrate the processes performed in the information processing system 1 of this embodiment. Figure 4(A) shows the target person 41 to whom the notification process will be performed, and the surrounding persons 42 who are located in the vicinity of this target person 41. In this embodiment, as shown in Figure 4(A), the device 200 is attached to the person 41 who is to be notified. In this embodiment, as shown in Figure 4(A), the display unit 217 provided on the device 200 shows the surrounding person 42. Furthermore, in this embodiment, as shown in Figure 4(A), individual microphones 600 are attached to the surrounding persons 42.
[0022] As described above, the device 200 of this embodiment is a pair of glasses. This pair of glasses is worn on the head of the subject 41. The subject 41 sees the surrounding persons 42 located around them through this device 200. In other words, the subject 41 sees the surrounding persons 42 located in front of them through this device 200. As described above, the device 200 is equipped with a device camera 214 (see Figure 3) and a display unit 217 capable of displaying the image obtained by the device camera 214. The subject 41 visually identifies the surrounding person 42 by referring to the surrounding person 42, which is captured by the equipment camera 214 and displayed on the display unit 217. Furthermore, if the device 200 is a transparent device as described above, the subject 41 will see the surrounding person 42 located behind the transparent display unit 217 through the transparent display unit 217.
[0023] In this embodiment, in the state shown in Figure 4(A), a CPU 11a, which is an example of a processor provided in the management server 300 (see Figure 2), acquires status information, which is information about the status of the peripheral person 42, who is located in the vicinity of the target person 41. Specifically, the CPU 11a acquires information about the surrounding person 42's situation based on video footage of the surrounding person 42 and audio information, which is information about the surrounding person 42's voice.
[0024] When acquiring situational information about the person in the vicinity 42 based on video footage of the person in the vicinity 42, the CPU 11a acquires the situational information of the person in the vicinity 42 based on video footage obtained from the device camera 214 (see Figure 3) installed on the device 200. In this embodiment, the video obtained by the equipment camera 214 is transmitted to the management server 300 via the communication line 400 (see Figure 1). The CPU 11a of the management server 300 analyzes this video to obtain information about the surrounding person 42. The CPU 11a obtains information about the surrounding person 42 based on the video in which the surrounding person 42 is visible.
[0025] When acquiring situational information about surrounding persons 42 based on their voice information, the CPU 11a acquires the situational information about surrounding persons 42 based on the voice information obtained from the individual microphones 600, which are microphones worn by each of the surrounding persons 42. In this embodiment, audio information obtained by the individual microphones 600 is transmitted to the management server 300 via the communication line 400. The CPU 11a of the management server 300 analyzes this audio information to obtain status information of the surrounding persons 42.
[0026] [Database Description] Figure 5 shows the database stored in the information storage unit 19 (see Figure 2) of the management server 300. In this embodiment, as shown in Figure 5, information about each person 42 is registered in the database stored in the information storage unit 19. In this embodiment, the database has been pre-registered with information such as an identification ID used to identify each of the surrounding persons 42, microphone identification information which is the identification information of the individual microphone 600 that each of the surrounding persons 42 possesses, and facial information of the surrounding persons 42. In this embodiment, the faces of surrounding persons 42 are photographed in advance. Then, facial information, which is the facial information of surrounding persons 42, is registered in the database. The facial information includes images of the faces of 42 people in the vicinity, as well as information about the facial features of these 42 people obtained by analyzing these images.
[0027] In this embodiment, as described above, video footage acquired by the equipment camera 214 and audio information obtained by the individual microphones 600 are transmitted to the management server 300. The CPU 11a of the management server 300 acquires this video and audio information, and based on this video and audio information, acquires status information of the surrounding persons 42. Specifically, when the CPU 11a of the management server 300 acquires video footage from the equipment camera 214, it identifies the person 42 in the video footage based on the image of the person's face in the video footage and the face information stored in the database. Furthermore, the CPU 11a of the management server 300 analyzes this video and obtains status information of the identified person in the vicinity 42.
[0028] Furthermore, when the CPU 11a of the management server 300 acquires voice information obtained by the individual microphone 600, it identifies the person 42 from whom the voice information was acquired by the individual microphone 600 based on the microphone identification information transmitted to the management server 300 along with the voice information and the microphone identification information pre-stored in the database. Furthermore, the CPU 11a of the management server 300 analyzes this audio information to obtain status information of the identified person in the vicinity 42.
[0029] Each of the individual microphones 600 stores microphone identification information to identify itself. The individual microphone 600 transmits the voice information acquired by the individual microphone 600, along with the microphone identification information, to the management server 300. The CPU 11a of the management server 300 identifies the person 42 from whom voice information was acquired by the individual microphone 600, based on this microphone identification information and the microphone identification information registered in the database.
[0030] In this embodiment, an individual microphone 600 (see Figure 4(A)) is provided for each of the surrounding persons 42. In this embodiment, each of the surrounding persons 42 has their own individual microphone 600, which is provided for each of the surrounding persons 42, and their respective voice information is acquired. In this embodiment, the audio information obtained by the individual microphones 600 is transmitted to the management server 300 along with the microphone identification information, as described above. The CPU 11a of the management server 300 identifies the person 42 in the vicinity based on microphone identification information, and further acquires situational information about this identified person 42 based on voice information.
[0031] More specifically, in this embodiment, when transmitting audio information from the individual microphone 600 to the management server 300, the individual microphone 600 selects audio information from the audio information obtained by the individual microphone 600 in which the sound pressure exceeds a predetermined threshold. Then, this selected audio information, along with microphone identification information, is transmitted from the individual microphone 600 to the management server 300. As a result, in this embodiment, the transmission of voice information from other peripherals 42, different from the peripheral 42 to which the individual microphone 600 is attached, to the management server 300 through the individual microphone 600 is suppressed.
[0032] Other persons 42 are located away from the person 42 who is equipped with an individual microphone 600, and the sound pressure of the voices of these other persons 42, which are normally acquired by the individual microphone 600, is low. In a configuration where audio information exceeding a predetermined sound pressure threshold is selected and transmitted to the management server 300, the transmission of audio information from other peripherals 42 to the management server 300 via the individual microphones 600 of the peripherals 42 wearing microphones is suppressed.
[0033] Furthermore, the selection of audio information whose sound pressure exceeds a predetermined threshold may be performed by the management server 300. In this case, the management server 300 selects audio information from the audio information transmitted from the individual microphones 600 whose sound pressure exceeds a predetermined threshold. The management server 300 then acquires this selected audio information as the audio information of the person 42 wearing the microphone.
[0034] In addition, the CPU 11a of the management server 300 may acquire voice information of each of the peripherals 42 based on voice information obtained from microphones installed in terminal devices owned by each of the peripherals 42, such as smartphones and tablet terminals. In addition, a common microphone may be provided, and the CPU 11a of the management server 300 may acquire the voice information of each of the surrounding persons 42 based on the voice information acquired by this common microphone.
[0035] If a common microphone is used, characteristic information, which is information about the voice characteristics of each of the surrounding individuals 42, should be registered in the database beforehand. The CPU 11a of the management server 300 identifies the voice information of each of the surrounding persons 42 based on this characteristic information registered in the database, and obtains status information for each of the surrounding persons 42 based on this voice information.
[0036] [Explanation of specific procedures] In this embodiment, as described above, the CPU 11a of the management server 300 acquires information about the surrounding persons 42 based on video footage acquired by the equipment camera 214 and audio information obtained by the individual microphones 600. Then, if the situation identified by the acquired situation information is a specific situation, the CPU 11a generates control information used to control the device 200 owned by the subject 41 (see Figure 4(A)). More specifically, the CPU 11a generates control information that ensures a predetermined notification is sent to the target person 41 via the device 200.
[0037] More specifically, the CPU 11a generates control information that notifies the target person 41 if the situation identified by the acquired situation information is a situation in which there is a possibility of a surrounding person 42 speaking. More specifically, the CPU 11a generates control information that causes the device 200 to receive a notification indicating that there is a possibility of speech from the surrounding person 42.
[0038] The CPU 11a determines that there is a possibility of a person in the vicinity 42 speaking if the situation identified by the acquired situational information is, for example, one of the following situations. • If there is a breath sound from 42 people nearby If a person in the vicinity 42 makes specific sounds such as "hmm," "um," "ah," or "eh" - When the facial expression of the person in the vicinity 42 changes to a specific expression, such as when the mouth of the person in the vicinity 42 opens wider. If a person in the vicinity 42 looks in the direction of the subject 41 for a predetermined period of time. • When the surrounding person 42 is facing the direction of the target person 41 If the person in the vicinity 42 nods, moves their hands close to their face, or straightens their posture, or if the person in the vicinity 42 performs a predetermined specific action.
[0039] In addition, situational information of surrounding persons 42 may be obtained based on the biometric information of surrounding persons 42. Specifically, the CPU 11a may acquire status information of the person 42 based on biometric information such as pulse, heart rate, and blood pressure obtained from sensors (not shown) attached to the person 42. When acquiring status information of surrounding persons 42 based on biometric information, the biometric information obtained by the sensor and sensor identification information for each sensor are transmitted from the sensor to the management server 300 via a communication line (not shown).
[0040] The CPU 11a of the management server 300 identifies the person in the vicinity 42 based on the sensor identification information, and if the situation identified by the transmitted biometric information is a specific situation, it determines that the identified person in the vicinity 42 may have spoken. Specifically, the CPU 11a of the management server 300 determines, for example, that if values such as pulse rate, heart rate, or blood pressure rise, it may be due to speech from a nearby person 42 identified based on sensor identification information.
[0041] In the processing example shown in Figure 4, the CPU 11a generates control information that causes a notification to be sent to the device 200 indicating that there is a possibility of speech from the person in the vicinity 42 (see Figure 4(A)). This control information causes a display image indicating the possibility of speech from the person in the vicinity 42 to be displayed on the display unit 217 of the device 200. The generated control information is transmitted to device 200, which then controls the display of the display unit 217 based on this control information. As a result, in this embodiment, as shown in Figure 4(B), a display image 45 indicating the possibility of a nearby person 42 speaking is displayed on the display unit 217 of the device 200.
[0042] This displayed image 45 is an image representing what is commonly known as a "speech bubble." In this embodiment, if it is determined that there is a possibility of a person in the vicinity 42 speaking, this display image 45, which consists of an image representing a speech bubble, is displayed on the display unit 217 of the device 200, as shown in Figure 4(B), before the display of the speech content, which will be described later, is performed. This allows the subject 41 to recognize the possibility that the utterance was made by a person in the vicinity 42. The display image 45 is not limited to a 2D (2-dimensional) image, but may also be a 3D (3-dimensional) image. If the display image 45 is a 3D (3-dimensional) image, images with different viewing angles for each eye are displayed on the display unit 217. In other words, if the display image 45 is a 3D (3-dimensional) image, multiple images with different viewing angles are displayed on the display unit 217 as the display image 45.
[0043] Furthermore, the generation of control information may be performed by a server other than the management server 300. In this embodiment, the management server 300 generates the control information, but the system is not limited to this, and the generation of control information may be performed by a device other than the management server 300. The generation of control information may be performed, for example, by device 200. When device 200 generates control information, device 200 determines whether the person in the vicinity 42 is in a specific situation, for example, based on video footage of the person in the vicinity 42 obtained by its own device camera 214 (see Figure 3) and audio information obtained by individual microphones 600.
[0044] The device 200 then generates control information that, if the person in the vicinity 42 is in a specific situation, will send a notification to the device 200 indicating that the person in the vicinity 42 may speak. Specifically, the device 200 generates control information to cause the display image 45 to be displayed on its display unit 217. As a result, the display image 45 is displayed on the display unit 217 of the device 200, similar to when the CPU 11a of the management server 300 generates control information.
[0045] The CPU 11a of the management server 300 generates control information to enable the display image 45 (see Figure 4(B)) to be displayed on the display unit 217. This control information enables the display image 45 to be displayed in a manner that associates it with the surrounding persons 42 visible on the display unit 217. As a result, as shown in Figure 4(B), the display image 45 is displayed in a way that it is associated with the surrounding person 42 that is visible on the display unit 217. More specifically, in this embodiment, the display image 45 is displayed in a manner that it is associated with the head of the person in the vicinity 42.
[0046] The CPU 11a of the management server 300 identifies each of the surrounding persons 42 shown on the display unit 217 based on the video footage acquired by the equipment camera 214 (see Figure 3). Specifically, the CPU 11a of the management server 300 identifies each of the surrounding persons 42 shown on the display unit 217 of the device 200 based on the video of the surrounding persons 42 shown in the video acquired by the device camera 214 and the facial information registered in the database.
[0047] The CPU 11a of the management server 300 then generates control information that allows the display image 45 to be associated with the peripheral persons 42 that have been identified and are judged to have the potential to speak (hereinafter sometimes referred to as "peripheral persons with the potential to speak 42"). Specifically, when generating this control information, the CPU 11a generates control information that includes position information, which is information about the display position of the display image 45.
[0048] The CPU 11a determines the position of the person 42 who may speak, as shown in the video acquired by the device camera 214, as the display position of the display image 45, and generates control information that includes position information, which is information about this display position. More specifically, the CPU 11a determines the position of the head of the person 42 who may speak in the video footage acquired by the device camera 214 as the display position for the display image 45. The CPU 11a then generates control information that includes position information, which is information about the determined display position.
[0049] In this embodiment, control information including this location information is transmitted to the device 200. The device 200 then performs display control so that the display image 45 is displayed at the location identified by this location information. As a result, as shown in Figure 4(B), the display unit 217 of the device 200 displays the display image 45 in a manner that associates it with the person 42 who has the potential to speak.
[0050] Furthermore, if the device 200 is the transparent type device 200 described above, the CPU 11a of the management server 300 generates control information so that the display image 45 is displayed on the portion of the display unit 217 of the device 200 that is located on the straight line connecting the eyes of the target person 41 and the person 42 who may speak nearby. In this case, the CPU 11a of the management server 300 first obtains the angle between the front direction of the device 200 and the direction from the device 200 toward the person 42 who may speak. Specifically, the CPU 11a of the management server 300 analyzes the video acquired by the equipment camera 214 to obtain the angle between the forward direction and the direction toward the person 42 who may be speaking.
[0051] Then, the CPU 11a determines the display position of the display image 45 on the display unit 217 based on this angle, and generates control information that includes information about this determined display position. As a result, even in the transparent device 200, the display image 45 is displayed in a way that associates it with the person 42 who has the potential to speak. In this case, the subject 41 will perceive both the person 42 who is likely to speak in the real space and the displayed image 45 on the display unit 217, which is positioned between the subject's eyes and the person 42 who is likely to speak.
[0052] Subsequently, in this processing example, as indicated by the symbol 4C in Figure 4(C), the actual utterance by the person in the vicinity 42 begins. In other words, the actual utterance by the person in the vicinity 42 who has the potential to utter begins. When the person in the vicinity 42 actually begins to speak, an image 46 indicating that the person in the vicinity 42 has begun speaking is displayed inside the display image 45 which is associated with the person in the vicinity 42, as shown in Figure 4(C). If an actual utterance occurs by a person in the vicinity 42 to whom the displayed image 45 is associated, an image 46 indicating that the utterance by the person in the vicinity 42 has begun will be displayed inside the displayed image 45.
[0053] In other words, in this embodiment, when an actual utterance occurs by a person in the vicinity 42 to whom the display image 45 is associated, an image indicating that the acquisition of voice information by the individual microphone 600 has started is displayed inside the display image 45. Whether or not an actual utterance occurred by the person in the vicinity 42 to whom the displayed image 45 is associated is determined, for example, based on the output from the individual microphone 600 worn by the person in the vicinity 42.
[0054] In this embodiment, an image 46 indicating that a person in the vicinity 42 has started speaking is displayed within the area enclosed by the displayed image 45. This image 46, indicating that the process has started, is displayed on the display unit 217 of the device 200 until the content of the surrounding person's speech 42 is acquired by the CPU 11a of the management server 300.
[0055] Acquiring the speech content by the CPU 11a takes time. In this embodiment, while waiting for the speech content to be acquired by the CPU 11a, an image 46 indicating that the person in the vicinity 42 has started speaking is displayed on the display unit 217. Note that the display of image 46, which indicates that the process has started, is not mandatory. The display shown in Figure 4(B) may be bypassed by the display shown in Figure 4(C), and the display shown in Figure 4(D), which will be explained next, may be displayed instead.
[0056] When the peripheral 42 actually speaks, the CPU 11a of the management server 300 retrieves the content of the peripheral 42's speech. The CPU 11a of the management server 300 analyzes the audio information transmitted from the individual microphone 600 worn by the person in the vicinity 42 to whom the display image 45 is associated, and obtains the content of the person in the vicinity 42's speech. The acquisition of speech content based on the audio information can be done using a known method.
[0057] Next, the CPU 11a of the management server 300 generates control information so that the acquired speech content is displayed on the display unit 217 in a manner that is associated with the surrounding person 42 who made the speech. Then, the CPU 11a of the management server 300 transmits this generated control information to the device 200.
[0058] As a result, in this embodiment, as shown in Figure 4(D), the content of the speech 48 of the person in the vicinity 42 is displayed at a predetermined display location 47 on the display unit 217 of the device 200. In this embodiment, when the person in the vicinity 42 who had the potential to speak actually speaks, the content of the person in the vicinity 42's speech 48 is displayed on the display unit 217 of the device 200. In this embodiment, when the speech content 48 is displayed, it is displayed inside a display image 45 which is an image representing a speech bubble. In this embodiment, the content of the speech 48 of the person 42 who is determined to be likely to speak is displayed inside the display image 45 that is associated with the person 42.
[0059] In the processing example described above, the CPU 11a first generates control information to display the display image 45 on the display unit 217 of the device 200, as described above, when there is a possibility of speech from the person in the vicinity 42. As a result, the display image 45 is displayed on the display unit 217, as shown in Figure 4(B). The CPU 11a generates control information to enable the display image 45 to be displayed, which is the display area 47 of the display unit 217 where the spoken content 48 (not shown in Figure 4(B)) is scheduled to be displayed, and the display image 45 is associated with this display image 45.
[0060] In this embodiment, the area inside the display image 45 (see Figure 4(B)) is the display area 47 where the spoken content 48 is to be displayed. The CPU 11a generates control information to enable the display image 45 to be displayed, such that the display image 45 is associated with the display area 47. More specifically, the CPU 11a generates control information to associate the display image 45 with the display area 47, such as the control information shown in Figure 4(B), which causes the display unit 217 to display an image representing a speech bubble that surrounds the display area 47.
[0061] In this embodiment, after this control information is generated, if the person in the vicinity 42 actually speaks, the content of the person in the vicinity 42's speech 48 is displayed within an area surrounded by a display image 45, which consists of an image representing a speech bubble, as shown in Figure 4(D). In other words, when a person in the vicinity 42 actually speaks, the content of the speech 48 is displayed in the display area 47 located inside the display image 45.
[0062] The CPU 11a of the management server 300 generates control information that enables the device 200 to display a display image 45, as control information that enables the device 200 to display a notification indicating the possibility of speech, thereby enabling the device 200 to display a notification that appeals to the subject 41's visual sense. As a result, in this embodiment, the device 200 provides a notification that the subject 41 can visually confirm.
[0063] [Notification Processing Method] Here, the display image 45 is not limited to an image having a shape that surrounds the display area 47 described above. The shape of the display image 45 is not particularly limited; any image that can be visually confirmed by the subject 41 is acceptable. Another example of the displayed image 45 is, for example, a dotted image. If this dot-shaped image is to be displayed on the display unit 217, then, similar to the image having a surrounding shape as described above, this dot-shaped image will be displayed in a manner that corresponds to the surrounding object 42. Furthermore, when this dot-shaped image is displayed on the display unit 217, the content of the speech 48 of the person in the vicinity 42 is displayed around this dot-shaped image.
[0064] In addition, the display image 45 shown on the display unit 217 of the device 200 may be an image of text indicating the possibility of speech, such as "Speech possible". Furthermore, the visual notification to the subject 41 is not limited to images; it may also be done by illuminating a light source (not shown) provided in the device 200. Furthermore, notifications that appeal to the visual sense of the target person 41 may be made by changing the overall color of the display screen shown on the display unit 217, or by changing the color of a part of the display screen, such as the edges of the display screen shown on the display unit 217.
[0065] In addition, control information that enables the device 200 to notify the device 200 of the possibility of a person in the vicinity 42 speaking may also be generated, for example, by generating control information that vibrates a vibration source (not shown) provided in the device 200. In this case, the target person 41 recognizes the possibility of a person in the vicinity 42 speaking based on the vibration of the device 200. In addition, control information may be generated, for example, to cause sound or voice to be emitted from the speaker 216 (see Figure 3) provided in the device 200. If the subject 41 is hearing impaired, notification by sound is difficult, but if the subject 41 is not hearing impaired, the subject 41 will recognize the possibility of a person 42 speaking through sound.
[0066] If the subject 41 is hearing impaired, when the speech content 48 is displayed as shown in Figure 4(D), the subject 41 can recognize that the person around them 42 is speaking. Here, the display of the utterance content 48 is delayed compared to the actual utterance by the person in the vicinity 42. If the display of the utterance content 48 is delayed, a situation may arise where the utterance content 48 is not yet displayed even though the actual utterance by the person in the vicinity 42 has already begun. In this case, the subject 41 may mistakenly believe that the person in the vicinity 42 is not speaking, and a situation may arise in which the subject 41 speaks while the person in the vicinity 42 is speaking.
[0067] In contrast, in this embodiment, the subject 41 is notified, as described above, that there is a possibility of speech being spoken by the surrounding person 42 before the actual speech is initiated. In this case, if subject 41 is notified that there is a possibility of speech, they will refrain from speaking. In this case, situations such as subject 41 starting to speak after a person nearby 42 has started speaking become less likely. Furthermore, the information processing system 1 of this embodiment also functions effectively when the target person 41 is not a person with a hearing impairment. Even if the subject 41 is not hearing impaired, if the subject 41 is notified that there is a possibility of speech from the person in the vicinity 42, it will become less likely that the subject 41 will speak after the person in the vicinity 42 has started speaking.
[0068] [Other display examples] Figure 6 shows another example of the display on the display unit 217 of the device 200. Figure 6 shows an example of a display in a situation where there is no utterance from the surrounding person 42 and there is no possibility of utterance from the surrounding person 42. The CPU 11a of the management server 300 may generate control information that causes the device 200 to send a notification indicating that there is no speech when the peripheral 42 does not speak. As a result, in this case, as shown in Figure 6, an image 51 indicating that there is no speech is displayed on the display unit 217 of the device 200.
[0069] As shown in Figure 6, image 51, which indicates the absence of speech, is composed of the words "No Speech". By displaying image 51, which indicates the absence of speech, the subject 41 recognizes that the person around them 42 is not speaking. Furthermore, image 51, which indicates the absence of speech, is not limited to images composed of characters; it may also be an image other than one composed of characters, such as a symbol or a graphic.
[0070] Here, let's consider a scenario where the situation shown in Figure 6 presents the possibility of a surrounding person 42 speaking. In this case, the CPU 11a of the management server 300 generates control information that causes the display on the device 200 to switch to the display shown in Figure 4(B). More specifically, the CPU 11a generates control information such that the image 51 indicating no speech is deleted, and the display image 45 shown in Figure 4(B) is displayed on the display unit 217. When the display image 45 shown in Figure 4(B) is displayed, the subject 41 recognizes that it may be a speech by a person nearby 42.
[0071] [Other processing examples] Figures 7(A) to 7(C) show other processing examples. This section explains how to handle situations where there is no actual utterance from the surrounding person 42. In this example, as described above, first, as shown in Figures 7(A) and (B), the possibility of a person in the vicinity 42 speaking arises, and in response, the display image 45 is displayed on the display unit 217 of the device 200. Subsequently, in this processing example, there was no actual utterance by the surrounding person 42. In this case, as shown in Figure 7(C), the display image 45 corresponding to the surrounding person 42 is erased.
[0072] If, after the CPU 11a has generated the control information described above to cause the display image 45 to be displayed on the display unit 217, there is no actual utterance by the person in the vicinity 42, the CPU 11a generates control information to cause the display image 45 displayed on the display unit 217 to be erased. More specifically, the CPU 11a generates control information to erase the display image 45 if, within a predetermined time elapsed since generating the control information to enable the display image 45 to be displayed, or within a predetermined time elapsed since the display image 45 was displayed on the display unit 217, there is no actual utterance by the person in the vicinity 42 corresponding to the displayed display image 45. In this case, the device 200 erases the display image 45 accordingly. As a result, the display image 45 that was displayed on the display unit 217 is erased, as shown in Figures 7(B) and (C).
[0073] [Other processing examples] Figures 8(A) to 8(C) show other processing examples. The CPU 11a of the management server 300 generates control information such that, when multiple peripherals 42 displayed on the display unit 217 of the device 200 are in a specific situation, the display image 45 is displayed in a manner that associates the display image 45 with each of the multiple peripherals 42. As a result, in this case, as shown in Figure 8(B), the display unit 217 of the device 200 displays the display image 45 in a manner that is associated with each of the multiple surrounding persons 42. In this case, the person 41 referring to the display unit 217 of the device 200 recognizes that there is a possibility of speech from multiple people in the vicinity 42.
[0074] Subsequently, when one of the surrounding persons 42 actually speaks, the spoken content 48 is displayed within the area surrounded by the display image 45, which is associated with the surrounding person 42 who actually spoke, as shown by the symbol 8D in Figure 8(C). Furthermore, in this processing example shown in Figure 8, for other peripheral persons 42, indicated by symbol 8E in Figure 8(C), who did not actually speak, the display image 45 that was associated with these other peripheral persons 42 is erased, as shown in Figures 8(B) and (C).
[0075] Although not shown in the diagram, in the state shown in Figure 8(C), if there is a possibility of another person in the vicinity 42, indicated by reference numeral 8E, the display image 45 corresponding to this other person in the vicinity 42 will be displayed again. In this embodiment, if another person 42 has the potential to speak while one of the surrounding persons 42, indicated by reference numeral 8F, is speaking, the display image 45 and speech content 48 corresponding to the former surrounding person 42 are displayed, and a new display image 45 corresponding to the other surrounding person 42 is displayed.
[0076] [Other processing examples] Figures 9(A) to 9(C) and 10 show other processing examples. Figure 10 shows the state when viewed from above in the vertical direction, with the equipment 200, surrounding persons 42, and target person 41. In the processing examples shown in Figures 9 and 10, as shown in Figure 10, some of the surrounding persons 42, indicated by reference numeral 10B, are outside the shooting range 10A of the equipment camera 214 (not shown in Figure 10) installed on the equipment 200. The shooting range 10A can also be considered as the field of view of the subject 41 who is viewing the area in front of them via the device 200. In the processing examples shown in Figures 9 and 10, some of the surrounding persons 42, indicated by reference numeral 10B, are excluded from this field of view. Hereafter, these individuals 42 will be referred to as "hidden individuals 42B". As shown in Figure 9(A), the display unit 217 of the device 200 shows two surrounding persons 42, indicated by code 9D, excluding the invisible surrounding person 42B, and the invisible surrounding person 42B is not visible on the display unit 217 of the device 200.
[0077] Furthermore, in this processing example shown in Figures 9 and 10, the invisible peripheral person 42B, which is not displayed on the display unit 217 of the device 200, is in a specific situation, and there is a possibility that the invisible peripheral person 42B may speak. In this case, the CPU 11a of the management server 300 generates control information to cause a display image 45 indicating the possibility of speech from the hidden peripheral 42B to be displayed on the display unit 217. In this embodiment, control information is generated to ensure that the display image 45 is displayed on the display unit 217 even if there is a possibility that the invisible peripheral person 42B is speaking.
[0078] As a result, in this processing example, as shown in Figure 9(B), a display image 45 corresponding to the invisible peripheral person 42B is displayed on the display unit 217 of the device 200. When subject 41 refers to the display unit 217 in the state shown in Figure 9(B), they recognize that there is a possibility that the speech is coming from a person in the vicinity 42 located outside their field of vision. When the hidden person 42B actually speaks, the content of the hidden person 42B's speech 48 is displayed in a manner that corresponds to the display image 45 corresponding to the hidden person 42B, as shown in Figure 9(C). In this example, the speech content 48 of the hidden person 42B is displayed within the area surrounded by the display image 45 corresponding to the hidden person 42B. In the processing examples shown in Figures 9 and 10, the speech content 48 of the invisible surrounding person 42B, which is not displayed on the display unit 217 of the device 200, is also displayed on the display unit 217 of the device 200.
[0079] In this processing example shown in Figures 9 and 10, in order to display an image 45 corresponding to the hidden person 42B on the display unit 217 of the device 200, a display is made so that the target person 41 can see the direction in which the hidden person 42B is located, as shown in Figure 9(B). In Figure 9(B), the undisplayed peripheral person 42B is located to the left of the front of the device 200, and the display image 45 corresponding to the undisplayed peripheral person 42B is also located to the left of the central part 217C of the display unit 217 in the figure. In this embodiment, the display position of the display image 45 corresponding to the hidden peripheral 42B changes according to the position of the hidden peripheral 42B.
[0080] Figure 11 shows the view of the display unit 217 and surrounding area 42 from the direction indicated by arrow XI in Figure 10. As shown in Figure 11, the CPU 11a generates control information to display a display image 45 corresponding to the hidden peripheral 42B in the portion of the display unit 217 located on the straight line 11L connecting the hidden peripheral 42B and the central portion 217C of the display unit 217. Please refer to Figure 10 for a detailed explanation. Here, we assume a virtual plane 10K along the display unit 217 of the device 200. Furthermore, we assume a line 10H connecting the center of the target person 41 to the non-displaying peripheral person 42B.
[0081] Furthermore, this example assumes the projection of an invisible peripheral object 42B onto a virtual plane 10K. More specifically, we consider the case where the hidden peripheral 42B is projected onto a virtual plane 10K, in the direction of the line 10H mentioned above, from the location where the hidden peripheral 42B is located. In this case, on the virtual plane 10K, the invisible peripheral 42B is located at the location indicated by the symbol 10M. Hereinafter, when the invisible peripheral object 42B is projected onto this virtual plane 10K along the display unit 217, this position of the invisible peripheral object 42B will be referred to as "plane position 10M".
[0082] The CPU 11a generates control information to ensure that the display image 45 (see Figure 11) of the hidden peripheral 42B is displayed on the display unit 217. This control information is generated so that the display image 45 is displayed in a direction that lies on the straight line 11L connecting the planar position 10M and the central part 217C of the display unit 217. As a result, the display image 45 is displayed on the display unit 217 of the device 200 at the location indicated by reference numeral 11X in Figure 11. The subject 41 can determine the direction in which the non-visible surrounding person 42B who may be speaking is located by referring to the display unit 217 shown in Figure 11.
[0083] Furthermore, when acquiring situational information about an unseen peripheral person 42B who is not visible on the display unit 217 of the device 200, based on video footage showing this unseen peripheral person 42B, this situational information about the unseen peripheral person 42B is acquired based on video footage obtained by the overall camera 500 (see Figure 10). Specifically, in this case, the hidden surrounding person 42B is identified based on the video footage obtained by the overall camera 500 and the facial information registered in the database (see Figure 5), and situational information of the identified hidden surrounding person 42B is obtained based on this video footage.
[0084] Furthermore, in order to display the image 45 on the display unit 217 of the device 200, it is necessary to identify the location of the identified non-visible person 42B. In this case, the CPU 11a of the management server 300 analyzes the video footage obtained from the overall camera 500 to determine the location of the hidden peripheral person 42B. Furthermore, in this case, the CPU 11a of the management server 300 analyzes the video obtained from the overall camera 500 to determine the position of the central part 217C of the display unit 217 of the device 200 worn by the subject 41, and the orientation of the device 200.
[0085] Then, the CPU 11a of the management server 300 determines the above-mentioned planar position 10M based on the position of the hidden peripheral 42B, the position of the central part 217C of the display unit 217 of the device 200, and the orientation of the device 200. Next, the CPU 11a of the management server 300 determines the display position of the display image 45 on the display unit 217 based on the identified plane position 10M and the position of the central part 217C of the display unit 217.
[0086] Then, the CPU 11a of the management server 300 generates control information that includes information about the determined display position. The device 200 performs display control on the display unit 217 according to this control information. As a result, the display image 45 is displayed on the display unit 217 of the device 200 on a straight line 11L connecting the planar position 10M and the central part 217C of the display unit 217, as shown in Figure 11.
[0087] In addition, the CPU 11a may generate control information such that, if the situation identified by the acquired situation information is one peripheral 42 looking at another peripheral 42, a notification is sent to the device 200 indicating that there is a possibility of speech being made by the other peripheral 42. In the above, we determined whether or not there was a possibility of utterance by the person in the vicinity 42 based on the situational information of that person in the vicinity 42. However, we are not limited to this, and we may also determine whether or not there is a possibility of utterance by other persons in the vicinity 42 based on the situational information of one person in the vicinity 42.
[0088] If the CPU 11a determines that there is a possibility of speech from this other peripheral 42, it generates control information, for example, that will associate the display image 45 with this other peripheral 42. More specifically, the CPU 11a generates control information that allows the displayed image 45 to be associated with other peripheral objects 42 that are displayed on the display unit 217.
[0089] Specifically, the CPU 11a determines, for example, that there is a possibility of speech from one of the surrounding persons 42, if one of the surrounding persons 42, who is visible in the video obtained by the equipment camera 214, continues to look at another surrounding person 42, who is visible in the same video, for a predetermined period of time. In this case, the CPU 11a generates control information to associate the display image 45 with the other peripheral device 42 and display it accordingly.
[0090] [Explanation of the processing flow] Figure 12 is a flowchart showing the flow of processes that are executed when the above notification process takes place. The following explains the sequence of processes described above. In this embodiment, first, the CPU 11a of the management server 300 determines whether the status of each of the surrounding persons 42 located around the target person 41 is in the specific situation described above (step S101). Then, if the CPU 11a determines that the situation of the peripheral 42 is in a specific state, it identifies the peripheral 42 in this specific state (step S102).
[0091] Subsequently, the CPU 11a generates control information to cause a display image 45 corresponding to the peripheral 42 in a specific situation to be displayed on the display unit 217 (step S103). As a result, a display image 45 indicating the possibility of speech is displayed on the display unit 217 of the device 200. Subsequently, the CPU 11a determines, based on the voice information of the identified person 42, whether or not the identified person 42 actually made a statement (step S104).
[0092] Then, if the CPU 11a does not determine that the identified peripheral person 42 actually made a statement, it generates control information to erase the display image 45 corresponding to the peripheral person 42 (step S105). As a result, the display image 45 displayed on the display unit 217 of the device 200 is erased. On the other hand, if the CPU 11a determines that the identified person 42 has actually spoken, it analyzes the audio information and obtains the speech content 48 corresponding to this person 42 (step S106).
[0093] Next, the CPU 11a generates control information to ensure that the spoken content 48 is displayed within the display image 45 (step S107). In this case, the CPU 11a generates control information to display the utterance content 48 within the display image 45 that is associated with the identified peripheral person 42. As a result, the spoken content 48 is displayed within the displayed image 45.
[0094] [Notification processing regarding the end of an utterance] Next, we will explain the notification process for the end of an utterance. The above describes the notification process regarding the possibility of speech. In addition, notifications indicating the possibility of the end of speech, or notifications indicating that speech has ended, may be sent to the subject 41 via the device 200. In this embodiment, as described above, the CPU 11a of the management server 300 acquires status information, which is information about the situation of the surrounding person 42 before they actually make a speech. Then, if there is a possibility that the surrounding person 42 may make a speech, the CPU 11a makes a notification indicating that there is a possibility of speech, as described above. Hereinafter, in this specification, this situational information, which is information about the situation of the surrounding person 42 before the actual utterance, will be referred to as "pre-utterance situational information."
[0095] Furthermore, in the process described below, after the peripheral 42 actually begins speaking, the CPU 11a of the management server 300 acquires status information, which is information about the status of the peripheral 42 that is speaking. Hereinafter, in this specification, this situational information about the circumstances of the person 42 in the vicinity who is speaking will be referred to as "speech situational information." Furthermore, the CPU 11a generates control information used to control the device 200 if the situation identified by the acquired speech situation information is a specific situation.
[0096] Specifically, the CPU 11a generates control information that causes the device 200 to issue a notification indicating that the utterance of the peripheral 42 may be ending (hereinafter referred to as the "termination possibility indication notification"). In other words, the CPU 11a generates control information that causes the device 200 to issue a termination indication notification indicating that the speech of the person 42, who is the subject of the display image 45, may be ending.
[0097] Furthermore, if the situation identified by the acquired speech status information is in a specific state, the CPU 11a generates control information that causes the device 200 to send a notification (hereinafter referred to as "termination notification") indicating that the speech of the person in the vicinity 42 has ended. In other words, the CPU 11a generates control information that causes the device 200 to send a termination notification indicating that the speech of the person 42 who is the subject of the display image 45 has ended.
[0098] [Notice suggesting possibility of termination] This explains the notification indicating the possibility of termination. The CPU 11a generates control information that, if the situation identified by the acquired speech status information indicates a possibility of the surrounding person 42 ending their speech, a notification suggesting the possibility of termination is issued by this device 200. The CPU 11a generates control information to cause a termination suggestion notification to be issued by the device 200 if the situation identified by the acquired speech status information is, for example, one of the following situations. • When the tone of voice of the person in the vicinity (42) drops, or the pitch of their speech decreases, or when the situation is identified by the audio information, it indicates that the person is in a specific situation. If the person in the vicinity 42 lowers their raised hand or performs any other predetermined action
[0099] [Termination Notice] Next, we will explain the termination notification. The CPU 11a generates control information to cause a termination notification to be sent to the device 200 if the situation identified by the acquired speech status information indicates that the surrounding person 42 has finished speaking. Specifically, the CPU 11a generates control information to cause a termination notification to be sent to the device 200 when the situation identified by the acquired speech status information is, for example, one of the following situations. • If voice information is no longer acquired If the facial expression of the person in the vicinity 42 becomes a specific state, such as when the mouth of the person in the vicinity 42 is closed,
[0100] The CPU 11a acquires information about the speaking status of the person 42 based on the audio information of the person 42 and the video footage of the person 42. More specifically, the CPU 11a acquires information about the speaking status of the surrounding person 42 based on audio information obtained from the individual microphones 600 and video footage obtained from the equipment camera 214 and the overall camera 500. Then, if the situation identified by this speech status information is a predetermined situation, the CPU 11a generates control information that causes the device 200 to issue a termination suggestion notification or a termination notification.
[0101] Furthermore, the information used to acquire pre-utterance situational information may be different from the information used to acquire in-utterance situational information. Specifically, for example, pre-speech situation information may be obtained based on video footage showing the person in the vicinity 42, and in-speech situation information may be obtained based on the voice information of the person in the vicinity 42.
[0102] When acquiring pre-speech situational information, the person in the vicinity 42 often does not make a clear statement. Therefore, acquiring pre-speech situational information based on video footage showing the person in the vicinity 42 tends to improve the accuracy of the judgment regarding the likelihood of them speaking. On the other hand, when acquiring information about the situation while someone is speaking, since the person in the vicinity 42 is actually speaking, acquiring information about the situation while speaking based on audio information tends to improve the accuracy of judging the possibility of speech ending and whether speech has ended compared to acquiring information about speech while speaking based on video.
[0103] The control information described above, which enables device 200 to issue termination indication notifications and termination notifications, is transmitted to device 200 in the same manner as described above. Accordingly, in this embodiment, the device 200 performs control based on this control information and sends a notification to the target person 41 indicating the possibility of termination or a termination notification. As a result, subject 41 recognizes that the person in the vicinity 42 is about to finish speaking, or that the person in the vicinity 42 has finished speaking.
[0104] [Specific examples of processing] Figures 13(A) to (D) show specific examples of the process. Figure 13(A) shows a situation in which a person in the vicinity 42 is speaking. In this embodiment, when the situation shown in Figure 13(A) is in place, the spoken content 48 is displayed within the area enclosed by the display image 45. As shown in Figure 13(A), the CPU 11a of the management server 300 acquires speech status information when the peripheral 42 is speaking. More specifically, the CPU 11a acquires information about the speaking status of the surrounding person 42 to whom the displayed image 45 is associated, based on video footage obtained from the overall camera 500 and the equipment camera 214, and audio information obtained from the individual microphone 600.
[0105] Then, if the situation identified by the acquired speech status information indicates a situation where the surrounding person 42 may be ending their speech, the CPU 11a generates control information that causes a notification suggesting the possibility of termination to be sent to the device 200. Furthermore, if the situation identified by the acquired speech status information indicates that the surrounding person 42 has finished speaking, the CPU 11a generates control information to cause a termination notification to be sent to this device 200.
[0106] Figure 13(B) shows the state of the display unit 217 of the device 200 when the situation identified by the speech status information indicates that the surrounding person 42 may have finished speaking. When the CPU 11a is in a situation where the peripheral 42's speech may be ending, it generates control information to ensure that a notification indicating the possibility of termination is sent to the device 200, as described above. In this example, the CPU 11a generates control information that changes the display image 45 displayed on the display unit 217 of the device 200, which is the control information that causes the device 200 to issue a notification indicating the possibility of termination. In this specification, the control information that enables the termination indication notification to be issued by the device 200 will be referred to as "first control information".
[0107] In this example, the CPU 11a generates control information as first control information, which changes the display image 45 that is displayed in association with the display area 47 of the speech content 48 of the person in the vicinity 42 (see Figure 13(A)). Here, the displayed image 45 can be understood as a corresponding displayed image that is displayed in association with the display location 47 of the speech content 48 of the person in the vicinity 42. In this embodiment, the displayed image 45 representing a speech bubble is displayed as this corresponding displayed image. The CPU 11a generates control information to change the display image 45, which represents this speech bubble and is an example of a corresponding display image, as first control information to enable the device 200 to display a notification indicating the possibility of termination.
[0108] More specifically, the CPU 11a generates control information to change the shape of the display image 45, which represents a speech bubble and is displayed surrounding the display area 47, as this first control information for changing the corresponding display image. Specifically, the CPU 11a generates control information, as first control information for changing the corresponding display image, such that the protruding portion 45G (see Figure 13(A)), which is provided as part of the display image 45 representing the speech bubble, is erased.
[0109] As a result, in this embodiment, the protruding portion 45G is eliminated, as shown in Figures 13(A) and (B). In this embodiment, the protrusion 45G provided on the display image 45, which is displayed in association with a nearby person 42 in a situation where speech may be ending, is erased. The subject 41 recognizes that the protruding part 45G has been erased, which may indicate that the person in the vicinity 42 has finished speaking.
[0110] In this embodiment, as described above, the CPU 11a generates control information such that a display image 45 representing a speech bubble is displayed on the display unit 217 provided in the device 200 when the situation identified by the pre-speech situation information is in a specific state. As a result, in this embodiment, first, as shown in Figure 13(A), the display image 45 representing a speech bubble is displayed on the display unit 217 of the device 200 in a manner that corresponds to the surrounding person 42. This display image 45 is provided with a protruding portion 45G that extends toward the surrounding object 42. In this embodiment, a person 42 who may be speaking, or a person 42 who is in the process of speaking, is positioned at the tip of the protruding portion 45G in the direction of protrusion.
[0111] The CPU 11a generates control information to change the display format of the display image 45 displayed on the display unit 217, as first control information to enable the device 200 to issue a termination indication notification. Specifically, the CPU 11a generates control information as this first control information, which causes the protruding portion 45G of the displayed image 45 to be hidden. As a result, in this embodiment, the protruding portion 45G of the displayed image 45 is hidden, as described above.
[0112] Furthermore, as described above, when the utterance of the peripheral 42 actually ends, the CPU 11a generates control information that causes the device 200 to send an termination notification, which is a notification indicating that the utterance has ended. Hereafter, this control information, which causes the termination notification to be issued by device 200, will be referred to as "second control information." As this second control information, the CPU 11a generates, in this case as well, control information to change the display image 45 displayed on the display unit 217 of the device 200. More specifically, the CPU 11a generates control information as second control information to change the display image 45 that is displayed in association with the display area 47 of the speech content 48 of the person in the vicinity 42.
[0113] More specifically, the CPU 11a generates control information as second control information, which causes the display format of the display image 45 to be changed by the first control information described above, and further changes the display format of the display image 45 after the display format has been changed by the first control information described above. More specifically, the CPU 11a generates control information as second control information, which causes the thickness of the lines that make up the display image 45 to change.
[0114] As a result, in this embodiment, as shown in Figures 13(B) and (C), the lines constituting the display image 45 displayed on the display unit 217 of the device 200 become thinner. In other words, in this embodiment, the lines that make up the display image 45, which is displayed in association with the person 42 speaking, become thinner. As a result, subject 41 recognizes that the person in the vicinity 42 has finished speaking.
[0115] In this embodiment, there is a time difference between the timing when the surrounding person 42 finishes speaking and the timing when the content of the speech 48 at the end of the surrounding person 42's speech is displayed on the display unit 217. In this embodiment, as shown in Figures 13(C) and (D), the display processing of the speech content 48 continues even after the timing of the end of the speech of the person in the vicinity 42, until the entirety of the speech content 48 is displayed. In other words, in this embodiment, even if the lines constituting the displayed image 45 become thinner, the display processing of the utterance content 48 does not end, and this display processing continues until the entirety of the utterance content 48 is displayed.
[0116] The subject 41 can also recognize the end of the display process for the utterance content 48, thereby recognizing the end of the utterance by the person in the vicinity 42. In this embodiment, however, even though the utterance of the person in the vicinity 42 has already finished, the display process continues until the entirety of the utterance content 48 is displayed. In this case, subject 41 is likely to mistakenly believe that person 42 is still speaking, even though person 42 has already finished speaking.
[0117] In this case, there is a tendency for a gap in speech to occur between the end of the utterance by the person in the vicinity 42 and the start of the utterance by the subject 41. In contrast, as in this embodiment, when a notification indicating the possibility of termination or a termination notification is given, the target person 41 can recognize the termination of the utterance by the person in the vicinity 42 at an earlier stage. In this case, subject 41 can begin speaking shortly after the person in the vicinity 42 has finished speaking.
[0118] Figures 14(A) to (I) show a series of steps in the display process. Figure 14(A) shows a situation where there is no utterance from the person in the vicinity 42, and there is no possibility of the person in the vicinity 42 uttering any utterance. In this case, no sound is detected. Also, in this case, the CPU 11a does not generate control information to display the display image 45, and the display image 45 is not displayed on the display unit 217 of the device 200. Figure 14(B) shows a situation in which speech may occur, in which case the display image 45 is displayed on the display unit 217 of the device 200.
[0119] Figures 14(C) to (F) show the situation while the person in the vicinity 42 is speaking. In this case, the CPU 11a acquires the utterance content 48 and generates control information to display this utterance content 48. As a result, the spoken content 48 is sequentially displayed on the display unit 217 of the device 200, as indicated by reference numeral 13X. More specifically, the spoken content 48 is sequentially displayed inside the display image 45 shown on the display unit 217 of the device 200. In this embodiment, there is a time difference between the timing at which the CPU 11a acquires voice information and the timing at which the spoken content 48 is displayed on the display unit 217 of the device 200. Therefore, in this embodiment, as shown by the arrow 14Y in Figure 14, the display of the spoken content 48 is delayed after the acquisition of voice information by the CPU 11a.
[0120] Figure 14(F) shows a situation in which the utterance of the surrounding person 42 may have ended. In this case, in this embodiment, the protruding portion 45G (see Figure 14(E)) which was displayed as part of the display image 45 is erased. Figures 14(G) and later show the situation after the person in the vicinity 42 has finished speaking. In this case, as shown in Figures 14(G) and (H), the lines that make up the displayed image 45 become thinner.
[0121] In this embodiment, when there is a possibility that the utterance is about to end, at least one of the shape, thickness, and color of the display image 45 is changed, and when the utterance has actually ended, at least one of the shape, thickness, and color of the display image 45 is further changed. The above example illustrates a case where the shape of the display image 45 is changed when there is a possibility that the utterance is about to end, and the thickness of the lines that make up the display image 45 is changed when the utterance has actually ended. In other words, the above example explained the case in which the shape of the displayed image 45 is changed first, and then the thickness of the lines that make up the displayed image 45 is changed.
[0122] As another example, for instance, the thickness of the lines that make up the display image 45 may be changed first, and then the shape of the display image 45 may be changed. Alternatively, the shape of the display image 45 may be changed first, and then the shape of this display image 45 may be changed further. Alternatively, the lines constituting the display image 45 may be made thinner first, and then even thinner. Or, the lines constituting the display image 45 may be made thicker first, and then even thicker. Alternatively, the color of the displayed image 45 may be changed first, and then the color of the displayed image 45 may be changed again.
[0123] [Notification Processing Method] In this embodiment, a notification that appeals to the subject 41's vision is given at the end of speech, similar to the notification at the start of speech. Specifically, in this embodiment, as a notification that appeals to the visual sense of the target person 41, the display image 45 is changed for the first time as described above, and then the display image 45 is changed for the second time thereafter. In addition to the visual notification to the subject 41 at the end of speech, other options include displaying images of text such as "Possible end of speech" or "End of speech" on the display unit 217 of the device 200.
[0124] Furthermore, visual notification to the subject 41 at the end of speech may also be provided, for example, by turning on or off a light source (not shown) provided in the device 200. Furthermore, a visual notification to the subject 41 at the end of speech may be made by changing the overall color of the display screen shown on the display unit 217, or by changing the color of a part of the display screen, such as the edges of the display screen shown on the display unit 217.
[0125] Furthermore, notification of the end of speech may be given, for example, by vibrating a vibration source (not shown) provided in the device 200. Furthermore, notification at the end of speech may be provided, for example, by emitting sound or voice from a speaker 216 (see Figure 3) provided in the device 200. If subject 41 is hearing impaired, it becomes difficult to provide notifications by sound. However, if subject 41 is not hearing impaired, notifications of possible termination or termination can be provided to subject 41 by sound.
[0126] [Explanation of the processing flow] Figure 15 is a flowchart showing the processing flow when notification processing is performed at the end of speech. Note that the processing in steps S201 to S207 in Figure 15 is the same as the processing in steps S101 to S107 shown in Figure 12. In this embodiment, first, as described above, the CPU 11a determines whether the situation identified by the pre-speech situation information for each of the surrounding persons 42 located around the target person 41 is the specific situation described above (step S201).
[0127] Then, if the CPU 11a determines that the situation identified by the pre-speech situation information is a specific situation, it identifies the surrounding person 42 who is in this specific situation (step S202). In other words, if the CPU 11a determines that the situation identified by the pre-speech situation information is a situation in which speech is possible, it identifies a nearby person 42 who may be likely to speak.
[0128] Next, the CPU 11a generates control information to cause the display image 45 corresponding to the peripheral 42, which has been identified as being in a specific situation, to be displayed on the display unit 217 (step S203). As a result, the display unit 217 of the device 200 displays the display image 45 shown in Figure 14(B) in a manner corresponding to the nearby person 42 who may be speaking. The displayed image 45 is provided with a protruding portion 45G.
[0129] Subsequently, the CPU 11a determines whether or not the person 42 actually spoke, based on the audio information of the person 42 that was the subject of the displayed image 45 (step S204). Then, the CPU 11a generates control information that causes the display image 45 to be erased if it does not determine that the peripheral person 42 has actually spoken (step S205). As a result, the display image 45 displayed on the display unit 217 of the device 200 is erased. On the other hand, if the CPU 11a determines that the peripheral 42 has actually spoken, it analyzes the audio information and obtains the content of the speech 48 (step S206).
[0130] Next, the CPU 11a generates control information to display the utterance content 48 within the display image 45 (step S207). As a result, the utterance content 48 is displayed within the display image 45, as shown by reference numeral 13X in Figure 14. In addition, the processing examples shown in Figures 13 and 14 above describe a case where a display image 45 composed of thick lines is displayed at the stage where there is a possibility of speech from the surrounding person 42, but the representation form of the display image 45 is not limited to this. When there is a possibility of the person in the vicinity 42 speaking, a display image 45 composed of thin lines may be shown. Then, when the person in the vicinity 42 actually starts speaking, the lines that make up the display image 45 may be made thicker.
[0131] Next, in step S208, the CPU 11a determines whether the utterance of the peripheral 42 is likely to end. Then, if the CPU 11a determines that the peripheral 42 may have finished speaking, it generates control information to erase the protruding portion 45G (step S209). As a result, the protruding portion 45G is erased as described above. Next, the CPU 11a determines whether or not the person in the vicinity 42 has finished speaking (step S210). If the CPU 11a determines that the person in the vicinity 42 has finished speaking, it generates control information to make the lines constituting the display image 45 thinner (step S211).
[0132] 〔others〕 The above describes the case in which three notification processes occur: a notification process regarding the possibility of utterance before utterance, a notification process regarding the possibility of utterance termination after utterance has started, and a notification process regarding the termination of utterance after utterance has started. It is not mandatory for all three processes to be performed; it is acceptable to perform only one or two of them. The above describes a case where, upon the end of an utterance, two notification processes are performed: one to notify the possibility of the utterance ending, and another to notify the system when the utterance has actually ended. However, upon the end of an utterance, only one of these two notification processes may be performed.
[0133] (Note) (((1))) Equipped with a processor, The aforementioned processor, By acquiring audio information from people located around the target person, Based on the aforementioned audio information, the content of the speech, which is the content of the speech of the person in the vicinity who is speaking, is displayed on the display unit of the device owned by the target person. When the person in the vicinity finishes speaking, or when there is a possibility that they will finish speaking, the display image corresponding to the display area of the spoken content shown on the display unit is changed. Information processing system. (((2))) The aforementioned processor, When the person in the vicinity finishes speaking, or when there is a possibility that they will finish speaking, the shape of the displayed image is changed. The information processing system described in (((1))). (((3))) The aforementioned processor, When the person in the vicinity finishes speaking, or when there is a possibility that they will finish speaking, the shape of the display image having a protrusion is changed so that the display image does not have the protrusion. The information processing system described in (((2))). (((4))) The aforementioned processor, When the person in the vicinity finishes speaking, or when there is a possibility that they will finish speaking, the color of the displayed image changes. An information processing system as described in any of (((1))) to (((3))). (((5))) The aforementioned processor, When the speaker in the vicinity finishes speaking, or when there is a possibility that they will finish speaking, the thickness of the lines that make up the displayed image is changed. An information processing system as described in any of (((1))) through (((4))). (((6))) The aforementioned processor, When the person in the vicinity finishes speaking, or when there is a possibility that they will finish speaking, the lines that make up the displayed image are made thinner. The information processing system described in (((5))). (((7))) A function to acquire voice information of people located around the target person, Based on the aforementioned audio information, a display function is provided to display the content of the speech, which is the content of the speech of the person in the vicinity who is speaking, on the display unit of the device owned by the person in question. A change function that changes the display image displayed on the display unit corresponding to the display location of the speech content when the speech of the person in the vicinity has finished, or is likely to finish, A program to make a computer realize this.
[0134] According to the information processing system described in (((1))) to (((6))), it is possible to more easily identify the timing of the utterance that the speaker is about to make, compared to a configuration in which no notification regarding utterances is given to the speaker. According to the program described in (((7))), compared to a configuration in which no notification regarding utterances is given to the speaker, it becomes easier to identify the timing of the utterance that the speaker is about to make. [Explanation of Symbols]
[0135] 1...Information processing system, 10K...Virtual plane, 11a...CPU, 11L...Straight line, 41...Target person, 42...Peripheral persons, 45...Displayed image, 47...Display area, 48...Speech content, 200...Device, 217...Display unit, 217C...Central part
Claims
1. Equipped with a processor, The aforementioned processor, By acquiring audio information from people located around the target person, Based on the aforementioned audio information, the content of the speech, which is the content of the speech of the person in the vicinity who is speaking, is displayed on the display unit of the device owned by the target person. When the person in the vicinity finishes speaking, or when the voice of the person speaking reaches a specific state, or when the person speaking performs a specific action, the display image corresponding to the display area of the speech content displayed on the display unit is changed before the entire content of the speech at the end of the speech is displayed on the display unit. Information processing system.
2. The aforementioned processor, When the person in the vicinity finishes speaking, or when the voice of the person speaking reaches the specified state, or when the person speaking performs the specified action, the shape of the displayed image changes. The information processing system according to claim 1.
3. The aforementioned processor, When the person in the vicinity finishes speaking, or when the voice of the person speaking reaches the specific state described above, or when the person speaking performs the specific action described above, the shape of the display image having a protrusion is changed so that the display image does not have the protrusion. The information processing system according to claim 2.
4. The aforementioned processor, When the person in the vicinity finishes speaking, or when the voice of the person speaking reaches the specified state, or when the person speaking performs the specified action, the color of the displayed image changes. The information processing system according to claim 1.
5. The aforementioned processor, When the person in the vicinity finishes speaking, or when the voice of the person speaking reaches the specified state, or when the person speaking performs the specified action, the thickness of the lines constituting the display image is changed. The information processing system according to claim 1.
6. The aforementioned processor, When the person in the vicinity finishes speaking, or when the voice of the person speaking reaches the specified state, or when the person speaking performs the specified action, the lines constituting the display image are made thinner. The information processing system according to claim 5.
7. A function to acquire voice information of people located around the target person, Based on the aforementioned audio information, a display function is provided to display the content of the speech, which is the content of the speech of the person in the vicinity who is speaking, on the display unit of the device owned by the person in question. When the person in the vicinity finishes speaking, or when the voice of the person speaking reaches a specific state or the person speaking performs a specific action, a change function is provided to change the display image displayed on the display unit that corresponds to the display location of the speech content displayed on the display unit before the entire content of the speech at the end of the speech is displayed on the display unit. A program to make a computer realize this.