Communication support system, communication support method, and program
Patent Information
- Application Number
- JP2025527871
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-06-15
- Filing Date
- 2024-06-06
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2044-06-06
AI Technical Summary
【0008】 本開示の一態様に係るコミュニケーション支援システム等によれば、より適切にコミュニケーションを支援することができる。
Smart Images

Figure 0007923480000001 
Figure 0007923480000002 
Figure 0007923480000003
Abstract
Description
[[Technical Field]]
[0001] The present disclosure relates to a communication support system, a communication support method, and a program. [[Background Art]]
[0002] When communication between people becomes lively (in other words, activated), many advantages are produced, such as smooth progress of discussions, exchange of various opinions, and acquisition of new knowledge combining multiple opinions. Therefore, when people communicate, the degree of activation is important. Conventionally, a communication support system that supports activation of communication is known (see Patent Document 1). [[Prior Art Literature]] [[Patent Documents]]
[0003] [[Patent Document 1]] Japanese Unexamined Patent Publication No. 2017-201479 [[Summary of Invention]] [[Problem to be Solved by the Invention]]
[0004] However, conventional communication support systems sometimes fail to provide appropriate support. Therefore, the present disclosure provides a communication support system and the like for more appropriately supporting communication. [[Means for Solving the Problem]]
[0005] A communication support system according to one aspect of this disclosure includes: a voice acquisition unit that acquires voice information spoken in communication between multiple people; a condition setting unit that sets estimation conditions for estimating the state of the communication based on the voice information acquired from the start of the communication until a predetermined period of time has elapsed; an estimation unit that estimates the state of the communication using the set estimation conditions; and an intervention control unit that controls equipment for providing intervention stimuli to the multiple people based on the estimated state of the communication.
[0006] Furthermore, a communication support method according to one aspect of the present disclosure is a communication support method executed on a server for a communication support system, and includes: a voice acquisition step of acquiring voice information spoken in communication between multiple people; a condition setting step of setting estimation conditions for estimating the state of the communication based on the voice information acquired from the start of the communication until a predetermined period has elapsed; an estimation step of estimating the state of the communication using the set estimation conditions; and an intervention control step of controlling equipment for providing intervention stimuli to the multiple people based on the estimated state of the communication.
[0007] Furthermore, a program relating to one aspect of this disclosure is a program that causes a computer to execute the communication support method described above. [Effects of the Invention]
[0008] According to one aspect of this disclosure, a communication support system, etc., can provide more appropriate support for communication. [Brief explanation of the drawing]
[0009] [Figure 1] Figure 1 is a block diagram showing the functional configuration of a communication support system according to an embodiment. [Figure 2] Figure 2 shows an example of the first database in the communication support system according to the embodiment. [Figure 3] Figure 3 shows an example of a second database in the communication support system according to the embodiment. [Figure 4] Figure 4 shows an example of a third database in the communication support system according to the embodiment. [Figure 5] Figure 5 is a flowchart showing an example of the operation of the communication support system according to the embodiment. [Figure 6A] Figure 6A is a flowchart showing a detailed example of the operation of the estimation unit according to one embodiment. [Figure 6B] Figure 6B is a flowchart showing a detailed example of the operation of the control unit according to one embodiment. [Figure 7A] Figure 7A is a flowchart showing a detailed example of the operation of the estimation unit according to another example of the embodiment. [Figure 7B] Figure 7B is a flowchart showing a detailed example of the operation of the control unit according to another example of the embodiment. [Figure 8A] Figure 8A is a flowchart showing a detailed example of the operation of the estimation unit according to yet another example of the embodiment. [Figure 8B] Figure 8B is a flowchart showing a detailed example of the operation of the control unit according to yet another example of the embodiment. [Modes for carrying out the invention]
[0010] The embodiments will be described in detail below with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, arrangement and connection configurations of components, steps, and the order of steps shown in the following embodiments are examples only and are not intended to limit the disclosure. Furthermore, components in the following embodiments that are not described in an independent claim will be described as optional components.
[0011] It should be noted that each drawing is a schematic diagram and is not necessarily strictly illustrated. In addition, in each drawing, substantially identical configurations are denoted by the same reference numerals, and overlapping descriptions may be omitted or simplified.
[0012] (Embodiment) Hereinafter, the overall configuration of the communication support system according to the embodiment will be described. FIG. 1 is a block diagram showing the functional configuration of the communication support system according to the embodiment.
[0013] In the communication support system according to the present disclosure, as an example of communication between people, communication between participants in a meeting in which the date, time, place, participants, etc. are determined in advance will be described; however, communication to which the communication support system can be applied is not limited to meetings, and all scenes involving conversation between people such as unexpected casual talks, presentations, seminars, etc. are conceivable.
[0014] As shown in FIG. 1, the communication support system 100 is implemented by using devices installed in a space 150 that is a meeting place (corresponding to a meeting room in this example) and a server 200 implemented by an information processing server such as a cloud server or an edge server.
[0015] The space 150 is provided with a device 151 capable of inputting and outputting information to and from the server 200, a selection device 152, an input device 153, a camera 154, and a microphone 155, respectively.
[0016] Each device 151 acts on a meeting participant and is a device that stimulates any of the participant's five senses. In other words, the device 151 is a device that applies an intervention stimulus to meeting participants. The device 151 has a configuration corresponding to the five-sense stimulus it applies. A plurality of devices 151 are provided in the space 150. Furthermore, in the present example, the device 151 is a device that serves both as an environmental device for controlling the environment of the space 150 and as a device that applies an intervention stimulus to meeting participants. In other words, the device 151 applies the intervention stimulus while performing environmental control. Note that an environmental device that controls the environment and a device that applies an intervention stimulus to meeting participants may be provided separately.
[0017] Devices 151 include acoustic devices, lighting devices, airflow generators, and scent generators.
[0018] The acoustic device is an electrical device that applies an intervention stimulus via an auditory stimulus. Therefore, the acoustic device generates sound that stimulates human hearing. The acoustic device outputs so-called environmental sounds into the space 150, for example, the rustle of trees, bird songs such as the chirping of small birds, insect chirps, the sound of a campfire, the sound of water such as the babble of a stream, or the sound of blowing wind such as a breeze. The acoustic device may further play music. In other words, the acoustic device may function as a music playback device. As the acoustic device, in addition to a sound output device such as a speaker, devices such as a signal converter and a gain amplifier can be used in combination.
[0019] A lighting device is an electrical device that provides interventional stimuli through visual stimulation. Therefore, a lighting device generates and emits light that stimulates human vision. Furthermore, a lighting device may generate images that stimulate vision through the emitted light. For example, a lighting device may emit green light, reminiscent of a forest, to various parts of space 150, reflecting the light to illuminate space 150 in the colors of a forest. Alternatively, a lighting device may emit light on the floor to recreate dappled sunlight, or display images showing trees and other plants, animals, or scenes from a forest including a river. The lighting device is preferably indirect lighting such as wall lighting and direct lighting such as spotlights, rather than ambient lighting such as base lights. As a lighting device, a light source device that emits light from a light source such as an LED, and a display device such as a projector (in other words, an image projection device) can be used. When displaying images with a lighting device, either still images or moving images may be displayed. Thus, lighting devices include devices that only provide illumination based on luminance and chromaticity, and devices that emit light that contains information within the light itself, such as images. In the following explanation, the term "video projection device" may be used when referring specifically to video projection devices among lighting equipment.
[0020] An airflow generator is an electrical device that provides interventional stimuli through tactile stimulation. The airflow generator generates wind that stimulates a person's sense of touch through the airflow it generates. For example, the airflow generator can generate wind like that blowing in a forest. A blower or the like can be used as the airflow generator. Alternatively, an air conditioner may be used as the airflow generator. This allows for adjustment of the temperature and humidity within space 150, so that wind with changes in temperature and humidity can be generated as a tactile stimulating element. The airflow generator may also generate airflow to diffuse fragrance components supplied from the fragrance generator (described later) into space 150. For this purpose, the airflow generator should be installed near the fragrance generator. Alternatively, the fragrance generator should be installed near the airflow generator. As a preferred example of such arrangement, the airflow generator and the fragrance generator may be configured as an integrated device. In this way, the fragrance components supplied by the fragrance generator are carried by the airflow and rapidly diffused into space 150, allowing for efficient olfactory stimulation by the fragrance components.
[0021] A fragrance generator is an electrical device that provides interventional stimuli through olfactory stimulation. The fragrance generator emits scents that stimulate a person's sense of smell. It supplies fragrance components into a space 150 that evoke the feeling of being in a forest (for example, the scent of plants). Aroma diffusers and the like can be used as fragrance generators.
[0022] Here, it was explained that these devices 151 may be used for controlling the environment or for providing intervention stimuli to one or more specific meeting participants. When used for controlling the environment, the devices 151 need to act on the entire environment of space 150. On the other hand, when used for providing intervention stimuli to meeting participants, they need to act locally on the location of the participant in space 150. Therefore, the devices 151 are configured to act on each unit location within space 150. Furthermore, the devices 151 can also act on the entirety of these unit locations.
[0023] The selection device 152 is a device that can be used by meeting participants to receive operations from the communication support system 100. The selection device 152 may be a device installed in space 150, or it may be an information terminal owned by at least one of the participants. The selection device 152 receives a selection regarding the type of meeting and transmits that selection to the server 200. Examples of meeting types include brainstorming meetings, decision-making meetings, and casual chat meetings. Such meeting types should be set appropriately according to the purpose that the meeting organizer wants to achieve through the meeting. The above examples of meeting types are just examples, and other types may be set.
[0024] For example, an appropriate method of intervention stimulation in one type of meeting may not be appropriate in another type of meeting. This is because the direction of activation, that is, the participants whose speech is suppressed or promoted by the intervention stimulation, changes depending on the objective to be achieved in the meeting. Therefore, in this embodiment, the type of meeting according to the objective is accepted in advance, and then appropriate intervention stimulation is provided for that type of meeting, so that the activation of the meeting is carried out appropriately.
[0025] Furthermore, depending on the type of meeting, appropriate environmental control tailored to the direction of activation, based on the objectives to be achieved by that meeting, can be effectively implemented. In other words, by suppressing and promoting speech in accordance with the overall progress of the meeting through environmental control, more efficient interaction can be achieved. Therefore, in this embodiment, the type of meeting according to its purpose is selected in advance, and appropriate environmental stimuli are applied to that type of meeting, so the meeting can proceed efficiently. In particular, since there is one or more different scenes depending on the type of meeting in terms of environmental control, it is possible to proceed with the meeting more efficiently by changing the environmental control over time according to the scene. For example, from the perspective of proceeding with the meeting more efficiently, intervention stimuli according to the type of meeting may not be applied, and only environmental control according to the type of meeting may be applied according to the scene.
[0026] Figure 2 shows an example of the first database in the communication support system according to the embodiment. Figure 2 shows a portion of the information stored in database 209. As shown in Figure 2, the environmental control changes over time (i.e., scenarios) for each type of meeting. For example, in a brainstorming meeting, the first environment is an icebreaker environment control. In the icebreaker environment control, the environment is controlled so that green lighting reminiscent of nature, scents such as marjoram pine, and sounds such as birdsong are provided to space 150 to help participants relax and become more comfortable with each other. Next, the second environment is a divergent environment control. In the divergent environment control, the environment is controlled so that bright warm-colored lighting, fluctuating lighting that periodically changes in illuminance and color, scents such as lavender with a relaxing effect, and music such as jazz with a relaxing effect are provided to space 150 to create a pleasant atmosphere so that the total number of utterances from all participants increases while guiding them towards topics related to the meeting agenda. Next, the third environment is a convergent environment control. In the environmental control of the convergence environment, the environment is controlled so that the space 150 is provided with cool-toned, slightly dim lighting that facilitates concentration in order to organize the content.
[0027] Next, as the fourth environment, environmental control of the divergent environment is performed. Next, as the fifth environment, environmental control of the convergent environment is performed. Next, as the sixth environment, environmental control of the divergent environment is performed. Next, as the seventh environment, environmental control of the convergent environment is performed. Thus, in the type of brainstorming meeting, a scenario is set in which environmental control of the divergent environment and the convergent environment are repeatedly performed after the environmental control of the icebreaker environment. If multiple environmental controls are included in one scenario, for example, each environmental control may be performed for a time equal to the number of environmental controls divided by the scheduled time from the start to the end of the meeting. In other words, all environmental controls may be performed evenly within the scheduled time. Alternatively, each environmental control may have a set duration. In other words, each environmental control may be performed for its respective duration, regardless of the scheduled time of the meeting.
[0028] In the decision-making meeting, the first environment, the reporting environment, is controlled. In the reporting environment, the environment is controlled so that the reporting participants speak first, the monitor screen is easy to see, and slightly bright, cool-toned lighting is provided in space 150. Next, the second environment, the divergent environment, is controlled. Next, the third environment, the convergent environment, is controlled.
[0029] In informal meetings, the initial environmental control involves icebreakers.
[0030] Returning to Figure 1, the selection device 152, as will be described in detail later, also has the function of accepting the type of team when the participants in the meeting are considered as a team, in addition to selecting the type of meeting.
[0031] The input device 153 is a device that can be used by meeting participants to allow the communication support system 100 to receive their input. The input device 153 may be a device installed in space 150, or it may be an information terminal owned by at least one of the participants. The input device 153 functions as a trigger to allow meeting participants to perform intervention stimuli at their own timing. Therefore, it is preferable that an input device 153 be provided for each meeting participant, and that each participant be able to operate it at their own discretion.
[0032] Camera 154 is a shooting device for capturing moving images within space 150. Camera 154 transmits the captured moving images to server 200.
[0033] Microphone 155 is a sound-collecting device for collecting sound within space 150. Microphone 155 transmits the collected sound to server 200 in a manner that allows for temporal positional correspondence between the sound and the video image captured by camera 154. Alternatively, if the video image can be temporally correlated with the sound collected by microphone 155, the sound does not need to be temporally correlated.
[0034] Server 200 processes information received from each device in space 150 and performs processing to enable device 151 to control the environment and provide intervention stimuli. Server 200 is realized by executing a predetermined program using a processor and memory. Server 200 comprises a selection unit 201, an acquisition unit 202, an audio acquisition unit 203, a video acquisition unit 204, an estimation unit 205, an input unit 206, a scenario determination unit 207, and a control unit 208. The estimation unit 205 also has a threshold condition setting unit 205a, and the control unit 208 has a database 209.
[0035] The selection unit 201 receives the selection of the conference type from the selection device 152. The selection unit 201 then passes the information regarding the received conference type selection to the control unit 208.
[0036] The acquisition unit 202 accepts the selection of the team type from the selection device 152. Alternatively, the acquisition unit 202 may estimate and determine the type of team as a participant in the meeting at that time based on at least one of the video and audio captured or recorded by the camera 154 and microphone 155, and acquire the team type based on the result of that determination. The acquisition unit 202 passes the acquired information regarding the team type to the control unit 208.
[0037] The voice acquisition unit 203 acquires the sound collected by the microphone 155. The voice acquisition unit 203 passes the voice information of the acquired sound to the estimation unit 205.
[0038] The video acquisition unit 204 acquires video footage captured by the camera 154. The video acquisition unit 204 passes the video information of the acquired video footage to the estimation unit 205.
[0039] The estimation unit 205 estimates the state of the meeting based on audio information and video information. Here, the estimation unit 205 processes the following for each participant in the meeting: the speaking time, which is the duration of the utterance by the person speaking at that time; the silence period, which is the period when all participants are silent (no utterances); the number of times the speaker has changed; and the utterance after a silence period of a certain length or longer.
[0040] These processes can all be carried out using existing technologies, as long as the above parameters can be calculated from the speech. Furthermore, while only speech information is needed to calculate the above parameters from speech, it may be difficult to determine the direction of the speaker relative to the microphone 155 because sounds such as those from environmental control are supplied into space 150. Therefore, in this example, the above parameters are accurately calculated by determining the direction of sound arrival using video information. If technological advancements allow for the accurate calculation of the above parameters from speech information alone, the camera 154 and video acquisition unit 204 may not be necessary.
[0041] Furthermore, the estimation unit 205 estimates the state of the meeting based on whether each of the parameters detected as described above satisfies a threshold condition (determination of whether it is greater than or less than the threshold). In other words, the state of the meeting is a concept indicated by whether each of one or more parameters is greater than or less than the threshold. The threshold here may be set experimentally or empirically in advance, but here, the threshold condition setting unit 205a dynamically sets the threshold based on audio information collected from the start of the meeting until a predetermined period has elapsed. The threshold condition setting unit 205a is an example of a condition setting unit. The setting of threshold conditions by the threshold condition setting unit 205a will be described later. The estimation unit 205 estimates the state of the meeting using the set threshold conditions and passes the estimation result to the control unit 208. Therefore, the threshold conditions can also be understood as estimation conditions for estimating the state of the meeting.
[0042] By dynamically setting threshold conditions using the threshold condition setting unit 205a, the appropriate state of the meeting is estimated according to how the meeting is actually progressing (what state of activity it is in, etc.), and intervention stimuli are provided based on the estimation results. This allows for more appropriate intervention stimuli to be given to participants for each meeting, thus enabling more effective communication support from the perspective of meeting support.
[0043] As described above, in this embodiment, the method of intervention stimulation is changed depending on the type of meeting. Figure 3 is a diagram showing an example of a second database in the communication support system according to this embodiment. As shown in Figure 3, the database used is one that is referenced when providing intervention stimulation, and it links information on the type of communication, the state of communication, and the control of the equipment for providing the intervention stimulation.
[0044] As shown in the diagram, if the meeting type is a brainstorming session, the goal is to create a pleasant atmosphere and encourage participation. Since it is necessary to encourage participation from all participants, for example, if the calculated parameters detect that the speaking time is longer than a threshold (detection of the presence of a participant with a long speaking time), spot lighting is used to provide an intervention stimulus to that participant. Alternatively, if a period of silence longer than a threshold is detected, airflow and scent are used to provide an intervention stimulus to all participants. Alternatively, if a change of speaker is detected (here, by comparing its presence or absence, i.e., with a threshold of 1), lighting effects are used to provide an intervention stimulus to all participants. Alternatively, if a participant who breaks the silence is detected (similarly, by comparing it with a threshold of 1), visual effects are used to provide an intervention stimulus to the entire group, including that participant. In this embodiment, such intervention stimuli are used to support the progress of the meeting so that communication among participants takes place according to the type of meeting.
[0045] Similarly, for other types of meetings, a database has been compiled detailing what kind of intervention stimuli are used under different meeting conditions. By referring to this database, it becomes possible to provide appropriate intervention stimuli to support meetings effectively.
[0046] Furthermore, Figure 4 shows an example of a third database in the communication support system according to the embodiment. The database shown in Figure 4 describes what kind of intervention stimulus to provide when providing an intervention stimulus for each type of team (i.e., team composition). Below, an example of when an intervention stimulus is provided will be explained with reference to Figures 3 and 4. For example, if the type of meeting is a brainstorming meeting and the type of team is a "quiet" team, if the calculated parameters detect that the speaking time is longer than the threshold, no intervention stimulus will be provided. On the other hand, if a silence period longer than the threshold is detected, an intervention stimulus will be subtly provided to all participants through airflow and scent. The intervention stimulus here will be provided with stimulus parameters that are relatively weaker than standard stimulus parameters, such as airflow and scent. Also, if one or more speaker changes, which is the threshold, are detected, an intervention stimulus will be provided to all participants through wall lighting to change the atmosphere. Furthermore, if one or more participants who break the silence are detected, which is the threshold, an intervention stimulus will be provided to the entire group, including the participant who broke the silence, through video effects. Thus, in Figures 3 and 4, if intervention stimuli are set to be applied to either, the intervention stimuli will be applied during that meeting. On the other hand, in Figures 3 and 4, if intervention stimuli are not set to be applied to either, the intervention stimuli will not be applied during that meeting.
[0047] Furthermore, the conditions (database) for intervention stimuli, which are tailored to the type of meeting and the type of team, can include not only the conditions for providing or not providing an intervention stimulus, but also a threshold for the parameter used to determine whether to provide an intervention stimulus (i.e., a threshold for estimating the state of the meeting in the estimation unit), the equipment 151 used to provide the intervention stimulus, and stimulus parameters such as stimulus intensity and duration when providing the intervention stimulus. Since many factors contribute to setting such conditions, for example, the intervention stimuli given during a meeting in advance can be evaluated by participants' subjective evaluations or creativity tests, and a large number of evaluation results can be collected and used to build an intervention control model that has been trained using machine learning along with a large number of intervention stimulus conditions. The conditions for intervention stimuli set by such an intervention control model can also be made more optimal through retraining, so it is desirable to configure the system to allow for model updates and updates to the intervention stimulus conditions (database).
[0048] The input unit 206 receives input from the participant via the input device 153 and interrupts the execution of the intervention stimulus. Therefore, the input unit 206 passes on to the control unit 208 that input has been received. The conditions under which input is received by the input unit 206 are also used to update the conditions of the intervention stimulus.
[0049] The scenario determination unit 207 determines the scenario for controlling the environment of space 150 according to the type of meeting. The scenario determination unit 207 determines the scenario by selecting a predetermined scenario according to the type of meeting. The scenario determination unit 207 passes the determined scenario to the control unit 208.
[0050] The control unit 208 is an example of an environmental control unit and an intervention control unit, and controls environmental equipment for controlling the environment of space 150 according to the type of meeting, and also controls equipment for providing intervention stimuli to meeting participants based on the type of meeting, the state of the meeting, and the type of team.
[0051] The operation of the communication support system, along with the operation of the control unit 208, will be described below with reference to Figure 5. Figure 5 is a flowchart showing an example of the operation of the communication support system according to the embodiment.
[0052] As shown in Figure 5, first, the control unit 208 receives and acquires the selection of the type of meeting (selection step S11). Next, the control unit 208 receives and acquires the selection of the type of team (acquisition step S12). The control unit 208 refers to the database 209 (referencing information as illustrated in Figures 2-4) and reads out the conditions for performing the intervention stimulus (step S13). Meanwhile, the scenario determination unit 207 determines the scenario (step S14) and passes it to the control unit 208.
[0053] Subsequently, as soon as the meeting begins, the control unit 208 starts environmental control according to the scenario (step S15).
[0054] If input is received by the input unit 206 during the meeting (Yes in step S16), the control unit 208 controls the device 151 to provide the intervention stimulus specified in the input or a pre-set stimulus (step S17). At that time, the control unit 208 updates the conditions for providing the intervention stimulus based on the conditions at the time of input (step S18). The control unit 208 then rereads the updated conditions (step S19), returns to step S16, and continues to determine whether or not to provide the intervention stimulus using the latest conditions.
[0055] On the other hand, if there is no input to the input unit 206 (No in step S16), the server 200 acquires audio information and video information (step S20, which includes the audio acquisition step and the video acquisition step). The estimation unit 205 estimates the state of the meeting by calculating various parameters (estimation step S21). If the estimated state of the meeting satisfies the conditions (Yes in step S22), the control unit 208 controls the device 151 to provide an intervention stimulus under the read conditions (intervention control step S23). If the state of the meeting does not satisfy the conditions (No in step S22), the intervention control step S23 is skipped and it is determined whether the meeting has ended or not (step S24). If the meeting has not ended (No in step S24), the process is repeated from step S16. On the other hand, if it is determined that the meeting has ended, for example, if there is silence in the space 150 for a certain period of time (Yes in step S24), the process is terminated.
[0056] Here, the threshold conditions set by the threshold condition setting unit 205a, the state of the meeting estimated by those threshold conditions, and some examples of intervention stimuli controlled by the device 151 based on such estimation results will be explained with reference to Figures 6A to 8B.
[0057] Figures 6A and 6B are flowcharts showing more detailed operation examples in one embodiment. Figure 6A shows an example of the estimation of the meeting state by the estimation unit 205, i.e., the estimation step S21. Figure 6B shows an example of the control of the device 151 by the control unit 208, i.e., the intervention control step S23.
[0058] As shown in Figure 6A, in this example, a fourth threshold is set for the length of speech duration in which at least one of the participants continues speaking. In other words, the threshold condition includes a fourth threshold for the length of speech duration. The fourth threshold is set to avoid prolonged speech by any one participant.
[0059] The threshold condition setting unit 205a sets a fourth threshold from a portion of the acquired audio information from the start of the meeting until a predetermined period has elapsed (step S21a). For example, if the threshold condition setting unit 205a finds that the speakers are relatively concentrated in some of the audio information and that one participant tends to speak for a long time, it decreases the fourth threshold. Conversely, if the speakers are relatively dispersed and that one participant tends to speak for a short time, it increases the fourth threshold. The predetermined period can be extended by continuing to acquire some of the audio information until the necessary number of samples for such a determination is included. Alternatively, if a necessary and sufficient period has been determined experimentally or empirically, the threshold condition may be set using only a portion of the audio information for a fixed predetermined period.
[0060] The estimation unit 205 determines whether the utterance time in the audio information is shorter than or equal to the set fourth threshold (step S21b), estimates the state of the meeting including the determination result, and generates an estimation result (step S21c).
[0061] Then, as shown in Figure 6B, if the control unit 208 indicates that the meeting status is such that the speaking time is longer than the fourth threshold, it controls the lighting device to shine a spotlight on the person (speaker) who is continuing to speak for a length of time longer than the fourth threshold (step S23a). The illumination time is, for example, a short enough time to give attention, such as 5 seconds. Furthermore, in order to encourage other participants to speak, the control unit 208 controls the airflow generator and the fragrance generator to generate an airflow and spray a fragrance, for example, towards the person who has spoken the least in the last 60 seconds (the person with the least speaking volume) (step S23b).
[0062] Figures 7A and 7B are flowcharts showing more detailed operation examples in another embodiment. Figure 7A shows an example of the estimation of the meeting state by the estimation unit 205, i.e., estimation step S21. Figure 7B shows an example of the control of the device 151 by the control unit 208, i.e., intervention control step S23.
[0063] As shown in Figure 7A, in this example, a first threshold is set for the length of silence periods when multiple people remain silent. In other words, the threshold condition includes a first threshold for the length of silence periods. The first threshold is set to avoid prolonged periods of silence.
[0064] The threshold condition setting unit 205a sets a first threshold from a portion of the acquired audio information from the start of the meeting until a predetermined period has elapsed (step S21d). For example, if the threshold condition setting unit 205a finds that there are relatively few speakers in some of the audio information, it decreases the first threshold, and conversely, if there is a tendency for there to be relatively many speakers, it increases the first threshold. The predetermined period can be extended by continuing to acquire some of the audio information until the number of samples necessary for such a determination is included. Alternatively, if a necessary and sufficient period has been determined experimentally or empirically, the threshold condition may be set using only a portion of the audio information for a fixed predetermined period.
[0065] The estimation unit 205 determines whether the silence period in the audio information is shorter than or longer than the set first threshold (step S21e), estimates the state of the meeting including the determination result, and generates an estimation result (step S21f).
[0066] Then, as shown in Figure 7B, if the control unit 208 indicates that the meeting state is in a period of silence longer than a first threshold, it controls the airflow generator and the fragrance generator to encourage participants to speak, for example, by generating an airflow and spraying a fragrance towards the person who has spoken the least in the last 60 seconds (the person with the least speaking volume) (step S23c). Furthermore, if the silence is broken, i.e., any participant speaks, as a result of this intervention stimulus, the control unit 208 provides an intervention stimulus to praise the participant and controls the system to encourage further speaking. Specifically, after step S23c, the control unit 208 determines whether any participant has spoken (step S23d). If a participant has spoken (Yes in step 23d), it controls the video projection device to project an image of praise onto the table or other surface surrounding the meeting participants in space 150 (step S23e). On the other hand, if no participant has spoken (No in step 23d), the control unit 208 does not execute step S23e and proceeds to the next process (step S24).
[0067] Figures 8A and 8B are flowcharts illustrating more detailed operation examples in yet another embodiment. Figure 8A shows an example of the estimation of the meeting state by the estimation unit 205, i.e., estimation step S21. Figure 8B shows an example of the control of the device 151 by the control unit 208, i.e., intervention control step S23.
[0068] As shown in Figure 8A, in this example, a first threshold is set for the length of silence periods in which multiple people remain silent, and a second threshold is set for the number of times one speaker takes over from another. In other words, the threshold conditions include a first threshold for the length of silence periods and a second threshold for the number of alternations. The first threshold is set to avoid prolonged periods of silence, and the second threshold is set to encourage participants to take turns speaking.
[0069] The threshold condition setting unit 205a sets a first threshold and a second threshold from some of the acquired audio information from the start of the meeting until a predetermined period has elapsed (step S21g). For example, if the threshold condition setting unit 205a observes a tendency for relatively few speakers in some of the audio information, it increases the first threshold, and conversely, if it observes a tendency for relatively many speakers, it decreases the first threshold. Also, if the threshold condition setting unit 205a observes a tendency for one participant to speak continuously with few changes in some of the audio information, it sets the second threshold to 5, and conversely, if it observes a tendency for speakers to be relatively dispersed with many changes, it sets the second threshold to 10. The predetermined period can be extended by continuing to acquire some of the audio information until the necessary number of samples for such determination are included. Alternatively, if a necessary and sufficient period has been determined experimentally or empirically, the threshold conditions can be set using only some of the audio information for a fixed predetermined period. The above examples of the number of times used as the second threshold are merely illustrative, and other numbers besides 5 and 10 may be used as the second threshold; the examples above are not particularly limiting.
[0070] The estimation unit 205 determines, with respect to a set first threshold, whether the silence period in the audio information is shorter than the first threshold or longer than or equal to the first threshold (step S21h). The estimation unit 205 also determines, with respect to a set second threshold, whether the number of alternations in the audio information is less than the second threshold or greater than or equal to the second threshold (step S21i). Then, the estimation unit 205 estimates the state of the meeting, including these determination results, and generates an estimation result (step S21j).
[0071] Then, as shown in Figure 8B, the control unit 208 first resets the number of consecutive changes to 0 (step S23f). Next, the control unit 208 determines whether there has been a change in the speaker (step S23g). If there has been a change in the speaker (Yes in step S23g), it adds 1 to the number of consecutive changes (step S23h). On the other hand, if there has been no change in the speaker (No in step S23g), it skips step S23h.
[0072] Here, if the conference state indicates a silence period of a length greater than or equal to the first threshold (Yes in step S23j), the control unit 208 changes the control mode of the light emitted from the lighting device. For example, it reduces the increased light irradiation area (step S23m) and resets the number of consecutive alternations (returns to step S23f). Alternatively, it may stop the irradiation of light by changing the control mode. If the conference state does not indicate a silence period of a length greater than or equal to the first threshold (No in step S23j), it further determines whether the number of consecutive alternations is greater than or equal to the second threshold (step S23k).
[0073] If the number of consecutive changes does not exceed the second threshold (No in step S23k), the process returns to step S23g and continues to determine whether or not a change has occurred. If the number of consecutive changes is equal to or greater than the second threshold (Yes in step S23k), the control unit 208 controls the lighting device to keep it lit (step S23l) and changes the lighting pattern of the lighting device to give a visual effect that praises the large number of consecutive changes.
[0074] Alternatively, the number of times multiple people nodded can be used as an alternative concept to the number of changes in participation. The number of times participants nodded corresponds to the number of utterances that give a positive impression, such as those that are easy to empathize with or agree with. Therefore, it is advisable to set threshold conditions so that the number of times multiple people nodded increases. For example, the threshold condition setting unit 205a sets a third threshold for the number of times multiple people nodded from some of the audio information acquired from the start of the meeting until a predetermined period has elapsed. Then, for example, in the explanations of Figures 8A and 8B, "second threshold" should be read as "third threshold," "number of changes in participation" as "number of nods," and "number of consecutive changes in participation" as "number of consecutive agreements," which indicates the number of times the number of consecutive nods was equal to or greater than the third threshold.
[0075] [Effects, etc.] As described above, the communication support system 100 according to the first aspect of this disclosure includes: a voice acquisition unit 203 that acquires voice information spoken in communication between multiple people; a condition setting unit (threshold condition setting unit 205a) that sets estimation conditions for estimating the state of communication based on the voice information acquired from the start of communication until a predetermined period has elapsed; an estimation unit 205 that estimates the state of communication using the set estimation conditions; and an intervention control unit (part of the function of the control unit 208) that controls a device 151 for providing intervention stimuli to multiple people based on the estimated state of communication.
[0076] According to this method, estimation conditions can be set to estimate the state of communication using audio information from the time communication actually begins. Therefore, since estimation conditions can be set according to the actual communication, more accurate estimations can be made. Thus, based on a more accurately estimated state of communication, more appropriate communication support can be provided.
[0077] Furthermore, the communication support system 100 according to the second aspect of this disclosure is the communication support system 100 described in the first aspect, wherein the estimation condition includes a first threshold for the length of the silence period in which multiple people remain silent, and the estimation unit 205 estimates the state of communication, including whether the silence period in communication is shorter than the first threshold or longer than or equal to the first threshold.
[0078] According to this, a first threshold can be set as an estimation condition for the length of silence periods when multiple people remain silent. Then, the state of communication, including whether the silence period in communication is shorter than the first threshold or longer than the first threshold, can be estimated.
[0079] Furthermore, the communication support system 100 according to the third aspect of this disclosure is the communication support system 100 described in the first aspect, wherein the device 151 includes at least one of a lighting device, an airflow generator, and a fragrance generator, and the intervention control unit, after indicating that the communication state is such that the silence period is longer than or equal to a first threshold, controls at least one of the multiple people to irradiate at least one of the multiple people who have started speaking, to irradiate light from the lighting device, to generate an airflow from the airflow generator, and to spray a fragrance from the fragrance generator.
[0080] According to this, when the state of communication indicates that the silence period is longer than or equal to a first threshold, and at least one of the multiple people begins to speak, at least one of the following controls can be performed: illuminating at least one of the multiple people who started speaking with light from an illumination device, generating an airflow from an airflow generator, and spraying a fragrance from a fragrance generator.
[0081] Furthermore, the communication support system 100 according to the fourth aspect of this disclosure is the communication support system 100 described in the first aspect, wherein the intervention control unit controls the projection of images from the lighting device into a space 150 where multiple people are present.
[0082] According to this, it is possible to control the projection of images from a lighting device into a space 150 where multiple people are present.
[0083] Furthermore, the communication support system 100 according to the fifth aspect of this disclosure is the communication support system 100 described in the first aspect, wherein the estimation conditions include a second threshold for the number of speaker changes in communication, and the estimation unit 205 estimates the state of communication, including whether the number of speaker changes in communication is less than the second threshold or greater than or equal to the second threshold.
[0084] According to this, a second threshold for the number of speaker changes in communication can be set as an estimation condition. Then, the state of communication, including whether the number of speaker changes in communication is less than the second threshold or more than the second threshold, can be estimated.
[0085] Furthermore, the communication support system 100 according to the sixth aspect of this disclosure is the communication support system 100 described in the first aspect, wherein the equipment 151 includes a lighting device, and the intervention control unit changes the control mode of the light emitted from the lighting device when the communication status indicates that the number of exchanges is greater than or equal to the second threshold.
[0086] According to this, if the communication status indicates that the number of alternations is greater than or equal to the second threshold, the control mode of the light emitted from the lighting device can be changed.
[0087] Furthermore, the communication support system 100 according to the seventh aspect of this disclosure is the communication support system 100 described in the first aspect, wherein the estimation conditions further include a first threshold for the length of the silence period in which multiple people remain silent, the estimation unit 205 estimates the state of communication, including whether the silence period in communication is shorter than the first threshold or longer than the first threshold, and the intervention control unit changes the control mode of the light emitted from the lighting device if the state of communication indicates that the number of alternations is greater than or equal to a second threshold, and stops the irradiation of light from the lighting device in the changed control mode if the state of communication indicates that the number of alternations is greater than or equal to a second threshold and then indicates that the silence period is longer than or equal to the first threshold.
[0088] According to this, estimation conditions can be set that further include a first threshold for the length of the silence period during which multiple people remain silent. The state of communication is estimated, including whether the silence period in communication is shorter than the first threshold or longer than the first threshold. If the state of communication indicates that the number of alternations is greater than or equal to the second threshold, the intervention control unit can change the control mode of the light emitted from the lighting device. If the state of communication indicates that the number of alternations is greater than or equal to the second threshold, and then indicates that the silence period is longer than or equal to the first threshold, the irradiation of light from the lighting device in the changed control mode can be stopped.
[0089] Furthermore, the communication support system 100 according to the eighth aspect of this disclosure is the communication support system 100 described in the first aspect, wherein the estimation condition includes a third threshold for the number of times multiple people nodded, and the estimation unit 205 estimates the state of communication, including whether the number of nods in the communication is less than the third threshold or more than the third threshold.
[0090] According to this, a third threshold can be set as an estimation condition for the number of times multiple people nodded. Then, the state of communication can be estimated, including whether the number of nods in the communication was less than the third threshold or more than the third threshold.
[0091] Furthermore, the communication support system 100 according to the ninth aspect of this disclosure is the communication support system 100 described in the first aspect, wherein the device 151 includes a lighting device, and the intervention control unit changes the control mode of the light emitted from the lighting device when the communication status indicates that the number of nods is greater than or equal to the third threshold.
[0092] According to this, if the state of communication indicates that the number of nods is greater than or equal to the third threshold, the control mode of the light emitted from the lighting device can be changed.
[0093] Furthermore, the communication support system 100 according to the tenth aspect of this disclosure is the communication support system 100 described in the first aspect, wherein the estimation condition includes a fourth threshold for the length of speech time in which at least one person among multiple people is continuing to speak, and the estimation unit 205 estimates the state of communication, including whether the speech time in the communication is shorter than the fourth threshold or longer than or equal to the fourth threshold.
[0094] According to this, a fourth threshold can be set as an estimation condition for the length of speech duration in which at least one person among multiple individuals continues speaking. Then, the state of communication can be estimated, including whether the speech duration in communication is shorter than or longer than the fourth threshold.
[0095] Furthermore, the communication support system 100 according to the 11th aspect of this disclosure is the communication support system 100 described in the first aspect, wherein the device 151 includes at least one of a lighting device, an airflow generator, and a fragrance generator, and the intervention control unit, when the state of communication indicates that the speaking time is longer than or equal to the fourth threshold, controls at least one of the following: irradiating light from the lighting device to the person who is continuing to speak among the multiple people, generating an airflow from the airflow generator, and spraying a fragrance from the fragrance generator.
[0096] According to this, if the state of communication indicates that the speech duration is longer than or equal to the fourth threshold, at least one of the following controls can be implemented: illuminating the person who is continuing to speak among multiple people with light from a lighting device, generating an airflow from an airflow generator, and spraying a fragrance from a fragrance generator.
[0097] Furthermore, the communication support system 100 according to the twelfth aspect of this disclosure is the communication support system 100 described in the first aspect, further comprising a video acquisition unit 204 that acquires video information of multiple people, and an estimation unit 205 that estimates the state of communication based on the video information and audio information.
[0098] According to this, video information can be used in addition to audio information when estimating the state of communication. Estimates that cannot be made with audio information alone become possible by using video information. For example, even if it is unclear from audio information alone who made the sound, the speaker can be determined from video information, and the state of communication can be estimated based on that.
[0099] Furthermore, the communication support method according to the 13th aspect of this disclosure is a communication support method executed on a server for a communication support system, and includes: a voice acquisition step (part of step S20) for acquiring voice information spoken in communication between multiple people; a condition setting step (steps S21a, S21d, and S21g, etc.) for setting estimation conditions for estimating the state of communication based on the voice information acquired from the start of the communication until a predetermined period has elapsed; an estimation step S21 for estimating the state of communication using the set estimation conditions; and an intervention control step S23 for controlling equipment for providing intervention stimuli to multiple people based on the estimated state of communication.
[0100] According to this, it will have the same effect as the communication support system 100 described above.
[0101] Furthermore, the program relating to the 14th aspect of this disclosure is a program that causes a computer to execute the communication support method described in the 13th aspect.
[0102] According to this, computers can achieve the same effect as the communication support methods described above.
[0103] (Other embodiments) Although embodiments have been described above, this disclosure is not limited to the embodiments described above.
[0104] For example, in the above embodiment, a process executed by a specific processing unit may be executed by another processing unit. Furthermore, the order of multiple processes may be changed, or multiple processes may be executed in parallel.
[0105] Furthermore, the communication support system in this disclosure may be implemented by multiple devices, each having a portion of the multiple components, or by a single device having all of the multiple components. Also, some of the functions of one component may be implemented as the functions of another component, and the functions may be distributed among the components in any way. Any configuration that substantially includes all the functions necessary to implement the communication support system in this disclosure is included in this disclosure.
[0106] Furthermore, in the above embodiment, each component may be realized by executing a software program suitable for each component. Each component may also be realized by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.
[0107] Furthermore, each component may be implemented by hardware. For example, each component may be a circuit (or integrated circuit). These circuits may form a single circuit as a whole, or they may be separate circuits. Also, each of these circuits may be a general-purpose circuit or a dedicated circuit.
[0108] Furthermore, the general or specific embodiments of this disclosure may be implemented as a system, apparatus, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM. They may also be implemented in any combination of systems, apparatus, methods, integrated circuits, computer programs, and recording media.
[0109] Furthermore, this disclosure may be implemented as a communication support method to be performed by a communication support system. This disclosure may be implemented as a program for causing a computer to perform such a communication support method, or as a computer-readable non-temporary recording medium on which such a program is recorded.
[0110] Furthermore, this disclosure also includes forms obtained by applying various modifications to the embodiments that a person skilled in the art could conceive, or forms realized by arbitrarily combining the components and functions of each embodiment without departing from the spirit of this disclosure. [Explanation of Symbols]
[0111] 100 Communication Support Systems 150 space 151 Equipment 152 Selected Devices 153 Input Devices 154 Cameras 155 Mike 200 servers 201 Selection Section 202 Acquisition Department 203 Voice Acquisition Unit 204 Video Acquisition Unit 205 Estimation section 205a Threshold condition setting unit 206 Input section 207 Scenario Decision Department 208 Control Unit 209 Databases
Claims
1. A speech acquisition unit that acquires spoken audio information in communication between multiple people, A condition setting unit sets estimation conditions for estimating the state of the communication based on the voice information obtained from the start of the communication until a predetermined period has elapsed, An estimation unit that estimates the state of communication using the set estimation conditions, The system includes an intervention control unit that controls equipment for providing intervention stimuli to the multiple persons based on the estimated state of communication, Communication support system.
2. The estimation condition includes a first threshold for the length of the silence period during which the multiple persons remained silent, The estimation unit estimates the state of the communication, including whether the silence period in the communication is shorter than the first threshold or longer than the first threshold. The communication support system according to claim 1.
3. The aforementioned equipment includes at least one of a lighting device, an airflow generator, and a fragrance generator. The intervention control unit, after the communication state indicates that the silence period is longer than or equal to the first threshold, and at least one of the multiple persons begins to speak, controls at least one of the following: irradiate the at least one of the multiple persons who began to speak with light from the lighting device, generate an airflow from the airflow generator, and spray a fragrance from the fragrance generator. The communication support system according to claim 2.
4. The intervention control unit controls the projection of images from the lighting device into the space where the multiple people are present. The communication support system according to claim 3.
5. The estimation conditions include a second threshold for the number of speaker changes in the communication, The estimation unit estimates the state of the communication, including whether the number of alternations in the communication is less than the second threshold or greater than or equal to the second threshold. The communication support system according to claim 1.
6. The aforementioned equipment includes a lighting device, The intervention control unit changes the control mode of the light emitted from the illumination device when the communication status indicates that the number of alternations is greater than or equal to the second threshold. The communication support system according to claim 5.
7. The estimation condition further includes a first threshold for the length of the silence period during which the multiple persons remained silent, The estimation unit estimates the state of the communication, including whether the silence period in the communication is shorter than the first threshold or longer than the first threshold. The intervention control unit, If the communication status indicates that the number of alternations is greater than or equal to the second threshold, the control mode of the light emitted from the illumination device is changed. If the communication state indicates that the number of alternations is greater than or equal to the second threshold, and the silence period is longer than or equal to the first threshold, the irradiation of light from the lighting device in the modified control mode is stopped. The communication support system according to claim 6.
8. The estimation conditions include a third threshold for the number of times the multiple people nodded, The estimation unit estimates the state of the communication, including whether the number of nods in the communication is less than the third threshold or more than the third threshold. The communication support system according to claim 1.
9. The aforementioned equipment includes a lighting device, The intervention control unit changes the control mode of the light emitted from the illumination device when the communication status indicates that the number of nods is greater than or equal to the third threshold. The communication support system according to claim 8.
10. The estimation condition includes a fourth threshold for the length of speech duration during which at least one of the multiple persons continues speaking. The estimation unit estimates the state of the communication, including whether the utterance time in the communication is shorter than the fourth threshold or longer than the fourth threshold. The communication support system according to claim 1.
11. The aforementioned equipment includes at least one of a lighting device, an airflow generator, and a fragrance generator. The intervention control unit, when the state of communication indicates that the speech duration is longer than or equal to the fourth threshold, controls at least one of the following: irradiating the person among the multiple people who is continuing to speak with light from the lighting device, generating an airflow from the airflow generator, and spraying a fragrance from the fragrance generator. The communication support system according to claim 10.
12. The system further includes a video acquisition unit that acquires video information of the aforementioned multiple people, The estimation unit estimates the state of the communication based on the video information and the audio information. A communication support system according to any one of claims 1 to 11.
13. A communication support method executed on a server for a communication support system, A speech acquisition step that acquires spoken audio information in communication between multiple people, A condition setting step of setting estimation conditions for estimating the state of the communication based on the voice information obtained from the start of the communication until a predetermined period has elapsed, An estimation step in which the state of communication is estimated using the set estimation conditions, Includes an intervention control step of controlling equipment for delivering intervention stimuli to the plurality of people based on the estimated state of communication, Communication support methods.
14. To cause a computer to execute the communication support method described in claim 13, program.
Citation Information
Patent Citations
Communication supporting system
JP2017201479A
Information processing system, information processing device, program, information processing method and room
JP2021135701A
Information processing device and program
JP2022049326A
Meeting support device, meeting support system, and meeting support method
WO2022085215A1