Program, information processing apparatus, and control method
The system addresses the challenge of passive listener responses in online communication by detecting speaker events, acquiring listener reactions, and outputting real-time feedback, thereby improving communication effectiveness.
Patent Information
- Application Number
- JP2024090849
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-04
- Publication Date
- 2025-12-16
AI Technical Summary
In online communication scenarios such as lectures and seminars, speakers often face challenges in obtaining responses from passive listeners, making it difficult to gauge audience engagement and adjust their presentation effectively.
A system that detects predetermined events caused by the speaker's actions, acquires listener reactions through various input methods, and outputs reaction information to the speaker, allowing for real-time feedback and adjustment based on listener engagement.
Enables speakers to obtain timely and relevant responses from listeners, enhancing their communication effectiveness by providing real-time feedback on audience engagement.
Smart Images

Figure 2025183008000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a program, an information processing device, and a control method. [Background technology]
[0002] As a technique for remote calls using a telepresence system, Patent Document 1 discloses a technique that aims to reduce stress for users while facilitating smooth communication. Patent Document 1 discloses that in one-to-one remote communication, the movements of the person in conversation are reflected in a telepresence robot. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 7106097 Summary of the Invention [Problem to be solved by the invention]
[0004] Online communication is not limited to one-to-one, but can also be one-to-many, such as in lectures and seminars. When a speaker talks one-sidedly, the listener tends to adopt a passive communication stance, making it difficult for the speaker to get a response from the listener.
[0005] Non-limiting examples of the present disclosure contribute to providing a program, an information processing device, and a control method that allow a speaker to obtain responses from listeners. [Means for solving the problem]
[0006] A program according to one embodiment of the present disclosure causes a computer to execute a listener information acquisition step of determining whether a specified event caused by a speaker's actions has been detected, and if the specified event has been detected, acquiring listener information indicating the listener's reaction, and an output step of outputting reaction information based on the acquired listener information.
[0007] A program according to one embodiment of the present disclosure causes a computer to execute a listener information generation step of generating listener information indicating a listener's reaction, and a listener information transmission step of transmitting the listener information to a reaction information generation device.
[0008] An information processing device according to one embodiment of the present disclosure includes a listener information acquisition unit that determines whether a predetermined event caused by a speaker's action has been detected, and acquires listener information indicating the listener's reaction when the predetermined event has been detected, and an output unit that outputs reaction information based on the listener information acquired from multiple listeners.
[0009] An information processing device according to an embodiment of the present disclosure includes a listener information generating unit that generates listener information indicating a listener's reaction, and a listener information transmitting unit that transmits the listener information to a reaction information generating device.
[0010] A control method according to one embodiment of the present disclosure is a control method for an information processing device, which determines whether a specified event caused by a speaker's action has been detected, and if the specified event has been detected, acquires listener information indicating the listener's reaction, and outputs reaction information based on the listener information acquired from multiple listeners.
[0011] A control method according to an embodiment of the present disclosure is a control method for an information processing device, which generates listener information indicating a listener's reaction, and transmits the listener information to a reaction information generation device. [Effects of the Invention]
[0012] A non-limiting example of the present disclosure allows a speaker to obtain a listener's response.
[0013] Further advantages and benefits of an embodiment of the present disclosure will become apparent from the specification and drawings. Such advantages and / or benefits may be provided by some of the embodiments and features described in the specification and drawings, respectively, but not necessarily all of them may be provided to obtain one or more identical features. [Brief explanation of the drawings]
[0014] [Figure 1] A diagram showing an example of the configuration of a communication system [Figure 2] FIG. 1 shows an example of the configuration of a speaker device. [Figure 3] FIG. 1 shows an example of the configuration of a listener device. [Figure 4] Reaction table diagram [Figure 5] Flowchart showing the flow of speaker processing [Figure 6] Flowchart showing the listener process [Figure 7] Flowchart showing the flow of information acquisition processing [Figure 8] Flowchart showing the flow of aggregation process A [Figure 9] Flowchart showing the flow of aggregation process B [Figure 10] Flowchart showing the flow of output process A [Figure 11] Diagram showing the adjustment table [Figure 12] Flowchart showing the flow of output process B DETAILED DESCRIPTION OF THE INVENTION
[0015] Hereinafter, preferred embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functions are designated by the same reference numerals, and redundant description will be omitted.
[0016] 1 is a diagram showing an example of the configuration of a communication system 10 according to this embodiment. The communication system 10 is made up of a speaker device 100 and a plurality of listener devices 200. The speaker device 100 and the listener devices 200 are connected via a network 20.
[0017] The communication system 10 is used for online lectures or conferences in which the speaker device 100 transmits the speaker's voice and video to the listener device 200. Therefore, the speaker device 100 is an information processing device such as a smartphone, tablet, laptop computer, or desktop computer that can acquire the speaker's voice and video. On the other hand, the listener device 200 is an information processing device such as a smartphone, tablet, laptop computer, or desktop computer that can output the speaker's voice and video transmitted from the speaker device 100.
[0018] 2 is a diagram showing an example of the configuration of the speaker device 100. The speaker device 100 is composed of an input unit 110, an event detection unit 120, a display unit 130, a communication unit 140, a listener information acquisition unit 150, a reaction information output unit 160, and an event database 170. The input unit 110 is composed of a microphone, a camera, a keyboard, and a mouse. Information such as audio and video input to the input unit 110 is output to the event detection unit 120.
[0019] The event detection unit 120 detects predetermined events caused by the speaker's actions from the speaker information input from the input unit 110. To detect predetermined events, the event detection unit 120 is equipped with a voice recognition function, a character recognition function, a character search function, and the like. In this embodiment, the predetermined events used are events caused by the speaker's mouse operation, the speaker's keyboard operation, and the speaker's utterance. Examples of events caused by the speaker's mouse operation include a slide switching operation in a presentation application, a video playback operation, an animation playback operation, and a pointer operation.
[0020] Events caused by keyboard operation include the display of words such as "Important," "Key Points," "Attention," "Issues," "Conclusion," and "Summary" on a slide. Note that the above words may also be displayed when a slide is changed by mouse operation, and this may also be considered a predetermined event.
[0021] Examples of events resulting from the speaker's utterances that were recognized include "Look here," "Pay attention here," "This is the important point," "Remember this," "Do you understand?" and "Did you get it?"
[0022] The event database 170 may be configured to allow predetermined events to be registered for each speaker. For example, a speaker may wish to see only the reaction when a certain phrase is uttered. Specifically, if the speaker is a lecturer, by registering the phrase "Did you understand?" as a predetermined event, the reaction of the listener when "Did you understand?" is uttered can be obtained. In this way, by registering only predetermined events corresponding to the reactions required by the speaker, unnecessary reactions are not output, thereby improving convenience for the speaker.
[0023] When the predetermined event described above is detected, the event detection unit 120 notifies the listener information acquisition unit 150. When the predetermined event is detected, the listener information acquisition unit 150 acquires listener information indicating the listener's reaction. The listener information acquisition unit 150 aggregates the acquired listener information and outputs the results to the reaction information output unit 160. Specific processing by the listener information acquisition unit 150 will be described later.
[0024] The communication unit 140 communicates with the listener devices 200. Specifically, the communication unit 140 transmits the speaker's voice and video to the multiple listener devices 200. The communication unit 140 also receives listener information from the multiple listener devices 200 and outputs it to the listener information acquisition unit 150.
[0025] The reaction information output unit 160 outputs reaction information based on the listener information to the display unit 130. In this embodiment, the reaction information is output to the display unit 130, so that the display unit 130 displays the reaction information. Specific processing by the reaction information output unit 160 will be described later.
[0026] 3 is a diagram showing an example configuration of the listener device 200. The listener device 200 is composed of an input unit 210, a listener information generation unit 220, and a communication unit 230. The input unit 210 differs slightly depending on whether the listener device 200 is a smartphone or a personal computer, but is composed of a microphone, a camera, a keyboard (software keyboard), a mouse (touchpad), and the like.
[0027] The listener information generating unit 220 generates listener information indicating the listener's reaction based on the information input from the input unit 210. The generation of listener information will be explained using Fig. 4. Fig. 4 is a diagram showing a reaction table in which types of input information are associated with three types of listener information.
[0028] There are three types of listener information: positive, negative, and attentive listening. Positive indicates a positive response. Negative indicates a negative response. Attentive listening indicates a response that is neither positive nor negative, or no response at all.
[0029] The types of input information include gestures, facial expressions, voice, icons, and text. The listener information generation unit 220 acquires gestures and facial expressions based on the video input from the camera. The listener information generation unit 220 acquires voice based on the sound input from the microphone. The acquired voice is converted into text using a voice recognition function. The listener information generation unit 220 acquires icons input via a keyboard or mouse. The listener information generation unit 220 also acquires text input via a keyboard or mouse.
[0030] The acquired input information is classified into three types: positive, negative, and attentive listening, as shown in Fig. 4. When the listener information generating unit 220 acquires input information, it refers to the reaction table to determine which listener information corresponds to the acquired input information, and transmits the listener information corresponding to the input information to the speaker device 100.
[0031] For example, if the input information is the gesture "nodding," the listener information generation unit 220 transmits a positive to the speaker device 100. If the input information is the text "I don't know," the listener information generation unit 220 transmits a negative to the speaker device 100. Note that if there is no input information corresponding to the input information in the reaction table, the listener information generation unit 220 does not transmit anything to the speaker device 100. Furthermore, the listener information generation unit 220 may generate listener information based on only one type, or may generate listener information based on multiple types of input information. If based on multiple types of input information, listener information corresponding to each type of input information may be transmitted, or if the listener information is contradictory, nothing may be transmitted. An example of a case where the listener information is contradictory is when the facial expression is "surprised" (positive) and the text is "I don't know" (negative).
[0032] The flow of processing performed by the speaker device 100 and the listener device 200 in the above-described configuration will be described using flowcharts. In the following description, processing performed by the speaker device 100 will be referred to as "speaker processing," and processing performed by the listener device 200 will be referred to as "listener processing." First, the speaker processing and listener processing will be described. Next, details of the speaker processing will be described.
[0033] FIG. 5 is a flowchart showing the flow of speaker processing. This speaker processing is executed from the start of communication between a speaker and a listener to the end of communication. In FIG. 5, the speaker device 100 executes information acquisition processing to detect a predetermined event caused by the speaker's action and acquire listener information indicating the listener's reaction (step S101). The speaker device 100 executes output processing to output reaction information based on the acquired listener information (step S102), and then returns to step S101. The information acquisition processing and output processing shown in FIG. 5 will be described in detail later.
[0034] Fig. 6 is a flowchart showing the flow of listener processing. This listener processing is executed from the start of communication between the speaker and listener to the end. In Fig. 6, the listener device 200 acquires input information (step S201). The listener device 200 refers to the reaction table and determines whether or not anything corresponding to the acquired input information exists in the reaction table (step S202). If nothing corresponding to the acquired input information exists in the reaction table (step S202: NO), the listener device 200 returns to step S201.
[0035] If the response table contains information corresponding to the acquired input information (step S202: YES), the listener device 200 generates listener information to which the acquired input information belongs (step S203). The listener device 200 transmits the listener information (step S204) and returns to step S201.
[0036] In this way, the listener device 200 repeatedly performs the process of transmitting listener information from the start to the end of communication between the speaker and the listener. On the other hand, the speaker device 100 acquires listener information only within a predetermined period after a predetermined event is detected from the start to the end of communication, as will be described with reference to the following FIG. 7, and does not use listener information received at any other time.
[0037] Next, the information acquisition process and output process in the speaker process will be described in detail. Fig. 7 is a flowchart showing the flow of the information acquisition process. In Fig. 7, when speaker information is input from the input unit 110 (step S301), the speaker device 100 determines whether or not a predetermined event caused by the speaker's action has been detected from the speaker information (step S302). If the predetermined event has not been detected (step S302: NO), the process returns to step S301.
[0038] If a predetermined event is detected (step S302: YES), the speaker device 100 acquires the output timing for outputting the reaction information (step S303). For example, the time when the predetermined event is detected is set as the current time, and time T t seconds after that is acquired. Therefore, listener information received within a predetermined period of time t seconds after the detection of the predetermined event is acquired, and reaction information based on this is output. Here, t is preferably, for example, about 2. This is because listener reactions, including relatively time-consuming reactions such as nodding "yes, yes," are usually within 2 seconds at most. If t is shorter than 2 seconds, relatively time-consuming reactions may not be picked up. Conversely, if t is longer than 2 seconds, reactions other than those corresponding to the predetermined event may be picked up. In either case, accuracy cannot be guaranteed. Furthermore, if t is longer than 2 seconds, the speaker may not know the reaction until it is too late, which may leave a bad impression.
[0039] Returning to the explanation of the flowchart, when the speaker device 100 acquires the output timing in step S303, it continues to acquire listener information until the output timing arrives. Specifically, it acquires listener information (step S304), performs a counting process (step S305), and determines whether time T has arrived (step S306). If time T has not arrived (step S306: NO), the process returns to step S304. On the other hand, if time T has arrived (step S306: YES), the speaker device 100 ends the information acquisition process.
[0040] Next, the details of the aggregation process in step S305 above will be described. In this embodiment, there are two types of aggregation processes, which are referred to as aggregation process A and aggregation process B. Aggregation process A is a process that simply aggregates the acquired listener information. On the other hand, aggregation process B is a process that aggregates only the listener information acquired from a specific listener out of the listener information acquired from multiple listeners.
[0041] Fig. 8 is a flowchart showing the flow of the counting process A. In Fig. 8, the speaker device 100 determines whether the received listener information indicates a positive result (step S401). If the listener information indicates a positive result (step S401: YES), the speaker device 100 increments the positive counter by 1 (step S402) and ends the process. This positive counter counts the number of positive results, and is set to 0 at the start of the counting process A.
[0042] If the listener information does not indicate a positive response (step S401: NO), the speaker device 100 determines whether the received listener information indicates a negative response (step S403). If the listener information indicates a negative response (step S403: YES), the speaker device 100 increments the negation counter by 1 (step S404) and ends the process. This negation counter counts the number of negative responses and is set to 0 at the start of the counting process A.
[0043] If the listener information does not indicate a negative response (step S403: NO), the speaker device 100 increments the listening counter by 1 (step S405) and ends the process. This listening counter counts the number of listening attempts, and is set to 0 at the start of the counting process A.
[0044] Next, the counting process B will be described. In the counting process B, the listener device 200 transmits specific information along with the listener information. This specific information is information for determining whether or not to count in the counting process B. The specific information is, for example, attribute information of the listeners. For example, suppose the speaker is a person giving a presentation on a new business, and the listeners are executives who have the decision-making authority to decide whether to grant permission to start the business, and employees who do not. In such a case, the attribute information is used as the executives and employees. Since the speaker wants to know the reactions of the executives, in the counting process B, only the listener information of the executives is counted, and the listener information of the employees is not counted. As a result, the person giving the presentation can know only the reactions of the executives, which allows the presentation to be made more effectively.
[0045] In addition, in lectures, listeners are divided into three groups based on their grades: top, middle, and bottom. Only the listener information of listeners in the bottom group is counted, while the listener information of listeners in other groups is not counted. As a result, the lecturer can only know the reactions of listeners with lower grades, making it possible to deliver lectures more effectively.
[0046] Fig. 9 is a flowchart showing the flow of the counting process B. In Fig. 9, the speaker device 100 determines whether the specific information received together with the received listener information is preset specific information, thereby determining whether the received listener information is to be counted (step S501). If the received listener information is not to be counted (step S501: NO), the speaker device 100 ends the process without doing anything.
[0047] If the received listener information is to be counted (step S501: YES), the speaker device 100 determines whether the received listener information indicates a positive state (step S502). If the listener information indicates a positive state (step S502: YES), the speaker device 100 increments the above-mentioned positive counter by 1 (step S503) and ends the process.
[0048] If the listener information does not indicate a positive response (step S502: NO), the speaker device 100 determines whether the received listener information indicates a negative response (step S504). If the listener information indicates a negative response (step S504: YES), the speaker device 100 increments the negation counter by one (step S505) and ends the process.
[0049] If the listener information does not indicate a negative response (step S504: NO), the speaker device 100 increments the listening counter by one (step S506) and ends the process. According to the counting process B, it is possible to count the listener's reaction information that the speaker wants to know.
[0050] When the above-mentioned counting process is completed, as explained in Figure 7, it is determined whether the output timing has arrived, and when the output timing has arrived, an output process is executed to output reaction information based on the acquired listener information (see Figure 5, step S102).
[0051] In this embodiment, there are two types of output processes, which are referred to as output process A and output process B. Output process A is a process that outputs information indicating any of affirmative, negative, attentive listening, and antagonism as reaction information. Antagonism is output when two or more of the affirmative counter, negative counter, and attentive listening counter reach the same maximum value, that is, when there are multiple counters that reach the maximum value.
[0052] Output process B is a process in which, for example, multiple levels are set for affirmation and negation according to the degree of affirmation or negation, such as strong affirmation, strong negation, weak affirmation, and weak negation, and reaction information is output with the level adjusted according to the purpose of communication between the speaker and the listener.
[0053] Fig. 10 is a flowchart showing the flow of output processing A. In Fig. 10, the speaker device 100 determines whether or not there are multiple counters that have reached the maximum value among the positive counter, negative counter, and listening counter (step S601). If there are multiple counters that have reached the maximum value (step S601: YES), the speaker device 100 outputs a antagonism image to the display unit 130 (step S602), and ends the processing. The antagonism image is an image that indicates that any two or more of positive, negative, and listening are in antagonism.
[0054] If there are no multiple counters that have reached the maximum value (step S601: NO), the speaker device 100 determines whether the positive counter is at the maximum or not (step S603). If the positive counter is at the maximum (step S603: YES), the speaker device 100 outputs a positive image to the display unit 130 (step S604) and ends the process. A positive image is an image that indicates a positive reaction.
[0055] If the positive counter is not at its maximum (step S603: NO), the speaker device 100 determines whether the negative counter is at its maximum (step S605). If the negative counter is at its maximum (step S605: YES), the speaker device 100 outputs a negative image to the display unit 130 (step S606) and ends the process. A negative image is an image that indicates a negative reaction.
[0056] If the negation counter is not at its maximum (step S605: NO), the speaker device 100 outputs an attentive listening image to the display unit 130 (step S607) and ends the process. The attentive listening image is an image indicating that the reaction is attentive listening.
[0057] This output process allows the speaker to obtain a response from the listener. Next, before explaining output process B, we will explain the adjustment table for adjusting the output. Figure 11 is a diagram showing the adjustment table. The adjustment table shows the purpose of communication and how to adjust the response (positive, negative, attentive listening) in that communication.
[0058] In FIG. 11, "promotion," "normal," and "suppression" indicate the type of adjustment made to the output. As described above, output process B has stages according to the degree of affirmation and negativity. For example, there are three stages. The affirmative stages include positive affirmation, normal affirmation, and negative affirmation, with the degree of affirmation being the strongest, followed by normal affirmation, and the weakest, negative affirmation. Similarly, the negative stages include positive negation, normal negation, and negative negation, with the degree of negation being the strongest, followed by normal negation, and the weakest, positive negation. The active listening stages include active listening, normal listening, and negative listening, with the degree of listening being the strongest, followed by normal listening, and the weakest, negative listening.
[0059] Active affirmation, active negation, and active listening are collectively referred to as active responses. Normal affirmation, normal negation, and normal listening are collectively referred to as normal responses. Negative affirmation, negative negation, and passive listening are collectively referred to as passive responses.
[0060] Based on the above, we will explain "promotion," "normal," and "inhibition" in Figure 11. "Promotion" indicates that the reaction indicated by the reaction information is adjusted to a positive reaction. "Normal" indicates that the reaction indicated by the reaction information is adjusted to a normal reaction. "Inhibition" indicates that the reaction indicated by the reaction information is adjusted to a negative reaction. Note that for normal, the stage does not change, but for convenience it is described as being adjusted.
[0061] For example, if the purpose of communication is a conference or academic presentation and the tally results are positive, the adjustment table is set to "promote," so the reaction indicated by the reaction information is adjusted to a more positive affirmative that emphasizes the reaction. On the other hand, if the purpose of communication is a conference or academic presentation and the tally results are negative, the adjustment table is set to "suppress," so the reaction indicated by the reaction information is adjusted to a more negative negative that suppresses the reaction. Also, if the purpose of communication is a class or lecture and the tally results are positive, the adjustment table is set to "normal," so the reaction indicated by the reaction information is adjusted to a normal positive that neither particularly emphasizes nor suppresses the reaction.
[0062] In this way, "promotion" indicates adjusting the response indicated by the response information in a positive direction, while "inhibition" indicates adjusting the response indicated by the response information in a negative direction. In the above example, there were three levels, but it may also be ten levels, for example. In this case, if the normal response is the fifth level, it may be specified that "promotion +2" adjusts the response two levels in a positive direction, or "inhibition -3" adjusts the response three levels in a negative direction.
[0063] Fig. 12 is a flowchart showing the flow of output process B. In Fig. 12, the speaker device 100 determines whether or not there are multiple counters that have reached the maximum value among the affirmative counter, negative counter, and listening counter (step S701). If there are multiple counters that have reached the maximum value (step S701: YES), the speaker device 100 outputs the above-mentioned antagonistic image to the display unit 130 (step S702), and ends the process.
[0064] If there are no multiple counters that have reached the maximum value (step S701: NO), the speaker device 100 acquires the purpose of the communication (step S703). The purpose of the communication may be a preset purpose set before the communication and acquired from a storage unit, or may be acquired from the listener device 200.
[0065] The speaker device 100 determines whether the affirmative counter is at its maximum (step S704). If the affirmative counter is at its maximum (step S704: YES), the speaker device 100 refers to the adjustment table and adjusts the level in accordance with the purpose (step S705). The speaker device 100 outputs a positive image corresponding to the adjustment result (for example, positive affirmation, normal affirmation, negative affirmation, etc.) to the display unit 130 (step S706), and ends the process. The positive image corresponding to the adjustment result is an image that indicates that the reaction is positive and the degree of that positive reaction. For example, it is an image that says "positive affirmation."
[0066] If the positive counter is not at its maximum (step S704: NO), the speaker device 100 determines whether the negative counter is at its maximum (step S707). If the negative counter is at its maximum (step S707: YES), the speaker device 100 refers to the adjustment table and adjusts the adjustment stepwise according to the purpose (step S708). The speaker device 100 outputs a negative image corresponding to the adjustment result (for example, positive negative, normal negative, negative negative, etc.) to the display unit 130 (step S709), and ends the process. The negative image corresponding to the adjustment result is an image that indicates that the reaction is negative and the degree of negativity. For example, it is an image that depicts "positive negative" or the like.
[0067] If the negation counter is not at its maximum (step S707: NO), the speaker device 100 refers to the adjustment table and makes a stepwise adjustment according to the purpose (step S710). The speaker device 100 outputs an attentive listening image corresponding to the adjustment result (for example, active listening, normal attentive listening, passive listening, etc.) to the display unit 130 (step S711), and ends the process. The attentive listening image corresponding to the adjustment result is an image that indicates that the reaction is attentive listening and the degree of attentive listening. For example, it is an image that depicts "active attentive listening" or the like.
[0068] In this way, the adjustment table allows the degree of positivity or negativity to be adjusted according to the purpose of communication, so that reaction information that emphasizes the speaker or, conversely, suppresses the speaker's reaction information can be freely output.
[0069] For example, in Figure 11, if the purpose of communication is a class or lecture and the aggregated results show a negative response, "promotion" is set in the adjustment table. In a class or lecture, a negative response is considered to mean that the listener does not understand what is being said. If the listener does not understand what is being said in a class or lecture, the purpose of the class or lecture will not be achieved, which is undesirable. Therefore, by outputting a positive negation to the speaker, it is possible to more strongly convey to the speaker that the listener does not understand, compared to when a normal negation is output.
[0070] The above-described two counting processes A and B and the two output processes A and B can be freely combined. That is, the speaker device 100 can execute any of the following combinations: a combination of counting process A and output process A; a combination of counting process A and output process B; a combination of counting process B and output process A; and a combination of counting process B and output process B.
[0071] <Modification> In the embodiment described above, the reaction information is output to the display unit 130, but this is not limiting. For example, the reaction information may be output to a speaker and conveyed to the speaker using voice. Alternatively, the reaction information may be output to a robot and conveyed to the speaker by the robot's movements or behavior.
[0072] In the above-described embodiment, the maximum counter value is output as the reaction information, but a percentage may also be output. For example, if the positive counter is 5, the negative counter is 3, and the listening counter is 2, information indicating 50% positive, 30% negative, and 20% listening may be output as the reaction information. In this case, the speaker can check the percentages other than the maximum reaction information.
[0073] In the above-described embodiment, the speaker device 100 executes the speaker processing. However, a server may be provided between the speaker device 100 and the listener device 200, and the server may execute the speaker processing. In this case, the listener device 200 transmits listener information to the server. The server executes the speaker processing shown in FIG. 5, and sets the output destination in the output processing to the speaker device 100 instead of the display unit 130. The speaker device 100 displays an image output from the server, or expresses reaction information using a robot as described above. In this embodiment, the processing load on the speaker device 100 can be reduced, so a device with relatively low processing power can be used as the speaker device 100.
[0074] In the above-described embodiment, there are three types of listener information: positive, negative, and attentive listening. However, there may be two types: positive and negative. In this case, in the reaction table of Fig. 4, the type corresponding to attentive listening may be deleted, or if it corresponds to the type corresponding to attentive listening, it may not be counted, and reaction information may be output based only on the aggregated results of positive and negative.
[0075] In the above-mentioned Figures 8 and 9, it is determined whether the listener information indicates positive or negative, and then it is determined whether the listener information indicates negative or positive. However, it is also possible to determine whether the listener information indicates negative or positive, and then it is also possible to add a process to determine whether the listener information indicates attentive listening. In this case, the order in which the listener information is determined to be positive, negative, or attentive listening does not matter.
[0076] 10 and 12, it is determined whether the positive counter is at its maximum, and then it is determined whether the negative counter is at its maximum. However, it may be determined whether the positive counter is at its maximum, after it is determined whether the negative counter is at its maximum. Furthermore, a listening counter that counts listening may be provided, and a process for determining whether the listening counter is at its maximum may be added. In this case, the order in which it is determined whether the positive counter is at its maximum, whether the negative counter is at its maximum, and whether the listening counter is at its maximum does not matter.
[0077] Although the embodiments of the present invention have been described above in detail with reference to the drawings, the functions of the above-described devices can be realized by a computer program.
[0078] A computer that realizes the functions of each of the above-mentioned devices by a program reads the program for realizing the functions of each of the above-mentioned devices from a recording medium on which the program is recorded and stores the program in a storage device. Alternatively, the computer communicates with a server device connected to a network and downloads the program for realizing the functions of each of the above-mentioned devices from the server device and stores the program in a storage device.
[0079] The CPU of the computer copies the program stored in the storage device to RAM, and sequentially reads and executes the instructions contained in the program from RAM, thereby realizing the functions of each of the devices.
[0080] <Summary of the embodiment> A program according to one embodiment of the present disclosure is a program for causing a computer to execute a listener information acquisition step (step S304) of determining whether a predetermined event caused by a speaker's action has been detected, and if the predetermined event has been detected (step S302: YES), acquiring listener information indicating the listener's reaction, and an output step (step S102) of outputting reaction information based on the acquired listener information.
[0081] An information processing device according to one embodiment of the present disclosure includes a listener information acquisition unit (listener information acquisition unit 150) that determines whether a predetermined event caused by a speaker's action has been detected, and acquires listener information indicating the listener's reaction when the predetermined event has been detected, and an output unit (reaction information output unit 160) that outputs reaction information based on the listener information acquired from multiple listeners.
[0082] A control method according to one embodiment of the present disclosure is a control method for an information processing device, which determines whether a predetermined event caused by a speaker's action has been detected, and if the predetermined event has been detected (step S302: YES), acquires listener information indicating the listener's reaction (step S304), and outputs reaction information based on the acquired listener information (step S102).
[0083] Although various embodiments have been described above with reference to the drawings, it goes without saying that the present disclosure is not limited to such examples. It is clear that a person skilled in the art can conceive of various modifications or alterations within the scope of the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure. Furthermore, the components of the above-described embodiments may be combined in any manner without departing from the spirit of the disclosure.
[0084] Although specific examples of the present disclosure have been described in detail above, these are merely examples and do not limit the scope of the claims. The technology described in the claims includes various modifications and alterations of the specific examples exemplified above. [Industrial Applicability]
[0085] An embodiment of the present disclosure is suitable for online communication. [Explanation of symbols]
[0086] 10. Communication System 20 Network 100 Speaker device 110 Input section 120 Event detection unit 130 Display section 140 Communications Department 150 Listener Information Acquisition Unit 160 Reaction information output unit 170 Event Database 200 Listener Device 210 Input section 220 Listener Information Generation Unit 230 Communications Department
Claims
1. On the computer, a listener information acquisition step of determining whether a predetermined event caused by a speaker's action has been detected, and acquiring listener information indicating a listener's reaction when the predetermined event has been detected; an output step of outputting reaction information based on the acquired listener information; A program to execute.
2. 2. The program according to claim 1, wherein the output step outputs the reaction information based on the listener information acquired from the listener within a predetermined period after the predetermined event is detected.
3. The program according to claim 1 , wherein the reaction information is information based on listener information acquired from a predetermined specific listener among the listener information acquired from a plurality of listeners.
4. 2. The program according to claim 1, wherein the reaction information includes positive information indicating a positive reaction and negative information indicating a negative reaction.
5. The positive information is provided with a plurality of stages according to the degree of affirmation, and the negative information is provided with a plurality of stages according to the degree of denial, 5. The program according to claim 4, wherein the reaction information is information in which the level of the positive information or the level of the negative information is adjusted depending on the purpose of communication between the speaker and the listener.
6. On the computer, a listener information generating step of generating listener information indicating a listener's reaction; a listener information transmitting step of transmitting the listener information to a reaction information generating device; A program to execute.
7. a listener information acquisition unit that determines whether a predetermined event caused by a speaker's action has been detected, and acquires listener information indicating a listener's reaction when the predetermined event has been detected; an output unit that outputs reaction information based on the listener information acquired from a plurality of listeners; An information processing device comprising:
8. a listener information generating unit that generates listener information indicating listener responses; a listener information transmitting unit that transmits the listener information to a reaction information generating device; An information processing device comprising:
9. A control method for an information processing device, comprising: determining whether a predetermined event caused by the speaker's action has been detected, and if the predetermined event has been detected, acquiring listener information indicating the listener's reaction; outputting reaction information based on the listener information acquired from a plurality of listeners; Control method.
10. A control method for an information processing device, comprising: Generate listener information indicating listener responses; transmitting the listener information to a reaction information generating device; Control method.
Citation Information
Patent Citations
Telepresence System
JP7106097B2