Speech desire estimation device, speech desire estimation method, and program
The speaking desire estimation device in remote conference devices calculates a user's desire to speak based on operation information, addressing the challenge of estimating speaking desire without video and audio, and enhancing communication by preventing speech collisions.
Patent Information
- Application Number
- JP2023561954
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-11-16
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-11-16
AI Technical Summary
In remote conferences, it is challenging to estimate a user's desire to speak without relying on video and audio information, especially when users turn off their cameras and microphones due to noise or line pressure.
A speaking desire estimation device is implemented in conference devices to generate operation information based on user interactions during a remote conference. This information is used to calculate a speaking desire degree, which is then transmitted to other conference devices, allowing users to determine the speaking desire of others without using video and audio.
The solution enables effective estimation of a user's speaking desire without video and audio, facilitating smoother communication by preventing speech collisions and allowing users to easily identify who wants to speak.
Smart Images

Figure 0007687433000001 
Figure 0007687433000002 
Figure 0007687433000003
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for estimating a user's desire to speak in a remote conference.
Background Art
[0002] In a remote conference such as a Web conference, it is difficult to grasp a person who wants to speak (a person with a desire to speak) compared to real face-to-face communication due to effects such as unclear video and network delay.
[0003] Patent Document 1 discloses a technique for acquiring the behavior of a user (a participant in a remote conference) from a camera and a microphone, calculating the degree of the user's desire to speak, and displaying it. According to this technique, each user can easily grasp who wants to speak.
[0004] However, in a remote conference, it is often done to prevent interference with communication due to line pressure and noise by turning off the camera and microphone, and there is a problem that it is difficult to perform estimation of the desire to speak using video and audio.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] An object of the present invention is to provide a technique for estimating a user's desire to speak without using video and audio information.
Means for Solving the Problems
[0007] The speaking desire estimation device according to one aspect of the present invention is provided in a first conference device among a plurality of conference devices used for a remote conference via a communication network, and generates operation information indicating an operation performed by a user on the first conference device during the remote conference. An operation information generation unit, a speaking desire degree calculation unit that calculates a speaking desire degree indicating the degree to which the user desires to speak based on the generated operation information, and information based on the calculated speaking desire degree is transmitted to a second conference device among the plurality of conference devices. And a communication unit.
Effect of the Invention
[0008] According to the present invention, a technique for estimating a user's speaking desire without using video and audio information is provided.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Mode for Carrying Out the Invention
[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings.
[0011] An embodiment relates to a conference system in which a plurality of users existing in different locations hold a remote conference using a plurality of conference devices connected to a communication network. In one embodiment, each conference device includes a speaking desire estimation device that estimates the speaking desire of the user who uses the conference device. The speaking desire estimation device calculates the degree of the user's speaking desire based on the operations performed by the user on the conference device during the remote conference, and transmits information based on the calculated degree of the speaking desire to other conference devices. The degree of the speaking desire indicates the degree (extent) to which the user desires to speak. Each conference device receives information indicating the degree of the speaking desire of other users from other conference devices, and presents the received information to the user. According to the conference system according to the embodiment, it is possible to estimate the speaking desire of each user without using video and audio information, and each user can easily determine whether or not other users desire to speak. As a result, it is possible to avoid collisions of speech.
[0012] <First Embodiment> [Configuration] FIG. 1 schematically shows a conference system 10 according to the first embodiment. As shown in FIG. 1, the conference system 10 includes a plurality of clients 11 each used by a plurality of users, and a server 12 connected to the clients 11 via a communication network 19. The communication network 19 may include the Internet, an intranet, or a combination of the Internet and an intranet. The server 12 relays data between the clients 11. For example, the server 12 receives data from the clients 11 via the communication network 19, and transfers the received data to other clients 11 via the communication network 19.
[0013] Each client 11 can be a computer such as a personal computer (PC). The client 11 corresponds to a conference device used for remote conferences via the communication network 19. In the present embodiment, the client 11 functions as a conference device by executing a remote conference application. In other embodiments, the client 11 may function as a conference device by accessing the server 12 using a browser.
[0014] The clients 11 can have the same or similar configurations to each other. Hereinafter, the configuration of one client 11 will be described as a representative.
[0015] FIG. 2 schematically shows the functional configuration of the client 11 according to the present embodiment. As shown in FIG. 2, the client 11 includes a control unit 21, an input unit 22, an output unit 23, a communication unit 24, an operation information generation unit 25, a speaking desire degree calculation unit 26, and a storage unit 29. The storage unit 29 includes an operation information storage unit 291 and a rule storage unit 292. The control unit 21, the operation information generation unit 25, and the speaking desire degree calculation unit 26 are collectively referred to as a processing unit 27. The control unit 21, the communication unit 24, the operation information generation unit 25, the speaking desire degree calculation unit 26, the operation information storage unit 291, and the rule storage unit 292 correspond to the speaking desire estimation device according to the present embodiment.
[0016] The control unit 21 controls the operation of the client 11. Specifically, the control unit 21 controls the input unit 22, the output unit 23, the communication unit 24, the operation information generation unit 25, the speaking desire degree calculation unit 26, and the storage unit 29.
[0017] The input unit 22 receives input from the user and sends the received input to the control unit 21. In the example shown in FIG. 2, the input unit 22 includes a mouse 221, a camera 222, and a microphone 223. The mouse 221 enables the user to operate the client 11. For example, the mouse 221 enables the user to operate the user interface provided by the remote conferencing application. Instead of or in addition to the mouse 221, a touch pad (track pad), a touch panel, a keyboard, etc. may be used. The camera 222 images the user and generates video data showing the video of the user. The camera 222 may be provided with a physical button for switching the camera 222 between on and off. The microphone 223 collects the sound uttered by the user and generates audio data indicating the user's voice. The microphone 223 may be provided with a physical button for switching the microphone 223 between on and off. The control unit 21 receives video data from the camera 222 and audio data from the microphone 223, and transmits the video data and the audio data to other clients 11 via the communication unit 24.
[0018] The output unit 23 outputs the information generated by the control unit 21 to the user. In the example shown in FIG. 2, the output unit 23 includes a display device 231 and a speaker 232. The display device 231 is a display such as a liquid crystal display device, and displays the image generated by the control unit 21. For example, the control unit 21 generates an image including the user interface provided by the remote conferencing application, and the display device 231 displays the image including the user interface. The user interface includes an area for displaying the video of other users. The control unit 21 receives the video data of other users from other clients 11 via the communication unit 24, and applies the received video data to the user interface in order to display the video of other users on the user interface. The speaker 232 emits sound according to the audio data supplied by the control unit 21. For example, the control unit 21 receives the audio data of other users from other clients 11 via the communication unit 24, and sends the received audio data to the speaker 232 so that the speaker 232 outputs the voice of other users.
[0019] Figure 3 schematically shows a user interface 30 for a remote meeting provided by a remote meeting application. In the example shown in Figure 3, the user interface 30 includes a video area 31 and a control bar 32. The video area 31 is an area for displaying the videos of other users. The control bar 32 includes a mute button 321, an audio setting button 322, a video button 323, and a video setting button 324.
[0020] The mute button 321 is a button for switching the audio input between on (enabled) and off (disabled). When the mute button 321 is clicked while the audio input is on, the audio input switches to off, and when the mute button 321 is clicked while the audio input is off, the audio input switches to on. When the audio input is on, the audio data obtained by the microphone 223 is sent to other clients 11, and when the audio input is off, the audio data obtained by the microphone 223 is not sent to other clients 11.
[0021] The audio setting button 322 is a button for displaying an audio-related list. The audio-related list includes a plurality of items such as microphone settings and speaker settings. When an item in the microphone settings is selected (clicked), a microphone setting screen for setting the microphone 223 is displayed. On the microphone setting screen, it is possible to adjust the volume of the microphone 223.
[0022] The video button 323 is a button for switching the video input between on and off. When the video button 323 is clicked while the video input is on, the video input switches to off, and when the video button 323 is clicked while the video input is off, the video input switches to on. When the video input is on, the video data obtained by the camera 222 is sent to other clients 11, and when the video input is off, the video data obtained by the camera 222 is not sent to other clients 11.
[0023] The video setting button 324 is a button for displaying a video-related list. The video-related list includes a plurality of items such as camera switching and camera settings. When an item of camera settings is selected, a camera setting screen for setting the currently used camera 222 is displayed. On the camera setting screen, the video obtained by the currently used camera 222 is displayed.
[0024] Referring to FIG. 2 again, the communication unit 24 communicates with other clients 11 via the communication network 19 and the server 12. The communication unit 24 transmits information related to the remote conference received from the control unit 21 to other clients 11. For example, the communication unit 24 transmits video data obtained by the camera 222 and audio data obtained by the microphone 223 to other clients 11. The communication unit 24 receives information related to the remote conference from other clients 11 and sends the received information to the control unit 21. For example, the communication unit 24 receives video data and audio data obtained by other clients 11 from other clients 11.
[0025] The operation information generation unit 25 generates operation information indicating the operations of the client 11 performed by the user during the remote meeting, and stores the generated operation information in the operation information storage unit 291. The operation information is information indicating the operations performed by the user on the client 11 during the remote meeting, specifically, information indicating the operations performed by the user on the user interface provided by the remote meeting application during the remote meeting. Examples of the operations to be recorded include cursor placement on the mute button 321, switching the voice input from off to on, displaying the microphone settings screen, displaying the speaker settings screen, displaying the camera settings screen, transitioning the remote meeting application to the foreground, transitioning the remote meeting application to the background, speaking, and the like. The state in which the remote meeting application is operating in the foreground refers to an active state in which the user can operate the remote meeting application. The state in which the remote meeting application is operating in the background refers to a state in which the remote meeting application is operating but the user cannot operate the remote meeting application. The operation information generation unit 25 receives, from the control unit 21, mouse operation information indicating the operations of the mouse 221 performed by the user and screen information indicating the image displayed on the display device 231. The operation information generation unit 25 can detect operations on the user interface from the operation information and the screen information. For example, the operation information generation unit 25 can detect the position of the cursor on the user interface from the operation information and the screen information. For example, the operation information generation unit 25 detects that the cursor has moved onto the mute button 321 and remains on the mute button 321, and generates operation information related to the operation of cursor placement on the mute button 321.
[0026] Figure 4 schematically shows an example of the operation information stored in the operation information storage unit 291. Each operation is managed by one record (entry). In the example shown in Figure 4, each record includes, as data items, an identifier (No.), an operation type, a start time, an end time, a duration, and an operation flag. The identifier indicates information for identifying the operation. For example, the identifier represents the order in which the operations are performed. The operation type indicates the type of the operation. The start time indicates the time when the operation is started. The end time indicates the time when the operation ends. The duration indicates the length of time the operation is performed. The operation flag indicates whether the operation is in progress. The operation flag "-" indicates that the operation has ended, and the operation flag "○" indicates that the operation is in progress.
[0027] Referring again to Figure 2, the speech desire degree calculation unit 26 calculates the user's speech desire degree based on the operation information stored in the operation information storage unit 291. In the present embodiment, a value in the range from 0 to 1 is taken, and the speech desire degree is defined such that the higher the degree to which the user desires to speak, the larger the value.
[0028] In the present embodiment, the speech desire degree is calculated based on a rule. The rule storage unit 292 stores predetermined speech desire estimation rules. The speech desire degree calculation unit 26 refers to the speech desire estimation rules stored in the rule storage unit 292 in order to calculate the user's speech desire degree. The speech desire estimation rules include information specifying the types of operations estimated to be speech desires. Examples of operations estimated to be speech desires include placing the cursor on the mute button, switching the voice input from off to on, displaying the microphone setting screen, displaying the camera setting screen, and shifting the remote conference application to the foreground.
[0029] Generally, when the user speaks from a state where the voice input and / or video input is off in a remote conference, the user often performs the following actions. (1) The user places the cursor on the mute button so that the voice input can be switched from off to on immediately after the current speaker's speech ends, and waits for the current speaker's speech to end. (2) The user clicks the mute button to switch the voice input from off to on, and then waits for the current speaker to finish speaking. (3) The user displays the microphone settings screen and checks the volume of the microphone. (4) The user displays the camera settings screen and checks the video shown on the camera. (5) The user resumes the remote conferencing application to the foreground.
[0030] The actions frequently performed before speaking (pre-speaking actions) as described above are adopted as operations estimated to be the speaking desire. Hereinafter, the operations estimated to be the speaking desire are also referred to as target operations. Placing the cursor on the mute button, displaying the microphone settings screen, and displaying the camera settings screen are continuous target operations, and switching the voice input from off to on and shifting the remote conferencing application to the foreground are instantaneous target operations. The speaking desire degree calculation unit 26 estimates that the user is in a speaking desire state when an operation matching the target operation occurs after the user's previous speech (when the user has not spoken yet, at the time of joining the remote conference or at the start of the remote conference).
[0031] The speaking desire degree calculation unit 26 calculates a score indicating the possibility that each operation performed by the user after the previous speech is a pre-speaking action, and calculates the speaking desire degree based on the calculated scores. The speaking desire estimation rule may include a reference time set for each continuous target operation. The reference time for each target operation is used to calculate the score of the operation. As an example, the reference time for placing the cursor on the mute button is set to 5 seconds, the reference time for displaying the microphone settings screen is set to 5 seconds, and the reference time for displaying the camera settings screen is set to 10 seconds.
[0032] When the operation is a continuous target operation such as placing the cursor on the mute button, the speech desire degree calculation unit 26 calculates the score of the operation from the duration of the operation and the reference time related to the target operation. For example, when the duration of the operation is equal to or longer than the reference time related to the target operation, the speech desire degree calculation unit 26 determines the score of the operation as 1. When the duration of the operation is less than the reference time related to the target operation, the speech desire degree calculation unit 26 calculates the score of the operation based on the difference or ratio between the duration of the operation and the reference time related to the target operation. Let the duration of the operation be D, the reference time related to the target operation that matches the operation be R, and the score of the operation be S. Then S = D / R. In this example, when the duration D is 2 seconds and the reference time R is 5 seconds, the score S is 0.4. Note that the score may be calculated by a function other than a linear function. For example, S = (D / R) 2 may also be used. In this example, when the duration D is 2 seconds and the reference time R is 5 seconds, the score S is 0.16.
[0033] For example, the user may click the mute button 321 immediately after moving the cursor to the mute button 321 to turn on voice input. When the user clicks the mute button 321 to turn on voice input, the speech desire degree calculation unit 26 may determine the score of the operation of placing the cursor on the mute button as 1 regardless of the duration of placing the cursor on the mute button.
[0034] When the operation is an instantaneous target operation such as the transition of the remote conference application to the foreground, the speech desire degree calculation unit 26 determines the score of the operation as 1.
[0035] When the operation is not any of the target operations, the speech desire degree calculation unit 26 determines the score of the operation as 0.
[0036] When there is an idle period in the operation room for a certain period of time or longer, the speech desire degree calculation unit 26 may consider that one operation (operation type "no operation") has occurred during that period and determine the score of that operation to be 0. The speech desire estimation rule may include information indicating the above-mentioned certain period of time.
[0037] The speech desire degree calculation unit 26 uses the average of the scores calculated for each operation as the speech desire degree. Alternatively, the speech desire degree calculation unit 26 may obtain the weighted average of the scores calculated for each operation as the speech desire degree. As an example, the weight for an operation that occurred from 30 seconds before the current time to the current time is 1, the weight for an operation that occurred from 60 seconds before the current time to 30 seconds before the current time is 0.9, the weight for an operation that occurred from 90 seconds before the current time to 60 seconds before the current time is 0.8, and so on. In other examples, the weight for the operation that the user is currently performing is 1, the weight for the previous operation is 0.9, the weight for the operation two operations before is 0.8, and so on.
[0038] The control unit 21 transmits user information based on the user's speech desire degree to other clients 11 via the communication unit 24. For example, the control unit 21 drives the communication unit 24 to transmit the user information to other clients 11. The user information may include the speech desire degree of the user itself. Alternatively, the user information may include information for notifying that the user has a speech desire. For example, when the speech desire degree calculated by the speech desire degree calculation unit 26 exceeds a predetermined threshold value, the control unit 21 notifies other clients 11 that the user has a speech desire.
[0039] The control unit 21 receives user information based on the degree of speaking desire of other users from other clients 11 via the communication unit 24. The control unit 21 applies the received user information to the user interface. In an example where the user information includes the degree of speaking desire, the control unit 21 may display the degree of speaking desire of each user in association with the video of each user. Alternatively, the control unit 21 may emphasize the video of the user whose degree of speaking desire exceeds a predetermined threshold. For example, the control unit 21 may surround the video of the user whose degree of speaking desire exceeds a predetermined threshold with a red frame, or may apply a mark to the video of the user whose degree of speaking desire exceeds a predetermined threshold.
[0040] FIG. 5 schematically shows an example of the hardware configuration of the client 11. As shown in FIG. 5, the client 11 includes a computer 50 as hardware components, in addition to the mouse 221, camera 222, microphone 223, display device 231, and speaker 232 shown in FIG. 2.
[0041] The computer 50 includes a CPU (Central Processing Unit) 51, a RAM (Random Access Memory) 52, a program memory 53, a storage device 54, an input / output interface 55, and a communication interface 56. The CPU 51 is communicably connected to the RAM 52, the program memory 53, the storage device 54, the input / output interface 55, and the communication interface 56.
[0042] The CPU 51 is an example of a processor. As the processor, other general-purpose circuits may be used, or dedicated circuits such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field-Programmable Gate Array) may be used.
[0043] The RAM 52 includes a volatile memory such as SDRAM (Synchronous Dynamic Random Access Memory). The RAM 52 is used by the CPU 51 as a working memory. The program memory 53 stores programs executed by the CPU 51, such as a remote conferencing application including a speech desire estimation program. The programs include computer-executable instructions. For example, a ROM (Read Only Memory) is used as the program memory 53. A partial area of the storage device 54 may be used as the program memory 53. The CPU 51 expands the programs stored in the program memory 53 into the RAM 52 and interprets and executes the programs. When the remote conferencing application is executed by the CPU 51, the CPU 51 is caused to perform a series of processes described with respect to the processing unit 27. In other words, the CPU 51 functions as the control unit 21, the operation information generation unit 25, and the speech desire degree calculation unit 26 according to the remote conferencing application. Note that the speech desire estimation program may be provided as a program separate from the remote conferencing application. When the speech desire estimation program is executed by the CPU 51, the CPU 51 is caused to perform a series of processes related to speech desire estimation.
[0044] The programs may be provided to the computer 50 in a state stored in a computer-readable recording medium. In this case, the computer 50 includes a drive for reading data from the recording medium and acquires the programs from the recording medium. Examples of the recording medium include magnetic disks, optical disks (such as CD-ROM, CD-R, DVD-ROM, DVD-R), magneto-optical disks (such as MO), and semiconductor memories. Also, the programs may be distributed through a network. Specifically, the programs may be stored in a server on the network and the computer 50 may download the programs from the server.
[0045] The storage device 54 includes a non-volatile memory such as a HDD (Hard Disk Drive) or an SSD (Solid State Drive). The storage device 54 stores data. The storage device 54 functions as a storage unit 29, specifically, an operation information storage unit 291 and a rule storage unit 292.
[0046] The input / output interface 55 is an interface for communicating with peripheral devices. The mouse 221, the camera 222, the microphone 223, the display device 231, and the speaker 232 are connected to the computer 50 by the input / output interface 55. In an example where the computer 50 is a notebook PC, the camera 222, the microphone 223, the display device 231, and the speaker 232 may be built into the computer 50.
[0047] The communication interface 56 is an interface for communicating with external devices (for example, the server 12 and other clients 11 shown in FIG. 1) connected to the communication network 19. The communication interface 56 includes a wired module and / or a wireless module. The communication interface 56 functions as a communication unit 24.
[0048] [Operation] FIG. 6 schematically shows a method for estimating the desire to speak executed by the client 11 shown in FIG. 2. Here, it is assumed that another user is speaking at the current time.
[0049] In step S61 of FIG. 6, the operation information generation unit 25 generates operation information indicating the operations performed by the user on the client 11 during the remote meeting, and stores the generated operation information in the operation information storage unit 291. Specifically, the operation information generation unit 25 generates operation information indicating the user's operations on the user interface provided by the conference application.
[0050] In step S62, the speech desire degree calculation unit 26 calculates the user's speech desire degree based on the operation information. For example, the speech desire degree calculation unit 26 identifies the operation performed by the user on the client 11 after the previous speech by the user during the remote conference from the operation information stored in the operation information storage unit 291, calculates a score for each operation, and calculates the speech desire degree from the calculated scores. When the operation is any of the target operations, the speech desire degree calculation unit 26 calculates the score of the operation based on the duration D of the operation and the reference time R for the target operation. When the duration D of the operation is greater than or equal to the reference time R for the target operation, the speech desire degree calculation unit 26 determines the score to be 1. When the duration D of the operation is less than the reference time R for the target operation type, the value obtained by dividing the duration D of the operation by the reference time R for the target operation type is used as the score of the operation. When the operation is not any of the target operations, the speech desire degree calculation unit 26 determines the score of the operation to be zero. When there is a certain time gap between operations, the speech desire degree calculation unit 26 regards the operation that does not correspond to the target operation as having been performed, and determines the score of the operation to be zero. Subsequently, the speech desire degree calculation unit 26 obtains the user's speech desire degree by averaging the scores calculated for each detected operation.
[0051] In step S63, the control unit 21 transmits user information including the user's speech desire degree obtained in step S62 to other clients 11 via the communication unit 24.
[0052] The process shown in step S61 may be executed periodically, for example, at intervals of 1 second, during the remote conference. The processes shown in steps S62 and S63 may be executed periodically, for example, at intervals of 1 second, during the remote conference and during the period when the user is not speaking.
[0053] With reference to the operation information shown in FIG. 4, the calculation of the speech desire degree will be described. Here, it is assumed that the reference time for cursor placement on the mute button is set to 5 seconds, the reference time for microphone setting screen display is set to 5 seconds, and the reference time for camera setting screen display is set to 10 seconds.
[0054] From 14:28:22 to 14:30:21 when the speech ended, the user did not perform any operations, and the speech desire degree was zero. At 14:29:22, since no operation occurred for 60 seconds, the speech desire degree calculation unit 26 determines that one operation has occurred and determines the score of the operation to be 0. The speech desire degree remains zero.
[0055] At 14:30:21, the user opens the microphone settings screen. At 14:30:22, the score for displaying the microphone settings screen is 0.2 (= 1 / 5), and the speech desire degree is 0.1 (= (0 + 0.2) / 2). The speech desire degree S becomes 0.2 at 14:30:23, 0.3 at 14:30:24, 0.4 at 14:30:25, and 0.5 from 14:30:26 to 14:30:27.
[0056] At 14:30:27, the user closes the microphone settings screen and opens the camera settings screen. At 14:30:27, the score for displaying the camera settings screen is 0.1 (= 1 / 10), and the speech desire degree is 0.37 (≈ (0 + 1 + 0.1) / 3). The speech desire degree becomes 0.4 at 14:30:27, 0.43 at 14:30:28,..., 0.63 at 14:30:36, and 0.67 from 14:30:37 to 14:31:05. At 14:30:42, the user closes the camera settings screen and does not perform any operations from 14:30:42 to 14:31:05.
[0057] At 14:31:05, the user operates the mouse 221 to move the cursor to the mute button 321. At 14:31:06, the score for placing the cursor on the mute button is 0.2 (= 1 / 5), and the speech desire degree is 0.55 (≈ (0 + 1 + 1 + 0.2) / 4). The speech desire degree becomes 0.6 at 14:31:07, 0.65 at 14:31:08, 0.7 at 14:31:09, and 0.75 from 14:31:10 to 14:31:13. At 14:31:13, the user clicks the mute button 321 to start speaking.
[0058] [Effect] In this embodiment, each of the clients 11 used for remote conferences via the communication network 19 generates operation information indicating an operation performed by the user on the client 11 during the remote conference, calculates the degree of the user's speaking desire based on the operation information, and transmits the calculated degree of the speaking desire to other clients 11. Operation information indicating an operation performed by the user on the client 11 is used to calculate the degree of the speaking desire. According to this configuration, it is possible to estimate the user's speaking desire without using voice and video information. Further, the calculated degree of the speaking desire is notified to other clients 11. According to this configuration, it is possible to display the degree of the speaking desire of other users on each client 11. As a result, the users of each client 11 can determine whether other users desire to speak, and can avoid speaking collisions.
[0059] The client 11 identifies an operation performed by the user on the client 11 after the previous speech by the user during the remote conference from the operation information, calculates a score indicating the possibility that the operation is a pre-speech action for each identified operation, and calculates the degree of the speaking desire from the calculated scores. According to this configuration, it is possible to evaluate whether the user has performed a pre-speech action, and it is possible to appropriately estimate the user's speaking desire.
[0060] When the operation is a continuous target operation, the client 11 may calculate the score of the operation based on a comparison between the duration of the operation and a reference time related to the target operation. According to this configuration, it is possible to calculate the score according to the duration of the operation.
[0061] The continuous target operation may include at least one of cursor placement on a mute button for switching voice input on and off, display of a microphone setting screen for setting a microphone, and display of a camera setting screen for setting a camera. These are typical examples of pre-speech actions, and thus it is possible to appropriately estimate the user's speaking desire.
[0062] <Second Embodiment> In the above-described first embodiment, the degree of speaking desire is calculated based on rules. In the second embodiment, the degree of speaking desire is calculated using a speaking desire estimation model obtained by machine learning. In the second embodiment, the description of the same components and processes as in the first embodiment will be omitted as appropriate.
[0063] [Configuration] FIG. 7 schematically shows a client 71 according to the second embodiment. The conference system according to the second embodiment is the same as that shown in FIG. 1, and the client 71 shown in FIG. 7 is used as an alternative to the client 11 shown in FIG. 1. In FIG. 7, the same reference numerals are given to the same components as those shown in FIG. 2, and the description thereof will be omitted as appropriate.
[0064] As shown in FIG. 7, the client 71 includes a control unit 21, an input unit 22, an output unit 23, a communication unit 24, an operation information generation unit 25, a speaking desire degree calculation unit 76, a learning unit 78, and a storage unit 79. The storage unit 79 includes an operation information storage unit 291 and a model storage unit 792. The control unit 21, the operation information generation unit 25, the speaking desire degree calculation unit 76, and the learning unit 78 are collectively referred to as a processing unit 77. The control unit 21, the communication unit 24, the operation information generation unit 25, the speaking desire degree calculation unit 76, the learning unit 78, the operation information storage unit 291, and the model storage unit 792 correspond to a speaking desire estimation device according to the second embodiment.
[0065] The learning unit 78 generates a speaking desire estimation model configured to receive, as an input, operation information indicating at least one operation on the client 71 by machine learning and output a numerical value representing the degree of speaking desire. The learning unit 78 learns the speaking desire estimation model using the operation information stored in the operation information storage unit 291 as learning data. The speaking desire estimation model may be a neural network, and learning is a process of determining parameters (weights and biases) constituting the neural network.
[0066] The learning unit 78 generates operation information leading to speech and operation information not leading to speech from the operation information stored in the operation information storage unit 291. For example, the learning unit 78 obtains the operation information in a predetermined period (for example, 60 seconds) immediately before each speech as the operation information leading to speech. Specifically, the learning unit 78 obtains the operation information from the time 60 seconds before the start time of each speech to the start time of the speech as the operation information leading to speech. The learning unit 78 obtains the operation information in a previous predetermined period (for example, 60 seconds) as the operation information not leading to speech. Specifically, the learning unit 78 obtains the operation information from the time 120 seconds before the start time of each speech to the time 60 seconds before the start time of the speech, the operation information from the time 180 seconds before the start time of each speech to the time 120 seconds before the start time of the speech, etc. as the operation information not leading to speech.
[0067] The learning unit 78 uses the operation information leading to speech and the operation information not leading to speech as inputs to the speech desire estimation model to perform machine learning of the speech desire estimation model. The model storage unit 792 stores the speech desire estimation model generated by the learning unit 78.
[0068] The speech desire degree calculation unit 76 calculates the degree of the user's speech desire based on the operation information stored in the operation information storage unit 291 using the speech desire estimation model. For example, the speech desire degree calculation unit 76 extracts the operation information in a predetermined period (for example, 60 seconds) from the operation information stored in the operation information storage unit 291. Specifically, the speech desire degree calculation unit 76 extracts the operation information indicating the operations performed by the user on the client 71 from the time 60 seconds before the current time to the current time after the previous speech by the user during the remote meeting from the operation information stored in the operation information storage unit 291. The speech desire degree calculation unit 76 inputs the extracted operation information into the speech desire estimation model and obtains the numerical value output from the speech desire estimation model as the speech desire degree.
[0069] When the range of the value output from the speech desire estimation model is not in the range from 0 to 1, the speech desire degree calculation unit 76 may perform normalization so that the value output from the speech desire estimation model is in the range from 0 to 1.
[0070] Note that until a certain amount of operation information is accumulated, the learning of the speech desire estimation model cannot be performed. For this reason, until a certain amount of operation information is accumulated, the speech desire degree calculation unit 76 may use a pre-prepared speech desire estimation model (a speech desire estimation model preset in the remote conference application). Alternatively, the speech desire degree calculation unit 76 may calculate the speech desire degree in the same manner as described in the first embodiment.
[0071] The client 71 can have the same hardware configuration as that shown in FIG. 5. When the remote conference application including the speech desire estimation program according to the present embodiment is executed by the CPU, the CPU is caused to perform a series of processes described with respect to the processing unit 77. In other words, the CPU functions as the control unit 21, the communication unit 24, the operation information generation unit 25, the speech desire degree calculation unit 76, and the learning unit 78 according to the remote conference application.
[0072] [Operation] A learning method executed by the client 71 will be described.
[0073] The operation information generation unit 25 generates operation information indicating an operation performed by the user on the client 71 during the remote conference, and stores the generated operation information in the operation information storage unit 291.
[0074] The learning unit 78 generates a plurality of samples including a plurality of first samples as operation information leading to speech and a plurality of second samples as operation information not leading to speech from the operation information stored in the operation information storage unit 291. Correct answer data is assigned to each sample. For example, when the output layer of the speech desire estimation model includes two nodes, a vector (1, 0) may be assigned as the correct answer data to each first sample, and a vector (0, 1) may be assigned as the correct answer data to each second sample.
[0075] The learning unit 78 randomly selects at least one sample from among the samples, for example. The learning unit 78 inputs each sample into the speech desire estimation model and obtains output data from the speech desire estimation model. The learning unit 78 updates the parameters of the speech desire estimation model so that the output data approaches the correct answer data. For example, cross-entropy error may be used as the objective function, and gradient descent method may be used as the optimization algorithm.
[0076] The learning unit 78 repeatedly executes the process from sample selection to parameter update. As a result, a speech desire estimation model adapted to the user using the client 71 is generated.
[0077] Next, a speech desire estimation method executed by the client 71 will be described. Here, it is assumed that the learning of the speech desire estimation model has been completed. Further, it is assumed that another user is speaking at the current time.
[0078] The operation information generation unit 25 generates operation information indicating the operations performed by the user on the client 71 during the remote conference, and stores the generated operation information in the operation information storage unit 291.
[0079] The speech desire degree calculation unit 76 calculates the speech desire degree of the user based on the operation information stored in the operation information storage unit 291 by using the speech desire estimation model stored in the model storage unit 792. For example, the speech desire degree calculation unit 26 extracts the operation information from the time 60 seconds before the current time to the current time from the operation information stored in the operation information storage unit 291, inputs the extracted operation information into the speech desire estimation model, and obtains the value output from the speech desire estimation model as the speech desire degree.
[0080] The control unit 21 transmits user information including the speech desire degree of the user calculated by the speech desire degree calculation unit 26 to other clients 11 via the communication unit 24.
[0081] [Effect] This embodiment can obtain the same effects as those described in the first embodiment. In this embodiment, the speech desire degree is calculated using the speech desire estimation model obtained by machine learning. According to this configuration, it can be expected that the speech desire of the user can be estimated more appropriately.
[0082] The client 71 learns the speech desire estimation model by using the operation information stored in the operation information storage unit 291 as learning data. According to this configuration, it becomes possible to obtain a speech desire estimation model adapted to the user, and it becomes possible to estimate the speech desire of the user more appropriately.
[0083] [Modification Example] In the above-described embodiment, the remote conference is implemented based on the client-server model. In other embodiments, the conference system may not include a server, and the remote conference may be performed peer-to-peer between clients.
[0084] Note that the present invention is not limited to the above-described embodiments, and various modifications can be made without departing from the gist thereof at the implementation stage. Also, the respective embodiments may be implemented in appropriate combination, and in that case, the combined effects can be obtained. Further, the above-described embodiments include various inventions, and various inventions can be extracted by combinations selected from the plurality of disclosed components. For example, even if some components are deleted from all the components shown in the embodiments, if the problem can be solved and the effects can be obtained, the configuration from which these components are deleted can be extracted as an invention.
Explanation of Reference Numerals
[0085] 10 … Conference system 11 … Client 12 … Server 19 … Communication network 21 … Control unit 22 … Input unit 221… Mouse 222… Camera 223… Microphone 23 … Output unit 231… Display device 232… Speaker 24 … Communication unit 25 … Operation information generation unit 26 … Calculation unit 27 … Processing unit 29 … Storage unit 291… Operation information storage unit 292… Rule storage unit 30 … User interface 31 … Video area 32 … Control bar 321… Mute button 322… Audio setting button 323… Video button 324… Video setting button 50 … Computer 51 … CPU 52 … RAM 53 … Program memory 54 … Storage device 55... Input / Output Interface 56... Communication Interface 71... Client 76... Calculation Unit 77... Processing Unit 78... Learning Unit 79... Memory Unit 792... Model Memory Unit
Claims
1. A speaking desire estimation device provided in a first conference device among a plurality of conference devices used for a remote conference via a communication network, an operation information generation unit that generates operation information indicating an operation performed by a user on the first conference device during the remote conference, the operation information including information representing the type of the operation and a duration that is the length of time the operation was performed, a speaking desire degree calculation unit that calculates a speaking desire degree indicating the degree to which the user desires to speak based on the information representing the type and duration included in the generated operation information, and a communication unit that transmits information based on the calculated speaking desire degree to a second conference device among the plurality of conference devices. The speaking desire estimation device comprising the above.
2. The speaking desire degree calculation unit identifies an operation performed by the user on the first conference device after the previous speech by the user during the remote conference from the generated operation information, calculates a score indicating the possibility that the identified operation is a pre-speech action for each identified operation, and calculates the speaking desire degree from the calculated scores. The speaking desire estimation device according to Claim 1.
3. When the identified operation matches a predetermined operation, the speaking desire degree calculation unit calculates the score of the identified operation based on a comparison between the duration of the identified operation and a reference time set for the predetermined operation. The speaking desire estimation device according to Claim 2.
4. The predetermined operation includes at least one of cursor placement on a mute button for switching voice input on and off, display of a microphone setting screen for setting a microphone, and display of a camera setting screen for setting a camera. The speaking desire estimation device according to Claim 3.
5. The device further includes a speaking desire estimation model configured to receive, as an input, operation information indicating at least one operation and output a numerical value representing the speaking desire degree. The speaking desire degree calculation unit extracts operation information indicating an operation performed by the user on the first conference device after the previous speech by the user during the remote conference from the generated operation information, inputs the extracted operation information into the speaking desire estimation model, and obtains the numerical value output from the speaking desire estimation model as the speaking desire degree. The speaking desire estimation device according to any one of Claims 1 to 4.
6. The speech desire estimation device according to claim 5, further comprising a learning unit that learns the speech desire estimation model using the generated operation information.
7. A speech desire estimation method executed by a first conference device among a plurality of conference devices used for a remote conference via a communication network, comprising: generating operation information indicating an operation performed by a user on the first conference device during the remote conference, the operation information including information representing a type of the operation and a duration of time during which the operation was performed; calculating a speech desire degree indicating a degree to which the user desires to speak based on the information representing the type and the duration included in the generated operation information; transmitting information based on the calculated speech desire degree to a second conference device among the plurality of conference devices; and a speech desire estimation method comprising the steps.
8. A program for causing a computer to function as the speech desire estimation device according to any one of claims 1 to 6.
Citation Information
Patent Citations
Conference device, conference method, and conference program
JP2012244285A
Conference device, conference method and conference program
JP2013183183A
Web conference system, information processing method, and program
JP2017111643A