Communication system, information processing device, information processing method, and program

The communication system addresses the inability to analyze and notify users of group call results by generating and analyzing fragmentary voice data within the system, thereby enhancing user interaction and awareness.

JP2025091714APending Publication Date: 2025-06-19BONX INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023207134
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-07
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing communication systems cannot analyze and notify users of the results of group calls, limiting user awareness and interaction within group communication settings.

Method used

A communication system that includes an information processing apparatus and multiple communication terminals, where each terminal generates and transmits fragmentary voice data of human voice portions, and the apparatus analyzes this data to output analysis results, enabling user notification of group call analysis.

Benefits of technology

Enables users to be informed of group call analysis results, enhancing user interaction and awareness within group communication environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025091714000001_ABST
    Figure 2025091714000001_ABST
Patent Text Reader

Abstract

To provide a communication system, an information processing device, an information processing method, and a program that can inform a user of analysis results of a group call.SOLUTION: A communication system 1000 includes a cloud 1300 and a plurality of mobile communication terminals 1200A and 1200B, and a group call is performed between the plurality of mobile communication terminals 1200A and 1200B. The cloud 1300 analyzes the group call based on multiple pieces of fragmentary voice data of a portion determined to be a human voice exchanged between the plurality of mobile communication terminals 1200A and 1200B in the group call, and outputs analysis results.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a communication system, an information processing apparatus, an information processing method, and a program.

Background Art

[0002] Conventionally, there has been known a communication system that registers a plurality of mobile communication terminals as a group and performs multi-party calls within the group by exchanging voice data obtained by encoding a user's voice among the plurality of mobile communication terminals within the group.

[0003] In addition, when making a call using a mobile communication network, in addition to making a call using a microphone and a speaker provided in the mobile communication terminal, a headset used by connecting to the mobile communication terminal by a short-range communication method such as Bluetooth (registered trademark) is also used (see, for example, Patent Document 1). By using a headset, even when the user is not holding the mobile communication terminal by hand, the mobile communication terminal can pick up the user's voice and convey the voice sent from another user's mobile communication terminal to the user.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the invention described in Patent Document 1, it was not possible to collect call data exchanged within the group and analyze the group call. For this reason, it was not possible to notify the user of the analysis result of the group call.

[0006] Therefore, an object of the present invention is to provide a communication system, an information processing apparatus, an information processing method, and a program capable of notifying a user of the analysis result of a group call.

Means for Solving the Problems

[0007] One aspect of the present invention is a communication system including an information processing apparatus and a plurality of communication terminals, in which a group call is performed among the plurality of communication terminals. Each of the plurality of communication terminals generates a plurality of fragmentary voice data of a portion determined to be human voice from the voice data acquired using a microphone and transmits the data to another communication terminal, and reproduces the plurality of fragmentary voice data received from the other communication terminal with a speaker. The information processing apparatus analyzes the group call based on the plurality of fragmentary voice data exchanged among the plurality of communication terminals in the group call and outputs an analysis result.

Effects of the Invention

[0008] According to one aspect of the present invention, it is possible to provide an information processing apparatus, an information processing method, and a program capable of notifying a user of the analysis result of a group call.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Embodiments for Carrying Out the Invention

[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0011] <1. Overall Configuration of the System> FIG. 1 is a schematic configuration diagram of a communication system according to an embodiment of the present invention. The communication system 1000 includes a first headset 1100A, a second headset 1100B, a first mobile communication terminal 1200A, a second mobile communication terminal 1200B, a cloud (information processing device) 1300, a computer 1400, a display unit 1410, and a mobile communication terminal 1500.

[0012] The first headset 1100A is worn on the user's ear and includes a button 1110A and a communication unit 1111A. The button 1110A functions as a manual switch. The communication unit 1111A includes a microphone as an audio input unit and a speaker as an audio output unit. The first headset 1100A includes a chip for wireless connection to the first mobile communication terminal 1200A. The same applies to the second headset 1100B.

[0013] The first mobile communication terminal 1200A is a mobile phone such as a smartphone or a tablet terminal, and is connected to a cloud 1300 that provides services. The first headset 1100A detects audio data indicating the voice uttered by the user using the microphone, and transmits the detected audio data to the first mobile communication terminal 1200A. The first mobile communication terminal 1200A transmits the audio data received from the first headset 1100A to the cloud 1300. Also, the first mobile communication terminal 1200A transmits the audio data received from the cloud 1300 and detected by the second headset 1100B to the first headset 1100A. The first headset 1100A plays back the audio data received from the first mobile communication terminal 1200A using the speaker.

[0014] The second mobile communication terminal 1200B is a mobile phone such as a smartphone or a tablet terminal, and is connected to a cloud 1300 that provides services. The second headset 1100B detects audio data indicating the voice uttered by the user using the microphone, and transmits the detected audio data to the second mobile communication terminal 1200B. The second mobile communication terminal 1200B transmits the audio data received from the second headset 1100B to the cloud 1300. Also, the second mobile communication terminal 1200B transmits the audio data received from the cloud 1300 and detected by the first headset 1100A to the second headset 1100B. The second headset 1100B plays back the audio data received from the second mobile communication terminal 1200B using the speaker.

[0015] The cloud 1300 may include a VoIP (Voice Over Internet Protocol) server for controlling voice communication, and an API (Application Programmable Interface) server for managing the connection with the first mobile communication terminal 1200A and the second mobile communication terminal 1200B and the allocation of the VoIP server. Also, the cloud 1300 may be composed of a single server or a plurality of servers.

[0016] The cloud 1300 collects a plurality of fragmented voice data from the first mobile communication terminal 1200A and the second mobile communication terminal 1200B, synthesizes the collected plurality of fragmented voice data to generate synthesized voice data, and holds the generated synthesized voice data for a predetermined period (for example, six months). The user can obtain the synthesized voice data from the cloud 1300 using the first mobile communication terminal 1200A or the second mobile communication terminal 1200B connected to the cloud 1300. Details of the synthesized voice data will be described later.

[0017] The computer 1400 is a desktop computer, but is not limited thereto. For example, the computer 1400 may be a notebook computer, or may be a tablet or a smartphone. The computer 1400 is connected to a display unit 1410. The display unit 1410 is a display device such as a liquid crystal display device. The display unit 1410 may be built in the computer 1400.

[0018] Computer 1400 is used by a user with administrative privileges. A user with administrative privileges is a user who can perform various settings of the communication system 1000 (for example, granting various permissions to users, changing accounts, and inviting users, etc.). A user with administrative privileges is, for example, a tenant administrator and a manager. The tenant administrator has the authority to manage the entire tenant and can perform user registration within the tenant, deletion of users, etc. A tenant is a contracting entity that concludes a system usage contract. The tenant administrator identifies users using, for example, email addresses, etc. A manager is a user who has the authority to create group call rooms and register terminals, etc. within a tenant. Similar to the tenant administrator, the manager also identifies users using email addresses, etc.

[0019] Users without administrative privileges are, for example, general users and shared users. A general user is an ordinary user who participates in group calls. The tenant administrator and the manager identify general users using email addresses, etc. On the other hand, a shared user is a user who participates in group calls, but the tenant administrator and the manager do not identify shared users using email addresses, etc.

[0020] The mobile communication terminal 1500 is used by a user with administrative privileges (such as a tenant administrator and a manager). The mobile communication terminal 1500 is a mobile phone such as a smartphone or a tablet terminal, etc., and is connected to the cloud 1300 that provides services.

[0021] A user with administrative privileges (such as a tenant administrator and a manager) can perform various settings of the communication system 1000 by using the computer 1400 or the mobile communication terminal 1500.

[0022] Hereinafter, when distinguishing between the first headset 1100A and the second headset 1100B is not necessary, it is simply denoted as headset 1100. Also, when distinguishing between the first mobile communication terminal 1200A and the second mobile communication terminal 1200B is not necessary, it is simply denoted as mobile communication terminal 1200.

[0023] FIG. 2 is a block diagram showing the configuration of a mobile communication terminal according to an embodiment of the present invention. The mobile communication terminal 1200 includes a group call management unit 1201, a group call control unit 1202, a noise estimation unit 1203, a speaking candidate determination unit 1204, a speakability determination unit 1205, a voice data transmission unit 1206, a reproduced voice data transmission unit 1207, a communication unit 1208, a short-range wireless communication unit 1209, a recorded data storage unit 1210, a voice data generation unit 1211, a display unit 1212, and a reproduction unit 1213.

[0024] A program for executing each function of the mobile communication terminal 1200 is stored in the memory within the mobile communication terminal 1200. The CPU (Central Processing Unit) within the mobile communication terminal 1200 can execute the above-described functions by reading out and executing the program from the memory.

[0025] The group call management unit 1201 exchanges information related to the management of group calls with the cloud 1300 via the communication unit 1208, and manages the start and end of group calls. Specifically, the group call management unit 1201 transmits various requests such as a group call start request, a client addition request, and a group call end request to the cloud 1300. Also, the group call management unit 1201 manages the group call by issuing a command to the group call control unit 1202, which will be described later, in response to the response from the cloud 1300 to the various requests.

[0026] Based on the instructions from the group call management unit 1201, the group call control unit 1202 controls the transmission and reception of voice data between the mobile communication terminals 1200 owned by other users participating in the group call and the transmission and reception of voice data between the headset 1100. In addition, the group call control unit 1202 performs voice detection and data quality control of the voice data related to the user's speech received from the headset 1100, which are carried out by the noise estimation unit 1203, the speech candidate determination unit 1204, and the speech activity determination unit 1205 described later.

[0027] The noise estimation unit 1203 estimates the average ambient noise from the voice data related to the user's speech received from the headset 1100. The voice data related to the user's speech received from the headset 1100 includes the user's speech and ambient noise. As a method for noise estimation by the noise estimation unit 1203, known methods such as minimum mean square error (MMSE) estimation, maximum likelihood method, and maximum a posteriori probability estimation may be used. For example, the noise estimation unit 1203 may sequentially update the power spectrum of the ambient noise based on the probability estimation of the presence of voice for each sample frame according to the MMSE criterion, and estimate the ambient noise, which is the noise, from the voice data using the power spectrum of the ambient noise.

[0028] Based on the estimation result of the ambient noise, which is the noise by the noise estimation unit 1203, the speech candidate determination unit 1204 determines a sound different from the average ambient noise in the voice data as a speech candidate. Specifically, the speech candidate determination unit 1204 compares the long-term spectrum variation in units of several frames with the power spectrum of the ambient noise estimated by the noise estimation unit 1203, and determines that the non-stationary part of the voice data is the voice data due to the user's speech.

[0029] The speechability determination unit 1205 determines a portion of the voice data that is estimated to be sudden environmental noise other than a human voice, for the portion determined by the speech candidate determination unit 1204 to be voice data due to speech by the user. Specifically, the speechability determination unit 1205 determines whether it is voice data based on the voice emitted from a human throat or the like by performing an estimation of the content ratio of spectral periodic components on the portion determined by the speech candidate determination unit 1204 to be voice data due to speech by the user. Further, the speechability determination unit 1205 estimates the degree of echo from the voice waveform to evaluate the distance from the speaker and whether it is a direct wave, and determines whether it is voice data based on the voice emitted by the speaker.

[0030] The voice data transmission unit 1206 encodes the voice data in the range excluding the portion determined by the speechability determination unit 1205 to be sudden environmental noise from the range determined by the speech candidate determination unit 1204 as a speech candidate, and transmits it to the cloud 1300. When encoding the voice data, the voice data transmission unit 1206 encodes the voice data with the encoding method and communication quality determined by the group call control unit 1202 based on a command from the cloud 1300.

[0031] The reproduced voice data transmission unit 1207 transmits the voice data received from the cloud 1300 via the communication unit 1208 and decoded to the headset 1100 via the short-range wireless communication unit 1209.

[0032] The communication unit 1208 controls communication via the mobile network. The communication unit 1208 is realized using a communication interface for a general mobile communication network or the like. The short-range wireless communication unit 1209 controls short-range wireless communication such as Bluetooth (registered trademark). The short-range wireless communication unit 1209 is realized using a general short-range wireless communication interface.

[0033] The recording data storage unit 1210 temporarily stores, as recording data, the voice data (voice data before synthesis) acquired by the headset 1100 that can communicate with the mobile communication terminal 1200. Based on the recording data stored in the recording data storage unit 1210, the voice data generation unit 1211 generates fragmentary voice data indicating the voice during the period when the user using the headset 1100 spoke. Although details will be described later, the voice data generation unit 1211 attaches the user ID, the start time of the speech, and the end time of the speech to the generated fragmentary voice data as metadata. Note that details of the fragmentary voice data generated by the voice data generation unit 1211 will be described later. The display unit 1212 is, for example, a touch panel display. The playback unit 1213 is, for example, a speaker that plays back voice data.

[0034] FIG. 3 is a diagram showing a schematic functional configuration of a headset in an embodiment of the present invention. The headset 1100 includes a voice detection unit 1101, a speech emphasis unit 1102, a playback control unit 1103, and a short-range wireless communication unit 1104.

[0035] The program for executing each function of the headset 1100 is stored in the memory within the headset 1100. The CPU within the headset 1100 can execute each of the above functions by reading out and executing the program from the memory.

[0036] The voice detection unit 1101 detects the speech of the user wearing the headset 1100 and converts it into voice data. The voice detection unit 1101 is composed of a microphone, an A / D conversion circuit, an encoder for voice data, etc. provided in the headset 1100. It is desirable to provide at least two microphones as the microphones constituting the voice detection unit 1101.

[0037] The speech emphasis unit 1102 emphasizes the speech of the user wearing the headset 1100 from among the voice data detected and converted by the voice detection unit 1101. The speech emphasis unit 1102 relatively emphasizes the speech of the user with respect to the ambient sound by using, for example, a known beamforming algorithm or the like. By the process performed by the speech emphasis unit 1102, the ambient sound included in the voice data is relatively suppressed with respect to the speech of the user. For this reason, it is possible to improve the sound quality, improve the performance of the subsequent signal processing, and reduce the computational load. The voice data converted by the voice emphasis unit 1102 is transmitted to the mobile communication terminal 1200 via the short-range wireless communication unit 1104.

[0038] The playback control unit 1103 plays back the voice data received from the mobile communication terminal 1200 via the short-range wireless communication unit 1104. The playback control unit 1103 is composed of, for example, a decoder for voice data, a D / A conversion circuit, a speaker, etc. provided in the headset 1100. When playing back the voice in the speech section in the voice data received from the mobile communication terminal 1200, the playback control unit 1103 plays back the voice data in a form that is easy for the user to hear based on the ambient sound detected by the microphone provided in the headset 1100. For example, the playback control unit 1103 may cancel the ambient sound by executing noise cancellation processing based on the ambient sound estimated by the voice detection unit 1101 to make the playback sound easier to hear, or may execute processing to increase the playback volume in conjunction with the magnitude of the ambient sound to relatively make the playback sound easier to hear.

[0039] The headset 1100 and the mobile communication terminal 1200 of the present embodiment can perform clear speech playback while reducing the size of the voice data transmitted through the communication path by performing multi-faceted voice data processing that correlates various estimation processes related to speech and ambient sound. Thereby, it is possible to save power consumption in the headset 1100 and the mobile communication terminal 1200 and significantly improve the UX (User Experience) of the call.

[0040] Figure 4 is a block diagram showing the configuration of the cloud in an embodiment of the present invention. Cloud 1300 is an information processing apparatus that provides voice data of conversations conducted using headset 1100. Cloud 1300 includes a communication unit 1301, a voice data synthesis unit 1302, a voice data storage unit 1303, an analysis unit 1304, and an output unit 1305.

[0041] Programs for executing the respective functions of cloud 1300 are stored in a memory within cloud 1300. The CPU within cloud 1300 can execute the above-described respective functions by reading the programs from the memory and executing them.

[0042] Communication unit 1301 communicates with first mobile communication terminal 1200A, second mobile communication terminal 1200B, computer 1400, and mobile communication terminal 1500. For example, communication unit 1301 transmits and receives voice data between first mobile communication terminal 1200A and second mobile communication terminal 1200B. Voice data synthesis unit 1302 generates synthesized voice data by synthesizing a plurality of fragmented voice data received from first mobile communication terminal 1200A and second mobile communication terminal 1200B. Details of the synthesized voice data generated by voice data synthesis unit 1302 will be described later. Voice data storage unit 1303 stores the synthesized voice data generated by voice data synthesis unit 1302.

[0043] Analysis unit 1304 analyzes a group call conducted among a plurality of users based on the voice data received by communication unit 1301. Output unit 1305 outputs the analysis result of the group call by analysis unit 1304. Details of the processing contents of analysis unit 1304 and output unit 1305 will be described later.

[0044] Hereinafter, the speech detection function and the voice reproduction control function of communication system 1000 having the above configuration will be described using a sequence chart showing the operation flow.

[0045] <2. Speech Detection Function> FIG. 5 is a sequence chart showing the flow of processes executed on a headset and a mobile communication terminal related to a speech detection function according to an embodiment of the present invention. ● [Step SA01] The voice detection unit 1101 detects the user's speech including ambient sound as voice and converts it into voice data. ● [Step SA02] The speech enhancement unit 1102 relatively enhances the user's speech sound included in the voice data converted in Step SA01 with respect to the ambient sound. ● [Step SA03] The short-range wireless communication unit 1104 transmits the voice data in which the user's speech sound is enhanced in Step SA02 to the first mobile communication terminal 1200A.

[0046] ● [Step SA04] The noise estimation unit 1203 analyzes the voice data received from the first headset 1100A and estimates the ambient sound which is the noise included in the voice data. ● [Step SA05] The speech candidate determination unit 1204 determines, based on the estimation result of the ambient sound which is the noise by the noise estimation unit 1203 in Step SA04, a sound different from the average ambient sound in the voice data as a speech candidate. ● [Step SA06] The speechability determination unit 1205 determines the user's speech part from the part of the voice data determined by the speech candidate determination unit 1204 in Step SA05 as a speech candidate by the user. Specifically, the speechability determination unit 1205 determines the voice data in the range excluding the part determined as a speech from a sudden ambient sound or a speech emitted from a position far from the microphone of the first headset 1100A in the range determined as a speech candidate in Step SA05 as the user's speech part. ● [Step SA07] The group call control unit 1202 encodes the voice data in the range determined as the user's speech part by the speechability determination unit 1205 in Step SA06 with the encoding method and communication quality determined in the interaction with the cloud 1300, and transmits the encoded voice data to the cloud 1300.

[0047] FIG. 6 is a diagram showing an image of conversion until voice data transmitted from the detected voice is generated according to the sequence chart of FIG. 5 in the embodiment of the present invention. As shown in FIG. 6, in the communication system 1000 of the present embodiment, since only the portion necessary for reproducing the utterance is extracted from the detected voice, the voice data encoded and transmitted to the cloud 1300 can be made smaller in size compared to the voice data transmitted in a normal communication system.

[0048] <3. Voice playback control function> FIG. 7 is a sequence chart showing the flow of processing executed on the headset and the mobile communication terminal related to the voice playback control function in the embodiment of the present invention. ● [Step SB01] The group call control unit 1202 receives voice data of only the utterance part from the cloud 1300. Further, the group call control unit 1202 decodes the received voice data by the encoding method determined in the communication with the cloud 1300. ● [Step SB02] The reproduced voice data transmission unit 1207 transmits the voice data decoded in step SB02 to the second headset 1100B.

[0049] ● [Step SB03] The voice detection unit 1101 detects environmental sound as voice and converts it into voice data. ● [Step SB04] The playback control unit 1103 plays back the voice data received from the second mobile communication terminal 1200B while performing processing to make the playback sound easier to hear with respect to the environmental sound detected in step SB03 in the utterance section of the voice data. In the present embodiment, the second headset 1100B is configured to cancel the environmental sound and play back the voice data received from the second mobile communication terminal 1200B, but the present invention is not limited to this. For example, the second headset 1100B may play back the voice data received from the second mobile communication terminal 1200B as it is without canceling the environmental sound.

[0050] <4. Connection between headset and mobile communication terminal> FIG. 8 is a diagram showing a connection status display screen for displaying the connection status between a headset (earphone) and a mobile communication terminal in an embodiment of the present invention. The connection status display screen shown in FIG. 8 is displayed on the display unit 1212 of the mobile communication terminal 1200. On the display unit 1212, identification information for identifying the headset 1100 connected to the mobile communication terminal 1200 (in FIG. 8, "xxxxxx") is displayed. The mobile communication terminal 1200 and the headset 1100 are connected using short-range wireless communication such as Bluetooth (registered trademark), and various control data is transmitted in addition to audio data.

[0051] <5. Login> FIG. 9 is a diagram showing a login screen displayed when logging in to the communication system in an embodiment of the present invention. A user can log in to the communication system 1000 by inputting login information (tenant ID, email address, and password) into the login screen shown in FIG. 9. The tenant ID is a symbol for identifying a tenant and is represented by N digits, letters, etc.

[0052] The communication system 1000 of this embodiment is a cloud service assuming business use. For this reason, a tenant selection key 1214 and a shared user login key 1215 are displayed on the login screen. When the user selects (taps) the tenant selection key 1214, a tenant change screen (FIG. 10) described later is displayed on the display unit 1212. On the other hand, when the user selects (taps) the shared user login key 1215, the user can log in as a shared user without inputting login information. Note that the shared user may be authenticated using code information (for example, a QR code (registered trademark)) provided to the shared user by a user having administrator authority. Details about the shared user will be described later.

[0053] FIG. 10 is a diagram showing a tenant change screen displayed when changing a tenant in an embodiment of the present invention. On the tenant change screen, a tenant list 1216 and a new tenant addition key 1218 are displayed. A check mark 1217 is displayed for tenant A, which is selected by default. Also, the user can reselect tenants B to E that have been selected in the past from the tenant list 1216.

[0054] Also, the user can add a new tenant by selecting (tapping) the new tenant addition key 1218. When a tenant is selected by the user, the login screen shown in FIG. 9 is displayed on the display unit 1212, and the user can log in by entering login information (ID, email address, password). A user who has logged in to a tenant can participate in a room. A room is a unit for managing calls for each group. Rooms are created and deleted by users with administrator privileges.

[0055] <6. Participation in a Room and Creation of a New Room> FIG. 11 is a diagram showing a room participation screen displayed when a user participates in a room in an embodiment of the present invention. On the room participation screen shown in FIG. 11, an invitation notice 1219, a room key input area 1220, a room participation key 1221, a room participation history 1222, and a new room creation key 1223 are displayed.

[0056] The invitation notice 1219 is displayed as new arrival information when an invitation to a room arrives from another user. An invitation is a notice for adding a user to a room. The user can directly participate in the invited room by selecting (tapping) the invitation notice 1219.

[0057] The room key input area 1220 is an area where a user inputs a room key for identifying a room. The room key is a unique key that uniquely determines a call connection destination. The user can participate in a room by inputting the room key into the room key input area 1220 and selecting (tapping) the room participation key 1221.

[0058] The room participation history 1222 is a list of rooms that the user has participated in the past. The user can participate in the selected room by selecting (tapping) the room displayed in the room participation history 1222.

[0059] When the user logs in using a user ID that has the authority to create a new room, the room creation key 1223 is displayed at the bottom of the room participation screen. When the user selects (taps) the room creation key 1223, the room creation screen (Figure 12) is displayed on the display unit 1212.

[0060] Figure 12 is a diagram showing the room creation screen displayed when the user creates a new room in the embodiment of the present invention. The room creation screen shown in Figure 12 displays a room name 1224, a room key 1225, a room URL 1226, and a member invitation key 1227.

[0061] The room name 1224 is automatically determined but may be changed by the user. The room key 1225 is authentication information required to participate in the newly created room and is represented by, for example, numbers. The room URL 1226 is information for specifying the location of the newly created room on the Internet. The member invitation key 1227 is a key for displaying a screen for selecting a user to invite to the room from the member list.

[0062] A user who has newly created a room can invite other users to the room using the room key or room URL, and can also select other users to invite to the room from the member list. The user who has newly created the room may notify other users to be invited to the room of the room key or room URL by email or orally. On the other hand, when a user selects another user from the member list, a push notification is sent to the mobile communication terminal 1200 owned by the selected other user, and an invitation notification 1219 shown in FIG. 11 is displayed on the display unit 1212. Note that the rooms that have been invited are displayed in the room participation history 1222 shown in FIG. 11. Therefore, even if a user invited to a room misses the invitation notification 1219, the user can join the room from the room participation history 1222.

[0063] <7. Conversation Function in Room> FIG. 13 is a diagram showing a call screen displayed when a user joins a room in an embodiment of the present invention. On the call screen shown in FIG. 13, an end call key 1228, a room member key 1229, a call key 1230, a recording key 1231, and a push / hands-free switch key 1232 are displayed.

[0064] When the user taps the end call key 1228, the call in the room ends. Note that instead of tapping the end call key 1228, the call may end in response to swiping the end call key 1228. The room member key 1229 displays the user name and the number of members participating in the room (in the example of FIG. 13, six members). When the user selects (taps) the room member key 1229, a room member screen (FIG. 14) described later is displayed on the display unit 1212.

[0065] When the user selects (taps) the call key 1230, the on / off state of the call key 1230 is toggled. When the call key 1230 is on, the mobile communication terminal 1200 transmits the voice data acquired from the headset 1100 to the cloud 1300. On the other hand, when the call key 1230 is off, the mobile communication terminal 1200 does not transmit the voice data acquired from the headset 1100 to the cloud 1300. Thus, by turning off the call key 1230, the user's speech can be prevented from being heard by the communication partner. Note that instead of the call key 1230, the on / off state may be toggled using the button 1110 of the headset 1100.

[0066] When the user selects (taps) the recording key 1231, recording of the voice data acquired based on the user's speech is started. The recorded voice data is stored as recording data in the recording data storage unit 1210. Also, by the user swiping the push / hands-free switching key 1232 in the vertical direction, the push and hands-free modes can be switched.

[0067] FIG. 14 is a diagram showing a room member screen displayed when the user in the embodiment of the present invention checks the room members. The room member screen displays a list of active members participating in the room and a list of members who have left the room.

[0068] Note that a user having administrator authority can select one of the users displayed on the room member screen and call or remove the selected user from the room. For example, when a user having administrator authority selects (taps) the selection key 1233 of user B from the room member screen, a pop-up window 1234 (FIG. 15) is displayed on the display unit 1212.

[0069] FIG. 15 is a diagram showing a pop-up window displayed on the room member screen in an embodiment of the present invention. A user with administrator authority can select to call user B or remove user B from the room from the pop-up window 1234. On the other hand, when a user with administrator authority selects (taps) the cancel key 1235, the pop-up window 1234 can be closed.

[0070] <8. User Settings Screen> FIG. 16 is a diagram showing a user settings screen in an embodiment of the present invention. FIG. 16 shows, as an example, the user settings screen of a user (manager A) with administrator authority. On the user settings screen, an account settings icon 1236, a talk settings icon 1237, and a management screen selection key 1238 are displayed.

[0071] When a user with administrator authority selects (taps) the account settings icon 1236, the account settings screen is displayed on the display unit 1212. A user with administrator authority can change the password and nickname on the account settings screen.

[0072] Also, as described above, in addition to users with administrator authority, there are general users and shared users. Similar to users with administrator authority, general users can change the password and nickname on the account settings screen. On the other hand, shared users can change the nickname on the account settings screen but cannot change the password. Note that the setting contents on the above-described account settings screen are only examples and are not limited thereto. For example, on the account settings screen, settings other than the password and nickname may be enabled. Also, shared users may be able to change the password on the account settings screen.

[0073] When the user selects (taps) the talk setting icon 1237, the talk setting screen is displayed on the display unit 1212. The user can change the volume, noise suppression level, etc. on the talk setting screen. Note that the setting contents on the above-described talk setting screen are only examples and are not limited thereto. For example, on the talk setting screen, settings other than the volume, noise, and suppression level may be enabled.

[0074] The management screen selection key 1238 is displayed on the user setting screen of a user having administrator authority, but is not displayed on the user setting screen of a general user and the user setting screen of a shared user. When a user having administrator authority selects (taps) the management screen selection key 1238, the management screen is displayed on the display unit 1212. The user can manage information about users within the tenant and recording data for each room on the management screen.

[0075] Note that the management screen may be displayed on the computer 1400 or the mobile communication terminal 1500 according to the user's operation. Thereby, the user can display the management screen on the computer 1400 or the mobile communication terminal 1500.

[0076] <9. Synthesis Process of Audio Data> FIG. 17 is a sequence diagram showing the synthesis process of audio data in an embodiment of the present invention. In the present embodiment, both user A who owns the first mobile communication terminal 1200A and user B who owns the second mobile communication terminal 1200B have administrator authority and can instruct the start of recording in the room. For simplicity of explanation, FIG. 17 shows an example where the number of users participating in the room is two, but the number of users may be three or more. Further, FIG. 17 shows an example where recording is started in response to the recording key 1231 (FIG. 13) being selected (tapped), but it may be set to record all conversations by default.

[0077] In FIG. 17, when user A who holds the first mobile communication terminal 1200A participates in a room where user B who holds the second mobile communication terminal 1200B is present, the first mobile communication terminal 1200A transmits a participation notice to the cloud 1300 (step S1). When the cloud 1300 receives the participation notice from the first mobile communication terminal 1200A, it transmits new member information (e.g., user ID) about user A to the second mobile communication terminal 1200B (step S2). As a result, user A can participate in the group call and can talk to user B within the room.

[0078] When user A selects (taps) the recording key 1231 (FIG. 13), the first mobile communication terminal 1200A transmits a recording start notice to the second mobile communication terminal 1200B (step S3). As a result, the first mobile communication terminal 1200A and the second mobile communication terminal 1200B start recording voice data.

[0079] The first headset 1100A transmits the voice data acquired using the microphone to the first mobile communication terminal 1200A. The voice data generation unit 1211 of the first mobile communication terminal 1200A generates a plurality of fragmented voice data of portions determined to be human voices by the utterance determination unit 1205 from the voice data received from the first headset 1100A. This fragmented voice data is the voice data of the portion where user A spoke. The first mobile communication terminal 1200A transmits this fragmented voice data to the second headset 1100B via the cloud 1300 and the second mobile communication terminal 1200B. As a result, the speaker of the second headset 1100B can reproduce only the voice data of the portion where user A spoke.

[0080] In addition, the recording data storage unit 1210 of the first mobile communication terminal 1200A stores a plurality of fragmented voice data generated by the voice data generation unit 1211 of the first mobile communication terminal 1200A as recording data. Note that user ID of user A, start time of speech, and end time of speech are attached as metadata to the fragmented voice data stored in the recording data storage unit 1210 of the first mobile communication terminal 1200A. If voice data before recording start instruction is stored in the recording data storage unit 1210 of the first mobile communication terminal 1200A, this voice data may be used as recording data.

[0081] On the other hand, the second headset 1100B transmits the voice data acquired using the microphone to the second mobile communication terminal 1200B. The voice data generation unit 1211 of the second mobile communication terminal 1200B generates a plurality of fragmented voice data of the parts determined to be human voice by the speech determination unit 1205 from the voice data received from the second headset 1100B. This fragmented voice data is the voice data of the parts where user B spoke. The second mobile communication terminal 1200B transmits this fragmented voice data to the first headset 1100A via the cloud 1300 and the first mobile communication terminal 1200A. As a result, the speaker of the first headset 1100A can reproduce only the voice data of the parts where user B spoke.

[0082] In addition, the recording data storage unit 1210 of the second mobile communication terminal 1200B stores a plurality of fragmented voice data generated by the voice data generation unit 1211 of the second mobile communication terminal 1200B as recording data. Note that user ID of user B, start time of speech, and end time of speech are attached as metadata to the fragmented voice data stored in the recording data storage unit 1210 of the second mobile communication terminal 1200B. If voice data before recording start instruction is stored in the recording data storage unit 1210 of the second mobile communication terminal 1200B, this voice data may be used as recording data.

[0083] The first mobile communication terminal 1200A transmits the voice data of user B received from the second mobile communication terminal 1200B via the cloud 1300 to the first headset 1100A, but does not store this voice data of user B in the recording data storage unit 1210 of the first mobile communication terminal 1200A. Also, the second mobile communication terminal 1200B transmits the voice data of user A received from the first mobile communication terminal 1200A via the cloud 1300 to the second headset 1100B, but does not store this voice data of user A in the recording data storage unit 1210 of the second mobile communication terminal 1200B.

[0084] User A exits the room by tapping the call end key 1228 (Fig. 13). At this time, the first mobile communication terminal 1200A transmits an exit notification to the cloud 1300 (step S4). When user A exits the room, the first mobile communication terminal 1200A ends the recording of voice data. After a predetermined time has elapsed since user A exited the room, the first mobile communication terminal 1200A reads out the plurality of fragmented voice data stored in the recording data storage unit 1210 and transmits it to the cloud 1300 (step S5).

[0085] On the other hand, user B exits the room by tapping the call end key 1228 (Fig. 13). At this time, the second mobile communication terminal 1200B transmits an exit notification to the cloud 1300 (step S6). When user B exits the room, the second mobile communication terminal 1200B ends the recording of voice data. After a predetermined time has elapsed since user B exited the room, the second mobile communication terminal 1200B reads out the plurality of fragmented voice data stored in the recording data storage unit 1210 and transmits it to the cloud 1300 (step S7).

[0086] Thereafter, the voice data synthesizing unit 1302 of the cloud 1300 generates synthesized voice data by synthesizing a plurality of fragmented voice data received from the first mobile communication terminal 1200A and the second mobile communication terminal 1200B. The voice data storage unit 1303 of the cloud 1300 stores the synthesized voice data generated by the voice data synthesizing unit 1302.

[0087] In the sequence diagram shown in FIG. 17, the mobile communication terminal 1200 buffers voice data, sends the voice data to the cloud 1300 when the user leaves the room, and the cloud 1300 synthesizes the voice data. However, the present invention is not limited to this. For example, the cloud 1300 may start storing the voice data received from the mobile communication terminal 1200 in response to receiving a recording start notification from the mobile communication terminal 1200. That is, instead of the mobile communication terminal 1200 buffering the voice data, the cloud 1300 may always accumulate the voice data while the conversation is in progress. Further, when all users leave the room (or stop recording), the voice data synthesizing unit 1302 of the cloud 1300 may generate synthesized voice data using the time stamps (speech start time and speech end time) attached to the voice data.

[0088] FIG. 18 is a diagram showing a plurality of fragmented voice data received by the cloud in the embodiment of the present invention. In FIG. 18, it is assumed that the user ID of user A is "ID001" and the user ID of user B is "ID002". In FIG. 18, voice data A, voice data C, and voice data E are a plurality of fragmented voice data (voice data of user A) received from the first mobile communication terminal 1200A, and voice data B, voice data D, and voice data F are a plurality of fragmented voice data (voice data of user B) received from the second mobile communication terminal 1200B.

[0089] Voice data A, voice data C, and voice data E are respectively attached with metadata D1, metadata D3, and metadata D5. Also, voice data B, voice data D, and voice data F are respectively attached with metadata D2, metadata D4, and metadata D6. Metadata D1 to D6 include a user ID, a speech start time, and a speech end time. The voice data synthesis unit 1302 of the cloud 1300 generates synthesized voice data by synthesizing a plurality of fragmented voice data based on the speech start time and the speech end time included in the metadata. The synthesized voice data generated by the voice data synthesis unit 1302 is recording data in which the conversation content in the room is recorded. The voice data storage unit 1303 stores the synthesized voice data generated by the voice data synthesis unit 1302 as recording data for a predetermined period (for example, six months).

[0090] According to the present embodiment, by extracting only the voice data of the parts where the user spoke, the size of the voice data can be reduced, and the data communication volume in the communication system 1000 can be reduced. Also, instead of synthesizing all the speech contents, the cloud 1300 may generate a dictation text by extracting only the voice of a specific user. Thereby, only the voice associated with each user can be extracted as the original data of the conversation data.

[0091] Note that when generating the synthesized voice data (conversation file), the voice data synthesis unit 1302 of the cloud 1300 may adjust the speech timing of the voice data of each user so that the voices of each user do not overlap. Thereby, it is possible to make it easier to hear the conversation of each user. Also, the cloud 1300 may use the voice data for each user for speech recognition. Furthermore, the cloud 1300 may search for the synthesized voice data using the user ID as a search key. Thereby, it is possible to efficiently identify the users participating in the conversation in the room.

[0092] <10. Analysis Processing of Group Calls> When the group call ends, the cloud 1300 performs analysis processing of the group call. As described above, the communication unit 1301 of the cloud 1300 receives a plurality of fragmented voice data of portions determined to be human voices from each of the plurality of mobile communication terminals 1200. The analysis unit 1304 of the cloud 1300 analyzes the group call based on the plurality of fragmented voice data received by the communication unit 1301. The output unit 1305 of the cloud 1300 outputs the analysis result of the group call by the analysis unit 1304. For example, the output unit 1305 may display the analysis result of the group call on a display device provided in the cloud 1300, may print it on paper with a printer, or may transmit it to the mobile communication terminal 1200, the computer 1400, or the mobile communication terminal 1500 for display. Hereinafter, with reference to FIGS. 19 to 24, the analysis result of the group call by the analysis unit 1304 will be described.

[0093] FIG. 19 is a diagram showing the daily transition of the number of users participating in the group call and the daily transition of the total speaking time of all users participating in the group call. In FIG. 19, the horizontal axis represents the date, the vertical axis of the bar graph represents DAU (Daily Active User), and the vertical axis of the line graph represents the total speaking time of all users. DAU indicates the number of users participating in the group call per day. Note that the analysis result shown in FIG. 19 may be output for each room used in the group call.

[0094] Based on the voice data (Figure 18) received by the cloud 1300, the analysis unit 1304 counts the number of users participating in the group call on a daily basis and calculates the total speaking time of all users participating in the group call on a daily basis. For example, the analysis unit 1304 may calculate the number of users participating in the group call on a daily basis by counting the number of user IDs included in the metadata attached to the voice data. Also, the analysis unit 1304 may calculate the total speaking time of all users participating in the group call on a daily basis based on the speaking start time and speaking end time included in the metadata attached to the voice data. At this time, the total speaking time of each user may be added. For example, even if there is a time when multiple users are speaking simultaneously, they may be added separately. Then, as shown in Figure 19, the output unit 1305 outputs, as analysis results, the daily transition of the number of users participating in the group call and the daily transition of the total speaking time of all users participating in the group call.

[0095] For example, when the communication system 1000 of the present embodiment is used by employees of a retail store, the user can visually grasp the periodicity of the workload, such as the fact that the conversation volume among employees increases as the number of customers increases on weekends, that is, Saturdays and Sundays, by checking the analysis results shown in Figure 19. Also, the user can visually grasp the change in the workload during a specific period, such as the increase in the conversation volume among employees during the busy season at the end of the year and the beginning of the year. In this example, the total speaking time from the 21st to the 23rd, which are weekdays in the fourth week, is longer than that of other weeks. Since the DAU does not change significantly, the speaking volume per person has increased. Thus, in a group where periodicity is observed by day of the week, when the speaking volume per person increases by 10% or more, more preferably 25% or more, compared to the same day of the week in other weeks, it is advisable to change the color of the line graph, display it thicker, or highlight it. This can indicate the change point to the analyst and draw attention.

[0096] FIG. 20 is a diagram showing the transition of the total speaking time of all users participating in a group call by time zone. In FIG. 20, the horizontal axis represents the time zone, and the vertical axis represents the total speaking time of all users. Note that the analysis results shown in FIG. 20 may be output for each room used in the group call.

[0097] Based on the voice data (FIG. 18) received by the cloud 1300, the analysis unit 1304 calculates the total speaking time of all users participating in the group call for each time zone. For example, the analysis unit 1304 may calculate the total speaking time of all users participating in the group call for each time zone based on the speaking start time and speaking end time included in the metadata attached to the voice data. Then, as shown in FIG. 20, the output unit 1305 outputs the transition of the total speaking time of all users participating in the group call by time zone as an analysis result.

[0098] For example, when the communication system 1000 of the present embodiment is used by store employees, the user can visually grasp the time zone when the conversation volume among employees peaks in a day and the degree of congestion of the store for each time zone by checking the analysis results shown in FIG. 20. The example of FIG. 20 is data for a day when the speaking volume per person increased in FIG. 19. Here, for example, in the time zones of 17:00 and 18:00, when the total speaking time increased by 20% or more, more preferably 30% or more, compared with the same day of other weeks, it is advisable to change the hatching, change the color, or highlight the bar graph. It is possible to alert the analyst that an event may have been held or an accident may have occurred during this time zone. If the sales data for this day has increased or decreased compared with other days and there is no clear reason such as a sales event, there may be something unexpected happening, and it may be worthwhile to investigate the cause regardless of whether the sales increase or decrease has occurred. Therefore, it is even better to link the sales data to the system of this embodiment.

[0099] FIG. 21 is a diagram showing the relationship between the total stay time and the total speaking time of rooms used in group calls for each user. In FIG. 21, the horizontal axis represents the total room stay time for each individual, and the vertical axis represents the total speaking time for each individual. Also, each of the plurality of points plotted on the graph of FIG. 21 corresponds to each of the plurality of users who participated in the group call. Note that the analysis results shown in FIG. 21 may be output for each room used in the group call.

[0100] The analysis unit 1304 derives the relationship between the total stay time for each individual and the total speaking time for each individual of the rooms used in the group call. For example, the analysis unit 1304 may calculate the time from when a participation notification (step S1 in FIG. 17) is received until an exit notification (step S4 in FIG. 17) is received as the stay time for each individual in the room per day. Also, the analysis unit 1304 may calculate the total of the stay times for each individual in the room for one month as the total stay time for each individual shown on the horizontal axis of FIG. 21. Further, the analysis unit 1304 may calculate the speaking time for each individual per day based on the speaking start time and the speaking end time included in the metadata attached to the voice data. Also, the analysis unit 1304 may calculate the total of the speaking times for each individual for one month as the total speaking time for each individual shown on the vertical axis of FIG. 21. Thereafter, as shown in FIG. 21, the output unit 1305 outputs the relationship between the total stay time for each individual and the total speaking time for each individual of the rooms used in the group call as an analysis result.

[0101] Figure 21 shows a plurality of regions A1 to A4. The users included in region A1 are those who have a short stay time in the room and a short speaking time. Such users cannot be said to effectively utilize group calls. The users included in region A2 are those who have a long stay time in the room but a short speaking time. Such users may be staying in the room but hardly having conversations with other users. The users included in region A3 are those who have a short stay time in the room but a long speaking time. Such users may leave the room after making many statements in a short time without giving other users a chance to speak. The users included in region A4 are those who have both a long stay time and a long speaking time in the room, and are considered to be ideally using group calls. Therefore, by checking the analysis results shown in Figure 21, the user can visually grasp the usage status of group calls for each of the multiple users who participated in the group call.

[0102] In FIG. 21, each dot represents each user, but usually it is not indicated which dot represents whom. For example, if a user who was normally included in area A2 moves to area A4, since the total speaking time per individual has increased, it is possible that the user is becoming more active in speaking within the team. Thus, when the area in which a user is included changes, for the changed dot, it is advisable to highlight it by enlarging the dot or changing its color. Also, for a dot that the analyst is interested in, when an action such as placing the cursor on it or tapping it in the case of a touch panel is performed, the individual represented by that dot may be displayed on the screen. This can draw the attention of the analyst, who is in a position to monitor employees who are changing more actively, and can also be used as a reference for, for example, selecting the most suitable candidates for promotion during performance evaluations or to higher positions. Conversely, for example, if a user who was normally included in area A4 moves to area A2, although the staying time has not changed, the total speaking time per individual has decreased. If the user returns to area A4 immediately, there is no problem, but if the user continues to stay in area A2, it is possible that there are physical or mental problems, or that there have been negative changes in the interpersonal relationships within the team. Therefore, by highlighting and drawing attention, it is also possible to observe whether care is needed.

[0103] FIG. 22 is a diagram showing the relationship between the average room staying time per person and the average speaking time per person for each room used in group calls. In FIG. 22, the horizontal axis represents the average room staying time per person, and the vertical axis represents the average speaking time per person. Also, each of the circled marks plotted on the graph in FIG. 22 corresponds to a room, and the number attached to the circled mark indicates the number of users who participated in the room. These circled marks are displayed larger as the number of users increases.

[0104] The analysis unit 1304 derives the relationship between the average room stay time per person and the average speaking time per person for each room used in the group call. For example, the analysis unit 1304 may calculate the individual stay time of each room per day by the method described above, and calculate the average value of the individual stay times of all users who participated in the room as the average room stay time per person shown on the horizontal axis of FIG. 22. Further, the analysis unit 1304 may calculate the individual speaking time per day by the method described above, and calculate the average value of the individual speaking times of all users who participated in the room as the average speaking time per person shown on the vertical axis of FIG. 22. Thereafter, as shown in FIG. 22, the output unit 1305 outputs the relationship between the average room stay time per person and the average speaking time per person for each room used in the group call as an analysis result.

[0105] FIG. 22 shows a plurality of regions A5 to A7. The rooms included in region A5 are rooms where the stay time of each user is short and the speaking time is also short. Such rooms cannot be said to be effectively utilized in group calls. The rooms included in region A6 are rooms where the stay time of each user is long but the speaking time is short. In such rooms, although a plurality of users stay for a long time, there is a possibility that they hardly talk to other users. The rooms included in region A7 are rooms where the stay time and the speaking time of each user are long, and it is considered that the group call is ideally utilized. Therefore, by checking the analysis result shown in FIG. 22, the user can visually grasp the usage status of the group call for each of the plurality of rooms in which the group call was made.

[0106] For example, it can be seen that the rooms included in area A6 have a large number of participants in group calls. Since each user tends to find it difficult to talk when the number of participants in a group call is too large, it is effective to disperse users among multiple rooms, such as by separating rooms for each floor of the store. By taking such measures, it becomes possible to bring each room closer to area A7, where group calls are considered to be ideally utilized. Also, for example, when a group that was normally located in area A7 moves to area A6 or area A5, it may be highlighted, such as by displaying it in a different color or showing it with a thick line. When such a change occurs, since there may be some trouble within the group and the average speaking time may have decreased, by highlighting it to draw attention, it is possible to monitor whether care is needed. Furthermore, in combination with the attendance information of each user belonging to the group, if the average speaking time per person of the entire group increases during the working hours of a specific user, it is conceivable that the user is activating the group. This can also serve as a reference for analysts in the position of supervising employees, such as when drawing attention and selecting the user as the most suitable candidate for promotion during performance evaluations or to higher positions. Conversely, if the average speaking time per person of the entire group becomes short during the working hours of a specific user, it is conceivable that the user is isolated within the group, and by drawing attention to analysts in the position of supervising employees, it is also possible to monitor whether care is needed. Therefore, it is even better to link the attendance data of employees to the system of this embodiment.

[0107] FIG. 23 is a diagram showing the results of key influencer analysis for each room used in group calls. Key influencer analysis is a process of analyzing the speaking time of each user and the status of the spread of statements among users. In FIG. 23, the circles marked A to I correspond to users A to I, respectively, and the larger the circle is displayed, the longer the total speaking time of each user is. Also, the lines connecting the users are lines indicating the spread of statements. For example, when user A makes a statement in the room and users B, D, F, G, and I are participating in the room, and furthermore, if the statement of the user continues immediately after the statement of user A, it is presumed that there is a high possibility of conversation, so user A is connected by a line to each of users B, D, F, G, and I. Although FIG. 23 shows only one connection line for simplicity of explanation, the number of connection lines may be increased according to the number of statements, or it may be highlighted by changing the thickness or color of the connection lines according to the number of statements.

[0108] The analysis unit 1304 performs key influencer analysis. For example, the analysis unit 1304 calculates the individual speaking time of each user by the method described above. Further, the analysis unit 1304 determines the users who participated in the group call based on the voice data (Figure 18) received by the cloud 1300. For example, the analysis unit 1304 determines the users who participated in the group call based on the user IDs included in the metadata attached to the voice data, and determines the spread status of the speech among a plurality of users. Thereafter, as shown in Figure 23, the output unit 1305 outputs the result of the key influencer analysis as an analysis result. In the case of the example in Figure 23, it shows the possibility that user A has a large amount of speech, is connected to more users, and is a key influencer who has a lot of communication. Conversely, user H and user I are only connected to user G and user A respectively, are shown as small circles, and have little speech. In such a case, since user H and user I may be isolated at work, by alerting an analyst who is in a position to supervise employees, it is also possible to watch out for whether care is necessary. Especially in the case of user H, the connection with user A, who is presumed to be a key influencer, is extremely low, and although it is connected to user G, who is connected to user A, it is not connected to user A, so it may require more care, so it is better to highlight it more strongly and alert.

[0109] By checking the analysis result shown in Figure 23, the user can grasp the influencers (users who are the core of communication) who become key persons in the group call and the spread status of the speech among a plurality of users.

[0110] FIG. 24 is a diagram showing the results of key interviewer analysis for each room used in a group call. The key interviewer analysis is a process of analyzing the number of reactions of each user and the targets to which the reactions were made. In FIG. 24, the circles A to I correspond to users A to I, respectively, and the larger the circle, the greater the number of reactions of each user. A reaction is a short utterance such as an interjection. Also, the lines connecting the users are lines indicating the propagation of reactions. For example, when user A reacts in the room and users C, D, E, and G are participating in that room, user A is connected by a line to each of users C, D, E, and G.

[0111] The analysis unit 1304 performs key interviewer analysis. For example, the analysis unit 1304 determines short utterances of less than a predetermined time as reactions based on the voice data (FIG. 18) received by the cloud 1300, and calculates the number of reactions of each user. Also, the analysis unit 1304 may determine the users who participated in the group call by the above-described method, and determine the propagation status of reactions. Thereafter, as shown in FIG. 24, the output unit 1305 outputs the results of the key interviewer analysis as analysis results. In the example of FIG. 24, since user D is connected to more users and the circle is large, there is a possibility that user D is an interviewer in this room. Here, more detailed analysis is possible if it is possible to analyze whether the interjection is positive such as "yes" or "so it is", or negative such as "but" or "however". Therefore, it is still better to link the system of this embodiment with a natural language processing system for voice.

[0112] By checking the analysis results shown in FIG. 24, the user can grasp the interviewers (users who become listeners to smooth communication) who are key persons in the group call and the propagation status of reactions among multiple users.

[0113] By continuously checking the analysis results of group calls shown in FIGS. 19 to 24, various problems in communication between users can be discovered. For example, in the analysis results of FIG. 21, even though many users are approaching region A4 where the ideal usage state is shown over time, if there are users who are not improved at all, there may be some problems with such users. According to the present embodiment, such problems can be easily discovered.

[0114] As described above, the communication system 1000 of the present embodiment includes a cloud (information processing device) 1300 and a plurality of mobile communication terminals 1200, and group calls are made between the plurality of mobile communication terminals 1200. Each of the plurality of mobile communication terminals 1200 generates a plurality of fragmentary voice data of portions determined to be human voices from the voice data acquired using a microphone and transmits it to other mobile communication terminals, and reproduces the plurality of fragmentary voice data received from other mobile communication terminals with a speaker. The cloud (information processing device) 1300 analyzes the group call based on the plurality of fragmentary voice data exchanged between the plurality of mobile communication terminals 1200 in the group call and outputs the analysis result. Thereby, the communication system 1000 of the present embodiment can notify the user of the analysis result of the group call.

[0115] In the above embodiment, a part of the functions of the headset 1100 may be provided in the mobile communication terminal 1200. For example, instead of the headset 1100, the mobile communication terminal 1200 may include the speech emphasis unit 1102 shown in FIG. 3. Also, in the above embodiment, a part of the functions of the mobile communication terminal 1200 may be provided in the headset 1100. For example, instead of the mobile communication terminal 1200, the headset 1100 may include all or part of the group call management unit 1201, group call control unit 1202, noise estimation unit 1203, speech candidate determination unit 1204, speech activity determination unit 1205, voice data transmission unit 1206, recording data storage unit 1210, and voice data generation unit 1211 shown in FIG. 2.

[0116] Further, in the above embodiment, part of the functions of the headset 1100 and the mobile communication terminal 1200 may be provided in the cloud 1300. For example, instead of the headset 1100, the cloud 1300 may include the speech emphasis unit 1102 described in FIG. 3. Further, instead of the mobile communication terminal 1200, the cloud 1300 may include all or part of the group call management unit 1201, the group call control unit 1202, the noise estimation unit 1203, the speech candidate determination unit 1204, the speechability determination unit 1205, the voice data transmission unit 1206, the recording data storage unit 1210, and the voice data generation unit 1211 described in FIG. 2.

[0117] As described above, some embodiments of the present invention have been described. However, these embodiments are merely examples and do not limit the technical scope of the present invention. The present invention can take various other embodiments, and various changes such as omissions and substitutions can be made without departing from the gist of the present invention. These embodiments and their modifications are included in the scope and gist of the invention described in this specification and the like, and are also included in the invention described in the claims and its equivalent scope.

[0118] In the conversation analysis of the present invention, it is also possible to record the user's speech, analyze the content of the conversation by combining language analysis processing, or combine sentiment analysis (analysis that quantifies emotions such as joy, anger, sorrow, and happiness). However, even employees during work may chat when they are off guard, and such chats are rather necessary to exert the necessary concentration when needed. In some cases, if the system can collect and analyze the content of conversations related to privacy, continuously wearing the receiver that collects such conversations will be a major stress factor, reducing the morale of individual employees and greatly hindering the smooth communication between employees. Therefore, in this embodiment, language analysis is performed only on the range that can be analyzed from the speech time or on a very short speech that is presumed to be a response, and the analysis is limited to the range that can be analyzed therefrom.

Explanation of Reference Numerals

[0119] 1000 Communication System 1100A First Headset 1100B Second Headset 1101 Voice Detection Unit 1102 Speech Emphasis Unit 1103 Playback Control Unit 1104 Short - Range Wireless Communication Unit 1200A First Portable Communication Terminal 1200B Second Portable Communication Terminal 1201 Group Call Management Unit 1202 Group Call Control Unit 1203 Noise Estimation Unit 1204 Speech Candidate Judgment Unit 1205 Speech Ability Judgment Unit 1206 Voice Data Transmission Unit 1207 Playback Voice Data Transmission Unit 1208 Communication Unit 1209 Short - Range Wireless Communication Unit 1210 Recording Data Storage Unit 1211 Voice Data Generation Unit 1212 Display Unit 1213 Playback Unit 1300 Cloud 1301 Communication Unit 1302 Voice Data Synthesis Unit 1303 Voice Data Storage Unit 1304 Analysis Unit 1305 Output Unit

Claims

1. A communication system comprising an information processing apparatus and a plurality of communication terminals, wherein a group call is performed among the plurality of communication terminals, Each of the plurality of communication terminals generates a plurality of fragmented voice data of a part determined to be human voice from voice data acquired using a microphone and transmits it to other communication terminals, and plays back the plurality of fragmented voice data received from the other communication terminals with a speaker, The information processing apparatus analyzes the group call based on the plurality of fragmented voice data exchanged among the plurality of communication terminals in the group call and outputs an analysis result. Communication system.

2. The information processing apparatus outputs, as the analysis result, the daily transition of the number of users who participated in the group call and the daily transition of the total speaking time of all users who participated in the group call. The communication system according to claim 1.

3. The information processing apparatus outputs, as the analysis result, the transition of the total speaking time of all users who participated in the group call by time zone. The communication system according to claim 1.

4. The information processing apparatus outputs, as the analysis result, the relationship between the total stay time and the total speaking time of the room used in the group call for each user. The communication system according to claim 1.

5. The information processing apparatus outputs, as the analysis result, the relationship between the average room stay time per person and the average speaking time per person for each room used in the group call. The communication system according to claim 1.

6. The information processing device performs key influencer analysis to analyze the speaking time of each user and the situation of the propagation of utterances between users for each room used in the group call, and outputs the result of the key influencer analysis as the analysis result. The communication system according to claim 1.

7. The information processing device performs key interviewer analysis to analyze the number of reactions of each user and the target of the reaction for each room used in the group call, and outputs the result of the key interviewer analysis as the analysis result. The communication system according to claim 1.

8. An information processing device communicably connected to a plurality of communication terminals where a group call is made, a communication unit that receives a plurality of fragmentary voice data of parts determined to be human voices from each of the plurality of communication terminals; an analysis unit that analyzes the group call based on the plurality of fragmentary voice data received by the communication unit; an output unit that outputs the analysis result of the group call by the analysis unit; An information processing device comprising:

9. An information processing device communicably connected to a plurality of communication terminals where a group call is made, receives a plurality of fragmentary voice data of parts determined to be human voices from each of the plurality of communication terminals, analyzes the group call based on the received plurality of fragmentary voice data, and outputs the analysis result of the group call. An information processing method.

10. In an information processing device communicably connected to a plurality of communication terminals where a group call is made, causes each of the plurality of communication terminals to receive a plurality of fragmentary voice data of parts determined to be human voices, Cause the group call to be analyzed based on the received plurality of fragmented voice data, Output the analysis result of the group call, Program.

Citation Information

Patent Citations

  • Method and device for bluetooth pairing

    JP2011182407A