Teleconferencing system, teleconferencing equipment, program, and method for determining the role of speaker in a teleconferencing conference.
Patent Information
- Application Number
- JP2022182179
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-11-14
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2042-11-14
AI Technical Summary
【0009】 本発明では、音声データ毎に、音声データの音声認識結果であるテキストデータおよびそのテキストデータに含まれている所定品質の語句に基づいて判断された参加者の役割が、受信開始時刻とともにこの音声データに紐付けられる。したがって、本発明によれば、会議における発言者それぞれの実際の役割を、会議の進行状況に合わせて把握することができるので、電話会議全体の流れを把握して、電話会議進行上の問題点および改善点等を検討することができる。
Smart Images

Figure 0007916756000001 
Figure 0007916756000002 
Figure 0007916756000003
Abstract
Description
Technical Field
[0001] The present invention relates to a telephone conference system, and particularly to a technology for determining the role of a speaker in a telephone conference.
Background Art
[0002] Patent Document 1 discloses a technology for notifying other participants of a speaker each time a speech is made during a telephone conference. In this technology, each participant presses the speech button of their telephone conference terminal when giving a speech, and the telephone conference terminal whose speech button has been pressed transmits speaker identification information to the telephone conference device. When the telephone conference device receives speaker identification information from any telephone conference terminal, it mixes audio data representing the participant associated with the speaker identification information with the speech audio data transmitted from said telephone conference terminal, and transmits the mixed data to each telephone conference terminal other than said telephone conference terminal.
Prior Art Literature
Patent Literature
[0003]
Patent Document 1
Summary of the Invention
Problem to be Solved by the Invention
[0004] Generally, a conference proceeds smoothly when participants broadly and appropriately fulfill one of the roles of explainer, questioner, or audience member. However, in a conference, the exchange of speeches by participants may hinder the progress of the conference, such as when the explainer gives insufficient explanation, and some questioners repeatedly ask tough questions without being able to understand the explainer's intention. In addition, the role of a participant may change from the initial role according to the progress of the conference; for example, an audience member who feels that the progress of the conference is hindered because the explainer's explanation is unclear may unavoidably act as an intermediary and explain to the questioner on behalf of the explainer.
[0005] Therefore, in order to determine whether a meeting proceeded smoothly without any problems, it is important to understand the actual role of each speaker in the meeting as the meeting progressed. However, the technology described in Patent Document 1 does not take this point into consideration at all.
[0006] This invention has been made in view of the above circumstances, and its purpose is to enable the understanding of the role of each speaker in a meeting as the meeting progresses. [Means for solving the problem]
[0007] To solve the above problems, in the present invention, each time the teleconferencing device receives audio (speech) data from a teleconferencing terminal, it stores this audio data linked to the source teleconferencing terminal and the start time of its reception (speech). Furthermore, it performs speech recognition processing on each of the audio data stored linked to the source teleconferencing terminal and the start time of reception to generate text data. It then performs text analysis processing, including morphological analysis, on the generated text data to extract words corresponding to predetermined parts of speech such as nouns, verbs, and adjectives from this text data, and links the text data and the extracted words to the corresponding audio data. Then, for each audio data stored linked to the source teleconferencing terminal and the start time of reception of the audio data, the teleconferencing device determines the role of the participant who made the speech (explainer, questioner, mediator) based on the extracted words linked to this audio data, and links the determined role to this audio data.
[0008] For example, the telephone conferencing system of the present invention is A telephone conferencing system comprising: a plurality of telephone conferencing terminals; and a telephone conferencing device for each telephone conferencing terminal that mixes audio data received from the plurality of telephone conferencing terminals excluding the said telephone conferencing terminal to generate telephone conferencing data and transmits it to the said telephone conferencing terminal, The aforementioned telephone conferencing device is A voice data storage means that stores each of the voice data received from the plurality of telephone conference terminals, linked to the originating telephone conference terminal and the start time of reception; For each audio data stored in the aforementioned audio data storage means, a speech recognition means performs speech recognition processing on the audio data to generate text data, and associates the generated text data with the audio data. A text analysis means performs text analysis on text data associated with each audio data stored in the audio data storage means, extracts words of a predetermined part of speech from the text data, and associates the extracted words with the audio data. A role determination means that, for each audio data stored in the audio data storage means, determines the role of the participant based on the phrase associated with the audio data, and associates the determined role with the audio data. For each audio data stored in the audio data storage means, an emotion determination means determines the emotion of the participant who spoke the audio data based on the acoustic characteristics of the audio data, including the volume level and speech pitch, and associates the determined emotion of the participant with the audio data. The aforementioned audio data storage means includes a disruption detection means that detects disruptions that occurred during a teleconference based on the emotions of participants associated with a plurality of audio data arranged chronologically in order of reception start time, and associates the detected disruptions with the plurality of audio data. It holds. [Effects of the Invention]
[0009] In this invention, for each audio data, the text data, which is the result of speech recognition of the audio data, and the roles of the participants, determined based on predetermined quality of words contained in that text data, are linked to the audio data along with the reception start time. Therefore, according to this invention, the actual roles of each speaker in a meeting can be grasped in accordance with the progress of the meeting, allowing for an understanding of the overall flow of the teleconference and enabling the consideration of problems and areas for improvement in the progress of the teleconference. [Brief explanation of the drawing]
[0010] [Figure 1] Figure 1 is a schematic diagram of a telephone conferencing system according to one embodiment of the present invention. [Figure 2] Figure 2 is a schematic diagram of the functional configuration of the teleconferencing device 1. [Figure 3] Figure 3 is a schematic diagram showing an example of the registered contents of the voice data storage unit 103. [Figure 4] Figure 4 is a schematic diagram showing an example of the registered contents in the analysis result storage unit 104. [Figure 5]FIG. 5 is a diagram schematically showing an example of registered content in a phrase list storage unit 105. [Figure 6] FIG. 6 is a diagram schematically showing an example of registered content in an obstacle information storage unit 106. [Figure 7] FIG. 7 is a flow diagram for explaining a telephone conference recording process performed by the telephone conference apparatus 1. [Figure 8] FIG. 8 is a flow diagram for explaining a telephone conference analysis process performed by the telephone conference apparatus 1. [Figure 9] FIG. 9 is a flow diagram for explaining a telephone conference analysis process performed by the telephone conference apparatus 1, which is a continuation of FIG. 8. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0011] An embodiment of the present invention will be described below.
[0012] FIG. 1 is a schematic configuration diagram of a telephone conference system according to the present embodiment.
[0013] As illustrated, the telephone conference system according to the present embodiment includes a plurality of telephone conference terminals 2-1 to 2-n (hereinafter, also simply referred to as telephone conference terminals 2), a telephone conference apparatus 1 that accommodates the plurality of telephone conference terminals 2 and provides telephone conference services to these telephone conference terminals 2, and a management terminal 3 that maintains and manages the telephone conference apparatus 1, which are connected to each other via a network 4 such as a WAN (Wide Area Network) or a LAN (Local Area Network).
[0014] For each telephone conference terminal 2, the telephone conference apparatus 1 mixes audio data received from a plurality of other telephone conference terminals 2 to generate telephone conference data, and transmits the generated telephone conference data to the telephone conference terminal 2 (telephone conference service). Furthermore, in a telephone conference held between a plurality of telephone conference terminals 2 via the telephone conference service, the telephone conference apparatus 1 recognizes the role and emotion of a speaker for each utterance, and based on these recognition results, determines whether an obstacle that hinders smooth progress of the conference (delay in the progress of the telephone conference such as complication or confusion) has occurred.
[0015] Next, the teleconference device 1 according to the present embodiment will be described. Note that existing telephone terminals such as key telephones can be used as the teleconference terminals 2. In addition, an existing network terminal such as a PC (Personal Computer) can be used as the management terminal 3. Therefore, detailed descriptions of these are omitted here.
[0016] FIG. 2 is a schematic functional configuration diagram of the teleconference device 1.
[0017] As illustrated, the teleconference device 1 comprises a network interface unit 100, a telephone control unit 101, a teleconference processing unit 102, a voice data storage unit 103, an analysis result storage unit 104, a phrase list storage unit 105, a trouble information storage unit 106, a voice recognition unit 107, a text analysis unit 108, a role determination unit 109, an emotion determination unit 110, a trouble detection unit 111, and a main control unit 112.
[0018] The network interface unit 100 is an interface for connecting to the network 4.
[0019] The telephone control unit 101 establishes a communication path with the teleconference terminals 2 via the network 4 and releases the established communication path in accordance with a call control protocol such as SIP (Session Initiation Protocol).
[0020] The teleconference processing unit 102 provides a teleconference service to a plurality of teleconference terminals 2 for which a communication path with the teleconference device 1 has been established by the telephone control unit 101 (that is, the plurality of teleconference terminals 2 participating in the teleconference). Specifically, for each teleconference terminal 2 participating in the teleconference, teleconference data is generated by mixing voice data received from a plurality of other teleconference terminals 2 excluding said teleconference terminal 2, and the generated teleconference data is transmitted to said teleconference terminal 2.
[0021] Furthermore, each time the teleconferencing processing unit 102 receives audio data from any teleconferencing terminal 2 participating in a teleconferencing conference, it stores this audio data in the audio data storage unit 103, associating it with the start time of reception, the originating teleconferencing terminal 2, and the teleconferencing conference.
[0022] The audio data storage unit 103 stores, for each conference call, the audio data received from the conference call terminals 2 participating in the conference call, linked to the start time of reception and the originating conference call terminal 2.
[0023] Figure 3 is a schematic diagram showing an example of the registered contents of the voice data storage unit 103.
[0024] As shown in the diagram, the audio data storage unit 103 stores a conference call table 1030 for each conference call, in which the audio data of the speeches during that conference call is recorded in chronological order, and this table is linked to the conference call's identification information (conference ID) and the conference start date and time.
[0025] The teleconferencing table 1030 stores a record 1031 of the audio data for each statement made during a teleconferencing session. The audio data record 1031 includes a field 1032 in which the start time of receiving (speaking) the audio data is registered, a field 1033 in which information about the speaker is registered (such as the phone number of the teleconferencing terminal 2 that sent the audio data or the name information of the participant associated with that phone number), and a field 1034 in which the audio data is registered.
[0026] The analysis result storage unit 104 stores the analysis results for the audio data of speeches during a teleconference for each teleconference table 1030 stored in the audio data storage unit 103.
[0027] Figure 4 is a schematic diagram showing an example of the registered contents in the analysis result storage unit 104.
[0028] As shown in the diagram, the analysis result storage unit 104 stores an analysis result table 1040 for each conference call, which records the analysis results of the audio data of the speeches made during that conference call in chronological order. This table is linked to the conference ID and start date and time of the conference call.
[0029] The analysis results table 1040 stores a record 1041 of the analysis results for each statement made during a conference call. The analysis result record 1041 includes a field 1042 in which the start time of receiving (speaking) the audio data is registered, a field 1043 in which speaker information (such as the phone number of the conference call terminal 2 that sent the audio data or the names of participants associated with that phone number) is registered, a field 1044 in which the text data, which is the result of speech recognition of the audio data, is registered, a field 1045 in which a word or phrase of a predetermined part of speech (extracted word) extracted from this text data is registered, a field 1046 in which the speaker's role (explainer, questioner, supplementator, etc.) determined based on the extracted word is registered, a field 1047 in which the speaker's emotion determined based on the volume level, pitch of speech, etc. of the audio data is registered, and a field 1048 in which one of the disruption pattern IDs (see Figure 6) described later is registered if the progress of the conference call is being disrupted.
[0030] The word list memory unit 105 stores a list of words (including nouns, verbs, and adjectives, which are words that fall under a specified part of speech) that may be included in statements made with the aim of fulfilling the role that a speaker plays in a meeting.
[0031] Figure 5 is a schematic diagram illustrating an example of the registered contents of the word list storage unit 105.
[0032] As shown in the diagram, the word list storage unit 105 stores a word list record 1050 for each role that a speaker plays in a meeting. The word list record 1050 has a field 1051 in which the speaker's role is registered, and a field 1052 in which a list of words (words corresponding to predetermined parts of speech, including nouns, verbs, and adjectives) that may be included in a statement intended to fulfill that role is registered.
[0033] The trouble information storage unit 106 stores trouble information for each type of trouble that may be anticipated during a teleconference.
[0034] Figure 6 is a schematic diagram showing an example of the registered contents of the malfunction information storage unit 106.
[0035] As shown in the figure, the trouble information storage unit 106 stores trouble information records 1060 for each trouble expected in a teleconference. The trouble information record 1060 has a field 1061 in which identification information of the trouble occurrence pattern (trouble occurrence pattern ID) is registered, a field 1062 in which the trouble occurrence pattern (the order in which the speakers' roles and emotions during the series of statements that caused the trouble are registered) is registered, and a field 1064 in which the content of the trouble is registered. In addition, the trouble occurrence pattern field 1062 has multiple subfields 1063 in which the roles and emotions of the speakers during the series of statements that caused the trouble are stored in the order in which they were made. In this embodiment, as an example, for three audio data files that are consecutive in time, a subfield 1063-1 is provided in which the role and emotions of the speaker of the first audio data file (first speaker) are registered, a subfield 1063-2 is provided in which the role and emotions of the speaker of the second audio data file (second speaker) are registered, and a subfield 1063-3 is provided in which the role and emotions of the speaker of the third audio data file (third speaker) are registered.
[0036] The speech recognition unit 107 refers to the speech data storage unit 103, reads the speech data from the record 1031 to be analyzed stored in the conference call table 1030 to be analyzed, performs speech recognition processing, and generates text data of the content of the speech represented by this speech data. Then, it registers the generated text data into record 1041 of the analysis result table 1040, which is stored in the analysis result storage unit 104 and is linked to the conference ID and conference start date and time common to the conference call table 1030 to be analyzed (record 1041 in which the reception start time and speaker common to the record 1031 to be analyzed are registered).
[0037] The text analysis unit 108 reads text data from record 1041 stored in the analysis result table 1040 of the analysis result storage unit 104, performs text analysis processing including morphological analysis on this data, and extracts words and phrases (words and phrases corresponding to predetermined parts of speech, including nouns, verbs, and adjectives) that make up this text data. Then, it registers the extracted words and phrases in the analysis result record 1041.
[0038] The role determination unit 109 identifies the record 1041 to be determined for a role from the analysis result table 1040 in the analysis result storage unit 104, and reads the extracted phrases from this record 1041. Then, it searches the phrase list storage unit 105 for the record 1050 that contains the phrase list with the most common phrases with the extracted phrases, and registers the speaker's role registered in this record 1050 in the record 1041 to be determined for a role.
[0039] The emotion judgment unit 110 identifies the record 1031 to be judged from the conference call table 1030 to be analyzed, reads the audio data from this record 1031, and judges the speaker's emotion (calm, excited, intimidated, etc.) based on the acoustic information such as the volume level and pitch of the audio data. Then, it refers to the analysis result storage unit 104 and registers the judged speaker's emotion in record 1041 of the analysis result table 1040, which is linked to the conference ID and conference start date and time common to the conference call table 1030 to be analyzed (analysis result record 1041 which has the same reception start time and speaker registered as record 1031 of the audio data to be judged).
[0040] The trouble detection unit 111 searches the analysis result storage unit 104 for an array of analysis result records 1041 (groups of analysis result records 1041 arranged in chronological order based on the reception start time) that correspond to the role and emotion of the speaker according to the trouble occurrence pattern registered in record 1060, for each record 1060 stored in the trouble information storage unit 1060. If an array of analysis result records 1041 corresponding to any trouble occurrence pattern is detected, the trouble occurrence pattern ID registered in the record 1060 corresponding to the trouble occurrence pattern is linked to this array of analysis result records 1041. Specifically, the corresponding trouble occurrence pattern ID is registered in one of the records 1041 (in this case, the last record 1041) that make up the detected array of analysis result records 1041.
[0041] The main control unit 112 comprehensively controls each of the parts 100 to 111 of the teleconferencing device 1. The main control unit 112 also transmits the registered contents of the voice data storage unit 103 and the analysis result storage unit 104 to the management terminal 3 in accordance with instructions received from the management terminal 3 via the network interface unit 100.
[0042] The functional configuration of the teleconferencing device 1 shown in Figure 2 may be implemented in hardware using integrated logic ICs such as ASICs (Application Specific Integrated Circuits) and FPGAs (Field Programmable Gate Arrays), or it may be implemented in software using a computer such as a DSP (Digital Signal Processor). Alternatively, it may be implemented as a process in a general-purpose computer such as a PC (Personal Computer) equipped with a CPU (Central Processing Unit), memory, auxiliary storage devices such as SSDs (Solid State Drives) and HDDs (Hard Disk Drives), and communication devices such as NICs (Network Interface Cards), by the CPU loading a predetermined program from the auxiliary storage device into memory and executing it.
[0043] Figure 7 is a flowchart illustrating the teleconferencing recording process of the teleconferencing device 1.
[0044] This process begins when the telephone control unit 101 establishes a call path with multiple teleconferencing terminals 2 via the network interface unit 100, and the teleconferencing processing unit 102 starts providing teleconferencing services to these teleconferencing terminals 2.
[0045] First, the teleconferencing processing unit 102 registers a new teleconferencing table 1030 in the audio data storage unit 103 to store the audio data of a series of statements in a newly scheduled teleconferencing session, and associates this teleconferencing table 1030 with the newly issued conference ID and conference start date and time (current date and time) (S200).
[0046] Subsequently, until the telephone control unit 101 terminates the provision of the teleconferencing service by releasing the call paths to all teleconferencing terminals 2 (NO in S203), the teleconferencing processing unit 102 adds a record 1031 to the newly registered teleconferencing table 1030 each time any participant speaks as a speaker during the teleconferencing, that is, each time it receives audio data from a teleconferencing terminal 2 via the telephone control unit 101 (YES in S201). In this record 1031, it registers the start time of reception of the audio data and information about the speaker (such as the number information of the teleconferencing terminal 2 that sent the audio data or the name information of the participant associated with that number information), and also registers the received audio data (S202).
[0047] Figures 8 and 9 are flowcharts illustrating the teleconferencing analysis process of the teleconferencing device 1.
[0048] This process is initiated when a request for conference call analysis is received from the management terminal 3 via the network interface unit 100.
[0049] First, the main control unit 112 checks whether there are any unanalyzed teleconferencing tables 1030 in the audio data storage unit 103 (S300). Specifically, it refers to the audio data storage unit 103 and the analysis result storage unit 104, and searches the audio data storage unit 103 for teleconferencing tables 1030 associated with a conference ID and conference start date and time that are not associated with any of the analysis result tables 1040 in the analysis result storage unit 104. If there are no unanalyzed teleconferencing tables 1030 in the audio data storage unit 103 (NO in S300), this flow is terminated.
[0050] On the other hand, if there are unanalyzed teleconferencing tables 1030 (YES in S300), the main control unit 112 selects one of the unanalyzed teleconferencing tables 1030 as the target for analysis, and registers a new analysis result table 1040 in the analysis result storage unit 104. The analysis result table 1040 is then linked to the conference ID and conference start date and time associated with the teleconferencing table 1030 to be analyzed (S301).
[0051] Next, the main control unit 112 checks whether there are any unanalyzed audio data records 1031 in the conference call table 1030 to be analyzed (S302). If there are no unanalyzed audio data records 1031 and all audio data records 1031 have been analyzed (NO in S302), the process proceeds to S309. On the other hand, if there are unanalyzed audio data records 1031 (YES in S302), the main control unit 112 selects the record 1031 containing the earliest reception start time from among the unanalyzed audio data records 1031 as the record 1031 to be analyzed (S303). Then, it adds a new record 1041 to the newly registered analysis result table 1040 and registers the reception start time and speaker registered in the record 1031 to be analyzed in this new record 1041 (S304).
[0052] Next, the main control unit 112 notifies the speech recognition unit 107 of the conference ID associated with the conference table 1030 to be analyzed, and the reception start time and speaker registered in the record 1031 to be analyzed, and instructs it to perform speech recognition processing on the audio data. In response, the speech recognition unit 107 refers to the audio data storage unit 103 and identifies the record 1031 to be analyzed, which is associated with the reception start time and speaker notified by the main control unit 112, from the conference table 1030 to be analyzed, which is associated with the conference ID notified by the main control unit 112. It then performs speech recognition processing on the audio data registered in this record 1031 to generate text data. Finally, referring to the analysis result storage unit 104, it registers this text data in the record 1041 of the analysis result table 1040, which is associated with the conference ID notified by the main control unit 112, and which contains the reception start time and speaker notified by the main control unit 112 (S305).
[0053] Next, the main control unit 112 notifies the text analysis unit 108 of the conference ID associated with the conference table 1030 to be analyzed, and the reception start time and speaker registered in the record 1031 to be analyzed, and instructs it to perform text analysis processing on the text data. In response, the text analysis unit 108 refers to the analysis result storage unit 104 and identifies the analysis result record 1041 associated with the reception start time and speaker notified by the main control unit 112 from the analysis result table 1040 associated with the conference ID notified by the main control unit 112, and performs text analysis processing, including morphological analysis, on the text data of this record 1041. As a result, it extracts words and phrases corresponding to predetermined parts of speech from the text data and registers the extracted words and phrases in the identified analysis result record 1041 (S306).
[0054] Next, the main control unit 112 notifies the role determination unit 109 of the conference ID associated with the conference table 1030 to be analyzed, and the reception start time and speaker registered in the record 1031 to be analyzed, instructing it to determine the role of the speaker (participant). In response, the role determination unit 109 refers to the analysis result storage unit 104 and identifies the analysis result record 1041 associated with the reception start time and speaker notified by the main control unit 112 as the target for role determination from the analysis result table 1040 associated with the conference ID notified by the main control unit 112. It reads the extracted phrases from this record 1041 and searches the phrase list storage unit 105 for the record 1050 that contains the phrase list with the most common phrases with the extracted phrases. Then, it registers the role of the speaker registered in the record 1050 of the retrieved phrase list in the analysis result record 1041 of the target for role determination (S307).
[0055] Next, the main control unit 112 notifies the emotion judgment unit 110 of the conference ID associated with the conference table 1030 to be analyzed, and the reception start time and speaker registered in the record 1031 to be analyzed, instructing it to make an emotion judgment about the speaker (participant). In response, the emotion judgment unit 110 refers to the audio data storage unit 103 and identifies the record 1031 associated with the reception start time and speaker notified by the main control unit 112 as the target for role judgment from the conference table 1030 to be analyzed, which is associated with the conference ID notified by the main control unit 112. Based on the acoustic information such as the volume level and speech pitch of the audio data registered in this record 1031, it determines the speaker's emotion (calm, excited, intimidated, etc.). Then, from the analysis result table 1040 stored in the analysis result storage unit 104 and linked to the conference ID notified by the main control unit 112, the system identifies the record 1041 that is linked to the same reception start time and speaker notified by the main control unit 112 and the speaker, and registers the judged emotion of the speaker in this record 1041 (S308). After that, the system returns to S302.
[0056] Furthermore, in S309, the main control unit 112 notifies the trouble detection unit 111 of the conference ID associated with the conference table 1030 to be analyzed, instructing it to detect troubles in the conference call. In response, the trouble detection unit 111 selects an unselected record 1060 from the trouble information storage unit 106 (S309), and searches the analysis result storage unit 104 for a sequence of records 1041 (a group of analysis result records 1041 in which the reception start times are arranged consecutively in chronological order) in which the trouble occurrence pattern registered in this record 1060 matches the role and emotion of the speaker and their arrangement (S310). If a sequence of records 1041 in which the role and emotion of the speaker are arranged according to the trouble occurrence pattern is detected (YES in S311), the trouble occurrence pattern ID registered in the selected record 1060 is registered in the last record 1041 of this sequence of records 1041 (S312).
[0057] Next, if there are unselected records 1060 in the problem information storage unit 106 (YES in S313), the problem detection unit 111 returns to S309. If all records 1060 in the problem information storage unit 106 have already been selected (NO in S313), it notifies the main control unit 112 of this fact and returns to S300.
[0058] One embodiment of the present invention has been described above.
[0059] In this embodiment, for each statement made during a conference call, the speaker's role is determined based on predetermined parts of speech contained in the text data, which is the result of speech recognition of the audio data of that statement. This speaker's role is then linked to the audio data of that statement, using the start time of reception of the audio data and its source (speaker) as keys. Therefore, according to this embodiment, the content of each speaker's statements and their actual roles in a conference call can be grasped chronologically in accordance with the progress of the conference call, allowing for an understanding of the overall flow of the conference call and enabling the consideration of problems and areas for improvement in the progress of the conference call.
[0060] Furthermore, in this embodiment, for each statement made during a conference call, the word list storage unit 105 searches for a word list containing the most words of a predetermined part of speech extracted from the text data, which is the speech recognition result of the audio data of that statement. The role of the participant associated with the searched word list is then linked to the audio data of that statement. Therefore, according to this embodiment, the role of only the participants who spoke during the meeting can be efficiently determined.
[0061] Furthermore, in this embodiment, for each statement made during a conference, the speaker's emotion is determined based on the acoustic characteristics, including the volume level and pitch of the audio data of that statement. This emotion is then linked to the audio data of that statement, using the start time of reception and the source of the audio data as keys. Based on the roles and emotions of the speakers, which are arranged chronologically in the order of their statements during the conference call, the occurrence of disruptions that hinder the progress of the conference call is detected. The detected disruptions are then linked to the audio data of the series of statements that caused the disruptions, using the start time of reception and the speaker who was the source of the audio data as keys. Therefore, according to this embodiment, it is possible to understand at what point in the conference call a disruption occurred, making it possible to grasp the overall flow of the conference call with greater accuracy and to efficiently consider problems and areas for improvement in the progress of the conference call.
[0062] Furthermore, in this embodiment, for each disruption pattern stored in the disruption information storage unit 106, the roles and emotions of speakers arranged chronologically according to the disruption pattern are retrieved from the analysis result storage unit 104. If the corresponding sequence of speaker roles and emotions is detected, it is determined that a disruption to the progress of the teleconferencing occurred in the conversation in which the series of audio data (a group of audio data arranged chronologically) linked to these speaker roles and emotions was recorded. The details of the disruption linked to the disruption pattern are then linked to the sequence of audio data of the conversation that caused the disruption to the progress of the teleconferencing, using the start time of reception of the audio data and its source as keys. Therefore, according to this embodiment, it is possible to understand at what point in the audio conference a disruption occurred, along with the details of the disruption, and problems and areas for improvement of the teleconferencing can be examined more efficiently.
[0063] It should be noted that the present invention is not limited to the embodiments described above, and numerous modifications are possible within the scope of its essence.
[0064] For example, in the above embodiment, the conference call table 1030 and the analysis result table 1040 are linked to each other by conference ID and stored in the audio data storage unit 103 and the analysis result storage unit 104, respectively. In addition, the audio data record 1031 stored in the conference call table 1030 and the analysis result record 1041 stored in the analysis result table 1040 are linked to each other by the start time of reception of the audio data and the speaker who is the source of the transmission. However, the present invention is not limited to this. The conference call table 1030 and the analysis result table 1040 may be integrated by combining the audio data record 1031 and the analysis result record 1041. In this case, one of the audio data storage unit 103 and the analysis result storage unit 104 can be omitted.
[0065] Furthermore, in the above embodiment, the word list storage unit 105 and / or the problem information storage unit 106 of the teleconferencing device 1 may be updatable by the management terminal 3. That is, in the teleconferencing device 1, the main control unit 112 updates the registered contents of the word list storage unit 105 and / or the problem information storage unit 106 according to instructions received from the management terminal 3 via the network interface unit 100. Specifically, it receives a word list record 1050 from the management terminal 3 via the network interface unit 100 and adds this record 1050 to the word list storage unit 105. It also adds a problem occurrence pattern record 1060 received from the management terminal 3 via the network interface unit 100 to the problem information storage unit 106. [Explanation of symbols]
[0066] 1: Teleconferencing equipment 2-1~2-n: Teleconferencing terminals 3: Management terminal 4: Network 100: Network Interface Unit 101: Telephone control unit 102: Telephone conference processing unit 103: Audio data storage unit 104: Analysis result storage unit 105: Word list memory unit 106: Problem information memory unit 107: Speech recognition unit 108: Text Analysis Unit 109: Role Judgment Unit 110: Emotion Judgment Unit 111: Fault detection unit 112: Main control unit
Claims
1. A telephone conferencing system comprising: a plurality of telephone conferencing terminals; and a telephone conferencing device for each telephone conferencing terminal that mixes audio data received from the plurality of telephone conferencing terminals excluding the said telephone conferencing terminal to generate telephone conferencing data and transmits it to the said telephone conferencing terminal, The aforementioned telephone conferencing device is A voice data storage means that stores each of the voice data received from the plurality of telephone conference terminals, linked to the originating telephone conference terminal and the start time of reception; For each audio data stored in the aforementioned audio data storage means, a speech recognition means performs speech recognition processing on the audio data to generate text data, and associates the generated text data with the audio data. A text analysis means performs text analysis on text data associated with each audio data stored in the audio data storage means, extracts words of a predetermined part of speech from the text data, and associates the extracted words with the audio data. For each audio data stored in the audio data storage means, a role determination means determines the role of the participant who spoke the audio data based on the words associated with the audio data, and associates the determined role with the audio data. For each audio data stored in the aforementioned audio data storage means, an emotion determination means determines the emotion of the participant who spoke the audio data based on the acoustic characteristics of the audio data, including the volume level and speech pitch, and associates the determined emotion of the participant with the audio data. The audio data storage means includes a disruption detection means that detects disruptions that occurred during a teleconference based on the emotions of participants associated with a plurality of audio data arranged chronologically in order of reception start time, and associates the detected disruptions with the plurality of audio data. A teleconferencing system characterized by the following features.
2. A telephone conferencing system according to claim 1, The aforementioned telephone conferencing device is The system further includes a word list storage means that stores a list of predetermined parts of speech that may be included in the statements of each participant in that role, The aforementioned role determination means is For each audio data stored in the audio data storage means, the word list storage means searches for the word list that contains the most words associated with the audio data, and the participant's role associated with the searched word list is determined to be the participant's role associated with the audio data. A teleconferencing system characterized by the following features.
3. A telephone conferencing system comprising: a plurality of telephone conferencing terminals; a telephone conferencing device for each telephone conferencing terminal that mixes audio data received from the plurality of telephone conferencing terminals excluding the telephone conferencing terminal to generate telephone conferencing data and transmits it to the telephone conferencing terminal; and a management terminal connected to the telephone conferencing device, The aforementioned telephone conferencing device is A voice data storage means that stores each of the voice data received from the plurality of telephone conference terminals, linked to the originating telephone conference terminal and the start time of reception; For each audio data stored in the aforementioned audio data storage means, a speech recognition means performs speech recognition processing on the audio data to generate text data, and associates the generated text data with the audio data. A text analysis means performs text analysis on text data associated with each audio data stored in the audio data storage means, extracts words of a predetermined part of speech from the text data, and associates the extracted words with the audio data. For each audio data stored in the audio data storage means, a role determination means determines the role of the participant who spoke the audio data based on the words associated with the audio data, and associates the determined role with the audio data. A word list storage means stores a list of words of a predetermined part of speech that may be included in the speech of a participant in that role, for each participant's role. The system includes a word list update means that updates the registered contents of the word list storage means in accordance with instructions received from the management terminal, The aforementioned role determination means is For each audio data stored in the audio data storage means, the word list storage means searches for the word list that contains the most words associated with the audio data, and the participant's role associated with the searched word list is determined to be the participant's role associated with the audio data. A teleconferencing system characterized by the following features.
4. A telephone conferencing system comprising: a plurality of telephone conferencing terminals; and a telephone conferencing device for each telephone conferencing terminal that generates telephone conferencing data by mixing audio data received from the plurality of telephone conferencing terminals excluding the telephone conferencing terminal and transmits it to the telephone conferencing terminal, The aforementioned telephone conferencing device is A voice data storage means that stores each of the voice data received from the plurality of telephone conference terminals, linked to the originating telephone conference terminal and the start time of reception; For each audio data stored in the aforementioned audio data storage means, a speech recognition means performs speech recognition processing on the audio data to generate text data, and associates the generated text data with the audio data. A text analysis means performs text analysis on text data associated with each audio data stored in the audio data storage means, extracts words of a predetermined part of speech from the text data, and associates the extracted words with the audio data. For each audio data stored in the audio data storage means, a role determination means determines the role of the participant who spoke the audio data based on the words associated with the audio data, and associates the determined role with the audio data. A word list storage means stores a list of words of a predetermined part of speech that may be included in the speech of a participant in that role, for each participant's role. For each audio data stored in the aforementioned audio data storage means, an emotion determination means determines the emotion of the participant who spoke the audio data based on the acoustic characteristics of the audio data, including the volume level and speech pitch, and associates the determined emotion of the participant with the audio data. The aforementioned audio data storage means includes a disruption detection means that detects disruptions that occurred during a teleconference based on the roles and emotional changes of the participants associated with a plurality of audio data arranged chronologically in order of reception start time, and associates the detected disruptions with the plurality of audio data. The aforementioned role determination means is For each audio data stored in the audio data storage means, the word list storage means searches for the word list that contains the most words associated with the audio data, and the participant's role associated with the searched word list is determined to be the participant's role associated with the audio data. A teleconferencing system characterized by the following features.
5. A telephone conferencing system according to claim 1, The aforementioned telephone conferencing device is The system further includes a problem information storage means that stores, for each problem anticipated in the aforementioned teleconference, the content of the problem, linked to a problem occurrence pattern that includes the roles and emotions of the participants involved in the occurrence of the problem. The aforementioned fault detection means is For each disruption pattern stored in the disruption information storage means, the sequence of roles and emotions of the participants that matches the disruption pattern is searched for in the sequence of audio data arranged chronologically in order of reception start time in the audio data storage means. If a sequence of audio data containing a sequence of roles and emotions that matches the disruption pattern is detected, it is determined that a disruption occurred in the conversation indicated by that sequence of audio data, and the content of the disruption associated with the disruption pattern is linked to that sequence of audio data. A teleconferencing system characterized by the following features.
6. A telephone conferencing system according to claim 4, The aforementioned telephone conferencing device is The system further includes a problem information storage means that stores, for each problem anticipated in the aforementioned teleconference, the content of the problem, linked to a problem occurrence pattern that includes the roles and emotions of the participants involved in the occurrence of the problem. The aforementioned fault detection means is For each disruption pattern stored in the disruption information storage means, the sequence of roles and emotions of the participants that matches the disruption pattern is searched for in the sequence of audio data arranged chronologically in order of reception start time in the audio data storage means. If a sequence of audio data containing a sequence of roles and emotions that matches the disruption pattern is detected, it is determined that a disruption occurred in the conversation indicated by that sequence of audio data, and the content of the disruption associated with the disruption pattern is linked to that sequence of audio data. A teleconferencing system characterized by the following features.
7. A telephone conferencing system according to claim 4 or 6, The system further includes a management terminal connected to the aforementioned teleconferencing device, The aforementioned telephone conferencing device is The system further includes a word list update means that updates the registered contents of the word list storage means in accordance with instructions received from the management terminal. A teleconferencing system characterized by the following features.
8. A telephone conferencing system according to claim 5 or 6, The system further includes a management terminal connected to the aforementioned teleconferencing device, The aforementioned telephone conferencing device is The system further includes a fault information updating means that updates the registered contents of the fault information storage means in accordance with instructions received from the management terminal. A teleconferencing system characterized by the following features.
9. A teleconferencing device that generates teleconferencing data by mixing audio data received from multiple teleconferencing terminals excluding the teleconferencing terminal, and transmits it to the teleconferencing terminal, A voice data storage means that stores each of the voice data received from the plurality of telephone conference terminals, linked to the originating telephone conference terminal and the start time of reception; For each audio data stored in the aforementioned audio data storage means, a speech recognition means performs speech recognition processing on the audio data to generate text data, and associates the generated text data with the audio data. A text analysis means performs text analysis on text data associated with each audio data stored in the audio data storage means, extracts words of a predetermined part of speech from the text data, and associates the extracted words with the audio data. For each audio data stored in the audio data storage means, a role determination means determines the role of the participant who spoke the audio data based on the words associated with the audio data, and associates the determined role with the audio data. For each audio data stored in the aforementioned audio data storage means, an emotion determination means determines the emotion of the participant who spoke the audio data based on the acoustic characteristics of the audio data, including the volume level and speech pitch, and associates the determined emotion of the participant with the audio data. The audio data storage means includes a disruption detection means that detects disruptions that occurred during a teleconference based on the emotions of participants associated with a plurality of audio data arranged chronologically in order of reception start time, and associates the detected disruptions with the plurality of audio data. A teleconferencing device characterized by the following features.
10. A program that causes a computer to function as a teleconferencing device, which mixes audio data received from multiple teleconferencing terminals (excluding the terminal in question) to generate teleconferencing data and transmits it to the teleconferencing terminal, A voice data storage means that stores each of the voice data received from the plurality of telephone conference terminals, linked to the originating telephone conference terminal and the start time of reception. For each audio data stored in the aforementioned audio data storage means, speech recognition means performs speech recognition processing on the audio data to generate text data, and associates the generated text data with the audio data. For each audio data stored in the audio data storage means, the text analysis means performs text analysis on the text data associated with the audio data, extracts words of a predetermined part of speech from the text data, and associates the extracted words with the audio data. For each audio data stored in the audio data storage means, a role determination means determines the role of the participant who spoke the audio data based on the words associated with the audio data, and associates the determined role with the audio data. For each audio data stored in the audio data storage means, an emotion determination means determines the emotion of the participant who spoke the audio data based on the acoustic characteristics including the volume level and speech pitch of the audio data, and associates the determined emotion of the participant with the audio data, and In the aforementioned audio data storage means, the computer functions as a disruption detection means that detects disruptions that occurred during a teleconference based on the emotions of participants associated with a plurality of audio data arranged chronologically in order of reception start time, and associates the detected disruptions with the plurality of audio data. A program characterized by the following features.
11. A method for determining the role of a speaker in a teleconferencing conference using a teleconferencing device that generates teleconferencing data by mixing audio data received from multiple teleconferencing terminals excluding the teleconferencing terminal in question, and transmits it to the teleconferencing terminal, Each of the audio data received from the aforementioned multiple teleconferencing terminals is stored in association with the originating teleconferencing terminal and the start time of reception. For each stored audio data, speech recognition processing is performed on that audio data to generate text data, and the generated text data is linked to and stored in relation to that audio data. For each stored audio data, text analysis is performed on the text data associated with that audio data to extract words of a predetermined part of speech from the text data, and these extracted words are stored in association with the audio data. For each stored audio data, the role of the participant who spoke the audio data is determined based on the aforementioned phrase associated with that audio data, and the determined role is stored in association with that audio data. For each stored audio data, the emotions of the participant who spoke the audio data are determined based on the acoustic characteristics of the audio data, including the volume level and speech pitch, and the determined emotions of the participant are associated with the audio data and stored in memory. Based on the emotions of participants associated with multiple stored audio data sets arranged chronologically in order of reception start time, the system detects disruptions that occurred during the conference call and stores the detected disruptions in association with those multiple audio data sets. A method for determining the role of a speaker in a teleconference, characterized by the following features.
Citation Information
Patent Citations
Conference support system, proceeding forming method, and computer program
JP2005277462A
Communication apparatus
JP2008141348A
Program, method, information processing device, and system
JP2022099335A