Utterance information extraction device and program thereof

The speech information extraction device optimizes speaker diarization by accounting for microphone-speaker relationships, enhancing accuracy and efficiency in speech extraction and visualization.

JP2025113752AActive Publication Date: 2025-08-04ACES CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024008068
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2025-08-04
Estimated Expiration
2044-01-23

AI Technical Summary

Technical Problem

Existing speaker diarization technologies do not account for the relationship between speakers and microphones, leading to inaccuracies in speech extraction and processing.

Method used

A speech information extraction device that acquires speech data, determines the number of microphones and speakers, segments the speech, associates the segments with speakers, and displays the associations, optimizing processing when the number of microphones matches speakers and minimizing unnecessary processing when it does not.

Benefits of technology

Improves processing accuracy and efficiency by clarifying the relationship between speakers and microphones, allowing for accurate speech extraction and visualization of speaker associations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025113752000001_ABST
    Figure 2025113752000001_ABST
Patent Text Reader

Abstract

To provide an utterance information extraction device and a program, which clarify a relationship between a speaker and a microphone and improve processing accuracy when utterance of the speaker is extracted from voice data.SOLUTION: An utterance information extraction device 1 includes: voice data acquisition means 3 for acquiring voice data 2; information acquisition means 4 for acquiring the number of microphones used in the voice data 2 and the number of speakers; utterance division means 5 for dividing the voice data 2 for each utterance into divided utterance; linking means 6 for linking divided utterance and the speaker; and display means 7 for displaying the speaker and a linked link result. When the number of microphones and the number of speakers are the same, processing for linking divided utterance recorded in each microphone with the single speaker is performed in the voice data 2.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an apparatus and a program for recognizing a speaker and extracting speech from voice data such as online meetings or lectures.

Background Art

[0002] A technique for estimating and extracting who spoke when from voice data such as online meetings or lectures is called speaker diarization technology. For example, for voice data of an online meeting, by recognizing the speaker and extracting the speech, it is possible to analyze the effectiveness of the speech in the online meeting afterwards.

[0003] Patent Document 1 discloses a speaker diarization method for a voice file, which uses a reference voice of a speaker to identify the speaker of the reference voice from the voice file, and performs speaker identification on the remaining unrecognized speech segments using clustering.

[0004] In the method of Patent Document 1, there is no description of how to handle the microphone that recorded the voice file. Therefore, the voice file is simply segmented for each speech, and it is identified which speaker the segmented speech belongs to.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] In the method of Patent Document 1, the same processing is performed regardless of whether the number of microphones that recorded the voice file is one or plural. However, the inventors of the present application have found that the relationship between the speaker and the microphone affects the accuracy of the processing when performing speaker diarization.

[0007] An object of the present invention is to provide a speech information extraction device and a program that clarify the relationship between a speaker and a microphone and improve the processing accuracy when extracting the speech of the speaker from speech data.

Means for Solving the Problems

[0008] In order to achieve the above object, a speech information extraction device of the present invention is a speech information extraction device that extracts the speech of a speaker from speech data, and includes speech data acquisition means for acquiring the speech data, information acquisition means for acquiring the number of microphones used in the speech data and the number of speakers, speech segmentation means for dividing the speech data into segments for each speech to obtain segmented speech, association means for associating the segmented speech with the speaker, and display means for displaying the association result associated with the speaker. When the number of microphones is the same as the number of speakers, the association means associates the segmented speech for each microphone with a single speaker, and when the number of microphones is less than the number of speakers, the association means associates the segmented speech with the speaker for each segmented speech.

[0009] The speech information extraction device of the present invention obtains the number of microphones and the number of speakers by means of information acquisition. When the number of microphones is less than the number of speakers, the voice data recorded by each microphone will include the speeches of multiple speakers. Therefore, the association with the speaker is performed for each divided speech. On the other hand, when the number of microphones is the same as the number of speakers, the voice data recorded by each microphone is only the speech by a single speaker. Therefore, for the voice data of each microphone, once it can be determined which speaker's speech it is, there is no need to determine whether it is the speech of other speakers afterwards. Thus, when the number of microphones is the same as the number of speakers, unnecessary processing is not performed, so there is no incorrect processing. It is only necessary to perform the association process when the number of microphones is less than the number of speakers, which improves the processing efficiency and the processing accuracy.

[0010] In the speech information extraction device of the present invention, when the number of the microphones is the same as the number of the speakers, the display means displays the speaker associated with each microphone. When there are a plurality of speakers associated with the microphone, the main speaker is displayed and a plurality of displays indicating that there are a plurality of speakers are performed. When there is an indication to disclose the plurality of displays, all of the plurality of speakers associated with the microphone may be displayed.

[0011] According to this configuration, when the voice data recorded by a certain microphone includes the speeches of a plurality of speakers, the user can visually confirm the main speaker and the fact that there are a plurality of speakers at a glance. Also, when the user wants to know who the other speakers using that microphone are, the speaker can be known by giving an indication to disclose. On the other hand, when there is no need to know the information of other speakers, or when the speakers are erroneously divided, the indication to disclose can be cancelled. Here, the main speaker may be the speaker with the longest speech time of the voice data recorded by that microphone, or the owner of the microphone may be the main speaker.

[0012] In the speech information extraction device of the present invention, the associating means further includes a speaker DB learning means for referring to a speaker DB (DB = database; the same applies hereinafter) learned by reference voices of a plurality of speakers and learning the speaker DB, and the speaker DB learning means may enable a user to edit parameters constituting the reference voice of the speaker through a management screen.

[0013] According to this configuration, for example, regarding the data of a certain speaker stored in the speaker DB, when looking at the meeting record recorded as a parameter and determining that the speech in the meeting on a specific day is inappropriate as a feature amount of the speaker, such as being a meeting attended by another speaker, etc., the data of the meeting on that specific day can be deleted. By this editing, the data in the speaker DB becomes the data of the original speaker, so that the reliability of the data in the speaker DB can be improved.

[0014] In this configuration, the parameter may be an event in which the reference voice is recorded. Here, an event includes an online meeting, a lecture, a panel discussion, etc. Thus, when the parameter is an event, when there is an event that should not be a reference voice in the speaker DB, editing such as deletion can be performed for each event. This editing may be performed by the user or automatically using a predetermined rule, algorithm, or the like.

[0015] Also, in this configuration, the speaker DB may include a feature amount of speech as a result of analyzing the reference voice of the speaker. According to this configuration, since the feature amount of the speaker is included as the data of the speaker DB, the speech of the speaker can be accurately extracted from the voice data. Here, the reference voice is the voice that should be used as a reference when extracting the features of the speaker, and includes not only ordinary conversations but also conversations with various emotions.

[0016] In the utterance information extraction device of the present invention, a correction means that can be corrected by the user is provided for the association result of the association means, and the speaker DB learning means causes the speaker DB to learn the correction result corrected by the user, and the association means may perform the association with reference to the learned speaker DB.

[0017] According to this configuration, even if there is an error in the processing by the association means, the user can correct the error and reflect it in the speaker DB. Therefore, for example, even if the user corrects an error in one association, since the speaker DB is updated, other similar errors will also be automatically corrected.

Effect of the Invention

[0018] According to the present invention, when extracting the utterance of the speaker from the voice data, it is possible to provide an utterance information extraction device that clarifies the relationship between the speaker and the microphone and realizes an improvement in processing accuracy.

Brief Description of the Drawings

[0019]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Embodiments for Carrying Out the Invention

[0020] Next, referring to FIGS. 1 to 6, a speech information extraction device and a speech information extraction program according to an embodiment of the present invention will be described. The speech information extraction device 1 according to the present embodiment is, for example, a device that extracts and displays voice data 2, which is the voice of a meeting recorded in an online meeting or the like, for each speaker S.

[0021] As a functional configuration, the speech information extraction device 1 according to the present embodiment includes, as shown in FIG. 1, a voice data acquisition unit 3 that acquires voice data 2, an information acquisition unit 4 that acquires the number of microphones M used in the voice data 2 and the number of speakers S, a speech segmentation unit 5 that divides the voice data 2 into segments for each speech to obtain segmented speech 10, an association unit 6 that associates the segmented speech 10 with the speaker S, and a display unit 7 that displays an association result 11 associated with the speaker S.

[0022] Further, in the speech information extraction device 1 according to the present embodiment, the association unit 6 is connected to a speaker DB 8. The speaker DB 8 is provided with a speaker DB learning unit 9 that performs learning of the reference voice of the speaker S. The speaker DB 8 may be connected to the association unit 6 via a network or the like, or may be stored in a storage unit inside the speech information extraction device 1.

[0023] The voice data 2 is, for example, the voice of a meeting recorded in an online meeting or the like. In addition, data recorded in a panel discussion or the like in which a plurality of panelists (speakers) participate is also included.

[0024] The voice data acquisition unit 3 is a functional unit that acquires the voice data 2 in the speech information extraction device 1. The data format of the voice data 2 may be data of only voice, or may be a voice track of video data. The voice data acquisition unit 3 can be acquired, for example, by uploading the voice data 2 to a computer that is the speech information extraction device 1. The upload of the voice data 2 may be performed via a network or the like, or may be performed via a recording medium such as an SD card.

[0025] The information acquisition means 4 is a means for acquiring the number of microphones M used in the voice data 2 and the number of the speakers. For example, in the case of an online meeting via a network 14 as shown in Fig. 2(B), at the base (d), a speaker Sd conducts a meeting in front of the meeting device 12d. The speaker Sd wears a headset 13d, and the microphone Md has a one-to-one relationship with the speaker Sd. On the other hand, at the base (e), one microphone Me is installed in front of the meeting device 12e, and a plurality of speakers Se1 to Se3 speak using the one microphone Me.

[0026] When extracting speaker information from the voice data 2 of such a meeting, the information acquisition means 4 acquires information on the number of microphones M and the number of speakers S based on data attached to the voice data 2, such as data on microphone IDs and speaker IDs. Specifically, a method of acquiring the IDs of users who participated in the meeting using an online meeting tool or a method of having the user input information using ACESMeet, which is an automated minutes system by AI provided by the applicant, can be used.

[0027] The utterance segmentation means 5 is a functional unit that segments the voice data 2 into individual utterances to obtain segmented utterances 10. As a segmentation method, for example, a method based on the power threshold of the voice data 2 can be used. In this method, the voice is segmented every certain time (e.g., 0.25 milliseconds), and if the total power in the segmented frame is equal to or greater than the threshold, the frame is regarded as a voice frame, and if it is less than the threshold, it is regarded as a non-voice frame. Then, the sections where the voice frames are continuous are grouped together as the segmented utterances 10. Note that the method for segmenting the voice data 2 is not limited to this method, and other known methods may also be used.

[0028] The linking means 6 is a functional unit that links the segmented utterance 10 segmented by the utterance segmentation means 5 with the speaker S. For example, in the case of an online meeting as shown in Fig. 2(A), at bases (a) to (c), there is a configuration where each speaker Sa to Sc has one microphone Ma to Mc respectively. In such a case, for the segmented utterance 10 of the voice data 2 of each microphone Ma to Mc, since it only consists of the voices of the speakers Sa to Sc who are the users of the microphones Ma to Mc, for the segmented utterance 10 of the microphones Ma to Mc, speakers other than the speakers Sa to Sc are not linked.

[0029] On the other hand, in the case of an online meeting as shown in Fig. 2(B), at base (e), since a plurality of speakers Se1 to Se3 are using the microphone Me, for the segmented utterance 10 of the microphone Me, the linking of which speaker's utterance it is is performed.

[0030] The display means 7 is a functional unit that displays the linking result 11, which is the segmented utterance 10 linked with the speaker S. The linking result 11 is, for example, as shown in Fig. 3(A), on the result display screen 20, for each microphone M, the name 21 of the speaker S who is speaking with that microphone M, and the segmented utterance 10 indicating the utterance made by the speaker at each microphone are displayed. Fig. 3(A) shows the linking result 11 during the online meeting in Fig. 2(A).

[0031] The linking result 11 during the online meeting in Fig. 2(B) is as shown in Fig. 3(B). In Fig. 3(B), for the number of microphones, the name 21 of the main speaker of that microphone and the segmented utterance 10 are displayed. Also, in Fig. 3(B), on the left side of the display of the name 21 of the microphone in the upper row, a multiple display 22 indicating that there are a plurality of speakers using that microphone is displayed.

[0032] As described above, in the utterance information extraction device 1 of the present embodiment, the display means 7 performs the plurality of displays 22 to notify the user that there are a plurality of speakers associated with the microphone. Further, by the user performing a touch operation or the like to give a disclosure instruction, all of the plurality of speakers associated with the microphone can be displayed in this plurality of displays 22.

[0033] The speaker DB 8 is a database in which the results learned by the reference voices of a plurality of speakers are stored. The data stored in the speaker DB 8 includes the names of the speakers, a plurality of reference voices of the speakers, and the feature amounts (vectors) of the utterances as a result of analyzing the reference voices.

[0034] As the speaker DB 8, for example, "d-vector", "i-vector", etc., which are known as text-dependent speaker authentication using a deep neural network, can be used. In addition, as the speaker DB 8, other databases such as "Pyannote.audio", which is an open-source framework in Python, may be used.

[0035] The speaker DB learning means 9 is a functional unit that can edit parameters by the management screen 23 shown in FIG. 5 and performs learning of the speaker DB 8 by changing the parameters. As shown in FIG. 5, the management screen 23 displays the currently registered meetings (events) in the item of "list of meetings registered in the database of Mr. ○○" as the speaker. In this case, the event that is the parameter is a meeting.

[0036] Next to the meeting, a delete button 24 is provided and used when the user wants to delete the meeting. Also, an add button 25 is provided at the last row where meetings are arranged. Further, below the management screen 23, a register button 26 for registering the edited content and a cancel button 27 for canceling are provided. In the present embodiment, the management screen 23 displays meeting data as parameters for learning the speaker DB 8. However, the present invention is not limited to this, and specific utterances during a meeting, for example, only a specific section of one-hour voice data can be specified, and voice files uploaded by the user himself / herself can be used as parameters.

[0037] Based on the data stored in the speaker DB 8, the association means 6 determines which speaker the certain segmented utterance 10 is from. As a method of determination, the association means 6 may use AI (Artificial Intelligence) for determination. When performing determination by AI, the input to the AI is the segmented utterance 10, and the output is the speaker S who is presumed to have made the utterance of the segmented utterance 10.

[0038] As another determination method, the voice of the segmented utterance 10 can be analyzed to calculate the feature amount of the utterance, and the feature amount can be compared with the data regarding the feature amount stored in the speaker DB 8, and the speaker who made an utterance with a similar feature amount can be used as the determination result. Also, as a determination method, a known method in the research fields of "speaker classification" and "speaker verification" may be adopted.

[0039] The utterance information extraction device 1 of the present embodiment is realized by operating an utterance information extraction program 1P using a computer (user terminal 1U) used by the user or the like. This utterance information extraction program 1P may be installed in a personal computer, or may be installed in a server and used on a client terminal. Further, it may be stored in a CD-ROM, DVD-ROM, etc., or may be uploaded onto a server and downloadable through a network.

[0040] The computer on which the utterance information extraction program 1P is executed includes a processor such as a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit), storage means such as a hard disk and a memory, connection means to various networks, a keyboard, a mouse, and a display, etc. (not shown).

[0041] Next, the operation of the utterance information extraction device 1 of the present embodiment will be described with reference to FIGS. 1 to 6. When the user activates the user terminal 1U and activates the utterance information extraction program 1P, an initial screen for recognizing the user is displayed on the screen of the user terminal 1U (not shown). When the user inputs necessary items to the initial screen, if there is a registered online meeting tool, voice data 2 is input from the video data of the online meeting tool by the voice data acquisition means 3.

[0042] Next, when the user designates the speaker DB 8 to be referred to and gives an execution instruction, first, a splitting process is performed in which the utterance splitting means 5 splits the voice data 2 into split utterances 10 (STEP1). As a splitting method, the method based on the power threshold of the voice data 2 described above is used.

[0043] Next, an association process is performed by the association means 6 to associate the split utterances 10 with the speaker S. In this association process, first, the association means 6 obtains information such as the number of microphones M, the number of speakers S, and the owners of the microphones M from the data attached to the voice data 2 (STEP2).

[0044] Here, it is determined whether the number of microphones M is the same as the number of speakers S (STEP3). When the number of microphones M is the same as the number of speakers S (Y in STEP3), for the divided utterances 10 recorded by each microphone M in the voice data 2, a process of associating them with a single speaker is performed (STEP4).

[0045] Specifically, for the divided utterances 10 recorded by each microphone M, the association between each microphone M and the speaker S is performed based on the information of the owner of the microphone M obtained in STEP2. The display means 7 displays the result of this association on the result display screen 20 (STEP5).

[0046] On the other hand, when the number of microphones M is less than the number of speakers S (N in STEP3), the association of the speaker S is performed for each divided utterance 10 (STEP6). The association process uses the divided utterance 10 generated from the voice data 2 as the input of the determination AI, and as the output, obtains the information of the divided utterance 10 associated with the speaker, and is displayed by the display means 7 (STEP7).

[0047] Here, in the voice data 2 obtained by the voice data acquisition means 3, if the voice data for each microphone M is not available and it is a single file, when the speaker S of each microphone M can be identified from the video data from which the voice data 2 is derived, the association with the speaker S is performed for each microphone M. On the other hand, when the speaker S of each microphone M cannot be identified, it is processed as if a plurality of speakers S are speaking in one microphone M.

[0048] Occasionally, the number of microphones M may be more than the number of speakers S. In this case, during the division process by the utterance division means 5, the microphone M with low power is identified, and it is determined that the microphone M is not in use, and the subsequent process may be performed excluding that microphone M. Alternatively, after performing the division process for each microphone M, the association process may be performed as usual.

[0049] The display by the display means 7 is performed by causing the user terminal 1U to display the result of the association by the association means 6 (STEP5, 7). The result of the association can be, for example, in the configuration shown in FIG. 3.

[0050] When the number of microphones M is the same as the number of speakers S (Y in STEP3), the display means 7 displays the result display screen 20 shown in FIG. 3(A) (STEP5). In FIG. 3(A), as a result of the association, the name (surname) of the main speaker S in each microphone and the segmented speech 10 in which the speaker S spoke in the voice data 2 obtained by that microphone are displayed. This segmented speech 10 is displayed in time series with the time stored in the voice data 2 on the horizontal axis. In the example of FIG. 3(A), it can be seen that Mr. Yamada, who is the speaker S, spoke first, and then Mr. Suzuki and Mr. Sato spoke in that order.

[0051] On the other hand, when the number of microphones M is less than the number of speakers S (N in STEP3), the display becomes the result display screen 20 shown in FIG. 3(B). In FIG. 3(B), as a result of the association, the main speaker S of each microphone is displayed. Specifically, the name of Mr. Yamada is displayed on the microphone in the first row, and the name of Mr. Sato is displayed on the microphone in the second row.

[0052] Also, in the result display screen 20 of FIG. 3(B), a plurality display 22 indicating that there are a plurality of speakers S using the microphone M is displayed to the left of the display of the microphone in the first row. Specifically, the plurality display 22 is in a state where one vertex of a triangle is directed toward Mr. Yamada, who is the speaker S.

[0053] The user can issue a disclosure instruction for the multiple display by performing operations such as clicking on the multiple display 22 with the mouse of the user terminal 1U or touching the screen. When receiving a disclosure instruction from the user (Y in STEP8), the display means 7 displays all of the multiple speakers associated with the microphone M (STEP9). At this time, the multiple display 22 is changed so that one vertex of the triangle faces downward. In FIG. 3(C), it can be seen that there are two speakers associated with Mr. Yamada's microphone M other than Mr. Yamada.

[0054] In FIG. 3(C), the speakers associated with Mr. Yamada's microphone M are "Yamada_1" and Mr. Suzuki. The display of "Yamada_1" is the split utterance 10 associated with Mr. Yamada's microphone M, indicating a state where it could not be associated by the association means 6. In this state, if the user clicks on the multiple display 22 again, the display of the split utterance 10 of speakers other than Mr. Yamada that was being displayed is folded up and returns to the state of FIG. 3(B).

[0055] Regarding the speaker S for which association could not be made, the user can play and confirm the content of the split utterance 10 by designating the split utterance 10 of that speaker and issuing a playback instruction. In the present embodiment, in such a case, the split utterance 10 with incorrect association can be corrected on the result display screen 20 (correction means in the present invention).

[0056] Here, when the user corrects the speaker S for the split utterance 10 (Y in STEP10), as shown in FIG. 3(D), the user can change the display of the split utterance 10 to the corrected content (STEP11). Here, "Yamada_1" is corrected to Mr. Saito.

[0057] In this way, when the association of the split utterance 10 is corrected by the user, the voice data 2 for which speaker extraction was performed this time is registered as data to be used for speaker association in the speaker DB8 (STEP12). As a result, the voice data 2 can be used for speaker association in subsequent processes.

[0058] Next, with reference to FIG. 5, a case where the user uses the speaker DB learning means 9, changes parameters using the management screen 23, and performs learning of the speaker DB 8 will be described. In FIG. 5, a list of meetings registered in the database of the speaker "Mr. ○○" is displayed. When the user views this meeting list and determines, for example, that Meeting 2 is not suitable for learning, the user can click the delete button 24 to the right of Meeting 2 to delete Meeting 2.

[0059] Also, when the user determines that they want to register another meeting, the user clicks the add button 25. Then, a screen (not shown) prompting the user to specify the meeting to be registered is displayed, so the user can specify the voice data 2 of the meeting that they consider necessary to register and register it in the speaker DB 8.

[0060] Next, the correction when there is an error in the segmented utterance 10 for which the association process has been performed will be described with reference to FIG. 6. For example, in the state shown in FIG. 6(A), when the user checks the segmented utterance 10 at the right end among the three segmented utterances 10 recorded by Mr. Yamada's microphone M, it is found that it is the utterance of Mr. Sato instead of Mr. Suzuki.

[0061] In this case, since the user needs to make a correction (Y in STEP10), the user can move the segmented utterance 10 at the right end from Mr. Suzuki to Mr. Sato by operating the mouse or touch, etc., and make a correction (STEP11). When this correction is confirmed, the corrected content is registered in the speaker DB 8 (STEP12).

[0062] Here, when the user wants to perform the association process again, by executing the process again, the association means 6 can refer to the learned speaker DB 8 after the correction and perform the association process again. As a result, for example, when it is found that the second segmented utterance 10 from the left of Mr. Suzuki in FIG. 6 is also the utterance of Mr. Sato, the display means 7 displays the corrected segmented utterance 10 as shown in FIG. 6(B).

[0063] As described above, according to the utterance information extraction device 1 of the present embodiment, when the number of microphones M is the same as the number of speakers S, the voice data 2 recorded by each microphone M is only the utterance by a single speaker. Therefore, for the voice data 2 from each microphone M, it is not necessary to perform the process of determining which speaker S the utterance is from. Thus, in this case, since unnecessary processes are not performed, incorrect processes are not carried out. It is only necessary to perform the association process when the number of microphones M is less than the number of speakers, and the processing efficiency and the processing accuracy are improved.

[0064] Also, when the voice data 2 recorded by a certain microphone M includes the utterances of a plurality of speakers S, on the result display screen 20, the user can visually confirm the main speaker S and the fact that there are a plurality of speakers S at a glance. Further, when the user wants to know who the other speaker S is who used that microphone M, the user can know that speaker S by giving a disclosure instruction.

[0065] On the other hand, when the user checks other speakers S who used the microphone M through the multiple display 22, there may be a case where it is found that although the speaker S is actually alone, it is misrecognized that two or more people are speaking. In this case, if the user clicks on the multiple display 22 again, the display of the plurality of divided utterances 10 that were being displayed can be folded and returned to the state of FIG. 3(B). With this configuration, even when the recognition of the speaker S is incorrect, the incorrect divided utterance 10 can be made non-displayed.

[0066] Also, for the data of a certain speaker S stored in the speaker DB 8, when looking at the recorded meeting records recorded as parameters and determining that they are not suitable as learning data, editing such as deleting the meeting data of that specific day can be performed. By this editing, the reliability of the data in the speaker DB 8 can be improved.

[0067] In addition, by making the parameters of the speaker DB8 editable, it becomes possible to associate a plurality of meetings with a single speaker S. Even for the same speaker S, the characteristics of their voice are affected by the type of microphone being used, the speaker's physical condition, variations in speech (such as laughter, whispering, etc.), the distance between the microphone and the speaker, and the acoustic environment characteristics related to the size and reverberation of the room. For this reason, if the feature amount is estimated only from a specific meeting for speaker S, the accuracy of determining speaker S in other acoustic environments will be low. However, by associating meetings in a plurality of acoustic environments, it is possible to improve the determination accuracy.

[0068] Furthermore, even if there is an error in the segmented speech 10 for which the association process has been performed, the user can make a correction, and the correction result is reflected in the speaker DB8. Therefore, if a single correction is made, other similar corrections can be automatically performed.

[0069] Note that in the above embodiment, on the result display screen 20, when the voice data 2 recorded by the microphone M has a plurality of speakers S, the main speaker S is set as the speaker S with the longest speaking time in the voice data 2 recorded by the microphone M. However, this is not limiting, and the speaker S registered as the owner of the microphone M may be set as the main speaker S.

[0070] Also, in the above embodiment, the input of the voice data 2 by the voice data acquisition means 3 is automatically performed from the video data of the registered online meeting tool. However, this is not limiting, and an input screen (not shown) may be displayed on the user terminal 1U, and the voice data 2 may be input from the input screen.

[0071] Also, in the above embodiment, the multiple display 22 is in the form of a triangle. However, this is not limiting, and as long as it can be recognized that there are multiple, other display methods may be used. For example, it can be an arbitrary display such as a figure like an arrow, or a figure that blinks and stops blinking when expanded.

[0072] Also, in the above-described embodiment, in STEP4, when performing the process of associating the divided utterances 10 recorded by each microphone M among the voice data 2 with a single speaker, for the divided utterances 10 recorded by each microphone M, the association between each microphone M and the speaker S is performed based on the information of the owner of the microphone M obtained in STEP2. However, it may be input to a determination AI that makes a determination by referring to the speaker DB8, and a speaker who is presumed to have made the utterance may be obtained as the output.

[0073] Also, in the above-described embodiment, as the speaker DB learning means 9, the parameters can be edited by the management screen 23 shown in FIG. 5, but it is not limited thereto, and learning may be automatically performed. For example, in STEP4, when performing the process of associating the divided utterances 10 recorded by each microphone M with a single speaker, additional learning may be performed in the speaker DB8 to update the database. Alternatively, a separate learning management screen (not shown) may be provided so that it is possible to set whether to learn the result of the association process in STEP4.

[0074] Also, in the above-described embodiment, the multiple display 22 is configured to display all of the multiple speakers associated with the microphone when the user performs a touch operation or the like to give a disclosure instruction. However, it is not limited thereto, and a setting screen (not shown) for setting whether to perform multiple displays may be provided, and it may be possible to set whether to perform all multiple displays according to the setting.

Explanation of Reference Numerals

[0075] 1... Utterance information extraction device 1P... Utterance information extraction program 2... Voice data 3... Voice data acquisition means 4... Information acquisition means 5... Utterance division means 6... Association means 7... Display means 8... Speaker DB 9... Speaker DB learning means 10... Divided utterance 11... Association result 20…Result display screen 22…Multiple display 23…Management screen M…Microphone S…Speaker

Claims

1. A speech information extraction device for extracting a speaker's speech from speech data, comprising: speech data acquisition means for acquiring the speech data; information acquisition means for acquiring the number of microphones used in the speech data and the number of speakers; speech segmentation means for segmenting the speech data into segments for each speech; association means for associating the segmented speech with the speaker; display means for displaying the association result associated with the speaker, wherein when the number of microphones is the same as the number of speakers, the association means associates the segmented speech for each microphone with a single speaker, and when the number of microphones is less than the number of speakers, the association means associates the speech with the speaker for each segmented speech. A speech information extraction device characterized by this.

2. The speech information extraction device according to Claim 1, wherein the display means when the number of microphones is the same as the number of speakers, displays the speaker associated with each microphone, when there are a plurality of speakers associated with a microphone, displays the main speaker and performs a plurality of displays indicating that there are a plurality of speakers, and when there is an instruction to disclose the plurality of displays, displays all of the plurality of speakers associated with the microphone. A speech information extraction device characterized by this.

3. The speech information extraction device according to Claim 1, wherein the association means refers to a speaker DB learned by the reference voices of a plurality of speakers, further comprising speaker DB learning means for learning the speaker DB, wherein the speaker DB learning means is characterized in that parameters for configuring the reference voice of the speaker can be edited by a management screen. A speech information extraction device.

4. The speech information extraction device according to Claim 3, wherein the parameter is an event in which the reference voice is recorded. A speech information extraction device characterized by this.

5. The speech information extraction device according to Claim 3, wherein the speaker DB includes feature amounts of speech as a result of analyzing the reference voice of the speaker. A speech information extraction device characterized by this.

6. The speech information extraction device according to Claim 1, comprising correction means by which a user can correct the association result of the association means, wherein the speaker DB learning means causes the speaker DB to learn the correction result corrected by the user, and the association means performs association with reference to the learned speaker DB. A speech information extraction device characterized by this.

7. A speech information extraction program for operating a computer as the speech information extraction device according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Speech processing device, speech processing system, and speech processing program

    JP2009139592A

  • Conversation support system, conversation support method, and program

    JP2022015775A

  • Information processing device and program

    JP2022109048A

  • Speaker diarization method, system, and computer program coupled with speaker identification

    JP2022109867A