Speaker group estimation device
The speaker group estimation device automates the identification of speaker groups by analyzing voice data from multiple microphones, addressing inefficiencies in manual pre-association and adapting to changing group compositions, thereby reducing instructor workload and ensuring accurate group estimation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-16
- Publication Date
- 2026-03-30
AI Technical Summary
Existing methods for identifying speaker groups in active learning environments require manual pre-association of microphones with students, leading to increased workload and inefficiency, especially when group compositions change during discussions.
A speaker group estimation device that records voice data from multiple microphones, performs spectral analysis, detects speech intervals, calculates cross-correlation coefficients, and clusters speakers into groups based on proximity, reducing the need for manual pre-association and adapting to dynamic group compositions.
Automates the identification of speaker groups, reducing the workload for instructors and accurately estimating group compositions without prior preparation, even in dynamic learning scenarios.
Smart Images

Figure 0007837148000013 
Figure 0007837148000014 
Figure 0007837148000015
Abstract
Description
[Technical Field]
[0001] The disclosure herein relates to a speaker group estimation device that estimates the group to which each speaker belongs from the voice data of each of the multiple speakers when multiple speakers are divided into multiple groups and speak. [Background technology]
[0002] In recent years, active learning, an educational method in which students think and learn for themselves, has been attracting attention. Active learning is a learning method designed to enable learners (children, students, etc.) to learn actively, rather than passively receiving instruction as in the past. To implement active learning, a common approach is to divide the learners into multiple groups of several people and have them engage in discussions within each group.
[0003] Furthermore, in order to visualize the learning process of the learners through the aforementioned discussions, the learners are assigned and fitted with recording microphones (hereinafter referred to as "microphones"), and the audio data recorded through these microphones is converted into text.
[0004] Since the aforementioned microphone picks up not only the voice of the student speaking but also the voices of other nearby students, both inside and outside the group, it is required to accurately estimate the speech interval of each speaker.
[0005] Conventionally, in order to solve the above problem, a multi-channel speech interval estimation device has been proposed that includes a storage unit for storing time information and microphone identification information for audio data from three or more microphones, and a processing unit, wherein the processing unit includes a calculation unit for calculating the value of the audio power point per unit time from each of the audio data to determine the audio power, a processing calculation unit for assigning a label a to points whose value is less than a threshold a, a calculation unit for calculating separation lines in pairs of one audio power and the other audio power, an estimation unit for using the separation lines to estimate whether each point is a non-utterance point of the corresponding speaker and assigning a label c to the points estimated to be non-utterance points, and an exclusion unit for excluding the interval of audio data corresponding to the points to which label information a and c have been assigned as a non-utterance interval of the corresponding speaker (see, for example, Patent Document 1 and Non-Patent Document 1). With this configuration, the difference between the feature quantities of speech and non-speech can be increased by improving the feature quantities for speech interval detection using the fluctuating components of long intervals, thereby improving the performance of the speech interval detection.
[0006] Furthermore, in group discussions, instead of assigning a headset microphone to each speaker, a method was proposed in which a single microphone array is placed in the center of the table, and a minimum variance distortionless response (MVDR) beamformer technology is used to improve speech recognition accuracy, taking into account reverberation, overlap of multiple speakers, noise, etc. (See Non-Patent Document 2). [Prior art documents] [Patent Documents]
[0007] [Patent Document 1] Patent No. 6887622 [Non-patent literature]
[0008] [Non-Patent Document 1] Kaito Nakano, Takahiro Nakayama, Hajime Shiramizu, and Osamu Ichikawa, "Multichannel VAD for Improving Speech Recognition Accuracy in Group Work in Classrooms," Proceedings of the 82nd National Convention of the Information Processing Society of Japan 2020(1), 173-174, February 20, 2020. [Non-Patent Document 2] S. Araki, M. Okada, T. Higuchi, A. Ogawa and T. Nakatani, "Spatial correlation model based observation vector clustering and MVDR beamforming for meeting recognition," Proceedings of International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2016, pp. 385-389, 2016. [Overview of the project] [Problems that the invention aims to solve]
[0009] Incidentally, for reasons such as ensuring the smooth progress of the lesson and preventing disparities between groups, the teacher or other instructor often pre-determines the groups and the students assigned to them. Typically, the microphones are distributed at the start of the discussion. Therefore, in order to transcribe the audio data, it is necessary to identify which student belonged to which group and which microphone they used.
[0010] Currently, the identification process by instructors requires, for example, distributing microphones with students' ID numbers (such as student ID numbers) attached, and creating a management ledger that pre-associates the groups with the ID numbers of the students belonging to those groups. This process is burdensome for instructors. Furthermore, if the group composition changes on the day of the discussion due to student absences or other reasons, it may be necessary to revise the management ledger, which adds to the workload for the instructors.
[0011] Therefore, in order to reduce the workload described above, it is necessary to devise a way to perform the identification based on the audio data acquired by the microphone, without requiring any prior preparation work such as creating the management ledger mentioned above.
[0012] Furthermore, by using a distributed microphone array with multiple microphones spatially arranged, the difference in arrival time from the sound source (learner) to each microphone is a function of the sound source's position, allowing for the aforementioned identification using sound source localization techniques. In other words, this method allows for the estimation of the positional relationship between each learner and the microphone. If the learners' positions are known, it may be possible to predict the group composition to some extent based on their physical proximity.
[0013] However, the sound localization technique requires that the position of each microphone and each learner remain constant throughout the session. This proved difficult to implement in a format like the one described above, where each learner would be wearing a microphone and facing in various directions.
[0014] Furthermore, in active learning, multiple sessions may be conducted within a single class, and group members may be rearranged when switching sessions. In such cases, the groups must be estimated for each session. To do this, it is necessary to know at what point in the audio track the rearrangement occurred, which has traditionally been determined by human listening. Automating this determination is highly desirable.
[0015] Furthermore, in active learning, multiple groups discuss in a relatively narrow area such as a classroom. The conventional technology based on sound source localization tacitly assumes that there is one group in one room and does not consider the case where voices are mixed between groups.
[0016] As described above, the discussion scenario in active learning has been taken as an example. However, not limited to active learning, when multiple speakers wear microphones or the like and discuss in multiple groups within a relatively narrow area, the same problems as described above may occur.
[0017] The disclosure in this specification is for solving the above problems, and an object thereof is to provide a speaker group estimation device for estimating which group each speaker belongs to from the voice data of each speaker when multiple speakers are divided into multiple groups.
Means for Solving the Problems
[0018] In order to solve the above problems, one aspect of the speaker group estimation device disclosed in this specification includes a recording unit that records voice data of multiple speakers separately as waveform data on a voice track via a microphone having a recording function assigned to each of the speakers; a spectrum analysis unit that cuts out the waveform data recorded on the voice track as frames composed of complex spectrum data in a predetermined time unit and acquires the cut-out frames as time-series data arranged in the time direction; a detection unit that detects, as a speaking section of a target speaker to whom the microphone is assigned, a range exceeding a predetermined threshold value among the voice amplitudes of the complex spectrum data in frame units constituting the time-series data; Each voice track having the speech section is made into a pair track in all pairs, and according to a predetermined condition, one of the pair tracks is set as a main voice track in which the waveform data of the target speaker is recorded, and the other is set as a slave voice track in which the waveform data of the target speaker is recorded in a microphone assigned to a proximity speaker close to the target speaker. A setting unit for setting; The speech section of the target speaker in the main voice track is defined as the main speech section, and the speech section of the target speaker who becomes the proximity speaker for the target speaker in the main voice track in the slave voice track corresponding in the time direction to the main speech section is defined as the slave speech section. A specifying unit that specifies, as the target speech section, a range in which the slave speech sections do not overlap among the main speech sections; talk area Among the intervals, the slave talk area An identification unit that identifies a range in which the intervals do not overlap as the target speech section; A calculation unit that calculates a maximum value in the target speech section of a cross-correlation coefficient between the main voice track corresponding to the main speech section and the slave voice track corresponding to the slave speech section; An adjacency matrix generation unit that generates an adjacency matrix of the pair tracks with the maximum value of the calculated cross-correlation coefficient as a component; Each voice track constituting the pair track of the adjacency matrix is used as a node corresponding to the speaker, the maximum value of the cross-correlation coefficient is used as an edge weight indicating the degree of proximity between the nodes, and the nodes existing within a predetermined range of the edge weight are divided into a plurality of groups. And a clustering processing unit that estimates a plurality of clusters composed of a plurality of the speakers.
[0019] <, According to this configuration, for example, by clustering M voice tracks into N groups, M speakers can be clustered into N groups. And as a result of clustering, the actual group configuration can be estimated.
[0020] The spectrum analysis unit may multiply each of the waveform data by a window function, cut out the frames, perform Fourier transform on each frame, and acquire the time-series data for each voice track.
[0021] The detection unit only needs to detect the speech interval of the target speaker by selecting the range of speech power obtained from the square of the speech amplitude that exceeds a predetermined threshold.
[0022] The predetermined condition of the setting unit may be to determine the audio track with the longer sum of the speech segments detected in each frame among the paired tracks as the main audio track.
[0023] The identifying unit can identify the target utterance section by moving the subordinate audio track in the time direction by a predetermined number of frames within an arbitrarily set time range.
[0024] The calculation unit may calculate the maximum value by using the average value of the maximum cross-correlation coefficients obtained per frame.
[0025] The calculation unit only needs to calculate the maximum value of the whitening cross-correlation coefficient for the target utterance section.
[0026] The clustering processing unit can form a network graph consisting of the nodes and edge weights, perform spectral clustering on the network graph, divide the network graph into multiple groups, and estimate multiple clusters.
[0027] The clustering processing unit may be configured to include a change detection unit that divides the time-series data into detection intervals having a predetermined time length, sets a sliding window of a predetermined time unit for each detection interval, detects the point in time when the cluster changes while sliding the sliding window in the time direction, and estimates multiple groups after the change. [Effects of the Invention]
[0028] The speaker group estimation device of the present invention can estimate which group each speaker belongs to from the audio data picked up by microphones distributed to each speaker when multiple speakers are divided into multiple groups. This has the effect of reducing the workload of creating a management ledger or similar document in advance to record which speaker's voice belongs to which group and which microphone is distributed. [Brief explanation of the drawing]
[0029] [Figure 1] Figure 1 is a schematic diagram of the system including the speaker group estimation device. [Figure 2] Figure 2 is a block diagram of the speaker group estimation device. [Figure 3] Figure 3 is a processing flow diagram of the speaker group estimation device. [Figure 4] Figure 4 is a conceptual diagram of waveform data recorded in an audio track. [Figure 5] Figure 5 is an explanatory diagram of the method for extracting waveform data used in spectral analysis. [Figure 6] Figure 6 is a conceptual diagram of complex spectral data obtained by performing a Fourier transform on waveform data. [Figure 7] Figure 7 is a conceptual diagram of the waveform data in which the speech interval was detected. [Figure 8] Figure 8 is an explanatory diagram of the method for identifying the target speech segment in an audio track. [Figure 9] Figure 9 is a processing flow diagram of the calculation unit. [Figure 10] Figure 10 is a process flow diagram for determining the maximum value of the whitening cross-correlation coefficient. [Figure 11] Figure 11 shows the adjacency matrix of paired tracks, with the cross-correlation coefficient as its component. [Figure 12] Figure 12 shows the clusters estimated from the network graph. [Figure 13] Figure 13 is an explanatory diagram of a method for detecting change points in a group. [Figure 14]Figure 14 shows that the cluster has been changed. [Modes for carrying out the invention]
[0030] Hereinafter, embodiments for carrying out the disclosures herein will be described with reference to the drawings. When a subsequent embodiment has components corresponding to an embodiment described earlier, the same reference numerals will be used and redundant descriptions will be omitted. Also, when only a part of the configuration is described in each embodiment, the reference numerals of the previously described embodiment may be used for the other parts of that configuration. Even if it is not explicitly stated in each embodiment that a combination is possible, it is possible to partially combine embodiments as long as there is no particular impediment to such combination.
[0031] Figure 1 is a schematic diagram of the system including the speaker group estimation device. In this embodiment, we will explain using the example of 12 students (hereinafter referred to as "speakers") in classroom A who are divided into three groups GR1, GR2, and GR3 for a discussion. Each speaker has a unique ID number (student ID number) from 1 to 12. Each student is assigned and wears a microphone MP. However, it is assumed that it is not possible to know which student is wearing which microphone MP at the time of the discussion.
[0032] The speaker's voice is recorded via microphone MP as voice data, i.e., waveform data x, on the voice track TR of the speaker group estimation device 1. The speaker group estimation device 1 processes the voice track TR (TR1~TR 12 Waveform data x (x1 to x) recorded in ) m ) Then, clustering of clusters CL1 to CL3 with the ID numbers corresponding to groups GR1 to GR3 is performed by predetermined processing.
[0033] In other words, according to the speaker group estimation device 1 disclosed in this embodiment, first, from 12 microphones MP assigned to 12 speakers, 12 waveform data x1 to x 12 Each of these was recorded on audio tracks TR1 to TR 12 This can be obtained. And since there is a one-to-one correspondence between microphone MP and speaker, the 12 audio tracks can be clustered into 3 groups, thereby clustering the 12 members into 3 groups. The number of groups obtained through this clustering is the actual number of groups to which each of the 12 members belongs. In other words, according to the speaker group estimation device 1, M audio tracks TR recorded by capturing M waveform data with M microphone MP assigned to M members can be clustered into N groups, and the group to which each of the M members belongs can be estimated from these clustered groups.
[0034] <Configuration of Speaker Group Estimation Device 1> Figure 2 is a block diagram of the speaker group estimation device 1. The speaker group estimation device 1 includes a CPU (Central Processing Unit), RAM (memory), ROM (storage) (none of which are shown), and an input / output unit 14 such as a mouse, keyboard, display, and speaker. The CPU reads data stored in the storage (or external storage device) according to instructions input from the input / output unit 14, reads a predetermined processing program into the memory, and controls the output of data generated by the predetermined processing to the output device (e.g., display) or other input / output unit 14.
[0035] The speaker group estimation device 1 disclosed herein consists of an adjacency matrix estimation processing unit 11 and a clustering processing unit 12.
[0036] The adjacency matrix estimation processing unit 11 is further composed of a recording unit 111, a spectral analysis unit 112, a detection unit 113, a setting unit 114, a specific unit 115, a calculation unit 116, and an adjacency matrix generation unit 117.
[0037] The recording unit 111 records the voices of multiple speakers as waveform data x on the audio track TR via a microphone MP having a recording function assigned to each of the speakers. The spectral analysis unit 112 extracts each waveform data x as a frame composed of complex spectral data of a predetermined time unit, and acquires the extracted frames as time-series data arranged in the time direction. The detection unit 113 detects the range of the speech amplitude of the complex spectral data of the frame units constituting the time-series data that exceeds a predetermined threshold as the speech interval of the target speaker to which the microphone MP is assigned. The setting unit 114 sets each audio track TR having the detected speech interval as a pair track for all possible combinations, and sets one of the pair tracks as the main audio track on which the waveform data x of the target speaker is recorded, and the other as the secondary audio track on which the waveform data x of the target speaker is recorded on a microphone MP assigned to a nearby speaker close to the target speaker, according to predetermined conditions. The identification unit 115 identifies the target utterance section by defining the utterance section of the target speaker in the main audio track as the main utterance section, and the utterance section of the secondary audio track corresponding to the main utterance section in the time direction as the secondary utterance section. The calculation unit 116 calculates the maximum value of the cross-correlation coefficient between the main audio track covering the main utterance section and the secondary audio track covering the preceding secondary utterance section. The adjacency matrix generation unit 117 generates an adjacency matrix of the paired tracks, with the calculated maximum value of the cross-correlation coefficient as its component.
[0038] On the other hand, the clustering processing unit 12 uses the audio track TR that constitutes each pair of tracks in the adjacency matrix as a node corresponding to the speaker, the maximum value of the cross-correlation coefficient as an edge weight indicating the degree of proximity between the nodes, and divides the nodes that are within a predetermined range of edge weights into a plurality of groups G to estimate a plurality of clusters CL composed of a plurality of speakers. The clustering processing unit 12 may also be configured to have a change detection unit 121 that divides the time series data into detection intervals having a predetermined time length, sets a sliding window of a predetermined time unit for each detection interval, and detects the point in time when the group changes while sliding the sliding window in the time direction, and estimates a plurality of groups after the change.
[0039] The following describes the components of the speaker group estimation device 1 in detail, following the overall processing flow of the speaker group estimation device 1 shown in Figure 3.
[0040] <Record Section 111> Multiple speakers are each assigned a microphone MP with recording capabilities, such as a close-contact microphone or a headset. When each speaker speaks in a predetermined group G, each speaker's voice is captured via the microphone MP. The captured voice is recorded as waveform data x on the audio track TR by the recording unit 111 (S1).
[0041] Waveform data x is recorded in the recording unit 111 via the communication interface 13 of the speaker group estimation device 1 from the microphone MP. As described above, the microphone MP has a recording function, and is typically an IC recorder. If the microphone MP is an IC recorder, the waveform data x may be manually transferred to the recording unit 111, but if the microphone MP is capable of transferring the waveform data x via a communication network, it may be configured to record in the recording unit 111 in real time. Furthermore, the recording unit 111 may be equipped with a multi-track recorder capable of simultaneously recording each waveform data x.
[0042] Figure 4 is a conceptual diagram of the data recorded in the recording unit 111, where each waveform data x is recorded on separate audio tracks TR. In this embodiment, the waveform data x acquired via the microphone MP is recorded on M audio tracks TR (audio track TR1 to audio track TR1). M The waveform data x is recorded in the recording unit 111. The waveform data x is stored in a column-by-column manner in the recording unit 111 for each audio track TR that has recorded the waveform data x according to the time index i. Here, if the audio track index is m, the waveform data of the i-th time of the m-th audio track TR is represented as x(m,i). Note that the waveform data x stored in the recording unit 111 is a mixture of waveform data x of the speaker to whom the microphone MP is assigned (target speaker) and waveform data x of speakers adjacent to the target speaker (nearby speakers).
[0043] <Spectral Analysis Unit 112> The spectral analysis unit 112 acquires complex spectral data by performing a Fourier transform on each waveform data x recorded by the recording unit 111 (S2). The complex spectral data can be acquired by extracting the waveform data x at predetermined time intervals using a predetermined method.
[0044] The aforementioned extraction can be performed, for example, by multiplying each waveform data x by a window function such as a Hamming window, extracting it in predetermined time units, and then performing a Fourier transform on each. Figure 5 is an explanatory diagram of the method for extracting waveform data x used in spectral analysis, and Figure 6 is a conceptual diagram of the complex spectral data obtained by performing the Fourier transform on the waveform data x.
[0045] In Figure 5, the horizontal axis represents the sound acquisition time of the waveform data x, and the vertical axis represents the amplitude of the waveform data x. The sine curve represents the window function used for sampling. Here, a frame of a predetermined size (e.g., 100 times / second) is set as the unit in the time direction, and the frames are shifted (frame shifted) at predetermined time units. Samples of the data size to be Fourier transformed are zero-padding. With this method, for example, if there are several hundred sample data, complex spectral data with several hundred dimensions of frequency resolution can be obtained by performing a Fourier transform of several hundred dimensions.
[0046] For example, let X be the complex spectral data, and let X(m,k,t) represent the complex spectral data of the m-th audio track TR and the t-th frame index. Here, k represents the frequency index, that is, it represents the k-th frequency component of the m-th audio track TR. Then, the complex spectral data X(X1~X M ) refers to the audio tracks TR1~TR in Figure 4. M Correspondingly, the configuration will be as shown in Figure 6.
[0047] If, for example, a Hamming window is applied as the aforementioned window function, and the window size is L, the frame shift is S, the input width after Fourier transformation is D, and the window function for cropping is w, and if the window size L is smaller than the input width D, then the value of the window function w at positions beyond L is set to 0, then the window function w can be expressed by the following formula.
[0048]
number
[0049] Furthermore, the waveform data extracted for each frame is as follows:
[0050]
number
[0051] From the above, the complex spectral data X(m,k,t) can be obtained by performing a Fourier transform (discrete Fourier transform) on the waveform data x extracted by the window function w for each frame, as shown in the following equation.
[0052]
number
[0053] <Detection unit 113> The detection unit 113 detects the speech interval by setting a predetermined threshold for the speech amplitude of the complex spectral data X (S3). In other words, it detects whether the speaker associated with the audio track TR is speaking. The detection of the speech interval can be processed by a known speech interval detection technique (VAD technique: Voice Activity Detection). VAD technique is a technique that determines between intervals containing speech signals (speech intervals) and intervals other than speech signals (non-speech intervals) from an observed signal that contains both speech and other signals.
[0054] In this embodiment, it is assumed that the speaker speaks close to the microphone MP, such as a close-contact microphone. Therefore, as described above, even if the sound of a nearby speaker is picked up, it can be inferred that a clearly louder sound is the sound of the target speaker to whom the microphone MP has been assigned. Accordingly, it is desirable to set the threshold value high so as to detect only the range in which it can be reliably determined that the target speaker is speaking.
[0055] The detection unit 113 should set a flag as VAD information indicating whether or not the speaker associated with the audio track TR is speaking. VAD information (V1~V M The configuration shown in Figure 7 corresponds to the audio track TR in Figure 4. The VAD information is represented by V(m,t), and if its value is 1, it means that the speaker of the audio recorded in the m-th audio track TR is speaking at frame index t, and if its value is 0, it means that the speaker is not speaking at frame index t.
[0056] The detection unit 113 may also be configured to detect, instead of the speech amplitude, the range of speech power obtained from the square of the speech amplitude that exceeds a predetermined threshold as the speech interval of the target speaker.
[0057] <Settings section 114> The following steps are performed in the setting unit 114, the identification unit 115, and the calculation unit 116 to calculate the cross-correlation coefficient (S4).
[0058] In the setting unit 114, for example, if there are M VAD pieces of information detected, all combinations can be set for a pair of audio tracks TR in M × (M-1) / 2 ways. The main audio track and secondary audio track set in the pair track correspond to the direction (vector) between the sound of the target speaker picked up by the microphone MP assigned to the target speaker and the sound picked up by the microphone MP assigned to the adjacent speaker, in the calculation of the whitening cross-correlation coefficient described later. In other words, the sound spoken by the target speaker in the main audio track is also mixed into the secondary audio track.
[0059] Specifically, the condition for setting the primary audio track and the secondary audio track is to determine the audio track TR with the longer sum of the speech segments detected in each frame as the primary audio track. That is, if the two audio tracks TR constituting the paired track are p and q, then if the following equation holds, the primary audio track can be set as p and the secondary audio track as q. If it does not hold, the primary audio track is set as q and the secondary audio track as p.
[0060]
number
[0061] However, the method for determining the primary and secondary audio tracks is not limited to this. For example, the conditions for the above setting may be weighted by the audio power of each audio track TR. That is, as described above, if the frame index is t and the audio track TR is m, and the audio power is Q(m,t), then p can be set as the primary audio track and q as the secondary audio track when the following equation holds.
[0062]
number
[0063] <Specific part 115> In the identification unit 115, the target utterance section is identified by moving the subordinate audio track frame by frame in the time direction within an arbitrarily set time range. This is because the actual recording of the audio is initiated manually by each speaker, resulting in a time difference. In this embodiment, the process of obtaining the cross-correlation coefficient uses the complex spectral data X obtained per frame. Therefore, the time difference within the frame can be absorbed by the operation of finding the maximum value of the cross-correlation coefficient. However, if a delay (or advance) exceeding the frame occurs, the cross-correlation coefficient cannot be calculated, so data with one of the frames shifted is prepared to calculate the cross-correlation coefficient assuming the delay (or advance). Therefore, if the cross-correlation coefficient is calculated using waveform data x instead of complex spectral data X, the time difference can be set without the size of the frame, and this shifting of the utterance section information becomes unnecessary. However, calculating the cross-correlation coefficient using waveform data x is computationally expensive and is not recommended. Therefore, the specific unit 115 anticipates cases where the frame count exceeds a certain limit and moves a predetermined number of frames, for example, by 1 frame "delay or advance", 5 frames "delay or advance", 10 frames "delay or advance", etc. Let this frame movement be r. The frame movement r varies within the range of -R to R. The subordinate audio track is moved by the number of frames r and compared with the main audio track.
[0064] Figure 8 is an explanatory diagram of the method for identifying the target utterance section in audio track TR. If the utterance section of the target speaker in the main audio track TR(p) is p1, and the utterance section of the target speaker (a speaker adjacent to the target speaker in the main audio track TR(p)) in the secondary audio track TR(q) is q1, then the target utterance section U can be identified as the target utterance sections U1 and U2 within the utterance section p1 that do not overlap with the utterance section q1. On the other hand, the utterance section p2 of the target speaker in the main audio track TR(p) completely overlaps with the utterance section q2 of the target speaker in the secondary audio track TR(q), so the target utterance section U cannot be identified.
[0065] Based on the above, if we let V(p,t) be the speech interval of the main audio track TR(p), r be the frame shift, and V(q,t+r) be the speech interval of the secondary audio track TR(q), then the target speech interval U(p,q,t,r) can be calculated using the following formula.
[0066]
number
[0067] In other words, in relation to the speech interval V(p,t) of the main speech track TR(p), the speech interval V(q,t+r) of the secondary speech track TR(q) where the target speaker (i.e., the adjacent speaker) in the secondary speech track TR(q) is not speaking can be said to be the interval in which the voice of the speaker of the main speech track TR(p) is picked up by the microphone MP of the adjacent speaker without superimposing on the voice of the adjacent speaker. Within this range, the main speech track TR(p) and the secondary speech track TR(q), which are the paired tracks, are identified as speech intervals to be compared.
[0068] <Calculation section 116> As described above, the calculation unit 116 calculates the cross-correlation coefficient between the main voice track TR(p) covering the main voice section and the secondary voice track TR(q) covering the secondary voice section for the identified target voice section U (S4). However, it is preferable to perform the calculation of the whitening cross-correlation coefficient (maximum value) for the target voice section U. Generally, the cross-correlation coefficient measures the similarity between two waveform data, so if the voice of the target speaker in the main voice track TR(p) is observed more clearly in the secondary voice track TR(q), it will take a large value. In particular, the whitening cross-correlation coefficient has the characteristic of being normalized by the amplitude of the cross-spectrum obtained by multiplying certain frequency components of the two voice tracks TR and averaging them. Therefore, the whitening cross-correlation coefficient is suitable for the calculation process in the calculation unit 116 because it can obtain a sharp peak at the time delay position between the two and obtain a sensitive and highly accurate correlation. The calculation process of the calculation unit 116 will be explained below using the calculation of the whitening cross-correlation coefficient as an example.
[0069] Figure 9 shows a subprocess of the calculation unit 116's processing (S4) in the processing flow (overall processing flow) of the speaker group estimation device 1 in Figure 3, and further, Figure 10 shows a subprocess of the calculation of the whitening cross-correlation coefficient (S42) in Figure 9.
[0070] As shown in the processing flow of Figure 9, the calculation unit 116 first selects one of all the pair tracks set in the setting unit 114 (S41), and calculates the whitening cross-correlation coefficient for the selected pair track (S42). The calculated whitening cross-correlation coefficient is set to the corresponding position in the adjacency matrix described later (S43). If there are any pair tracks that have not been processed from S41 to S43, the processing from S41 to S43 is repeated for those unprocessed pair tracks (N in S44), and when the processing from S41 to S43 is completed for all pair tracks (Y in S44), the process returns to S5 in the overall processing flow of Figure 3.
[0071] Furthermore, as shown in the processing flow of Figure 10, the processing in S42 of Figure 9 first involves setting the main audio track TR(p) and the secondary audio track TR(q) in the setting unit 114 (S421). Before the identification unit 115 identifies the target utterance section U, the frame movement r is processed in the range from -R to +R as described above (S422). That is, the frame movement r is set to -R as the initial value and the maximum value Cmax to 0 (S422), and the frame movement r is executed to identify the target utterance section U in the identification unit 115 (S423).
[0072] For the identified target speech segment U, the whitening cross-correlation value is calculated. The whitening cross-correlation coefficient φ can be determined by the following formula.
[0073]
number
[0074] Here, p is the speech interval of the main audio track, q is the speech interval of the secondary audio track, t is the frame index, and r is the frame shift. Also, the index d of the whitening cross-correlation coefficient φ is the dimension generated from the frequency index k by the inverse Fourier transform IDFT on the right-hand side, and indicates that it has D dimensions in the time direction.
[0075] Generally, frequency components tend to have a larger proportion of low-frequency sound components, so when calculating the correlation coefficient using conventional methods, it is more heavily influenced by low-frequency sound. On the other hand, in equation 7 above, normalization is performed by the amplitude for each frequency index, thereby performing an inner product operation that makes the length of the vector equal to 1. This eliminates the influence of the loudness of sound at each frequency when calculating the correlation coefficient.
[0076] The calculated whitening cross-correlation coefficient φ represents the degree of correlation between the primary audio track TR(p) and the secondary audio track TR(q) for each frame index t. The maximum value C in the frame-level range of the whitening cross-correlation coefficient φ in the aforementioned D dimension is obtained by the following formula.
[0077]
number
[0078] This shows the similarity between the set primary audio track TR(p) and secondary audio track TR(q) in the identified target utterance section U, using a correlation coefficient. Furthermore, the average value is calculated by aggregating the maximum value C per frame for the frame index t of the target utterance section U and dividing by the total number of frames using the following formula (S424).
[0079]
number
[0080] In addition, the target utterance interval U in equation 9 represents the 1 / 0 flag, as explained above. That is, it indicates that the aggregation is performed to take the average value for the portion of the target utterance interval U that has the 1 flag.
[0081] If the (maximum) value C′ of the whitening cross-correlation coefficient after frame shift r, obtained in S424 and equation 9, is greater than the provisional maximum value Cmax, then substitute it into Cmax using the following formula (S425).
[0082]
number
[0083] Once the assignment of the maximum value Cmax for the frame movement r is complete, it is incremented (S426), and the process from S423 to S426 is repeated until the frame movement r exceeds the aforementioned +R (S427).
[0084] <Adjacency matrix generation unit 117> The adjacent matrix generation unit 117 generates an adjacent matrix of the pair of tracks with the whitened cross-correlation coefficient calculated by Equation 10 as a component (S5). FIG. 11 is a diagram showing an example of an adjacent matrix α of the pair of tracks with the cross-correlation coefficient as a component. The adjacent matrix α takes values from the audio track TR1 to the audio track TR M respectively in the matrix. When the main audio track TR(p) and the subordinate audio track TR(q) are obtained with the calculated whitened cross-correlation coefficient, numerical values are entered at two locations, α(p,q) and α(q,p). In FIG. 11, for example, for the track TR p and the track TR q , the whitened cross-correlation coefficient 0.2 obtained by the calculation is entered at two locations. Note that 0 is entered for the same audio track (for example, between audio tracks TR1) in both matrices.
[0085] The adjacent matrix α indicates the proximity of the two audio tracks TR. That is, it can be estimated that the higher the similarity between the main audio track TR(p) and the subordinate audio track TR(q) constituting the pair of tracks, the higher the value of the whitened cross-correlation coefficient and the closer the physical distance. Therefore, the value of the whitened cross-correlation coefficient serves as an index (degree of proximity) of the distance between two audio tracks.
[0086] <Clustering processing unit 12> The clustering processing unit 12 generates a network graph with the main audio track TR(p) and the subordinate audio track TR(q) respectively set in the matrix in the adjacent matrix α as nodes corresponding to the speaker, and with the value of the cross-correlation coefficient as an edge weight indicating the degree of proximity between the nodes (S6).
[0087] The upper part of Figure 12 shows a network graph β consisting of nodes ND from 1 to 10 and edges EG connecting each node ND, based on the adjacency matrix α (undirected graph). The lower part of Figure 12 shows clusters CL1, CL2, and CL3 estimated from the network graph β by non-hierarchical clustering. In Figure 12, nodes ND belonging to cluster CL1 are 1, 3, 4, and 8; nodes ND belonging to cluster CL2 are 5, 6, and 7; and nodes ND belonging to cluster CL3 are 2, 9, and 10.
[0088] In this embodiment, spectral clustering, which allows specifying the target number of clusters, will be used as the clustering process.
[0089] In a network graph β, if there are M nodes ND, (M-1) edges EG are subtracted from each node ND, resulting in an undirected graph where the edges EG have no direction. Spectral clustering targets this undirected graph, and its ultimate goal is to solve for eigenvalues called graph Laplacian. That is, clusters are created from the magnitude of the values assigned to each node ND.
[0090] In the adjacency matrix α explained in Figure 11, α(p,q) represents the weight of the edge EG, which corresponds to the strength of the connection between the primary audio track TR(p) and the secondary audio track TR(q). The adjacency matrix α is a symmetric matrix, and its diagonal terms are 0 (zero). Also, as mentioned above, there are M nodes ND, so the matrix is M × M dimensional. Here, we first estimate the graph Laplacian φ from the adjacency matrix α using the following formula (here, we take an unnormalized graph Laplacian matrix as an example).
[0091]
number
[0092] Here, ν represents a matrix of order, and is given by the following equation.
[0093]
number
[0094] Find the eigenvalues of the graph Laplacian φ, select N eigenvalues in ascending order, and create N corresponding eigenvectors (γ1, γ2…γ N ) We find the answer. By arranging these N vectors in the column direction, we obtain the matrix Γ (a matrix with M-dimensional rows and N-dimensional columns).
[0095] Take the matrix Γ in the row direction and obtain M N-dimensional vectors (ρ1, ρ2…ρ M Let's assume that these M N-dimensional vectors correspond to M speakers. By dividing these into N groups using a vector clustering method such as k-means, we can cluster the M audio tracks, i.e., the M speakers, into N groups.
[0096] Incidentally, multiple sessions are conducted within a single class, and group members may be rearranged when switching sessions. In this case, the clustering processing unit 12 must estimate the groups for each session. To do this, it is necessary to know at what point in the audio track TR the rearrangement occurred, and traditionally this was determined by a human listening. There is a need to automate this determination.
[0097] In such cases, as shown in Figure 13, the change detection unit 121 of the clustering processing unit 12 processes the complex spectral data X1, X2, X3…X M To achieve this, a detection interval I having a predetermined time length is set, and the sliding window SW is slid in the time direction by the width of the detection interval I. Each time, the point in time when the cluster changes is detected, and the multiple clusters are estimated. By performing the above detection process with the change detection unit 121, it is also possible to estimate the changed clusters CL4, CL5, and CL6 from clusters CL1, CL2, and CL3, as shown in Figure 14.
[0098] The technology disclosed in this specification is not limited to the embodiments described above. That is, it encompasses the exemplary embodiments and variations thereof by those skilled in the art. It also encompasses the substitution or combination of parts, elements between one embodiment and another. Furthermore, the scope of the disclosed technology is not limited to the descriptions of the embodiments. The scope of the disclosed technology is indicated by the claims and further includes all modifications within the meaning and scope equivalent to the claims. [Explanation of Symbols]
[0099] 1. Speaker Group Estimation Device 11 Adjacent Matrix Estimation Processing Unit 12. Clustering Processing Unit 13 Communication Interface 14 Input / output section 111 Records Department 112 Spectral Analysis Department 113 Detection unit 114 Settings Section 115 Specific section 116 Calculation Section 117 Adjacent Matrix Generation Unit 121 Change detection unit Classroom A CL cluster Cmax is the maximum value of the whitening cross-correlation coefficient over frame shift. C' (Maximum) Whitening Cross-Correlation Coefficient (Integrated from Multiple Frames) C Maximum whitening cross-correlation coefficient per frame D: Input width after Fourier transform d Index of the whitening cross-correlation coefficient φ EG Edge GR Group i Time Index k frequency index L Window size S Frame Shift Size m Audio track index MP Microphone ND node r Frame movement t Frame Index TR audio track U Target utterance section V VAD information w window function X complex spectral data x Waveform data α adjacency matrix β Network Graph
Claims
1. A recording unit that records the voices of multiple speakers as waveform data on separate audio tracks via microphones with recording functions assigned to each speaker, A spectral analysis unit extracts the waveform data recorded in the audio track as frames composed of complex spectral data for a predetermined time unit, and acquires each of the extracted frames as time-series data arranged in the time direction. A detection unit detects the range of speech amplitude in the complex spectral data of the frame units constituting the time-series data that exceeds a predetermined threshold as the speech interval of the target speaker to whom the microphone is assigned. A setting unit sets each audio track having the aforementioned speech interval as a pair track for all possible combinations, and sets one of the pair tracks as the main audio track on which the waveform data of the target speaker is recorded, and the other as the secondary audio track on which the waveform data of the target speaker is recorded on a microphone assigned to a nearby speaker adjacent to the target speaker, according to predetermined conditions. A selection unit identifies the portion of the main utterance that does not overlap with the secondary utterance, where the speech interval of the target speaker in the main audio track is defined as the main utterance, and the speech interval of the target speaker in the secondary audio track that corresponds to the main utterance in the time direction is defined as the secondary utterance, where the secondary utterance is the nearest speaker to the target speaker in the main audio track is defined as the secondary utterance, and the portion of the main utterance that does not overlap with the secondary utterance is defined as the target utterance. A calculation unit that calculates the maximum value of the cross-correlation coefficient between the main audio track relating to the main utterance section and the secondary audio track relating to the secondary utterance section within the target utterance section, An adjacency matrix generation unit generates an adjacency matrix of the paired tracks, with the maximum value of the calculated cross-correlation coefficient as an element. A speaker group estimation device comprising: a clustering processing unit that uses the audio tracks constituting each pair of tracks in the adjacency matrix as nodes corresponding to the utterances, the maximum value of the cross-correlation coefficient as an edge weight indicating the degree of proximity between the nodes, and divides the nodes that are mutually within a predetermined range of edge weights into a plurality of clusters to estimate a plurality of groups composed of a plurality of utterances.
2. The speaker group estimation device according to claim 1, wherein the spectral analysis unit multiplies each of the waveform data by a window function, extracts it into frames, performs a Fourier transform on each frame, and obtains the time-series data for each of the audio tracks.
3. The speaker group estimation device according to claim 1 or 2, wherein the detection unit detects a range of speech power obtained from the squared value of the speech amplitude that exceeds a predetermined threshold as the speech interval of the target speaker.
4. The speaker group estimation device according to any one of claims 1 to 3, wherein the predetermined condition of the setting unit is to determine the audio track with the longer sum of the speech segments detected in each frame among the paired tracks as the main audio track.
5. The speaker group estimation device according to any one of claims 1 to 4, wherein the identifying unit moves the subordinate audio track in the time direction by a predetermined number of frames within an arbitrarily set time range to identify the target utterance section.
6. The speaker group estimation device according to any one of claims 1 to 5, wherein the calculation unit calculates the maximum value by the average value of the maximum values of the cross-correlation coefficients obtained per frame.
7. The speaker group estimation device according to any one of claims 1 to 6, wherein the calculation unit calculates the maximum value of the whitening cross-correlation coefficient of the target speech interval.
8. The speaker group estimation device according to any one of claims 1 to 7, wherein the clustering processing unit forms a network graph consisting of the nodes and the edge weights, performs spectral clustering on the network graph, and divides the network graph into a plurality of clusters to estimate the plurality of groups.
9. The speaker group estimation device according to any one of claims 1 to 8, wherein the clustering processing unit divides the time series data into detection intervals having a predetermined time length, sets a sliding window of a predetermined time unit for each detection interval, and detects the point in time when the group has changed while sliding the sliding window in the time direction, and estimates a plurality of groups after the change.
Citation Information
Patent Citations
Apparatus, method and program detecting interaction
JP2008242318A
Information presentation device, method and program
JP2010128633A
Information processing device, information processing method, program, and information processing system
JP2012155374A
Voice processing device, voice processing method and voice processing program
JP2017062307A
Multi-channel speech segment estimation device
JP6887622B1