Online speaker successive discrimination method, online speaker successive discrimination device, and online speaker successive discrimination system

By updating the voice buffer, using time-continuous time frame blocks to store voice signals and speaker probability information, the accuracy of voice sequence distinction in the multi-speaker environment is solved, and online voice sequence distinction is achieved for infinite multi-speaker users.

JP7674231B2Active Publication Date: 2025-05-09HITACHI LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021208075
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-22
Publication Date
2025-05-09
Estimated Expiration
2041-12-22

AI Technical Summary

Technical Problem

Existing online voice sequence distinction techniques are difficult to accurately identify speakers when multiple speakers speak at the same time, and accuracy may be reduced when the number of speakers is not fixed.

Method used

By updating the voice buffer, using time-continuous time frame blocks to store voice signals and corresponding speaker probability information, realizing online voice sequence distinction for infinite number of speakers.

Benefits of technology

It realizes accurate identification of speakers when multiple speakers speak at the same time, and can handle the situation where the number of speakers is not fixed, improving the accuracy and flexibility of online voice sequence distinction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007674231000013
    Figure 0007674231000013
  • Figure 0007674231000014
    Figure 0007674231000014
  • Figure 0007674231000015
    Figure 0007674231000015
Patent Text Reader

Abstract

To provide online speaker sequentially discriminating means which can generate highly accurate speaker sequential discrimination even if a plurality of speakers speak at the same time without limiting the number of speakers, which can be outputted.SOLUTION: An on-line speaker sequential discrimination method includes the steps of: inputting a first voice signal; generating a first probability value indicating utterance probability for each speaker on a first time frame of the first voice signal; storing the first time frame and the first probability value in a selective utterance information buffer; inputting a second voice signal; generating a coupled voice signal from the first and second voice signals; generating a second probability value indicating utterance probability for each speaker on a second time frame of the coupled voice signal; determining a correspondence relation of the speakers in the first and second voice signals; selecting a subset of the second time frame which temporally continues and a subset of the second probability value and storing them in the selective utterance information buffer; and generating a speaker sequential discrimination result for identifying each speaker of user interaction.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an online speaker discrimination method, an online speaker discrimination device and an online speaker discrimination system. [Background technology]

[0002] In recent years, there has been an increasing need to utilize speech recognition in environments with multiple speakers, such as for information retrieval from news programs, transcribing meetings and telephone conversations, etc. In such environments, by solving a problem called speaker diarization, which determines "who spoke when," it is possible to assign various metadata labels, such as speaker order, characteristics (gender, age), and changes in audio source, to speech signals, making indexing and searching easier.

[0003] There are two types of speaker sequential differentiation: offline speech differentiation, in which the speech signal has been acquired and the entire speech signal to be analyzed is available from start to finish, and online speech differentiation, in which analysis is performed in real time as the speech signal is acquired. Several proposals have been made for these offline and online speech discrimination methods.

[0004] For example, Patent Document 1 discloses a technology in which "the speaker discrimination system 30 includes a memory unit 42 that stores speaker GMMs 74-78, a voice activity detection unit 30 that segments speech data, a novelty determination unit 34 that determines whether the current segment belongs to any of the speaker GMMs 74-78, a new model generation unit 40 that generates a new speaker GMM and labels the current segment with the new speaker GMM when the current segment does not belong to any of the speaker GMMs 74-78, a speaker identification unit 44 that identifies a speaker and labels the current segment with the speaker when the current segment belongs to one of the speaker GMMs 74-78, a training unit 48 that trains the speaker GMM using the current segment, and a merging unit 46 that merges segment labels according to the sequence of segments output by the voice activity detection unit 30."

[0005] Furthermore, Patent Document 2 discloses a technology in which "the online speaker successive distinction method uses a speech information buffer that stores important information (information useful for determining who spoke and when) extracted from previous voice data segments, thereby making it possible to generate accurate voice successive distinction results in real time even when voice data segments in which multiple different users' speech overlap occur in a dialogue. The information stored in the speech information buffer includes a specific time position in the voice data segment and the probability that a specific speaker is speaking, corresponding to that time position. In addition, the information stored in the speech information buffer may be determined by a predetermined selection method based on attention weight, absolute value of probability, pseudorandom number, etc."

[0006] In addition, Non-Patent Document 1 discloses a technology that "recently proposed a new speaker diarization method called end-to-end neural diarization vector clustering (EEND vector clustering), which integrates clustering-based and end-to-end neural network-based diarization approaches into one framework. The proposed method combines the advantages of both frameworks, namely high diarization performance and handling of overlapping speech based on EEND, and robust handling of long recordings with any number of speakers based on the clustering-based approach. However, this method has so far only been evaluated on simulated two-speaker conference-like data. In this paper, (1) we report recent progress we have made on this framework, including a newly introduced robust constrained clustering algorithm, and (2) we experimentally show that this method can significantly outperform competing diarization methods such as Encoder-Decoder Attractor (EDA)-EEND on CALLHOME data, which contains real conversational speech data with overlapping speech and any number of speakers."

[0007] In addition, Non-Patent Document 2 discloses a technology that "Attractor-based end-to-end diarization has achieved accuracy comparable to that of carefully tuned conventional clustering-based methods for difficult datasets. However, its main drawback is that it cannot handle the case where the number of speakers is larger than the number observed during training. This is because it relies on supervised learning. In this paper, we propose an unsupervised clustering process incorporated into attractor-based end-to-end diarization. First, we divide the sequence of frame-wise embeddings into short subsequences, and then perform attractor-based diarization for each subsequence. Given the diarization results for each subsequence, the speaker correspondence between subsequences is obtained by unsupervised clustering of vectors calculated from attractors from all subsequences. This allows us to generate diarization results for a large number of speakers across the entire recording, even when the number of output speakers for each subsequence is limited."

[0008] In addition, Non-Patent Document 3 discloses a technology that states, "In this paper, we propose a new end-to-end neural network-based speaker diarization method. Unlike most existing methods, our method does not use separate modules for speaker expression extraction and clustering. Instead, our model has a single neural network that directly outputs speaker diarization results. To realize such a model, we formulate the speaker diarization problem as a multi-label classification problem and introduce a permutation-free objective function to directly minimize the diarization error without suffering from speaker and label permutation problems. In addition to its end-to-end simplicity, our method has the advantage of being able to explicitly handle overlapping speech during training and inference. With this advantage, the model can be easily trained / adapted with real recorded multi-speaker conversations by simply inputting the corresponding multi-speaker segment labels."

[0009] In addition, Non-Patent Document 4 discloses a technology that "We propose a streaming diarization method that handles overlapping speech with a flexible number of speakers based on an end-to-end neural diarization (EEND) model. In previous research, a speaker trace buffer (STB) mechanism was proposed to realize chunk-wise streaming diarization using a pre-trained EEND model. STB traces speaker information of the previous chunk to map speakers in the new chunk. However, STB only worked with recordings of two speakers. In this paper, we propose FLEX-STB as an extended STB for a flexible number of speakers. In this method, the difference in the number of speakers between the buffer and the current chunk is mitigated by using zero padding followed by speaker tracing. We also consider a buffer update strategy to select important frames for tracing multiple speakers."

[0010] Furthermore, Non-Patent Document 5 discloses a technology that "We propose a new deep learning model that supports permutation invariant training (PIT) for speaker-independent multi-speaker speech separation, commonly known as the cocktail party problem. Unlike many conventional techniques that treat speech separation as a multiclass regression problem and deep clustering methods that consider it as a segmentation (or clustering) problem, our model optimizes the separation regression error while ignoring the order of the mixed sources. This strategy elegantly solves the long-standing label permutation problem that has hindered the progress of deep learning-based methods for speech separation. Experiments on an equal energy mixture setting of the Danish corpus confirm the effectiveness of PIT. We believe that improvements built on PIT will eventually solve the cocktail party problem and enable applications such as automatic speech-to-text transcription of meetings and multi-party human-computer interaction, where speech overlap is common." [Prior art documents] [Patent documents]

[0011]

Patent Document 1

Patent document 2

Non-licensed literature

[0012] [Non-licensed document 1] Kinoshita et al., “Advances in integration of end-to-end neural and clustering-based diarization for real conversational speech,” in Proc. INTERSPEECH 2021.

Non-licensed Document 2

Non-licensed Document 4

Non-licensed Document 5

[0013] The above-mentioned Patent Document 1 discloses a method for performing online speaker diarization by detecting speech segments online and estimating whether the detected speech segments are the speech of a speaker who has spoken in the past or a new speaker.

[0014] Furthermore, Patent Document 2 and Non-Patent Document 4 disclose a method of performing online diarization using a neural network that estimates speech and non-speech of multiple people for each time frame from speech features. Since the order of speakers output by the neural network is not fixed, it is necessary to prevent the order of speakers output from being changed during online processing. For this reason, Patent Document 2 and Non-Patent Document 4 disclose a method of correcting the output order of speakers to be consistent by using a buffer that stores previously input speech features and inference results.

[0015] In addition, Non-Patent Documents 1 and 2 disclose a method for estimating diarization results for an unlimited number of speakers as a whole by performing estimation for each short interval and then integrating the diarization results performed for each short interval, which is a problem that the number of speakers that can be output is limited in a neural network that estimates speech / non-speech of multiple people for each time frame from speech features.

[0016] However, in Patent Document 1, since each detected voice is assigned to one of the speakers, there is a problem that when multiple speakers speak simultaneously, the speaker of a specific utterance cannot be correctly estimated.

[0017] According to Patent Document 2, by using a buffer that stores previously input speech features and inference results, accurate speech discrimination results can be generated in real time even when multiple speakers speak simultaneously. However, since the buffer length cannot be infinite, it is necessary to select the time frame to store in the buffer using some method.

[0018] Patent Document 2 and Non-Patent Document 4 describe a method for selecting a time frame based on a first-in-first-out buffer, or attention weights or Kullback-Leibler divergence, and Non-Patent Document 4 discloses a method for selecting a time frame based on Kullback-Leibler divergence. However, since the neural network used in the diarization method of Non-Patent Document 4 has a limited number of output speakers, there is a possibility that the accuracy will decrease if it is used in an environment where the number of speakers is not known in advance. In addition, the neural network of Patent Document 2 also has a similar problem since the number of output speakers is limited.

[0019] In Non-Patent Documents 1 and 2, by allowing the number of output speakers to be limited for short sections such as 5 seconds, it is possible to estimate diarization results for an unlimited number of speakers while dealing with simultaneous speech by multiple speakers. However, since online processing cannot be performed, these methods cannot be used when real-time speaker diarization is required.

[0020] If we consider combining the diarization method of Non-Patent Document 1 or Non-Patent Document 2, which can handle an unlimited number of speakers, with the buffer disclosed in Patent Document 2 to realize online diarization for an unlimited number of speakers, the time frames stored in the buffer are selected on a time frame basis in Patent Document 2, so the time frames stored in the buffer are discontinuous. Therefore, the precondition of Non-Patent Documents 1 and 2 that "the number of output speakers can be limited for short sections" cannot be used.

[0021] Therefore, the present disclosure has been made in consideration of the above-mentioned problems, and aims to provide a means for sequentially distinguishing online speakers with an unlimited number of speakers by updating a speech information buffer in block units consisting of temporally consecutive time frames. [Means for solving the problem]

[0022] In order to solve the above problem, one representative online speaker successive discrimination method of the present invention includes the steps of: inputting a first speech signal corresponding to a user dialogue and consisting of a first set of time frames; generating a first speech processing result for the first speech signal, the first speech signal including a set of first probability values ​​indicating a probability that a specific speaker is speaking in each time frame of the first set of time frames; storing the set of first time frames and the set of first probability values ​​corresponding to the set of first time frames in a selective speech information buffer; inputting a second speech signal corresponding to the user dialogue; combining speech features corresponding to the set of first time frames stored in the selective speech information buffer with the second speech signal to generate a combined speech signal; generating a second speech processing result including a set of second probability values ​​indicating a probability that a speaker of the first speech signal is speaking in each time frame of the set of second time frames; determining correspondence information indicating a correspondence between a speaker in the first speech signal and a speaker in the second speech signal based on the first speech processing result and the second speech processing result; selecting a subset of second time frames that are part of the set of second time frames and are consecutive in time and a subset of second probability values ​​that are part of the set of second probability values ​​and correspond to the subset of second time frames based on a predetermined selection technique, and storing them in the selectable speech information buffer according to the correspondence information; and generating a speaker sequential discrimination result that identifies each of the speakers of the user dialogue based on the second speech processing result and the correspondence information. Effect of the Invention

[0023] According to the present disclosure, by updating the speech information buffer in units of blocks each consisting of a series of time frames that are consecutive in time, it is possible to provide a means for sequentially distinguishing online speakers with an unlimited number of speakers. Other objects, configurations and effects will become apparent from the following description of the preferred embodiment of the invention. [Brief description of the drawings]

[0024] [Figure 1] FIG. 1 illustrates a computer system for implementing an embodiment of the present disclosure. [Diagram 2] FIG. 2 is a diagram illustrating an example of a configuration of an online speaker successive discrimination system according to an embodiment of the present disclosure. [Diagram 3] FIG. 3 is a flowchart showing the flow of the online speaker successive discrimination method according to the first embodiment of the present disclosure. [Figure 4] FIG. 4 is a diagram illustrating a configuration example of the selective utterance information buffer according to the first embodiment of the present disclosure. [Diagram 5] FIG. 5 is a diagram illustrating an example of a method for updating the selective utterance information buffer according to the first embodiment of the present disclosure. [Figure 6] FIG. 6 is a flowchart illustrating a flow of the online speaker successive discrimination method according to the second embodiment of the present disclosure. [Figure 7] FIG. 7 is a diagram illustrating a configuration example of a selective utterance information buffer and a first-in first-out utterance information buffer according to the second embodiment of the present disclosure. [Figure 8] FIG. 8 is a diagram illustrating an example of a method for updating the selective utterance information buffer and the first-in first-out utterance information buffer according to the second embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0025] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. Note that the present invention is not limited to the embodiment. In addition, in the description of the drawings, the same parts are denoted by the same reference numerals. (Explanation of terms and symbols)

[0026] First, before describing the embodiments of the present disclosure, symbols and terms used in the present disclosure will be explained.

[0027] In this disclosure, the time-series audio signal

number

number

number

[0028] In addition, the posterior probability indicating whether the speaker is speaking in each time frame is

number

number

[0029] In addition, in this disclosure, the method described in Non-Patent Document 2 will be used for explanation as diarization for an unlimited number of speakers, so definitions and explanations of symbols will be provided here in advance. However, the application of this disclosure is not limited to Non-Patent Document 2, and may be, for example, the diarization method described in Non-Patent Document 1 or a method for tasks other than diarization, such as sound source separation described in Non-Patent Document 5. This method converts the time-series speech signal of [Equation 1] into a time-series embedding vector

number

number

[0030] Where:

number

number

number

[0031] Where:

number

number

[0032] However, if the length of the short interval T / L is sufficiently short, the number of people who will speak is likely to be small. l Finally, the diarization results P1, ...P L is integrated to obtain the diarization result for all frames as shown in [Equation 4]. For this integration, a clustering method such as those described in Non-Patent Documents 1 and 2 may be used. In this case, the estimated number of speakers S is S l Since the value can be larger than the above, Non-Patent Document 2 can output diarization results with no limit on the number of speakers.

[0033] Next, referring to FIG. 1, a computer system 100 for implementing the embodiments of the present disclosure will be described. The mechanisms and devices of the various embodiments disclosed herein may be applied to any suitable computing system. The main components of the computer system 100 include one or more processors 102, memory 104, terminal interface 112, storage interface 113, I / O (input / output) device interface 114, and network interface 115. These components may be interconnected via a memory bus 106, an I / O bus 108, a bus interface unit 109, and an I / O bus interface unit 110.

[0034] Computer system 100 may include one or more general purpose programmable central processing units (CPUs) 102A and 102B, collectively referred to as processors 102. In some embodiments, computer system 100 may include multiple processors, and in other embodiments, computer system 100 may be a single CPU system. Each processor 102 executes instructions stored in memory 104 and may include an on-board cache.

[0035] In some embodiments, memory 104 may include random access semiconductor memory, storage devices, or storage media (either volatile or non-volatile) for storing data and programs. Memory 104 may store all or part of the programs, modules, and data structures that implement the functions described herein. For example, memory 104 may store an online speaker discrimination application 150. In some embodiments, online speaker discrimination application 150 may include instructions or instructions for executing on processor 102 the functions described below.

[0036] In some embodiments, the online speaker discrimination application 150 may be implemented in hardware via semiconductor devices, chips, logic gates, circuits, circuit cards, and / or other physical hardware devices instead of or in addition to a processor-based system. In some embodiments, the online speaker discrimination application 150 may include data other than instructions or descriptions. In some embodiments, a camera, sensor, or other data input device (not shown) may be provided to communicate directly with the bus interface unit 109, the processor 102, or other hardware of the computer system 100.

[0037] Computer system 100 may include a bus interface unit 109 that provides communication between processor 102, memory 104, display system 124, and I / O bus interface unit 110. I / O bus interface unit 110 may couple to an I / O bus 108 for transferring data to and from various I / O units. I / O bus interface unit 110 may communicate via I / O bus 108 with multiple I / O interface units 112, 113, 114, and 115, also known as I / O processors (IOPs) or I / O adapters (IOAs).

[0038] Display system 124 may include a display controller, a display memory, or both. The display controller may provide video, audio, or both data to display device 126. Computer system 100 may also include one or more sensors or other devices configured to collect data and provide the data to processor 102.

[0039] For example, computer system 100 may include biometric sensors to collect heart rate data, stress level data, etc., environmental sensors to collect humidity data, temperature data, pressure data, etc., and motion sensors to collect acceleration data, movement data, etc. Other types of sensors may also be used. Display system 124 may be connected to a display device 126, such as a standalone display screen, a television, a tablet, or a handheld device.

[0040] The I / O interface unit provides the ability to communicate with various storage or I / O devices. For example, the terminal interface unit 112 may be attached to user I / O devices 116, such as user output devices, such as a video display, a television with speakers, and user input devices, such as a keyboard, a mouse, a keypad, a touchpad, a trackball, buttons, a light pen, or other pointing device. A user may use a user interface to enter input data or instructions to the user I / O devices 116 and the computer system 100, and receive output data from the computer system 100, by manipulating the user input devices. The user interface may be displayed on a display, played through speakers, or printed via a printer, for example, via the user I / O devices 116.

[0041] Storage interface 113 may be attached to one or more disk drives or direct access storage device 117 (usually a magnetic disk drive storage device, but may also be an array of disk drives or other storage devices configured to appear as a single disk drive). In one embodiment, storage device 117 may be implemented as any secondary storage device. Contents of memory 104 may be stored in storage device 117 and retrieved from storage device 117 as needed. I / O device interface 114 may provide an interface to other I / O devices such as printers, fax machines, etc. Network interface 115 may provide a communications path to allow computer system 100 and other devices to communicate with each other. This communications path may be, for example, network 130.

[0042] In some embodiments, computer system 100 may be a device that receives requests from other computer systems (clients) without a direct user interface, such as a multi-user mainframe computer system, a single-user system, or a server computer. In other embodiments, computer system 100 may be a desktop computer, a portable computer, a laptop, a tablet computer, a pocket computer, a telephone, a smartphone, or any other suitable electronic device.

[0043] Next, with reference to FIG. 2, a configuration of an online speaker successive discrimination system according to an embodiment of the present disclosure will be described.

[0044] Fig. 2 is a diagram illustrating an example of a configuration of an online speaker discriminator system 200 according to an embodiment of the present disclosure. As shown in Fig. 2, the online speaker discriminator system 200 mainly includes a voice signal acquiring device 210, a client terminal 220, a communication network 225, and an online speaker discriminator 230. The voice signal acquiring device 210, the client terminal 220, and the online speaker discriminator 230 are connected to each other via the communication network 225. The communication network 225 may include, for example, a local area network (LAN), a wide area network (WAN), a satellite network, a cable network, a WiFi network, or any combination thereof. Also, the connection between the voice signal acquiring device 210, the client terminal 220, and the online speaker successive distinguishing device 230 may be wired or wireless.

[0045] The voice signal acquiring device 210 is a device for acquiring a voice signal to be processed in a speaker successive discrimination method described later. The voice signal acquiring device 210 may be, for example, a computing device equipped with a microphone such as a smartphone or a personal computer, or a recorder. The voice signal acquired by the voice signal acquiring device 210 may be directly transmitted to the online speaker successive discrimination device 230 via a communication network 225, or may be transmitted to the online speaker successive discrimination device 230 via a client terminal 220.

[0046] The client terminal 220 is a terminal that transmits the voice signal acquired by the above-mentioned voice signal acquiring device 210 to the online speaker successive distinction device 230 via the communication network 225. In addition, the client terminal 220 may provide a GUI (Graphical User Interface) or the like for confirming the speaker successive distinction result transmitted from the online speaker successive distinction device 230. The client terminal 220 may be a terminal used by an individual, or may be a terminal shared by an organization such as a private company. The client terminal 220 may be any device, such as a desktop computer, a laptop computer, a tablet, a smartphone, etc. In some embodiments, the client terminal 220 and the audio signal capturing device 210 may be the same computing device.

[0047] The data storage unit 227 is a storage unit for storing the voice signal to be processed by the speaker successive distinction method, which is transmitted from the client terminal 220 or the voice signal acquiring device 210 via the communication network 225. This data storage unit 227 may be, for example, a local storage such as a hard disk drive (HDD) or a solid state drive (SSD), or may be a cloud-type storage area accessible to the online speaker successive distinction device 230. The data storage unit 227 may also store information to be stored in, for example, an utterance information buffer described later.

[0048] The online speaker discrimination device 230 is a device for implementing the processing of the speaker discrimination method according to an embodiment of the present disclosure. The online speaker discrimination device 230 can generate a speaker discrimination result for identifying each of the speakers of the user dialogue by processing the speech signals stored in the data storage unit 227. As described above, the speaker discrimination result may be transmitted to the client terminal 220. It should be noted that the user dialogue here refers to oral exchange between at least two speakers, and includes, for example, conversations, meetings, telephone calls, interviews, and the like.

[0049] Also, as shown in FIG. 2, the online speaker successive distinction device 230 includes a signal input unit 232, a voice processing result determination unit 234, a signal combining unit 236, a buffer management unit 238, and a result output unit 240 as functional units for implementing the speaker successive distinction method according to an embodiment of the present disclosure.

[0050] The signal input unit 232 is a functional unit that receives an input of a voice signal (for example, a first voice signal, a second voice signal, etc.) from the voice signal acquiring device 210 or the client terminal 220.

[0051] The speech processing result determination unit 234 is a functional unit that generates a speech processing result indicating a probability (for example, a first set of probability values ​​or a second set of probability values) that a specific speaker is speaking for each time frame (for example, a first set of time frames or a second set of time frames) in a speech signal to be processed. The speech processing result determination unit 234 may be, for example, an end-to-end speaker successive differentiation network (EEND) using a self-attention method, or may generate a speech processing result using a so-called speaker diarization method or sound source separation method.

[0052] The signal combining unit 236 is a functional unit that combines a plurality of audio signals (for example, a first audio signal and a second audio signal) to generate a combined audio signal.

[0053] The buffer management unit 238 is a functional unit for managing information stored in the speech information buffer according to the embodiment of the present disclosure. In the present disclosure, the process of storing new information in the speech information buffer is also referred to as "buffer update." As will be described later, in the embodiment of the present disclosure, two types of speech information buffers are used: a "selective speech information buffer" and a "first-in, first-out speech information buffer." Hereinafter, the selective speech information buffer and the first-in, first-out speech information buffer are collectively referred to as "speech information buffer." As will be described later, the buffer management unit 238 can also determine correspondence information indicating the correspondence between speakers in a plurality of audio signals.

[0054] The result output unit 240 is a functional unit that generates a speaker successive discrimination result for identifying each speaker in a user dialogue, based on the voice processing result generated by the voice processing result determination unit 234 and the speaker correspondence relationship information determined by the buffer management unit 238. This speaker successive discrimination result is information indicating the speech probability of each speaker for each time frame of the voice signal.

[0055] Each functional unit included in the online speaker successive discrimination device 230 may be a software module constituting the online speaker successive discrimination application 150 shown in Fig. 1, or may be an independent dedicated hardware device. Moreover, the above-mentioned functional units may be implemented in the same computing environment or in distributed computing environments. Furthermore, the embodiments of the present disclosure are not limited to the configuration of the online speaker discrimination system 200 described with reference to FIG. EXAMPLES

[0056] Next, an online speaker successive discrimination method according to a first embodiment of the present disclosure will be described with reference to Figs. 3 to 5. In the first embodiment, speech signals such as a first speech signal X1 and a second speech signal X2 arrive in units of a block length T / L (for example, in the case of Fig. 4, in units of 6 frames).

[0057] Fig. 3 is a flowchart showing the flow of an online speaker successive discrimination method 300 according to the first embodiment of the present disclosure. The online speaker successive discrimination method 300 shown in Fig. 3 is a method for performing online speaker successive discrimination capable of handling an unlimited number of speakers by using a selective speech information buffer that stores temporally consecutive time frames selected from a speech signal by a predetermined selection method, and may be executed by various functional units of the online speaker successive discrimination device 230 shown in Fig. 2, for example.

[0058] First, in step S305, the signal input unit 232 inputs a first audio signal X1 as an input. This first input signal X1 may be, for example, waveform data recorded using the above-mentioned audio signal acquisition device 210, or may be an audio signal that has been subjected to preprocessing such as dereverberation in advance. In addition, the first audio signal X1 may be a feature amount obtained by converting a waveform into a time-frequency domain using a Fourier transform or the like. The first input signal X1 is composed of a first set of time frames, which are sub-segments of the audio signal divided into any time units (1 ms, 10 ms, 1 sec, etc.). For example, dividing a 1 second audio signal into 100 ms sub-segments results in ten 100 ms sub-segments, each corresponding to a different time frame (1, 2, 3...10).

[0059] Next, in step S310, the speech processing result determination unit 234 generates a first speech processing result including a first set of probability values ​​P1 indicating the probability that a specific speaker is speaking for each time frame in the first speech signal X1 acquired in step S305. The calculation of the probability is performed, for example, by in1 This may be done by the trained EEND network described above based on More specifically, the first set of probability values ​​P1 represents the probability that a speaker, of any number of speakers participating in the user dialogue, is speaking for each time frame constituting the first speech signal X1. This probability may be expressed, for example, as a value between 0 and 1. In the following, the probability that a speaker is speaking is also referred to as "utterance probability."

[0060] Next, in step S315, the buffer management unit 238 stores the first audio signal X1 and a set of first probability values ​​P1 calculated for each of the sets of first time frames constituting the first audio signal X1 calculated in step S310 in a selective speech information buffer. The selective utterance information buffer according to the embodiment of the present disclosure is a storage area for temporarily storing information that is included in the speech signal and is useful for speaker sequential discrimination processing, selected by a predetermined selection method. More specifically, the information stored in the selective utterance information buffer includes information on a time frame in the speech signal (e.g., the first speech signal X1) and information on the speech probability calculated for the time frame.

[0061] Since the size of the selective speech information buffer is limited, in order to reduce the amount of information stored in the selective speech information buffer, it is desirable to store only a part of the information selected by a predetermined selection method, rather than the information of all time frames of the speech signal, depending on the length of the speech signal. This makes it possible to generate good speaker successive discrimination results while reducing the size of the selective speech information buffer. However, at the stage of step S315, since the selective utterance information buffer has not yet stored any information and is empty, it is possible to store all time frames in the first audio signal X1 until the selective utterance information buffer is filled without using the above-mentioned specified selection method. A selection method for selecting information to be stored in the selective utterance information buffer (that is, a method for updating the selective utterance information buffer) will be described later.

[0062] Next, in step S320, the signal input unit 232 inputs a second audio signal X2 as an input. This second audio signal X2 may be, for example, waveform data recorded using the above-mentioned audio signal acquisition device 210, similar to the first input signal X1, or may be an audio signal that has been subjected to preprocessing such as dereverberation, or may be a feature amount obtained by converting a waveform into a time-frequency domain using a Fourier transform or the like. The second voice signal X2 may be, for example, a voice signal following the first input signal X1 input in step S305 in a user dialogue.

[0063] Next, in step S325, the signal combining unit 236 combines the speech features of the first speech signal X1 stored in the selective speech information buffer in step S315 with the second speech signal X2 input in step S320 to generate a combined speech signal [X 1、 X2]. The combining of the audio signals here may be done by any known means.

[0064] Next, in step S330, the audio processing result determination unit 234 determines whether the combined audio signal [X 1、 a second set of probability values ​​[P m1、 P m2 ]. Now, the second set of probability values ​​[P m1、 P m2 ] is the first group of probability values ​​P m1 and a second group of probability values ​​P, which are the utterance probabilities calculated for X2. m2 The calculation of the utterance probabilities may be performed by a trained EEND network, similar to the first set of probability values ​​described above. Here, X1 is a set of time-continuous T / L speech signals, and X2 is also a time-continuous T / L speech signal, so the combined speech signal [X 1、X2] is also a set of speech signals with continuous time T / L. Therefore, the speaker diarization described in Non-Patent Document 1 and Non-Patent Document 2, which does not limit the number of output speakers, can be applied, and as a result, the second set of probability values ​​[P m1、 P m2 The second speech processing result including ] also does not impose a restriction on the number of output speakers.

[0065] Next, in step S335, the buffer management unit 238 determines correspondence information indicating the correspondence between the speaker in the first audio signal and the speaker in the second audio signal based on the first audio processing result and the second audio processing result. Here, in step S330, the combined signal [X 1、 Since the second speech processing result generated for the combined speech signal [X1, X2] and the first speech processing result generated for the first speech signal in step S310 are generated independently, the order of speakers is not fixed, and the order of speakers output during online processing may be interchanged. For this reason, the correspondence between the speakers in the first speech signal X1 and the speakers in the second speech signal X2 may become unclear. In other words, 1、 X2] corresponds to speaker A of the first speech signal X1 (i.e., the same person) or speaker B (i.e., a different person). This is known as the so-called "permutation problem" in neural networks.

[0066] Therefore, to produce consistent speaker-by-speaker discrimination results, the combined speech signal [X 1、 It is necessary to associate the speaker of the first speech signal X1 with the speaker of the second speech signal X2. 1、In order to solve the above-mentioned permutation problem by associating the speaker of the first speech signal X1 with the speaker of the first speech signal X2, a selective speech information buffer according to an embodiment of the present disclosure is used. A first set of probability values ​​P1, which are speech probabilities calculated for a previous speech signal (e.g., the first speech signal X1) stored in the selective speech information buffer, and a second set of probability values ​​[P m1、 P m2 ], it is possible to match the speaker.

[0067] More specifically, the buffer management unit 238 divides the first probability value set P1 corresponding to each time frame in the first speech signal X1 stored in the selective utterance information buffer into the first probability value group P2 generated in step S330. m1 and comparing the subset of first probability values ​​corresponding to the first speaker in the first set of probability values ​​P1 with the group of first probability values ​​P m1 If the correlation with the subgroup of the first probability value corresponding to the second speaker satisfies a predetermined correlation standard, correspondence relationship information indicating that the first speaker and the second speaker are the same speaker can be generated. In other words, the buffer manager 238 divides the first set of probability values ​​P1 and the first group of probability values ​​P m1 In both cases, speakers having similar probability values ​​can be determined to be the same person. Here, the buffer management unit 238 may determine the correspondence between speakers using a correlation coefficient described in Non-Patent Document 4, for example.

[0068] Next, in step S340, the buffer management unit 238 calculates the combined audio signal [X 1、A subset of second time frames which are part of the set of second time frames constituting [X2] and are consecutive in time, and a subset of second probability values ​​which are part of the set of second probability values ​​and correspond to the subset of second time frames are selected based on a predetermined selection method, and stored in a selective utterance information buffer in accordance with the correspondence relationship information determined in step S335. Here, the expression "contiguous in time" means that each time frame is connected in order without any gap between them. For example, when an audio signal is divided into 1-second time frames, the set of time frames "1, 2, 3, 4" can be said to be contiguous in time, but the sets of time frames "1, 2, 4" and "1, 2, 3, 5" cannot be said to be contiguous in time.

[0069] More specifically, as mentioned above, since the size of the selective speech information buffer is limited, in order to reduce the amount of information stored in the buffer, depending on the length of the speech signal, it is desirable to store only a portion of the information selected by a predetermined selection method, rather than the information of all time frames of the speech signal. 1、 X2] and a second set of probability values ​​[P m1、 P m2 ] is longer than the buffer length L of the selective utterance information buffer, a subset of time frames to be stored in the buffer (a subset of second time frames) is selected from the set of second time frames, and the selected subset of second time frames and a subset of second probability values ​​corresponding to the subset of second time frames are stored in the selective utterance information buffer. Incidentally, the method of updating the utterance information buffer will be described with reference to FIG. 5, and therefore the description thereof will be omitted here.

[0070] In addition, here, the expression "storing in a selective speech information buffer according to the correspondence information" may include modifying the order or labels of speakers to match the speakers corresponding to the second subset of time frames and the second subset of probability values ​​with the speaker determined to be the same person based on the correspondence information.

[0071] Next, in step 345, the result output unit 240 generates a speaker successive discrimination result for identifying each speaker in the second audio signal X2 based on the second audio processing result and the correspondence relationship information. More specifically, this speaker successive discrimination result is consistent with previously output audio processing results (i.e., the same identification label is assigned to the same speaker) by correcting the order and labels of the speakers in the second audio processing result based on the correspondence relationship information.

[0072] In the above-described first embodiment, the description has been given on the assumption that the audio signal arrives in units of block length T / L, but when the audio signal is longer than the block length, the present embodiment can be applied as is by regarding multiple blocks as arriving simultaneously. Also, when the audio signal arrives in units shorter than the block length, the embodiment of the present disclosure can be applied by waiting until a signal of the block length arrives before performing processing.

[0073] According to the online speaker sequential distinction method 300 according to the first embodiment of the present disclosure described above, past speech signals and the speech probabilities generated for the speech signals are stored in a selective speech information buffer in units of blocks consisting of temporally consecutive time frames, so that speaker diarization can be performed without an upper limit on the number of speakers even when multiple speakers speak simultaneously.

[0074] Next, a configuration of the selective utterance information buffer according to the first embodiment of the present disclosure will be described with reference to FIG.

[0075] 4 is a diagram showing a configuration example of the selective utterance information buffer 400 according to the first embodiment of the present disclosure. As described above, the selective utterance information buffer 400 according to the embodiment of the present disclosure stores in the data storage unit 227 the first speech signal X1410 and the set of first probability values ​​P1420, which are the utterance probabilities for each speaker S calculated for the first speech signal X1410.

[0076] 4, the selective utterance information buffer 400 is divided into a set of blocks 425 (blocks 1 to 4) that store temporally consecutive time frames of the speech signal, and the block length of each block is set to, for example, the length T / L of the short interval used in Non-Patent Document 2. The selective utterance information buffer 400 shown in FIG. 4 shows a case where the buffer length T=24, the block length T / L=6, and the number of speakers S=3.

[0077] The speech signals and corresponding speech probabilities in each block of the selective utterance information buffer 400 are assumed to be continuous in time, but the blocks do not have to be continuous in time. The method of updating the selective utterance information buffer 400 will be described later.

[0078] Next, a selective utterance information buffer updating method according to the first embodiment of the present disclosure will be described with reference to FIG.

[0079] 5 is a diagram illustrating an example of a selective utterance information buffer updating method 500 according to the first embodiment of the present disclosure. As described above, in the selective utterance information buffer updating according to the first embodiment, the buffer is updated in units of blocks each consisting of a time frame that is continuous in time, thereby making it possible to distinguish an unlimited number of online speakers one by one. In the example of the selective utterance information buffer update method 500 shown in Fig. 5, a case will be described in which the selective utterance information buffer already stores four blocks and a new voice signal is input. In this case, four blocks to be stored in the selective utterance information buffer are selected from the four blocks (block 1 to block 4) originally existing in the selective utterance information buffer and block 5 of the newly input voice signal.

[0080] At this time, in order to select blocks to be stored in the selective speech information buffer, the buffer management unit 238 can calculate an evaluation value for each time frame by using, for example, attention weights (see Patent Document 2) or Kullback-Leibler divergence (see Non-Patent Document 4), and then calculate an evaluation value for each block by averaging the evaluation values ​​for each time frame within the block.

[0081] Thereafter, the buffer management unit 238 stores the blocks that satisfy the predetermined evaluation threshold in the selective utterance information buffer, and does not store the blocks that do not satisfy the predetermined evaluation threshold in the selective utterance information buffer. For example, referring to FIG. 5, if blocks 1, 3, 4, and 5 meet a predetermined evaluation threshold and block 2 does not meet the predetermined evaluation threshold, blocks 1, 3, 4, and 5 are stored in the selective speech information buffer and block 2 is removed from the selective speech information buffer.

[0082] According to the selective speech information buffer updating method 500 described above, the buffer is updated in units of blocks consisting of consecutive time frames, so that the condition that "the number of output speakers can be limited for short sections" is satisfied, and speaker diarization can be performed without an upper limit on the number of speakers. EXAMPLES

[0083] Next, a method for sequentially distinguishing online speakers according to a second embodiment of the present disclosure will be described with reference to FIGS.

[0084] In the online speaker successive discrimination method according to the first embodiment of the present disclosure described above, it is assumed that the speech signal arrives in units of block length or in units longer than that. Also, when the speech signal arrives in units shorter than the block length, it is assumed that the method waits until the number of samples of the block length is accumulated before executing the method. In this case, the time it takes to wait until the number of samples of the block length is accumulated may become the delay time of the entire online speaker successive discrimination system. However, for example, Non-Patent Document 1 shows that the highest accuracy is achieved when the block length is set to 30 seconds, and Non-Patent Document 2 shows that the block length is set to 5 seconds, and these values ​​are by no means short as a processing unit for the online speaker successive discrimination method. On the other hand, as Non-Patent Document 1 also shows, the longer the block length itself is, the higher the accuracy of online speaker successive discrimination is.

[0085] In view of the above, a method for sequentially distinguishing online speakers according to a second embodiment of the present disclosure can provide a method for sequentially distinguishing online speakers with reduced delay while maintaining a long block length. In the second embodiment, the speech signals such as the first speech signal X1 and the second speech signal X2 arrive in units T / LN (where N>1) that are smaller than the block length T / L units (for example, in the case of FIG. 7, in units of 6 frames). For example, in the case of FIG. 6, when N=2, the speech signals arrive in units of 3 frames, and when N=3, the speech signals arrive in units of 2 frames. Even if T / LN is not an integer, the online speaker successive discrimination method according to the second embodiment of the present disclosure can be implemented by, for example, allowing the expansion and contraction of the first-in, first-out speech information buffer. However, in the following, for convenience of explanation, the case where T / LN is an integer will be explained as an example.

[0086] The system configuration for implementing the online speaker successive distinction method according to the second embodiment of the present disclosure is similar to the online speaker successive distinction system 200 shown in the first embodiment described above, and therefore the description thereof will be omitted.

[0087] 6 is a flowchart showing a flow of an online speaker successive discrimination method 600 according to a second embodiment of the present disclosure. The online speaker successive discrimination method 600 shown in FIG. 6 is a method for performing online speaker successive discrimination with reduced delay by using a first-in, first-out utterance information buffer in addition to the above-mentioned selective utterance information buffer, and may be executed by various functional units of the online speaker successive discrimination device 230 shown in FIG. 2, for example. For convenience of explanation, the online speaker sequential distinction method 600 will be described on the assumption that the first voice signal X1 and the first probability value set P1 have already been stored in the selective utterance information buffer. The process for storing the first voice signal X1 in the selective utterance information buffer has been described in steps S305 to S315 shown in FIG. 3, and therefore will not be described here.

[0088] First, in step S605, the signal input unit 232 inputs a second speech signal X2 having a length T / LN as an input. This second speech signal X2 may be, for example, waveform data recorded using the above-mentioned speech signal acquisition device 210, like the first input signal X1, or may be a speech signal that has been subjected to preprocessing such as dereverberation, or may be a feature amount obtained by converting a waveform into a time-frequency domain using a Fourier transform or the like. Also, the second voice signal X2 may be, for example, a voice signal subsequent to the first voice signal X1 that has already been stored in the selective utterance information buffer.

[0089] Next, in step S610, the buffer management unit 238 stores the second voice signal X2 in the first-in first-out speech information buffer. In addition, the first-in, first-out speech information buffer is a so-called FIFO buffer that maintains the chronological order of information, processing and outputting information that arrives earlier first and information that arrives later than the earlier information, so the second audio signal X2 stored in the first-in, first-out speech information buffer is continuous in time.

[0090] Next, in step S615, the signal combining unit 236 combines the speech features of the first speech signal X1 stored in the selective speech information buffer with the second speech signal X2 stored in the first-in-first-out speech information buffer in step S610 to generate a combined speech signal [X 1、 X2]. The combining of the audio signals here may be done by any known means. After that, the audio processing result determination unit 234 calculates the combined audio signal [X 1、a second set of probability values ​​[P m1、 P m2 ], where the second set of probability values ​​[P m1、 P m2 ] is a group P of first probability values, which are utterance probabilities calculated for the first speech signal X1 stored in the selective utterance information buffer. m1 and a second group P of probability values, which are utterance probabilities calculated for the second speech signal X2 stored in the first-in-first-out utterance information buffer. m2 Includes.

[0091] Next, in step S620, the buffer management unit 238 determines correspondence information indicating the correspondence between the speaker in the first audio signal X1 and the speaker in the second audio signal X2 based on the first audio processing result and the second audio processing result. More specifically, the buffer management unit 238 divides the first probability value set P1 corresponding to the first time frame of the first voice signal X1 stored in the selective utterance information buffer into a first probability value set P2 corresponding to the first time frame of the first voice signal X1 stored in the selective utterance information buffer, and the first probability value group P3 corresponding to the first time frame of the first voice signal X1 stored in the selective utterance information buffer, m1 and comparing the subset of first probability values ​​corresponding to the first speaker in the first set of probability values ​​P1 with the group of first probability values ​​P m1 If the correlation with the subgroup of the first probability value corresponding to the second speaker satisfies a predetermined correlation standard, correspondence relationship information indicating that the first speaker and the second speaker are the same speaker can be generated. In other words, the buffer manager 238 divides the first set of probability values ​​P1 and the first group of probability values ​​P m1 In both cases, speakers having similar probability values ​​can be determined to be the same person.

[0092] Next, in step S625, the buffer management unit 238 determines whether or not to update the selective utterance information buffer by the above-mentioned selection method based on whether or not a predetermined update condition is satisfied for the selective utterance information buffer. The update condition here is a condition for managing the timing of updating the selective utterance information buffer. In one aspect of the present disclosure, the update condition may be, for example, a condition that specifies that an update is permitted once for N inputs. In another aspect of the present disclosure, the update condition may be a condition that permits an update only when the voice signal stored in the selective utterance information buffer and the voice signal stored in the first-in-first-out utterance information buffer fall below a predetermined overlapping degree standard (for example, the overlapping degree is 5% or less). This makes it possible to prevent signals and utterance probability values ​​of the same time frame from being stored in the selective utterance information buffer in a duplicated manner. If the update condition is met, the process proceeds to step S630; if the update condition is not met, the process proceeds to step S635.

[0093] Even if the update condition is not satisfied in step S625, only the first set of probability values ​​P1 stored in the selective utterance information buffer may be updated. For example, the first set of probability values ​​P1 may be updated by dividing the first set of probability values ​​P1 and the first group of probability values ​​P m1 By overwriting with the average value of and, the accuracy of the first set of probability values ​​P1 calculated for the speech signal in the selective speech information buffer is improved due to the effect of data augmentation during general inference, thereby improving the accuracy of determining the speaker correspondence.

[0094] In step S630, the buffer management unit 238 calculates the combined audio signal [X 1、 X2], and a subset of the second time frames that are contiguous in time and that constitute the second set of time frames [P m1、 P m2] and a second subset of probability values ​​corresponding to the second subset of time frames are selected based on a predetermined selection method and stored in the selective utterance information buffer according to the correspondence relationship information determined in step S620. Incidentally, the method of updating the selective utterance information buffer here has been explained with reference to, for example, FIG. 5 above, and therefore the explanation thereof will be omitted here.

[0095] In step S635, the result output unit 240 generates a speaker successive discrimination result for identifying each speaker in the second audio signal X2 based on the second audio processing result and the correspondence relationship information. More specifically, this speaker successive discrimination result is consistent with previously output audio processing results (i.e., the same identification label is assigned to the same speaker) by correcting the second audio processing result based on the correspondence relationship information.

[0096] In the online speaker successive discrimination method according to the second embodiment of the present disclosure described above, by introducing a first-in, first-out speech information buffer consisting of one block, it is possible to reduce delays in online speaker successive discrimination for speech signals in units shorter than the block length. In addition, in the online speaker successive distinction method according to Example 2 of the present disclosure, similar to Example 1 of the present disclosure described above, the selective utterance information buffer that stores past speech signals and utterance probabilities is updated in units of blocks consisting of consecutive time frames, making it possible to perform online speaker successive distinction without an upper limit on the number of speakers, even when multiple speakers speak simultaneously.

[0097] Next, configurations of a selective utterance information buffer and a first-in first-out utterance information buffer according to the second embodiment of the present disclosure will be described with reference to FIG.

[0098] FIG. 7 is a diagram illustrating a configuration example of the selective utterance information buffer 720 and the first-in first-out utterance information buffer 740 according to the second embodiment of the present disclosure. As described above, the second embodiment of the present disclosure is configured to use two types of utterance information buffers, the selective utterance information buffer 720 and the first-in first-out utterance information buffer 740, unlike the first embodiment which uses only a selective utterance information buffer.

[0099] Similar to the selective utterance information buffer 400 described above in the first embodiment, the selective utterance information buffer 720 is selected by a predetermined selection method and updated in units of blocks each made up of temporally continuous time frames. However, the first-in, first-out speech information buffer 740 is an utterance information buffer that maintains the time sequence of information, processing and outputting information that arrives earlier first, and processing and outputting information that arrives later than the earlier information.

[0100] More specifically, in the example shown in Fig. 7, the selective utterance information buffer 720 includes three blocks, blocks 1 to 3. These three blocks, blocks 1 to 3, are updated by a sampling method based on, for example, attention weights (see Patent Document 2) or Kullback-Leibler divergence (see Non-Patent Document 4). Meanwhile, the first-in-first-out speech information buffer 740 contains one block, block 4. This block 4 is updated in a first-in-first-out manner.

[0101] In Figure 7, a selective speech information buffer 720 including three blocks and a first-in, first-out speech information buffer 740 including one block are described as examples, but as described above, the buffer length and block length of each speech information buffer may be set appropriately. Moreover, a method for updating the selective utterance information buffer 720 and the first-in first-out utterance information buffer 740 will be described later.

[0102] 7 shows an example in which the utterance probability (P2) calculated for a time frame is not held in the first-in-first-out utterance information buffer 740, but the utterance probability calculated for a time frame may be stored in the first-in-first-out utterance information buffer 740, as in the selective utterance information buffer 720. This allows the results of the first T(N-1) / LN frames in the first-in-first-out utterance information buffer 740 to be used when determining correspondence information indicating the correspondence between speakers (for example, step S620 in FIG. 6), thereby improving the accuracy of determining the correspondence between speakers.

[0103] Next, a method for updating the selective utterance information buffer and the first-in first-out utterance information buffer according to the second embodiment of the present disclosure will be described with reference to FIG.

[0104] 8 is a diagram showing an example of a method 800 for updating a selective utterance information buffer and a first-in, first-out utterance information buffer (hereinafter, also referred to as an "utterance information buffer updating method 800") according to a second embodiment of the present disclosure. As described above, in the utterance information buffer updating method 800 according to the second embodiment, by using the first-in, first-out utterance information buffer 740 in addition to the selective utterance information buffer 720, it is possible to perform online speaker sequential discrimination with reduced delay.

[0105] In an example of the speech information buffer updating method 800 shown in Fig. 8, a case will be described in which the selective utterance information buffer 720 already stores three blocks, and the first-in, first-out utterance information buffer 740 stores one block. In this case, three blocks to be saved in the selective utterance information buffer 720 are selected from the three blocks originally existing in the selective utterance information buffer 720 and the one block existing in the first-in, first-out utterance information buffer 740, and stored.

[0106] At this time, in order to select blocks to be stored in the selective utterance information buffer 720, the buffer management unit 238 uses, for example, attention weights (see Patent Document 2) or Kullback-Leibler divergence (see Non-Patent Document 4) to calculate an evaluation value for each time frame in the selective utterance information buffer 720 and the first-in, first-out utterance information buffer 740, and then calculates an evaluation value for each block by averaging the evaluation values ​​for each time frame within the block.

[0107] Then, the buffer management unit 238 stores in the selective utterance information buffer 720 those blocks that satisfy a predetermined evaluation threshold from the blocks stored in the selective utterance information buffer 720 and the blocks stored in the first-in, first-out utterance information buffer 740, and does not store in the selective utterance information buffer 720 those blocks that do not satisfy the predetermined evaluation threshold. For example, referring to FIG. 7, if blocks 1, 3, and 4 meet a predetermined evaluation threshold and block 2 does not meet the predetermined evaluation threshold, blocks 1, 3, and 4 are stored in the selective speech information buffer 720 and block 2 is removed from the selective speech information buffer 720.

[0108] According to the speech information buffer updating method 800 described above, the buffer is updated in units of blocks consisting of consecutive time frames, so that the condition that "the number of output speakers can be limited for short sections" is satisfied, and speaker diarization can be performed without an upper limit on the number of speakers. Furthermore, in the speech information buffer updating method 800, by using the first-in, first-out speech information buffer 740 in addition to the selective speech information buffer 720, it is possible to reduce delays in online speaker successive discrimination for speech signals in units shorter than the block length.

[0109] Although the embodiment of the present invention has been described above, the present invention is not limited to the above-described embodiment, and various modifications are possible without departing from the gist of the present invention. [Explanation of symbols]

[0110] 200 Online speaker discrimination system 210 Audio signal acquisition device 220 Client terminals 225 Communication Network 227 Data Storage Unit 230 On-line speaker discrimination device 232 Signal input section 234 Audio processing result judgment unit 236 Signal Combination Section 238 Buffer Management Unit 240 Result output section

Claims

1. 1. A method for online speaker discrimination, comprising: inputting a first speech signal corresponding to a user interaction and comprising a first set of time frames; generating, for the first speech signal, a first speech processing result comprising a first set of probability values ​​indicative of a probability that a particular speaker is speaking in each time frame of the first set of time frames; storing the first set of time frames and a first set of probability values ​​corresponding to the first set of time frames in a selective utterance information buffer; inputting a second speech signal corresponding to the user interaction; combining speech features stored in the selective speech information buffer corresponding to the first set of time frames with a second speech signal to generate a combined speech signal; generating, for a second set of time frames in the combined speech signal, a second speech processing result comprising a second set of probability values ​​indicative of a probability that a particular speaker is speaking in each time frame of the second set of time frames; determining correspondence information indicating a correspondence between a speaker in the first audio signal and a speaker in the second audio signal based on the first audio processing result and the second audio processing result; selecting a subset of second time frames that are part of the set of second time frames and are consecutive in time, and a subset of second probability values ​​that are part of the set of second probability values ​​and correspond to the subset of second time frames based on a predetermined selection method, and storing the selected utterance information in the selective utterance information buffer according to the correspondence information; generating a speaker discrimination result for identifying each of the speakers of the user dialogue based on the second speech processing result and the correspondence information; 13. A method for online speaker discrimination comprising:

2. The second set of probability values ​​comprises: a first group of probability values ​​corresponding to speech features of the first speech signal; a second group of probability values ​​corresponding to speech features of the second speech signal; The step of determining correspondence information indicating a correspondence between speakers includes: generating correspondence information indicating that the first speaker and the second speaker are the same speaker when a correlation between a subset of first probability values ​​corresponding to a first speaker among the set of first probability values ​​and a subgroup of first probability values ​​corresponding to a second speaker among the group of first probability values ​​satisfies a predetermined correlation criterion; The method of claim 1, further comprising:

3. The step of storing the selective utterance information in the selective utterance information buffer includes: calculating an evaluation value for each block including the second set of time frames and the second set of probability values ​​that are temporally consecutive, using a predetermined evaluation method; If the evaluation value calculated for a particular block satisfies a predetermined evaluation threshold, the particular block is stored in the selective utterance information buffer. The method of claim 1, further comprising:

4. The evaluation method includes: and calculating, for each of the second set of time frames, an attention weight of a neural network as the evaluation value. The method of claim 3, further comprising:

5. The evaluation method includes: A method of calculating a Kullback-Leibler divergence as the evaluation value for each of the second probability values. The method of claim 3, further comprising:

6. The online speaker successive discrimination method includes: storing the second speech signal in a first-in, first-out speech information buffer; The step of generating the combined audio signal comprises: combining speech features corresponding to the first set of time frames stored in the selective speech information buffer with the second speech signal stored in the first-in-first-out speech information buffer; The method of claim 1, wherein the speaker is a first speaker.

7. The step of storing the selective utterance information in the selective utterance information buffer includes: This is performed only when the voice signal stored in the selective speech information buffer and the voice signal stored in the first-in first-out speech information buffer are below a predetermined overlapping degree criterion. The method of claim 6, wherein the speaker discrimination is performed sequentially.

8. The first audio processing result and the second audio processing result are Generated using speaker diarization techniques, The method of claim 1, further comprising:

9. The first audio processing result and the second audio processing result are Generated using a sound source separation method, The method of claim 1, further comprising:

10. An online speaker discrimination apparatus comprising: a signal input unit for inputting a voice signal corresponding to a user dialogue; a speech processing result determination unit that generates, for the speech signal, a speech processing result including a set of probability values ​​indicating a probability that a specific speaker is speaking in each time frame included in the speech signal; a buffer management unit for managing information stored in the selective utterance information buffer and the first-in, first-in, first-out utterance information buffer; a signal combining unit for combining a plurality of audio signals; a result output unit for generating a speaker-sequential discrimination result for identifying each of the speakers of the user dialogue; The signal input unit includes: inputting a first speech signal corresponding to the user interaction and comprising a first set of time frames; The audio processing result determination unit generating, for the first speech signal, a first speech processing result comprising a first set of probability values ​​indicative of a probability that a particular speaker is speaking in each time frame of the first set of time frames; The buffer management unit storing the first set of time frames and a first set of probability values ​​corresponding to the first set of time frames in a selective speech information buffer; The signal input unit includes: inputting a second audio signal corresponding to the user interaction; The signal coupling unit combining speech features stored in the selective speech information buffer corresponding to the first set of time frames with a second speech signal to generate a combined speech signal; The audio processing result determination unit generating, for a second set of time frames in the combined speech signal, a second speech processing result comprising a second set of probability values ​​indicative of a probability that a particular speaker is speaking in each time frame of the second set of time frames; The buffer management unit determining correspondence information indicating a correspondence between a speaker in the first audio signal and a speaker in the second audio signal based on the first audio processing result and the second audio processing result; selecting a subset of second time frames that are part of the set of second time frames and are continuous in time, and a subset of second probability values ​​that are part of the set of second probability values ​​and correspond to the subset of second time frames based on a predetermined selection method, and storing the selected subset of second probability values ​​in the selective utterance information buffer according to the correspondence information; The result output unit is generating a speaker discrimination result for identifying each of the speakers of the user dialogue based on the second speech processing result and the correspondence relationship information; 13. An online speaker discrimination device comprising:

11. an audio signal acquisition device that acquires an audio signal; A client terminal; An online speaker successive discrimination system connected to an online speaker successive discrimination device that performs online speaker successive discrimination via a communication network, The online speaker discrimination device comprises: a signal input unit for inputting a voice signal corresponding to a user dialogue from the voice signal acquisition device; a speech processing result determination unit that generates, for the speech signal, a speech processing result including a set of probability values ​​indicating a probability that a specific speaker is speaking in each time frame included in the speech signal; a buffer management unit for managing information stored in the selective utterance information buffer and the first-in, first-in, first-out utterance information buffer; a signal combining unit for combining a plurality of audio signals; a result output unit that generates a speaker sequential discrimination result for identifying each of the speakers of the user dialogue and transmits the result to the client terminal; The signal input unit includes: inputting a first speech signal corresponding to the user interaction and comprising a first set of time frames; The audio processing result determination unit generating, for the first speech signal, a first speech processing result comprising a first set of probability values ​​indicative of a probability that a particular speaker is speaking in each time frame of the first set of time frames; The buffer management unit storing the first set of time frames and a first set of probability values ​​corresponding to the first set of time frames in a selective utterance information buffer; The signal input unit includes: inputting a second audio signal corresponding to the user interaction; The signal coupling unit combining speech features stored in the selective speech information buffer corresponding to the first set of time frames with a second speech signal to generate a combined speech signal; The audio processing result determination unit generating, for a second set of time frames in the combined speech signal, a second speech processing result comprising a second set of probability values ​​indicative of a probability that a particular speaker is speaking in each time frame of the second set of time frames; The buffer management unit determining correspondence information indicating a correspondence between a speaker in the first audio signal and a speaker in the second audio signal based on the first audio processing result and the second audio processing result; selecting a subset of second time frames that are part of the set of second time frames and are continuous in time, and a subset of second probability values ​​that are part of the set of second probability values ​​and correspond to the subset of second time frames based on a predetermined selection method, and storing the selected subset of second probability values ​​in the selective utterance information buffer according to the correspondence information; The result output unit is generating a speaker discrimination result for identifying each of the speakers of the user dialogue based on the second speech processing result and the correspondence relationship information; An online speaker discrimination system comprising:

Citation Information

Patent Citations

  • System for sequentially distinguishing online speaker and computer program thereof

    JP2009109712A

  • On-line speaker sequential distinguishing method, on-line speaker sequential distinguishing device, and on-line sequential speaker distinguishing system

    JP2021131524A

  • Voice analysis device, voice classification method, and voice classification program

    WO2008126627A1