Speaker splitting method and system

By identifying a central feature vector with the highest confidence level and repeatedly processing residual signals, the method addresses speaker overlap and noise issues in voice recognition, achieving accurate speaker segmentation.

JP7870404B2Active Publication Date: 2026-06-04NAVER CORP

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NAVER CORP
Filing Date
2023-12-19
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing voice recognition systems face challenges in accurately segmenting speakers when noise is present or when multiple speakers overlap, leading to contaminated speaker feature vectors and errors in segmentation results.

Method used

A method and system that extracts a central feature vector with the highest confidence level from a set of speaker feature vectors, allowing for accurate identification of a specific speaker's speech, even in noisy or overlapping conditions, by repeatedly performing speaker feature extraction and central feature vector determination on residual signals.

Benefits of technology

This approach reduces the number of overlapping speaker sections and enhances the accuracy of speaker recognition by extracting reliable speaker feature vectors, ensuring high-quality speaker segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007870404000001
    Figure 0007870404000001
  • Figure 0007870404000002
    Figure 0007870404000002
  • Figure 0007870404000003
    Figure 0007870404000003
Patent Text Reader

Abstract

The present disclosure relates to a speaker segmentation method performed by at least one processor, the speaker segmentation method including the steps of: extracting a first set of speaker feature vectors based on speech segments included in an input signal; extracting a first central feature vector having the highest reliability based on the extracted first set of speaker feature vectors; and extracting speech of a first speaker associated with the first central feature vector from the input signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0006] , ,

[0001] The present disclosure relates to a speaker segmentation method and system. Specifically, it relates to a speaker segmentation method and system that extracts a central feature vector having the highest reliability among speaker feature vectors and extracts the voice of a specific speaker associated with the extracted central feature vector.

Background Art

[0002] Due to the development of IT technology and voice recognition technology, various voice recognition-related services are provided. For example, there are provided a meeting record generation service that automatically generates a meeting record based on a meeting recording file, and an automatic subtitle generation service that automatically generates subtitles for dramas, movies, videos, etc.

[0003] In services including such a voice recognition function, when the recognition target voice data includes voices of multiple speakers, speaker segmentation may be performed to distinguish and recognize the utterance sections of the multiple speakers. However, there is a problem that if noise is mixed in the recognition target voice or there is a section where the voices of multiple speakers overlap, the speaker feature vector will be contaminated, and an error may occur in the speaker segmentation result.

Summary of the Invention

Problems to be Solved by the Invention

[0004] The present disclosure provides a speaker segmentation method, a computer-readable non-temporary recording medium recording instruction words, and an apparatus (system) for solving the above problems.

Means for Solving the Problems

[0005] The present disclosure may be implemented in various ways including a method, an apparatus (system), or a computer-readable non-temporary recording medium recording instruction words.

[0006] According to one embodiment of the present disclosure, a speaker segmentation method performed by at least one processor includes the steps of: extracting a first set of speaker feature vectors based on speech segments included in an input signal; extracting a first central feature vector having the highest confidence level based on the extracted first set of speaker feature vectors; and extracting the speech of a first speaker associated with the first central feature vector from the input signal.

[0007] A computer-readable non-temporary recording medium is provided, which stores instructions for executing a method according to one embodiment of the present disclosure on a computer.

[0008] An information processing system according to one embodiment of the present disclosure includes a memory and at least one processor connected to the memory and configured to execute at least one computer-readable program contained in the memory, the at least one program including instructions for extracting a first set of speaker feature vectors based on speech segments contained in an input signal, extracting a first central feature vector having the highest confidence based on the extracted first set of speaker feature vectors, and extracting the speech of a first speaker associated with the first central feature vector from the input signal. [Effects of the Invention]

[0009] According to some embodiments of this disclosure, by extracting a central feature vector that is highly reliable in identifying a specific speaker from among the speaker feature vectors extracted from the speech segment, it is possible to extract a representative speaker feature vector that accurately reflects the speech characteristics of a specific speaker, even if a portion of the speaker feature vectors is contaminated.

[0010] According to some embodiments of this disclosure, by repeatedly performing a series of processes—extracting the central feature vector and extracting the associated speech of other specific speakers—on the residual signal from which the speech of a specific speaker associated with the central feature vector has been removed from the input signal, the number of sections in which the speeches of multiple speakers overlap or are closely spaced during the speaker recognition and speech extraction process is reduced, and as a result, a speaker feature vector can be extracted that can identify one or more speakers with higher confidence.

[0011] The effects of this disclosure are not limited to those mentioned above, and any other effects not mentioned above would be clearly understood by a person with ordinary skill in the art to which this disclosure pertains ("ordinary skill") from the claims. [Brief explanation of the drawing]

[0012] Embodiments of the present disclosure will be described below with reference to the accompanying drawings, in which similar reference numerals indicate similar elements, but are not limited thereto. [Figure 1] This figure shows an example of a speaker splitting method according to one embodiment of the present disclosure. [Figure 2] This is a schematic diagram showing an information processing system, according to one embodiment of the present disclosure, in which an information processing system is connected to multiple user terminals in a communicative manner. [Figure 3] This is a block diagram showing the internal configuration of a user terminal and an information processing system according to one embodiment of the present disclosure. [Figure 4] This figure illustrates an example of extracting multiple speaker feature vectors based on speech segments included in an input signal, according to one embodiment of the present disclosure. [Figure 5] This figure shows an example of performing clustering on speaker feature vectors according to one embodiment of the present disclosure. [Figure 6] This figure shows an example of extracting a central feature vector based on a plurality of extracted speaker feature vectors according to one embodiment of the present disclosure. [Figure 7] This figure shows an example of extracting the voice of a specific speaker from an input signal according to one embodiment of the present disclosure. [Figure 8] This figure shows an example of a target voice extraction model according to one embodiment of the present disclosure. [Figure 9] This figure shows an example of a method for training a target speech extraction model according to one embodiment of the present disclosure. [Figure 10] This flowchart illustrates an example of a speaker splitting method according to one embodiment of the present disclosure. [Modes for carrying out the invention]

[0013] The specific details for implementing this disclosure will be described below with reference to the attached drawings. However, in the following explanation, if there is a risk of obscuring the essence of this disclosure, specific explanations of widely known functions and configurations will be omitted.

[0014] In the attached drawings, identical or corresponding components are denoted by the same reference numeral. Furthermore, in the following descriptions of embodiments, redundant descriptions of identical or corresponding components may be omitted. However, the omission of a description of a component does not mean that that component is not included in any of the embodiments.

[0015] The advantages and features of the disclosed embodiments, as well as methods for achieving them, will become clear from the embodiments described below, together with the accompanying drawings. However, this disclosure is not limited to the embodiments disclosed below and may be embodied in various other forms; these embodiments are merely meant to complete the disclosure, and the disclosure is provided to fully inform a person of the ordinary skill of the scope of the invention.

[0016] The terms used in this specification will be briefly explained, and the disclosed embodiments will be specifically described. The terms used in this specification are selected as general terms that are currently widely used as much as possible while considering the functions in this disclosure. However, this may change depending on the intentions or precedents of those skilled in the relevant art, the emergence of new technologies, etc. Also, in certain cases, there are terms arbitrarily selected by the applicant, and in such cases, the meaning thereof will be described in detail in the part explaining the corresponding invention. Therefore, the terms used in this disclosure should be defined based not on the simple name of the terms but on the meaning the terms have and the entire content of this disclosure.

[0017] The singular expressions in this specification include plural expressions unless specifically identified as singular in the context. Also, plural expressions include singular expressions unless specifically identified as plural in the context. Throughout the specification, when a certain part is said to include a certain component, this means, unless otherwise stated, not to exclude other components but rather to possibly further include other components.

[0018] Furthermore, the terms “module” or “part” as used in this specification refer to components of software or hardware, and a “module” or “part” may perform either role. However, this does not mean that a “module” or “part” is limited to software or hardware. A “module” or “part” may be configured to reside on an addressable storage medium, or to regenerate one or more processors. Thus, as an example, a “module” or “part” may include components such as software components, object-oriented software components, class components, and task components, and at least one of the following: processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, or variables. The components and the functions provided within a “module” or “part” may be combined with a smaller number of components and “modules” or “parts,” or further separated into additional components and “modules” or “parts.”

[0019] According to one embodiment of the present disclosure, a "module" or "unit" may be embodied as a processor and a memory. The "processor" should be broadly interpreted to include general-purpose processors, central processing units (CPUs), microprocessors, digital signal processors (DSPs), controllers, microcontrollers, state machines, and the like. In some environments, the "processor" may refer to an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field-programmable gate array (FPGA), or the like. The "processor" may refer to a combination of processing devices such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors coupled with a DSP core, or a combination of any other such configuration. Also, the "memory" should be broadly interpreted to include any electronic component capable of storing electronic information. The "memory" may refer to various types of processor-readable media such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage devices, registers, and the like. If the processor can read information from and / or record information in the memory, the memory is in an electronic communication state with the processor. Memory integrated with the processor is in an electronic communication state with the processor.

[0020] In the present disclosure, a "system" may include, but is not limited to, at least one of a server device and a cloud device. For example, the system may be composed of one or more server devices. As another example, the system may be composed of one or more cloud devices. As yet another example, the system may be configured and operated with both a server device and a cloud device.

[0021] In this disclosure, “machine learning model” may include any model used to infer an answer to a given input. According to one embodiment, the machine learning model may include an artificial neural network model comprising an input layer, a plurality of hidden layers, and an output layer, where each layer may contain a plurality of nodes. In this disclosure, the machine learning model may mean an artificial neural network model, and the artificial neural network model may mean a machine learning model.

[0022] In this disclosure, “Display” may mean any display device associated with a computing device, for example, any display device capable of displaying any information / data controlled by or provided by a computing device.

[0023] In this disclosure, “each of the A's” or “each of the A's” may refer to each of all the components included in the A's, or to each of some of the components included in the A's.

[0024] In this disclosure, “confidence” may refer to the probability that a particular speaker feature vector is estimated to represent the speech characteristics of a particular speaker. Therefore, in this disclosure, the speaker feature vector with the highest confidence may refer to the speaker feature vector that is estimated to best represent the speech characteristics of a particular speaker, or the speaker feature vector that is estimated to best reflect the speech characteristics of a particular speaker. The confidence of a speaker feature vector may be measured by various measures. According to one embodiment, in a vector space where multiple speaker feature vectors extracted from a particular input signal are located, a particular speaker feature vector may have a higher confidence the more densely packed the surrounding region of that speaker feature vector is.

[0025] In this disclosure, “central feature vector” may refer to the speaker feature vector with the highest confidence level among a plurality of speaker feature vectors extracted from a particular input signal.

[0026] Figure 1 shows an example of a speaker splitting method according to one embodiment of the present disclosure. The speaker splitting method can analyze an input signal 110 containing the voices of multiple speakers, divide and recognize the speech intervals of each speaker, and extract the voice of the respective speaker. First, the information processing system can receive the input signal 110. Here, the input signal 110 may include the voices of multiple speakers and background noise.

[0027] Subsequently, the information processing system can detect speech segments from the input signal 110 (120). Here, a speech segment may be a segment in the input signal 110 that contains the speaker's utterance. If a speech segment is detected, the information processing system can extract multiple speaker feature vectors based on at least a portion of the input signal 110 corresponding to the detected speech segment (130). The process by which the information processing system detects speech segments from the input signal 110 and extracts multiple speaker feature vectors will be described in detail later with reference to Figure 4.

[0028] The information processing system can extract the central feature vector with the highest confidence level based on the extracted speaker feature vectors (140). Here, the specific speaker may refer to any one speaker who uttered the speech associated with the central feature vector, rather than a predetermined speaker from among the multiple speakers whose utterances are included in the input signal 110. Therefore, the central feature vector with the highest confidence level among the multiple speaker feature vectors may refer to a speaker feature vector that is estimated to best represent the speech features of any one speaker, even though it is not possible to specifically identify or recognize which speaker it is from among the multiple speakers whose utterances are included in the input signal 110.

[0029] According to one embodiment, if information related to the speaker can be obtained, a speaker identification (SID) process may be performed further. For example, if some speakers' voice samples, or speaker feature information extracted from some speakers' voice samples, are obtained in advance, the information processing system can estimate which speaker the speaker feature vector extracted from the input signal 110 is associated with by performing a further speaker identification process. In this case, the central feature vector may be used to refer to the speaker feature vector that is estimated to best represent the voice features of a speaker that has been determined or identified in advance from among the multiple speakers included in the input signal 110. However, in the following description, "specific speaker" will be used to refer to any one speaker among the multiple speakers whose utterances are included in the input signal. The process by which the information processing system extracts the central feature vector with the highest confidence level based on the multiple speaker feature vectors extracted will be described in detail later with reference to Figure 6. As described above, by extracting the most reliable central feature vector based on multiple speaker feature vectors, it is possible to extract a representative speaker feature vector that accurately reflects the speech characteristics of a particular speaker, even if some of the speaker feature vectors are contaminated.

[0030] An information processing system that has extracted a central feature vector can extract the speech of a specific speaker associated with the central feature vector 152 from the input signal 110 (150). For example, the information processing system can extract the speech of a specific speaker associated with the central feature vector from the input signal 110 using a target speech extraction model. The process by which the information processing system extracts the speech of a specific speaker associated with the central feature vector from the input signal 110 will be described in detail later with reference to Figures 7 and 8.

[0031] Furthermore, the information processing system can determine whether or not a speech segment remains in the residual signal 154, which is the remaining signal after removing the speech 152 of a specific speaker extracted from the input signal 110 (160). If it is determined that a speech segment remains in the residual signal 154, the speech of other speakers included in the input signal 110 can be extracted by repeating the above-described series of processes.

[0032] For example, an information processing system can extract the voices of other speakers by repeatedly performing processes such as speaker feature vector extraction 130, central feature vector extraction 140, and specific speaker voice extraction 150 on the remaining voice segments of the input signal 110, excluding the segment containing the extracted voice 152 of a specific speaker.

[0033] As another example, an information processing system can extract the voices of other speakers by repeatedly performing processes such as central feature vector extraction 140 and specific speaker voice extraction 150 on some of the speaker feature vectors extracted based on the input signal 110, excluding the segment containing the voice 152 of a specific speaker.

[0034] As yet another example, the information processing system can take the residual signal 154 as the input signal 110 and repeatedly perform the series of processes described above (120, 130, 140, 150, etc.).

[0035] When the process described above is repeated until there are no more speech segments remaining in the residual signal 154, the speech of all speakers whose utterances are included in the input signal 110 can be separated and recognized.

[0036] As described above, first, by using the most reliable speaker feature vector to extract the speech 152 of a specific speaker from the input signal 110, the speech 152 of a specific speaker can be separated with high accuracy. Subsequently, by repeating a series of processes on the residual signal 154 from which the speech 152 of a specific speaker has been removed from the input signal 110, the sections in which the speeches of multiple speakers overlap or are closely spaced gradually decrease during this process. As a result, highly reliable speaker feature vectors can be extracted even for the remaining speakers whose speech is included in the residual signal 154, enabling high-quality speaker segmentation.

[0037] The speaker segmentation method of this disclosure will be described in more detail below with reference to Figures 2 to 10. For the sake of explanation, the speaker associated with the speech from which the central feature vector is first extracted from the input signal 410 will be referred to as the first speaker. Similarly, the speaker associated with the speech from which the nth central feature vector is extracted will be referred to as the nth speaker.

[0038] In the description above, it is stated that the speaker splitting method is performed by an information processing system, but it is not limited to this, and at least some of the steps included in the speaker splitting method of this disclosure may be performed by a user terminal. For example, a series of processes related to the speaker splitting method of this disclosure (120, 130, 140, 150, etc.) may be performed by a user terminal. However, for the sake of explanation, the speaker splitting method of this disclosure will be described below assuming that it is performed by an information processing system.

[0039] Figure 2 is a schematic diagram showing a configuration in which an information processing system 230 is communicably connected to a plurality of user terminals 210_1, 210_2, 210_3 according to one embodiment of the present disclosure. As shown in the figure, the plurality of user terminals 210_1, 210_2, 210_3 may be connected via a network 220 to an information processing system 230 that can provide speaker splitting services, speech recognition services, or various applications using speaker splitting and / or speech recognition functions (e.g., an automatic meeting transcript generation application, an automatic subtitle playback application, etc.). Here, the plurality of user terminals 210_1, 210_2, 210_3 may include terminals of users to whom speaker splitting services and / or speech recognition services are provided. In one embodiment, the information processing system 230 may include one or more server devices and / or databases, or one or more cloud computing service-based distributed computing devices and / or distributed databases, that can store, provide and execute computer executable programs (e.g., downloadable applications) and data related to speaker splitting services and / or speech recognition services, etc.

[0040] As an alternative, user terminals 210_1, 210_2, and 210_3 can also provide speaker splitting services and / or speech recognition services to users using computer-executable programs (e.g., downloadable applications) related to speaker splitting services and speech recognition services, without using the network 220.

[0041] The speaker splitting service and / or speech recognition service provided by the information processing system 230 may be provided to users through speaker splitting applications, speech recognition applications, automatic meeting record generation applications, automatic subtitle generation applications, speech editing applications, mobile browser applications, or web browsers installed on each of the multiple user terminals 210_1, 210_2, and 210_3. For example, the information processing system 230 can provide information corresponding to speaker splitting requests, speech recognition requests, meeting record generation requests, subtitle generation requests, etc., received from user terminals 210_1, 210_2, and 210_3 through the speaker splitting application, and perform corresponding processing.

[0042] Multiple user terminals 210_1, 210_2, and 210_3 can communicate with the information processing system 230 via the network 220. The network 220 may be configured to enable communication between the multiple user terminals 210_1, 210_2, and 210_3 and the information processing system 230. Depending on the installation environment, the network 220 may consist of wired networks such as Ethernet (registered trademark), Power Line Communication, telephone line communication equipment and RS-serial communication, mobile communication networks, wireless networks such as WLAN (Wireless LAN), Wi-Fi (registered trademark), Bluetooth (registered trademark), and ZigBee (registered trademark), or a combination thereof. The communication method is not limited, and not only communication methods that utilize communication networks that the network 220 may include (for example, mobile communication networks, wired internet, wireless internet, broadcasting networks, satellite networks, etc.) may also be included, as may short-range wireless communication between user terminals 210_1, 210_2, and 210_3.

[0043] In Figure 2, a mobile phone terminal 210_1, a tablet terminal 210_2, and a PC terminal 210_3 are shown as examples of user terminals, but the user terminals 210_1, 210_2, and 210_3 may be any computing device capable of wired and / or wireless communication, on which applications such as speaker splitting, speech recognition, automatic meeting transcript generation, automatic caption generation, audio editing, mobile browser applications, or web browsers can be installed and run. For example, user terminals may include AI speakers, smartphones, mobile phones, navigation systems, computers, laptops, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablet PCs, game consoles, wearable devices, IoT (Internet of Things) devices, VR (virtual reality) devices, AR (augmented reality) devices, set-top boxes, and the like. Furthermore, although Figure 2 shows that three user terminals 210_1, 210_2, and 210_3 communicate with the information processing system 230 via the network 220, the system is not limited to this, and other numbers of user terminals may be configured to communicate with the information processing system 230 via the network 220.

[0044] According to one embodiment, the information processing system 230 can receive input signals containing utterances from multiple speakers from multiple user terminals 210_1, 210_2, and 210_3. Subsequently, the information processing system 230 can extract multiple speaker feature vectors based on the speech segments included in the input signal, and extract the central feature vector with the highest confidence level based on the extracted speaker feature vectors. Subsequently, the information processing system 230 can extract the speech of a specific speaker associated with the central feature vector from the input signal. The information processing system 230 can separate the speech of each speaker included in the input signal by repeating the above process multiple times. Additionally, the information processing system 230 can transmit the separated speech of each speaker and / or the results of further processing of each speaker's speech to the user terminal 210 as the result of speaker segmentation. For example, the information processing system can generate meeting minutes or subtitles based on the separated voices of each speaker, using methods such as automatic speech recognition (ASR) or speech to text (STT), and transmit them to the user terminal 210.

[0045] Figure 3 is a block diagram showing the internal configuration of a user terminal 210 and an information processing system 230 according to one embodiment of the present disclosure. The user terminal 210 may refer to any computing device capable of running applications such as speaker splitting applications, speech recognition applications, automatic meeting transcript generation applications, automatic subtitle generation applications, audio editing applications, mobile browser applications, or web browsers, and capable of wired / wireless communication. For example, it may include the mobile phone terminal 210_1, tablet terminal 210_2, and PC terminal 210_3 shown in Figure 2. As shown, the user terminal 210 may include a memory 312, a processor 314, a communication module 316, and an input / output interface 318. Similarly, the information processing system 230 may include a memory 332, a processor 334, a communication module 336, and an input / output interface 338. As shown in Figure 3, the user terminal 210 and the information processing system 230 may be configured to communicate information and / or data via a network 220 using their respective communication modules 316 and 336. Furthermore, the input / output device 320 may be configured to input information and / or data to the user terminal 210 via the input / output interface 318, or to output information and / or data generated from the user terminal 210.

[0046] The memories 312,332 may include any non-temporary computer-readable recording medium. According to one embodiment, the memories 312,332 may include permanent mass storage devices such as RAM (random access memory), ROM (read-only memory), disk drives, SSDs (solid-state drives), and flash memory. As another example, permanent mass storage devices such as ROM, SSDs, flash memory, and disk drives may be included in the user terminal 210 or information processing system 230 as separate permanent storage devices distinct from the memories. The memories 312,332 may also store an operating system and at least one program code (for example, code installed in the user terminal 210 for a speaker splitting application).

[0047] Such software components may be loaded from a computer-readable storage medium separate from memory 312,332. Such a separate computer-readable storage medium may include a storage medium that can be directly connected to such a user terminal 210 and information processing system 230, but may also include computer-readable storage media such as floppy disks, disks, tapes, DVD / CD-ROM drives, and memory cards. As another example, software components may be loaded into memory 312,332 via a communication module rather than a computer-readable storage medium. For example, at least one program may be loaded into memory 312,332 based on a computer program installed by a file provided via the network 220 by a developer or a file distribution system that distributes application installation files.

[0048] Processors 314,334 may be configured to process computer program instructions by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to processors 314,334 by memory 312,332 or by communication modules 316,336. For example, processors 314,334 may be configured to execute instructions received by program code stored in a recording device such as memory 312,332.

[0049] Communication modules 316 and 336 can provide configurations or functions for the user terminal 210 and the information processing system 230 to communicate with each other via the network 220, and can provide configurations or functions for the user terminal 210 and / or the information processing system 230 to communicate with other user terminals or other systems (for example, a separate cloud system). For example, requests or data (e.g., speaker splitting requests, speech recognition requests, meeting transcript generation requests, subtitle generation requests, etc.) generated by program code stored in a recording device such as memory 312 by the processor 314 of the user terminal 210 may be transmitted to the information processing system 230 via the network 220 under the control of the communication module 316. Conversely, control signals and commands provided under the control of the processor 334 of the information processing system 230 may be received by the user terminal 210 via the communication module 336 and the network 220 through the communication module 316 of the user terminal 210. For example, the user terminal 210 can receive, via the communication module 316, the audio of each speaker separated from the input signal and / or the results of additional processing on each speaker's audio (e.g., meeting minutes, subtitles, etc.) as a result of speaker splitting from the information processing system 230.

[0050] The input / output interface 318 may be a means for interface with the input / output device 320. For example, the input device may include a camera including an audio sensor and / or image sensor, a keyboard, a microphone, a mouse, etc., and the output device may include a display, a speaker, a haptic feedback device, etc. As another example, the input / output interface 318 may be a means for interface with a device in which the configuration or function for input and output is integrated into one, such as a touchscreen. For example, a service screen, which is configured using information and / or data provided by the information processing system 230 or another user terminal when the processor 314 of the user terminal 210 processes instructions of a computer program loaded into memory 312, may be displayed on the display via the input / output interface 318. Figure 3 shows an example in which the input / output device 320 is not included in the user terminal 210, but is not limited to this, and may be configured as a single device with the user terminal 210. Furthermore, the input / output interface 338 of the information processing system 230 may be means for interface with an input or output device (not shown) that is connected to or included by the information processing system 230. In Figure 3, the input / output interfaces 318 and 338 are shown as elements configured separately from the processors 314 and 334, but the system is not limited to this, and the input / output interfaces 318 and 338 may be configured to be included in the processors 314 and 334.

[0051] The user terminal 210 and the information processing system 230 may include more components than those shown in Figure 3. However, it is not necessary to clearly illustrate most of the conventional components. In one embodiment, the user terminal 210 may be implemented to include at least some of the input / output devices 320 described above. The user terminal 210 may also include other components such as a transceiver, a GPS (Global Positioning system) module, a camera, various sensors, and a database. For example, if the user terminal 210 is a smartphone, it may include components that are generally included in a smartphone, such as an accelerometer, a gyroscope, a camera module, various physical buttons, buttons using a touch panel, input / output ports, and a vibrator for vibration. In one embodiment, the processor 314 of the user terminal 210 may be configured to run an application that provides speaker splitting services. In this case, code related to the application and / or program may be loaded into the memory 312 of the user terminal 210.

[0052] While a program for a speaker splitting application or the like is running, the processor 314 can receive input or selected text, images, videos, audio, and / or actions through input devices such as a touchscreen, keyboard, camera including audio sensors and / or image sensors, and microphones, which are connected to the input / output interface 318. The received text, images, videos, audio, and / or actions can be stored in memory 312 or provided to the information processing system 230 via the communication module 316 and network 220. For example, the processor 314 can receive user input requesting speaker splitting and provide it to the information processing system 230 via the communication module 316 and network 220. As another example, the processor 314 can receive input indicating a user's selection to an input signal containing utterances from multiple speakers and provide it to the information processing system 230 via the communication module 316 and network 220.

[0053] The processor 314 of the user terminal 210 may be configured to manage, process, and / or store information and / or data received from the input device 320, other user terminals, the information processing system 230, and / or multiple external systems. The information and / or data processed by the processor 314 may be provided to the information processing system 230 via the communication module 316 and the network 220. The processor 314 of the user terminal 210 can transmit and output information and / or data to the input / output device 320 via the input / output interface 318. For example, the processor 314 can display the received information and / or data on the screen of the user terminal.

[0054] The processor 334 of the information processing system 230 may be configured to manage, process, and / or store information and / or data received from multiple user terminals 210 and / or multiple external systems. The information and / or data processed by the processor 334 can be provided to the user terminals 210 via the communication module 336 and the network 220. According to one embodiment, the information processing system 230 can extract multiple speaker feature vectors based on the speech segments contained in the input signal received from the user terminals 210, and extract the central feature vector with the highest confidence level based on the extracted speaker feature vectors. Subsequently, the information processing system 230 can extract the speech of a specific speaker associated with the central feature vector from the input signal. The information processing system 230 can separate the speech of each speaker contained in the input signal by repeating the above process multiple times. In addition, the information processing system 230 can provide the user terminals 210 with the speech of each separated speaker and / or the results of additional processing of each speaker's speech as speaker segmentation. For example, the information processing system can generate meeting minutes or subtitles based on the separated voices of each speaker by performing automatic speech recognition (ASR) or speech-to-telephony (STT), and transmit them to the user terminal 210.

[0055] The processor 334 of the information processing system 230 may be configured to output processed information and / or data through output devices 320 such as a display output device (e.g., touchscreen, display, etc.) or an audio output device (e.g., speaker) of the user terminal 210. For example, the processor 334 of the information processing system 230 may be configured to provide the separated voices of each speaker to the user terminal 210 via the communication module 336 and network 220 as a result of speaker splitting, and to output the voices of each speaker through the sound output device of the user terminal 210. As another example, the processor 334 of the information processing system 230 may be configured to provide the user terminal 210 via the communication module 336 and network 220 with a meeting transcript or subtitles generated by performing automatic speech recognition (ASR) or STT, etc., based on the separated voices of each speaker, and to output them through the display output device of the user terminal 210.

[0056] Figure 4 shows an example of extracting multiple speaker feature vectors based on speech segments included in an input signal according to one embodiment of the present disclosure. According to one embodiment, an information processing system or a processor of an information processing system can receive an input signal 410. Here, the input signal 410 may include the speech of multiple speakers (and ambient noise). Here, the speaker may be a concept that includes not only humans, but also virtual people or characters capable of synthesized speech, software or hardware modules capable of generating and outputting sounds including speech, etc. For example, the input signal 410 may include speech recorded during a meeting in which many people participated. As another example, the input signal 410 may include speech extracted from video such as a drama or movie.

[0057] The input signal 410 may be sound in various forms. For example, the input signal 410 may be sound in waveform form or sound in spectrogram form. According to one embodiment, the information processing system can receive sound in waveform form, perform frequency conversion based on the received sound in waveform form, and use the sound converted into spectrogram form as the input signal 410.

[0058] Subsequently, the information processing system can detect a speech segment from the input signal 410 (420). Here, the speech segment may be the segment in the input signal 110 that contains the speaker's utterance. For example, the information processing system can detect the speech segment included in the input signal 410 by performing End Point Detection (EPD) to detect the start and end points of the utterance.

[0059] When a speech segment is detected, the information processing system can extract multiple speaker feature vectors 432 based on the detected speech segment (430). The speaker feature vectors may include information about the characteristics of the spoken speech contained in the speech segment. That is, the speaker feature vectors may include information about the speaker who uttered the speech.

[0060] According to one embodiment, the information processing system can extract multiple speaker feature vectors 432 based on detected speech segments by extracting one speaker feature vector for each unit speech segment. For example, assuming an embodiment where the length of a unit speech segment is 0.5 seconds, if the length of the detected first speech segment 422 is 1 second, the information processing system can extract two speaker feature vectors (E1 and E2 in the illustrated example) based on the first speech segment 422. Also, if the length of the detected second speech segment 424 is 1.5 seconds, the information processing system can extract three speaker feature vectors (E3, E4, and E5 in the illustrated example) based on the second speech segment 424.

[0061] The information processing system can extract multiple speaker feature vectors 432 from the detected speech segments using a speech segment detection and speaker feature vector extraction method based on unit speech segments of various lengths.

[0062] According to one embodiment, the information processing system can extract a plurality of speaker feature vectors 432 based on speech segments using a speaker embedding extractor. Here, the speaker embedding extractor may be a machine learning model that has been trained to extract speaker feature vectors based on speech segments. For example, the speaker embedding extractor may be a machine learning model that has been trained to extract identical / similar speaker feature vectors from speech segments containing the speech of the same speaker, and to extract speaker feature vectors that make it easy to distinguish the feature differences of each speaker from speech segments containing the speech of different speakers.

[0063] According to another embodiment, the information processing system can extract multiple speaker feature vectors 432 by processing each speech segment (e.g., unit speech segment) detected from the input signal 410. For example, the information processing system can extract multiple speaker feature vectors 432 by calculating the average value of each speech segment in spectrogram form.

[0064] Figure 5 shows an example of performing clustering on speaker feature vectors according to one embodiment of the present disclosure. According to one embodiment, the information processing system can perform clustering on a plurality of speaker feature vectors extracted from speech segments included in the input signal. When clustering is performed on a plurality of speaker feature vectors, similar feature vectors among the plurality of speaker feature vectors may be grouped. For example, the information processing system can perform clustering using any clustering method such as K-means clustering or spectral clustering. According to one embodiment, the information processing system can estimate the number of clusters from the extracted plurality of speaker feature vectors and group the plurality of speaker feature vectors into the estimated number of clusters.

[0065] Conventional speaker segmentation methods require a clustering process to separate the speech of each speaker included in the input signal; however, according to the speaker segmentation method of the present disclosure, the clustering process may be performed selectively.

[0066] Figure 5 shows the results of clustering two sets of speaker feature vectors extracted based on two input signals containing the speech of three speakers. The first clustering result 520 is the result of clustering the first set of speaker feature vectors 510 extracted from the speech segments included in the first input signal containing the speech of three speakers. In explaining the first clustering result 520, it can be confirmed that the first set of speaker feature vectors 510 were grouped into three clusters 522, 524, and 526. Speaker feature vectors belonging to the same cluster can be presumed to have been extracted from speech segments containing the speech of the same speaker.

[0067] The second clustering result 540 is the result of clustering a second set of speaker feature vectors 530 extracted from speech segments included in a second input signal containing the speech of three speakers. The second input signal may include an overlapping speech segment 532 in which the speech of the first and second speakers overlap. In this case, the speaker feature vector 534 extracted based on the overlapping speech segment 532 may be contaminated and may not accurately reflect the characteristics of any one of the three speakers. For example, the speaker feature vector 534 extracted based on the overlapping speech segment 532 may be located between the region where the speaker feature vector extracted from the speech segment containing only the speech of the first speaker is located and the region where the speaker feature vector extracted from the speech segment containing only the speech of the second speaker is located. As a result, the speaker feature vector 534 extracted based on the overlapping speech segment 532 may obscure the distinction between the region of the speaker feature vector corresponding to the first speaker and the region of the speaker feature vector corresponding to the second speaker.

[0068] The second clustering result 540, which is the result of clustering the second set of speaker feature vectors 530 extracted based on the second input signal, can be explained as follows: Although the second input signal actually contains the utterances of three speakers, the number of speakers was incorrectly estimated, and the second set of speaker feature vectors 530 was grouped into two clusters 542 and 544.

[0069] As described above, if the input signal contains duplicate speech segments 532, the duplicate speech segments 532 can contaminate the speaker feature vector, potentially leading to errors in the speaker segmentation results. Furthermore, contamination of the speaker feature vector can occur not only when the input signal contains duplicate speech segments 532, but also when the input signal contains noise such as ambient noise.

[0070] According to one embodiment of this disclosure, a representative speaker feature vector for a particular speaker can be extracted by extracting the central feature vector with the highest confidence level based on a plurality of extracted speaker feature vectors. Therefore, speaker segmentation can be performed without errors even if there are errors in the clustering results, or even if clustering is not performed. The process of extracting the central feature vector with the highest confidence level based on a plurality of extracted speaker feature vectors will be explained with reference to Figure 6.

[0071] Figure 6 shows an example of extracting a central feature vector based on a plurality of extracted speaker feature vectors according to one embodiment of the present disclosure. Figure 6 shows an example of a first set of speaker feature vectors 600 extracted based on an input signal containing the speech of three speakers (speaker A, speaker B, and speaker C). The first set of speaker feature vectors 600 may include a first group 610 containing a plurality of speaker feature vectors extracted based on an audio segment containing only speaker A's speech, a second group 620 containing a plurality of speaker feature vectors extracted based on an audio segment containing only speaker B's speech, a third group 630 containing a plurality of speaker feature vectors extracted based on an audio segment where the speech of speakers A and B overlaps, and a fourth group 640 extracted based on an audio segment containing only speaker C's speech. Thus, when the input signal includes a section where the speech of two or more speakers overlaps, clustering the speaker feature vectors may result in the first group 610, the second group 620, and the third group 630 all being classified into the same cluster. For this reason, if speaker segmentation is performed based solely on the clustering results using conventional methods, errors may occur in the speaker segmentation results.

[0072] According to one embodiment of the present disclosure, an information processing system can extract a central feature vector with the highest confidence level based on a first set of extracted speaker feature vectors 600. According to one embodiment, the information processing system can use the density of speaker feature vectors as a measure of confidence level.

[0073] For example, an information processing system can first determine the central region as the area in the vector space where speaker feature vectors are most densely concentrated. As a concrete example, the information processing system can determine the central region as the area containing the largest number of speaker feature vectors among multiple areas of the same size centered on each of the first set of speaker feature vectors 600 extracted in the vector space. In the illustrated example, the first region 652, centered on the first feature vector 650, contains six speaker feature vectors excluding the first feature vector 650. Also, the second region 662 (an area of ​​the same shape and size as the first region 652), centered on the second feature vector 660, contains five speaker feature vectors excluding the second feature vector 660. Therefore, in this case, the first region 652 may be determined as the central region.

[0074] Subsequently, the information processing system can determine the central feature vector based on one or more speaker feature vectors included in the central region. For example, the information processing system can determine the average of one or more speaker feature vectors included in the central region as the central feature vector. According to this embodiment, the information processing system can determine the average of seven speaker feature vectors included in the first region 652 determined as the central region as the central feature vector. As another example, the information processing system can extract the speaker feature vector located at or closest to the center of the central region as the central feature vector. According to this embodiment, the information processing system can extract the first feature vector 650 located at the center of the first region 652 determined as the central region as the central feature vector. The first feature vector 650 extracted as the central feature vector may be estimated to best represent speaker A's speech, as it is extracted based on the densest region among the first group 610, which includes multiple speaker feature vectors extracted based on speech segments containing only speaker A's speech.

[0075] In one embodiment, if clustering is performed on the first set of speaker feature vectors 600 before extracting the central feature vector, the information processing system can extract the central feature vector from the cluster containing the largest number of speaker feature vectors as a result of the clustering. For example, as a result of clustering the first set of speaker feature vectors 600, they may be classified into a first cluster containing the first group 610, the second group 620, and the third group 630, and a second cluster containing the fourth group 640. In this case, a central region can be determined in the space corresponding to the first cluster containing the largest number of speaker feature vectors, and the central feature vector can be extracted based on one or more speaker feature vectors included in the central region.

[0076] As described above, by extracting the central feature vector with the highest confidence level based on multiple speaker feature vectors extracted from the input signal, a representative speaker feature vector for a specific speaker can be extracted. Therefore, even if there are errors in the clustering results, or even if clustering is not performed, speaker segmentation can be performed with high accuracy.

[0077] Figure 7 shows an example of extracting a specific speaker's voice 720 from an input signal 710 according to one embodiment of the present disclosure. The information processing system can extract a specific speaker's voice 720 associated with a central feature vector from the input signal 710. For example, the information processing system can extract a first speaker's voice associated with a first central feature vector from the input signal using a target speech extraction model 700. Specific examples of the target speech extraction model 700 used to extract a specific speaker's voice 720 and its learning method will be described in detail later with reference to Figures 8 and 9.

[0078] After extracting the first speaker's voice, the information processing system can determine whether or not a voice segment remains in the residual signal excluding the first speaker's voice extracted from the input signal 710. If it is determined that a voice segment remains in the residual signal, at least some of the processes described above can be repeated, referring to Figures 4 to 7.

[0079] For example, the information processing system can extract multiple speaker feature vectors based on the remaining speech segments from the input signal 710, excluding the speech segment containing the extracted speech of a specific speaker 720; extract a second central feature vector with the highest confidence level based on the extracted speaker feature vectors; and extract the speech of a second speaker associated with the second central feature vector extracted from the input signal 710 or the residual signal.

[0080] As another example, the information processing system can extract a second central feature vector with the highest confidence level based on a first subset of speaker feature vectors extracted from a first set of speaker feature vectors extracted based on the input signal 710, excluding the speech segment containing the speech of a specific speaker 720, and then extract the speech of a second speaker associated with the second central feature vector extracted from the input signal 710 or the residual signal.

[0081] As yet another example, the information processing system can extract multiple speaker feature vectors based on the speech segments contained in the residual signal, extract a second central feature vector with the highest confidence level based on the extracted speaker feature vectors, and extract the speech of a second speaker associated with the second central feature vector extracted from the input signal 710 or the residual signal.

[0082] If the above-described series of processes is repeated until no speech segments remain in the residual signal, the speech of all speakers whose utterances are included in the input signal 710 can be separated.

[0083] Furthermore, the information processing system can transmit the separated audio of each speaker and / or the results of additional processing of each speaker's audio to the user terminal as a result of speaker segmentation. For example, based on the separated audio of each speaker, the information processing system can generate meeting minutes or subtitles by performing automatic speech recognition (ASR) or speech to text (STT) and transmit them to the user terminal.

[0084] Figure 8 shows a specific example of a target speech extraction model 800 according to one embodiment of the present disclosure. According to one embodiment, the target speech extraction model 800 may be a model that receives a mixed speech 810 and a target speaker feature vector 820 as inputs and outputs a mask 830 associated with the target speaker feature vector 820.

[0085] Here, the mixed speech 810 may be a speech obtained by mixing the voice of the target speaker with the voices of other speakers and / or noise. For example, the mixed speech 810 may be a speech in waveform form or a speech in spectrogram form. According to one embodiment, by performing frequency conversion on a speech in waveform form to convert it into a spectrogram form, a speech in spectrogram form can be used as the mixed speech 810.

[0086] The target speaker feature vector 820 may be a vector containing information about the target speaker. The mask 830 associated with the target speaker feature vector 820 may be a mask for removing the voices of other speakers and noise from the input signal, excluding the target speaker. Using the output mask 830 and the mixed voice 810, the target speaker's voice 840 can be extracted. For example, the target speaker's voice 840 can be extracted by applying the mask 830 to the mixed voice 810.

[0087] In one embodiment, when a spectrogram-like speech is used as the mixed speech 810, the extracted speech 840 of the target speaker may also be in a spectrogram-like form. In this case, the speech of the target speaker in spectrogram form can be converted into a waveform-like speech by performing an inverse frequency conversion.

[0088] According to one embodiment, the information processing system inputs the input signal as a mixed speech 810 and the extracted central feature vector as a target speaker feature vector 820 to the target speech extraction model 800, estimates a mask 830 associated with the central feature vector, and extracts the speech of a specific speaker by applying the mask 830 to the input signal. Additionally or alternatively, the information processing system may also use other speaker feature vectors extracted by a method similar to or different from the central feature vector extraction method described above as the target speaker feature vector 820.

[0089] The target speech extraction model 800 shown in Figure 8 is merely an example, and in other embodiments, the target speech extraction model 800 may be implemented differently from the illustrated model. For example, the target speech extraction model 800 may be implemented to immediately output the target speaker's voice 840 instead of outputting a mask 830 associated with the target speaker feature vector 820.

[0090] Figure 9 shows an example of a method for training a target speech extraction model 800 according to one embodiment of the present disclosure. According to one embodiment, the target speech extraction model 800 may be a machine learning model trained to receive a mixed speech 810 and a target speaker feature vector 820 as inputs and output a mask 830 associated with the target speaker feature vector 820. For example, the target speech extraction model 800 may be an artificial neural network model including at least one of a CNN (convolutional neural network), an LSTM (long short-term memory), or an FC (fully connected) layer.

[0091] According to one embodiment, the information processing system can acquire a first target speaker reference voice 822, a second target speaker reference voice 812, and other speaker reference voices 814 as training data for training the target speech extraction model 800.

[0092] The first target speaker reference voice 822 and the second target speaker reference voice 812 may be voices that include only the utterances of the target speaker. The first target speaker reference voice 822 and the second target speaker reference voice 812 may be the same as or different from each other. In addition, the other speaker reference voice 814 may be voices that include utterances of speakers other than the target speaker and / or noise. The information processing system can generate a mixed voice 810 by mixing the second target speaker reference voice 812 with the other speaker reference voice 814. Each of the voices 810, 812, 814, and 822 described above may be in waveform form or in spectrogram form. According to one embodiment, frequency conversion can be performed on the waveform form voice to convert it into a spectrogram form voice for use.

[0093] First, the information processing system can extract the target speaker feature vector 820 based on the first target speaker reference speech 822. The process of extracting the target speaker feature vector 820 based on the first target speaker reference speech 822 may be the same as or similar to the method described above, with reference to Figure 4.

[0094] Subsequently, the information processing system inputs the mixed speech 810 and the target speaker feature vector 820 into the target speech extraction model 800, and can estimate a mask 830 associated with the target speaker feature vector 820. Then, by applying the estimated mask 830 to the mixed speech 810, the target speaker's speech 840 can be extracted.

[0095] The information processing system can calculate a predicted loss 850 based on a comparison between the extracted target speaker's voice 840 and the second target speaker's reference voice 812. Subsequently, the calculated predicted loss 850 can be used to update the weight values ​​corresponding to the connections between multiple layers or nodes included in the target speech extraction model 800.

[0096] Figure 9 and the learning method described above are merely illustrative examples, and in other embodiments, the target speech extraction model 800 may be implemented differently from the illustrated model or may be learned by other methods.

[0097] Figure 10 is a flowchart illustrating an example of a speaker segmentation method 1000 according to one embodiment of the present disclosure. According to one embodiment, the speaker segmentation method 1000 may be initiated by a processor (e.g., at least one processor in an information processing system or user terminal) extracting a first set of speaker feature vectors based on speech segments included in an input signal (S1010). For example, the processor can detect speech segments from the input signal and extract a first set of speaker feature vectors based on the detected speech segments. Here, the first set of speaker feature vectors may include a plurality of speaker feature vectors. According to one embodiment, the processor can extract a plurality of speaker feature vectors based on each of a plurality of unit speech segments included in the input signal.

[0098] Subsequently, the processor can extract the first central feature vector with the highest confidence level based on the extracted first set of speaker feature vectors (S1020).

[0099] For example, the processor can first determine a first central region in the vector space based on the first set of speaker feature vectors that have been extracted. In one embodiment, the processor can determine the region in the vector space where the speaker feature vectors are most densely concentrated as the first central region. For example, the processor can determine the region containing the largest number of speaker feature vectors as the first central region from among several regions of the same size centered on each of the first set of speaker feature vectors extracted in the vector space.

[0100] After the first central region is determined, the processor can extract the first central feature vector based on one or more speaker feature vectors contained within the first central region. For example, the processor can determine the average of one or more speaker feature vectors contained within the first central region as the first central feature vector. Alternatively, the processor can extract the speaker feature vector located at or closest to the center of the first central region as the first central feature vector.

[0101] According to one embodiment, the processor can perform clustering on the extracted first set of speaker feature vectors before extracting the first central feature vector. If clustering is performed, the processor can extract the first central feature vector from the cluster containing the largest number of speaker feature vectors as a result of the clustering. For example, the processor can determine the first central region in the vector space corresponding to the cluster containing the largest number of speaker feature vectors as a result of the clustering, and extract the first central feature vector based on one or more speaker feature vectors included in the first central region.

[0102] Subsequently, the processor can extract the speech of the first speaker associated with the first central feature vector from the input signal (S1030). According to one embodiment, the processor can extract the speech of the first speaker associated with the first central feature vector from the input signal using a target speech extraction model. For example, the processor can use a target speech extraction model to estimate a mask associated with the first central feature vector based on the input signal and the first central feature vector, and then extract the speech of the first speaker based on the input signal and the estimated mask.

[0103] Furthermore, the processor can determine whether or not a speech segment remains in the residual signal after removing the first speaker's voice extracted from the input signal. If it is determined that a speech segment remains in the residual signal, the above-described series of processes can be repeated to extract the voices of other speakers included in the input signal or residual signal.

[0104] For example, the processor can return to step S1010 and repeat the series of processes. Specifically, the processor can extract a second set of speaker feature vectors based on the speech segments contained in the residual signal. Then, based on the extracted second set of speaker feature vectors, it can extract a second central feature vector with the highest confidence level and extract the speech of a second speaker associated with the second central feature vector from the input signal or residual signal.

[0105] As another example, the processor can repeat the series of processes by returning to step S1020. Specifically, the processor can extract a second central feature vector with the highest confidence level from the first set of speaker feature vectors, based on the speaker feature vectors extracted based on the speech segments included in the residual signal, and then extract the speech of the second speaker associated with the second central feature vector from the input signal or residual signal.

[0106] As yet another example, the processor can extract a second central feature vector with the highest confidence level based on the residual speech segments of the input signal, excluding the segments containing the first speaker's speech, and then extract the second speaker's speech associated with the second central feature vector from the input signal or residual signal.

[0107] The flowchart and the description above are merely illustrative examples, and the scope of this disclosure is not limited thereto. For example, according to other embodiments, some steps may be added / modified / deleted, and the order of the steps may be changed.

[0108] The methods described above may be provided as computer programs stored on a computer-readable recording medium for execution on a computer. The medium may continuously store computer executable programs or temporarily store them for execution or download. The medium may be various recording or storage means in the form of a single or several hardware devices combined, and is not limited to a medium directly connected to a computer system, but may be distributed on a network. Examples of mediums include magnetic media such as hard disks, floppy disks (registered trademarks), and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical mediums such as floptical disks; and ROMs, RAMs, flash memories, etc., which may be configured to store program instructions. Other examples of mediums include recording media or storage media managed by app stores that distribute applications or other sites, servers that supply or distribute various software.

[0109] The methods, operations, or techniques described herein may be embodied by a variety of means. For example, such techniques may be embodied by hardware, firmware, software, or a combination thereof. It will be understood by the ordinary art that various exemplary logical blocks, modules, circuits, and algorithmic stages described in conjunction with the disclosure may be embodied by electronic hardware, computer software, or a combination of both. To clearly illustrate such interchangeability of hardware and software, various exemplary components, blocks, modules, circuits, and stages have been generally described above in terms of their functional aspects. Whether such functions are embodied as hardware or as software depends on the design requirements given to the particular application and the overall system. The ordinary art may also embodied the functions described in various ways for each particular application, but such embodiments should not be construed as deviating from the scope of this disclosure.

[0110] In hardware implementation, the processing unit used to perform the method may be implemented in one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to have the functions described herein, computers, or combinations thereof.

[0111] Accordingly, the various exemplary logic blocks, modules, and circuits described in conjunction with this disclosure may be embodied or executed by any combination of general-purpose processors, DSPs, ASICs, FPGAs or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or anything designed to have the functions described herein. The general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be embodied as a combination of computing devices, for example, a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other combination of configurations.

[0112] In the embodiment of firmware and / or software, the technique may be embodied as instructions stored on a computer-readable medium such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, compact disc (CD), or magnetic or optical data storage devices. The instructions may be executable by one or more processors, and the processors may be made to perform specific embodiments of the functions described herein.

[0113] When embodied as software, the above methods may be stored on or transmitted through a computer-readable medium as one or more instructions or codes. The computer-readable medium includes both computer storage and communication media, including any medium that facilitates the transmission of computer programs from one location to another. The storage medium may be any available medium accessible to a computer. As a non-limiting example, such a computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium accessible to a computer that is used to transport or store desired program code in the form of instructions or data structures. Any connection may also be appropriately referred to as a computer-readable medium.

[0114] For example, when software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, stranded wire, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, these are included within the definition of a medium. The terms "disk" and "disc" as used in this application include CD, laserdisc, optical disc, DVD (digital versatile disc), floppy disk (registered trademark), and Blu-ray disc, where typically a disk reproduces data magnetically and a disc reproduces data optically using a laser. The above combinations should also be included within the scope of computer-readable media.

[0115] The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, portable disk, CD-ROM, or any other known form of storage medium. An exemplary storage medium may be coupled to the processor so that the processor can read information from or record information to the storage medium. Alternatively, the storage medium may be integrated into the processor. The processor and storage medium may reside within an ASIC. The ASIC may reside within a user terminal. Alternatively, the processor and storage medium may exist as separate components in the user terminal.

[0116] Although the embodiments described above are described as utilizing aspects of the currently disclosed subject matter in one or more standalone computer systems, the disclosure may be embodied in conjunction with any computing environment, such as a network or a distributed computing environment, without limiting it. Furthermore, aspects of the subject matter may be embodied in multiple processing chips or devices, and storage may be similarly affected across multiple devices. These devices may include PCs, network servers, and portable devices.

[0117] While this disclosure has been described in relation to some embodiments, various modifications and alterations are permitted as long as they do not deviate from the scope of this disclosure as understandable to a person of ordinary skill in the art to which the invention of this disclosure pertains. Furthermore, such modifications and alterations should be considered to fall within the scope of the claims appended to this specification.

Claims

1. A speaker splitting method performed by at least one processor, The steps include: detecting multiple speech segments contained in the input signal, extracting a speaker feature vector for each speech segment, and extracting a first set of speaker feature vectors; The first step involves extracting the speaker feature vector with the highest confidence level from the extracted first set of speaker feature vectors as the first central feature vector. The steps include: extracting the speech of the first speaker corresponding to the first central feature vector from the input signal using the target speech extraction model; Includes, The step of extracting the first central feature vector is as follows: The confidence level is evaluated based on the density of the speaker feature vectors in the first set, and the first central feature vector is extracted. The speaker segmentation method is characterized in that the confidence level is the probability that the speaker feature vector is estimated to represent the speech features of a particular speaker.

2. The step of extracting the first central feature vector is as follows: The steps include determining a first central region in the vector space based on the density of the first set of speaker feature vectors extracted, A step of extracting the first central feature vector based on one or more speaker feature vectors included in the first central region, The speaker splitting method according to claim 1, including the method described in claim 1.

3. The step of determining the first central region is: The step of determining the region in the vector space where the speaker feature vectors are most densely concentrated as the first central region. The speaker splitting method according to claim 2, including the method described in claim 2.

4. The step of determining the first central region is: The step of determining the first central region as the region containing the largest number of speaker feature vectors from among multiple regions of the same size centered on each of the first set of speaker feature vectors extracted within the vector space. The speaker splitting method according to claim 2, including the method described in claim 2.

5. The step of extracting a first central feature vector based on one or more speaker feature vectors contained in the first central region is as follows: Steps to determine the first central feature vector as the average of one or more speaker feature vectors included in the first central region. The speaker splitting method according to claim 2, including the method described in claim 2.

6. The step of extracting a first central feature vector based on one or more speaker feature vectors contained in the first central region is as follows: Step 1: Extract the speaker feature vector located at the center of the first central region or closest to the center as the first central feature vector. The speaker splitting method according to claim 2, including the method described in claim 2.

7. The step of performing clustering on the first set of speaker feature vectors extracted. The speaker splitting method according to claim 1, further comprising:

8. The step of extracting the first central feature vector is as follows: The step of extracting the first central feature vector from the cluster containing the largest number of speaker feature vectors as a result of the clustering. The speaker splitting method according to claim 7, including the following:

9. The step of extracting the first central feature vector is as follows: In the vector space, within the space corresponding to the cluster containing the largest number of speaker feature vectors as a result of the clustering, a first central region is determined based on the density of speaker feature vectors. Steps to extract the first central feature vector based on one or more speaker feature vectors included in the first central region. The speaker splitting method according to claim 7, including the following:

10. This step involves determining whether or not an audio segment remains in the residual signal after removing the first speaker's voice extracted from the input signal. The speaker splitting method according to claim 1, further comprising:

11. If it is determined that a speech segment remains in the residual signal, the step of extracting a second central feature vector with the highest confidence level from the first set of speaker feature vectors, based on the speaker feature vector extracted based on the speech segment included in the residual signal, The steps include extracting the voice of a second speaker associated with the second central feature vector from the input signal or the residual signal, The speaker splitting method according to claim 10, further comprising:

12. The steps include: extracting a second set of speaker feature vectors based on the speech segments included in the residual signal; Based on the extracted second set of speaker feature vectors, the step of extracting the second central feature vector with the highest confidence level, The steps include extracting the voice of a second speaker associated with the second central feature vector from the input signal or the residual signal, The speaker splitting method according to claim 10, further comprising:

13. If it is determined that an audio segment remains in the residual signal, the step of extracting a second central feature vector with the highest confidence level based on the residual audio segments included in the input signal, excluding the audio segment containing the voice of the first speaker; The steps include extracting the voice of a second speaker associated with the second central feature vector from the input signal or the residual signal, The speaker splitting method according to claim 10, further comprising:

14. The step of extracting the first set of speaker feature vectors is: Steps to extract the first set of speaker feature vectors based on each of the plurality of speech segments included in the input signal. The speaker splitting method according to claim 1, including the method described in claim 1.

15. The step of extracting the speech of the first speaker corresponding to the first central feature vector from the input signal is as follows: Steps include extracting the speech of the first speaker corresponding to the first central feature vector from the input signal using the aforementioned target speech extraction model. The speaker splitting method according to claim 1, including the method described in claim 1.

16. A computer-readable non-temporary recording medium that stores instruction words for executing the speaker splitting method according to any one of claims 1 to 15.

17. An information processing system, Memory and A processor connected to the memory and configured to execute at least one computer-readable program contained in the memory, Includes, The aforementioned at least one program, Multiple speech segments are detected within the input signal, and a speaker feature vector is extracted for each speech segment to extract a first set of speaker feature vectors. The confidence level is evaluated based on the density of the speaker feature vectors in the first set, and the speaker feature vector with the highest confidence level among the speaker feature vectors in the first set is extracted as the first central feature vector. Using the target speech extraction model, the speech of the first speaker corresponding to the first central feature vector is extracted from the input signal. An information processing system comprising an instruction word for performing a process configured such that the confidence level is the probability that the speaker feature vector is estimated to represent the speech features of a particular speaker.