Apparatus and method for recognizing user speech through speaker separation in speech data

The device automatically separates and classifies user voices in voice data, addressing cumbersome registration issues and privacy concerns by employing speaker separation and clustering techniques.

WO2025183426A1PCT designated stage Publication Date: 2025-09-04RETURN ZERO INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/002588
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-26
Filing Date
2025-02-25
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing voice recognition technologies require cumbersome user registration and periodic re-registration to distinguish multiple speakers, incurring management costs and violating privacy laws.

Method used

A user voice recognition device and method that automatically collects voice data, separates speakers based on vocal differences, calculates similarity, and classifies the user's voice without separate registration, using a processor to perform embedding and clustering.

Benefits of technology

Enables quick and easy recognition of the user's voice among multiple speakers, reducing registration burdens and management costs while complying with privacy regulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025002588_04092025_PF_FP_ABST
    Figure KR2025002588_04092025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to an apparatus and a method for recognizing user speech through speaker separation in speech data, the apparatus comprising: a communication module for performing communication with an external device; a memory for storing at least one process for performing a user speech recognition operation through speaker separation in speech data; and a processor for performing the user speech recognition operation through speaker separation in the speech data according to the process, wherein the processor collects speech data in which a plurality of speakers, including a user, participate on the basis of a user terminal, separates each speaker on the basis of each piece of speech data and performs embedding, and performs clustering on all combinations between each embedding vector on the basis of similarity so as to classify the speech of the user from among the plurality of speakers.
Need to check novelty before this filing date? Find Prior Art

Description

Device and method for user voice recognition through speaker separation in voice data

[0001] The present disclosure relates to a speech recognition device and method, and more particularly, to a user speech recognition device and method through speaker separation in speech data.

[0002] As demand for non-face-to-face services has increased, technology development for telecommunications-based communication is actively underway in various fields.

[0003] One such technology has been developed that converts voice data into text and provides it. However, this technology requires user registration through the device to distinguish multiple speakers based on vocal differences or to recognize a specific user's voice.

[0004] Not only is this process extremely cumbersome, but it also requires periodic re-registration. Furthermore, recognizing and storing additional personal information can incur management costs to meet the requirements of the Personal Information Protection Act, potentially creating a financial burden on users.

[0005] Therefore, there is a need to develop a technology that can quickly and easily recognize only the user's voice based on voice data containing the voices of multiple speakers participating in a conversation without a separate registration procedure to recognize / distinguish the user's voice.

[0006] The present disclosure provides a user voice recognition device and method through speaker separation in voice data, which automatically collects voice data including the user's voice and calculates the similarity between each voice data, thereby enabling the user's voice to be recognized and classified quickly and easily among the speakers in the voice data.

[0007] The problems to be solved by the present disclosure are not limited to the problems mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the description below.

[0008] According to an aspect of the present disclosure for achieving the above-described technical problem, a user voice recognition device through speaker separation in voice data comprises: a communication module for performing communication with an external device; a memory storing at least one process for performing a user voice recognition operation through speaker separation in voice data; and a processor for performing a user voice recognition operation through speaker separation in the voice data according to the process, wherein the processor collects voice data in which a plurality of speakers including a user have participated based on a user terminal, separates each speaker based on each voice data and performs embedding, and performs clustering based on similarity for all combinations between each embedding vector to classify the user's voice among the plurality of speakers, wherein each voice data includes the voices of two speakers, and the processor selects only one embedding vector among two embedding vectors existing for each of the voice data to generate all possible combinations, calculates a similarity for each combination, and selects a combination whose sum of the similarities is the largest, thereby classifying the user's voice among the plurality of speakers.

[0009] In addition, a method for user voice recognition through speaker separation in voice data according to one aspect of the present disclosure includes the steps of collecting voice data in which a plurality of speakers, including a user, participated based on a user terminal; performing embedding by separating each speaker based on each voice data; selecting only one embedding vector among two embedding vectors existing for each of the voice data according to the embedding to generate all possible combinations; calculating a similarity for each combination; and selecting a combination in which the sum of the similarities is the largest to classify the user's voice among the plurality of speakers, wherein each of the voice data may include the voices of two or more speakers.

[0010] In addition, a computer program stored in a computer-readable recording medium for executing a method for implementing the present disclosure may be further provided.

[0011] In addition, a computer-readable recording medium recording a computer program for executing a method for implementing the present disclosure may be further provided.

[0012] According to the aforementioned problem solving means of the present disclosure, voice data including the user's voice is automatically collected, and the similarity between each voice data is calculated, thereby enabling the user's voice to be recognized and classified quickly and easily among speakers in the voice data.

[0013] The effects of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.

[0014] FIG. 1 is a diagram showing the network structure of a system for providing a user voice recognition service through speaker separation in voice data according to one embodiment of the present disclosure.

[0015] FIG. 2 is a diagram showing the configuration of a service server for user voice recognition through speaker separation in voice data according to one embodiment of the present disclosure.

[0016] FIG. 3 is a diagram illustrating a user voice recognition method through speaker separation in voice data according to one embodiment of the present disclosure.

[0017] FIG. 4 is a diagram showing a specific operation of generating an embedding vector in a user voice recognition method through speaker separation in voice data according to one embodiment of the present disclosure.

[0018] FIG. 5 is a diagram showing a specific operation of performing clustering in a user voice recognition method through speaker separation in voice data according to one embodiment of the present disclosure.

[0019] The advantages and features of the present disclosure, and methods for achieving them, will become clearer with reference to the embodiments described below in detail with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided solely to ensure that the present disclosure is complete and to fully inform those skilled in the art of the scope of the present disclosure, and the present disclosure is defined solely by the scope of the claims.

[0020] The terminology used herein is for the purpose of describing embodiments only and is not intended to limit the present disclosure. In this specification, the singular also includes the plural unless specifically stated otherwise. As used herein, the terms "comprises" and / or "comprising" do not exclude the presence or addition of one or more other components in addition to the mentioned components. Like reference numerals refer to like components throughout the specification, and "and / or" includes each and any combination of one or more of the mentioned components. Although "first", "second", etc. are used to describe various components, these components are not limited by these terms. These terms are only used to distinguish one component from another. Therefore, it should be understood that a first component mentioned below may also be a second component within the technical spirit of the present disclosure.

[0021] Unless otherwise defined, all terms (including technical and scientific terms) used herein may be used in their common sense to those of ordinary skill in the art to which this disclosure pertains. Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise.

[0022] Throughout this disclosure, the same reference numerals denote the same components. This disclosure does not describe all elements of the embodiments, and general contents in the technical field to which this disclosure belongs or contents overlapping between embodiments are omitted. The term "part" or "module" as used in the specification means a hardware component such as a software or FPGA or ASIC, and the "part" or "module" performs certain roles. However, the "part" or "module" is not limited to software or hardware. The "part" or "module" may be configured to be in an addressable storage medium and may be configured to play one or more processors. Thus, as an example, the "part" or "module" includes components such as software components, object-oriented software components, class components, and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functionality provided within the components and “sub-components” or “modules” may be combined into a smaller number of components and “sub-components” or “modules” or further separated into additional components and “sub-components” or “modules”.

[0023] Throughout the specification, when a part is said to be "connected" to another part, this includes not only direct connection but also indirect connection, and indirect connection includes connection via a wireless communication network.

[0024] Additionally, when a part is said to "include" a component, this does not mean that it excludes other components, but rather that it may include other components, unless otherwise specifically stated.

[0025] Throughout the specification, when we say that an element is "on" another element, this includes not only cases where the element is in contact with the other element, but also cases where another element exists between the two elements.

[0026] The terms first, second, etc. are used to distinguish one component from another, and the components are not limited by the aforementioned terms.

[0027] Singular expressions include plural expressions unless the context clearly indicates otherwise.

[0028] The identification codes for each step are used for convenience of explanation and do not describe the order of each step. Each step may be performed in a different order than specified unless the context clearly indicates a specific order.

[0029] The terms used in the following explanation are defined as follows.

[0030] Although described herein as a "service server," this device provides user voice recognition services through speaker separation within voice data, and may include various devices capable of performing computational processing. Furthermore, this service server may be connected / interconnected with separate servers, computers, and / or mobile devices to allow for configuration changes or to collect, analyze, and provide information. The service server's type and form are not limited.

[0031] Here, the computer may include, for example, a notebook, desktop, laptop, tablet PC, slate PC, etc. equipped with a web browser.

[0032] The above server is a server that processes information by communicating with external devices, and may include an application server, a computing server, a database server, a file server, a game server, a mail server, a proxy server, and a web server.

[0033] The above portable terminal may include, for example, a wireless communication device that ensures portability and mobility, and may include all kinds of handheld-based wireless communication devices such as a PCS (Personal Communication System), GSM (Global System for Mobile communications), PDC (Personal Digital Cellular), PHS (Personal Handyphone System), PDA (Personal Digital Assistant), IMT (International Mobile Telecommunication)-2000, CDMA (Code Division Multiple Access)-2000, W-CDMA (W-Code Division Multiple Access), WiBro (Wireless Broadband Internet) terminal, a smart phone, and a wearable device such as a watch, a ring, a bracelet, an anklet, a necklace, glasses, contact lenses, or a head-mounted device (HMD).

[0034] The operating principle and embodiments of the present disclosure are described below with reference to the attached drawings.

[0035] FIG. 1 is a diagram showing the network structure of a system for providing a user voice recognition service through speaker separation in voice data according to one embodiment of the present disclosure.

[0036] Referring to FIG. 1, a system (hereinafter referred to as a “service providing system”) (10) for providing a user voice recognition service through speaker separation in voice data according to one embodiment of the present disclosure may be configured to include at least one of a service server (100) and a user terminal (200).

[0037] The service server (100) is a device for providing a user voice recognition service to a user (i.e., a user voice recognition device through speaker separation in voice data). For this purpose, a separate web page and / or platform (application) can be provided, and the user terminal (200) can provide or receive various information / data for the user voice recognition service based on the web page or platform.

[0038] When a user runs a separate web page or platform through a user terminal (200), this service server (100) collects voice data in which multiple speakers, including the user, participate, separates each speaker based on each voice data, performs embedding, and performs clustering based on similarity for all combinations between each embedding vector, thereby classifying the user's voice among multiple speakers.

[0039] Meanwhile, the user terminal (200) is a terminal possessed by a user who wishes to receive a user voice recognition service through the service server (100), and may include at least one or more terminals.

[0040] Each user may access a web page provided by the service server (100) or install and provide a platform (application) to receive user voice recognition service through his / her user terminal (200).

[0041] Accordingly, when a user requests provision of a service to the service server (100) by executing a separate web page or platform through a user terminal (200), the service server (100) recognizes only the user's voice from specific voice data or video data in response, thereby providing the recognition data in at least one form.

[0042] At this time, the user terminal (200) may be a computer, UMPC (Ultra Mobile PC), workstation, netbook, PDA (Personal Digital Assistants), portable computer, web tablet, wireless phone, mobile phone, smart phone, pad, smart watch, wearable terminal, e-book, PMP (portable multimedia player), portable game console, navigation device, black box or digital camera, other mobile communication terminal, etc., which can install and execute multiple application programs (i.e., applications) desired by the user. That is, the user terminal (200) may be provided in various forms, and its form is not limited.

[0043] Although only one user terminal (200) is illustrated in FIG. 1, this is only for convenience of explanation, and there may be at least one or more, and the number and type are not limited.

[0044] As described above, the service provision system according to the present disclosure can be implemented through data / information transmission and reception between a network-based service server (100) and a user terminal (200).

[0045] FIG. 2 is a diagram showing the configuration of a service server for user voice recognition through speaker separation in voice data according to one embodiment of the present disclosure.

[0046] Referring to FIG. 2, a service server (100) according to one embodiment of the present disclosure may be configured to include at least one of a communication module (110), a memory (120), and a processor (130).

[0047] The communication module (110) can communicate with at least one of various terminals (devices), external storage (e.g., database (140)), external servers, and cloud servers.

[0048] Meanwhile, an external server or cloud server may be configured to perform at least a portion of the role of the processor (130). That is, data processing or data operations, etc. may be performed on an external server or cloud server, and the present disclosure does not place any particular limitations on this method.

[0049] Meanwhile, the communication module (110) can support various communication methods according to the communication standards of the communicating target (e.g., electronic device, external server, device, etc.).

[0050] For example, the communication module (110) may be configured to communicate with a communication target using at least one of WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed ​​Downlink Packet Access), HSUPA (High Speed ​​Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), 5G (5th Generation Mobile Telecommunication), Bluetooth™, RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra-Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi Direct, and Wireless USB (Wireless Universal Serial Bus) technologies.

[0051] Meanwhile, the memory (120) may be configured to store various information related to the present disclosure. In the present disclosure, the memory (120) may be provided in the device itself according to the present disclosure. Alternatively, at least a portion of the memory (120) may refer to at least one of a database (DB, 140) and a cloud storage (or cloud server). That is, the memory (120) may be sufficient as long as it is a space where information necessary for the device and method according to the present disclosure is stored, and it may be understood that there are no restrictions on the physical space. Accordingly, in the following, the memory (120), the database (140), the external storage, and the cloud storage (or cloud server) will not be separately distinguished, and will all be referred to as the memory (120).

[0052] This memory (120) can store a plurality of application programs (or applications) running on the service server (100), data for the operation of the service server (100), and commands. At least some of these application programs can be downloaded from an external server via wireless communication. Meanwhile, the application programs can be stored in at least one memory provided in the memory (120), installed on the service server (100), and driven to perform operations (or functions) by at least one processor stored in the memory (120) via the processor (130).

[0053] Meanwhile, at least one memory may include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, and an optical disk. In addition, the memory may store information temporarily, permanently, or semi-permanently, and may be provided as a built-in or removable type.

[0054] Next, the processor (130) may be configured to control the overall operation of the device related to the present disclosure. The processor (130) may process signals, data, information, etc. input or output through the components discussed above, or provide or process appropriate information or functions to the user.

[0055] The processor (130) includes at least one CPU (Central Processing Unit) and can perform functions according to the present disclosure.

[0056] Specifically, when the processor (130) runs a separate web page or platform, it collects voice data in which multiple speakers, including the user, participated, separates each speaker based on each voice data and performs embedding, and performs clustering based on similarity for all combinations between each embedding vector, thereby classifying the user's voice among multiple speakers. Here, the voice data may be generated by recording a call made based on the user terminal (200). However, this is only limited as an example for the convenience of explanation, and may be voice data generated by recording a meeting, conference, conversation, etc., or may use various data recorded / taped in other ways, and is not limited thereto. At this time, the processor (130) may collect voice data for a preset period of time, or may collect voice data until a preset threshold (number) is satisfied.

[0057] At this time, the processor (130) may perform an operation to provide a user voice recognition service immediately when the user runs a web page or platform based on the setting information of the user terminal (200) (first case), or may perform an operation to provide a user voice recognition service according to a separate service provision request received by the user based on the web page or platform (second case).

[0058] In the first case, the processor (130) collects voice data involving multiple speakers, including the user, based on the user terminal (200). Furthermore, in the second case, the processor (130) searches for and collects voice data previously stored in the user terminal (200) through the memory of the user terminal (200). While the voice data is being collected, the web page or platform may be executed in a hidden state without being displayed on the user terminal (200).

[0059] Thereafter, when performing embedding, the processor (130) separates each speaker from each voice data, generates an embedding vector, and groups each speaker. At this time, the embedding vector can be generated for each word. Thereafter, the processor (130) calculates the mean of all embedding vectors in each group formed by the grouping, normalizes it, and extracts a representative embedding vector for each group.

[0060] Meanwhile, when performing clustering, the processor (130) selects only one embedding vector for each voice data, generates all possible combinations, and calculates the similarity for each combination. Here, the similarity can be calculated through the inner product of the different embedding vectors of each combination.

[0061] Accordingly, the processor (130) can select the combination that maximizes the sum of the calculated similarities and classify it as the user's voice. Here, the sum of the similarities can be calculated through a preset method or algorithm.

[0062] In addition, the specific operation of the processor (130) will be described below based on each drawing.

[0063] At least one component may be added or deleted to correspond to the performance of the components illustrated in FIG. 2. Furthermore, it will be readily apparent to those skilled in the art that the relative positions of the components may be altered to correspond to the performance or structure of the device.

[0064] FIG. 3 is a diagram illustrating a user voice recognition method through speaker separation in voice data according to one embodiment of the present disclosure.

[0065] Referring to FIG. 3, the service server (100) collects voice data in which multiple speakers, including the user, participate based on the user terminal (S110), and performs embedding by separating each speaker based on each voice data (S120).

[0066] Next, the service server (1000) performs clustering based on the similarity for all combinations between each embedding vector generated according to the result of the embedding performed by step S120, and classifies the user's voice among the plurality of speakers by calculating the sum of the similarities for each combination (S140).

[0067] Below, each of the above-described steps S120 and S130 will be described in detail based on FIGS. 4 and 5.

[0068] FIG. 4 is a drawing showing a specific operation of generating an embedding vector in a method for user voice recognition through speaker separation in voice data according to one embodiment of the present disclosure, and shows step S120 of FIG. 3 in more detail.

[0069] Referring to FIG. 4, the service server (100) separates each speaker from each voice data to generate an embedding vector (S121), and groups each speaker by generating the embedding vector (S122). For example, if a single voice data contains the voices of speaker A and speaker B, grouping is performed as speaker A and speaker B.

[0070] Next, the service server (100) calculates the mean of all embedding vectors in each group, normalizes them (S123), and extracts a representative embedding vector of each group (S124).

[0071] That is, when embedding is performed on n pieces of voice data, a total of 2n embedding vectors can be generated. At this time, the embedding vectors can be determined by the number of speakers or the number of reference voice data.

[0072] FIG. 5 is a drawing showing a specific operation of performing clustering in a user voice recognition method through speaker separation in voice data according to one embodiment of the present disclosure, and shows step S130 of FIG. 3 in more detail.

[0073] Referring to FIG. 5, the service server (100) selects only one embedding vector for each voice data and generates all possible combinations (S131).

[0074] That is, when using n voice data, a total of 2 n There are 2 possible combinations. For example, when n = 10, all possible combinations are 2 10 There are a total of 1024.

[0075] Next, the service server (100) calculates the similarity for each combination (S132). Specifically, when the embedding vector selected for each voice data by step S131 is N and there are a total of 10 voice data, N 2 The nC2 similarities of the branches are calculated. That is, 50 similarities out of 100 are calculated.

[0076] Next, based on the similarity for each combination produced by step S132, the sum of the similarities for each combination is calculated, and the combination with the largest sum is selected (S133). That is, if there are 10 voice data, the sum of the similarities for all 1024 possible combinations is calculated, and the combination with the largest sum is selected.

[0077] Accordingly, the combination selected by step S133 can be classified as the user's voice.

[0078] Below, we describe the operation of calculating the sum of similarities for each combination.

[0079] First, m is the shape of the embedding vector, which can vary depending on the embedding model.

[0080] To define a lookup table, a matrix of (1, m) and its corresponding vectors are concatenated to create a vector of dimension (2, m).

[0081] For example, when extracting n voice data containing the user's voice, it can be expressed as a vector of dimensions (n, 2, m), which is named v-vector.

[0082] If we change the v-vector resources I, j, and k to their abbreviated notations, silver This becomes the following. Split the above vector again into two from (n, 2, m) to (n, 1, m), and stack the 0th 2nd and 1st 2nd again. Finally, an (n, 4, m) matrix is ​​generated. Meanwhile, Is , becomes. If the number of speakers changes from two to three, it becomes a duplicate combination. n H c , 2 people (0, 1), (0, 1), (1, 0), (1, 1), and (0, 0, 0) make 8.

[0083] one side, and For each (n, 4, m), we can transpose (n, 4, m) to (4, n, m) and repeat the values ​​of the second dimension to generate two vectors of dimensions (4, n, m). By doing this, we can perform matrix multiplication (inner product). × , that is, the similarity can be calculated.

[0084] At this time, Einstein notation can be used for the inner product.

[0085] Specifically, we define a probability table (prob_table) for all dimensions, and use it as a set of n duplicate combinations of 0 and 1. n H c After creating all combinations, the sum is calculated. Here, n is the number of speakers, and c can be the number of voice data / number of recent calls.

[0086] This can be expressed as in <Mathematical Formula 1> below.

[0087]

[0088] Accordingly, according to the present disclosure, it is possible to quickly and easily recognize only the user's voice based on voice data including the voices of multiple speakers participating in a conversation without a separate registration procedure for recognizing / distinguishing the user's voice.

[0089] The above-described program may include codes coded in a computer language, such as C, C++, JAVA, or machine language, that can be read by the processor (CPU) of the computer through the device interface of the computer, so that the computer reads the program and executes the methods implemented as a program. Such codes may include functional codes related to functions that define functions necessary for executing the methods, and may include control codes related to execution procedures necessary for the processor of the computer to execute the functions according to a predetermined procedure. In addition, such codes may further include memory reference-related codes regarding which location (address address) of the internal or external memory of the computer should reference additional information or media necessary for the processor of the computer to execute the functions. In addition, if the processor of the computer needs to communicate with any other computer or server located remotely in order to execute the functions, the code may further include communication-related code regarding how to communicate with any other computer or server located remotely using the communication module of the computer, and what information or media to send and receive during communication.

[0090] The above storage medium refers to a medium that stores data semi-permanently and can be read by a device, rather than a medium that stores data for a short period of time, such as a register, cache, or memory. Specifically, examples of the storage medium include, but are not limited to, ROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage device. That is, the program can be stored in various recording media on various servers that the computer can access or in various recording media on the user's computer. In addition, the medium can be distributed across network-connected computer systems, so that computer-readable code can be stored in a distributed manner.

[0091] The steps of a method or algorithm described in connection with the embodiments of the present disclosure may be implemented directly in hardware, implemented as a software module executed by hardware, or implemented by a combination thereof. The software module may reside in a random access memory (RAM), a read only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a flash memory, a hard disk, a removable disk, a CD-ROM, or any other form of computer-readable recording medium well known in the art to which the present disclosure pertains.

[0092] While the embodiments of the present disclosure have been described above with reference to the attached drawings, those skilled in the art will appreciate that the present disclosure can be implemented in other specific forms without altering the technical spirit or essential features thereof. Therefore, the embodiments described above should be understood to be illustrative in all respects and not restrictive.

Claims

1. Communication module that communicates with external devices; A memory storing at least one process for performing a user voice recognition operation by separating speakers in voice data; and It includes a processor that performs user voice recognition operation by separating speakers in the voice data according to the above process, The above processor collects voice data in which multiple speakers, including the user, participate based on the user terminal, separates each speaker based on each voice data, performs embedding, and performs clustering based on similarity for all combinations between each embedding vector to classify the user's voice among the multiple speakers. Each of the above voice data includes the voices of two speakers, The processor is characterized in that it selects only one embedding vector among the two embedding vectors existing for each of the voice data, generates all possible combinations, calculates the similarity for each combination, and selects the combination with the largest sum of the similarities to classify the user's voice among the plurality of speakers. A user voice recognition device that separates speakers from voice data.

2. In paragraph 1, The above processor, When performing the above embedding, each speaker is separated, an embedding vector is created, and each speaker is grouped. Characterized in that the mean of all embedding vectors in each group is calculated and normalized, and a representative embedding vector of each group is extracted. A user voice recognition device that separates speakers from voice data.

3. In paragraph 2, The above processor, Characterized in that the above embedding vector is generated for each word, A user voice recognition device that separates speakers from voice data.

4. In paragraph 1, The above processor, Characterized in that the similarity is calculated through the inner product of different embedding vectors of each combination above. A user voice recognition device that separates speakers from voice data.

5. In paragraph 4, The sum of the above similarities is It is characterized by being expressed as in <Mathematical Formula 1> based on the Einstein notation. A user voice recognition device that separates speakers from voice data. <Mathematical Formula 1> Here, silver , silver , n represents the number of speakers, and c represents the number of voice data.

6. In paragraph 1, The number of all combinations above is, 2 above n It's a dog, Here, n represents the number of voice data, and 2 represents the number of speakers. A user voice recognition device that separates speakers from voice data.

7. In paragraph 1, The above voice data is, Characterized in that it is generated by recording a call made based on the user terminal above, A user voice recognition device that separates speakers from voice data.

8. In paragraph 7, The above processor, Characterized in that the voice data is collected based on the user terminal until a preset threshold is satisfied. A user voice recognition device that separates speakers from voice data.

9. In a method for user voice recognition through speaker separation in voice data performed by a device, A step of collecting voice data in which multiple speakers, including the user, participate based on the user terminal; A step of performing embedding by separating each speaker based on each voice data; A step of generating all possible combinations by selecting only one embedding vector among the two embedding vectors existing for each of the above voice data according to the above embedding; A step of calculating the similarity for each combination; and A step of classifying the user's voice among the plurality of speakers by selecting a combination that maximizes the sum of the similarities, The above voice data is characterized in that it includes the voices of two or more speakers. A method for user voice recognition by separating speakers from voice data.

Citation Information

Patent Citations

  • System and method for voice recognition

    KR1020170135133A

  • Automated removable suction piles installation apparaturs

    KR1020220088673A

  • Composite material and manufacturing method of this

    KR1020230017520A

  • Transparent coating removal through laser ablation

    KR1020230023589A

  • Deep Learning Based Topology Reconstruction Method for Digitalization of Image Format Piping and Instrumentation Diagram

    KR1020240171279A