Electronic device, method, and non-transitory computer-readable storage medium for determining utterance section of speaker from audio data

By managing vector storage and computation through periodic removal of vectors, the electronic device optimizes resource usage for speaker separation, enabling efficient diarization and speech-to-text processing of long audio data on devices with limited resources.

WO2025146956A1PCT designated stage expired Publication Date: 2025-07-10SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/019205
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-02
Filing Date
2024-11-28
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing electronic devices face challenges in efficiently determining speaker segments from audio data due to the exponential increase in computational resources and memory requirements as the length of audio data increases, particularly when performing speaker separation and diarization.

Method used

The electronic device employs a method to manage and reduce the number of vectors by periodically removing or discarding vectors used in calculating similarities, maintaining the total number of vectors to be less than or equal to a specified number, thereby optimizing computational resources and memory usage.

Benefits of technology

This approach allows for high-speed speaker separation of long audio data on devices with limited resources by maintaining consistent computational and memory requirements, enabling efficient speaker diarization and speech-to-text processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024019205_10072025_PF_FP_ABST
    Figure KR2024019205_10072025_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device according to an embodiment of the present invention may obtain a plurality of frames by dividing audio data while obtaining the audio data using a microphone. The electronic device may obtain first vectors respectively corresponding to the plurality of frames. The electronic device may determine speakers respectively corresponding to the first vectors using groups which are obtained by grouping the first vectors and second vectors stored in a memory, and in which the first vectors are respectively included. The electronic device may store, in the memory, information indicating a speaker of at least one time section of the audio data, the information being determined using the speakers respectively corresponding to the first vectors. The electronic device may remove at least one of the first vectors and the second vectors from the memory on the basis of the total number of the first vectors and the second vectors in order to reduce the total number to at most a specified number.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device, method, and non-transitory computer-readable storage medium for determining a speaker's speech segment from audio data

[0001] The disclosure relates to an electronic device, a method, and a non-transitory computer-readable storage medium for determining a speaker's speech segment from audio data.

[0002] Natural language (or ordinary language) refers to the language used in human daily life. Electronic devices that process natural language are being developed. For example, electronic devices can detect user speech from audio data containing speech, or generate text expressing the detected speech.

[0003] The above information is disclosed solely as background information to aid understanding of the disclosure. No determination or assertion is made as to whether any of the above is prior art to the disclosure.

[0004] Aspects of the disclosure are intended to at least address the problems and / or shortcomings mentioned above and to provide at least the advantages described below. Accordingly, aspects of the disclosure are intended to provide an electronic device, method, and non-transitory computer-readable storage medium for determining a speaker's speech segment from audio data.

[0005] Additional aspects will be set forth in part in the description which follows, and in part will be apparent from the specification, or may be learned by the practice of the disclosed embodiments.

[0006] According to one aspect of the disclosure, an electronic device is provided. The electronic device may include a microphone, at least one processor including a processing circuit, and a memory including one or more storage media storing instructions. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to segment audio data to acquire a plurality of frames while acquiring audio data using the microphone. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to acquire first vectors respectively corresponding to the plurality of frames. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine a speaker corresponding to each of the first vectors using groups each including the first vectors, obtained by grouping the first vectors and second vectors stored in the memory. The grouping of the first vectors and the second vectors may be performed by the at least one processor. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to store, in the memory, information determined using a speaker corresponding to each of the first vectors, the information representing a speaker of at least one time segment of the audio data. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to store, in the memory, the first vectors based on a total number of the first vectors and the second vectors being less than or equal to a specified number.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to remove at least one of the first vectors and the second vectors from the memory to adjust the total number of the first vectors and the second vectors to less than or equal to the specified number, based on a total number of the first vectors and the second vectors exceeding the specified number.

[0007] According to another aspect of the disclosure, a method of an electronic device including a microphone is provided. The method may include an operation of acquiring audio data using the microphone, while segmenting the audio data to acquire a plurality of frames. The method may include an operation of acquiring first vectors corresponding to each of the plurality of frames. The method may include an operation of determining a speaker corresponding to each of the first vectors using groups each of the first vectors obtained by grouping the first vectors and second vectors stored in a memory of the electronic device. The grouping of the first vectors and the second vectors may be performed by at least one processor of the electronic device. The method may include an operation of storing, in the memory, information determined using the speaker corresponding to each of the first vectors, the information indicating the speaker of at least one time section of the audio data. The method may include an operation of storing the first vectors in the memory based on a total number of the first vectors and the second vectors being less than or equal to a specified number. The method may include an operation of removing at least one of the first vectors and the second vectors from the memory, based on a total number of the first vectors and the second vectors exceeding the specified number, to adjust the total number to be less than or equal to the specified number.

[0008] In one embodiment, an electronic device may include a microphone, at least one processor including a processing circuit, and a memory including one or more storage media storing instructions. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain first vectors each corresponding to a plurality of frames included in a first time interval of audio data obtained by controlling the microphone. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform clustering of the first vectors and second vectors obtained in a second time interval of the audio data prior to the first time interval to determine a speaker of each of the plurality of frames. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine at least one vector, from among the first vectors and the second vectors, to be stored in memory for identifying at least one speaker associated with a third time segment of the audio data after the first time segment.

[0009] In one embodiment, a non-transitory computer-readable storage medium storing instructions may be provided. The instructions, when executed by an electronic device including a microphone, may cause the electronic device to obtain first vectors each corresponding to a plurality of frames included in a first time interval of audio data acquired by controlling the microphone. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform clustering of the first vectors and second vectors acquired in a second time interval of the audio data prior to the first time interval to determine a speaker of each of the plurality of frames. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine, from among the first vectors and the second vectors, at least one vector to be stored in a memory for identifying at least one speaker associated with a third time interval of the audio data subsequent to the first time interval.

[0010] One or more non-transitory computer-readable storage media storing one or more computer programs may be provided. The one or more programs may include computer-readable instructions that, when individually or collectively executed by one or more processors of an electronic device, cause the electronic device to perform an operation of obtaining a plurality of frames by separating audio data while obtaining audio data using a microphone of the electronic device. The one or more programs may include computer-readable instructions that, when individually or collectively executed by one or more processors of the electronic device, cause the electronic device to perform an operation of determining a speaker corresponding to each of the first vectors by using groups including the first vectors obtained by grouping first vectors and second vectors stored in a memory of the electronic device, wherein the operation of grouping the first vectors and the second vectors may be performed by at least one processor of the electronic device. The one or more programs may include computer-readable instructions that, when individually or collectively executed by one or more processors of the electronic device, cause the electronic device to perform an operation of storing, in the memory, information determined using a speaker corresponding to each of the first vectors, the information representing a speaker of at least one time segment of the audio data.The one or more programs may include computer-readable instructions that, when individually or collectively executed by one or more processors of the electronic device, cause the electronic device to perform an operation of removing at least one of the first vectors and the second vectors from the memory to adjust the total number of the first vectors and the second vectors to be equal to or less than the specified number, based on a total number of the first vectors and the second vectors being equal to or greater than the specified number.

[0011] Other aspects, advantages, and key features of the disclosure will become apparent to those skilled in the art from the following detailed description, which sets forth various embodiments of the disclosure, taken in conjunction with the accompanying drawings.

[0012] The above-described and other aspects, features, and advantages of some embodiments of the disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings, in which:

[0013] FIG. 1 illustrates the operation of an electronic device for recognizing one or more speakers from audio data, according to various embodiments;

[0014] FIG. 2 illustrates a block diagram of an electronic device according to one embodiment of the disclosure;

[0015] FIG. 3 illustrates a flow diagram of an electronic device according to one embodiment of the disclosure;

[0016] FIG. 4 illustrates the operation of an electronic device for detecting at least one voice segment from audio data, according to one embodiment of the disclosure;

[0017] FIG. 5 illustrates the operation of an electronic device for segmenting a plurality of frames from audio data, according to one embodiment of the disclosure;

[0018] FIG. 6 illustrates the operation of an electronic device for generating vectors each corresponding to a plurality of frames of audio data, according to one embodiment of the disclosure;

[0019] FIG. 7 illustrates the operation of an electronic device for calculating similarities between vectors associated with audio data, according to one embodiment of the disclosure;

[0020] FIG. 8 illustrates the operation of an electronic device performing clustering of vectors associated with audio data, according to one embodiment of the disclosure;

[0021] FIGS. 9A, 9B, 9C, and 9D illustrate the operation of an electronic device for filtering vectors associated with audio data according to various embodiments of the disclosure;

[0022] FIG. 10 illustrates a flow diagram of an electronic device according to one embodiment of the disclosure;

[0023] FIG. 11A and FIG. 11B illustrate a user interface (UI) displayed by an electronic device according to various embodiments of the disclosure;

[0024] Figure 12 is a block diagram of an electronic device within a network environment according to a start-up execution screen.

[0025] It should be noted that throughout the drawings, similar drawing symbols are used to describe the same or similar elements, features, and structures.

[0026] The following description, with reference to the accompanying drawings, is provided to facilitate a comprehensive understanding of the various embodiments of the disclosure defined by the claims and their equivalents. While it includes numerous specific details to aid understanding, it is intended to be considered exemplary only. Accordingly, those skilled in the art will recognize that various modifications and variations of the various embodiments described herein may be made without departing from the scope of the disclosure. Furthermore, descriptions of well-known functions and structures may be omitted for clarity and brevity.

[0027] The singular forms "a," "an," and "the" should be understood to include plural references unless the context explicitly dictates otherwise. Thus, "a component surface" includes reference to one or more such surfaces.

[0028] It is apparent that each block of the flowchart, and combinations of flowcharts, can be performed by one or more computer programs comprising computer-executable instructions. The entirety of the one or more computer programs may be stored in a single memory device, or the one or more computer programs may be divided into different portions stored in different memory devices.

[0029] Any of the functions or operations disclosed herein may be processed by a single processor or a combination of processors. A single processor or a combination of processors may include a circuit that performs processing, and / or an application processor (AP) (e.g., a central processing unit (CPU)), a communication processor (CP) (e.g., a modem), a graphical processing unit (GPU), a neural processing unit (NPU)9 (e.g., an artificial intelligence (AI) chip), a wireless-fidelity (Wi-Fi) chip, or a Bluetooth TM A circuit including a chip, a global positioning system (GPS) chip, a near field communication (NFC) chip, a connectivity chip, a sensor controller, a touch controller, a fingerprint sensor controller, a display driver integrated circuit (IC), an audio CODEC chip, a universal serial bus (USB) controller, a camera controller, an image processing IC, a microprocessor unit (MPU), a system on chip (SoC), an IC, or the like.

[0030] FIG. 1 illustrates the operation of an electronic device (101) for recognizing one or more speakers from audio data (110) according to various embodiments of the disclosure. Referring to FIG. 1, an electronic device (101) having the appearance of a mobile phone is illustrated. The electronic device (101) may have various form factors, such as a laptop PC (personal computer) (101-1), smartphones (e.g., a bar-type smartphone (101-2), a foldable-type smartphone (101-3), or a sliderable (or rollable) type smartphone (101-4)), a tablet PC (101-5), a head-mounted display (HMD) device (101-6), a headset (101-7) (or headphones), a ring (101-8), a watch (101-9), and other similar computing devices (not shown). The electronic device (101) may be referred to as a mobile device, a user equipment (UE), a multi-function device, a portable communication device, and / or a portable device. The form factor of the electronic device (101) is not limited to the form factors illustrated in FIG. 1. For example, the electronic device (101) may be included as an electronic control unit (ECU) in a vehicle (e.g., an electric vehicle (EV)). The electronic device (101) may have a form factor that is wearable by a user, such as a headset (101-7), a ring (101-8), and / or a watch (101-9) having the appearance of a watch, or may have a form factor that is implantable on a body part of a user. The embodiment is not limited thereto, and the electronic device (101) may have a form factor of an earbud and / or an earphone. An example of a hardware configuration included in the electronic device (101) is described with reference to FIG. 2.

[0031] According to one embodiment, the electronic device (101) can perform speaker diarization. By performing speaker diarization, the electronic device (101) can detect or identify one or more speakers associated with audio data (110). By performing speaker diarization, the electronic device (101) can detect or determine at least one time interval corresponding to a particular speaker within the entire time interval of the audio data (110). By performing speaker diarization, the electronic device (101) can detect or determine a plurality of time intervals corresponding to each of the plurality of speakers within the entire time interval of the audio data (110). For example, the electronic device (101) can allocate or match the entire time interval of the audio data (110) to each of the plurality of speakers. For example, an electronic device (101) that determines that a portion of audio data (110) corresponding to a specific time interval includes speech from a specific speaker may determine the specific time interval, and / or match the specific speaker to the specific time interval. To perform speaker separation, the electronic device (101) may recognize one or more speakers from the audio data (110) (e.g., speaker recognition). The operation of the electronic device (101) that performs speaker separation is described with reference to FIG. 3.

[0032] Referring to FIG. 1, an exemplary portion of information (120) generated by an electronic device (101) that performs speaker separation related to audio data (110) is illustrated. The information (120) may include an identifier (ID, identity) (e.g., 0) indicating a speaker of an utterance recorded in a specific time interval of the audio data (110). Referring to FIG. 1, the information (120) may indicate that an utterance of a speaker having an identifier of 0 is recorded in a first time interval from a time point of 200 milliseconds (msec) to a time point of 1000 milliseconds. The information (120) may indicate that an utterance of a speaker having an identifier of 3 is recorded in a second time interval from a time point of 1200 milliseconds to a time point of 3000 milliseconds. Information (120) may indicate that the utterance of a speaker with an identifier of 0 is recorded in a third time interval from a time point of 4100 milliseconds to a time point of 7340 milliseconds. Information (120) may indicate that the utterance of a speaker with an identifier of 5 is recorded in a fourth time interval from a time point of 8770 milliseconds to a time point of 10230 milliseconds.

[0033] According to one embodiment, the electronic device (101) may determine a speaker who uttered a sound recorded in a portion of a time interval of audio data (110) (e.g., a frame having a specified length in the time domain) by using information representing the characteristics of the sound recorded in the portion. The information may be referred to as feature information, a feature vector, an embedding vector, and / or an acoustic feature. The electronic device (101) that divides the audio data (110) into a plurality of frames may compare vectors corresponding to each of the plurality of frames to determine a speaker associated with each of the plurality of frames. The electronic device (101) may obtain or determine a speech interval of a specific speaker by grouping a plurality of frames that are adjacent to each other in the time domain and are associated with a specific speaker. The speech interval may mean a time interval in which the speech of a specific speaker is determined to have been recorded. The operation of an electronic device (101) that performs speaker separation using information corresponding to each of the frames divided from audio data (110) is described with reference to FIGS. 4 to 7.

[0034] The number of vectors obtained by the electronic device (101) from the audio data (110) may be related to the length of the audio data (110) because the vectors correspond to each of a plurality of frames of the audio data (110). For example, the number of vectors may be generally proportional to the length of the audio data (110). The electronic device (101) may calculate the similarities of the vectors to match the vectors to one or more speakers, or group the vectors into groups each corresponding to the speakers. When the electronic device (101) determines the similarities of two vectors, for a number of vectors, the electronic device (101) aWe can determine C2 similarities. For example, since the number of similarities is proportional to the square of a, As the number of vectors increases, the amount of computation performed for grouping and / or the amount of memory used to store the similarities may increase exponentially. Since the number of vectors increases when the length of the audio data (110) increases, the number of similarities may also increase as the length of the audio data (110) increases.

[0035] Since the number of similarities increases with the length of the audio data (110), the amount of computation required to calculate the similarities (or the resources of the electronic device (101) occupied to calculate the similarities) may increase with the length of the audio data (110). According to one embodiment, the electronic device (101) may periodically (or repeatedly) remove or discard vectors used in calculating the similarities in order to reduce or maintain the time and / or resources (e.g., the amount of computation) required to calculate the similarities. For example, the electronic device (101) may manage or compress the number of vectors to reduce the amount of computation required to perform speaker separation or to improve the performance of speaker separation. The operation of an electronic device (101) for managing (e.g., storing, and / or removing) vectors obtained from audio data (110) is described with reference to FIG. 8, FIG. 9a to FIG. 9d, and / or FIG. 10.

[0036] In one embodiment, a time interval obtained from audio data (110) by performing speaker separation and corresponding to one of a plurality of speakers may be used to execute a function related to the audio data (110). For example, using information (120), the electronic device (101) may extract or determine a portion of the audio data (110) to be used for performing speech-to-text (STT). For example, the electronic device (101) may perform STT on a speech segment of a specific speaker obtained by performing speaker separation, thereby generating or obtaining text representing an utterance (e.g., one or more natural language sentences) of the specific speaker in the speech segment. The text may be stored in association with the speech segment within the information (120).

[0037] As described above, according to one embodiment, the electronic device (101) can analyze audio data (110) independently from an external electronic device, such as a server (e.g., on-device analysis). Analysis of the audio data (110) performed solely by the electronic device (101) may include speaker separation. The electronic device (101), which has more limited resources than the server, may periodically (or iteratively) remove information necessary for speaker separation (e.g., vectors corresponding to frames segmented from the audio data (110)) in order to reduce or maintain the amount of computation required to perform speaker separation even as the length of the audio data (110) increases. For example, the electronic device (101) can perform speaker separation of relatively long (e.g., exceeding 10 hours) audio data at high speed in a relatively short time.

[0038] Below, the hardware configuration of the electronic device (101) of FIG. 1 is described with reference to FIG. 2.

[0039] FIG. 2 illustrates a block diagram of an electronic device (101) according to one embodiment of the disclosure. According to one embodiment, the electronic device (101) may include at least one of a processor (210), a memory (215), a display (220), or a microphone (225). The processor (210), the memory (215), the display (220), or the microphone (225) may be electronically and / or operably coupled with each other by an electronic component, such as a communication bus (202). Hereinafter, operably coupled electronic components may mean that a direct connection or an indirect connection is established between the electronic components, either wired or wireless, such that a first electronic component among the electronic components controls a second electronic component. Although illustrated based on different blocks, the embodiment is not limited thereto, and some of the electronic components of FIG. 2 (e.g., at least a portion of the processor (210) and the memory (215)) may be included in a single integrated circuit such as a system on a chip (SoC). The type and / or number of electronic components included in the electronic device (101) is not limited to those illustrated in FIG. 2. For example, the electronic device (101) may include only some of the electronic components illustrated in FIG. 2.

[0040] A processor (210) of an electronic device (101) according to one embodiment may include a circuit (e.g., a processing circuit) for processing data based on one or more instructions. The circuit for processing data may include, for example, an arithmetic and logic unit (ALU), a floating point unit (FPU), a field programmable gate array (FPGA), a central processing unit (CPU), a graphic processing unit (GPU), a neural processing unit (NPU), and / or an application processor (AP). For example, the number of processors (210) may be one or more. A processing circuit of a processor that loads (or fetches) instructions and performs calculations corresponding to the loaded instructions may be referred to as or referred to as a core circuit (or core). For example, the processor may have a multi-core processor structure including a plurality of core circuits, such as a dual core, a quad core, a hexa core, or an octa core. In one embodiment having a multi-core processor structure, the core circuits included in the processor (210) may be classified into a big core circuit (or performance core circuit) that processes instructions relatively quickly and a little core circuit (or efficiency core circuit) that processes instructions relatively slowly, depending on speed (e.g., clock frequency), power consumption, and / or cache memory. The functions and / or operations described with reference to the disclosure may be individually or collectively performed by one or more processing circuits included in the processor (210).

[0041] According to one embodiment, the memory (215) of the electronic device (101) may include a circuit for storing data and / or instructions input to and / or output from the processor (210). The memory may include, for example, volatile memory such as random-access memory (RAM) and / or non-volatile memory such as read-only memory (ROM). The non-volatile memory may be referred to as storage. The volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). The non-volatile memory may include, for example, at least one of programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, hard disk, compact disc, solid state drive (SSD), and embedded multimedia card (eMMC). The processor (210) of the electronic device (101) can execute instructions of the memory (215) within the electronic device (101) to perform functions and / or operations indicated by the instructions. For example, when the electronic device (101) includes at least one processor, the at least one processor can be configured to collectively or individually execute the instructions.

[0042] Referring to FIG. 2, the electronic device (101) may further include a display (220). The embodiment is not limited thereto, and the display (220) may be omitted depending on the form factor of the electronic device (101). The display (220) of the electronic device (101) may output visualized information to the user. For example, the display (220) may be controlled by a controller such as a GPU (graphic processing unit) and / or a processor (210) to output visualized information to the user. The display (220) may include a liquid crystal display (LCD), a plasma display panel (PDP), and / or one or more light emitting diodes (LEDs). The LEDs may include organic LEDs (OLEDs). The display (220) may include a flat panel display (FPD) and / or electronic paper. The embodiment is not limited thereto, and the display (220) may have an at least partially curved shape or a deformable shape. A display (220) having a deformable shape may be referred to as a flexible display.

[0043] In one embodiment, the electronic device (101) may include a sensor (e.g., a touch sensor panel (TSP)) for detecting an external object (e.g., a user's finger) on the display (220). For example, using the TSP, the electronic device (101) may detect an external object that is in contact with the display (220) or floating on the display (220). In response to detecting the external object, the electronic device (101) may execute a function associated with a specific visual object among visual objects displayed on the display (220) that corresponds to a location of the external object on the display (220).

[0044] In one embodiment, the electronic device (101) may include a microphone (225) that outputs an electrical signal representing vibration of the atmosphere. For example, the electronic device (101) may output audio data (e.g., audio data (110) of FIG. 1) (or an audio signal) including a user's speech using the microphone (225). The user's speech included in the audio signal may be converted into information in a format recognizable by the processor (210) of the electronic device (101) based on a speech recognition model and / or a natural language understanding model. For example, the electronic device (101) may recognize the user's speech and execute one or more functions among a plurality of functions that can be provided by the electronic device (101).

[0045] Although not illustrated, in one embodiment, the electronic device (101) may include output means for outputting information in a form other than a visualized form. For example, the electronic device (101) may include a speaker for outputting an acoustic signal. For example, the electronic device (101) may include a motor for providing haptic feedback based on vibration.

[0046] In one embodiment, one or more instructions (or commands) representing operations and / or actions to be performed on data by the processor (210) may be stored in the memory (215) of the electronic device (101). A set of one or more instructions may be referred to as firmware, an operating system, a process, a routine, a sub-routine, a program, and / or a software application (hereinafter, “application”). For example, the electronic device and / or the processor may perform at least one of the operations of FIG. 3 and / or FIG. 10 when a set of a plurality of instructions distributed in the form of an operating system, firmware, driver, and / or application is executed. Hereinafter, the fact that an application is installed on an electronic device may mean that one or more instructions provided in the form of an application are stored in the memory of the electronic device, and that the one or more applications are stored in a format (e.g., a file having an extension specified by an operating system of the electronic device) that is individually or collectively executable by at least one processor (e.g., processor (210)) of the electronic device.

[0047] Referring to FIG. 2, a recording application (230) may be installed in the memory (215) of the electronic device (101). Instructions included in the recording application (230) may include subroutines, such as a voice activity detector (231) and / or a speaker segmenter (232), which are classified according to the functions performed by the instructions. The processor (210) may execute the recording application (230) to perform speaker segmentation of audio data (e.g., audio data (110) of FIG. 1) obtained from a microphone (225) and / or stored in the memory (215).

[0048] In order to perform speaker separation of audio data, the processor (210) may execute a voice activity detector (231) using the audio data. By executing the voice activity detector (231), the processor (210) may detect or determine a time interval (e.g., a voice interval and / or a voice section) in which a voice generated by a human exists. The operation of the processor (210) executing the voice activity detector (231) is described with reference to FIG. 4. By executing the voice activity detector (231), the processor (210) may execute a speaker segmenter (232) using at least one voice interval obtained from the audio data.

[0049] By executing the speaker segmenter (232), the processor (210) can perform speaker separation in at least a portion of audio data (e.g., at least one voice segment). The processor (210) can segment the voice segment to obtain a plurality of frames. An operation of the processor (210) to obtain the plurality of frames is described with reference to FIG. 5. All of the plurality of frames can have a specified length and can at least partially overlap each other in the time domain. The processor (210) can execute the feature vector determiner (241) using the plurality of frames.

[0050] The processor (210) executing the feature vector determiner (241) may generate or determine information corresponding to each of the plurality of frames. The information may include a vector including a plurality of numerical values ​​representing acoustic characteristics of the frame as elements. The information generated based on the execution of the feature vector determiner (241) may be generated to verify a speaker (e.g., text independent speaker verification (TISV)). The vector obtained by executing the feature vector determiner (241) may include implicit information for specifying a speaker recorded (or captured) in the frame as information generated for speaker separation. The operation of the processor (210) executing the feature vector determiner (241) is described with reference to FIG. 6. The processor (210) may execute a speaker cluster determiner (242) using the information corresponding to each of the frames.

[0051] The processor (210) executing the speaker cluster determiner (242) may perform clustering on the frames using information (e.g., vectors) corresponding to each frame. For example, the processor (210) may cluster the vectors of the frames to determine one or more speakers and the relationship between the vectors. Using the determined relationship, the processor (210) may match at least a portion of the audio data (e.g., a portion corresponding to a group of frames) with one or more speakers, thereby generating information about a point in time (or time) within the audio data at which the speaker spoke. For example, each of one or more groups of vectors generated by the clustering may correspond to one or more speakers. Using the one or more groups, the processor (210) may obtain or generate information (e.g., information (120) of FIG. 1) representing the result of performing speaker separation from the audio data.

[0052] In one embodiment, clustering using the speaker cluster determiner (242) may be performed when the number of feature vectors to be clustered is greater than or equal to a specified number. While receiving audio data in real time, the processor (210) may accumulate and store feature vectors in the memory (215). When the number of feature vectors stored in the memory (215) is greater than or equal to a specified number, the processor (210) may determine to perform clustering on at least some of the feature vectors. In one embodiment, the operation of comparing the number of feature vectors with the specified number and determining whether to perform the clustering may be performed in units of voice sections detected by the voice activity detector (231).

[0053] In one embodiment, the processor (210) executing the speaker cluster determiner (242) may store information indicating the result of performing speaker separation in the memory (215). The processor (210) may store vectors used to perform speaker separation in the vector storage area (233). The vector storage area (233) may be formed in at least a portion of the memory (215) (e.g., volatile memory and / or non-volatile memory) by the processor (210) executing the recording application (230) (or the speaker segmenter (232)). For example, the vector storage area (233) may be formed in the volatile memory or may be maintained while the processor (210) executes the recording application (230) to perform speaker separation.

[0054] As described above with reference to FIG. 1, as the length of audio data increases, the number of vectors generated from the audio data may increase. While performing speaker separation of audio data, the size of the vector storage area (233) may be gradually increased to store the vectors. In one embodiment, the processor (210) may maintain the number of vectors stored in the vector storage area (233) below a specified number (or threshold number).

[0055] For example, the processor (210) may select or extract one or more representative vectors associated with one or more recognized speakers to perform speaker separation. The processor (210) may store only the one or more representative vectors among vectors obtained from audio data in the vector storage area (233). While performing speaker separation gradually from the time of audio data, the processor (210) may filter out other vectors that are different from the representative vectors among the vectors accumulated in the vector storage area (233), thereby maintaining the number of vectors stored in the vector storage area (233) below a threshold number.

[0056] Since the number of vectors stored in the vector storage area (233) is maintained below a critical number, the computational effort required to calculate the similarities of the vectors can be maintained constant for each time interval (e.g., voice interval) from the start to the end of audio data. For example, an electronic device (101) with limited resources can complete speaker separation for audio data more quickly.

[0057] Below, with reference to FIG. 3, the operation of the processor (210) for speaker separation, described with reference to FIG. 2, is described.

[0058] FIG. 3 illustrates a flowchart of an electronic device according to an embodiment of the disclosure. The electronic device (101) of FIGS. 1 and 2, and / or the processor (210) of FIG. 2, may perform operations of the electronic device described with reference to FIG. 3. The order in which the operations of FIG. 3 are performed is not limited to the order illustrated in FIG. 3, and may be performed substantially simultaneously or in a different order than the order illustrated in FIG. 3.

[0059] Referring to FIG. 3, in operation (310), according to one embodiment, a processor of an electronic device may detect a voice segment in which a voice appears to have been recorded. The processor, which detects audio data received via a microphone (e.g., microphone (225) of FIG. 2) of the electronic device and / or stored in a memory (e.g., memory (215) of FIG. 2), may perform operation (310). While receiving audio data via the microphone (e.g., while recording using the microphone is being performed), the processor may perform operation (310) to detect voice activity indicated by the audio data. Operation (310) may be performed by a processor that executes voice activity detector (231) of FIG. 2. By performing operation (310), the processor may detect one or more voice segments in the entire time interval of the audio data (e.g., the time interval between the time point at which the audio data is received via the microphone and the present time point).

[0060] While receiving audio data, the processor may repeatedly perform the operations of FIG. 3. For example, the processor may newly detect a voice segment from a portion of audio data received between the last time point at which a voice segment was detected by performing operation (310) and the current time point. For the newly detected voice segment, the processor may further perform the operations following operation (310).

[0061] Referring to FIG. 3, in operation (320), a processor of an electronic device according to an embodiment may obtain vectors corresponding to each frame in a voice section of operation (310). The processor may perform operation (320) by executing the speaker segmenter (232) (or feature vector determiner (241)) of FIG. 2. The frames of operation (320) may be generated or determined by segmenting the voice section of operation (310). The processor may perform operation (320) to obtain vectors that match the frames one-to-one. The vector may be a data structure that includes a one-dimensional array having a specified number of elements. By performing the TISV algorithm, the processor may obtain or generate the vectors of operation (320).

[0062] In one embodiment, in which a processor acquires vectors included in a voice segment of audio data acquired by controlling a microphone, the processor may acquire vectors of operation (320) while acquiring audio data using the microphone. Since operation (320) is performed based on detecting a voice segment of operation (310), if the audio data includes a plurality of voice segments, the processor may repeatedly acquire vectors of operation (320) when each of the plurality of voice segments is detected.

[0063] Referring to FIG. 3, in operation (330), according to one embodiment, a processor of an electronic device may perform clustering on vectors of operation (320). Operation (330) may be performed by executing a speaker segmenter (232) (e.g., a speaker cluster determiner (242)) of FIG. 2. The processor may perform clustering of operation (330) based on the number of vectors obtained by performing operation (320) (e.g., when the number exceeds a specified number). When acquiring audio data in real time, the processor may perform operation (330) when detecting a voice segment of operation (310) from audio data acquired in real time.

[0064] When performing clustering on the first voice segment, the processor may perform clustering on first vectors obtained in the first voice segment and second vectors obtained in at least one second voice segment prior to the first voice segment. Clustering on the second voice segment may be completed before performing clustering on the first voice segment. At least one of the second vectors may be stored in a memory of the electronic device (e.g., memory (215) and / or vector storage area (233) of FIG. 2) when performing clustering on the first voice segment. By performing clustering on the first vectors and the second vectors, the processor may determine a speaker of each of the frames included in the first voice segment.

[0065] Referring to FIG. 3 , in operation (340), a processor of an electronic device according to an embodiment may determine speakers corresponding to each of the groups of vectors determined by clustering. The processor may assign an identifier (e.g., an ID and / or an index value) uniquely assigned to the speaker to each of the groups of operation (340). When performing operation (340) for a first voice segment, the processor may load identifiers of one or more speakers that were detected by performing speaker separation for at least one second voice segment prior to the first voice segment. The loaded identifiers may be used to determine a speaker corresponding to each of the frames of the first voice segment. Since the vectors of operation (340) correspond to each of the frames included in the voice segment detected by operation (310), the processor performing operation (340) may determine speakers for each of the frames.

[0066] In one embodiment, the processor may assign a speaker identifier to a voice segment. If frames included in a voice segment correspond to different speakers, the processor may determine the speaker corresponding to the largest number of frames as the representative speaker of the voice segment. The processor may assign the identifier of the determined representative speaker to the voice segment.

[0067] The result of determining the speakers of the operation (340) may be stored in the memory of the electronic device. For example, the processor may store in the memory information (e.g., information (120) of FIG. 1) indicating the speakers of at least one time segment of audio data (e.g., a time segment corresponding to a group of vectors grouped by clustering). When performing the operation (340) for the first voice segment, the processor may combine or merge the information determining one or more speakers associated with the first voice segment with information indicating one or more speakers associated with at least one second voice segment preceding the first voice segment.

[0068] Referring to FIG. 3, in operation (350), a processor of an electronic device according to an embodiment may selectively store vectors using a number associated with resources of the electronic device occupied for clustering. The vectors selected based on operation (350) may be used for, or stored for, clustering of vectors that is additionally performed after the clustering of operation (330). In operation (350), a processor of an electronic device according to an embodiment may selectively store vectors for clustering. When performing operation (350) for a first voice segment, the processor may perform filtering on first vectors obtained from the first voice segment and second vectors associated with at least one second voice segment prior to the first voice segment. For example, the processor may extract or select representative vectors from among the first vectors and the second vectors to be used for speaker separation in a third voice segment following the first voice segment. The selected representative vectors can be clustered together with the third vectors obtained from the third voice section when performing speaker separation in the third voice section.

[0069] For example, while acquiring audio data, in order to maintain a number of vectors associated with the audio data and stored in a memory at a specified number, the processor may compare the specified number with the number of the first vectors and the second vectors. Based on the number of the first vectors and the second vectors being less than or equal to the specified number, the processor may store the first vectors in the memory. Based on the number of the first vectors and the second vectors exceeding the specified number, the processor may store the specified number of vectors among the first vectors and the second vectors in the memory.

[0070] As described above, the operations of FIG. 3 may be repeatedly performed while receiving audio data using a microphone. In response to an input indicating a cessation of audio data acquisition, the processor may cease repeatedly performing the operations of FIG. 3 (e.g., cessation and / or completion of speaker separation). The processor, which has ceased speaker separation, may store or output the results of determining one or more speakers from the audio data, and time intervals (e.g., utterance intervals) corresponding to each of the one or more speakers. For example, the processor may display a screen related to information representing the results on a display (e.g., display (220) of FIG. 2). An example of such a screen is described with reference to FIG. 11A and / or FIG. 11B.

[0071] Below, each of the operations of FIG. 3 is exemplarily explained with reference to FIGS. 4 to 8.

[0072] FIG. 4 illustrates an operation of an electronic device for detecting at least one voice segment from audio data (410), according to various embodiments of the disclosure. The electronic device (101) of FIGS. 1 and 2, and / or the processor (210) of FIG. 2, may perform the operation of the electronic device described with reference to FIG. 4. The operation of the electronic device described with reference to FIG. 4 may be related to at least one of the operations of FIG. 3 (e.g., operation (310)).

[0073] Referring to FIG. 4, an electronic device that receives audio data (410) (e.g., from a microphone (225) of FIG. 2) can detect voice segments (V1, V2, V3, V4) in which a voice appears to have been recorded from the audio data (410). Referring to FIG. 4, the entire time segment of the audio data (410) can be classified into voice segments (V1, V2, V3, V4) and segments (N1, N2, N3, N4) in which a voice does not appear to have been recorded. The classification can be performed by an electronic device that executes the voice activity detector (231) of FIG. 2. Each of the segments (N1, N2, N3, N4) can be referred to as a non-voice segment and / or a noise segment.

[0074] An electronic device that detects voice segments (V1, V2, V3, V4) from the entire time interval of audio data (410) can obtain or generate information indicating start points (e.g., starting point detection (SPD)) and end points (e.g., end point detection (EPD)) of the voice segments (V1, V2, V3, V4). For example, using a timestamp starting from the start point of the audio data (410), the electronic device can obtain or determine timestamps corresponding to each of the start points and end points of the voice segments (V1, V2, V3, V4).

[0075] In order to detect the voice segments (V1, V2, V3, V4), the electronic device may perform a voice activity detection (VAD) (or speech activity detection, and / or speech detection) algorithm. The VAD algorithm may include a computational model for noise reduction and / or extracting characteristic information of the audio data (410) (e.g., information related to the waveform of the audio data (410), such as a spectrogram). By performing the calculations indicated by the VAD algorithm, the electronic device may identify the start and end points of the voice segments (V1, V2, V3, V4).

[0076] An electronic device that detects at least one of the voice segments (V1, V2, V3, V4) can acquire frames by dividing the detected at least one voice segment. For example, at each of the time points at which each of the voice segments (V1, V2, V3, V4) is detected, the electronic device can perform operation (320) of FIG. 3 for the detected voice segment. Hereinafter, with reference to FIG. 5, an operation of an electronic device that divides a plurality of frames from a specific voice segment will be described.

[0077] FIG. 5 illustrates an operation of an electronic device for segmenting audio data (410) into a plurality of frames (510, 520, 530) according to various embodiments of the disclosure. The electronic device (101) of FIGS. 1 and 2, and / or the processor (210) of FIG. 2, may perform the operation of the electronic device described with reference to FIG. 5. The operation of the electronic device described with reference to FIG. 5 may be related to at least one of the operations of FIG. 3 (e.g., operation (320)).

[0078] Referring to FIG. 5, the operation of an electronic device for dividing a voice section (V1) of audio data (410) into multiple frames is illustrated. The embodiment is not limited thereto, and the electronic device may divide frames of other voice sections (V2, V3, V4). The operation of the electronic device for dividing frames may be referred to as a segment.

[0079] Within the voice segment (V1), the electronic device can divide frames (510, 520, 530) having a specified size (or window size) (e.g., 1.5 seconds). For example, the frame (510) may correspond to a time segment having a length of 1.5 seconds from the start of the voice segment (V1). Other frames (520, 530) within the voice segment (V1) different from the frame (510) may also have a length of 1.5 seconds. The specified size of the frames (510, 520, 530) is not limited to the 1.5 seconds exemplified above.

[0080] The electronic device may segment frames (510, 520, 530) spaced apart by a specified interval (or step size, and / or offset) from the voice segment (V1). For example, the difference between the start time of a frame (510) and the start time of a frame (520) following the frame (510) may be S = 0.75 seconds. Similarly, the difference between the start time of a frame (520) and the start time of a frame (530) following the frame (520) may also be S = 0.75 seconds. Referring to FIG. 5, since the length of each of the frames (510, 520, 530) is longer than the interval (S) between the frames (510, 520, 530), the frames (510, 520, 530) may at least partially overlap each other in the time domain.

[0081] Referring to FIG. 5, three frames (510, 520, 530) divided from a voice segment (V1) are exemplarily illustrated. Depending on the length of the voice segment (V1), the number of frames divided from the voice segment (V1) may vary. For example, an electronic device may obtain or generate 15 frames by dividing a voice segment (V1) having a length of 12 seconds into frames having a size of 1.5 seconds and an interval of 0.75 seconds.

[0082] An electronic device that divides a voice section (V1) into multiple frames (e.g., frames (510, 520, 530)) can obtain or generate feature information (e.g., vectors) for each of the multiple frames. Hereinafter, with reference to FIG. 6, the operation of the electronic device that generates vectors from the multiple frames is described.

[0083] FIG. 6 illustrates an operation of an electronic device for generating vectors (610, 620, 630) corresponding to a plurality of frames (510, 520, 530) of audio data, respectively, according to various embodiments of the disclosure. The electronic device (101) of FIGS. 1 and 2, and / or the processor (210) of FIG. 2, may perform the operation of the electronic device described with reference to FIG. 6. The operation of the electronic device described with reference to FIG. 6 may be related to at least one of the operations of FIG. 3 (e.g., operation (320)).

[0084] Referring to FIG. 6, a state of an electronic device that obtains vectors (610, 620, 630) from a plurality of frames (e.g., frames (510, 520, 530)) segmented from a voice segment (V1) of FIGS. 4 and 5 is illustrated. The embodiment is not limited thereto, and the electronic device may generate or determine vectors from one or more voice segments detected from audio data (e.g., audio data (110) of FIG. 1 and / or audio data (410) of FIGS. 4 and 5) and a plurality of frames segmented from the one or more voice segments.

[0085] In one embodiment, the electronic device may execute a TISV model using each of the frames (510, 520, 530). The TISV model may be a computational model trained to generate feature information, such as vectors, from the frames. For example, the TISV model may include an artificial neural network. The artificial neural network is a computational model for simulating neural activity (e.g., reasoning, recognition, and / or classification) of a living organism, including a human, and may include instructions for performing a plurality of calculations represented by the computational model, and resources used for the plurality of calculations. The resources may include a plurality of coefficients (e.g., weights and / or filter matrices) used for executing the artificial neural network. In one embodiment, by executing the TISV model trained to distinguish speakers, the electronic device may obtain or generate vectors (610, 620, 630) corresponding to each of the frames (510, 520, 530).

[0086] In one embodiment, the electronic device can execute a TISV model using raw features obtained from frames (510, 520, 530). For example, the electronic device can obtain or generate acoustic information (e.g., information representing a spectrogram, a cepstrum, a spectrum, a pitch, a zero-crossing rate) of the frame (510). Using the TISV model into which the acoustic information is input, the electronic device can obtain a feature vector (610) corresponding to the frame (510). The acoustic information can include information on frequency characteristics of a waveform signal of the frame (510) over time. The acoustic information can be used as an input for a TISV model based on an artificial neural network. Output data of the TISV model into which the acoustic information is input can include the feature vector. The TISV model can be trained to reduce distances in vector space between feature vectors of the same speaker and to increase distances in vector space between feature vectors of different speakers.

[0087] The electronic device can determine vectors (610, 620, 630) corresponding to the frames (510, 520, 530), respectively. The vectors (610, 620, 630) can have a specified dimension and / or a specified size. All of the vectors (610, 620, 630) can include one-dimensional numeric values ​​(e.g., floating point numbers and / or integers) as elements. The electronic device obtaining any one of the vectors (610, 620, 630) can include an operation of storing an array of elements included in the vector in a memory (e.g., the memory (215) of FIG. 2 and / or the vector storage area (233)). The vectors (610, 620, 630) obtained by executing the TISV model may include information about the speakers of each of the frames (510, 520, 530) corresponding to the vectors (610, 620, 630).

[0088] An electronic device that obtains vectors (610, 620, 630) of frames (510, 520, 530) can perform clustering on the vectors (610, 620, 630). Hereinafter, with reference to FIG. 7, the operation of the electronic device that performs clustering of frames corresponding to each of the frames is described.

[0089] FIG. 7 illustrates the operation of an electronic device for calculating similarities between vectors associated with audio data (710) according to one embodiment of the disclosure. The electronic device (101) of FIGS. 1 and 2 and / or the processor (210) of FIG. 2 may perform the operation of the electronic device described with reference to FIG. 7. The operation of the electronic device described with reference to FIG. 7 may be related to at least one of the operations of FIG. 3 (e.g., operation (330)).

[0090] Referring to FIG. 7, a state of an electronic device that generates k vectors from audio data (710) is illustrated. The audio data (710) may correspond to one voice segment detected by the electronic device. The k vectors (e.g., v1 - vk) generated from the audio data (710) may each correspond to k frames segmented from the audio data (710). According to one embodiment, the electronic device may calculate similarities between the k vectors. The similarities may include cosine similarity (or cosine distance), Euclidean distance, and / or Manhattan distance of the vectors.

[0091] In one embodiment, the electronic device can calculate similarities between any two vectors among k vectors. The electronic device can obtain or generate a similarity matrix (720) including the similarities as elements. The size of the similarity matrix (720) can be k Х k. The elements of the x row and y column of the similarity matrix (720) can represent the similarity between the x-th vector and the y-th vector among the k vectors. Since the elements of the x row and y column and the elements of the y row and x column are identical to each other, the similarity matrix (720) can be a symmetric matrix.

[0092] Using the similarity matrix (720), the electronic device can perform clustering of k vectors. The clustering of vectors can include an operation of determining groups (731, 732, 733, 734) of vectors. The electronic device can create or obtain the groups (731, 732, 733, 734) by grouping vectors that are adjacent in the time domain and have a similarity exceeding a threshold similarity by the similarity matrix (720). Each of the groups (731, 732, 733, 734) can indicate that a time interval including frames corresponding to vectors corresponding to the group includes utterances of a specific speaker. Each of the groups (731, 732, 733, 734) can be determined as a time interval in which utterances of a specific speaker are detected. For example, groups (731, 732, 733, 734) may correspond to each of the time intervals included in the information (120) of FIG. 1.

[0093] Although the operation of the electronic device for calculating similarities between vectors, such as cosine similarity, has been described, the operation of clustering vectors is not limited thereto. For example, the electronic device may perform a clustering algorithm, such as K-Means, K-Medoids, CLARA (Clustering Large Applications), CLARANS (Clustering Large Applications based on RANdomized Search), and / or partitioning, to determine groups (731, 732, 733, 734) of vectors.

[0094] As described above with reference to FIG. 1, as the length of the audio data (710) increases, the number k of vectors divided from the audio data (710) may increase. As the number k of vectors increases, the size of the similarity matrix (720) may increase. Since the size of the similarity matrix (720) is k Х k, the size of the similarity matrix (720) may increase according to the number of vectors. The electronic device may maintain the size of the similarity matrix (720), which is at least temporarily stored in the memory, to be less than a critical size, and may at least partially remove vectors generated from the audio data (710) in order to maintain or reduce the amount of computation required to calculate the similarity matrix (720).

[0095] Hereinafter, with reference to FIG. 8, the operation of an electronic device that performs speaker separation using groups of vectors (731, 732, 733, 734) and stores vectors included in the groups (731, 732, 733, 734) is described.

[0096] FIG. 8 illustrates an operation of an electronic device that performs clustering of vectors (810) associated with audio data, according to one embodiment of the disclosure. The electronic device (101) of FIGS. 1 and 2, and / or the processor (210) of FIG. 2, may perform the operation of the electronic device described with reference to FIG. 8. The operation of the electronic device described with reference to FIG. 8 may be related to at least one of the operations of FIG. 3 (e.g., operations (340, 350)).

[0097] Referring to FIG. 8, vectors (810) obtained from audio data (e.g., audio data (110) of FIG. 1, audio data (410) of FIGS. 4 and 5, and / or audio data (710) of FIG. 7) are illustrated. Using the similarities of the vectors (810) (e.g., elements of the similarity matrix (720) of FIG. 7), the electronic device can identify or detect multiple speakers. The electronic device can assign an identifier of one of the identified multiple speakers to each of the vectors (810). The textures of each of the vectors (810) of FIG. 8 can represent the speaker (or ID of the speaker) corresponding to each of the vectors (810). For example, in a state where four speakers are detected, the electronic device may assign one of the IDs of the speakers (e.g., ID 0, ID 3, ID 5, and / or ID 7) to each of the vectors (810).

[0098] Referring to Fig. 8, clusters of vectors (821, 822, 823, 824) corresponding to each speaker are illustrated. Groups of vectors (731, 732, 733, 734) of Fig. 7 may correspond to consecutive utterance segments of a specific speaker in the time domain. Clusters (821, 822, 823, 824) may be obtained or generated by grouping vectors (810) according to speakers.

[0099] According to one embodiment, the electronic device may selectively store one or more vectors in each of the clusters (821, 822, 823, 824), thereby maintaining the number of vectors stored in the memory and / or the amount of computation required for clustering the vectors, while maintaining the performance of speaker separation of audio data to be additionally received.

[0100] For example, the electronic device can extract the centroid of each of the clusters (821, 822, 823, 824). The electronic device can optionally store vectors adjacent to the extracted centroid. For example, within a cluster (821) corresponding to a speaker having an identifier of ID 0, the electronic device can calculate the sum of the similarities of each of the m vectors within the cluster (821). For example, the sum of the similarities of the k-th vector within the cluster (821) can be the sum of the similarities of the k-th vector and each of the m - 1 remaining vectors. Similarly, the electronic device can calculate or obtain the sums of the similarities of the m vectors within the cluster (821).

[0101] For example, as the distance between the k-th vector and the m - 1 remaining vectors becomes closer, the sum of the similarities of the k-th vector may increase. According to one embodiment, the electronic device may store only n vectors (e.g., a natural number n smaller than m) having the largest sum of similarities within the cluster (821) in the memory (e.g., the memory (215) of FIG. 2 and / or the vector storage area (233)). When n is a specified value and the number of vectors included in the cluster (821) is smaller than n, the processor may store all vectors included in the cluster (821) in the memory. The electronic device may perform a similar operation on vectors included in other clusters (822, 823, 824) that are different from the cluster (821).

[0102] Referring to FIG. 8, vectors included in each of the clusters (821, 822, 823, 824) are shown sorted according to the sum of the similarities corresponding to each of the vectors. In each of the clusters (821, 822, 823, 824), vectors shown on the left side of the drawing may have relatively large sums of similarities. In each of the clusters (821, 822, 823, 824), when selectively storing n vectors having large sums of similarities, the electronic device may store only the vectors included in the set (830). Among the vectors (810), other vectors not included in the set (830) may be removed from the memory of the electronic device (e.g., the memory (215) of FIG. 2, and / or the vector storage area (233)). In one embodiment of FIG. 8, where four speakers are detected, the electronic device may selectively store 4 Х n vectors.

[0103] Although the operation of the electronic device for selectively storing vectors having a large sum of similarities has been described, the embodiment is not limited thereto. For example, the electronic device may randomly select vectors less than or equal to a specified number from among the vectors (810). For example, the electronic device may store the center vector of vectors included in a cluster corresponding to a specific speaker (e.g., one of the clusters (821, 822, 823, 824)) in the memory of the electronic device. The vectors selectively stored in the memory and associated with each of the clusters (821, 822, 823, 824) may be referred to as representative vectors for each of the clusters (821, 822, 823, 824).

[0104] The operation of the electronic device filtering vectors or selectively storing vectors, as described with reference to FIG. 8, may be conditionally performed based on the number of vectors accumulated in the electronic device (e.g., stored in a memory of the electronic device). For example, if the number of vectors stored in the electronic device exceeds a threshold number, the electronic device may perform the operation described with reference to FIG. 8 to store vectors equal to or less than the product of the number of speakers associated with the audio data and an upper limit (e.g., n) of the number of vectors associated with a specific speaker. The embodiment is not limited thereto, and the electronic device may repeatedly filter vectors or selectively store vectors at specified intervals while acquiring audio data.

[0105] In one embodiment, the threshold number used to adjust the number of vectors stored in the electronic device may be adjusted according to the state of the electronic device. For example, the electronic device may determine whether to adjust the threshold number by using the period required to perform clustering of vectors less than or equal to the threshold number. The electronic device may schedule, predict, or determine the amount of computation related to clustering and / or the period required to perform calculations related to clustering by using the state of the processor (e.g., the activity state of each of the big core circuit and / or the little core circuit, and / or the temperature) used for clustering (e.g., the processor 210 of FIG. 2). The electronic device may decrease the threshold number when the period required to perform calculations related to clustering increases (e.g., when the period is predicted to exceed a specified period). When the threshold number is decreased, the number of vectors maintained in the memory of the electronic device may be reduced, and the amount of computation required to calculate the similarity of the vectors may be reduced.

[0106] While an embodiment has been described in which the threshold number is adjusted based on the state of the processor, the embodiment is not limited thereto. The electronic device may increase or decrease the threshold number depending on the state of charge (SOC) of the battery and / or the temperature of the electronic device. For example, if the SOC of the battery is below a threshold (e.g., a threshold for operation in a low power state), the electronic device may decrease the threshold number and / or at least temporarily suspend speaker separation. For example, if the battery is charged or the SOC of the battery exceeds a threshold, the electronic device may resume speaker separation, maintain the threshold number, or increase the threshold number. For example, if the temperature of the electronic device exceeds a threshold (e.g., a threshold set for throttling), the electronic device may decrease the threshold number and / or at least temporarily suspend speaker separation. For example, if the temperature of the electronic device is below a threshold, the electronic device may resume speaker separation, maintain the threshold number, or increase the threshold number.

[0107] In one embodiment, an electronic device may communicate with a server configured to perform speaker separation to request speaker separation. In the example above, if a threshold number decreases to a specified lower limit or speaker separation is at least temporarily suspended, the electronic device may request speaker separation for audio data from the server. If the electronic device receives audio data via a microphone (e.g., microphone 225 of FIG. 2), the electronic device may transmit a signal related to the audio data (e.g., a signal including a bitstream for streaming the audio data) to the server along with the request. To perform speaker separation using the server, at least one communication link may be established between the electronic device and the server.

[0108] In one embodiment, the electronic device may apply a threshold number for retaining vectors differently depending on the plurality of speakers associated with the audio data. For example, among the vectors of the cluster (821) of the user having an identifier of ID 0, the electronic device may store, filter, or select vectors less than or equal to a first threshold number. In the above example, among the vectors of the cluster (822) having an identifier of ID 3, the electronic device may selectively store a second threshold number of vectors determined independently from the first threshold number. The first threshold number and / or the second threshold number may vary depending on the distribution characteristics of the feature vectors corresponding to each of the speakers within the vector space. For example, if the vectors included in the cluster (821) have relatively dense locations within the vector space, the electronic device may determine the first threshold number to be less than the second threshold number. For example, if vectors associated with a user having an identifier of ID 0 corresponding to a cluster (821) are relatively far apart from other clusters (822, 823, 824) within the vector space, the electronic device may adjust the first threshold number to a value smaller than another threshold number (e.g., a second threshold number).

[0109] In one embodiment, the electronic device may set a threshold number corresponding to a primary user of the electronic device (e.g., a user identified by account information logged into the electronic device) to be smaller than a threshold number used to maintain vectors of other speakers. For example, the electronic device may perform speaker separation using vectors pre-stored in memory and corresponding to the primary user, and may maintain vectors corresponding to the primary user among vectors identified from audio data in a number smaller than the threshold number corresponding to other speakers. In this case, the vectors used to distinguish the primary user may be stored in a relatively small number in the memory of the electronic device.

[0110] As described above with reference to the similarity matrix (720) of FIG. 7, since the number of vectors is proportional to the number of similarities between vectors, the amount of computation required to calculate similarities can change relatively significantly depending on changes in the threshold number. The electronic device can effectively adjust the amount of computation required for clustering vectors based on similarities by adjusting the threshold number.

[0111] Hereinafter, with reference to FIGS. 9A to 9D, changes in the number of vectors stored or maintained in an electronic device while receiving audio data are described.

[0112] FIGS. 9A, 9B, 9C, and 9D illustrate the operation of an electronic device for filtering vectors associated with audio data according to various embodiments of the disclosure. The electronic device (101) and / or the processor (210) of FIGS. 1 and 2 may perform the operations of the electronic device described with reference to FIGS. 9A to 9D . The operations of the electronic device described with reference to FIGS. 9A to 9D may be related to the operations of FIG. 3 .

[0113] Referring to FIGS. 9A to 9D , states (901, 902, 903, 904) of vectors stored in an electronic device at different points in time while receiving audio data are illustrated. Referring to FIGS. 9A to 9D , a simplified distribution of vectors in a two-dimensional vector space is illustrated. The dimensions of the vectors and / or the vector space are not limited to the embodiments of FIGS. 9A to 9D .

[0114] Referring to Fig. 9a, the state (901) of vectors clustered by an electronic device at a first point in time when audio data is received is illustrated. Since the vectors are generated by a model for distinguishing speakers (e.g., a TISV model), the distance between vectors within the vector space may be related to whether each of the vectors corresponds to a speaker.

[0115] In state (901) of FIG. 9A, where a voice segment and / or an end point of the voice segment is detected, the electronic device may perform clustering of vectors. For example, the electronic device may identify or generate clusters (911, 912, 913) corresponding to three speakers, respectively. Frames corresponding to vectors included in a first cluster (911) (e.g., vectors expressed in the shape of a triangle (△)) may be determined to include utterances of a first speaker associated with the first cluster (911). Frames corresponding to vectors included in a second cluster (912) (e.g., vectors expressed in the shape of a square (□)) may be determined to include utterances of a second speaker associated with the second cluster (912). Frames corresponding to vectors included in the third cluster (913) (e.g., vectors expressed in the shape of a circle (○)) can be determined to include utterances of a third speaker related to the third cluster (913).

[0116] In state (901) of Fig. 9a, in a state where multiple speakers for multiple frames corresponding to vectors are determined, the electronic device may determine whether to store each of the vectors by using similarities between vectors determined using groups (e.g., clusters (911, 912, 913)) within the vector space corresponding to each of the multiple speakers. The electronic device may determine whether to store each of the vectors by using a distribution of vectors within the vector space based on a number of vectors exceeding a threshold number.

[0117] Referring to Fig. 9a, an electronic device that detects vectors exceeding a threshold number can selectively store vectors in memory. Based on the operation described with reference to Fig. 8, the electronic device can selectively store vectors having relatively large sums of similarities in each of the clusters (911, 912, 913). For example, if it is set to selectively store 6 vectors in each of the clusters (911, 912, 913), the electronic device can extract 18 vectors. Referring to Fig. 9b, a state (902) in which vectors having relatively large sums of similarities are selected is illustrated. Among the vectors in Fig. 9a, vectors having relatively small sums of similarities can be deleted from the memory of the electronic device. Referring to Fig. 9b, center vectors (c1, c2, c3) of each of the clusters (911, 912, 913) are illustrated. Vectors with relatively large sums of similarities can be positioned relatively close to the center vectors (c1, c2, c3) in each of the clusters (911, 912, 913).

[0118] In one embodiment, the electronic device may store the center vectors (c1, c2, c3) of the clusters (911, 912, 913) in memory as representative vectors to be used for clustering. In one embodiment, the electronic device may optionally store vectors adjacent to the center vector (e.g., center vectors (c1, c2, c3)) in each of the clusters (911, 912, 913). To identify the vectors adjacent to the center vector, the electronic device may calculate or determine the distance between the center vector and the vector.

[0119] In the state (902) of FIG. 9b corresponding to the second point in time after the first point in time of FIG. 9a, vectors stored by the electronic device can be used to separate speakers after the second point in time. For example, in the state (902) of FIG. 9b, vectors included in each of the clusters (911, 912, 913) can be stored in memory to identify at least one speaker in the portion of audio data corresponding to the second point in time after the second point in time.

[0120] While acquiring audio data using a microphone (e.g., microphone (225) of FIG. 2), the number of vectors acquired from the audio data may increase because the length of the audio data increases. Referring to FIG. 9C, a state (903) of the vector space at a third point in time after the second point in time corresponding to the state (902) of FIG. 9B is illustrated. As the electronic device generates additional frames from the audio data and vectors corresponding to the frames, vectors (e.g., Vectors expressed in the form of (a) may be added. Within the state (903) of FIG. 9c, the electronic device may perform clustering on vectors within the vector space to generate or determine clusters (911, 912, 913) corresponding to each of the plurality of speakers.

[0121] For example, included in the first cluster (911), Vectors expressed in the form can be determined to correspond to frames containing the utterance of the first speaker of the first cluster (911). For example, included in the second cluster (912), Vectors expressed in the form can be determined to correspond to frames containing the utterance of the second speaker of the second cluster (912). For example, included in the third cluster (913), The vectors expressed in the form can be determined to correspond to frames containing the utterances of the third speaker of the third cluster (913). In state (903) of FIG. 9c, the electronic device can further store in memory one or more time intervals from the second time point to the third time point, and information indicating relationships between multiple speakers.

[0122] In one embodiment, the electronic device may designate or assign identifiers of the first speaker to the third speaker assigned to the first cluster (911) to the third cluster (913) as vectors included in the first cluster (911) to the third cluster (913). Within the state (903) of FIG. 9c, vectors added to the vector space (e.g., Each of the vectors expressed in the form of (a) can be assigned one of the identifiers of the first speaker to the third speaker.

[0123] Referring to FIG. 9c, when detecting vectors exceeding a threshold number, the electronic device may perform an operation of filtering vectors in each of the clusters (911, 912, 913). The electronic device may maintain a specified number of vectors in each of the clusters (911, 912, 913) and discard the remaining vectors. Referring to FIG. 9d, a state (904) is illustrated after the state (903) of FIG. 9c, in which a specified number of vectors are selected in each of the clusters (911, 912, 913).

[0124] In state (904) of Fig. 9d, since the electronic device maintains vectors having relatively large sums of similarities in each of the clusters (911, 912, 913), only vectors adjacent to each of the center vectors (c1', c2', c3') of the clusters (911, 912, 913) can be stored, and the remaining vectors can be removed from the memory of the electronic device (e.g., the memory (215) of Fig. 2, and / or the vector storage area (233)). In state (904) of Fig. 9d, the vectors stored in the memory can be used for speaker separation of audio data received after state (904). Vectors stored in memory within state (904) can be used to identify one or more speakers (e.g., speakers corresponding to each of clusters (911, 912, 913)) that are commonly associated with a portion of audio data corresponding before state (904) and another portion of audio data corresponding after state (904).

[0125] For example, when 3,500 vectors are stored, the electronic device can perform filtering of the vectors as described with reference to FIG. 8 and / or FIGS. 9A-9D. For example, the electronic device can retain 150 vectors within a cluster of vectors corresponding to a particular speaker and discard or remove the remaining vectors. In the above example, if the electronic device has identified 10 clusters corresponding to each of 10 speakers, the electronic device can reduce the number of vectors stored in the memory to 1,500 at the point where 3,500 vectors are stored. At this point, the number of vectors stored in the memory can be reduced by approximately 57%.

[0126] Before reducing the number of vectors, in order to calculate the similarities between 3500 vectors, the electronic device calculates the similarities between two vectors among 3500 vectors 6,123,250 times ( = 3500 C2) It can be performed repeatedly. After reducing the number of vectors, in order to calculate the similarities between 1500 vectors (for convenience of explanation, the number of vectors will increase from 1500 as audio data is acquired), the electronic device performs the operation of calculating the similarities between two vectors among the 1500 vectors 1,124,250 times ( = 1500 C2) It can be performed repeatedly. For example, the number of times the operation of calculating similarities is performed can be reduced by approximately 82%. In other words, the speed of calculating similarities can be increased.

[0127] Hereinafter, the operation of the electronic device described with reference to FIGS. 1 to 8 and FIGS. 9a to 9d is described with reference to FIG. 10.

[0128] FIG. 10 illustrates a flowchart of an electronic device according to an embodiment of the disclosure. The electronic device (101) of FIGS. 1 and 2, and / or the processor (210) of FIG. 2, may perform operations of the electronic device described with reference to FIG. 10. The operations of FIG. 10 may be related to at least one of the operations of FIG. 3. The order in which the operations of FIG. 10 are performed is not limited to the order illustrated in FIG. 10, and may be performed substantially simultaneously, or may be performed in an order different from the order illustrated in FIG. 10.

[0129] Referring to FIG. 10 , in operation (1010), a processor of an electronic device according to an embodiment may obtain first vectors corresponding to a plurality of frames of audio data (e.g., audio data (110) of FIG. 1 , audio data (410) of FIGS. 4 and 5 , and / or audio data (710) of FIG. 7 ). For example, while obtaining audio data using a microphone (225) of FIG. 2 , the processor may segment the audio data to obtain a plurality of frames. The processor may obtain first vectors of operation (1010) corresponding to the plurality of frames, respectively. Operation (1010) of FIG. 10 may be performed similarly to operations (310, 320) of FIG. 3 .

[0130] Referring to FIG. 10 , in operation (1020), according to one embodiment, a processor of an electronic device may perform clustering of first vectors and / or second vectors stored in a memory to determine a speaker of each of a plurality of frames. The second vectors may have been obtained by performing speaker separation on another portion of audio data that was previously obtained prior to the portion of audio data that includes the plurality of frames of operation (1010). If operation (1010) is performed for the first time after acquisition of audio data using a microphone is initiated, no second vectors may be stored in the memory. In this case, the processor may perform operation (1030) following operation (1020) using only the first vectors of operation (1010). Operation (1020) of FIG. 10 may be performed similarly to operations (330, 340) of FIG. 3 .

[0131] Operation (1020) of FIG. 10 may include operations of the electronic device described with reference to FIGS. 7 and 8, and / or FIGS. 9A to 9D. For example, the processor may group the first vectors of operation (1010) and the second vectors of operation (1020) to obtain or generate one or more groups (e.g., clusters (911, 912, 913) of FIGS. 9A to 9D) that include at least one of the first vectors and the second vectors. Using the groups each of the first vectors includes, the processor may determine a speaker corresponding to each of the first vectors.

[0132] Referring to FIG. 10, in operation (1030), according to one embodiment, a processor of an electronic device may store information indicating a speaker of at least one time segment of audio data. The information of operation (1030) may include information (120) of FIG. 1. For example, the processor may store, in a memory, information indicating a speaker of at least one time segment of audio data, determined using a speaker corresponding to each of the first vectors of operation (1010). The information stored in operation (1030) may include a combination of a numerical value (e.g., an identifier) ​​indicating a specific speaker, a start time of a specific time segment of audio data in which speech of the specific speaker is recorded, and an end time of the specific time segment.

[0133] Referring to FIG. 10, in operation (1040), according to one embodiment, a processor of an electronic device may determine whether the number of first vectors and second vectors exceeds a specified number. For example, the processor may determine whether the total number of first vectors and second vectors, and / or the sum of the number of first vectors and the number of second vectors exceeds a specified number in operation (1040). The specified number in operation (1040) may correspond to the threshold number described with reference to FIG. 8 and / or FIGS. 9A to 9D. If the number of first vectors and second vectors exceeds the specified number (1040—Yes), the processor may perform operation (1050). If the number of first vectors and second vectors is less than or equal to the specified number (1040—No), the processor may perform operation (1060).

[0134] Referring to FIG. 10 , in operation (1050), according to an embodiment, a processor of an electronic device may store vectors less than or equal to a specified number of operation (1040) among first vectors and / or second vectors. Based on a total number of first vectors and second vectors exceeding the specified number of operation (1040), the processor may remove at least one of the first vectors of operation (1010) and the second vectors of operation (1020) from the memory to adjust the total number to less than or equal to the specified number. Operation (1050) of FIG. 10 may be performed similarly to the operation of the electronic device described with reference to FIG. 8 and / or FIGS. 9A to 9D . The processor may filter vectors in each of the groups formed in the vector space by the clustering of operation (1020) and selectively store vectors less than or equal to a specified number. The number of vectors stored in memory may be reduced by the processor performing operation (1050).

[0135] Referring to FIG. 10 , in operation (1060), a processor of an electronic device according to an embodiment may store first vectors and second vectors. Based on a total number of first vectors and second vectors less than or equal to a specified number of operation (1040), the processor may perform operation (1060). For example, the processor may store the first vectors of operation (1010) in a memory (e.g., memory (215) and / or vector storage area (233) of FIG. 2 ). In operation (1060), the electronic device may store the first vector of operation (1010) in the memory together with the second vectors stored in the memory. After performing any one of operations (1050, 1060) of FIG. 10 , the processor may repeatedly perform the operation of FIG. 10 for the remaining portion of continuously received audio data. For example, the processor may repeatedly perform the operation of FIG. 10 until acquisition of audio data is completed.

[0136] Hereinafter, with reference to FIG. 11a and / or FIG. 11b, the operation of an electronic device for displaying information obtained by performing speaker separation according to one embodiment is described.

[0137] FIGS. 11A and 11B illustrate a user interface (UI) displayed by an electronic device (101) according to various embodiments of the disclosure. The electronic device (101) of FIGS. 1 and 2 and / or the processor (210) of FIG. 2 may perform the operations of the electronic device described with reference to FIGS. 11A and 11B. The operations of FIGS. 11A and 11B may be related to at least one of the operations of FIGS. 3 and / or 10.

[0138] Referring to FIGS. 11A and 11B , states (1101, 1102, 1103, 1104) of an electronic device (101) executing a recording application (e.g., recording application (230) of FIG. 2) that supports functions related to audio data are illustrated. In response to an input indicating execution of the recording application, the electronic device (101) may display a screen provided from the recording application on the display (220). Referring to state (1101) of FIG. 11A , the electronic device (101) may display a screen including a visual object (1116) for initiating acquisition of audio data through a microphone (e.g., microphone (225) of FIG. 2) of the electronic device (101) on the display (220). The visual object (1116) may have a shape of a red circle. The visual object (1116) may be referred to as a recording button.

[0139] In state (1101) of FIG. 11A, the electronic device (101) may display a plurality of visual objects (1117, 1118, 1119) for switching screens displayed on the display (220). An area adjacent to the bottom of the display (220) including the plurality of visual objects (1117, 1118, 1119) may be referred to as a navigation bar. A visual object (1117) corresponding to a function for displaying a list of one or more screens displayed on the display (220) may be referred to as a multitasking button. A visual object (1118) corresponding to a function for displaying a designated screen (e.g., a screen referred to as a home screen and / or a launcher screen) may be referred to as a home button. A visual object (1119) corresponding to a function for sequentially displaying at least one screen that was displayed prior to the screen displayed within the current state (e.g., state (1101)) may be referred to as a back button.

[0140] Within the state (1101) of FIG. 11a, the electronic device (101) may display a screen including visual objects (1111, 1112, 1113, 1114, 1115, 1116, 1121) corresponding to different functions related to recording. The visual object (1111) may correspond to a function for displaying a list of one or more audio files including audio data recorded by the electronic device (101) in a different screen from the screen displayed within the state (1101) of FIG. 11a. The visual object (1112) may correspond to a function for changing a setting value of a software application (e.g., a recording application (230) of FIG. 2) executed by the electronic device (101).

[0141] The visual object (1113) may correspond to a first mode for performing recording. The visual object (1114) may correspond to a second mode for recording speech generated from a specific direction using multiple microphones. The visual object (1115) may correspond to a third mode for performing speaker separation to obtain a transcript (or a recording) corresponding to audio data recorded by the electronic device (101). The information obtained in the third mode may include the information (120) of FIG. 1 . Referring to state (1101) of FIG. 11A , the electronic device (101) may display a screen related to the third mode. Within state (1101), the electronic device (101) may emphasize the visual object (1115) corresponding to the third mode in comparison with the visual objects (1113, 1114).

[0142] Within the state (1101) of FIG. 11A, the electronic device (101) may display a screen including an area (1120). Within the area (1120), the electronic device (101) may display a visual object (1121) for selecting a language pack (e.g., a language package and / or a language file) to be used for analysis of audio data (e.g., STT). In response to an input indicating selection of the visual object (1121), the electronic device (101) may display a pop-up window (or menu) on the display (220) for adjusting the language pack. The language pack may include information to be used for processing or analyzing audio data, such as STT. Referring to the visual object (1121), the electronic device (101) may use any one of a plurality of language packs classified according to a category (e.g., language) and / or type of natural language to analyze the audio data. For example, after a language pack corresponding to Korean is selected (1101), the electronic device (101) can analyze audio data using information included in the language pack to generate or display text expressed in Korean.

[0143] Within the state (1101) of FIG. 11a, in response to an input indicating selection of a visual object (1116), the electronic device (101) may initiate acquisition of audio data. For example, the electronic device (101) may control a microphone to initiate acquisition of audio data. In response to the input, the electronic device (101) may transition from the state (1101) of FIG. 11a to the state (1102) of FIG. 11a.

[0144] Referring to FIG. 11A, a state (1102) of a screen displayed by an electronic device (101) that acquires audio data using a microphone is illustrated. While acquiring audio data, the electronic device (101) may display text (1124) indicating the length of the audio data (e.g., 15 seconds 9 in the state (1102) of FIG. 11A). While acquiring audio data, a numerical value included in the text (1124) may gradually increase. The electronic device (101) may display a visual object (1123) mapped to an attribute of the audio data being acquired (e.g., an attribute indicating whether it has been added to a favorites list) in the state (1102).

[0145] Referring to FIG. 11A, in a state (1102) of acquiring audio data, the electronic device (101) may display visual objects (1127, 1128, 1129) for controlling acquisition of audio data. The electronic device (101) may display a visual object (1128) corresponding to a function for temporarily stopping acquisition of audio data, and / or a visual object (1129) mapped to a function for terminating acquisition of audio data and saving a file corresponding to the audio data. The visual object (1128) may be referred to as a pause button. The visual object (1129) may be referred to as a stop button. The electronic device (101) may display a visual object (1127) mapped to a function for playing audio data acquired in the state (1102). The visual object (1127) may be referred to as a play button.

[0146] Within state (1102) of FIG. 11A, while acquiring audio data, the electronic device (101) may deactivate the visual object (1127). For example, despite a touch input on the visual object (1127), the electronic device (101) may not perform any function associated with the touch input. In response to an input indicating selection of the visual object (1128), the electronic device (101) may at least temporarily cease acquiring audio data, deactivate the visual object (1128), and / or activate the visual object (1127). In response to an input indicating selection of an activated visual object (1127), the electronic device (101) may resume acquiring audio data, deactivate the visual object (1127), and / or activate the visual object (1128).

[0147] Within the state (1102) of FIG. 11a, the electronic device (101) may display, within the region (1120), the result of performing speaker separation of audio data being acquired within the state (1102). For example, the electronic device (101) may display, from the top to the bottom of the region (1120), speakers sequentially separated in the time domain (e.g., speakers of each of the visual objects (e.g., icons) including A, B, and C of FIG. 11a) and text representing the utterance of the speakers. While acquiring audio data, the electronic device (101) may display, within the region (1120), one or more texts and a visual object representing at least one speaker corresponding to the one or more texts, according to the result of performing speaker separation (e.g., speaker separation described with reference to FIGS. 1 to 8, FIGS. 9a to 9d, and FIG. 10). Since the area (1120) has a limited size, the electronic device (101) can perform scrolling on the area (1120) to perform speaker separation so that the last acquired text is displayed in the area (1120).

[0148] Referring to FIG. 11A, within an area (1120), the electronic device (101) may display a visual object (1122) mapped to a function for adding a bookmark to a current point in time within the entire time interval of audio data. Within a state (1102) of FIG. 11A, the electronic device (101) that has received an input indicating selection of the visual object (1122) may store information indicating that the point in time (e.g., 15 seconds 9) at which the input was received has been set as a bookmark by the user within the entire time interval of audio data. The electronic device (101) that has received the input may further display a visual object and / or an icon indicating the setting of the bookmark on a timeline displayed through an area (1125).

[0149] In one embodiment, speaker separation may be repeatedly performed by the electronic device (101) while acquiring audio data. According to one embodiment, the electronic device (101) may display the result of performing speaker separation (e.g., information (120) of FIG. 1) on the display (220) before acquiring audio data. For example, the electronic device (101) may display the result of performing speaker separation within an area (1120) of the display (220).

[0150] While acquiring audio data, the electronic device (101) can visualize the audio data acquired by the electronic device (101) (e.g., a sound waveform related to the audio data) within an area (1125). While acquiring audio data, the electronic device (101) can display a sound waveform related to the audio data on the left side of the area (1125) based on an indicator (1126) indicating the current time.

[0151] Within the state (1102) of FIG. 11A, the electronic device (101) continuously acquiring audio data may determine whether to stop acquiring the audio data based on the available capacity of the memory (e.g., the memory (215) of FIG. 2). For example, if the size of the audio data acquired within the state (1102) exceeds the available capacity of the memory or exceeds a specified percentage of the available capacity, the electronic device (101) may stop acquiring the audio data. For example, in response to an input indicating selection of a visual object (1129), the electronic device (101) may stop acquiring the audio data. The electronic device (101) that has stopped acquiring the audio data may display the result of performing speaker separation within the area (1120) or store it in the memory.

[0152] Within the state (1101) of FIG. 11a, the electronic device (101) that receives an input indicating selection of a visual object (1111) may switch to the state (1103) of FIG. 11b. Within the state (1103) of FIG. 11b, the electronic device (101) may display a list of a plurality of audio files (e.g., files including audio data) stored in the electronic device (101). The list may be displayed within an area (1133) of the screen. Within the scrollable area (1133), the electronic device (101) may display items corresponding to each of the plurality of files. The electronic device (101) may display visual objects (1131, 1132) for controlling an option for displaying audio files in the area (1133) together with the area (1133). The visual object (1131) may be matched to a function of displaying a list of all audio files stored in the electronic device (101) in an area (1133). The visual object (1132) may be matched to a function of displaying a list of at least one audio file corresponding to a specific category among the audio files stored in the electronic device (101) in an area (1133).

[0153] Referring to state (1103) of FIG. 11b, within a region (1133) of the screen, the electronic device (101) may display an item object corresponding to a specific audio file. The item object may include an icon (e.g., a visual object referred to as a play button) that causes playback of the audio file. Referring to FIG. 11b, within an item object (1134), the electronic device (101) may display, together with the icon, the name of the audio file (e.g., "Recording_20231111"), the length of the audio file (e.g., 10 hours 3 minutes 12 seconds 24), and the date the audio file was saved (e.g., "11 NOV 2023" indicating November 11, 2023). Within the item object (1134), the electronic device (101) may display an indicator (1135) for indicating dialogue information for the audio file corresponding to the item object (1134). An indicator (1135) may be referred to as an icon, an image, and / or a visual object.

[0154] Within the state (1103) of FIG. 11b, in response to an input indicating selection of an item object (1134), the electronic device (101) may switch to a state (1104) for playing an audio file corresponding to the item object (1134). Within the state (1104), the electronic device (101) may display a screen for controlling playing of the audio file. Referring to FIG. 11b, within the state (1104), the electronic device (101) may display a visual object (1141) for adjusting a property of the audio file (e.g., a property indicating whether the audio file has been added to a favorites list), a visual object (1142) for changing a name of the audio file (e.g., a file name), and a visual object (1143) for displaying a menu including options related to playing of the audio file.

[0155] Within the state (1104) of FIG. 11b, the electronic device (101) may display an area (1146) for visualizing a waveform of an audio file. The area (1146) may be referred to as a timeline. Within the state (1104), the electronic device (101) may display text (1145) indicating the length of the audio file (e.g., 10 hours 3 minutes 12 seconds 24). Within the state (1104), the electronic device (101) may display text (1144) indicating the time point of the audio file currently being played. Within the state (1104) of FIG. 11b, the electronic device (101) may play a portion corresponding to 1 second 21 of the audio file.

[0156] Within the state (1104), the electronic device (101) may display visual objects (1151, 1152, 1153, 1154, 1155, 1156) for controlling playback of an audio file. The visual object (1151) may be mapped to a function for muting. The visual object (1152) may be mapped to a function for repeat playback of a section of the audio file. The visual object (1153) may be mapped to a function for adjusting the playback speed of the audio file. The visual object (1154) may be mapped to a function for playing a portion of the audio file corresponding to a time point in the past (e.g., 3 seconds before) the currently playing time point. The visual object (1156) may be mapped to a function for playing a portion of the audio file corresponding to a time point after the currently playing time point (e.g., 3 seconds after). A visual object (1155) may be mapped to a function for at least temporarily pausing playback of an audio file.

[0157] Within state (1104), the electronic device (101) may display the results of detecting speech segments of one or more speakers from an audio file obtained by performing speaker separation through an area (1147). Referring to FIG. 11B , an area (1147) containing texts corresponding to utterances of three speakers (e.g., A, B, C) is exemplarily illustrated. As the audio file is played, the electronic device (101) may change the text displayed through the area (1147) according to the utterances of the speakers at the time the audio file is played. The text may be generated using information obtained by performing speaker separation, for example, the operations described with reference to FIGS. 1 to 8 , FIGS. 9A to 9D , and FIG. 10 (e.g., information (120) of FIG. 1 ). The electronic device (101) can highlight text displayed through the area (1147) at least partially to indicate a location within the text corresponding to a portion of an audio file currently being played through the electronic device (101).

[0158] Within state (1104), the electronic device (101) may use region (1146) to visualize speech segments corresponding to each speaker associated with the audio file. For example, the electronic device (101) may change the color of the waveform displayed in region (1146) at least partially according to the speech segment. For example, the color of the waveform at a specific point in time may have a color corresponding to one of the speakers associated with the audio file. Referring to FIG. 11B , within region (1146), the electronic device (101) may include visual objects (1161, 1162, 1163) in the form of bars representing speech segments of each speaker associated with the audio file. The visual objects (1161, 1162, 1163) are illustrated as horizontal arrows extending from the names of the speakers (e.g., A, B, C), but the embodiment is not limited thereto.

[0159] Each of the bar-shaped visual objects (1161, 1162, 1163) may correspond to each of the speech segments detected from the audio file. For example, the visual object (1161) may represent a time segment in which speaker A's speech (e.g., "Hello") was recorded. The positions of the start and end points of the visual object (1161) within the region (1146) may correspond to the start and end points of the time segment, respectively. Similarly, the visual object (1162) may represent a time segment in which speaker B's speech (e.g., "Is the time okay?") was recorded. The start, end, and length of the visual object (1162) within the region (1146) may represent the start, end, and length of the time segment in which speaker B's speech was recorded. As the audio file plays, the waveform and / or visual objects (1161, 1162, 1163) may scroll within the region (1146). For example, the waveform and / or visual objects (1161, 1162, 1163) may scroll in a direction from the right edge to the left edge of the region (1146).

[0160] As described above, according to one embodiment, the electronic device (101) can perform speaker separation on an audio file (or audio data). To reduce the amount of computation required to perform speaker separation or prevent an exponential increase in the amount of computation, the electronic device (101) can iteratively reduce information used for speaker separation (e.g., vectors corresponding to each frame of audio data). The reduction of information can be performed to ensure speaker separation performance.

[0161] FIG. 12 is a block diagram of an electronic device (1201) within a network environment (1200) according to one embodiment of the disclosure. Referring to FIG. 12 , in the network environment (1200), the electronic device (1201) may communicate with the electronic device (1202) via a first network (1298) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (1204) or the server (1208) via a second network (1299) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (1201) may communicate with the electronic device (1204) via the server (1208). According to one embodiment, the electronic device (1201) may include a processor (1220), a memory (1230), an input module (1250), an audio output module (1255), a display module (1260), an audio module (1270), a sensor module (1276), an interface (1277), a connection terminal (1278), a haptic module (1279), a camera module (1280), a power management module (1288), a battery (1289), a communication module (1290), a subscriber identification module (1296), or an antenna module (1297). In some embodiments, the electronic device (1201) may omit at least one of these components (e.g., the connection terminal (1278)), or may have one or more other components added. In some embodiments, some of these components (e.g., sensor module (1276), camera module (1280), or antenna module (1297)) may be integrated into a single component (e.g., display module (1260)).

[0162] The processor (1220) may control at least one other component (e.g., hardware or software component) of the electronic device (1201) connected to the processor (1220) by executing, for example, software (e.g., program (1240)), and may perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (1220) may store commands or data received from other components (e.g., sensor module (1276) or communication module (1290)) in volatile memory (1232), process the commands or data stored in volatile memory (1232), and store result data in non-volatile memory (1234). According to one embodiment, the processor (1220) may include a main processor (1221) (e.g., a central processing unit or an application processor) or an auxiliary processor (1223) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (1221). For example, when the electronic device (1201) includes the main processor (1221) and the auxiliary processor (1223), the auxiliary processor (1223) may be configured to use less power than the main processor (1221) or to be specialized for a given function. The auxiliary processor (1223) may be implemented separately from the main processor (1221) or as a part thereof.

[0163] The auxiliary processor (1223) may control at least a portion of functions or states associated with at least one component (e.g., a display module (1260), a sensor module (1276), or a communication module (1290)) of the electronic device (1201), for example, on behalf of the main processor (1221) while the main processor (1221) is in an inactive (e.g., sleep) state, or together with the main processor (1221) while the main processor (1221) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (1223) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (1280) or a communication module (1290)). In one embodiment, the auxiliary processor (1223) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (1201) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (1208)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0164] The memory (1230) can store various data used by at least one component (e.g., the processor (1220) or the sensor module (1276)) of the electronic device (1201). The data can include, for example, software (e.g., the program (1240)) and input data or output data for commands related thereto. The memory (1230) can include a volatile memory (1232) or a non-volatile memory (1234).

[0165] The program (1240) may be stored as software in memory (1230) and may include, for example, an operating system (1242), middleware (1244), or an application (1246).

[0166] The input module (1250) can receive commands or data to be used in a component of the electronic device (1201) (e.g., a processor (1220)) from an external source (e.g., a user) of the electronic device (1201). The input module (1250) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0167] The audio output module (1255) can output audio signals to the outside of the electronic device (1201). The audio output module (1255) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0168] The display module (1260) can visually provide information to an external party (e.g., a user) of the electronic device (1201). The display module (1260) may include, for example, a display, a holographic device, or a projector, and a control circuit for controlling the device. In one embodiment, the display module (1260) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0169] The audio module (1270) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (1270) can acquire sound through the input module (1250), output sound through the sound output module (1255), or an external electronic device (e.g., electronic device (1202)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (1201).

[0170] The sensor module (1276) can detect the operating status (e.g., power or temperature) of the electronic device (1201) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (1276) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0171] The interface (1277) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (1201) with an external electronic device (e.g., the electronic device (1202)). In one embodiment, the interface (1277) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0172] The connection terminal (1278) may include a connector through which the electronic device (1201) may be physically connected to an external electronic device (e.g., the electronic device (1202)). In one embodiment, the connection terminal (1278) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0173] The haptic module (1279) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (1279) may include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0174] The camera module (1280) can capture still images and videos. According to one embodiment, the camera module (1280) may include one or more lenses, image sensors, image signal processors, or flashes.

[0175] The power management module (1288) can manage the power supplied to the electronic device (1201). According to one embodiment, the power management module (1288) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).

[0176] A battery (1289) may power at least one component of the electronic device (1201). In one embodiment, the battery (1289) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0177] The communication module (1290) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (1201) and an external electronic device (e.g., electronic device (1202), electronic device (1204), or server (1208)), and the performance of communication through the established communication channel. The communication module (1290) may operate independently from the processor (1220) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1290) may include a wireless communication module (1292) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (1294) (e.g., a local area network (LAN) communication module, or a power line communication module). Any of these communication modules may communicate with an external electronic device (1204) via a first network (1298) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (1299) (e.g., a long-range communication network such as a legacy cellular network, a fifth generation (5G) network, a next-generation communication network, the Internet, or a computer network (e.g., a local area network or a wide area network)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1292) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (1296) to verify or authenticate the electronic device (1201) within a communication network such as the first network (1298) or the second network (1299).

[0178] The wireless communication module (1292) can support a 5G network and next-generation communication technologies following the 4th generation (4G) network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency communications (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (1292) can support, for example, a high-frequency band (e.g., millimeter wave (mmWave) band) to achieve a high data transmission rate. The wireless communication module (1292) may support various technologies for securing performance in high-frequency bands, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (1292) may support various requirements specified in the electronic device (1201), an external electronic device (e.g., the electronic device (1204)), or a network system (e.g., the second network (1299)). According to one embodiment, the wireless communication module (1292) may be configured to achieve a peak data rate (e.g., 20 Gbps or more) for eMBB implementation, a loss coverage (e.g., 164 dB or less) for mMTC implementation, or a U-plane latency (e.g., 0 for downlink (DL) and uplink (UL) respectively) for URLLC implementation.It can support latency of 5ms or less, or round trip of 1ms or less.

[0179] The antenna module (1297) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (1297) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (1297) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (1298) or the second network (1299), may be selected from the plurality of antennas, for example, by the communication module (1290). A signal or power may be transmitted or received between the communication module (1290) and an external electronic device via the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (1297).

[0180] According to various embodiments, the antenna module (1297) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.

[0181] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0182] According to one embodiment, commands or data may be transmitted or received between the electronic device (1201) and an external electronic device (1204) via a server (1208) connected to a second network (1299). Each of the external electronic devices (1202, or 704) may be the same or a different type of device as the electronic device (1201). According to one embodiment, all or part of the operations executed in the electronic device (1201) may be executed in one or more of the external electronic devices (1202, 704, or 708). For example, when the electronic device (1201) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (1201) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (1201). The electronic device (1201) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (1201) may provide an ultra-low latency service using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (1204) may include an Internet of Things (IoT) device. The server (1208) may be an intelligent server utilizing machine learning and / or a neural network.According to one embodiment, an external electronic device (1204) or server (1208) may be included within the second network (1299). The electronic device (1201) may be applied to intelligent services (e.g., smart homes, smart cities, smart cars, or healthcare) based on 5G communication technology and IoT-related technology.

[0183] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.

[0184] The various embodiments of the disclosure and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. In this document, each of the phrases "A or B," "at least one of A and B," "at least one of A or B," "A, B, or C," "at least one of A, B, and C," and "at least one of A, B, or C" may include any one of the items listed together with the corresponding phrase among the phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish the corresponding component from other corresponding components and do not limit the corresponding components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as being “coupled” or “connected” to another component (e.g., a second component), with or without the terms “functionally” or “communicatively,” it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0185] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0186] Various embodiments of the present document may be implemented as software (e.g., a program (1240)) including one or more instructions stored in a storage medium (e.g., an internal memory (1236) or an external memory (1238)) readable by a machine (e.g., an electronic device (1201)). For example, a processor (e.g., a processor (1220)) of the machine (e.g., an electronic device (1201)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0187] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0188] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and arranged in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added. The electronic device (1201) of FIG. 12 may be an example of the electronic device (101) described with reference to FIGS. 1 to 8, FIGS. 9A to 9D, FIG. 10, and / or FIGS. 11A and 11B. For example, the memory (1230) of FIG. 12 may correspond to the memory (215) of FIG. 2. For example, the display module (1260) of FIG. 12 may correspond to the display (220) of FIG. 2. For example, the processor (1220) of FIG. 12 may correspond to the processor (210) of FIG. 2. For example, the audio module (1270) and / or the sensor module (1276) of FIG. 12 may at least partially include the microphone (225) of FIG. 2.

[0189] In one embodiment, a method may be required to efficiently manage or reduce the resources (e.g., memory occupancy, computational load, and / or computational time) of an electronic device occupied for speaker diarization. As described above, according to an embodiment, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (1201) of FIG. 12) may include a microphone (e.g., microphone (225) of FIG. 2), at least one processor including a processing circuit (e.g., processor (210) of FIG. 2), and a memory including one or more storage media for storing instructions (e.g., memory (215) of FIG. 2). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain audio data (e.g., audio data (110) of FIG. 1 , audio data (410) of FIGS. 4 and 5 , and / or audio data (710) of FIG. 7 ) using the microphone, while segmenting the audio data to obtain a plurality of frames (e.g., frames (510, 520, 530) of FIGS. 5 and 6 ). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain first vectors (e.g., vectors (610, 620, 630) of FIG. 6 ) respectively corresponding to the plurality of frames. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine a speaker corresponding to each of the first vectors using groups each of the first vectors containing the first vectors, obtained by grouping the first vectors and the second vectors stored in the memory. The grouping of the first vectors and the second vectors may be performed by the at least one processor.The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to store, in the memory, information (e.g., information (120) of FIG. 1) determined using a speaker corresponding to each of the first vectors, the information indicating a speaker of at least one time segment of the audio data. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to store the first vectors in the memory based on a total number of the first vectors and the second vectors being less than or equal to a specified number. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to remove at least one of the first vectors and the second vectors from the memory to adjust the total number to less than or equal to the specified number, based on a total number of the first vectors and the second vectors exceeding the specified number. According to one embodiment, an electronic device can efficiently manage the number of vectors used for speaker separation, and / or the size of the entire vectors. According to one embodiment, the electronic device can manage or reduce the computational load for clustering vectors by keeping the number of vectors used for speaker separation below a specified number.

[0190] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine whether to adjust the specified number of vectors using a period of time required to group the specified number of vectors.

[0191] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to decrease the specified number in response to the period exceeding the specified period.

[0192] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine whether to store each of the first vectors and the second vectors, based on a distribution of the first vectors and the second vectors within a vector space, based on the total number of the first vectors and the second vectors exceeding the specified number.

[0193] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine whether to store each of the first vectors and the second vectors using similarities between the first vectors and the second vectors determined using the groups in the vector space corresponding to each of the plurality of speakers, in a state where the plurality of speakers for the plurality of frames have been determined.

[0194] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine whether to store each of the first vectors and the second vectors using distances between the center vectors of the groups in the vector space corresponding to each of the plurality of utterances, with the plurality of utterances for the plurality of frames determined.

[0195] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device, in response to detecting from the audio data a voice segment in which a voice appears to have been recorded, to segment the voice segment to obtain the plurality of frames.

[0196] For example, the lengths of the plurality of frames may be the same. The plurality of frames may at least partially overlap each other in the time domain.

[0197] For example, the electronic device may include a display (e.g., display (220) of FIG. 2). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display, on the display, a screen related to the information in response to an input indicating a cessation of acquisition of the audio data.

[0198] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to compare a specified number of vectors associated with the audio data and stored in the memory to the specified number of vectors, while acquiring the audio data, to maintain the specified number.

[0199] As described above, in one embodiment, a method of an electronic device including a microphone may be provided. The method may include an operation of acquiring audio data using the microphone, while segmenting the audio data to acquire a plurality of frames. The method may include an operation of acquiring first vectors corresponding to each of the plurality of frames (e.g., operation 1010 of FIG. 10 ). The method may include an operation of determining a speaker corresponding to each of the first vectors using groups each of the first vectors obtained by grouping the first vectors and second vectors stored in a memory of the electronic device (e.g., operation 1020 of FIG. 10 ). The grouping of the first vectors and the second vectors may be performed by at least one processor of the electronic device. The method may include an operation of storing, in the memory, information indicating a speaker of at least one time segment of the audio data, the information determined using the speaker corresponding to each of the first vectors. The method may include an operation of storing the first vectors in the memory (e.g., operation (1060) of FIG. 10) based on a total number of the first vectors and the second vectors being less than or equal to a specified number. The method may include an operation of removing at least one of the first vectors and the second vectors from the memory (e.g., operation (1050) of FIG. 10) based on a total number of the first vectors and the second vectors exceeding the specified number to adjust the total number to less than or equal to the specified number.

[0200] For example, the method may include an operation of determining whether to adjust the specified number of vectors by using a period of time required to group the specified number of vectors.

[0201] For example, the action of determining whether to adjust the specified number may include an action of decreasing the specified number in response to the period exceeding the specified period.

[0202] For example, the operation of storing the specified number of vectors may include an operation of determining whether to store each of the first vectors and the second vectors, based on the total number of the first vectors and the second vectors exceeding the specified number, using a distribution of the first vectors and the second vectors within a vector space.

[0203] For example, the operation of storing the specified number of vectors may include an operation of determining whether to store each of the first vectors and the second vectors using similarities between the first vectors and the second vectors determined using the groups in the vector space corresponding to each of the plurality of speakers, in a state where a plurality of speakers for the plurality of frames are determined.

[0204] For example, the operation of storing the specified number of vectors may include an operation of determining whether to store each of the first vectors and the second vectors by using distances between the center vectors of the groups in the vector space corresponding to each of the first vectors and the second vectors and the plurality of speakers, in a state where a plurality of speakers for the plurality of frames are determined.

[0205] For example, the method may include determining a first set of first vectors of a first group from among the groups in the vector space. The method may include determining a second set of first vectors of the first group. The first set of first vectors may have higher similarities to the center vector of the first group than the second group of first vectors of the first group.

[0206] For example, in order to adjust the total number to be less than or equal to the specified number, the operation of removing at least one of the first vectors and the second vectors from the memory may include the operation of removing the second set of the first vectors of the first group from the memory.

[0207] For example, the method may include, in response to detecting from the audio data a voice segment in which a voice appears to have been recorded, an operation of segmenting the voice segment to obtain the plurality of frames.

[0208] For example, the lengths of the plurality of frames may be the same. The plurality of frames may at least partially overlap each other in the time domain.

[0209] For example, the method may include, in response to an input indicating a cessation of acquisition of the audio data, displaying a screen related to the information on a display of the electronic device.

[0210] For example, the method may include, while acquiring the audio data, comparing the number of vectors associated with the audio data and stored in the memory to the specified number with the total number of the first vectors and the second vectors to maintain the number of vectors to the specified number.

[0211] In one embodiment, as described above, a non-transitory computer-readable storage medium storing instructions may be provided. The instructions, when executed by an electronic device including a microphone, may cause the electronic device to segment audio data while acquiring audio data using the microphone to acquire a plurality of frames. The instructions, when executed by the electronic device, may cause the electronic device to acquire first vectors respectively corresponding to the plurality of frames. The instructions, when executed by the electronic device, may cause the electronic device to determine a speaker corresponding to each of the first vectors using groups each including the first vectors obtained by grouping the first vectors and second vectors stored in a memory. The grouping of the first vectors and the second vectors may be performed by the at least one processor of the electronic device. The instructions, when executed by the electronic device, may cause the electronic device to store, in the memory, information determined using a speaker corresponding to each of the first vectors, the information representing a speaker of at least one time segment of the audio data. The instructions, when executed by the electronic device, may cause the electronic device to remove, from the memory, at least one of the first vectors and the second vectors, based on a total number of the first vectors and the second vectors exceeding a specified number, to adjust the total number to be less than or equal to the specified number.

[0212] According to one embodiment, an electronic device as described above may include a microphone, at least one processor including a processing circuit, and a memory including one or more storage media storing instructions. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain first vectors each corresponding to a plurality of frames included in a first time interval of audio data obtained by controlling the microphone. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform clustering of the first vectors and second vectors obtained in a second time interval of the audio data prior to the first time interval, thereby determining a speaker of each of the plurality of frames. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine at least one vector, from among the first vectors and the second vectors, to be stored in memory for identifying at least one speaker associated with a third time segment of the audio data after the first time segment.

[0213] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to compare a specified number of vectors associated with the audio data and stored in the memory to a specified number of the first vectors and the second vectors while acquiring the audio data.

[0214] For example, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to, in response to detecting from the audio data the first time interval in which a voice appears to have been recorded, segment the first time interval to obtain the plurality of frames.

[0215] In one embodiment, as described above, a non-transitory computer-readable storage medium storing instructions may be provided. The instructions, when executed by an electronic device including a microphone, may cause the electronic device to obtain first vectors each corresponding to a plurality of frames included in a first time interval of audio data acquired by controlling the microphone. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform clustering of the first vectors and second vectors acquired in a second time interval of the audio data prior to the first time interval to determine a speaker of each of the plurality of frames. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine, from among the first vectors and the second vectors, at least one vector to be stored in a memory for identifying at least one speaker associated with a third time interval of the audio data subsequent to the first time interval.

[0216] For example, the instructions, when executed by the electronic device, may cause the electronic device to compare a specified number of vectors associated with the audio data and stored in the memory to a specified number of the first vectors and the second vectors while acquiring the audio data.

[0217] For example, the instructions, when executed by the electronic device, may cause the electronic device to, in response to detecting from the audio data the first time interval in which a voice appears to have been recorded, segment the first time interval to obtain the plurality of frames.

[0218] As described above, one or more non-transitory computer-readable storage media storing one or more computer programs may be provided. The one or more programs may include computer-readable instructions that, when individually or collectively executed by one or more processors of an electronic device, cause the electronic device to perform an operation of obtaining a plurality of frames by separating audio data while obtaining audio data using a microphone of the electronic device. The one or more programs may include computer-readable instructions that, when individually or collectively executed by one or more processors of the electronic device, cause the electronic device to perform an operation of grouping first vectors and second vectors stored in a memory of the electronic device, and determining a speaker corresponding to each of the first vectors using groups including the first vectors, wherein the operation of grouping the first vectors and the second vectors may be performed by at least one processor of the electronic device. The one or more programs may include computer-readable instructions that, when individually or collectively executed by one or more processors of the electronic device, cause the electronic device to perform an operation of storing, in the memory, information determined using a speaker corresponding to each of the first vectors, the information representing a speaker of at least one time segment of the audio data.The one or more programs may include computer-readable instructions that, when individually or collectively executed by one or more processors of the electronic device, cause the electronic device to perform an operation of removing at least one of the first vectors and the second vectors from the memory to adjust the total number of the first vectors and the second vectors to be equal to or less than the specified number, based on a total number of the first vectors and the second vectors being equal to or greater than the specified number.

[0219] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and one or more software applications running on the operating system. The processing device may also access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.

[0220] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.

[0221] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and digital versatile discs (DVDs), magneto-optical media such as floptical disks, and those configured to store program commands, including ROM, RAM, and flash memory. In addition, examples of other media may include recording or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.

[0222] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.

[0223] While the disclosure has been shown and described with reference to various embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the scope of the disclosure as defined by the appended claims and their equivalents.

Claims

1. In electronic devices, mike; At least one processor comprising a processing circuit; and A memory comprising one or more storage media storing instructions, wherein the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: While acquiring audio data using the above microphone, the audio data is divided to acquire multiple frames, Obtain first vectors corresponding to each of the above multiple frames, By using groups each of the first vectors included in the first vectors obtained by grouping the first vectors and the second vectors stored in the memory, a speaker corresponding to each of the first vectors is determined, and the grouping of the first vectors and the second vectors is performed by at least one processor of the electronic device. In the above memory, information determined using a speaker corresponding to each of the first vectors is stored, the information indicating a speaker of at least one time section of the audio data, and Based on the total number of the first vectors and the second vectors exceeding a specified number, causing at least one of the first vectors and the second vectors to be removed from the memory to adjust the total number to be less than or equal to the specified number. Electronic devices.

2. In claim 1, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Further causing a determination of whether to adjust the specified number of vectors by using the period required to group the specified number of vectors. Electronic devices.

3. In claim 2, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: In response to the above period exceeding the specified period, causing the above specified number to be reduced, Electronic devices.

4. In claim 1, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Further causing, based on the total number of the first vectors and the second vectors exceeding the specified number, to determine whether to store each of the first vectors and the second vectors by using a distribution of the first vectors and the second vectors in the vector space. Electronic devices.

5. In claim 4, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: In a state where multiple speakers for the multiple frames are determined, using the similarities between the first vectors and the second vectors determined using the groups in the vector space corresponding to each of the multiple speakers, it is further caused to determine whether to store each of the first vectors and the second vectors. Electronic devices.

6. In claim 4, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: In a state where multiple speakers for the multiple frames are determined, determining whether to store each of the first vectors and the second vectors by using distances between the center vectors of the groups in the vector space corresponding to each of the multiple speakers and the first vectors and the second vectors is further caused. Electronic devices.

7. In claim 1, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: In response to detecting a voice segment in which a voice appears to have been recorded from the above audio data, further causing the voice segment to be segmented to obtain the plurality of frames. Electronic devices.

8. In claim 7, The lengths of the above multiple frames are equal to each other, and The above multiple frames, at least partially overlapping each other in the time domain, Electronic devices.

9. In claim 1, Including more displays, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: In response to an input indicating a cessation of acquisition of said audio data, further causing said display to display a screen related to said information. Electronic devices.

10. In claim 1, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Further, while acquiring said audio data, to maintain the number of vectors associated with said audio data and stored in said memory to the specified number, comparing said specified number with the total number of said first vectors and said second vectors, Electronic devices.

11. A method of an electronic device including a microphone, An operation of acquiring audio data using the above microphone, dividing the audio data and acquiring a plurality of frames; An operation of obtaining first vectors each corresponding to each of the plurality of frames; An operation of determining a speaker corresponding to each of the first vectors by using groups each of the first vectors included in the groups obtained by grouping the first vectors and the second vectors stored in the memory of the electronic device, wherein the grouping of the first vectors and the second vectors is performed by at least one processor of the electronic device; An operation of storing, in the memory, information determined using a speaker corresponding to each of the first vectors, the information indicating a speaker of at least one time section of the audio data; An operation of removing at least one of the first vectors and the second vectors from the memory, based on a total number of the first vectors and the second vectors exceeding a specified number, to adjust the total number to be less than or equal to the specified number. method.

12. In claim 11, Further comprising an operation of determining whether to adjust the specified number of vectors by using a period of time required to group the specified number of vectors. method.

13. In claim 12, the operation of determining whether to adjust the specified number is: In response to said period exceeding a specified period of time, comprising an action of decreasing said specified number; method.

14. In claim 11, the operation of storing the specified number of vectors comprises: An operation including determining whether to store each of the first vectors and the second vectors, based on the total number of the first vectors and the second vectors exceeding the specified number, by using a distribution of the first vectors and the second vectors within the vector space. method.

15. In claim 14, the operation of storing the specified number of vectors comprises: In a state where multiple speakers for the multiple frames are determined, an operation of determining whether to store each of the first vectors and the second vectors by using the similarities between the first vectors and the second vectors determined by using the groups in the vector space corresponding to each of the multiple speakers is included. method.

Citation Information

Patent Citations

  • Vacuum Griper Device for Separating Ultra-Wide Insert Products from CVD Coating Tray

    KR102416681B1

  • Water pump

    KR102442322B1

  • Connected accessory for a voice-controlled device

    US11443739B1

  • Audio output control

    US20230367546A1

  • A system and a method for low latency speaker detection and recognition

    US20230386476A1