Apparatus and method for separating audio object

By introducing a frame delay and buffer management for mask data application in neural network-based audio object separation, the challenge of real-time accurate separation of audio objects in complex environments is addressed, enhancing performance metrics like SDR.

WO2026106034A1PCT designated stage Publication Date: 2026-05-21SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2025-08-07
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Existing audio object separation technologies struggle with accurate real-time separation of audio objects in complex environments without relying on cloud connections, particularly due to delays in applying mask data to frequency data, leading to degraded performance.

Method used

Implementing a frame delay in applying mask data to frequency data and using a buffer structure to ensure accurate alignment, combined with a neural network model for audio object separation, allowing for precise separation and overlap of audio objects on an electronic device.

Benefits of technology

Enhances audio object separation performance by improving signal-to-distortion ratio (SDR) through delayed mask application and buffer management, ensuring accurate and efficient separation of audio objects in real-time environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025011955_21052026_PF_FP_ABST
    Figure KR2025011955_21052026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are an electronic device and a method for separating an audio object. The electronic device receives first input data including audio data for a first frame, obtains first frequency data by converting the first input data into a frequency domain, applies an appropriate delay corresponding to a processing time related to mask data generation to the first frequency data, and obtains first frequency object data by applying first mask data generated for audio object separation. The electronic device obtains first object data by inverse-transforming the first frequency object data into a time domain.
Need to check novelty before this filing date? Find Prior Art

Description

Device and method for separating audio objects

[0001] The present disclosure relates to an electronic device and method for separating audio objects.

[0002] Recently, audio object separation technology has been evolving toward separating and analyzing various objects in real-time within increasingly complex environments. Among these, on-device real-time audio object separation technology refers to a technique that enables users to separate audio objects directly on the device itself, without the need for a separate cloud connection. This technology plays a crucial role in various applications in real-time environments, such as speech recognition, noise reduction, speech augmentation, and background noise suppression.

[0003] In these audio object separation techniques, methods that analyze audio signals using artificial intelligence models, particularly deep learning-based neural network models, are primarily utilized. Neural network models learn complex audio patterns in the time-frequency domain from sample audio data and can separate one or more independent audio objects from an audio signal.

[0004] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.

[0005] According to one embodiment of the present disclosure, an electronic device may be provided. The electronic device may include: an audio input unit comprising a circuitry; a memory comprising at least one storage medium storing at least one instruction; and at least one processor configured to execute the at least one instruction individually and / or collectively. The at least one processor may be configured to individually and / or collectively cause the electronic device to: convert a first input data comprising audio data for a first frame into the frequency domain to obtain a first frequency data; generate a first mask data for audio object separation using the first frequency data; delay the first frequency data by a first frame delay; obtain a first frequency object data by applying the delayed first frequency data to the first mask data; and convert the first frequency object data into the time domain to obtain a first object data. The first frame delay may be a number of frames related to the time taken to generate and apply the mask data from the frequency data.

[0006] According to one embodiment, at least one processor may individually and / or collectively cause the electronic device to: convert second input data, which is audio data input at a later time than the first input data and includes input audio data for the first frame, into the frequency domain to obtain second frequency data. The at least one processor may be configured to generate second mask data for audio object separation using the second frequency data, obtain second frequency object data by delaying the second frequency data by the first frame delay and applying it to the second mask data, and obtain second object data by inversely converting it into the time domain, and cause the object data portion for the first frame included in the first object data to overlap with the object data portion for the first frame included in the second object data to obtain overlapping object data for the first frame.

[0007] According to one embodiment, at least one processor may be configured to cause the electronic device to: obtain superimposed object data for a first frame by superimposing the audio object data portion for a first frame included in the second object data and the audio object data portion for a first frame included in the first object data. The second object data may be data including audio object data for a first frame as object data obtained prior to the first object data.

[0008] According to one embodiment, at least one processor may be configured to cause the electronic device to perform overlap by applying at least one window function to the object data portion for the first frame included in the first object data and the object data portion for the first frame included in the second object data, and to obtain overlapped object data for the first frame.

[0009] According to one embodiment, the first object data includes data stored in a buffer having a size that is an integer multiple of the audio samples per frame, and includes data containing audio object data for the first frame. The second object data includes data stored in a buffer having a size that is an integer multiple of the audio samples per frame, and may include data containing audio object data for a second frame following the first frame and at least one frame preceding the second frame, stored in frame order. The at least one frame preceding the second frame may include the first frame.

[0010] According to one embodiment, at least one processor may be configured to cause the electronic device to: delay input video data for a first frame by a number of frames corresponding to the total number of frames delayed from when audio data for the first frame is input until when it is output.

[0011] According to one embodiment, at least one processor may be configured to cause the electronic device to: store first frequency data in a first storage space of a first-in, first-out structure in memory, and to acquire first frequency object data by applying first mask data at the time when the first frequency data is output from the first storage space. The first storage space includes a storage space of a size such that a time equal to the first frame delay is required from the time when the first frequency data is input until it is output, and the first frequency data may be data stored in a buffer having a size that is an integer multiple of the frequency component per frame, and may be data in which data including frequency components for the first frame and at least one frame preceding the first frame is stored in frame order.

[0012] According to one embodiment, at least one processor may be configured to cause the electronic device to: separate the voice from the first input data for a first frame of input audio data containing voice, obtain first object data which is voice object data, and convert the first object data into corresponding text.

[0013] According to one embodiment, at least one processor may be configured to cause the electronic device to: convert first object data obtained by separating the voice into text in a first language, convert the text in the first language into text in a second language through machine translation, generate voice object data in a second language from the text in the second language using a text-to-speech (TTS) model, and cause the volume of the first object data in the first input data to be reduced and the voice object data in the second language to be added.

[0014] According to one embodiment, at least one processor may be configured to cause the electronic device, individually and / or collectively: to separate a voice object from first input data containing voice to obtain first object data, and to adjust or remove the volume of the first object data from the first input data. Additionally, at least one processor may be configured to cause the electronic device, individually and / or collectively: to adjust or remove the remaining volume excluding the first object data from the first input data.

[0015] According to one embodiment, at least one processor may be configured to cause the electronic device to: distinguish and separate audio data of multiple objects by object in input data comprising audio data of multiple objects combined.

[0016] According to one embodiment, at least one processor may be configured to cause the electronic device to implement spatial audio by individually and / or collectively transmitting different outputs for audio object data to spatially separated multiple audio output units, respectively, by taking into account the position information of the audio object. At least one processor may be configured to cause the electronic device to implement spatial audio by individually and / or collectively recognizing the user's position and head direction, and applying a head-related transfer function (HRTF) to the audio object data to generate output audio data that takes into account the user's position and head direction and the position information of the audio object.

[0017] According to one embodiment of the present disclosure, a method of operation of an electronic device may be provided. The method of operation of the electronic device may include at least one of: receiving audio data from an external source; converting the received audio data into a frequency domain; obtaining mask data for audio object separation using the frequency domain data; applying a delay to the frequency domain data; performing audio object separation by applying mask data corresponding to the frequency domain data to which the delay has been applied; inversely converting the audio object data in the frequency domain into a time domain; and overlapping a plurality of audio object data for the same audio frame.

[0018] According to one embodiment, a method of operating an electronic device may include an operation of overlapping an audio object data portion for a first frame included in first object data and an audio object data portion for a first frame included in second object data. The second object data may be object data obtained prior to the first object data and may include audio object data for a first frame.

[0019] According to one embodiment, the method of operating an electronic device may include at least one of the operation of storing first frequency data in a first storage space and the operation of acquiring first frequency object data by applying first mask data at the time when the first frequency data is output from the first storage space.

[0020] According to one embodiment, the method of operating an electronic device may include at least one of the operation of receiving video data for a first frame and the operation of outputting the input video data for the first frame after delaying it by a second frame delay.

[0021] According to one embodiment, a method of operating an electronic device may include at least one of the following operations: generating text in a first language corresponding to first object data; generating text in a second language through machine translation from text in a first language; generating voice object data in a second language using a TTS model from text in a second language; and reducing the volume of the first object data in the first input data and adding voice object data in a second language.

[0022] According to one embodiment, the method of operating an electronic device may include the operation of adjusting or removing the sound level of a first object data from a first input data.

[0023] The term "audio object" or "acoustic object" described in this disclosure may refer to, for example, an individual audio element that can be separated into a single independent sound source unit, such as a component of a specific sound or audio signal. Additionally, the term "audio object" or "acoustic object" may also be referred to as an "audio source."

[0024] In the present disclosure including the drawings, the same or similar reference numerals may be used for identical or similar components. Other aspects, features, and advantages of the foregoing description and specific embodiments of the present disclosure will become more apparent through the detailed description that follows with reference to the accompanying drawings.

[0025] FIG. 1 is a block diagram illustrating an exemplary configuration of an electronic device of the present disclosure according to various embodiments.

[0026] FIGS. 2a and 2b are drawings illustrating exemplary operations in which an electronic device performs audio object separation according to various embodiments.

[0027] FIG. 3 is a diagram illustrating the use of a buffer for storing input audio data in an electronic device according to various embodiments.

[0028] FIG. 4 is a diagram illustrating an example of a method for an electronic device to apply a delay in a data processing process according to various embodiments.

[0029] FIGS. 5A and 5B are drawings illustrating exemplary methods for an electronic device to apply superposition in a data processing process according to various embodiments.

[0030] FIGS. 6a and 6b are flowcharts for illustrating exemplary operation processes of an electronic device according to various embodiments.

[0031] FIG. 7 is a diagram illustrating an exemplary operation in which an electronic device applies a delay to input audio and video data according to various embodiments.

[0032] FIGS. 8a, 8b, 8c, 8d, and 8e are drawings illustrating examples of an electronic device separating and using a voice object according to various embodiments.

[0033] FIG. 9 is a diagram illustrating an exemplary operation for separating various audio objects by object according to various embodiments.

[0034] FIGS. 10a, 10b, and 10c are drawings for illustrating the operation of an electronic device implementing spatial acoustics according to various embodiments.

[0035] Hereinafter, various exemplary embodiments of the present disclosure are described in detail with reference to the drawings. However, the present disclosure may be embodied in various different forms and is not limited to the various exemplary embodiments described herein, and should be understood to include various modifications, equivalents, or substitutions of the various embodiments. The present disclosure may be modified by those skilled in the art without departing from the essence of the present disclosure, including the claims, and such modifications should be understood to be within the technical spirit or scope of the present disclosure.

[0036] In the following disclosure, functions, configurations, technical terms, and technical details well known in the art to which this disclosure pertains may be omitted. This is intended to convey the essentials of this disclosure more clearly and concisely by minimizing or reducing unnecessary detailed descriptions.

[0037] In the drawings, each block of the flowcharts and combinations of the flowcharts may be performed by at least one instruction. The instruction may be loaded into a processor of a computer or other programmable data processing equipment to generate means for performing the functions described in the drawings. The instruction may also provide steps for performing the functions described in the drawings by being executed on a computer or other programmable data processing equipment.

[0038] Various elements and regions in the drawings are depicted schematically, and the technical concept of the present disclosure is not limited by the relative sizes, spacing, or arrangements depicted in the attached drawings. The electronic device of the present disclosure is not limited to the configuration and / or operation shown in the drawings and may include all other configurations capable of performing the same or similar functions.

[0039] The individual components depicted in the drawings are not necessarily physically separated but are separated to aid in the explanation and understanding of the present disclosure. The present disclosure may include configurations in which the individual components shown in the drawings are merged, modified, or have some components deleted and / or added. Likewise, the operations depicted in the drawings are illustrative for the purpose of explanation and understanding, and the present disclosure may be modified by merging, changing the order of, or deleting and / or adding parts of the operations depicted in the drawings. For example, two or more operations depicted consecutively in the drawings may be performed substantially simultaneously or, if necessary, in reverse order.

[0040] FIG. 1 illustrates an exemplary block configuration of an electronic device according to various embodiments.

[0041] The electronic device (100) of FIG. 1 may be a smartphone, tablet PC, PC, smart TV, mobile phone, PDA (personal digital assistant), laptop, media player, micro server, digital broadcasting terminal, navigation, kiosk, home appliance, and other mobile or non-mobile computing devices, but is not limited thereto. Additionally, the electronic device (100) may perform various computing functions such as real-time video viewing and communication. The various exemplary embodiments of the electronic device (100) of the present disclosure described below may be equally applicable to other electronic devices having audio processing functions.

[0042] According to one embodiment, the electronic device (100) may include at least one processor (110) (e.g., including a processing circuit), an audio input unit (120) (e.g., including an audio input circuit), and a memory (130).

[0043] According to one embodiment, the memory (130) includes a storage medium used by the electronic device (100) and can store data such as at least one instruction (132) corresponding to at least one program or setting information. The program may include an operating system (OS) program and various application programs. When the at least one instruction (132) stored in the memory (130) is executed by at least one processor (110), it can cause the electronic device (100) to perform at least one operation.

[0044] According to one embodiment, the memory (130) may include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), RAM (random access memory, RAM), SRAM (static random access memory), ROM (read only memory, ROM), EEPROM (electrically erasable programmable ROM), PROM (programmable ROM), magnetic memory, a magnetic disk, and an optical disk.

[0045] According to one embodiment, an artificial intelligence model (131) may be stored in the memory (130). The artificial intelligence model (131) may be, for example, a multi-layered computing system, and each layer may include a neural network model comprising multiple basic units, such as neurons or nodes. The artificial intelligence model (131) may include data that can be used to learn specific patterns or characteristics from input data and to analyze or process new data based thereon. The artificial intelligence model (131) may be, for example, a neural network model that generally includes an input layer, a hidden layer, and an output layer, and may process input data hierarchically to convert it into an output. In this case, each neuron may be used to perform calculations on input values ​​through weights, biases, and / or activation functions, and may be used to transmit the calculation results to the next layer. Through this structure, the artificial intelligence model (131) can be used to learn the relationships between input data.

[0046] According to one embodiment, the artificial intelligence model (131) may include an object separation network that receives an audio signal containing a mixture of multiple objects as input and aims to separate specific object signals to be separated. The object separation network may be used to learn weights in a direction that minimizes / reduces the separated signals from the designated target signal.

[0047] According to one embodiment, a buffer may be statically and / or dynamically allocated on the memory (130). The buffer includes a storage space for temporarily storing data, and there may be a specific relationship between the order in which data is input and the order in which it is output. For example, it may have a first-in-first-out (FIFO) structure or a last-in-first-out (LIFO) structure, but is not limited thereto.

[0048] According to one embodiment, data at each step of the audio object separation process may be stored in the buffer in units of audio samples. For example, at least one of input audio data, data obtained by converting the input audio data into the frequency domain, data in the frequency domain after object separation, or data obtained by inversely converting the data in the frequency domain after object separation into the time domain may be stored in the buffer for each audio sample.

[0049] According to one embodiment, data for the same audio sample may be copied multiple times as needed and stored in multiple different buffers within the memory (130). The different buffers may refer, for example, to buffers allocated at different locations within the storage space included in the memory (130).

[0050] According to one embodiment, the audio input unit (120) includes various circuits and can receive audio data through a tuner, an input / output unit (e.g., including a circuit) and / or a communication unit (e.g., including a communication circuit). The audio input unit (120) may include at least one of the tuner and the input / output unit. The tuner can select the frequency of a broadcast channel to be received by the electronic device (100) from among many radio wave components by tuning through amplification, mixing, resonance, etc., of a broadcast signal received via wired or wireless means. The broadcast signal may include audio and additional data. The input / output unit may include at least one of an audio jack, an audio input port, and a USB input port capable of receiving audio data from an external device. The communication unit includes various communication circuits and can transmit and receive audio data from an external server and / or other electronic device via a wired and / or wireless network. The communication unit may include a function capable of streaming or downloading audio data in real time.

[0051] According to one embodiment, the audio input unit (120) may receive an analog audio signal from an external source and sample it. The sampling may refer, for example, to measuring an analog signal at regular time intervals. At this time, the number of times the analog signal is measured per second is called the sampling rate, and for example, a sampling rate of 48 kHz may refer to measuring an analog signal at regular time intervals of, for example, about 20.83 μs (microsecond). The audio input unit (120) may quantize the sampled data to convert it into discrete values ​​according to a predetermined bit depth, and convert it into a digital code to finally obtain digital audio data.

[0052] According to one embodiment, the audio input unit (120) may directly receive digital audio data from an external source. In this case, the audio input unit (120) may resample the input digital data to change the existing sampling rate. For example, the audio input unit (120) may receive audio data of 48 kHz and convert it to 44.1 kHz, or conversely, convert data of 44.1 kHz to 48 kHz.

[0053] According to one embodiment, at least one processor (110) includes various processing circuits and can perform control, computation, and / or data processing of at least a part of an electronic device (100) by executing at least one instruction (132) stored in memory (130).

[0054] According to one embodiment, at least one processor (110) may include at least one processing circuit and / or multiple processors. One or more of the at least one processor (110) may be configured to perform various functions described in the present disclosure individually and / or collectively. Where in the present disclosure, "processor," "at least one processor," or "one or more processors" are described as being configured to perform various functions, these terms may cover, for example, a situation in which one processor performs some of the cited functions and other processor(s) perform other parts of the cited functions, and may also cover, but are not limited to, a situation in which a single processor can perform all of the cited functions. Additionally, at least one processor (110) may include a combination of processors performing the cited / disclosed various functions, for example, in a distributed manner. At least one processor (110) may execute program instructions to achieve or perform various functions.

[0055] According to one embodiment, at least one processor (110) may include at least one of a central processing unit (CPU), a graphics processing unit (GPU), a neural network processing unit (NPU), a digital signal processor (DSP), a microcontroller unit (MCU), a sensor hub, a supplementary processor, a communication processor, an application processor, an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA), and may have a plurality of cores.

[0056] According to one embodiment, at least one processor (110) may include an audio DSP and an NPU. The audio DSP is a microprocessor specialized in the digital processing of audio signals and can effectively process operations such as filtering or fast Fourier transform (FFT) that require high-speed computation. The NPU may include a processor specialized in neural network computation and may be a processor optimized for processing parallel computation of machine learning and / or deep learning models.

[0057] According to one embodiment, at least one processor (110) can convert an input audio signal in the time domain into the frequency domain. In this case, the separated object signal can be inversely converted back to the original time domain.

[0058] According to one embodiment, the transformation may be performed using various mathematical transformation operations. For example, the transformation may be performed using a discrete Fourier transform (DFT), a short-time Fourier transform (STFT), and / or a fast Fourier transform (FFT). According to one embodiment, the inverse transformation may be performed through the inverse operation of the transformation, for example, through an inverse DFT (IDFT), an inverse STFT (ISTFT), and / or an inverse FFT (IFFT), but is not limited thereto.

[0059] According to one embodiment, at least one processor (110) can perform preprocessing on data in which an input signal has been converted into the frequency domain, such as performing operations such as filtering or normalization. At least one processor (110) can generate mask data for audio object separation from the preprocessed data using an artificial intelligence model (131) and / or an object separation network included in the artificial intelligence model (131).

[0060] According to one embodiment, the mask data may be data that serves as a filter used to emphasize a specific object or suppress another object in an audio object separation operation. The mask data may be one of a binary mask that indicates whether a corresponding signal component corresponds to a target signal using 0 and 1, or a soft mask that expresses the degree of correspondence to the target signal using a continuous value between 0 and 1. The mask data may be data in the form of a numeric vector that is applied, for example, by multiplying it by audio data in the frequency domain.

[0061] According to one embodiment, at least one processor (110) can use an artificial intelligence model (131) to analyze the characteristics of various objects from given various sample audio signals, distinguish object data included in each sample audio signal, learn the relationships, and optimize and / or improve the artificial intelligence model (131).

[0062] According to one embodiment, in the learning process, at least one processor (110) adds audio signals and corresponding object information to the artificial intelligence model (131) as learning data, analyzes patterns regarding the frequency, time, and acoustic characteristics of each sample audio, and can optimize the weights and biases of the artificial intelligence model (131). The artificial intelligence model (131) can be improved through this learning process to identify previously learned objects in new input audio signals or to separate them by object. The separation may refer, for example, to extract each object as an independent signal so that they can be used individually in subsequent processing.

[0063] According to one embodiment, at least one processor (110) can perform audio object separation by applying the mask data to audio data in the frequency domain. In this case, the at least one processor (110) may not immediately apply the mask data to audio data in the frequency domain for a specific frame, but may apply a delay of a specified number of frames and apply the mask data at a later point in time equal to the specified number of frames.

[0064] FIGS. 2a and 2b are drawings for illustrating exemplary operation according to whether or not a delay is applied to an electronic device (e.g., the electronic device (100) of FIG. 1) according to various embodiments.

[0065] In the following, for the sake of brevity, some expressions are written as follows:

[0066] - 'Input data' refers to audio data received from an audio input unit (e.g., audio input unit (120) of FIG. 1) and transmitted to a subsequent module, and refers to a signal converted from an external analog signal into a digital signal and / or a signal resampled from an external digital signal. When the input data is stored in a buffer within the memory (130), the buffer is referred to as an 'input buffer'.

[0067] - 'Frequency data' refers to data obtained by converting input data into the frequency domain. When the above frequency data is stored in a buffer within the memory (130), the buffer is called a 'frequency buffer'.

[0068] - 'Frequency object data' refers to data in the frequency domain obtained by performing object separation by applying mask data to the frequency data.

[0069] - 'Object data' refers to data obtained by inversely converting frequency object data into the time domain. When the object data is stored in a buffer within memory (130), the buffer is called an 'object buffer'.

[0070] In the following description, for the sake of brevity, the words defined above are followed by a frame number to indicate that they correspond to a specific audio frame. For example, @k is appended to indicate that it corresponds to the k-th frame. For instance, frequency data@k refers to frequency data corresponding to the k-th frame. If data for multiple frames is stored in a buffer, the buffer is designated by attaching the number of the last frame among the multiple frames. For instance, if data for the k-th, k-1st, and k-2nd frames is stored in a frequency buffer, that frequency buffer is designated as frequency buffer@k.

[0071] The NPU (2140) illustrated in FIGS. 2a and 2b is an example of a processing unit that may be included in at least one processor (e.g., at least one processor (110) of FIG. 1), but the present disclosure is not limited thereto. The NPU (2140) may be replaced with at least one other processing unit capable of generating mask data for audio object separation using an artificial intelligence model (e.g., the artificial intelligence model (131) of FIG. 1).

[0072] Referring to FIG. 2a, an electronic device (100) according to one embodiment of the present disclosure may perform at least one of the following exemplary operations:

[0073] - Operation (o2210) of converting input data@i(2110) for the i-th frame to obtain frequency data@i(2121),

[0074] - Operation (o2230) of preprocessing frequency data@i(2121) and transmitting it to NPU(2130),

[0075] - An operation (o2240) of generating mask data using an object separation network included in an artificial intelligence model (131) in an NPU (2130) and transmitting it to a subsequent module, and

[0076] - An operation (o2251) to perform corresponding object separation by applying mask data@id (2140) immediately to frequency data@i (2121) without delay processing (o2221).

[0077] At this time, the above d is the number of frames corresponding to the processing time (e.g., the total time required for operation o2230 and operation o2240) until mask data is obtained from frequency data, and if mask data is applied directly to frequency data @i (2121), mask data @id (2140) can be applied (o2251). The reason why frequency data @i (2121) cannot be matched with mask data @i is that obtaining mask data from frequency data requires processing time, such as computation time, time to input data into memory (130), or time to read stored data. For the same reason, mask data @i generated from frequency data @i (2121) can subsequently be matched with frequency data @i+d. Accordingly, since matching for the same frame is not achieved, a degradation in audio object separation performance may occur.

[0078] According to one embodiment, the processing time to obtain mask data from frequency data (e.g., the total time required for operations o2230 and o2240) is not always constant but may vary depending on various factors such as the load state of the relevant processor (e.g., NPU (2130)) and the complexity of the frequency data. Accordingly, to set an appropriate frame delay to optimize audio object separation performance, a frame delay d can be set using a statistical representative value after obtaining statistics on the processing time for numerous samples. The frame delay d may correspond to the number of frames in which the frequency data is delayed, or a frame-based offset for delaying the frequency data. According to one embodiment, the statistical representative value for setting the frame delay d may be a value obtained by at least one of the following exemplary methods:

[0079] (i) A value converted into frames after averaging the processing times, or the average of the values ​​converted into frames. Since processing times are continuous values ​​and the number of frames is a natural number, 'value converted into frames' here and below may mean 'a value obtained by dividing by the time per frame and rounding up, up, or down to the first decimal place.'

[0080] (ii) The median of the processing times converted into frames, or the median of the processing times converted into frames.

[0081] (iii) The mode of the values ​​converted from processing times into frames.

[0082] (iv) A representative value obtained by any one of (i) to (iii) above, after excluding some of the higher values ​​among the processing times. For example, a representative value obtained by any one of (i) to (iii) for the 95th percentile and lower values ​​of the processing times.

[0083] According to one embodiment, the frame delay d may be a number of frames predetermined by any one of the methods (i) to (iv) above. According to one embodiment, the frame delay d may be updated based on additional statistics even after it has been initially set.

[0084] Referring to FIG. 2b, an electronic device (100) according to one embodiment of the present disclosure may perform at least one of the following exemplary operations:

[0085] - Operation (o2210) of converting input data@i(2110) for the i-th frame to obtain frequency data@i(2121),

[0086] - Operation (o2230) of preprocessing frequency data@i(2121) and transmitting it to NPU(2130),

[0087] - An operation (o2240) of generating mask data for audio object separation using an object separation network included in an artificial intelligence model (131) in an NPU (2130) and transmitting it to a subsequent module,

[0088] - Operation to apply a delay of d frames to frequency data (o2222), and

[0089] - An operation (o2252) to perform corresponding object separation by applying mask data@id(2140) to frequency data@id(2122) with a delay of d frames applied from a point in time d frames earlier.

[0090] That is, in FIG. 2b, mask data of the same frame is applied to the frequency data so that object separation can be performed (2252). Accordingly, audio object separation can be performed more accurately.

[0091] Table 1 below shows, for one embodiment of the present disclosure, the object separation performance according to the k value, e.g., the difference in the number of frames between the frequency data and the mask data, when the frequency data @i (2121) and the mask data @ik (e.g., mask data @id (2140)) are matched, as an indicator of the signal-to-distortion ratio (SDR). SDR is an indicator that measures the distortion between the separated audio signal and the original signal, and a higher SDR value indicates better audio object separation performance.

[0092] kSDR (dB)05.8315.7025.3834.8944.4254.02

[0093] According to one embodiment related to Table 1, when no delay was applied as in FIG. 2a, d=2, i.e., referring to row k=2 of Table 1, an SDR of 5.38dB was obtained. When a delay of 2 frames was applied as in FIG. 2b, resulting in accurate matching (k=0), an SDR of 5.83dB was obtained by referring to Table 1. In other words, it was confirmed that the audio object separation performance was improved by applying a delay of d frames.

[0094] According to one embodiment, the operations illustrated in FIGS. 2a and FIGS. 2b can be performed by at least one processor (110) executing at least one instruction (132) stored in memory (130).

[0095] According to one embodiment, among the operations illustrated in FIGS. 2a and 2b, the operation of converting input data into frequency data (o2210), the operation of preprocessing and transmitting it to an NPU (o2230), and / or the operation of applying mask data to frequency data with or without delay (o2222) to perform object separation (o2252 and / or o2251) may be performed by a separate processor other than the NPU (2130). For example, it may be performed by an audio DSP.

[0096] According to one embodiment, the operation (o2230) of preprocessing frequency data in a separate processor and sending it to the NPU (2130) and / or the operation (o2240) of receiving mask data back can be performed by the two processors directly exchanging data or by exchanging data through a shared memory instead of directly exchanging data. According to one embodiment, the shared memory may be a random access memory (RAM).

[0097] FIG. 3 illustrates an example in which input data is stored frame by frame in an input buffer in memory (130) in an electronic device according to various embodiments (e.g., the electronic device (100) of FIG. 1). FIG. 3 illustrates an example in which the size of the input buffer is four times the number of audio samples per frame for ease of understanding.

[0098] According to one embodiment, as described above, the electronic device (100) may use a mathematical transformation operation to convert input data into the frequency domain. In this case, the frequency interval is inversely proportional to the number of audio samples being converted, and thus, as the number of audio samples increases, the frequency resolution increases, making precise frequency component analysis possible. Therefore, rather than converting each frame individually, if data from multiple frames is stored in an input buffer and then the data in the input buffer is converted, a larger number of audio samples can be converted, thereby enabling more precise frequency component analysis.

[0099] According to one embodiment, the number of audio samples that can be stored in the input buffer may be an integer multiple of the number of audio samples per frame. That is, if the number of audio samples per frame is F and the number of audio samples that can be stored in the input buffer is B, then for a natural number n, B may be equal to nF. For example, as shown in the example of FIG. 3, n may be equal to 4 and B may be equal to 4F.

[0100] According to one embodiment, with reference to FIG. 3, when n=4, at the time when input data@i (314) is stored in input buffer@i (310), input data (311, 312, 313, and 314) of the i-3, i-2, i-1, and i-th frames may be stored in order in input buffer@i (310). The input buffer@i (310) may have a first-in, first-out structure, for example, at a point in time after one frame, the oldest input data@i-3 (311) may be deleted, the input data of the remaining three frames may be moved forward, and new input data@i+1 (324) may be added to the last empty part. As a result, it may become input buffer@i+1 (320).

[0101] FIG. 4 is a diagram illustrating a state in which a plurality of frequency buffers are stored in a memory (130) at each time point, storing frequency data frame by frame, in an electronic device (e.g., the electronic device (100) of FIG. 1) according to various embodiments.

[0102] According to one embodiment, with reference to FIG. 4, d+1 frequency buffers of the same size for consecutive frames may be stored simultaneously in the memory (130). The d may be the number of frames corresponding to the processing time until mask data is obtained from frequency data. With reference to FIG. 4, for example, at frame@i time (410), frequency buffer@i, frequency buffer@id, and frequency buffers of the frames in between may be stored in the memory (130). At each passing frame, the oldest frequency buffer may be deleted and a new frequency buffer may be added.

[0103] According to one embodiment, with reference to FIG. 4, when the oldest frequency buffer among the stored frequency buffers (e.g., frequency buffer @i (431) at frame @i+d) is matched with mask data at the same time point, the mask data for the same frame can be accurately matched. That is, the delay application operation (e.g., operation o2222 of FIG. 2b) and the object separation operation through accurate matching (e.g., operation o2252 of FIG. 2b) can be performed by storing and using multiple frequency buffers as in FIG. 4. For example, at frame @i time point (410), frequency buffer @id (411) and mask data @id can be matched, at frame @i+1 time point (420), frequency buffer @i-d+1 (421) and mask data @i-d+1 can be matched, and at frame @i+d time point (430), frequency buffer @i (431) and mask data @i can be matched.

[0104] According to one embodiment, delay can be applied by storing and using more than d+1 buffers in a manner similar to FIG. 4. For example, k frequency buffers can be stored for k > d+1, and delay can be applied by reading and utilizing a portion of the required frame (e.g., the portion of frequency buffer@id at frame@i) from the storage space.

[0105] FIG. 5a is a diagram illustrating examples in which object data according to various embodiments is stored frame by frame in an object buffer in memory (130). FIG. 5a is illustrated for ease of understanding in an example where the size B of the object buffer is four times the number F of audio samples per frame, for example, B=4F.

[0106] Referring to FIG. 5a, if each object buffer of size B is represented as a [1:B] section, then object data for the i-th frame is stored in the [3F+1:B] section (511) of object buffer@i (510), the [2F+1:3F] section (521) of object buffer@i+1 (520), the [F+1:2F] section (531) of object buffer@i+2 (530), and the [1:F] section (541) of object buffer@i+3 (540). According to one embodiment, by overlapping all or part of this data, an overlapping object data@i, which is a more stable audio object separation result for the i-th frame, can be obtained.

[0107] FIG. 5b is a diagram illustrating an example in which object data according to various embodiments is stored in memory (130). FIG. 5b illustrates an example in which the size B of the object buffer is four times the number F of audio samples per frame.

[0108] Referring to FIG. 5b, according to one embodiment, by deleting object data that has already been overlapped and output, only object data for the frame to be overlapped and subsequent frames may be left in memory (130). For example, at the time when object data for frame i+3 is to be overlapped (top figure of FIG. 5b), object buffer @i+3 (540), object buffer @i+2 (530) with object data for frame @i-1 deleted, object buffer @i+1 (520) with object data for frame @i-1 and frame @i-2 deleted, and object buffer @i (510) with object data for frame @i-3 to frame @i-1 deleted may be stored.

[0109] Referring to FIG. 5b, in operation o542, an electronic device (e.g., the electronic device (100) of FIG. 1) can overlap and output object data for frame @i. As object data for frame i is output, object data for frames @i+1 through @i+3 may remain in the object buffer @i+3 (540).

[0110] Referring to FIG. 5b, in operation o552, the electronic device (100) can overlap and output object data for frame @i+1. Accordingly, object data for frames @i+2 through @i+4 may remain in object buffer @i+4 (550), and object data for frames @i+2 and @i+3 may remain in object buffer @i+3 (530) according to operations o542 and o552.

[0111] Referring to FIG. 5b, in operation o562, the electronic device (100) can overlap and output object data for frame @i+2. Accordingly, object data for frames @i+3 through @i+5 may remain in object buffer @i+5 (560), and object data for frames @i+3 and @i+4 may remain in object buffer @i+4 (550) according to operations o552 and o562. And only object data for frame @i+3 may remain in object buffer @i+3 (540) according to operations o542, o552 and o562.

[0112] According to one embodiment, when B=nF, up to n object data can be superimposed in the manner shown in FIG. 5a. In this case, the total storage space occupied for superposition is nB=n*nF, and therefore the total storage space occupied is It can be proportional to.

[0113] According to one embodiment, when B=nF, up to n object data can be superimposed in the manner shown in FIG. 5b. At this time, the size of the total storage space occupied for superposition is, and therefore, the size of the total storage space occupied is It can be proportional to. In this case, for natural numbers n > 1 Therefore, when performing nesting in the manner shown in Fig. 5b, memory can be saved compared to when performing nesting in the manner shown in Fig. 5a.

[0114] According to one embodiment, when multiple different object data for the same frame are superimposed, the influence of noise can be reduced by canceling out random noise compared to the case where only one object data is obtained. Accordingly, the sound of the desired object can be obtained more clearly.

[0115] According to one embodiment, when multiple different object data for the same frame are superimposed, compared to the case where only one object data is obtained, the inaccuracy of each mask data calculation algorithm, such as overshoot and / or undershoot caused by applying overweight or underweight to a specific frequency, can be compensated for. That is, different object data obtained as a result of applying different mask data can produce mutually complementary results.

[0116] According to one embodiment, when multiple different object data obtained at different frame times for the same frame are superimposed, the degree of variation in mask data between adjacent frames can be reduced. Accordingly, the consistency of object separation between adjacent frames is increased, and problems of audio interruption and / or unnatural transitions between frames can be improved. For example, when object data for frame @i+1 is superimposed in the example of FIG. 5a, the corresponding parts in object buffer @i+1 (520), object buffer @i+2 (530), object buffer @i+3 (540), and object buffer @i+4 (550) are superimposed. Meanwhile, when object data for frame @i is superimposed in the example of FIG. 5a, the corresponding parts in object buffer @i+1 (520), object buffer @i+2 (530), and object buffer @i+3 (540) are superimposed as described above. That is, in the example of Fig. 5a, each overlapping object data is a result obtained by overlapping the results of applying four different mask data, and overlapping object data@i and overlapping object data@i+1 may have three of the four mask data identically. Accordingly, the inter-frame variability of object separation is reduced, and when transitioning from overlapping object data@i to overlapping object data@i+1, a natural flow of object sound can be obtained.

[0117] According to one embodiment, with reference to FIGS. 5a and 5b, for n=B / F, there may be an additional delay of n-1 frames to overlap object data. For example, as in FIG. 5, when n=4, there may be an additional delay of 3 frames to overlap four pieces of data. If a delay of d frames is applied prior to the overlap as in FIG. 2b (2222), the final delay may be d+n-1 frames.

[0118] According to one embodiment, the superposition may be performed by applying at least one window function in addition to simply calculating the arithmetic mean of each signal. The at least one window function is a function used to apply weights to specific intervals of a signal during signal processing, and may include, for example, at least one of a Hann window, a Hamming window, and a rectangular window.

[0119] FIG. 6a is a flowchart illustrating an exemplary process for separating an audio object from an electronic device (100) and outputting object data for a first frame according to various embodiments.

[0120] Referring to FIG. 6a, in operation 610, the electronic device (100) can obtain first input data for a first frame from an external audio input unit (e.g., audio input unit (120) of FIG. 1). According to one embodiment, the first input data may be stored in an input buffer.

[0121] Referring to FIG. 6a, in operation 620, the electronic device (100) can convert the first input data into the frequency domain to obtain the first frequency data. According to one embodiment, the first frequency data can be stored in a frequency buffer.

[0122] Referring to FIG. 6a, in operation 631, the electronic device (100) can obtain first mask data from first frequency data using an object separation network. According to one embodiment, operation 631 can be performed by preprocessing the first frequency data in a first processor (e.g., audio DSP) and sending it to a second processor (e.g., NPU), then performing calculations on the first mask data using an object separation network in the second processor and then sending it back to the first processor.

[0123] Referring to FIG. 6a, in operation 632, the electronic device (100) may apply a delay of a first delay (681) to the first frequency data. The first delay (681) is a delay corresponding to the processing time to obtain mask data from the frequency data, for example, the time required for operation 631 in FIG. 6, and may be, for example, a delay (2222) of d frames in FIG. 2b. The method of applying a delay of the first delay (681) may be, for example, as described in FIG. 4, but is not limited thereto.

[0124] Referring to FIG. 6a, in operation 640, the electronic device (100) can obtain first frequency object data by applying first mask data to first frequency data to which a delay has been applied and performing object separation.

[0125] Referring to FIG. 6a, in operation 650, the electronic device (100) can obtain first object data by, for example, inversely converting first frequency object data into the time domain. The first object data can be stored in an object buffer.

[0126] According to one embodiment, the electronic device (100) can output first object data. At this time, the time required for operations 610, 620, 640, and 650 may be negligible compared to the time per frame, and accordingly, the total time required from the input of audio data for the first frame to the output of object data may be equal to the first delay. The first delay may be a delay of d frames, for example, as exemplified in FIG. 2b.

[0127] FIG. 6b is a flowchart illustrating an exemplary operation to obtain superimposed object data by superimposing the plurality of audio data after obtaining different plurality of audio data for a first frame in an electronic device (100) according to various embodiments.

[0128] Referring to FIG. 6b, in operation 660, the electronic device (100) may obtain first superimposed object data by superimposing multiple object data for a first frame, including first object data. The method of performing superimposition may be, for example, as described in FIG. 5a and / or FIG. 5b, but is not limited thereto. Additional delay may occur due to superimposition, and said additional delay may be n-1 frames, for example, as described in FIG. 5a and FIG. 5b.

[0129] According to one embodiment, the electronic device (100) may output first superimposed object data. At this time, the time required for operations 610, 620, 640, and 650 may be negligible compared to the time per frame, and accordingly, the total time required from the input of audio data for the first frame to the output of superimposed object data may be equal to the sum of the time required for the first delay and the time required for operation 660, for example, a delay equal to the second delay. The second delay may be a delay equal to d+n-1 frames, for example, as exemplified in FIG. 2b, FIG. 5a and FIG. 5b.

[0130] The operations described in FIGS. 6a and 6b can be performed by executing at least one instruction (e.g., at least one instruction (132) of FIG. 1) in at least one processor (e.g., at least one processor (110) of FIG. 1).

[0131] FIG. 7 is a diagram illustrating examples in which audio data and video data are input and output in an electronic device (e.g., the electronic device (100) of FIG. 1) according to various embodiments.

[0132] As described above in FIGS. 6a and 6b, according to one embodiment, the total delay from the time when audio data is input to the electronic device (100) (710) to the time when it is output (712) may be equal to the second delay (682).

[0133] Referring to FIG. 7, according to one embodiment, the electronic device (100) can output video by applying a second delay (682) from the time (720) when video data is input in order to synchronize video and audio (722).

[0134] FIGS. 8a, 8b, 8c, 8d and 8e are drawings illustrating various examples in which an electronic device (e.g., the electronic device (100) of FIG. 1) performs object separation and utilizes audio containing voice according to various embodiments.

[0135] According to one embodiment, in FIGS. 8a, 8b, 8c, 8d and 8e (which may be referred to as FIGS. 8a through 8e), the electronic device (100) receives a sound (810) containing voice and can perform more accurate audio object separation by applying a delay to the frequency data and then applying mask data, as described in FIG. 2b, for example.

[0136] According to one embodiment, the electronic device (100) in FIGS. 8a to 8e receives sound (810) containing voice, and, for example as described in FIG. 5a or FIG. 5b, acquires multiple different audio object data for the same frame and superimposes them to obtain a more stable audio object separation result.

[0137] Referring to FIG. 8a, according to one embodiment, an electronic device (100) may receive sound (810) containing voice and separate only the voice object (830). According to one embodiment, the electronic device (100) may improve the sound quality of video calls, VoIP calls, and / or hearing aids by removing sound (820) other than voice and outputting only the voice object (830).

[0138] Referring to FIG. 8b, according to one embodiment, an electronic device (100) can separate a voice object (830) from a sound containing voice (810) and then output only a sound other than voice (820) from the input audio. For example, if the input audio is a song mixed with voice and accompaniment, only the accompaniment (MR) can be output.

[0139] Referring to FIG. 8c, according to one embodiment, an electronic device (100) can separate a voice object (830) from a sound containing voice (810) and then add the voice object (830) to the sound containing voice (810) to produce output. Accordingly, audio containing a voice (840) that is amplified compared to the input audio can be produced. Similarly, according to one embodiment, only the volume of the voice in the input audio can be increased or decreased by adding or subtracting data that adjusts the volume of the voice object (830) in the sound containing voice (810).

[0140] Referring to FIG. 8d, according to one embodiment, an electronic device (100) can separate a voice object (830) from a sound (810) containing voice, and then apply a speech-to-text (STT) model to the voice object to generate a corresponding text (850).

[0141] According to one embodiment, the STT model may include an acoustic model and / or a language model. The acoustic model may be a model that receives a speech signal and converts it into phonemes. The language model may be a sentence that generates words or sentences by combining the phonemes.

[0142] According to one embodiment, the STT model may include a model utilizing natural language processing (NLP). The NLP may be used to generate contextually appropriate and natural sentences based on the grammatical structure or linguistic rules of a sentence. For example, it may be used for homonym processing, context understanding, and / or the application of accurate grammar.

[0143] Referring to FIG. 8e, according to one embodiment, an electronic device (100) receives an audio signal (811) containing a voice in a first language, separates a voice object (831) in the first language, generates text (851) in the first language by applying a speech-to-text (STT) model to the voice object (831) in the first language, obtains text (852) in the second language from the text (851) in the first language using machine translation, generates a voice object (832) in the second language by applying a text-to-speech (TTS) model to the text (852) in the second language, removes the original voice object (831) from the input audio signal (811), and adds the voice object (832) in the second language to obtain an audio signal (812) containing a voice in the second language. As a result, audio can be output in which only the voice is translated while maintaining background sounds, etc., in the input audio.

[0144] According to one embodiment, the TTS model is a model that converts text data into a voice signal and may be a model based on an artificial intelligence model (e.g., the artificial intelligence model (131) of FIG. 1). The electronic device (100) can use the TTS model based on the artificial intelligence model (131) to learn the acoustic characteristics of a voice signal sample and convert text into natural voice by reflecting the learned intonation, pronunciation, rhythm, etc.

[0145] FIG. 9 is a diagram illustrating that an electronic device (e.g., the electronic device (100) of FIG. 1) according to various embodiments receives audio data (910) in which a plurality of voice objects are mixed and / or merged and performs object separation. The electronic device (100) can use an artificial intelligence model (131) to distinguish and separate the plurality of voice objects individually. For example, orchestral music sounds (910) can be separated by instrument sounds to increase or decrease the sound of specific instruments.

[0146] FIGS. 10a, 10b, and 10c are drawings for explaining that an electronic device (e.g., the electronic device (100) of FIG. 1) according to various embodiments separates at least one audio object by object and uses this to simulate spatial audio. To simulate spatial audio, the electronic device (100) may consider location information of at least one object separated individually. The location information may be location information included in input audio data (e.g., a recording file recorded by recording location information with a plurality of microphone arrays), or may be location information that is virtually generated and / or assigned.

[0147] According to one embodiment, the spatial acoustics may be a technology for processing sound so that the sound is heard by the user as if it originated at a specific location in a real or virtual three-dimensional space. The electronic device (100) can process the sound of a single audio object using spatial acoustics technology so that it is heard as if it originated at a specific location, and can process the sounds of multiple audio objects individually so that they are heard as if they originated at the same or different locations. Spatial acoustics may be implemented in hardware, in software, or by using both hardware and software methods simultaneously.

[0148] Referring to FIG. 10a, according to one embodiment, an electronic device (100) can implement spatial sound in hardware by using location information of each object and a plurality of spatially separated audio output units after individually separating at least one audio object. For example, the electronic device (100) can implement spatial sound with a surround sound system such as a 5.1 channel or 7.1 channel system.

[0149] Referring to FIG. 10b, according to one embodiment, an electronic device (100) can implement spatial acoustics in software using position information and a head-related transfer function (HRTF) of each object after separating at least one audio object individually.

[0150] According to one embodiment, the head transfer function is a mathematical function that indicates how sound changes due to body structure when it reaches the ear from a specific location in space, and can be used to implement the sound as if it were heard from a virtual specific location.

[0151] According to one embodiment, with reference to FIG. 10b, the head transfer function is two transfer functions, class It can be considered by dividing it into. The above class It is a function that acts as a filter for sound reaching the left and right ears, respectively, and can be a function that links how sound changes and arrives compared to an omnidirectional source depending on the sound frequency, azimuth angle, and elevation angle. Humans can accurately determine the location of a sound source in space through auditory cues such as the difference in sound arrival times between the two ears, sound changes due to occlusion or diffraction at the head, and sound changes due to the asymmetrical shape of the outer ear. The above class The above auditory cues used by a person to determine the location of a sound within space may be functions that can be used to implement virtual object location information (1010) by mathematically converting the above auditory cues.

[0152] According to one embodiment, with reference to FIG. 10c, 3D spatial sound can be realized using only stereo audio (e.g., a 2-channel headset) that physically has only two audio outputs by using HRTF. For example, by modeling the virtual locations (1021) of audio objects through differences in the arrival time, frequency, amplitude, and / or waveform of the sound at both ears, the sound can be made to feel as if it is actually coming from a specific location (1022) in space.

[0153] The various exemplary embodiments described above may be implemented in software containing instructions stored on a device-readable storage medium, included in a computer program product in the form of a device-readable storage medium or distributed online through an application store, or implemented within a recording medium readable by a computer or similar device using software, hardware, or a combination thereof.

[0154] Each component according to the various embodiments described above may be composed of a single or multiple entities, and some auxiliary components may be omitted or additionally included. Some components may be integrated into a single entity to perform the same or similar functions as those performed by each corresponding component prior to integration.

[0155] The operations according to the various embodiments described above may be executed sequentially, in parallel, iteratively, or heuristically, or at least some operations may be executed in a different order, omitted, or other operations may be added.

[0156] Although the present disclosure has been described and illustrated with reference to various exemplary embodiments, it will be understood that such various exemplary embodiments are written for illustrative purposes only and are not intended to be limiting. Furthermore, it will be understood by those skilled in the art that various modifications, alternatives, and / or variations of the various exemplary embodiments are possible without departing from the true technical spirit and full technical scope of the present disclosure, including the appended claims and their equivalents. It will also be understood that any embodiment(s) described in the present disclosure may be used in conjunction with any other embodiment(s) described in the present disclosure.

Claims

1. In an electronic device, Memory for storing at least one instruction; and at least one processor including a processing circuit; comprising, At least one processor individually and / or collectively causes the electronic device: A first input data including input audio data for a first frame is converted into the frequency domain to obtain first frequency data, and Using the above first frequency data, first mask data for audio object separation is generated, and The first frequency object data is obtained by delaying the first frequency data by a first frame delay and applying the delayed first frequency data to the first mask data, wherein the first frame delay is the number of frames related to the time taken to generate and apply the mask data from the frequency data. Configured to cause the first object data to be acquired by converting the first frequency object data into the time domain. Electronic device.

2. In Paragraph 1, The above-mentioned at least one processor is, By overlapping the audio object data portion for the first frame included in the first object data and the audio object data portion for the first frame included in the second object data, a first overlapping object data is obtained, and The second object data above includes object data acquired prior to the first object data, and includes audio object data for the first frame. Electronic device.

3. In Paragraph 2, At least one processor individually and / or collectively causes the electronic device: In the audio object data portion for the first frame included in the first object data and the audio object data portion for the first frame included in the second object data, Configured to perform the nesting by applying at least one window function, thereby causing to obtain nested object data for the first frame, Electronic device.

4. In Paragraph 2 or 3, The first object data is data stored in a buffer having a size that is an integer multiple of the audio sample per frame, and includes data in which data containing audio object data for the first frame is stored. The second object data includes data stored in a buffer of an integer multiple of the size of the audio samples per frame, and includes data stored in frame order including audio object data for a second frame following the first frame and at least one frame preceding the second frame, wherein the at least one frame preceding the second frame includes the first frame. Electronic device.

5. In Paragraph 1, The first frequency data includes data stored in a buffer having a size that is an integer multiple of the frequency component per frame, and includes data in which frequency components for the first frame and at least one frame preceding the first frame are stored in frame order. At least one processor, individually and / or collectively, causes the electronic device: The first frequency data is stored in a first storage space having a first-in, first-out structure within the memory, and the first storage space is a storage space of a size that takes time equal to the first frame delay from the time the first frequency data is input until it is output. Configured to cause the acquisition of first frequency object data by applying the first mask data at the time when the first frequency data is output from the first storage space, Electronic device.

6. In Paragraph 1, Paragraph 2, or Paragraph 3, At least one processor individually and / or collectively causes the electronic device: It is configured to cause the input video data for the first frame to be output with a delay of the second frame delay, and The second frame delay is a number of frames corresponding to the total number of frames delayed from when audio data regarding the first frame is input until it is output, Electronic device.

7. In Paragraph 1, The above first input data includes data including voice, and The above first object data includes voice object data separated from the voice, and At least one processor individually and / or collectively causes the electronic device: Configured to cause the generation of text corresponding to the first object data above, Electronic device.

8. In Paragraph 1, The above first input data includes data containing speech in a first language, and The above first object data includes voice object data obtained by separating speech in the above first language, and The above-mentioned at least one processor, individually and / or collectively, causes the electronic device: Generate text in the first language corresponding to the first object data, and A text in a second language is generated from the text in the first language through machine translation, and A text-to-speech model is used to generate speech object data in the second language from the text in the second language, and Configured to cause the sound volume of the first object data in the first input data to be reduced and the voice object data in the second language to be added, Electronic device.

9. In Paragraph 1, The above first input data includes data including voice, and The above first object data includes voice object data separated from the voice, and At least one processor individually and / or collectively causes the electronic device: Configured to cause the sound level of the first object data to be adjusted or removed from the first input data, Electronic device.

10. In Paragraph 1, At least one processor individually and / or collectively causes the electronic device: Configured to cause the sound level of the remaining part of the first input data, excluding the first object data, to be adjusted or removed. Electronic device.

11. In Paragraph 1, The above first input data includes data in which audio data of multiple objects are combined, and The above first object data individually includes audio data of multiple objects, and At least one processor individually and / or collectively causes the electronic device: Configured to cause audio data of multiple objects included in the first object data to be distinguished and separated by object, Electronic device.

12. In Paragraph 1, The above electronic device is, Implementing spatial audio in hardware and / or software, The above hardware implementation is: At least one processor individually and / or collectively considers location information of audio objects and transmits different outputs for the first object data to spatially separated plurality of audio output units, respectively. The above software implementation is: At least one processor individually and / or collectively recognizes the user's position and head direction, and generates output audio data that takes into account the user's position and head direction and the position information of the audio object by using a head-related transfer function (HRTF) on the first object data. Electronic device.

13. In a method of operating an electronic device, Operation of receiving audio data for the first frame; An operation of converting first input data, including input audio data for a first frame, into the frequency domain to obtain first frequency data; The operation of generating first mask data for audio object separation using the above first frequency data; An operation to obtain first frequency object data by delaying the first frequency data by a first frame delay and applying the delayed first frequency data to the first mask data; and The operation of converting the first frequency object data into the time domain to obtain the first object data; is included, The above first frame delay is the number of frames related to the time taken to generate and apply mask data from frequency data, method.

14. In Paragraph 13, The operation of overlapping the audio object data portion for the first frame included in the first object data and the audio object data portion for the first frame included in the second object data; further comprising, The second object data above includes object data acquired prior to the first object data, and includes audio object data for the first frame. method.

15. In Paragraph 13 or 14, The first object data includes data stored in a buffer having a size that is an integer multiple of the audio sample per frame, and includes data in which data including audio object data for the first frame is stored. The second object data includes data stored in a buffer of an integer multiple of the size of the audio samples per frame, and includes data stored in frame order including audio object data for a second frame following the first frame and at least one frame preceding the second frame, wherein the at least one frame preceding the second frame includes the first frame. method.