Technique for removing vocal component from audio source data on basis of plurality of processing units

Multiple processing units with different sound source separation models enhance the speed and accuracy of vocal component removal, addressing the limitations of single-unit methods and improving musical quality.

WO2025211672A1PCT designated stage Publication Date: 2025-10-09ANAI CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/004157
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-01
Filing Date
2025-03-31
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing sound source separation technologies face challenges in accurately and efficiently removing vocal components from audio data due to the complexity and diversity of frequency components, particularly when using a single processing unit, which fails to reflect musical nuances and leads to listening issues.

Method used

A method utilizing multiple processing units, each employing distinct sound source separation models, to parallelly process audio data, followed by mixing the outputs to enhance accuracy and speed, with one unit having superior vocal component removal performance but slower processing, and another faster but with lesser accuracy.

Benefits of technology

Improves the speed and accuracy of vocal component removal from audio data while enhancing musical quality, as validated by professional evaluation using quantitative indicators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025004157_09102025_PF_FP_ABST
    Figure KR2025004157_09102025_PF_FP_ABST
Patent Text Reader

Abstract

A method according to several embodiments of the present disclosure comprises steps in which: a first processing unit inputs N-th audio source data into a first audio source separation model so as to acquire first audio source data from which a vocal component is removed; a second processing unit differing from the first processing unit inputs the N-th audio source data into a second audio source separation model differing from the first audio source separation model, so as to acquire second audio source data from which the vocal component is removed; and a third processing unit differing from the first processing unit and the second processing unit mixes the first audio source data and the second audio source data so as to generate N-th mixing data, where N can be a natural number greater than 0.
Need to check novelty before this filing date? Find Prior Art

Description

A technique for removing vocal components from sound source data based on multiple processing units

[0001] The present disclosure relates to a technique for removing vocal components from sound source data based on a plurality of processing units.

[0002] Typically, stereo sound source (AR: ALL Recorded) data output from speakers of karaoke machines, karaoke accompaniment machines, MP3 players, and other audio devices includes vocal components (Vocal recorded) which are the singing components of a singer, and non-vocal components (MR: Music recorded) which are accompaniment music using at least one instrument.

[0003] Recently, MR signals have been widely used as accompaniment music among consumers. To generate MR signals from sound source data, vocal components must be effectively removed from the sound source data. While sound source separation technology is a long-standing research field in the field of audio signal processing, the vocal components contained in sound source data span a certain length of frequency range, making it technically very difficult to completely remove them without loss. In particular, the frequency components of vocals vary significantly from person to person depending on factors such as gender and race compared to other instruments, and most of the frequency components overlap with those of general instruments. Therefore, the technical difficulty of perfectly separating only the vocal component is very high.

[0004] Meanwhile, techniques for removing vocal components from audio data primarily rely on a single processing unit. This approach has several limitations in terms of separation speed and data processing accuracy. In particular, using a single processing unit fails to adequately reflect the complexity and diversity encountered during audio data processing.

[0005] Furthermore, in the past, engineering experts attempted improvements based solely on technical aspects, without musical criteria. While this resulted in objectively efficient results, it lacked sufficient evaluation and analysis of the complexity and nuances of music expression. Consequently, significant listening issues often arose. For example, even when a sound was rated fast or objectively high in quality, it was difficult to fully grasp the complexity of the audio data without the listening evaluation of actual music experts.

[0006] These problems remain important challenges in the development of sound source separation technology, and new approaches are needed to solve them.

[0007] [Prior Art Literature]

[0008] (Patent Document 1) Republic of Korea Patent Registration No. 10-2018286 (registered on August 29, 2019)

[0009] The present disclosure aims to address the aforementioned and other issues. The technical challenge of some embodiments of the present disclosure is to enhance the speed and accuracy of sound source separation technology while also improving musical quality through the use of multiple processing units.

[0010] The technical problems to be achieved in the present disclosure are not limited to the technical problems mentioned above, and other technical problems not mentioned can be clearly understood by a person having ordinary skill in the technical field of the present disclosure from the description below.

[0011] A method according to some embodiments of the present disclosure includes the steps of: inputting N-th order sound source data into a first sound source separation model by a first processing unit to obtain first sound source data from which a vocal component has been removed; inputting the N-th order sound source data into a second sound source separation model different from the first sound source separation model by a second processing unit different from the first processing unit to obtain second sound source data from which a vocal component has been removed; and mixing the first sound source data and the second sound source data by a third processing unit different from the first processing unit and the second processing unit to generate N-th mixed data; wherein N may be a natural number greater than 0.

[0012] According to some other embodiments of the present disclosure, the step of obtaining the first sound source data may be performed in parallel with the step of obtaining the second sound source data.

[0013] According to some other embodiments of the present disclosure, the first sound source separation model may have better vocal component removal performance compared to the second sound source separation model, and may take a longer time to remove vocal components when removing vocal components from sound source data using the same processing unit, and the first processing unit may take a shorter time to remove vocal components when removing vocal components using the same sound source separation model compared to the second processing unit.

[0014] According to some other embodiments of the present disclosure, the step of generating the Nth mixing data by the third processing unit may include: adjusting the first sound source data to have a first decibel value by the third processing unit; adjusting the second sound source data to have a second decibel value lower than the first decibel value; and mixing the first sound source data having the first decibel value and the second sound source data having the second decibel value to generate the Nth mixing data.

[0015] According to some other embodiments of the present disclosure, the Nth-order sound source data may be generated by dividing the original sound source data into preset time units by the third processing unit.

[0016] According to some other embodiments of the present disclosure, N may have a value from 1 to M, and M may be a natural number greater than 1 and may be a value determined based on the length of the original sound source data.

[0017] According to some other embodiments of the present disclosure, when the Nth mixing data is generated, the method may further include a step of generating the N+1st mixing data using the N+1st sound source data while the Nth mixing data is played.

[0018] According to some other embodiments of the present disclosure, the step of generating the N+1st mixing data may include: a step of inputting the N+1st sound source data into the first sound source separation model by the first processing unit to obtain third sound source data from which a vocal component has been removed; a step of inputting the N+1st sound source data into a second sound source separation model having a different performance from the first sound source separation model by a second processing unit having a different computational capability from the first processing unit to obtain fourth sound source data from which a vocal component has been removed; and a step of mixing the third sound source data and the fourth sound source data to generate N+1st mixing data.

[0019] A computer program stored in a computer-readable storage medium according to some other embodiments of the present disclosure may perform steps of removing a vocal component from sound source data, the steps including: inputting N-th order sound source data into a first sound source separation model by a first processing unit to obtain first sound source data from which the vocal component has been removed; inputting the N-th order sound source data into a second sound source separation model different from the first sound source separation model by a second processing unit different from the first processing unit to obtain second sound source data from which the vocal component has been removed; and mixing the first sound source data and the second sound source data by a third processing unit different from the first processing unit and the second processing unit to generate N-th mixed data; wherein N may be a natural number greater than 0.

[0020] According to some other embodiments of the present disclosure, a device for removing a vocal component from sound source data includes: a storage unit storing a first sound source separation model and a second sound source model different from the first sound source separation model; a first processing unit; a second processing unit different from the first processing unit; and a third processing unit different from the first processing unit and the second processing unit; wherein the first processing unit inputs N-th order sound source data into the first sound source separation model to obtain first sound source data from which the vocal component has been removed, the second processing unit inputs the N-th order sound source data into the second sound source separation model to obtain second sound source data from which the vocal component has been removed, and the third processing unit mixes the first sound source data and the second sound source data to generate N-th order mixed data, wherein N may be a natural number greater than 0.

[0021] The technical solutions obtainable in the present disclosure are not limited to the solutions mentioned above, and other solutions not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below.

[0022] The effect of a method and device for removing vocal components from sound source data using a plurality of processing units according to the present disclosure is described as follows.

[0023] Some embodiments of the present disclosure may improve the speed and accuracy of removing vocal components from audio data, and may also improve the musical quality of the final product.

[0024] The effects that can be obtained through the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below.

[0025] Various embodiments of the present disclosure are described with reference to the drawings, wherein like reference numerals are used to refer to similar elements generally. In the following examples, for purposes of explanation, numerous specific details are set forth to provide a comprehensive understanding of one or more embodiments. However, it will be apparent that such embodiments may be practiced without these specific details.

[0026] FIG. 1 is a block diagram illustrating a device for removing vocal components from sound source data according to some embodiments of the present disclosure.

[0027] FIG. 2 is a flowchart illustrating an example of a method for removing vocal components from original sound source data according to some embodiments of the present disclosure.

[0028] FIG. 3 is a flowchart illustrating an example of a method for removing vocal components from sound source data containing vocal components using a plurality of processing units according to some embodiments of the present disclosure.

[0029] Hereinafter, various embodiments of a vocal component removal device and method according to the present disclosure will be described in detail with reference to the drawings. Regardless of the drawing symbols, identical or similar components are given the same reference numerals and redundant descriptions thereof will be omitted.

[0030] The purpose and effects of the present disclosure, as well as the technical configurations for achieving them, will become clearer with reference to the embodiments described in detail below, along with the accompanying drawings. In describing one or more embodiments of the present disclosure, if a detailed description of a related known technology is deemed to obscure the gist of at least one embodiment of the present disclosure, the detailed description will be omitted.

[0031] The terms of this disclosure are defined in consideration of the functions of this disclosure, and may vary depending on the intention or custom of the user or operator. In addition, the attached drawings are merely intended to facilitate easy understanding of one or more embodiments of this disclosure, and the technical spirit of this disclosure is not limited by the attached drawings, and should be understood to include all modifications, equivalents, or substitutes included in the spirit and technical scope of the present invention.

[0032] The suffixes “module” and “part” used for components in the following description are given or used interchangeably only for the convenience of writing the present disclosure, and do not have distinct meanings or roles in themselves.

[0033] While terms including ordinal numbers, such as "first" and "second," may be used to describe various components, these components are not limited by these terms. These terms are used solely to distinguish one component from another. Accordingly, a "first" component referred to below may also be a "second" component within the technical scope of the present disclosure.

[0034] Singular expressions include plural expressions unless the context clearly indicates otherwise. That is, unless otherwise specified or the context clearly indicates a singular form, the singular should generally be construed as meaning "one or more" in this disclosure and claims.

[0035] In this disclosure, it should be understood that terms such as “comprising,” “including,” or “having” are intended to specify the presence of a feature, number, step, operation, component, part, or combination thereof described in this disclosure, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0036] The term "or" in this disclosure should be understood as an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from context, "X utilizes A or B" is intended to mean either of the natural inclusive permutations. That is, if X utilizes A; X utilizes B; or X utilizes both A and B, "X utilizes A or B" can apply to any of these cases. Furthermore, the term "and / or" as used herein should be understood to refer to and encompass all possible combinations of one or more of the associated listed items.

[0037] The terms “information” and “data” as used herein may be used interchangeably.

[0038] Unless otherwise defined, all terms (including technical and scientific terms) used in this disclosure may be used with the meaning commonly understood by those of ordinary skill in the art of this disclosure. Furthermore, terms defined in commonly used dictionaries should not be overly interpreted unless specifically defined otherwise.

[0039] However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various other forms. These specific embodiments are provided solely to fully inform those skilled in the art of the present disclosure of the scope of the present disclosure, and the present disclosure is defined solely by the scope of the claims. Therefore, such definitions should be based on the overall content of the present disclosure.

[0040] Techniques for removing vocal components from audio data typically rely on a single processing unit. However, the present disclosure utilizes multiple processing units, rather than a single processing unit, to remove vocal components from audio data. A detailed description of this will be provided with reference to FIGS. 1 through 3.

[0041] FIG. 1 is a block diagram illustrating a device for removing vocal components from sound source data according to some embodiments of the present disclosure.

[0042] The device (100) described in the present disclosure may include any device that performs at least one of transmitting, receiving, and outputting data, content, services, and applications.

[0043] The device (100) of the present disclosure can be paired or connected with other devices, external servers, etc. via a wired / wireless network, and can transmit / receive predetermined data through this. In this case, the data transmitted / received via the device (100) can be converted before being transmitted / received.

[0044] The device (100) of the present disclosure may include, for example, a standing device such as a personal computer (PC), a microprocessor, a mainframe computer, a digital processor, a device controller, a network TV, a hybrid broadcast broadband TV (HBBTV), a smart TV, an internet protocol TV (IPTV), a digital TV, a digital signage, etc., a mobile device (mobile device or handheld device) such as a smart phone, a tablet PC, a notebook, etc., and a server such as an application server, a computing server, a database server, a file server, a web server, etc., but is not limited thereto.

[0045] When the term “device (100)” is used in this disclosure, it may refer to a computer system or a computer device, a fixed device, a mobile device, or a server depending on the context, and may be used to mean all of them unless specifically mentioned.

[0046] In some embodiments of the present disclosure, the device (100) can remove vocal components from original sound source data and reproduce the sound in real time. Here, the original sound source data may be sound source data uploaded to an external server and playable in real time. That is, the device (100) can provide sound source data uploaded to an external server (streaming server) and reproduced by removing vocal components in real time. However, the present disclosure is not limited thereto.

[0047] Referring to FIG. 1, the device (100) may include a first processing unit (110), a second processing unit (120), a third processing unit (130), a communication unit (140), and a storage unit (150). The components illustrated in FIG. 1 are not essential for implementing the device (100), and thus, the device (100) described in the present disclosure may have more or fewer components than the components listed above.

[0048] The first processing unit (110), the second processing unit (120), and the third processing unit (130) can process signals, data, information, etc. input or output through components of the device (100) or run application programs stored in the storage unit (150) to provide or process appropriate information or functions.

[0049] The first processing unit (110), the second processing unit (120), and the third processing unit (130) may be different processing units. Accordingly, when the first processing unit (110), the second processing unit (120), and the third processing unit (130) each process the same task, the processing times may be different from each other.

[0050] The first processing unit (110) may perform a task of removing a vocal component from specific sound source data using a first sound source separation model. In addition, the second processing unit (120) may perform a task of removing a vocal component from specific sound source data using a second sound source separation model that is different from the first sound source separation model. Here, the first sound source separation model and the second sound source separation model may be stored in the storage unit (150) as pre-trained artificial intelligence models to remove vocal components from input sound source data. However, the present disclosure is not limited thereto.

[0051] In the present disclosure, the first sound source separation model may refer to an artificial intelligence model that has better vocal component removal performance than the second sound source separation model and takes longer to remove vocal components from sound source data using the same processing unit. Here, the vocal component removal performance may be determined based on the degree of separation of sound source data from which the vocal component has been removed, i.e., how well the vocal component has been removed from the sound source data.

[0052] The vocal component removal performance of the first sound source separation model and the second sound source separation model can be determined by a music expert verifying sound source data with vocal components removed, output from each of the first sound source separation model and the second sound source separation model. However, the present disclosure is not limited thereto.

[0053] The first processing unit (110) may be a processing unit that takes a shorter time to remove vocal components when removing vocal components using the same sound source separation model compared to the second processing unit (120). Here, the sound source separation model may have a neural network structure (or deep neural network structure) that removes vocal components from input sound source data. In other words, the first processing unit (110) may be a processing unit that can remove vocal components from input sound source data at a faster rate using the same deep learning and / or machine learning algorithm (model).

[0054] As a result, the first processing unit (110) having a better processing speed and computational ability than the second processing unit (110) can perform a task of removing a vocal component from the input sound source data using the first sound source separation model having a better vocal component removal performance. In addition, the second processing unit (120) having a lower processing speed and computational ability than the first processing unit (110) can perform a task of removing a vocal component from the input sound source data using the second sound source separation model having a worse vocal component removal performance than the first sound source separation model. In this case, the first processing unit (110) and the second processing unit (120) can simultaneously and in parallel remove vocal components from the same sound source data to generate sound source data from which vocal components have been removed with different degrees of separation, and the time required to obtain the sound source data from which vocal components have been removed may not differ significantly.

[0055] For example, when comparing the processing speed of a task of removing vocal components from sound data using the same sound source separation model, let's assume that the processing speed of a GPU (Graphics Processing Unit) is faster than that of an NPU (Neural Processing Unit). In this case, the GPU can remove vocal components from the sound data using the first sound source separation model as the first processing unit (110), and the NPU can remove vocal components from the sound data using the second sound source separation model as the second processing unit (120). However, the above-described example is only an example and is not limited thereto, and various types of processing units can be used as the first processing unit (110) and the second processing unit (120).

[0056] Meanwhile, the processing units selected as the first processing unit (110) and the second processing unit (120) are not determined as specific processing units, and may be changed when the processing speed changes due to technological development. For example, if the processing speed of the NPU becomes faster than that of the GPU due to continuous development of the NPU, the NPU as the first processing unit (110) may remove the vocal component from the sound source data using the first sound source separation model, and the GPU as the second processing unit (120) may remove the vocal component from the sound source data using the second sound source separation model.

[0057] Compared to the first processing unit (110) and the second processing unit, the third processing unit (130) may be the processing unit that takes the longest time to remove vocal components when removing vocal components using the same sound source separation model. Therefore, the third processing unit (130) may not be used when removing vocal components using the sound source separation model, but may be used when performing other tasks.

[0058] The third processing unit (130) can control at least some of the components of the device (100) to drive an application program stored in the storage unit (150). Furthermore, the third processing unit (130) can operate at least two or more of the components included in the device (100) in combination to drive the application program.

[0059] For example, the third processing unit (130) may be a central processing unit (CPU). Accordingly, the third processing unit (130) may be a processing unit that primarily uses more complex logic processing than the first processing unit (110) and the second processing unit (120). However, the present disclosure is not limited thereto.

[0060] The storage unit (150) can store data supporting various functions of the device (100). The storage unit (150) can store a plurality of application programs (or applications) running on the device (100), data for the operation of the device (100), commands, and at least one program command. At least some of these application programs can be downloaded from an external server via wireless communication. In addition, at least some of these application programs may exist on the device (100) from the time of shipment for the basic functions of the device (100).

[0061] The storage unit (150) can store any form of information generated or determined by the first processing unit (110), the second processing unit (120), and the third processing unit (130) and any form of information received through the communication unit (140).

[0062] The storage unit (150) may include at least one type of storage medium among a flash memory type, a hard disk type, an SSD (Solid State Disk type), an SDD (Silicon Disk Drive type), a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, and an optical disk. The device (100) may also be operated in relation to web storage that performs the storage function of the storage unit (150) on the Internet.

[0063] According to some embodiments of the present disclosure, the storage unit (150) may store a first sound source separation model and a second sound source separation model. The first sound source separation model and the second sound source separation model may be pre-trained artificial intelligence models that remove vocal components from input sound source data.

[0064] The first sound source separation model and the second sound source separation model can be trained by updating the weights of the first sound source separation model and the second sound source separation model by backpropagating the difference value (error) between the label data labeled in the training data and the predicted data output from the first sound source separation model and the second sound source separation model.

[0065] Sound source data including vocal components can be input data for learning, and sound source data with vocal components removed can be label data for the input data for learning.

[0066] According to some embodiments of the present disclosure, the input data for learning may include multiple audio data, each of which may include vocal components of various individuals with different genders, races, etc. Each of the multiple audio data may include vocal components of only one individual, or may include vocal components of multiple individuals (the vocal component of the main vocal and the vocal component of the chorus). However, the present disclosure is not limited thereto.

[0067] The first and second sound source separation models may be comprised of a set of interconnected computational units, generally referred to as nodes. These nodes may also be referred to as neurons. A neural network may be comprised of at least one node. The nodes (or neurons) comprising the neural networks may be interconnected by one or more links.

[0068] According to some embodiments of the present disclosure, the number of nodes included in the first sound source separation model may be greater than the number of nodes included in the second sound source separation model. In this case, the first sound source separation model may be applied with more resources than the second sound source separation model, and when performing operations of the first sound source separation model and the second sound source separation model using the same processing unit, the operation speed of the first sound source separation model may be slower than that of the second sound source separation model. However, the present disclosure is not limited thereto.

[0069] Within the first sound source separation model and the second sound source separation model, one or more nodes connected via links can form a relationship between input nodes and output nodes. The concept of input nodes and output nodes is relative, and any node that is in an output node relationship with respect to one node can also be in an input node relationship with respect to another node, and vice versa. As described above, the relationship between input nodes and output nodes can be created based on links. One input node can be connected to one output node via a link, and vice versa.

[0070] In a relationship between input and output nodes connected via a single link, the value of the output node's data can be determined based on the data input to the input node. Here, the link interconnecting the input and output nodes can have a weight. The weight can be variable and can be adjusted by the user or an algorithm to allow the neural network to perform the desired function.

[0071] For example, if one or more input nodes are interconnected to one output node by respective links, the output node can determine the output node value based on the values ​​input to the input nodes connected to the output node and the weights set to the links corresponding to each input node.

[0072] As described above, the first sound source separation model and the second sound source separation model may have one or more nodes interconnected through one or more links to form input node and output node relationships within a neural network. The characteristics of the first sound source separation model and the second sound source separation model may be determined based on the number of nodes and links within the first sound source separation model and the number of associations between the nodes and links, and the weight values ​​assigned to each link.

[0073] The first sound source separation model and the second sound source separation model may be composed of a set of one or more nodes. A subset of the nodes constituting the first sound source separation model and the second sound source separation model may constitute a layer. Some of the nodes constituting the first sound source separation model and the second sound source separation model may constitute a layer based on distances from the initial input node. For example, a set of nodes that are n in distance from the initial input node may constitute n layers. The distance from the initial input node may be defined by the minimum number of links that must be traversed to reach the node from the initial input node. However, this definition of a layer is arbitrary for the purpose of explanation, and the order of layers within the first sound source separation model and the second sound source separation model may be defined in a different manner than described above. For example, a layer of nodes may be defined by its distance from the final output node.

[0074] The initial input node may refer to one or more nodes to which data is directly input without going through a link in a relationship with other nodes among the nodes in the first sound source separation model and the second sound source separation model. Alternatively, in the first sound source separation model and the second sound source separation model, the initial input node may refer to nodes that do not have other input nodes connected by a link in a relationship between nodes based on a link. Similarly, the final output node may refer to one or more nodes that do not have an output node in a relationship with other nodes among the nodes in the first sound source separation model and the second sound source separation model. In addition, the hidden node may refer to nodes that constitute the first sound source separation model and the second sound source separation model other than the initial input node and the final output node.

[0075] A neural network according to some embodiments of the present disclosure may be a neural network in which the number of nodes in an input layer may be the same as the number of nodes in an output layer, and the number of nodes decreases and then increases as it progresses from the input layer to the hidden layer. In addition, a neural network according to another embodiment of the present disclosure may be a neural network in which the number of nodes in an input layer may be less than the number of nodes in an output layer, and the number of nodes decreases as it progresses from the input layer to the hidden layer. In addition, a neural network according to still some other embodiments of the present disclosure may be a neural network in which the number of nodes in an input layer may be greater than the number of nodes in an output layer, and the number of nodes increases as it progresses from the input layer to the hidden layer. A neural network according to still another embodiment of the present disclosure may be a neural network in the form of a combination of the neural networks described above.

[0076] According to some embodiments of the present disclosure, the first sound source separation model and the second sound source separation model may have a deep neural network structure.

[0077] A deep neural network (DNN) can refer to a neural network that includes multiple hidden layers in addition to input and output layers. Using DNN, one can identify latent structures in data.

[0078] Deep neural networks may include convolutional neural networks (CNNs), recurrent neural networks (RNNs), autoencoders, restricted Boltzmann machines (RBMs), deep belief networks (DBNs), Q-networks, U-networks, Siamese networks, and generative adversarial networks (GANs). The description of the above-described deep neural networks is merely exemplary and the present disclosure is not limited thereto.

[0079] The neural network can be trained using at least one of supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. Training of the neural network can be a process of applying knowledge to the neural network for performing a specific operation. Training of the first sound source separation model and the second sound source separation model can be a process of applying knowledge to the neural network for performing an operation of removing vocal components from sound source data.

[0080] The first sound source separation model and the second sound source separation model can be trained in a direction to minimize output errors. In training the first sound source separation model and the second sound source separation model, training data is repeatedly input into the first sound source separation model and the second sound source separation model, and the errors of the outputs of the first sound source separation model and the second sound source separation model and the label data for the training data are calculated, and the errors of the first sound source separation model and the second sound source separation model are backpropagated from the output layer of the first sound source separation model and the second sound source separation model toward the input layer in a direction to reduce the errors, thereby updating the weights of each node of the first sound source separation model and the second sound source separation model.

[0081] The amount of change in the connection weights of each node being updated can be determined by the learning rate. The calculation of the first and second sound source separation models for the input data and the backpropagation of the error can constitute a learning cycle (epoch). The learning rate can be applied differently depending on the number of repetitions of the learning cycle of the first and second sound source separation models. For example, in the early stage of training of the first and second sound source separation models, a high learning rate can be used so that the first and second sound source separation models can quickly secure a certain level of performance, thereby increasing efficiency. In the later stage of training, a low learning rate can be used to increase accuracy.

[0082] In training the first and second sound source separation models, the training data may be a subset of the actual data. Therefore, there may be a learning cycle where errors on the training data decrease but errors on the actual data increase. Overfitting is a phenomenon where excessive training on the training data leads to increased errors on the actual data.

[0083] Overfitting can increase errors in machine learning algorithms. Various optimization methods can be used to prevent overfitting. These include increasing the training data, regularization, dropout (inactivating some network nodes during the learning process), and the use of batch normalization layers.

[0084] According to some embodiments of the present disclosure, the first sound source separation model and the second sound source separation model may have a structure that combines a diffusion model and a separation model. In this case, the accuracy of sound source separation may be improved. However, the present disclosure is not limited thereto.

[0085] The communication unit (140) may include one or more modules that enable wired / wireless communication between the device (100) and a wired / wireless communication system, between the device (100) and another device, or between the device (100) and an external server. In addition, the communication unit (140) may include one or more modules that connect the device (100) to one or more networks.

[0086] The communication unit (140) refers to a module for wired / wireless Internet access and can be built into or external to the device (100). The communication unit (140) can be configured to transmit and receive wired / wireless signals.

[0087] The communication unit (140) can transmit and receive wireless signals with at least one of a base station, an external terminal, and a server on a mobile communication network constructed according to technical standards or communication methods for mobile communication (e.g., GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), CDMA2000 (Code Division Multi Access 2000), EV-DO (Enhanced Voice-Data Optimized or Enhanced Voice-Data Only), WCDMA (Wideband CDMA), HSDPA (High Speed ​​Downlink Packet Access), HSUPA (High Speed ​​Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), etc.).

[0088] Wireless Internet technologies may include, for example, WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed ​​Downlink Packet Access), HSUPA (High Speed ​​Uplink Packet Access), LTE (Long Term Evolution), and LTE-A (Long Term Evolution-Advanced). However, the communication unit (140) may transmit and receive data according to at least one wireless Internet technology, including Internet technologies not listed above.

[0089] According to some embodiments of the present disclosure, a device (100) may be connected to an external server via a wired / wireless network through a communication unit (140). Here, the term "wired / wireless network" refers to a communication network that supports various communication standards or protocols for pairing and / or data transmission and reception between the device (100) and other devices, and between the device (100) and an external server. Such a wired / wireless network includes all communication networks currently or to be supported in the future by standards, and can support one or more communication protocols therefor.

[0090] In the present disclosure, the communication unit (140) can obtain original sound source data from an external server under the control of the third processing unit (130). Here, the original sound source data may be sound source data that can be separated into preset time units and may include vocal components. A detailed description thereof will be provided below with reference to FIG. 2.

[0091] FIG. 2 is a flowchart illustrating an example of a method for removing vocal components from original sound source data according to some embodiments of the present disclosure. FIG. 3 is a flowchart illustrating an example of a method for removing vocal components from sound source data containing vocal components using a plurality of processing units according to some embodiments of the present disclosure. With respect to FIGS. 2 and 3, any overlapping content with that described above in FIG. 1 will not be described again, and the differences will be primarily described below.

[0092] Referring to FIG. 2, the third processing unit (130) can obtain original sound source data from an external server through a communication unit (140) (S110).

[0093] Specifically, the third processing unit (130) can control the communication unit (140) to transmit a request-get signal to an external server. In this case, the third processing unit (130) can receive a link address to which audio data is uploaded as a return value through the communication unit (140). Here, the external server may be a streaming server that plays audio data in real time. However, the present disclosure is not limited thereto.

[0094] The third processing unit (130) can convert the streaming URL (Uniform Resource Locator) included in the return value into a file of a specific format (e.g., an mp3 file, a wav (Waveform audio format) file, an AAC (Advanced Audio Coding) file) using a codec (e.g., an ffmpeg codec) and a library. As a result, the third processing unit (130) can obtain original sound source data from an external server through the communication unit (140).

[0095] When the third processing unit (130) acquires original sound source data (S110), it can determine the M value based on the length of the original sound source data. Here, the M value may be a natural number greater than 1 and may be a value indicating how many pieces the original sound source data should be divided into when dividing it into preset time units.

[0096] When the M value is determined in step (S120), the third processing unit (130) can divide the original sound source data into preset time units to generate M sound source data (S130).

[0097] For example, if the length of the sound source data is 34 seconds and the preset time is 15 seconds, the M value can be determined as 3, and the third processing unit (130) can divide the original sound source data into 15-second units to generate three sound source data (two 15-second sound source data and one 4-second sound source data). However, the present disclosure is not limited thereto.

[0098] Meanwhile, when M sound source data are generated (S130), the first processing unit (110), the second processing unit (120), and the third processing unit (130) can generate N-th mixing data using N-th sound source data in step (S140). Here, N can be a natural number greater than 0.

[0099] In the present disclosure, N-th sound source data may indicate which sound source data it is among M sound source data. For example, among three sound source data, the first sound source data may be expressed as primary sound source data, the second sound source data may be expressed as secondary sound source data, and the third sound source data may be expressed as tertiary sound source data.

[0100] Meanwhile, the Nth mixing data can indicate which sound source data the mixing data was generated based on. For example, the mixing data generated using the first sound source data among three sound source data can be expressed as the first mixing data, the mixing data generated using the second sound source data can be expressed as the second mixing data, and the mixing data generated using the third sound source data can be expressed as the third mixing data.

[0101] The process of generating Nth mixing data using Nth sound source data in step (S140) is described in more detail with reference to FIG. 3.

[0102] Referring to FIG. 3, the first processing unit (110) can input N-th order sound source data into the first sound source separation model to obtain first sound source data (S141). In addition, the second processing unit (120) can input N-th order sound source data into the second sound source separation model to obtain second sound source data (S142).

[0103] In the present disclosure, the Nth-order sound source data may have the form of two-dimensional numeric data that can be input and processed into the first sound source separation model and the second sound source separation model.

[0104] Specifically, the third processing unit (130) can convert data loaded as waveform information when loading a sound source into two-dimensional numeric data using STFT (Short-Time Fourier Transform), ISTFT (Inverse Short-Time Fourier Transform), etc. The Nth-order sound source data can be input into the first sound source separation model by the first processing unit (110) as two-dimensional numeric data obtained in this manner and input into the second sound source separation model by the second processing unit (120).

[0105] In the present disclosure, steps (S141, S142) can be performed in parallel. Therefore, the time required to generate sound source data with different separation degrees using different sound source separation models can be shortened.

[0106] When acquiring the first sound source data and the second sound source data in steps (S141, S142), the third processing unit (130) may wait without performing any work. However, when generating the N+1st mixing data after the Nth mixing data is generated, the third processing unit (130) may perform a work of reproducing the Nth mixing data.

[0107] The first processing unit (110) may be a processing unit that takes a shorter time to remove a vocal component when removing a vocal component using the same sound source separation model compared to the second processing unit (120). In addition, the first sound source separation model may have better vocal component removal performance and may take a longer time to remove a vocal component when removing a vocal component from sound source data using the same processing unit compared to the second sound source separation model. Therefore, when the first processing unit (110) that has better processing speed and computational capabilities than the second processing unit (120) acquires first sound source data using the first sound source separation model that has better performance than the second sound source separation model and takes a longer time to remove a vocal component than the second sound source separation model, and the second processing unit (120) acquires second sound source data using the second sound source separation model, the first sound source data and the second sound source data may be acquired at similar times.

[0108] When the first processing unit (110) and the second processing unit (120) each obtain the first sound source data and the second sound source data using the first sound source separation model and the second sound source separation model, the third processing unit (130) can mix the first sound source data and the second sound source data to generate Nth mixing data.

[0109] Specifically, the third processing unit (130) can adjust the first sound source data to have a first decibel value, adjust the second sound source data to have a second decibel value that is smaller than the first decibel value, and mix the first sound source data having the first decibel value and the second sound source data having the second decibel value to generate Nth mixing data.

[0110] The first sound source data is used as the main sound source because it is sound source data generated by the first sound source separation model that has better vocal component removal performance than the second sound source separation model, and the second sound source data is sound source data generated by the second sound source separation model that has worse vocal component removal performance than the first sound source separation model, and thus can be used as a sub sound source. Accordingly, the decibel value of the first sound source data can be set higher than the decibel value of the second sound source data.

[0111] In the present disclosure, the first decibel value may be -3.9 dB, and the second decibel value may be -8.8 dB. However, the present disclosure is not limited thereto, and the first decibel value and the second decibel value may have different values.

[0112] When sound source data generated from two different sound source separation models is mixed and provided, it can be confirmed that the sound quality is improved and the vocal component removal performance is also improved.

[0113] Specifically, by utilizing the professional techniques and quantitative indicators (SDR (Signal-to-Distortion Ratio), SAR (Signal-to-Artifact Ratio), SIR (Signal-to-Interference Ratio)) of music experts, we evaluated the quality of sound source separation of mixed data. As a result, we were able to confirm that the quality of sound source separation (vocal component removal performance) was improved when vocal components were removed from sound source data using a single sound source separation model. The above-described quantitative indicators can play an important role in measuring the efficiency and accuracy of sound source separation and the signal quality of processed sound source data, and it was confirmed that good quality was maintained in the quantitative indicators of sound source data from which vocal components were removed generated by some embodiments of the present disclosure.

[0114] Referring back to FIG. 2, when the generation of the Nth mixing data is completed in step (S140), the third processing unit (130) can adjust the pitch of the Nth mixing data through various pitch normalization techniques. Since the pitch normalization technique utilizes conventional technology, a detailed description thereof will be omitted.

[0115] When pitch normalization is performed on Nth-order mixed data and then played back, listeners who listen to Nth-order mixed data with vocal components removed can hear music with adjusted pitch. Consequently, listeners perceive a greater improvement in sound quality.

[0116] Meanwhile, the third processing unit (130) can check whether the N value corresponds to the M value (whether the N value matches the M value) after the generation of the Nth mixing data is completed (S150).

[0117] If the third processing unit (130) recognizes that the N value does not correspond to the M value (S150, No), i.e., if the generation of mixing data using all sound source data is not completed, the device (100) can be controlled to generate the next mixing data using the next sound source data.

[0118] Specifically, when the M value is 2, the third processing unit (130) can control the device (100) to generate secondary mixing data using secondary sound source data because the N value is 1, not 2, after the primary mixing data is generated using the primary sound source data.

[0119] Meanwhile, the third processing unit (130) can control the device (100) to reproduce the Nth mixing data when the generation of the Nth mixing data is completed. In addition, the third processing unit (130) can control the device (100) to generate the N+1st mixing data using the N+1st sound source data while the Nth mixing data is being reproduced.

[0120] When generating N+1 mixing data, the same operations can be performed as when generating Nth mixing data.

[0121] Specifically, the first processing unit (110) may input N+1 sound source data into the first sound source separation model to obtain third sound source data from which the vocal component has been removed, and in parallel therewith, the second processing unit (120) may input N+1 sound source data into the second sound source separation model to obtain fourth sound source data from which the vocal component has been removed. In addition, the third processing unit (130) may mix the third sound source data and the fourth sound source data to generate N+1 mixing data. Here, when mixing the third sound source data and the fourth sound source data to generate the N+1 mixing data, the third processing unit (130) may adjust the decibel value of the third sound source data to have the first decibel value and adjust the decibel value of the fourth sound source data to have the second decibel value smaller than the first decibel value, and mix the third sound source data and the fourth sound source data with the adjusted decibel values.

[0122] When a user uses a karaoke service, if the user waits for more than 10 seconds, the user's motivation to use the service may decrease. According to some embodiments of the present disclosure, the loading time that occurs during the audio source separation process (the process of removing vocal components from audio data) can be minimized to improve service satisfaction. That is, the device (100) can divide the original audio source data into M audio source data, process the M audio source data sequentially, and play the audio source data from which the vocal component has been removed in real time during the processing of the M audio source data, thereby minimizing the loading time.

[0123] For example, let's assume that the task of removing vocal components from the primary sound source data after dividing the original sound source data into multiple sound source data takes 3 seconds to load. The primary sound source data with the vocal components removed is played for 15 seconds, during which time the vocal component removal can be performed on the secondary sound source data. The task of removing vocal components from the secondary sound source data can be completed before the playback of the primary sound source data is completed. Meanwhile, the task of removing vocal components from the tertiary sound source data can be performed when the secondary sound source data is played immediately after the playback of the primary sound source data is completed.

[0124] If the vocal component is removed from the original audio data as a whole without dividing it into multiple audio data, the loading time may be longer than 3 seconds.

[0125] According to some embodiments of the present disclosure, the user may perceive a loading time of only three seconds. Therefore, loading time can be minimized, thereby improving user satisfaction.

[0126] Meanwhile, in the present disclosure, the number of times to perform sound source separation can be determined based on the length of the original sound source data. That is, sound source separation can be performed M times.

[0127] If the third processing unit (130) recognizes that the N value corresponds to the M value (S150, Yes), it can terminate additional generation of mixing data while playing the currently generated mixing data.

[0128] According to some embodiments of the present disclosure, if the third processing unit (130) recognizes that the N value corresponds to the M value, it can combine all the mixing data in chronological order to generate a single sound source data with the vocal component removed and store the combined sound source data in the storage unit (150) in a file format (e.g., an mp3 file, a wav (Waveform audio format) file, an AAC (Advanced Audio Coding) file). Accordingly, if a user wants to remove the vocal component from the same sound source data in the future, the third processing unit (130) can search for and provide the corresponding sound source data from the storage unit (150).

[0129] According to at least one of the embodiments of the present invention described above, the speed and accuracy of removing vocal components from sound source data can be improved, and the musical quality of the final product can also be improved.

[0130] In the present disclosure, the device (100) is not limited to the configuration and method of some of the embodiments described above, but some of the embodiments may be configured by selectively combining all or some of the embodiments so that various modifications can be made.

[0131] The various embodiments described in the present disclosure may be implemented in a recording medium readable by a computer or similar device, for example, using software, hardware, or a combination thereof.

[0132] In a software implementation, some embodiments, such as the procedures and functions described in this disclosure, may be implemented as separate software modules. Each of the software modules may perform one or more functions, tasks, and operations described in this disclosure. The software code may be implemented as a software application written in a suitable programming language. Here, the software code may be stored in the storage unit (150) and executed by the device (100). That is, at least one program command is stored in the storage unit (150), and at least one program command may be executed by the device (100).

[0133] The method of removing vocal components from sound source data according to some embodiments of the present disclosure can be implemented as codes readable by a plurality of processing units on a recording medium readable by a plurality of processing units provided in the device. The recording medium readable by the plurality of processing units includes all types of recording devices that store data readable by the plurality of processing units. Examples of recording media readable by at least one processor include Read Only Memory (ROM), Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage devices, etc.

[0134] Meanwhile, while the present disclosure has been described with reference to the attached drawings, this is merely an example and is not limited to specific embodiments. Various modifications that can be implemented by those skilled in the art to which the invention pertains are also within the scope of the claims. Furthermore, such modifications should not be understood separately from the technical spirit of the present invention.

Claims

1. A step of inputting N-th sound source data into a first sound source separation model by a first processing unit to obtain first sound source data with vocal components removed; A step of inputting the Nth-order sound source data into a second sound source separation model different from the first sound source separation model by a second processing unit different from the first processing unit to obtain second sound source data with the vocal component removed; and A step of generating Nth mixing data by mixing the first sound source data and the second sound source data by a third processing unit different from the first processing unit and the second processing unit; Including, The above N is, A natural number greater than 0, method.

2. In paragraph 1, The step of acquiring the first sound source data is performed in parallel with the step of acquiring the second sound source data. method.

3. In paragraph 1, The above first sound source separation model is, Compared to the above second sound source separation model, the vocal component removal performance is better, and the time required to remove the vocal component is longer when removing the vocal component from the sound source data using the same processing unit. The above first processing unit, Compared to the second processing unit, the time required for vocal component removal is shorter when vocal component removal is performed using the same sound source separation model. method.

4. In paragraph 3, The step of generating the Nth mixing data by the third processing unit is: A step of adjusting the first sound source data to have a first decibel value by the third processing unit; A step of adjusting the second sound source data to have a second decibel value lower than the first decibel value; and A step of generating the Nth mixing data by mixing the first sound source data having the first decibel value and the second sound source data having the second decibel value; Including. method.

5. In paragraph 1, The above Nth sound source data is, The original sound source data is generated by dividing it into preset time units by the third processing unit. method.

6. In paragraph 5, The above N is, With values ​​from 1 to M, The above M is, A value determined based on the length of the original sound source data as a natural number greater than 1, method.

7. In paragraph 1, When the Nth mixing data is generated, a step of generating the N+1st mixing data using the N+1st sound source data while the Nth mixing data is played; including more, method.

8. In paragraph 7, The steps for generating the above N+1 mixing data are: A step of inputting the N+1-th order sound source data into the first sound source separation model by the first processing unit to obtain third sound source data with the vocal component removed; A step of inputting N+1-order sound source data into a second sound source separation model having different performance from the first sound source separation model by a second processing unit having different computational capabilities from the first processing unit to obtain fourth sound source data with vocal components removed; and A step of generating N+1 mixing data by mixing the third sound source data and the fourth sound source data; including, method.

9. A computer program stored in a computer-readable storage medium, which performs steps of removing vocal components from sound source data, wherein the steps are: A step of inputting N-th sound source data into a first sound source separation model by a first processing unit to obtain first sound source data with vocal components removed; A step of inputting the Nth-order sound source data into a second sound source separation model different from the first sound source separation model by a second processing unit different from the first processing unit to obtain second sound source data with the vocal component removed; and A step of generating Nth mixing data by mixing the first sound source data and the second sound source data by a third processing unit different from the first processing unit and the second processing unit; Including, The above N is, A natural number greater than 0, A computer program stored on a computer-readable storage medium.

10. In a device for removing vocal components from sound source data, the device: A storage unit storing a first sound source separation model and a second sound source model different from the first sound source separation model; First processing unit; a second processing unit different from the first processing unit; and A third processing unit different from the first processing unit and the second processing unit; Including, The above first processing unit, By inputting Nth-order sound source data into the first sound source separation model, first sound source data with vocal components removed is obtained, The second processing unit, By inputting the Nth sound source data into the second sound source separation model, second sound source data with the vocal component removed is obtained, The third processing unit, By mixing the first sound source data and the second sound source data, Nth mixing data is generated, The above N is, A natural number greater than 0, device.

Citation Information

Patent Citations

  • Sound source separating method and system therefor, and speech recognizing method and system therefor

    JP2005077731A

  • Apparatus and method for cancelling vocal

    KR100636248B1

  • Method and Apparatus for Removing Speech Components in Sound Source

    KR102018286B1

  • Optimization system for AI-based potato cultivation and its optimization method

    KR1020240162724A

  • Method and system for mixing multi sound source

    KR102242010B1