Device for detecting music data from video content and control method thereof

By using artificial intelligence models to detect and process music data in video content, the problem of difficulty in efficient extraction and processing of music data in the prior art is solved, and automated processing is realized, improving editing efficiency and convenience.

CN115735360BActive Publication Date: 2025-05-09COCHL INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180036982.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-19
Filing Date
2021-05-18
Publication Date
2025-05-09
Estimated Expiration
2041-05-18

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently extract and process music data from video content, especially when a video platform processes a large amount of data, the method of manually checking copyrighted works data is inefficient.

Method used

The artificial intelligence model is used to detect music data from the audio stream and delete or replace the music data of copyrighted works in the video content to achieve automated processing.

Benefits of technology

It improves the convenience and efficiency of video content editing, significantly reduces the cost of video editing, and enhances the convenience of video content owners or distributors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115735360B_ABST
    Figure CN115735360B_ABST
Patent Text Reader

Abstract

The data processing method according to the present invention comprises the following steps: receiving an input of video content including a video stream and an audio stream; detecting music data from the audio stream; and filtering the audio stream to delete the music data detected in the audio stream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for processing audio data mixed with music and speech. Background Art

[0002] The sound source separation technology divides the audio stream composed of various sounds into multiple audio data according to specific criteria. For example, the sound source separation technology can be used to extract only the singer's voice from stereo music, or to separate two or more audio signals recorded with a microphone. In addition, the sound source separation technology can also be used for noise elimination in vehicles, mobile phones, etc.

[0003] Recently, methods of introducing artificial intelligence into sound source separation technology have been introduced. Typically, there is a method of performing speech separation using pre-trained speech, noise pattern, or statistical data information. In this way, speech separation can be achieved even in a rapidly changing noise environment.

[0004] On the other hand, as the video content market develops, issues related to the copyright of data contained in video content arise. In particular, in the case where video content contains music without the permission of the copyright owner, the distribution of the relevant video content is restricted, and therefore, the need to separate copyrighted work data from video content is increasing.

[0005] That is, an operation is required to confirm whether copyrighted work data is included in video content, or to separate or remove copyrighted work data from original video content, or to change copyrighted work data to license-free data.

[0006] However, according to the previous video editing process, there is the trouble that the editor needs to confirm the above operation while directly playing the video. Considering the amount of data processed on the video platform recently, there is a problem that it is difficult to check a sufficient amount of video content through the existing method that requires users to manually check the copyrighted work data. Summary of the invention

[0007] Technical issues

[0008] The object of the present invention is to provide a data processing device and a control method thereof which can extract music data from any audio stream.

[0009] In addition, an object of the present invention is to provide a data processing device and a control method thereof that can use an artificial intelligence model to determine whether music data exists in any audio stream that does not include a separate label, a tag indicating the classification of audio data, or log information.

[0010] Furthermore, an object of the present invention is to provide a data processing device and a control method thereof which are capable of detecting music data from an original file of video content composed of an audio stream and a video stream, and deleting the detected music data from the original file.

[0011] In addition, an object of the present invention is to provide a data processing device and a control method thereof that can use an artificial intelligence model to detect whether music data exists in an audio stream and can detect the time domain in which music data exists.

[0012] Furthermore, an object of the present invention is to provide a data processing device and a control method thereof that can determine whether music data corresponding to a copyrighted work is included in an audio stream.

[0013] Solutions to the problem

[0014] In order to achieve the above-mentioned purpose, the present invention provides a data processing method, which includes the following steps: receiving an input of video content including a video stream and an audio stream; detecting music data from the above-mentioned audio stream; and filtering the above-mentioned audio stream to delete the above-mentioned music data detected in the above-mentioned audio stream.

[0015] Effects of the Invention

[0016] According to the present invention, music data contained in video content can be detected even if the user does not directly scan the video content, so the convenience of the user who edits the video content can be improved.

[0017] In addition, since music data can be detected for a large amount of video content in a short time, the cost of video editing can be significantly reduced.

[0018] In addition, according to the present invention, since the data processing apparatus deletes music data corresponding to a copyrighted work included in input video content or replaces it with substitute music, the convenience of the owner or distributor of the video content can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 The diagram is a conceptual diagram related to the data processing method according to the present invention.

[0020] Figure 2 is a block diagram showing components of a data processing apparatus according to the present invention.

[0021] Figure 3 The flowchart is a diagram showing an embodiment of a data processing method according to the present invention.

[0022] Figure 4 The flowchart is a diagram showing an embodiment of a data processing method according to the present invention. DETAILED DESCRIPTION

[0023] Best Mode for Carrying Out the Invention

[0024] A data processing method, characterized in that it includes the following steps: receiving an input of video content including a video stream and an audio stream; detecting music data from the above-mentioned audio stream; and filtering the above-mentioned audio stream to delete the above-mentioned music data detected in the above-mentioned audio stream.

[0025] Modes for carrying out the invention

[0026] Hereinafter, the embodiments disclosed in this specification will be described in detail in conjunction with the accompanying drawings. It should be noted that the technical terms used in this specification are only used to illustrate specific embodiments and are not intended to limit the technical ideas disclosed in this specification.

[0027] first, Figure 1 A conceptual diagram related to a data processing method according to the present invention is shown. In the following, video content 1 is defined as a moving image file including an audio stream 2 and a video stream 3. In addition, the audio stream may be composed of music data and / or non-music data.

[0028] As used herein, the term "music" may refer to any type of sound that may be characterized by one or more elements of rhythm (e.g., beat, meter, and articulation), pitch (e.g., melody and harmony), dynamics (e.g., volume of sounds or notes), etc., and may include the sounds of musical instruments, human voices, etc. In addition, the term "copyright work" may refer herein to a unique or distinct musical work or composition, and may include the creation or reproduction of such a musical work or composition in sound or audio form (e.g., a song, a tune, etc.). In addition, the term "audio stream" may refer to a sequence of one or more electrical signals or data representing one or more portions of a sound stream that may include multiple musical works, ambient sounds, speech, noise, etc.

[0029] Reference Figure 1 , the data processing device 100 according to the present invention can scan the audio stream included in the video content to identify whether the audio stream includes music data.

[0030] Specifically, the data processing device 100 may discern whether music data is included in the audio stream using an external server or an artificial intelligence model installed in the data processing device 100. In this case, the artificial intelligence model may be composed of an artificial neural network that performs deep learning or machine learning.

[0031] Figure 2 FIG. 2 is a block diagram showing a data processing device according to an embodiment of the present invention. Figure 2The data processing device 100 of the present invention includes an input unit 110 , an output unit 120 , a memory 130 , a communication unit 140 , a control unit 180 , and a power supply unit 190 .

[0032] More specifically, among the above components, the communication unit 140 may include one or more modules for implementing wireless communication between the data processing device 100 and a wireless communication system, or between the data processing device 100 and another data processing device 100, or between the data processing device 100 and an external server. In addition, the above communication unit 140 may include one or more modules for connecting the data processing device 100 to one or more networks.

[0033] The input unit 110 may include a camera or an image input unit for inputting an image signal, a microphone or an audio input unit for inputting an audio signal, a user input unit (e.g., a touch key, a key (mechanical key), etc.) for receiving information input from a user. Voice data or image data collected by the input unit 110 may be analyzed and processed by a user's control command.

[0034] The output unit 120 is used to generate output associated with vision, hearing or touch, and may include at least one of a display unit, an audio output unit, a tactile module and a light output unit. The display unit may be configured to form a mutual hierarchical structure or an integral body with the touch sensor, thereby enabling a touch screen to be implemented. The above-mentioned touch screen is used as a user input device that provides an input interface between the data processing device 100 and the user, and may also provide an output interface between the data processing device 100 and the user.

[0035] The memory 130 stores data that supports various functions of the data processing device 100. The memory 130 can store a plurality of applications or applications driven by the data processing device 100 and data fragments and commands for the operation of the data processing device 100. At least a part of these applications can be downloaded from an external server via wireless communication. In addition, for the basic functions of the data processing device 100 (e.g., receiving a call, making a call, receiving a message, sending a message, etc.), at least a part of these applications can exist on the data processing device 100 from the factory. On the other hand, the application is stored in the memory 130, is set on the data processing device 100, and can be driven by the control unit 180 to perform the operations (or functions) of the above-mentioned electronic device control device.

[0036] In addition to the operations related to the above-mentioned application programs, the control unit 180 generally controls the overall operation of the data processing device 100. The control unit 180 can provide or process information or functions suitable for the user by processing signals, data, information, etc. input or output by the various components mentioned in the foregoing description or driving the application programs stored in the memory 130.

[0037] In order to drive the application programs stored in the memory 130, the control unit 180 can control Figure 2 Furthermore, in order to drive the above application, the control unit 180 may combine at least two components included in the data processing device 100 to operate together.

[0038] The power supply unit 190 receives external power or internal power under the control of the control unit 180 to supply power to each component included in the data processing device 100. The power supply unit 190 includes a battery, and the battery may be configured to be embedded in the body or configured to be detachable from the body.

[0039] At least a portion of each of the above components can cooperate with each other to facilitate the operation, control or control method of the electronic device control device according to the various embodiments to be described below. In addition, the operation, control or control method of the electronic device control device can be implemented in the electronic device control device by driving at least one application stored in the above memory 130.

[0040] In one example, the data processing device 100 may be implemented in the form of a separate terminal, that is, a terminal such as a desktop computer, a digital television, or a mobile terminal such as a mobile phone, a laptop computer, a PDA, a tablet computer, a laptop computer, a wearable device, etc.

[0041] In the following, reference will be made to Figure 3 and Figure 4 The present invention describes the music data filtering method based on artificial intelligence.

[0042] First, the input unit 110 may receive an input of information related to video content including at least one of an audio stream and a video stream (S300). The input unit 110 may also receive an input of information related to the audio stream.

[0043] In addition, the communication unit 140 may receive information related to video content including at least one of an audio stream and a video stream from an external server or an external terminal.

[0044] That is, the video content or the audio stream may be a file directly uploaded by the user or a file received from an external server.

[0045] The control unit 180 may detect music data from an audio stream contained in input video content (S301). Figure 4 As shown, the step of detecting the above-mentioned music data (S301) includes a process of dividing an audio stream into music data and voice data (S311) and a process of detecting a segment containing music data from the above-mentioned audio stream (S321).

[0046] Specifically, the process of dividing the audio stream into music data and voice data (S311) can be performed by a pre-trained artificial intelligence model. That is, the control unit 180 can use the artificial intelligence model to divide the input audio stream into music data and voice data.

[0047] For example, the artificial intelligence model may receive an audio stream input and output a probability corresponding to music data and a probability corresponding to voice data for each preset unit segment of the input audio stream. That is, the control unit 180 may use the output of the artificial intelligence model to identify whether the audio of each unit segment of the input audio stream corresponds to music data or voice data.

[0048] At this time, the control unit 180 may variably set the unit segment based on the physical characteristics of the audio stream or the physical characteristics of the video content. In addition, the control unit 180 may variably set the unit segment based on the user input applied to the input unit 110. For example, the user input may be related to at least one of accuracy, performance, and processing speed.

[0049] In another example, the artificial intelligence model can output a variable energy distribution based on the sequence of the input audio stream. In this case, the energy distribution can be related to the probability that a portion of the audio stream is music and / or the probability that a portion of the audio stream is speech.

[0050] As another example, the control unit 180 divides the input audio stream into music data and non-music data using the first artificial intelligence model, and divides the divided non-music data into voice data and non-voice data using the third artificial intelligence model.

[0051] At this time, non-speech data refers to audio data that does not correspond to human speech, such as knocking sounds or animal calls, etc. In addition, the first artificial intelligence model can be an artificial neural network for detecting whether there is music, and the third artificial intelligence model can be an artificial neural network for identifying what kind of environmental sound the input audio is.

[0052] Of course, if necessary, the first artificial intelligence model and the third artificial intelligence model can be integrated and configured. In this case, the integrated artificial intelligence model outputs probability values ​​corresponding to multiple categories or tags including music for audio input.

[0053] Next, the control unit 180 may discriminate whether music is included in the target section while sequentially shifting the target section.

[0054] For example, the length of the target segment may be set to 1 second. In addition, the control unit 180 may move the target segment by 0.5 seconds so that the current segment overlaps the previous segment, while discerning whether the target segment includes music.

[0055] Compared with the above-mentioned division process (S311), the difference of the detection process (S321) is that a segment in which both voice and music exist can be detected. In addition, the control unit 180 can perform the detection process (S321) by using a second artificial intelligence model different from the first artificial intelligence model used to perform the division process (S311).

[0056] For example, the first artificial intelligence model used in the division process (S311) can be configured to perform learning using training data labeled as music data and speech data.

[0057] Different from this, the second artificial intelligence model used in the detection process (S321) can be configured to perform learning using training data marked as data including music and data not including music. More specifically, the second artificial intelligence model used in the detection process (S321) can be configured to perform learning using training data marked as data including music in a proportion greater than or equal to a reference value, data including music in a proportion less than or equal to a reference value, and data not including music at all.

[0058] As described above, the control unit 180 can detect music data from the audio stream by using at least one of the execution result of the division process (S311) and the execution result of the detection process (S321). On the other hand, when the accuracy of the division process (S311) is greater than or equal to the reference value, the control unit 180 can omit the detection process (S321).

[0059] In one embodiment, the control unit 180 may perform the above-mentioned detection process (S321) only on a portion of the input audio stream that is classified as music through the division process (S311).

[0060] In another embodiment, the control unit 180 may determine a target for performing the detection process ( S321 ) based on a probability of each unit segment output through the division process ( S311 ) in the input audio stream.

[0061] In another embodiment, similar to the dividing process ( S311 ), the control unit 180 may also perform the above detection process ( S321 ) on the entire input audio stream.

[0062] On the other hand, the control unit 180 uses at least one of the division process (S311) and the detection process (S321) to detect whether each unit segment of the audio stream is music data, and then based on the segment continuity of the detection result, a part of the audio stream can be detected as music data.

[0063] In addition, the control unit 180 can detect the variation pattern of the detected music data and divide the one music data into a plurality of music data based on the detected variation pattern. For example, when different music is continuously streamed and detected as a music data segment, the control unit 180 can divide the above music data into a plurality of music data by monitoring the variation pattern of the music data.

[0064] As described above, when music data is detected (S301), the control unit 180 may perform filtering on the audio stream in a manner of removing the detected music data from the audio stream (S302).

[0065] Specifically, the control unit 180 may delete a portion detected as music data in the audio stream.

[0066] As another example, the control unit 180 may change a portion detected as music data in the audio stream into substitute music data different from the above-described music data.

[0067] In one embodiment, the control unit 180 may determine whether the detected music data is equivalent to a copyrighted work, and perform the above filtering step (S302) according to the determination result. That is, even if music data is detected, if the detected music data is not equivalent to a copyrighted work, the control unit 180 may exclude it from the filtering object. When a plurality of different music data are detected from the audio stream, the control unit 180 may determine whether each music data is a copyrighted work.

[0068] In the process of performing the filtering step S302, in order to consider whether it is a copyrighted work, the memory of the data processing device 100 may store a copyrighted work database composed of information related to the copyrighted work. That is, the control unit 180 may determine whether the detected music data is a copyrighted work by using the copyrighted work database pre-stored in the memory. In addition, when it is determined that the detected music data is a copyrighted work, the control unit 180 may filter the audio stream to delete the music data.

[0069] On the other hand, the control unit 180 may determine the substitute music data in consideration of the characteristics of the detected music data, such as the characteristics related to at least one of genre, atmosphere, composition, tempo, volume, and sound source length.

[0070] In one embodiment, the control unit 180 may analyze information related to the genre and / or atmosphere of the detected music data using a fourth artificial intelligence model, and may select substitute music data based on the analysis result.

[0071] That is, the control unit 180 can detect information related to at least one of the genre and atmosphere of the detected music data by using a fourth artificial intelligence model designed to analyze the genre or atmosphere of the music. In particular, the fourth artificial intelligence model can be configured to perform learning through training data labeled as what genre or atmosphere the music is. In this case, the information obtained by the fourth artificial intelligence model can be in the form of a feature vector.

[0072] In addition, the control unit 180 can calculate the similarity between the detected music data and the alternative music candidate group by comparing the feature vector of the alternative music candidate group with the feature vector of the detected music data. In addition, the control unit 180 can select any one of the multiple alternative music data based on the calculated similarity and change the detected music data to the selected alternative music data.

[0073] In another embodiment, the control unit 180 may convert the alternative music data based on the volume of the detected music data. Specifically, the control unit 180 may calculate the energy level of each reset unit segment for the detected music data. For example, the control unit 180 may set the second unit segment to a segment shorter than the first unit segment applied in the division process (S311), and calculate the energy level of the music data detected for each of the above second unit segments. In one example, the second unit segment may be 0.2 seconds.

[0074] The control unit 180 may apply a low-pass filter defined by a vector composed of the calculated energy levels to the substitute music data, and change the existing music data into the above-mentioned application result.

[0075] On the other hand, the control unit 180 may analyze a portion of the video stream corresponding to the detected music data, and may determine substitute music data based on the analysis result.

[0076] Specifically, the control unit 180 can identify at least one object by performing image recognition on a portion of the video stream, and can determine the replacement music data based on the characteristics of the identified object. At this time, the characteristics of the object may include at least one of the number of objects, the label of each object, and the moving speed of the object.

[0077] Furthermore, the control unit 180 may determine the substitute music data by analyzing the color of each section of the above-mentioned part and the degree of color change.

[0078] Furthermore, after performing the filtering step (S302), the control unit 180 may output the filtered audio stream (S303).

[0079] The data processing device 100 according to the present invention can output the video content including the filtered audio stream in the form of a file stored in a memory, or can directly output the video content to a display. Alternatively, the data processing device 100 can send the filtered audio stream to an external server or an external terminal.

[0080] For example, the data processing device 100 according to the present invention can be installed on a server of a live video platform. In this case, when a user uploads video content to the relevant platform, the data processing device 100 performs a filtering step (S302) on the uploaded video content, and then transmits the filtering result to the platform control device, so that the filtering result is output on the platform.

[0081] In another example, the control unit 180 may control the output unit 120 to delete the music data detected from the original audio stream to output the video content including the changed audio stream. In addition, the control unit 180 may output information related to the segment where the music data is deleted from the original audio stream together with the changed video content.

[0082] For example, a text file separate from the changed video content file may be output. In another example, the control unit 180 may output information about the deleted segment of the music data by using a log provided by the video platform, and control the output of the changed video content on the above platform.

[0083] In another example, the control unit 180 may control the output unit 120 to parse the original video content based on the segment where the detected music data is located, and output it as a plurality of video contents.

[0084] According to the present invention, music data contained in video content can be detected even if the user does not directly scan the video content, so the convenience of the user who edits the video content can be improved.

[0085] In addition, since music data can be detected for a large amount of video content in a short time, the cost of video editing can be significantly reduced.

[0086] In addition, according to the present invention, since the data processing apparatus deletes music data corresponding to a copyrighted work included in input video content or replaces it with substitute music, the convenience of the owner or distributor of the video content can be improved.

[0087] Industrial Applicability

[0088] According to the present invention, music data contained in video content can be detected even if the user does not directly scan the video content, so the convenience of the user who edits the video content can be improved.

[0089] In addition, since music data can be detected for a large amount of video content in a short time, the cost of video editing can be significantly reduced.

[0090] In addition, according to the present invention, since the data processing apparatus deletes music data corresponding to a copyrighted work included in input video content or replaces it with substitute music, the convenience of the owner or distributor of the video content can be improved.

Claims

1. A data processing method, characterized in that: The steps include: receiving an input of video content including a video stream and an audio stream; Detecting music data from the above audio stream; and filtering the audio stream to delete the music data detected in the audio stream, The steps of detecting music data from the above audio stream include: A division process for dividing the audio stream into music data and voice data; as well as a detection process for detecting a segment containing music data from the above audio stream, The above division process is performed by a pre-trained first artificial intelligence model. The first artificial intelligence model is configured to learn using training data labeled as music or speech, Furthermore, the first artificial intelligence model is configured to output a probability that each preset unit segment of the audio stream corresponds to music data and a probability that each preset unit segment of the audio stream corresponds to voice data based on the learning result. The above detection process is performed by a pre-trained second artificial intelligence model. The above-mentioned second artificial intelligence model is configured to learn using training data recorded as data including a proportion of music greater than or equal to a reference value, data including a proportion of music less than or equal to a reference value, and data not including music.

2. The data processing method according to claim 1, characterized in that: The step of filtering the audio stream includes determining whether the detected music data is a copyrighted work based on the copyright information of the detected music data and filtering the audio stream according to whether the detected music data is a copyrighted work.

3. The data processing method according to claim 1, characterized in that: The method further includes the step of changing the detected music data into substitute music data different from the music data.

Citation Information

Patent Citations

  • Digital recorder for selectively storing only a music section out of radio broadcasting contents and method thereof

    CN1633690A

  • Content processor and distribution system

    JP2005071090A