Method for detecting music plagiarism, electronic device, and recording medium
By converting music data into quantized form and analyzing instrument-specific sound sources and musical structure, the method addresses the limitations of existing plagiarism detection technologies, enabling effective plagiarism checks on music data without score data.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- MIPPIA INC
- Filing Date
- 2024-12-06
- Publication Date
- 2026-05-07
AI Technical Summary
Existing music plagiarism detection technologies are limited by the need for sheet music data and struggle to effectively identify plagiarism in music data lacking score data, especially when using deep learning, which is restricted to laboratory-created data and fails to handle actual music data without sheet music.
The method converts music data into quantized data, comprising notes and musical score data, and analyzes elements like melody, chords, and rhythms to identify similarities, using models to separate instrument-specific sound sources, tempo, and musical structure, enabling plagiarism checks even on music data without score data.
Enables comprehensive plagiarism detection across various forms of music data, including those without score data, by converting audio into human-readable quantized data and analyzing musical patterns, providing detailed comparison results.
Smart Images

Figure KR2024019953_07052026_PF_FP_ABST
Abstract
Description
Music plagiarism detection method, electronic device and recording media
[0001] The present disclosure relates to a method for checking music plagiarism, an electronic device, and a recording medium.
[0002] Music plagiarism refers to the unauthorized use of another person's musical work, either in whole or in part. It can manifest in various forms, ranging from cases where the sound is merely similar to cases where the mood is alike or where the creator's ideas contained within the music are stolen. Such instances of music plagiarism occur frequently worldwide, leading to growing interest in technologies to resolve or prevent them.
[0003] The present disclosure aims to convert music data into quantized data composed of notes, etc., that can be understood by humans.
[0004] The present disclosure aims to convert music data into quantized data and store quantized data or musical score data corresponding to numerous music data in a database.
[0005] The present disclosure aims to analyze elements such as melody, chords, and rhythms of music data stored in a database to identify music data similar to music data entered by a user, and to provide information related thereto to the user.
[0006] The present disclosure aims to provide music structure information, harmony, and melody information to a user by analyzing music data.
[0007] The present disclosure aims to provide information on the results of comparison with second music data for each music segment of first music data.
[0008] The present disclosure aims to identify in the second music data a music segment whose similarity to the music segment included in the first music data is above a certain standard, and to visually provide comparison result information by comparing the musical pattern of the first music data and the musical pattern of the second music data.
[0009] The problems that this disclosure aims to solve are not limited to those described above, and other unmentioned problems will be clearly understood by a person skilled in the art from the description below.
[0010] A method for performing a music plagiarism check according to one embodiment of the present disclosure may include: receiving music data; obtaining sound source data for each instrument based on the music data; obtaining at least one of the tempo, beat, or strong beat of the sound source data for each instrument based on the sound source data for each instrument; obtaining melody data based on the sound source data for each instrument; obtaining music structure information based on at least one of the sound source data for each instrument, the tempo, beat, or strong beat; and obtaining quantized data corresponding to the music data based on the music structure information and the melody data.
[0011] A method according to one embodiment of the present disclosure may further include the steps of: obtaining metadata included in music data based on music data; normalizing the metadata according to a predetermined rule; and storing the normalized metadata and quantized data.
[0012] The sound source data for each instrument according to one embodiment of the present disclosure includes at least one of a first sound source data corresponding to vocals, a second sound source data corresponding to bass, a third sound source data corresponding to drums, or a fourth sound source data corresponding to piano, and the step of acquiring the sound source data for each instrument may include the step of acquiring the sound source data for each instrument by applying music data to a first model trained to separate the sound source data for each instrument.
[0013] The step of obtaining at least one of the tempo, beat, or strong beat of the instrument-specific sound source data according to one embodiment of the present disclosure may include: the step of obtaining the beat and strong beat by applying the instrument-specific sound source data to a second model trained to output the beat and strong beat from the instrument-specific sound source data; and the step of obtaining the tempo based on the beat.
[0014] The step of obtaining music structure information according to one embodiment of the present disclosure may include applying at least one of instrument-specific sound source data, tempo, beat, or strong beat to a third model to obtain music structure information including a musical segment and a pattern in which musical features are repeated in the musical segment.
[0015] A step of acquiring melody data according to one embodiment of the present disclosure may include: acquiring first melody data corresponding to a vocal melody based on first sound source data; acquiring second melody data corresponding to a bass melody based on second sound source data; acquiring third melody data corresponding to a drum melody based on third sound source data; acquiring fourth melody data corresponding to a piano melody based on fourth sound source data; and acquiring chord data based on music data.
[0016] A step of acquiring quantized data according to one embodiment of the present disclosure may include: a step of determining the starting point of a musical section in music data based on at least one of tempo, beat, or beat; a step of determining the theme of the musical section; a step of acquiring sub-melody data corresponding to the theme in melody data; and a step of acquiring quantized data based on notes included in the sub-melody data and position data of notes in the musical section.
[0017] A method according to one embodiment of the present disclosure may further include the steps of: receiving first music data to be subject to plagiarism check; determining second music data which is one of a plurality of stored music data; obtaining a similarity between the first music data and the second music data; adding to a list of songs similar to the first music data based on a determination that the similarity is greater than or equal to a certain standard; and determining third music data which is one of a plurality of music data based on a determination that the similarity is less than or equal to a certain standard.
[0018] A step of obtaining similarity according to one embodiment of the present disclosure may include: obtaining at least one of melody similarity or overall song similarity between first music data and second music data; and obtaining similarity by assigning a weight to at least one of melody similarity or overall song similarity.
[0019] A method for providing music plagiarism check result information according to one embodiment of the present disclosure may include: receiving first music data to be subject to music plagiarism check; identifying second music data having a similarity to the first music data greater than or equal to a certain standard; generating page information that displays the music plagiarism check result between the first music data and the second music data; and transmitting the page information to a terminal.
[0020] Page information according to one embodiment of the present disclosure may include at least one of plagiarism result information in the entire music section between the first music data and the second music data or comparison result information for each music section.
[0021] Plagiarism result information according to one embodiment of the present disclosure may be determined based on the similarity between a plurality of music segments included in first music data and a plurality of music segments included in second music data.
[0022] The comparison result information for each music segment according to one embodiment of the present disclosure may include comparison result information between a first music segment, which is one of a plurality of music segments included in the first music data, and a second music segment, which is one of a plurality of music segments included in the second music data.
[0023] The comparison result information for each music section according to one embodiment of the present disclosure may include at least one of a sound figure of a first music section, a harmony of a first music section, a sound figure of a second music section, a harmony of a second music section, a link for playing first music data, or a link for playing second music data.
[0024] The step of generating page information according to one embodiment of the present disclosure may include the step of generating page information such that in a two-dimensional graph in which the first axis is a bit and the second axis is a musical pattern, the musical pattern of the first music section is displayed in a first color and the musical pattern of the second music section is displayed in a second color.
[0025] A method according to one embodiment of the present disclosure may further include the step of generating page information such that the musical pattern of a second musical section, which is identical to the musical pattern of a first musical section in a two-dimensional graph, is displayed in a third color.
[0026] A step of generating page information according to one embodiment of the present disclosure may include: a step of generating page information such that a first object is included in a portion of a page displayed on a terminal screen to display comparison result information for all of a plurality of music segments included in a first music data; and a step of generating page information such that a second object is included in a portion of a portion of a page that displays one or more music segments included in the first music data, wherein the similarity between one of the plurality of music segments included in the first music data and one of the plurality of music segments included in the second music data is greater than or equal to a certain standard.
[0027] According to one embodiment of the present disclosure, the second music data is one music data determined by user input among a plurality of third music data having a similarity to the first music data above a certain standard, and the step of generating page information may include the step of generating page information such that a music list including a plurality of third music data is displayed in a first area of a page displayed on the screen of a terminal.
[0028] The step of generating page information according to one embodiment of the present disclosure may include the step of generating page information such that at least one of plagiarism result information in the entire music segment or comparison result information per music segment is displayed in a second area located next to a first area.
[0029] An electronic device according to another embodiment of the present disclosure comprises a memory; and one or more processors, wherein the one or more processors receive first music data, acquire instrument-specific sound source data based on the music data, acquire at least one of the tempo, beat, or strong beat of the instrument-specific sound source data based on the instrument-specific sound source data, acquire melody data based on the instrument-specific sound source data, acquire music structure information based on at least one of the instrument-specific sound source data, tempo, beat, or strong beat, and acquire quantized data corresponding to the music data based on the music structure information and the melody data.
[0030] One or more processes according to an embodiment of the present disclosure identify second music data having a similarity to first music data that is greater than or equal to a certain standard, generate page information that displays a music plagiarism check result between the first music data and the second music data, transmit the page information to a terminal, and the page information may include at least one of plagiarism result information for the entire music segment between the first music data and the second music data or comparison result information for each music segment.
[0031] The present disclosure can convert music data into quantized data composed of notes, etc., that can be understood by humans.
[0032] The present disclosure converts music data into quantized data and can store quantized data or musical score data corresponding to numerous music data in a database.
[0033] The present disclosure can identify music data similar to music data entered by a user by analyzing elements such as melody, chords, and rhythms of music data stored in a database.
[0034] The present disclosure can analyze music data and provide music structure information, harmony, and melody information to the user.
[0035] The present disclosure can identify and provide to a user second music data having a similarity to first music data above a certain standard.
[0036] The present disclosure may provide plagiarism result information in the entire music section between the first music data and the second music data, or comparison result information for each music section.
[0037] The present disclosure may provide information on the result of comparison with second music data for each music section of first music data.
[0038] The present disclosure identifies a music segment in the second music data that has a similarity to a music segment included in the first music data above a certain standard, and can visually provide comparison result information by comparing the musical pattern of the first music data and the musical pattern of the second music data.
[0039] The effects according to the present disclosure are not limited to those described above, and other unmentioned effects will be clearly understood by a person skilled in the art from the description below.
[0040] FIG. 1 is a drawing illustrating a system for performing music plagiarism checks in one embodiment of the present disclosure.
[0041] FIG. 2 is a diagram illustrating a method for obtaining music structure information in one embodiment of the present disclosure.
[0042] FIG. 3 is a diagram illustrating a method for obtaining quantized data in one embodiment of the present disclosure.
[0043] FIG. 4 is a diagram illustrating a method for identifying music data similar to a first music data in one embodiment of the present disclosure.
[0044] FIG. 5 is a diagram illustrating a method for calculating melody similarity in one embodiment of the present disclosure.
[0045] FIG. 6 is a drawing for explaining an electronic device according to one embodiment of the present disclosure.
[0046] FIG. 7 is a flowchart illustrating a method for an electronic device to perform music plagiarism checks according to one embodiment of the present disclosure.
[0047] FIG. 8 is a drawing illustrating a page displaying plagiarism result information in an entire music section in one embodiment of the present disclosure.
[0048] FIG. 9 is a drawing illustrating a page in which comparison result information between a first music section and a second music section is displayed in one embodiment of the present disclosure.
[0049] FIG. 10 is a drawing illustrating a page in which comparison result information between a first music section and a second music section is displayed in another embodiment of the present disclosure.
[0050] FIG. 11 is a drawing illustrating a window for uploading first music data in one embodiment of the present disclosure.
[0051] FIG. 12 is a flowchart illustrating a method for an electronic device to provide a music plagiarism check result to a user, according to one embodiment of the present disclosure.
[0052] Hereinafter, exemplary embodiments according to the present invention will be described in detail with reference to the contents described in the attached drawings. However, the present invention is not limited or restricted by exemplary embodiments. Unless otherwise defined, all terms used in this specification (including technical and scientific terms) shall be used in a meaning that is commonly understood by those skilled in the art to which this disclosure belongs, but this may vary depending on the intent of those skilled in the art, case law, the emergence of new technology, etc.
[0053] Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise. In certain cases, terms have been selected at the applicant's discretion, and in such cases, their meanings will be described in detail in the relevant explanatory sections. Accordingly, terms used in this disclosure should be defined not merely by their names, but based on their meanings and the content throughout this disclosure.
[0054] Throughout this specification, when a part is described as "comprising" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Furthermore, the singular form used in this specification includes the plural form unless specifically stated otherwise. Additionally, the expression "at least one of a, b, and / or c" as used throughout this specification may encompass 'a alone', 'b alone', 'c alone', 'a and b', 'a and c', 'b and c', or 'a, b, and c all'.
[0055] Meanwhile, terms such as "first and / or second" used in this specification may be used to describe various components, but they are used solely for the purpose of distinguishing one component from another and are not intended to limit the scope to the components referred to by such terms. For example, without departing from the scope of the present invention, the first component may be named the second component, and the second component may also be named the first component.
[0056] Additionally, terms such as “…part,” “…module,” etc., as described in this specification refer to a unit that processes at least one function or operation, which may be implemented in hardware or software, or a combination of hardware and software. Furthermore, embodiments of this disclosure may be represented in this specification by functional block configurations and various processing steps. These functional blocks may be implemented by various numbers of hardware and / or software configurations that execute specific functions. For example, embodiments of this disclosure may employ integrated circuit configurations such as memory, processing, logic, look-up tables, etc., which can execute various functions under the control of one or more microprocessors or other control devices.
[0057] In an embodiment according to the present disclosure, functions related to artificial intelligence may be implemented through a processor and memory. In this case, the processor may be any one of a general-purpose processor such as a CPU (Center Processing Unit), AP (Application Processor), DSP (Digital Signal Processor), a graphics-dedicated processor such as a GPU (Graphic Processing Unit) or VPU (Vision Processing Unit), and an artificial intelligence-dedicated processor such as an NPU (Neural Network Processing Unit). The processor may process input data according to predefined operation rules or artificial intelligence models stored in memory. Alternatively, if the processor is an artificial intelligence-dedicated processor, the artificial intelligence-dedicated processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model. In some embodiments according to the present disclosure, functions related to artificial intelligence may be implemented through a plurality of processors.
[0058] In an embodiment according to the present disclosure, a predefined operation rule or artificial intelligence model may be configured to perform machine learning. Here, being configured to perform machine learning means that the predefined operation rule or artificial intelligence model is configured to perform a desired characteristic (or objective) by learning using a plurality of training data based on a learning algorithm. Such learning may be performed on the device itself in which the artificial intelligence according to the present disclosure is implemented, or it may be performed through a separate server and / or system.
[0059] Artificial intelligence models can be implemented as neural networks (or artificial neural networks) and can operate based on statistical learning algorithms that mimic biological neurons in machine learning and cognitive science. A neural network can refer to a model in which artificial neurons (nodes), which form a network through synaptic connections, change the strength of synaptic connections through learning to possess problem-solving capabilities. A neural network can be composed of multiple neural network layers; for example, a neural network may include an input layer, a hidden layer, and an output layer. Each of the multiple neural network layers may include at least one node and at least one weight, and neural network operations can be performed through operations between the results of previous (precious) layers and the weights. At least one weight possessed by the multiple neural network layers may be optimized based on the learning results of the artificial intelligence model. For example, at least one weight may be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. Neural networks can infer a result to be predicted from an arbitrary input.
[0060] The learning methods of artificial intelligence models can be classified according to the learning approach into supervised learning, where input and output data are provided as training data and the correct answer (output data) corresponding to the problem (input data) is predetermined; unsupervised learning, where only input data is provided without output data and the correct answer (output data) corresponding to the problem (input data) is not predetermined; and reinforcement learning, where a reward is granted whenever an action is taken from the current state and learning proceeds in a direction that maximizes this reward. Alternatively, they can be classified according to the architecture, which is the structure of the learning model.
[0061] In the embodiments of the present disclosure, the artificial intelligence model is a Convolutional Neural Network (CNN) such as GoogleNet, AlexNet, VGG Network, Region with Convolutional Neural Network (R-CNN), Region Proposal Network (RPN), Recurrent Neural Network (RNN), Stacking-based Deep Neural Network (S-DNN), State-Space Dynamic Neural Network (S-SDNN), Deconvolution Network, Deep Belief Network (DBN), Restructured Boltzmann Machine (RBM), Fully Convolutional Network, Long Short-Term Memory Network (LSTM), Classification Network, Generative Modeling, eXplainable AI, Continual AI, Representation Learning, AI for Material Design, BERT, SP-BERT, MRC / QA for Natural Language Processing, Text Analysis, Dialog System, GPT-3, GPT-4, Visual Analytics, Visual Understanding, Video Synthesis for Vision Processing, Anomaly Detection, Prediction, Time-Series Forecasting, Optimization, Recommendation for ResNet Data Intelligence, At least one of various artificial intelligence structures and algorithms, such as data creation, may be used. The examples described above are merely examples of artificial intelligence structures and algorithms used according to the embodiments of the present disclosure and do not limit the artificial intelligence structures and algorithms used according to the embodiments of the present disclosure.
[0062] Hereinafter, various embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In describing the embodiments, technical details that are well known in the art to which the present invention pertains and are not directly related to the present invention will be omitted. This is to ensure that the essence of the present invention is conveyed more clearly without obscuring it by omitting unnecessary explanations. For the same reason, some components in the accompanying drawings may be exaggerated, omitted, or schematically depicted. Furthermore, the size of each component does not entirely reflect its actual size. Throughout this specification, the same reference numerals may refer to the same or corresponding components.
[0063] FIG. 1 is a drawing illustrating a system for performing music plagiarism checks in one embodiment of the present disclosure.
[0064] As the quality of music generated by artificial intelligence improves dramatically, there is a growing trend of users engaging in commercial activities using AI-generated music. Issues related to music plagiarism include not only cases where an artist intentionally plagiarizes another's work, but also instances where AI unintentionally generates music containing ideas identical to those of others. Therefore, as plagiarism can occur in various forms, professionals in the music industry may seek to prevent the risk of plagiarism in advance by performing plagiarism checks on the music they compose.
[0065] Therefore, as the forms of music plagiarism diversify, interest in technologies for detecting music plagiarism is increasing. For example, this goal can be achieved by utilizing deep learning technology. However, even when deep learning technology is used, the input data is limited to laboratory data intentionally created by experimenters to resemble the target music data, rather than actual music data in circulation. Furthermore, there is a technical limitation in that deep learning can only determine whether plagiarism has occurred under restricted circumstances where sheet music data, rather than actual music data, is input. Due to these technical limitations, music plagiarism checks cannot be performed using deep learning for music data for which sheet music data is unavailable.
[0066] To overcome these technical limitations, the present disclosure obtains quantized data from music data that lacks score data, obtains score data, and then performs music plagiarism checks using the quantized data and / or score data. By doing so, the present disclosure can perform music plagiarism checks extensively even on music data that lacks score data. Quantized data is data in a human-readable form and is generated based on music data. Music data may be in the form of an audio signal or a spectrogram. A spectrogram is two-dimensional data that visually represents frequency information over time. Quantized data may be expressed in forms such as, for example, a musical note (e.g., a sixteenth note) or the position of a note within a musical section (e.g., four measures). Score data is a format that visually records music, expressing musical elements such as notes, rests, beats, intervals, and rhythms using symbols and symbol systems. Score data may be in a digital form and may be, for example, in MusicXML, MIDI, ABC notation, or JSON format.
[0067] In one embodiment, the server (110) may be a device that performs music plagiarism checks. The server (110) may be a device that provides music plagiarism check services through a website and / or application. A user terminal (130) may access the website and / or application to upload music data and receive a music plagiarism check for the said music data. The user terminal (130) may be a terminal of a user who uses the music plagiarism check service. The server (110) may provide music plagiarism check results for the uploaded music data. More specifically, the server (110) according to an embodiment of the present disclosure may identify other music data from a database that may raise the possibility of plagiarism regarding the music data subject to plagiarism check, and provide plagiarism result information between the music data subject to plagiarism check and the music data identified from the database. The plagiarism result information may include analysis results regarding how similar the first music data subject to plagiarism check and the second music data identified as similar to the first music data are.
[0068] The server (110) can receive music data and obtain quantized data and / or musical score data of the music data. The server (110) can store the quantized data and / or musical score data in a database (150). The database (150) may be a component included in the server (110) or a separate device independent of the server (110).
[0069] A server (110), a database (150), and a user terminal (130) can be connected through a network (20). The network (20) refers to a connection structure that enables information exchange between each node, such as devices, terminals, and servers. Examples of such a network (20) include, but are not limited to, a 3GPP (3rd Generation Partnership Project) network, an LTE (Long Term Evolution) network, a 5G network, a WIMAX (World Interoperability for Microwave Access) network, the Internet, a LAN (Local Area Network), a Wireless LAN (Wireless Local Area Network), a WAN (Wide Area Network), a PAN (Personal Area Network), a Wi-Fi network, a Bluetooth network, a satellite broadcasting network, an analog broadcasting network, and a DMB (Digital Multimedia Broadcasting) network.
[0070] In one embodiment, the server (110) may receive music data. The server (110) may receive music data through web crawling or may receive it from a user terminal (130).
[0071] In one embodiment, the server (110) may obtain sound source data for each instrument based on music data. The sound source data for each instrument may be the sound produced by each instrument separated from music data composed of the performance of various instruments. For example, if the music data consists of vocals, drums, and bass, the sound source data for each instrument may include vocal sound source data, drum sound source data, and bass sound source data.
[0072] In one embodiment, the server (110) may obtain at least one of the tempo, beat, or strong beat of the instrument-specific sound source data based on the instrument-specific sound source data. The tempo may indicate the speed at which the music is played. The tempo may be measured in units of Beats Per Minute (BPM). The beat defines how the beats are divided and grouped. The beat is indicated at the beginning of the score and is expressed in fractional form. The strong beat indicates the first beat of a measure and may be the most strongly emphasized beat. The strong beat corresponds to the beginning of a measure, and the listener may feel the flow of the music through the strong beat. For example, if the instrument-specific sound source data includes vocal sound source data, drum sound source data, and bass sound source data, the server (110) may obtain the tempo, beat, and strong beat of the vocal sound source data, obtain the tempo, beat, and strong beat of the drum sound source data, and obtain the tempo, beat, and strong beat of the bass sound source data.
[0073] In one embodiment, the server (110) can acquire melody data based on sound source data for each instrument. Melody data is a representation of elements constituting a melody in music in a form that can be digitized or analyzed. A melody refers to the core melody of a song and is a flow of notes played continuously over time. Melody data is data that organizes information on the flow of notes and may include information related to pitch, note length, rhythm pattern, range, scale, and tonality.
[0074] In one embodiment, the server (110) may obtain music structure information based on at least one of instrument-specific sound source data, tempo, beat, or beat. The music structure information may describe how various parts constituting the entire piece of music are arranged and connected. The music structure serves as a kind of blueprint for how the piece begins, develops, and ends, and defines the form and flow of the piece. For example, the music structure information may include an introduction, verse, chorus, bridge, interlude, outro, etc. For example, the music structure information may include musical sections (e.g., measures) in the music data and patterns in which musical features are repeated within the musical sections.
[0075] In one embodiment, the server (110) can obtain quantized data corresponding to music data based on music structure information and melody data. The server (110) can divide the music data into multiple musical segments using the music structure information. The server (110) can then calculate the theme of each musical segment and match it with melody data corresponding to that theme. Since the melody data includes notes and note lengths in specific musical segments, the server (110) can place notes in specific musical segments using the music structure information and melody data. Through this, the server (110) can obtain quantized data composed of notes that are understandable to humans.
[0076] In one embodiment, the server (110) can obtain metadata included in the music data based on the music data. Metadata is data that describes various information related to the music data, and is data that includes additional information related to the music data rather than the audio signal itself. For example, metadata may include title, artist, album, track number, release year, composer, lyricist, producer, genre, file format, sampling rate, copyright information, etc. The server (110) can verify the metadata to determine whether the information related to the music data is valid data. If it is determined that specific information is invalid, the server (110) can modify the information. If it is determined that the metadata is valid as a result of verification, the server (110) can normalize the metadata according to a predetermined rule. Normalization means converting the format of the metadata according to a rule set by the server (110). For example, the artist in music data entered in Korean and the artist in music data entered in English may be the same. The server (110) may determine, by a predetermined rule, that the artist name should be expressed in Korean, and accordingly, perform normalization to convert the artist name entered in English into Korean. The server (110) may store the normalized metadata in the database (150). Upon determining that the storage of the metadata in the database (150) has failed, the server (110) may attempt to store the metadata again.
[0077] In one embodiment, the server (110) can cluster quantized data (or melody data, sheet music data) stored in the database (150) to cluster quantized data with high sound source similarity. Through this, when performing music plagiarism checks, it is possible to quickly determine which cluster to search intensively based on the similarity between the quantized data included in a specific cluster and the music data subject to plagiarism check.
[0078] In one embodiment, the server (110) may be a device that performs a music plagiarism check and provides plagiarism result information for the entire music segment and / or comparison result information for each music segment to a user terminal (130) (hereinafter referred to as 'terminal (130)'). The plagiarism result information for the entire music segment may include the possibility (e.g., plagiarism rate) that the first music data overall plagiarizes the second music data by comparing all music segments included in the first music data subject to plagiarism check with all music segments included in the second music data identified in the database (150). The comparison result information for each music segment may include comparison result information between a first music segment, which is one of a plurality of music segments included in the first music data, and a second music segment, which is one of a plurality of music segments included in the second music data. The comparison result information for each music segment may include the possibility that one music segment included in the first music data plagiarizes one music segment included in the second music data.
[0079] The server (110) may be a device that provides a music plagiarism check service through a website and / or application. A terminal (130) may access the website and / or application to upload music data and receive a music plagiarism check for the music data. The terminal (130) may be a terminal of a user using the music plagiarism check service. The server (110) may provide a music plagiarism check result for the uploaded music data.
[0080] In one embodiment, the server (110) may transmit page information to the terminal. The terminal (130) may receive the page information and display the results of a music plagiarism check. For example, the terminal (130) may display at least one of the plagiarism result information for the entire music segment between the first music data and the second music data included in the page information, or the comparison result information for each music segment.
[0081] FIG. 2 is a diagram illustrating a method for obtaining music structure information in one embodiment of the present disclosure.
[0082] In one embodiment, the server (110) may obtain sound source data for each instrument by applying music data (210) to a first model (220) trained to separate sound source data for each instrument. The first model (220) may be trained to separate sound source data for each instrument from the music data (210).
[0083] In one embodiment, the sound source data for each instrument may include at least one of a first sound source data (231) corresponding to vocals, a second sound source data (232) corresponding to bass, a third sound source data (233) corresponding to drums, or a fourth sound source data (234) corresponding to piano. The sound source data for each instrument may further include a fifth sound source data (235) corresponding to other instruments. Other instruments may be instruments other than vocals, bass, drums, and piano.
[0084] For example, the first model (220) may be trained based on training data in which the input is music data (210) and the outputs are first sound source data (231), second sound source data (232), third sound source data (233), fourth sound source data (234), and fifth sound source data (235). The first model (220) may include a Hybrid Transformer. A Hybrid Transformer is a model that combines a U-net and a Transformer. A U-net can recognize local patterns in the frequency domain of music data and extract frequency features. The Transformer can separate sounds (sound sources) by instrument by utilizing the extracted frequency features and time-series modeling capabilities. However, the types of the first model (220) are not limited to the examples described above.
[0085] In one embodiment, the server (110) can obtain beats and strong beats by inputting the instrument-specific sound source data to a second model (240) trained to output beats and strong beats from the instrument-specific sound source data. The server (110) can obtain a tempo based on the beats. For example, the tempo can be determined based on 60 / (time difference of two beats). For example, the server (110) can obtain a first beat (252) and a first strong beat (253) from the second model (240). And the server (110) can obtain a first tempo (251) from the first beat (252). The second model (240) may be a strong beat extraction model including a transformer. When the strong beat extraction model receives new music data, it calculates the probability of a strong beat occurring for each musical segment. A point in time with a high probability is predicted as a beat, and based on this, the rhythm and structure of the music can be analyzed. The server (110) can train a second model (240) using a label that includes the exact time when the beat appears. Additionally, the second model (240) can be trained to output beats from instrument-specific sound source data by utilizing training data that labels beats.
[0086] In another example, the second model (240) can output the tempo, beat, and strong beat from the instrument-specific sound source data. For example, the server (110) can obtain the second tempo (254), the second beat (255), and the second strong beat (256) by applying the second sound source data (232) to the second model (240).
[0087] In one embodiment, the server (110) may apply at least one of the instrument-specific sound source data (230), tempo (251, 254, 257, 260, 263), beat (252, 255, 258, 261, 264), or strong beat (253, 256, 259, 262, 265) to the third model (270) to obtain music structure information (280) including musical segments and patterns in which musical features are repeated within the musical segments. The third model (270) may include Dilated Neighborhood Attention (DNA). Dilated Neighborhood Attention is one of the techniques for effectively identifying the patterns and structures of a song in music structure analysis. This technique is a neural network structure and can be used to analyze musical context or time-series data. The main concept of extended neighborhood attention is to recognize important patterns or structures over a wider range by focusing attention on information in the "neighbors" or "nearby" through extended dilation. Through this, the server (110) can grasp the musical structure over a wider musical range.
[0088] FIG. 3 is a diagram illustrating a method for obtaining quantized data in one embodiment of the present disclosure.
[0089] In one embodiment, the server (110) can obtain first melody data (321) corresponding to a vocal melody based on first sound source data (231). The server (110) can obtain first melody data (321) by applying the first sound source data (231) to a vocal melody extraction model. The vocal melody extraction model may include an EfficientNet-based vocal transcription model. The EfficientNet-based vocal transcription model may convert audio data into a spectrogram form and then input it into the EfficientNet model to output a vocal melody.
[0090] In one embodiment, the server (110) may obtain second melody data (322) corresponding to a bass melody based on second sound source data (232). The server (110) may obtain the second melody data (322) by inputting the second sound source data (232) into a bass melody extraction model. The bass melody extraction model may include a jukebox and / or a transformer-based instrument melody transcription model. The jukebox may analyze music and generate music using deep learning technology. The transformer-based instrument melody transcription model may be a deep learning model that performs instrument-specific melody transcription. The instrument-specific melody transcription model may transcribe the melody of a specific instrument and convert elements such as pitch, duration, and timing into text or MIDI format.
[0091] MIDI (Musical Instrument Digital Interface) is one of the most common ways to represent melodic data digitally. The MIDI format includes information such as pitch, duration, start time, and tempo (dynamics), allowing it to be played back on instruments or software. MIDI can include note information, channels, and program change messages. Note information can include the start (note on) and end (note off) timings of each note, pitch, and tempo. Note information can determine which notes to play, when, and how loudly on an instrument. MIDI has 16 independent channels, each of which can be assigned to a different instrument or track. Settings such as the instrument, volume, and panning can be controlled based on each channel. Program change messages are commands that change the instrument. For example, a piano sound can be changed to a violin sound.
[0092] In one embodiment, the server (110) can obtain third melody data (323) corresponding to a drum melody based on third sound source data (233). The server (110) can obtain the third melody data (323) by inputting the third sound source data (233) into a drum melody extraction model. The drum melody extraction model may include a CNN-based drum transcription model. The CNN-based drum transcription model can analyze music data in which drums are played to determine when various components of the drum (e.g., kick drum, snare drum, hi-hat, etc.) were played and convert this into text or MIDI format. The CNN-based drum transcription model can receive a spectrogram as input. Since drum sounds generally appear strongly in a specific frequency range, the CNN can learn this effectively.
[0093] In one embodiment, the server (110) can obtain fourth melody data (324) corresponding to a piano melody based on the fourth sound source data (234). The server (110) can obtain the fourth melody data (324) by applying the fourth sound source data (234) to a piano melody extraction model. The piano melody extraction model may be a transformer-based instrument melody transcription model. The piano melody extraction model may be a model specialized in extracting piano melodies from instruments.
[0094] In one embodiment, the server (110) may obtain fifth melody data (325) based on fifth sound source data (235). The fifth melody data may be related to the melody of a guitar instrument.
[0095] The aforementioned melody data may follow the MIDI format.
[0096] In one embodiment, the server (110) can obtain chord data based on music data. The server (110) can obtain chord data by applying music data to a Harmony Transformer. The Harmony Transformer learns the relationship between notes over time using a transformer structure and can identify chords based on this.
[0097] In one embodiment, the server (110) may determine the starting point of a musical segment in the music data based on at least one of the tempo, beat, or beat. For example, the server (110) may determine the starting point of a musical segment (e.g., 4 measures) based on music structure information (280) and determine 4 measures from the starting point as one musical segment. The server (110) may determine the theme of the determined musical segment. The theme may be determined based on the musical characteristics of the determined musical segment. The server (110) may obtain sub-melody data corresponding to the theme in the melody data. The sub-melody data may be at least a part of the melody data and may correspond to the determined musical segment. The melody data may include notes and the lengths of the notes. Thus, the position data of the notes listed sequentially may be determined according to the length of each note. Through this, the server (110) can obtain quantized data (340) based on the position data of the notes in the musical section and the notes included in the sub-melody data.
[0098] FIG. 4 is a diagram illustrating a method for identifying music data similar to a first music data in one embodiment of the present disclosure.
[0099] In one embodiment, the server (110) may receive first music data (410) that is subject to plagiarism check. The first music data (410) may be input by a user. The server (110) may determine second music data (420) which is one of a plurality of music data stored in the database (150).
[0100] The server (110) can obtain a similarity (430) between the first music data (410) and the second music data (420). For example, the server (110) can obtain a similarity (430) based on the quantized data of the first music data (410) and the quantized data of the second music data (420). As another example, the server (110) may obtain a similarity between a feature vector obtained by applying the spectrogram of the first music data (410) to a CNN and a feature vector obtained by applying the spectrogram of the second music data (420) to a CNN.
[0101] Similarity (430) may be determined based on at least one of melody similarity (431) or overall song similarity (432). Melody similarity (431) may be the similarity between the melody data of the first music data (410) and the melody data of the second music data. Melody similarity (431) may include chord similarity. Chord similarity may be the similarity between the chord data of the first music data (410) and the chord data of the second music data (420). Overall song similarity (432) may be the similarity between the first music data and the second music data over the entire section. For example, overall song similarity (432) may be the similarity between the music structure information of the first music data (410) and the music structure information of the second music data (420).
[0102] In one embodiment, the server (110) may obtain at least one of the melody similarity (431) and the overall song similarity (432) between the first music data (410) and the second music data (420). The server (110) may determine the similarity (430) by assigning a weight to at least one of the melody similarity (431) and the overall song similarity (432).
[0103] In one embodiment, the server (110) can compare the similarity (430) with a threshold (440). Based on the determination that the similarity (430) is greater than or equal to the threshold, the server (110) can add the second music data (420) to a list of songs similar to the first music data (450). The server (110) can determine other music data in the cluster containing the second music data (420) as the third music data to obtain the similarity with the first music data (410). Through this, the server (110) can identify other similar music within the same cluster.
[0104] Conversely, based on the determination that the similarity (430) is below a threshold, the server (110) may determine a third music data, which is one of a plurality of music data, to identify a song similar to the first music data (460). The server (110) may determine one of the music data included in a different cluster from the second music data (420) as the third music data. By doing so, the server (110) can reduce the search time for identifying similar music data.
[0105] FIG. 5 is a diagram illustrating a method for calculating melody similarity in one embodiment of the present disclosure.
[0106] In one embodiment, the server (110) may obtain a final melody similarity (550) by applying a weight (530) to at least one of rhythm cosine similarity (510), pitch similarity (511), pitch variety (512), or piano roll similarity (513). The similarities described above are merely examples and the present disclosure is not limited thereto. In another embodiment, the server (110) may obtain a final melody similarity (550) by applying a weight to at least one of rhythm cosine similarity (510), pitch similarity (511), pitch variety (512), piano roll similarity (513), BPM similarity, or chord similarity.
[0107] Rhythm cosine similarity (510) is a method for measuring similarity between two rhythm patterns and can be used in the field of music analysis. Rhythm cosine similarity is a method that applies the concept of cosine similarity to rhythm patterns and can quantitatively evaluate how similar two rhythms are. For example, in music, a rhythm pattern can be represented as an array of beats separated by regular time intervals, and each rhythm pattern can be converted into a vector containing timing and intensity information. For example, in the case of a 16-bit pattern, whether a specific beat exists in each time interval can be expressed as 1 or 0.
[0108] Pitch similarity (511) is a concept that evaluates how similar two or more notes are or how similar the pitches used in different musical sections are. For example, pitch similarity can be calculated by converting pitch data into a vector and then calculating the similarity between two melodies or chords through methods such as cosine similarity, Euclidean distance, and DTW (Dynamic Time Warping).
[0109] Pitch variety (512) is a concept that evaluates how many different pitches are used in music. By analyzing how wide and diverse the pitches appear in a song, it measures how creative and original the song is. For example, a server (110) can evaluate pitch variety by measuring the range between the lowest and highest pitches used in a song.
[0110] Piano roll similarity (513) is a method of evaluating how similar two pieces of data are compared in the form of a piano roll. A piano roll is a two-dimensional table that visually represents time and pitch information. The horizontal axis represents time, and each time interval indicates the point in time when a note is played. The vertical axis represents pitch, and each row indicates a specific note. For example, the server (110) can acquire the piano roll itself as an image and obtain piano roll similarity based on image similarity (e.g., Mean Square Error (MSE), Structural Similarity Index Measure (SSIM)). As another example, the server (110) can convert the two piano rolls into vectors to calculate cosine similarity in order to obtain piano roll similarity.
[0111] In one embodiment, the server (110) may divide music data into a plurality of musical segments and then provide the similarity between the first music data and the second music data in each musical segment to the user terminal (130). The user of the user terminal (130) can easily determine which segment of the first music data is at risk of plagiarism of the second music data by looking at the similarity in the musical segments.
[0112] FIG. 6 is a drawing for explaining an electronic device according to one embodiment of the present disclosure.
[0113] An electronic device (600) according to one embodiment may be a server (110), a database (150), or a user terminal (130) (e.g., a mobile device, a desktop, a laptop, a personal computer, etc.). Referring to FIG. 6, an electronic device (600) according to one embodiment may include a user interface (610), a processor (630), a display (650), and a memory (670). The user interface (610), the processor (630), the display (650), and the memory (670) may be connected to each other via a communication bus (605).
[0114] A user interface (610) includes everything that enables interaction between a person and a machine. This can enable a user to manipulate and control a system, software, application, website, etc. For example, a user interface may include a graphical user interface, a text-based interface, a voice user interface, a natural user interface (e.g., gestures, touch, etc.).
[0115] The display (650) can output information determined by the processor (630).
[0116] The memory (670) can store various information generated during the processing of the processor (630) described above. In addition, the memory (670) can store various data and programs. The memory (670) may include volatile memory or non-volatile memory. The memory (670) may store various data by being equipped with a large-capacity storage medium such as a hard disk.
[0117] Additionally, the processor (630) may perform at least one method or an algorithm corresponding to at least one method described above through FIGS. 1 to 5. The processor (630) may be a data processing device implemented in hardware having a circuit having a physical structure for executing desired operations. For example, the desired operations may include code or instructions included in a program. The processor may be composed of, for example, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or a NPU (Neural Network Processing Unit). For example, the electronic device implemented in hardware may include a microprocessor, a central processing unit, a processor core, a multi-core processor, a multiprocessor, an ASIC (Application-Specific Integrated Circuit), or a FPGA (Field Programmable Gate Array).
[0118] The processor (630) can execute a program and control an electronic device. The program code executed by the processor (630) can be stored in memory.
[0119] FIG. 7 is a flowchart illustrating a method for an electronic device to perform music plagiarism checks according to one embodiment of the present disclosure.
[0120] In one embodiment, the electronic device (600) can receive music data (710).
[0121] In one embodiment, the electronic device (600) can acquire sound source data for each instrument based on music data (720).
[0122] In one embodiment, the electronic device (600) can obtain at least one of the tempo, beat, or beat of the instrument-specific sound source data based on the instrument-specific sound source data (730).
[0123] In one embodiment, the electronic device (600) can acquire melody data based on sound source data for each instrument (740).
[0124] In one embodiment, the electronic device (600) can obtain music structure information based on at least one of instrument-specific sound source data, tempo, beat, or beat (750).
[0125] In one embodiment, the electronic device (600) can acquire quantized data corresponding to music data based on music structure information and melody data (760).
[0126] In one embodiment, the electronic device (600) can store quantized data in a database (770).
[0127] In one embodiment, the electronic device (600) can obtain a similarity between the first music data subject to plagiarism check and the second music data, which is one of the music data stored in the database (780).
[0128] FIG. 8 is a drawing illustrating a page displaying plagiarism result information in an entire music section in one embodiment of the present disclosure.
[0129] In one embodiment, the page (800) may display plagiarism result information for the entire music section of the first music data. When the plagiarism check result tab (810) is selected on the page (800), the page (800) may be displayed on the screen of the terminal (130). The page (800) may include a first area (820) that displays a music list and a second area (830) that displays plagiarism result information.
[0130] The music list may include multiple third music data. The server (110) may identify multiple third music data from the database (150) that have a similarity to the first music data above a certain standard. The second music data may be one of the multiple third music data determined by user input.
[0131] For example, in the first area (820), third music data such as Miina (821), Because of you, and What I want may be displayed, and among these, the server (110) may receive user input for the music data called Miina (821). Miina (821) may be second music data.
[0132] The second area (830) may display information on the result of plagiarism between the first music data, music data A, and the second music data, music data, Miina (821). For example, the information on the result of plagiarism in the entire music section may include the total plagiarism rate (840). The total plagiarism rate (840) may indicate the probability that the first music data will plagiarize the second music data. For example, in the second area (830), the probability that the first music data, music data A, will plagiarize the second music data, Miina (821), may be displayed as 37.95%.
[0133] Information on the plagiarism results in the entire music segment can be determined based on the similarity between multiple music segments included in the first music data and multiple music segments included in the second music data. For example, the similarity may increase as the number of multiple music segments included in the first music data that have a similarity to multiple music segments included in the second music data above a certain standard increases. For example, the similarity between multiple music segments included in the first music data and multiple music segments included in the second music data may be expressed as the total plagiarism rate (840).
[0134] The second area (830) can display comparison result information for each music segment. The comparison result information for each music segment may include comparison result information between a first music segment, which is one of a plurality of music segments included in the first music data, and a second music segment, which is one of a plurality of music segments included in the second music data. For example, the server (110) can compare the first music segment with each of the plurality of music segments included in the second music data. For example, the server (110) can compare the first music segment with music segment A included in the second music data, music segment B included in the first music data, music segment C included in the first music data, etc., and can identify the music segment included in the second music data that has a similarity to the first music segment above a certain standard.
[0135] The comparison result information for each music section may be in the form of displaying a summary of the comparison result information for all music sections or in the form of displaying the comparison result information for one music section among all music sections. The comparison result information (850) for each music section displayed in the second area (830) may be in the form of displaying a summary of the comparison result information for all music sections.
[0136] The comparison result information for each music section may include at least one of a sound figure of the first music section, a harmony of the first music section, a sound figure of the second music section, a harmony of the second music section, a link to play the first music data, or a link to play the second music data.
[0137] For example, the comparison result information for each music segment may include comparison result information between music segment A (851) included in the first music data and music segment B (852) and music segment C (853) included in the second music data. The server (110) may generate page information to visualize a plurality of music segments (e.g., 4 measures) included in the first music data as a first bar and to visualize a plurality of music segments included in the second music data as a second bar. The first bar may be divided according to the number of music segments included in the first music data and the second bar may be divided according to the number of music segments included in the second music data. The server (110) may display a specific color (e.g., blue) for music segments in the divided first bar where the similarity is above a certain standard. As another example, the server (110) may determine that the color displayed in the first bar is displayed darker as the similarity is higher. The server (110) may display a color different from the color displayed in the divided first bar (e.g., green) for music segments in the divided second bar that have a similarity level above a certain standard. Through this, the user can easily identify which segment of the first music data within the entire section is likely to be plagiarized, and can also easily identify music segments included in the second music data that are similar to the segments likely to be plagiarized. For example, music segment A (851) in the first bar may be displayed in blue, and music segment B (852) and music segment C (853) in the second bar may be displayed in green.
[0138] The server (110) can display lines connecting two music segments that are being compared. For example, the server (110) can display lines connecting music segment A (851) and music segment B (852) and lines connecting music segment A (851) and music segment C (853). Through this, the user can easily identify which music segment of the second music data the first music segment of the first music data is similar to.
[0139] In one embodiment, upon receiving user input for a detail view object (860), the page (800) may be switched to a page (900) that displays comparison result information between the first music segment and the second music segment of FIG. 9. Alternatively, the page (800) may be switched to a page (900) that displays comparison result information when user input for at least one music segment (at least one of 851, 852 and 853) is received.
[0140] FIG. 9 is a drawing illustrating a page in which comparison result information between a first music section and a second music section is displayed in one embodiment of the present disclosure.
[0141] In one embodiment, the server (110) may generate page information to include a first object (912) for displaying comparison result information for all of the plurality of music segments included in the first music data in a portion of the page displayed on the screen of the terminal (130). For example, the portion of the page may be a third area (910) and may be displayed in the left area of the page (900). The third area (910) may display the plurality of music segments included in the first music data. The third area (910) may display comparison result information for each of the plurality of music segments. For example, the comparison result information for each of the plurality of music segments may include the similarity (e.g., plagiarism rate) between a specific music segment included in the first music data and a music segment included in the second music data. For example, the comparison result information for bar 8-11 (913), which is a music segment included in the first music data, may include that the similarity between bar 8-11 (913) and the music segment included in the second music data is 43.61%. As another example, the comparison result information of the 12-15 measure (914), which is a music section included in the first music data, may include that the similarity between the 12-15 measure (914) and the music section included in the second music data is 62.99%.
[0142] The first object (911) may be such that all music segments included in the first music data are displayed in the third area (910).
[0143] In one embodiment, the server (110) may generate page information to include a second object (912) for displaying one or more music segments included in the first music data that have a similarity level greater than a certain standard with one of the multiple music segments included in the second music data among the multiple music segments included in the first music data in a portion of the area. The server (110) may use the second object (912) to provide an option to display only the music segments of the first music data that have a similarity level greater than a certain standard through the third area (910). For example, depending on the determination that the certain standard is a similarity level of 1%, only one or more music segments included in the first music data that have a similarity level of 1% or more may be displayed in the third area (910). One or more music segments included in the first music data that have a similarity level of less than 1% may not be displayed in the third area (910).
[0144] A fourth area (930) may be displayed next to the third area (910). For example, the fourth area (930) may be displayed in the right area of the third area (910). The fourth area (930) may display comparison result information for one music segment included in the first music data determined by user input in the third area (910).
[0145] For example, the fourth region (930) may include a two-dimensional graph (940) in which the first axis is the beat (or measure, musical section) and the second axis is the pitch. For example, the two-dimensional graph (940) may have the x-axis as the beat and the y-axis as the pitch, and the pitch according to the beat may be displayed in the form of bars. For example, in the two-dimensional graph (940) for musical sections measures 8-11 (913), the x-axis may represent the interval between measures 8-11.
[0146] In one embodiment, the server (110) may generate page information such that the musical notes of the first music section in the two-dimensional graph (940) are displayed in a first color and the musical notes of the second music section are displayed in a second color. For example, the server (110) may decide to display the musical notes of the first music section, displayed in bar form in the two-dimensional graph (940), in a first color (e.g., blue), and the musical notes of the second music section, displayed in bar form, in a second color (e.g., green). For example, in the two-dimensional graph (940), the musical notes of measures 8-11, which is the first music section, may be displayed in a first color and the musical notes of measures 24-27, which is the second music section, may be displayed in a second color.
[0147] In one embodiment, the server (110) may generate page information such that the musical pattern of a second music section that is identical to the musical pattern of a first music section in a two-dimensional graph is displayed in a third color. For example, based on the determination that the pitch between 7.5 beats and 8 beats in measures 8-11 of the first music section in the two-dimensional graph (940) is identical to the pitch between 7.5 beats and 8 beats in measures 24-27 of the second music section, the server (110) may generate page information such that the bars between 7.5 beats and 8 beats are displayed in a third color (e.g., purple). This allows the user to identify at a glance which sections have identical musical patterns.
[0148] In one embodiment, the page (900) may include a graph view option. The graph view option may be a selection of bars to be displayed in the two-dimensional graph (940). For example, the graph view option may include a full view option (951), a view only of the first music section pattern option (952), a view only of the second music section pattern option (953), and / or a view of the same section option (954). The full view option (951) may be an option to display both the first music section pattern and the second music section pattern in the two-dimensional graph (940). The view of the same section option (954) may be an option to display only the second music section pattern that is identical to the first music section pattern in the two-dimensional graph (940). The more the second music section pattern that is identical to the first music section pattern, the greater the similarity between the first music section and the second music section may be.
[0149] FIG. 10 is a drawing illustrating a page in which comparison result information between a first music section and a second music section is displayed in another embodiment of the present disclosure.
[0150] The comparison result information of the first music section and the second music section can be displayed on the page (800) in a configuration different from the page.
[0151] The page (1000) may include a two-dimensional graph (940), a first music section (1010), a second music section (1020), a harmony of the first music section (1030), a harmony of the second music section (1040), a link (1050) to the first music data and / or a link (1060) to the second music data. For example, the first music section (1010) may be 10-13 measures and the second music section (1020) may be 33-36 measures. That is, the page (1000) may display information on the comparison result between the 10-13 measures of the first music data and the 33-36 measures of the second music data. By displaying the harmony of the first music section (1030) and the harmony of the second music section (1040), the user can check how similar the harmonies are.
[0152] FIG. 11 is a drawing illustrating a window for uploading first music data in one embodiment of the present disclosure.
[0153] Through the first music data upload window (1100), the user can transmit music data subject to plagiarism check to the server (110). The first music data upload window (1100) may include an area (1110) for receiving a file and an area (1120) for receiving a music data link. The user can upload the first music data by moving the music data to the first music data upload window (1100) by dragging and dropping, or by entering a music data link. When the first music data is uploaded, the page (800) of FIG. 8 may be displayed.
[0154] FIG. 12 is a flowchart illustrating a method for an electronic device to provide a music plagiarism check result to a user, according to one embodiment of the present disclosure.
[0155] In one embodiment, the electronic device can receive first music data to be subject to music plagiarism check (1210).
[0156] In one embodiment, the electronic device can identify second music data that has a similarity to first music data above a certain standard (1220).
[0157] In one embodiment, the electronic device can generate page information that displays the results of a music plagiarism check between the first music data and the second music data (1230).
[0158] In one embodiment, the electronic device can transmit page information to a terminal (1240).
[0159] Meanwhile, the embodiments disclosed in this specification may be implemented in the form of a recording medium that stores instructions executable by a computer. The instructions may be stored in the form of program code and, when executed by a processor, may generate a program module to perform the operations of the disclosed embodiments. The recording medium may be implemented as a computer-readable recording medium. A computer-readable recording medium may include all types of recording media that store instructions decipherable by a computer. Examples include ROM, RAM, magnetic tape, magnetic disk, flash memory, optical data storage devices, etc.
[0160] The above descriptions are specific embodiments for carrying out the present disclosure. The present disclosure will include not only the embodiments described above, but also embodiments that can be simply modified or easily modified. Furthermore, the present disclosure will include technologies that can be easily modified and implemented using the embodiments described above. Accordingly, the scope of the present disclosure should not be limited to the embodiments described above, but should be defined by the claims set forth below as well as equivalents to the claims of the present disclosure.
Claims
1. Regarding the method of performing music plagiarism checks, Step of receiving music data; A step of acquiring sound source data for each instrument based on the above music data; A step of obtaining at least one of the tempo, beat, or strong beat of the sound source data for each instrument based on the sound source data for each instrument; A step of acquiring melody data based on the sound source data for each instrument above; A step of obtaining music structure information based on at least one of the above-mentioned instrument-specific sound source data, the above-mentioned tempo, the above-mentioned beat, or the above-mentioned strong beat; and A method comprising the step of obtaining quantized data corresponding to the music data based on the music structure information and the melody data.
2. In Paragraph 1, A step of obtaining metadata included in the music data based on the music data; A step of normalizing the above metadata with a predetermined rule; and A method further comprising the step of storing the normalized metadata and the quantized data.
3. In Paragraph 1, The above sound source data for each instrument is, It includes at least one of a first sound source data corresponding to vocals, a second sound source data corresponding to bass, a third sound source data corresponding to drums, or a fourth sound source data corresponding to piano, and The step of acquiring sound source data for each instrument mentioned above is: A method comprising the step of obtaining the sound source data for each instrument by applying the music data to a first model trained to separate sound source data for each instrument.
4. In Paragraph 1, The step of obtaining at least one of the tempo, beat, or strong beat of the sound source data for each instrument is: A step of obtaining the beat and the beat by applying the instrument-specific sound source data to a second model trained to output the beat and the beat from the instrument-specific sound source data; and A method comprising the step of obtaining the tempo based on the beat.
5. In Paragraph 1, The step of acquiring the above music structure information is, A method comprising the step of obtaining music structure information including a musical segment and a pattern in which musical features are repeated in the musical segment by applying at least one of the above-mentioned instrument-specific sound source data, the above-mentioned tempo, the above-mentioned beat, or the above-mentioned strong beat to a third model.
6. In Paragraph 1, The step of acquiring the above melody data is, A step of obtaining first melody data corresponding to a vocal melody based on first sound source data; A step of obtaining second melody data corresponding to a bass melody based on second sound source data; A step of obtaining third melody data corresponding to a drum melody based on third sound source data; A step of obtaining fourth melody data corresponding to a piano melody based on fourth sound source data; and A method comprising the step of obtaining chord data based on the above music data.
7. In Paragraph 1, The step of acquiring the above quantization data is, A step of determining the starting point of a musical section in the music data based on at least one of the tempo, beat, or beat; A step of determining the theme of the above musical section; A step of obtaining sub-melody data corresponding to the theme from the above melody data; and A method comprising the step of obtaining the quantization data based on the notes included in the sub-melody data and the position data of the notes in the musical section.
8. In Paragraph 1, A step of receiving first music data subject to plagiarism check; A step of determining a second music data, which is one of a plurality of stored music data; A step of obtaining similarity between the first music data and the second music data; A step of adding to a list of songs similar to the first music data, based on a determination that the similarity is above a certain standard; and A method further comprising the step of determining a third music data, which is one of the plurality of music data, based on a determination that the similarity is less than a certain standard.
9. In Paragraph 8, The step of obtaining the above similarity is, A step of obtaining at least one of melody similarity or overall song similarity between the first music data and the second music data; and A method comprising the step of obtaining the similarity by assigning a weight to at least one of the melody similarity or the overall song similarity.
10. Regarding the method of providing information on music plagiarism check results, A step of receiving first music data subject to music plagiarism check; A step of identifying second music data having a similarity to the first music data above a certain standard; A step of generating page information that displays the results of a music plagiarism check between the first music data and the second music data; and The method includes the step of transmitting the above page information to a terminal, The above page information A method comprising at least one of plagiarism result information in the entire music segment between the first music data and the second music data or comparison result information for each music segment.
11. In Paragraph 10, The above plagiarism result information is, A method determined based on the similarity between a plurality of music segments included in the first music data and a plurality of music segments included in the second music data.
12. In Paragraph 10, The above comparison result information by music section is, A method comprising comparison result information between a first music segment, which is one of a plurality of music segments included in the first music data, and a second music segment, which is one of a plurality of music segments included in the second music data.
13. In Paragraph 12, The above information on the comparison results by music section A method comprising at least one of a sound figure of the first music section, a harmony of the first music section, a sound figure of the second music section, a harmony of the second music section, a link for playing the first music data, or a link for playing the second music data.
14. In Paragraph 13, The step of generating the above page information is, A method comprising the step of generating page information such that in a two-dimensional graph in which the first axis is a beat and the second axis is a musical pattern, the musical pattern of the first music section is displayed in a first color and the musical pattern of the second music section is displayed in a second color.
15. In Paragraph 14, A method further comprising the step of generating page information such that the musical pattern of the second music section, which is identical to the musical pattern of the first music section in the above two-dimensional graph, is displayed in a third color.
16. In Paragraph 10, The step of generating the above page information is, A step of generating page information such that a first object is included in a portion of the page displayed on the screen of the terminal to display comparison result information for all of the plurality of music segments included in the first music data; and A method comprising the step of generating page information such that the similarity between one of the plurality of music segments included in the first music data and one of the plurality of music segments included in the second music data in the aforementioned partial area is greater than or equal to a certain standard, and including a second object for displaying one or more music segments included in the first music data.
17. In Paragraph 10, The above second music data is, It is one music data determined by user input among a plurality of third music data having a similarity to the first music data above a certain standard, and The step of generating the above page information is, A method comprising the step of generating page information such that a music list including the plurality of third music data is displayed in a first area of a page displayed on the screen of the terminal.
18. In Paragraph 17, The step of generating the above page information is, A method comprising the step of generating page information such that at least one of the plagiarism result information in the entire music segment or the comparison result information for each music segment is displayed in a second area located next to the first area.
19. Electronic devices Memory; and One or more processors; Includes, The above one or more processors, Receive music data, and Based on the above first music data, sound source data for each instrument is obtained, and Based on the above-mentioned sound source data for each instrument, at least one of the tempo, beat, or strong beat of the above-mentioned sound source data for each instrument is obtained, and Based on the sound source data for each instrument mentioned above, melody data is obtained, and Music structure information is obtained based on at least one of the above-mentioned instrument-specific sound source data, the above-mentioned tempo, the above-mentioned beat, or the above-mentioned strong beat, and An electronic device that acquires quantized data corresponding to the music data based on the music structure information and the melody data.
20. In Paragraph 19, The above one or more processors, Identifying second music data that has a similarity to the first music data above a certain standard, and Generate page information displaying the results of a music plagiarism check between the first music data and the second music data, and Transmit the above page information to the terminal, and The above page information An electronic device comprising at least one of plagiarism result information in the entire music segment between the first music data and the second music data or comparison result information for each music segment.