New sound source production platform based on user input

WO2026164372A1PCT designated stage Publication Date: 2026-08-06WHIRIK AI INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
WHIRIK AI INC
Filing Date
2025-12-12
Publication Date
2026-08-06

Smart Images

  • Figure KR2025021579_06082026_PF_FP_ABST
    Figure KR2025021579_06082026_PF_FP_ABST
Patent Text Reader

Abstract

A sound source production method of a sound source production service server, according to an embodiment of the present invention, may be configured to comprise the steps of: receiving a purpose of use and an example sound source from a user requesting sound source generation; extracting, as a main element, an element causing an emotion corresponding to the purpose of use from the example sound source; requesting at least one composer terminal to compose a candidate result around the extracted main element; and confirming a final result according to selection by the user when the candidate result is received from the at least one composer terminal.
Need to check novelty before this filing date? Find Prior Art

Description

New user input-based music production platform

[0001] The present invention relates to a method for producing sound based on sentiment analysis using AI.

[0002] This invention is a research result of the Technology Innovation Program for Startups (TIPS) conducted with support from the Ministry of SMEs and Startups. [Project Title: Automatic Composition Data Collection System for Rapid Production of Content-Customized Background Music and Multimodal Composition Copilot AI Model, Project No.: RS-2024-00512934, Lead Agency: Whirlic AI Co., Ltd., Specialized Agency: Korea Technology Information Promotion Agency for SMEs, Period: 2024.11.01-2026.10.31]

[0003] Music production is a core process of the music industry and creative activities, and its importance is becoming increasingly prominent with the advancement of digital technology. In the past, music production relied on the collaboration of professional studios, composers, and sound engineers; however, recent advancements in digital platforms and artificial intelligence have created an environment where individual users can easily participate in music production. While these technological advancements are revolutionizing the production process, limitations still exist in producing music that accurately reflects the user's intentions or emotions.

[0004] In particular, when users sought to express specific emotions or produce audio suitable for a particular purpose, existing audio production tools faced the problem of being unable to concretely reflect the user's intentions. Generally, users express their desires by describing the direction of the audio they want or providing example tracks; however, this approach showed limitations in enabling composers or production tools to effectively interpret the user's emotions and reflect them in the audio. Consequently, issues such as a lack of communication between users and composers and low satisfaction with the final audio output frequently occurred during the production process.

[0005] Recently, methods of music production utilizing artificial intelligence technology have been gaining attention to address these issues. AI models can contribute to music production by learning from vast amounts of audio and emotion data to analyze specific elements that trigger user emotions in particular tracks and leveraging this analysis. Furthermore, this technology is useful for effectively incorporating user feedback during the composition process or for creating customized tracks by combining various compositions. However, current technology still fails to fully satisfy detailed user requirements in the production process, necessitating the development of more sophisticated and user-friendly music production systems.

[0006] To solve these technical problems, the present invention proposes a method of receiving the purpose of sound production and example sound samples from a user, analyzing the detailed elements of the example sound samples using an artificial intelligence model, producing candidate sound samples based on this analysis, and determining the final sound sample according to user feedback. Through this, it is expected that it will be possible to produce sound samples that more accurately reflect the user's intentions and emotions.

[0007] <Prior Art Literature>

[0008] <Patent Literature>

[0009] (Patent Document 1) Republic of Korea Published Patent No. 10-2024-0020625 (Artificial intelligence system that analyzes big data-based videos and provides background music and sound tailored to the user's preferences)

[0010] This invention was devised to overcome the existing limitations of failing to accurately reflect user intent during the sound production process. Specifically, this invention is designed to propose a method for providing an optimal sound source by precisely analyzing user requests and emotions evoked by example sound sources, producing candidate sound sources based on this analysis, and incorporating user feedback or combining candidate sound sources.

[0011] The purposes of the present disclosure are not limited to those mentioned above, and other purposes and advantages of the present disclosure not mentioned may be understood from the following description and will be more clearly understood from the embodiments of the present disclosure. Furthermore, it will be readily apparent that the purposes and advantages of the present disclosure can be realized by the means and combinations thereof set forth in the claims.

[0012] A method for producing a sound source of a sound source production service server according to an embodiment of the present invention may include the steps of: receiving a purpose of use and an example sound source from a user requesting the production of a sound source; extracting a key element that induces an emotion corresponding to the purpose of use from the example sound source; requesting at least one composer terminal to compose a candidate result based on the extracted key element; and, when the candidate result is received from the at least one composer terminal, determining a final result according to the user's selection.

[0013] And at this time, in the step of extracting key elements, the server may perform the steps of: identifying a plurality of detailed elements constituting the example sound source through sound source data analysis of the example sound source; identifying an emotion matching each of the plurality of detailed elements based on a plurality of sound source data related to each of the plurality of detailed elements and emotion data related to each of the plurality of sound source data; and extracting at least one key element corresponding to an emotion corresponding to the purpose of use among the plurality of detailed elements.

[0014] Additionally, when the server identifies a plurality of detailed elements constituting the example sound source, it may perform the steps of separating the example sound source into a plurality of audio signals corresponding to the sound of each instrument based on the plurality of audio signals, generating an electronic score for the sound of each instrument based on the plurality of audio signals, and classifying each audio signal corresponding to an individual instrument into a plurality of detailed elements according to a change in a pattern related to at least one of pitch and frequency based on the electronic score.

[0015] At this time, the above detailed elements may be characterized as corresponding to elements of an audio signal distinguished according to at least one of sound repeatability, frequency band, rhythmic regularity, sound intensity and temporal composition of sound, and sound diversity.

[0016] Furthermore, when the above-mentioned server extracts the above-mentioned key elements, it can extract the above-mentioned key elements based on the operation of an artificial intelligence model. At this time, the artificial intelligence model is characterized by being trained based on data regarding the correlation between audio data and the corresponding user emotion, and deriving user emotion corresponding to the detailed elements extracted from the input audio source.

[0017] And when the server confirms the final result according to the user selection, it may include the steps of receiving a request from the user for selection and modification of one of the candidate results, transmitting the user's modification request to the composer terminal that provided the selected candidate result, and, when the modified candidate result is received from the composer terminal, requesting the user to confirm the modified candidate result.

[0018] And, in the step of determining the final result according to the user selection, the server may perform the steps of receiving a request from the user for selection and combination of at least one candidate result, receiving information from the user regarding combination requirements for the selected at least one candidate result, generating a modified candidate result by combining the selected at least one candidate result in response to the combination requirements, and requesting the user to confirm the modified candidate result.

[0019] The present invention utilizes an artificial intelligence model to analyze example audio provided by a user and extracts the emotional elements of the audio as key elements, thereby enabling the production of audio that accurately reflects the user's intention.

[0020] By generating candidate sound sources in conjunction with a composer's terminal based on detailed elements of example sound sources received from the user, and supporting the user to select and combine them, it is possible to support faster and more efficient sound source production than before.

[0021] FIG. 1 is a diagram illustrating the communication operation between a server and a peripheral terminal according to an embodiment of the present invention.

[0022] FIG. 2 is a diagram illustrating the configuration of a server according to an embodiment of the present invention.

[0023] FIGS. 3 and FIGS. 4 are flowcharts illustrating the sequence of operations of a server according to an embodiment of the present invention.

[0024] FIG. 5 is a diagram illustrating an operation related to the selection of candidate results according to an embodiment of the present invention.

[0025] <Explanation of Symbols>

[0026] 100 : Server

[0027] 110 : Memory

[0028] 120 : Communications Department

[0029] 130 : Processor

[0030] 200 : User terminal

[0031] 300 : Composer's Terminal

[0032] The advantages and features of the present invention and the methods for achieving them will become clear by referring to the embodiments described below in detail together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below but can be implemented in various different forms. These embodiments are provided merely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the present invention, and the present invention is defined only by the scope of the claims.

[0033] The terms used in this specification are for describing embodiments and are not intended to limit the invention. In this specification, the singular form includes the plural form unless specifically stated otherwise in the text. The terms "comprises" and / or "comprising" used in this specification do not exclude the presence or addition of one or more other components in addition to the components mentioned. Throughout the specification, the same reference numerals refer to the same components, and "and / or" includes each of the mentioned components and all combinations of one or more. Although terms such as "first," "second," etc., are used to describe various components, these components are not limited by these terms. These terms are used merely to distinguish one component from another. Therefore, the first component mentioned below may be the second component within the technical scope of the invention.

[0034] Unless otherwise defined, all terms used herein (including technical and scientific terms) may be used in a meaning commonly understood by those skilled in the art to which the present invention pertains. Additionally, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise.

[0035] The terms “part” or “module” as used in the specification refer to software or hardware components, such as FPGAs or ASICs, and “part” or “module” perform certain roles. However, “part” or “module” is not limited to software or hardware. “Part” or “module” may be configured to reside in an addressable storage medium or configured to run on one or more processors. Thus, by example, “part” or “module” includes components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functions provided within the components and “parts” or “modules” may be combined into a smaller number of components and “parts” or “modules,” or further separated into additional components and “parts” or “modules.”

[0036] Spatially relative terms such as "below," "beneath," "lower," "above," and "upper" may be used to facilitate the description of the relationship between one component and other components as illustrated in the drawings. Spatially relative terms should be understood as encompassing different orientations of components during use or operation, in addition to the orientations depicted in the drawings. For example, if a component depicted in a drawing is inverted, a component described as "below" or "beneath" of another component may be placed "above" of that component. Therefore, the exemplary term "below" may encompass both the lower and upper directions. Components may also be oriented in other directions, and accordingly, spatially relative terms may be interpreted according to the orientation.

[0037] Embodiments of the present invention will be described below with reference to the attached drawings.

[0038] FIG. 1 is a diagram illustrating the communication operation between a server and a peripheral terminal according to an embodiment of the present invention.

[0039] As illustrated in FIG. 1, the server (100) can receive information about the purpose of use and example sound source along with a request for sound source generation from the user side (e.g., user terminal (200)).

[0040] The above server (100) can use an artificial intelligence model to extract key elements that induce emotions corresponding to the purpose of use from example sound sources, and can transmit a request for composition of candidate results to a composer terminal (300) based on the extracted key elements.

[0041] The composer terminal (300) may correspond to a terminal device of a composer that performs composition according to a request, but may also correspond to a terminal, server, virtual server, etc. that includes a digital entity that is not a person, such as an AI composer.

[0042] The above server (100) may provide candidate results to a user terminal (200) after receiving them from a composer terminal (300). At this time, at least one candidate result may be received.

[0043] And the server (100) receives a selection from the user terminal (200) for one item (e.g., an entire specific song, a part constituting a portion of a song, etc.) among at least one candidate result, and accordingly, can confirm the final result or manage the modification process. And the server (100) can transmit the finally confirmed final result to the user terminal (200).

[0044] FIG. 2 is a diagram illustrating the configuration of a server according to an embodiment of the present invention.

[0045] The above server (100) may be configured to include a memory (110), a communication unit (120), and a processor (130), as shown in FIG. 2.

[0046] The memory (110) can store usage purpose and example sound source data received from the user. The memory (110) can also store an artificial intelligence model. The memory (110) can temporarily store data related to detailed elements and key elements extracted from the example sound source.

[0047] The above communication unit (120) can receive a request for sound source creation, a purpose of use, and an example sound source through communication with a user terminal (200). It can transmit a composition request through communication with a composer terminal (300) and receive candidate results from the composer terminal.

[0048] The above communication unit (120) can transmit a list of candidate results to a user terminal (200) and receive a request for selection, modification, or combination from the user.

[0049] The above communication unit (120) can transmit a modification request or a combination request including combination requirements to a composer terminal and receive a modified result.

[0050] And the communication unit (120) can transmit the final result to the user terminal.

[0051] The above processor (130) can perform all operations necessary for the sound source generation process according to an embodiment of the present invention.

[0052] Specifically, the processor (130) first receives a request for sound source generation from a user terminal (200) and can verify data regarding the purpose of use and example sound source included in the request. At this time, the processor (130) can perform a data verification operation for the received example sound source. The data verification operation may refer to an operation to determine validity by checking, for example, file format, size, etc. If verification fails, the processor (130) can send an error message to the user terminal.

[0053] And the above processor (130) can perform data analysis operations on example sound sources.

[0054] Specifically, the processor (130) can receive data regarding the example sound source and perform an instrument-specific sound separation operation based on an artificial intelligence model. For example, the instrument-specific sound separation operation may include separating only the sound of Violin 1 from an orchestral work. Furthermore, to perform such instrument-specific sound separation, the processor (130) can convert the example sound source file into a frequency spectrum and separate the sound of each instrument based on an algorithm that separates the sound of a specific instrument in each frequency band. The data separated by instrument sound in this way can be configured to be stored separately and listened to separately.

[0055] At this time, the above artificial intelligence model can separate the sounds of each instrument by converting the sound source data into spectrograms and then using a deep learning-based sound separation model (e.g., U-Net architecture) to separate the sounds of each instrument.

[0056] Subsequently, the processor (130) can perform the operation of generating an electronic score (MIDI) based on the separated instrument-specific data. The processor (130) can convert the sound source data separated by each instrument into an electronic score format based on time-based pitch, length, and intensity data.

[0057] When generating an electronic score, the processor (130) can separate the data of each instrument and store it in individual tracks. For example, the piano can be stored separately in Track 1 and the violin in Track 2. Additionally, when generating an electronic score, the processor (130) can additionally perform the operation of inserting messages such as tempo (overall tempo of the song, tempo that changes over time), dynamics (changes in intensity such as crescendo and decrescendo), and expression symbols (playing techniques such as legato, staccato, and trill).

[0058] Subsequently, the processor (130) can extract detailed elements of a sound source based on features including sound repeatability (the degree to which a specific sound is repeated at regular intervals; the higher the repeatability, the more monotonous the rhythm), frequency band (the proportion occupied by a specific frequency range), rhythmic regularity (the degree to which the time interval of sound occurrence or beat is regular), sound intensity (e.g., volume intensity of a specific frequency range), temporal composition (sound length, identified based on the start and end times of the sound), and sound diversity (diversity of sounds occurring in a single instrument, presence or absence of chords, etc.). The detailed elements of the sound source may be extracted as the composition of an audio signal corresponding to each of the above features, or may correspond to the composition of an audio signal having a combination of at least two of the above features. In addition, the detailed elements may refer to elements of an audio signal distinguished based on information regarding additional information (e.g., crescendo, staccato, etc.) written in the electronic score, in addition to the aforementioned features.

[0059] For example, the processor (130) can analyze the cycle of a specific pattern repeating and quantify the value of repeatability. In addition, the processor (130) can identify the bandwidth of the sound by calculating the frequency ratio of each pitch range.

[0060] Additionally, when the processor (130) extracts a detailed element corresponding to an audio signal in which two or more features are combined, it can extract an audio signal that includes both frequency band and sound intensity features as a detailed element of the sound source. For example, the processor (130) can identify a sound element in which the intensity (decibel) in the mid-range (e.g., 250Hz to 2kHz) is maintained above a reference value as a sound corresponding to a 'strong mid-range' and extract it as a detailed element.

[0061] Alternatively, the processor (130) can extract detailed elements by combining repetition and sound intensity. For example, in a drum pattern, if a constant rhythm is repeated and the intensity changes regularly, it can be identified as a sound corresponding to a 'rhythm in which intensity appears repeatedly' and extracted as a detailed element.

[0062] According to one embodiment, the processor (130) can divide each audio signal corresponding to an individual instrument into a plurality of detailed elements based on the electronic musical score, according to a change in a pattern associated with at least one of pitch (frequency, sound intensity) and frequency (sound repeatability).

[0063] And the processor (130) can identify a plurality of detailed elements for an example sound source in this way, and identify an emotion that matches each of the plurality of detailed elements based on a plurality of sound source data associated with each of the plurality of detailed elements and emotion data associated with each of the plurality of sound source data.

[0064] And the processor (130) can extract at least one major element corresponding to an emotion corresponding to the purpose of use among a plurality of detailed elements.

[0065] Specifically, the processor (130) may extract data regarding the relationship between the extracted detailed elements and user emotions, and define the detailed elements as elements that trigger the corresponding emotions. Furthermore, the processor (130) may set the elements that trigger emotions corresponding to the user's input purpose as the main elements among the elements that trigger user emotions.

[0066] In this case, the emotion corresponding to the above purpose of use may be based on the content directly specified in the purpose of use entered by the user. For example, if the purpose of use is 'horror movie', the emotion corresponding to the purpose of use may be identified as 'fear'.

[0067] Meanwhile, in cases where no terms directly related to emotion are listed in the above-mentioned purpose of use, the processor (130) can identify the type of emotion with the highest probability corresponding to the purpose of use and select it as the target emotion when the value of the specificity indicator (e.g., number of characters, number of words, complexity of sentence structure) for the purpose of use entered by the user is greater than or equal to a threshold value (e.g., 10 characters or more, 3 or more words, including subordinate clauses, etc.).

[0068] For example, when a purpose of use such as 'music to play in a cafe on a rainy day' is input, the processor (130) can analyze the text corresponding to the purpose of use and extract keywords associated with emotions. The processor (130) can extract keywords for 'rainy day' and 'cafe', and based on an artificial intelligence model, compare the extracted keywords with an emotion database to calculate the probability of each type of emotion corresponding to each keyword. For example, the processor (130) can derive probabilities of each type of emotion such as 'depression' 60% and 'calmness' 40% corresponding to 'rainy day', and probabilities of each type of emotion such as 'calmness' 70% and 'joy' 30% corresponding to 'cafe'.

[0069] And the processor (130) combines the probabilities of each emotion to identify the emotion with the highest probability, and accordingly, the processor (130) can determine the 'emotion corresponding to the purpose of use'.

[0070] In this case, the method for aggregating the probabilities of each emotion can be based on the maximum value-first selection method, the weighted average method, etc.

[0071] For example, the processor (130) can determine the probability value of each emotion by selecting it with priority of the maximum value, and can select the emotion with the highest probability based on the probability value. For example, in a situation where depression is 60%, calmness is 40%, calmness is 70%, and joy is 30%, the processor (130) can select 70% as the maximum value for calmness is 70%, and determine 'calmness', which is the emotion with the highest probability value, as the emotion corresponding to the purpose of use.

[0072] Additionally, the processor (130) can calculate the final probability value by multiplying the probability value of each keyword by a weight according to the importance of the keyword and then dividing it by the sum of the weights.

[0073] For example, a weight of 0.6 may be assigned to the keyword corresponding to 'rainy day', and a weight of 0.4 may be assigned to the keyword corresponding to 'cafe'.

[0074] And as in the previous example, if the probability of calmness on a 'rainy day' is 40% and the probability of calmness on a 'cafe' is 70%, the final probability value for calmness can be calculated as follows by applying weights.

[0075] (Final probability value for calmness) = (0.6*40)+(0.4*70) / (0.6+0.4) = 52% and the processor (130) can identify the type of emotion with the highest probability value based on the final probability value calculated for each emotion in this way and determine it as the emotion corresponding to the purpose of use.

[0076] Furthermore, the processor (130) can determine the ‘emotion corresponding to the purpose of use’ based on the overall mood of the example music and the degree of matching of the keyword analysis results for the purpose of use.

[0077] For example, the processor (130) can derive probability values ​​for each type of emotion based on the overall mood of the sound source based on elements such as the frequency band, rhythm, and intensity of the example music, and can identify the type of emotion with the highest probability by combining the emotion probability for the purpose of use and the emotion probability based on the mood of the example music. The processor (130) can then set the type of emotion identified in this manner as the target emotion.

[0078] And the processor (130) can extract key elements based on the relationship between the data on the extracted detailed elements and the target emotion (emotion corresponding to the purpose of use).

[0079] The above processor (130) can determine the association between the detailed element extracted during the process of extracting the main element and the target emotion (e.g., fear, comfort) based on an artificial intelligence model, and determine the detailed element corresponding to the target emotion as the main element.

[0080] In this case, the above-mentioned artificial intelligence model can be trained based on data regarding the correlation between audio data and corresponding user emotions (e.g., large-scale audio data and corresponding user emotion data). Furthermore, the above-mentioned artificial intelligence model can learn the relationship mapping detailed elements of each audio source to corresponding emotions (e.g., fear, joy, calmness), and through this, analyze the input detailed elements to predict probability values ​​for each type of emotion and identify the detailed elements most suitable for the target emotion.

[0081] Accordingly, the artificial intelligence model can derive user emotions corresponding to detailed elements extracted from the input sound source. Based on this, the processor (130) can identify detailed elements corresponding to the target emotion and determine them as key elements.

[0082] Subsequently, the processor (130) can generate composition request data by organizing information about the main elements into a list form. At this time, the composition request data may be configured to include descriptive data about the main elements and sound data corresponding to the main elements. Alternatively, the composition request data may include only sound data corresponding to the main elements.

[0083] Subsequently, the processor (130) transmits data regarding the extracted key elements to at least one composer terminal (300) and may request a composition to include the key elements. The processor (130) may further include information corresponding to the purpose of use of the sound source received from the user (type of content to which the sound source is applied, e.g., movies, games, etc.) in the data requesting a composition to the composer terminal (composition request data). Alternatively, the processor (130) may extract information regarding a preset sound source format corresponding to the purpose of use and include the extracted information regarding the sound source format in the composition request data.

[0084] After a composition request including the above key elements is transmitted to the composer terminal, and the composition is completed from the composer terminal (300) and a candidate result of the sound source is received, the processor (130) can perform a processing operation on the candidate result.

[0085] First, the processor (130) can determine suitability by primarily analyzing the metadata of the candidate results. For example, the processor (130) can determine suitability by reviewing basic characteristics such as the length and sound quality of the candidate results.

[0086] The processor (130) can determine whether the length of the result corresponds to the purpose of use received from the user. The processor (130) can derive in advance a range of standard sound source length values ​​corresponding to the purpose of use received from the user, or receive and verify in advance a range of sound source lengths requested by the user. Accordingly, the processor (130) can determine whether the length of the candidate result deviates from the standard range and determine its suitability based on this. In addition, the processor (130) can detect the proportion of factors that may affect sound quality, such as noise, based on sampling, and evaluate the quality of the sample sound source based on this. If the sound quality of the sample sound source is below a standard value, the processor (130) can determine that it is unsuitable. In this way, the processor (130) can simply perform a primary suitability evaluation of the candidate results received from the composer terminal and provide only the candidate results that have passed the suitability judgment to the user terminal.

[0087] In this way, the processor (130) can select a candidate result whose suitability has been verified among the candidate results received from the composer terminal, and can provide at least one candidate result whose suitability has been verified to the user terminal.

[0088] Afterward, when selection information regarding a specific candidate result is received from the user terminal (200), the data is selected as the final result, and a corresponding process (e.g., providing guidance information to the composer terminal of the final result and a reward step) can be carried out.

[0089] Meanwhile, according to various embodiments, the processor (130) may receive not only a selection of candidate results but also a modification request or a combination request from a user terminal, and execute a process corresponding thereto.

[0090] The operation will be explained with reference to FIG. 5. FIG. 5 is a diagram illustrating an operation related to the selection of candidate results according to an embodiment of the present invention.

[0091] First, as shown in 510 of FIG. 5, the user terminal can provide the user with candidate results such as A, B, and C received from the server, and the user can select A among the candidate results. Accordingly, the processor (130) can determine the selected A as the final result.

[0092] Meanwhile, as in 520, the user may select multiple candidate results (A and B) and request a combination of the selected multiple candidate results, and accordingly, the server (100) may proceed with a subsequent process to perform a combination of the multiple candidate results selected by the user. For example, the server (100) may perform processes such as selecting a composer terminal to perform the combination work, requesting a combination work to the selected composer terminal, and receiving the combined candidate results from the composer terminal and providing them to the user terminal. However, the process is not limited to a process including a combination request operation to the composer terminal as described above. When the server (100) receives a combination request from the user terminal, it may request the user terminal to additionally input combination conditions for the candidate results to be combined, and according to the received combination conditions (e.g., combining by setting the time before the reference point of the sound source as A and the time after the reference point as B), the processor (130) of the server (100) may independently perform a combination operation to generate a modified candidate result.

[0093] Meanwhile, even if a single candidate result is selected, a new combination of each part constituting a portion of the song can be performed to generate a modified candidate result.

[0094] Additionally, as shown in 530 of FIG. 5, the user may select one of the presented candidate results and then send a modification request (including information regarding modifications) to the server. Accordingly, the server (100) may send a modification request for the candidate result, along with instructions regarding modifications, to the composer terminal that produced the candidate result selected by the user. Afterward, when the modified candidate result is received from the composer terminal, it is provided back to the user terminal, and when the user makes a final selection regarding it, the server (100) may determine it as the final result.

[0095] FIGS. 3 and FIGS. 4 are flowcharts illustrating the sequence of operations of a server according to an embodiment of the present invention.

[0096] As illustrated in FIG. 3, the server (100) may perform step 310 of receiving the purpose of use of the sound source (e.g., including the type or genre of content to which the sound source will be applied) and an example sound source from a user requesting the creation of a sound source.

[0097] Afterwards, the server (100) can perform step 320 of extracting key elements from the example sound source.

[0098] Afterwards, the above server (100) can perform step 330, which requests at least one composer terminal to compose a candidate result centered on the main element.

[0099] Afterwards, when the server (100) receives candidate results from at least one composer terminal, it can perform step 340 of selecting a final result according to user input.

[0100] At this time, the above 320 steps can be further subdivided more specifically, as shown in FIG. 4.

[0101] As illustrated in FIG. 4, the above server (100) may be configured to include, in the step of extracting key elements from an example sound source, step 410 of separating sounds by instrument for the example sound source, step 420 of generating an electronic score (MIDI) for the sounds separated by instrument, and step 430 of identifying and extracting key elements that induce emotions corresponding to the purpose of use from the example sound source. At this time, step 430 can identify detailed elements for the example sound source based on the electronic score generated in step 420, and can extract the corresponding elements as key elements by making a judgment on whether they induce the emotion required by the user based on the correspondence between these detailed elements and the user's emotion.

[0102] In summary, a method for producing a sound source of a sound source production service server according to an embodiment of the present invention may include the steps of: receiving a purpose of use and an example sound source from a user requesting the production of a sound source; extracting a key element that induces an emotion corresponding to the purpose of use from the example sound source; requesting at least one composer terminal to compose a candidate result based on the extracted key element; and, when the candidate result is received from the at least one composer terminal, determining a final result according to the user's selection.

[0103] And at this time, in the step of extracting key elements, the server may perform the steps of: identifying a plurality of detailed elements constituting the example sound source through sound source data analysis of the example sound source; identifying an emotion matching each of the plurality of detailed elements based on a plurality of sound source data related to each of the plurality of detailed elements and emotion data related to each of the plurality of sound source data; and extracting at least one key element corresponding to an emotion corresponding to the purpose of use among the plurality of detailed elements.

[0104] Additionally, when the server identifies a plurality of detailed elements constituting the example sound source, it may perform the steps of separating the example sound source into a plurality of audio signals corresponding to the sound of each instrument based on the plurality of audio signals, generating an electronic score for the sound of each instrument based on the plurality of audio signals, and classifying each audio signal corresponding to an individual instrument into a plurality of detailed elements according to a change in a pattern related to at least one of pitch and frequency based on the electronic score.

[0105] At this time, the above detailed elements may be characterized as corresponding to elements of an audio signal distinguished according to at least one of sound repeatability, frequency band, rhythmic regularity, sound intensity and temporal composition of sound, and sound diversity.

[0106] Furthermore, when the above-mentioned server extracts the above-mentioned key elements, it can extract the above-mentioned key elements based on the operation of an artificial intelligence model. At this time, the artificial intelligence model is characterized by being trained based on data regarding the correlation between audio data and the corresponding user emotion, and deriving user emotion corresponding to the detailed elements extracted from the input audio source.

[0107] And when the server confirms the final result according to the user selection, it may include the steps of receiving a request from the user for selection and modification of one of the candidate results, transmitting the user's modification request to the composer terminal that provided the selected candidate result, and, when the modified candidate result is received from the composer terminal, requesting the user to confirm the modified candidate result.

[0108] And, in the step of determining the final result according to the user selection, the server may perform the steps of receiving a request from the user for selection and combination of at least one candidate result, receiving information from the user regarding combination requirements for the selected at least one candidate result, generating a modified candidate result by combining the selected at least one candidate result in response to the combination requirements, and requesting the user to confirm the modified candidate result.

[0109] A server (100) according to an embodiment of the present invention may include a processor, memory, a communication unit, etc.

[0110] Memory can store various programs and data necessary for the operation of electronic devices. Memory can be implemented as non-volatile memory, volatile memory, flash memory, hard disk drives (HDD), or solid-state drives (SSD).

[0111] The communication unit can communicate with external devices. In particular, the communication unit may include various communication chips such as Wi-Fi chips, Bluetooth chips, wireless communication chips, NFC chips, and low-power Bluetooth chips (BLE chips). In this case, the Wi-Fi chip, Bluetooth chip, and NFC chip perform communication using LAN, Wi-Fi, Bluetooth, and NFC methods, respectively. When using a Wi-Fi chip or a Bluetooth chip, various connection information such as SSID and session key is transmitted and received first, and after establishing a communication connection using this information, various information can be transmitted and received. A wireless communication chip refers to a chip that performs communication according to various communication standards such as IEEE, Zigbee, 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), and LTE (Long Term Evolution).

[0112] The processor can control the overall operation of a user device using various programs stored in memory. The processor may be composed of RAM, ROM, a graphics processing unit, a main CPU, first to n interfaces, and a bus. At this time, the RAM, ROM, graphics processing unit, main CPU, first to n interfaces, etc., may be connected to each other through a bus.

[0113] RAM stores the operating system and application programs. Specifically, when an electronic device boots up, the operating system is stored in RAM, and various application data selected by the user can be stored in RAM.

[0114] The ROM stores a set of instructions for booting the system, etc. When a turn-on command is input and power is supplied, the main CPU copies the O / S stored in memory (200) to RAM according to the instructions stored in the ROM, and runs the O / S to boot the system. When booting is complete, the main CPU copies various application programs stored in memory to RAM, and runs the application programs copied to RAM to perform various operations.

[0115] The main CPU accesses memory and performs operations, including booting and execution, using the OS stored in memory. Additionally, the main CPU performs various operations using various programs, content, and data stored in memory.

[0116] The first to n interfaces are connected to the various components described above. One of the first to n interfaces may be a network interface connected to an external device through a network.

[0117] Furthermore, the processor can control the artificial intelligence model. In this case, the control unit may, of course, include a dedicated graphics processor (e.g., GPU) for controlling the artificial intelligence model.

[0118] The processor may include one or more cores (not shown) and a graphics processing unit (not shown) and / or a connection channel (e.g., a bus, etc.) for transmitting and receiving signals with other components.

[0119] A processor according to one embodiment performs the method described in connection with the present invention by executing one or more instructions stored in memory.

[0120] Meanwhile, the processor may further include RAM (Random Access Memory, not shown) and ROM (Read-Only Memory, not shown) for temporarily and / or permanently storing signals (or data) processed within the processor. Additionally, the processor (130) may be implemented in the form of a system-on-chip (SoC) that includes at least one of a graphics processing unit, RAM, and ROM.

[0121] Memory can store programs (one or more instructions) for processing and controlling the processor. Programs stored in the storage unit can be divided into multiple modules according to their function.

[0122] The steps of the method or algorithm described in connection with embodiments of the present invention may be implemented directly in hardware, implemented as a software module executed by hardware, or implemented by a combination thereof. The software module may reside in RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), Flash Memory, a hard disk, a removable disk, a CD-ROM, or any form of computer-readable recording medium well known in the art to which the present invention belongs.

[0123] The components of the present invention may be implemented as a program (or application) and stored on a medium to be executed in combination with a computer, which is hardware. The components of the present invention may be implemented as software programming or software elements, and similarly, embodiments may be implemented in programming or scripting languages ​​such as C, C++, Java, assembler, etc., including various algorithms implemented as combinations of data structures, processes, routines, or other programming configurations. Functional aspects may be implemented as algorithms executed on one or more processors.

[0124] Although the present invention has been described in detail with reference to the examples above, those skilled in the art may make modifications, changes, and variations to these examples without departing from the scope of the invention. In short, it should be noted that in order to achieve the intended effect of the present invention, it is not necessary to separately include all functional blocks shown in the drawings or follow all sequences shown in the drawings exactly as shown, and that such matters may fall within the technical scope of the invention as described in the claims.

Claims

1. Regarding the method of producing sound files for a sound file production service server, A step of receiving the purpose of use and example sound source from a user requesting sound source generation; A step of extracting key elements that induce emotions corresponding to the above purpose of use from the above example sound source; A step of requesting at least one composer terminal to compose a candidate result centered on the extracted key elements; and A method comprising the step of determining a final result according to user selection when a candidate result is received from at least one composer terminal.

2. In Paragraph 1, The step of extracting the above key elements is A step of identifying a plurality of detailed elements constituting the example sound source through sound source data analysis of the example sound source; A step of identifying an emotion matching each of the plurality of detailed elements based on a plurality of sound source data associated with each of the plurality of detailed elements and emotion data associated with each of the plurality of sound source data; and A method comprising the step of extracting at least one major element corresponding to an emotion corresponding to the purpose of use among the plurality of detailed elements above.

3. In Paragraph 2, The step of identifying multiple detailed elements constituting the above example sound source is A step of separating the above example sound source into multiple audio signals corresponding to the sounds of each instrument; A step of generating an electronic score for the sound of each instrument based on the plurality of audio signals above; and A method comprising the step of classifying each audio signal corresponding to an individual instrument into a plurality of detailed elements based on the above electronic score according to a change in a pattern associated with at least one of pitch and frequency.

4. In Paragraph 3, The above detailed elements A method characterized by corresponding to an element of an audio signal distinguished according to at least one of sound repeatability, frequency band, rhythmic regularity, sound intensity and temporal composition of sound, and sound diversity.

5. In Paragraph 2, The step of extracting the above key elements is Extracting the above key elements based on the operation of the artificial intelligence model, The above artificial intelligence model A method characterized by being learned based on data regarding the association between sound source data and the corresponding user emotion, and deriving a user emotion corresponding to a detailed element extracted from an input sound source.

6. In Paragraph 1, The step of confirming the final result according to the above user selection is A step of receiving a request from a user to select and modify any one of the above candidate results; A step of transmitting the user's modification request to the composer terminal that provided the selected candidate result; and A method comprising the step of requesting user confirmation of the modified candidate result when a modified candidate result is received from the composer terminal.

7. In Paragraph 1, The step of confirming the final result according to the above user selection is A step of receiving a request from a user for the selection and combination of at least one candidate result; A step of receiving information regarding combination requirements for at least one selected candidate result from the user; A step of generating a modified candidate result by combining at least one selected candidate result in response to the above combination requirements; and A method comprising the step of requesting user confirmation of the above-mentioned modified candidate result.