Method and system for generating overseas expansion strategy by artist characteristics based on artificial intelligence
Patent Information
- Application Number
- KR1020250186211
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-11-28
Smart Images

Figure 112025134651315-PAT00005_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to the fields of artificial intelligence (AI), music data analysis, and entertainment industry technology.
[0002] More specifically, the invention relates to a system and a method of operation thereof that analyzes an artist's music content to extract unique characteristics, comprehensively analyzes musical preferences and trend data of a target overseas market, utilizes an artificial intelligence (AI) model to virtually localize content to align with local market preferences while maintaining the artist's characteristics, and simulates market reactions based on this to generate an optimal overseas expansion strategy. Background Technology
[0004] With the recent surge in global popularity of Korean popular culture content, including K-pop, artists across various genres are actively attempting to enter overseas markets. Traditionally, decisions regarding international expansion have relied on the accumulated experience and know-how of management companies or production firms. While effective decision-making is possible when there is an understanding of a specific market, artists face significant challenges when attempting to expand into new regions or genres, particularly regarding identifying market characteristics, analyzing local preferences, and assessing content suitability. This is especially true for domains long monopolized by artists from specific countries, or for markets with high external entry barriers, such as the Japanese music market, where there have been virtually no means to verify the feasibility of entry in advance.
[0005] Meanwhile, the application of AI technology in the music industry has primarily focused on composition support, music recommendations, and genre classification, while attempts to virtually transform a specific artist's music to suit the preferences of a target market or predict market reactions based on localized content have been insufficient.
[0006] In addition, it may be difficult to quantitatively analyze musical similarities and suggest optimal localization directions for markets that share emotional codes similar to Korean adult pop music, such as the Southeast Asian market, or for areas where new opportunities may exist, such as the European world music market. The problem to be solved
[0008] Conventional methods for formulating strategies for artists' overseas expansion have the problem of failing to quantitatively analyze target market preferences and relying solely on experience and intuition, making it impossible to verify the likelihood of success in advance, as well as failing to create content optimized for the local market while maintaining the artist's unique characteristics.
[0009] The present invention aims to extract unique characteristics such as timbre, vocal technique, and rhythm patterns from an artist's music content in the form of multidimensional vectors, derive locally preferred styles by analyzing market data of a target country, and then use an artificial intelligence model to generate virtual localized content that reflects local music trends while maintaining the artist's identity.
[0010] In addition, the present invention aims to quantitatively calculate a market fit index by comparing generated virtual localized content with reference data of the target market, and based on this, to determine whether immediate entry, conditional entry, or content correction is necessary.
[0011] The present invention aims to improve localization accuracy by constructing a local music grammar model through learning chord progression patterns, scale characteristics, rhythm density, etc., from music chart data of a target country, and by automatically identifying sections in the original song that are inconsistent with the local style and suggesting optimal alternative chords or melody variations. means of solving the problem
[0013] An AI-based system for generating overseas expansion strategies based on artist characteristics may include memory and a processor for storing instructions. The system receives content data of at least one artist from a user terminal, analyzes market data of a target country to be entered to derive preference characteristics of that country, generates virtual localized content by modifying the content data to reflect the derived preference characteristics using an AI model, and controls the system to simulate expected reactions in the target country based on the generated virtual localized content. Effects of the invention
[0015] The AI-based method and system for generating overseas expansion strategies based on artist characteristics according to the present invention quantitatively analyzes the musical uniqueness of an artist and generates virtual localized content by matching it with the preference characteristics of the target market, thereby having the effect of simulating the possibility of success and minimizing risk prior to actual market entry.
[0016] In addition, the present invention utilizes an artificial intelligence model to automatically perform localization that reflects the chord progression, rhythm patterns, and vocal characteristics of the target country while preserving the artist's timbre and style, thereby significantly reducing time and costs compared to manual work.
[0017] The present invention can assist in localization by objectively identifying dissonant sections in the original song that may feel alien to local listeners and suggesting optimal alternatives, based on a local music grammar model constructed by analyzing actual chart data and popular music tracks of the target market. Brief explanation of the drawing
[0019] FIG. 1 is a drawing for explaining a system utilizing an artificial intelligence model according to one embodiment. FIG. 2 is a diagram illustrating the learning of a neural network according to one embodiment. FIG. 3 is a diagram illustrating the configuration of an artificial intelligence model according to one embodiment. Figure 4 is a block diagram showing the configuration of an AI-based system for generating overseas expansion strategies based on artist characteristics according to one embodiment. FIG. 5 is a flowchart illustrating a method for generating an AI-based overseas expansion strategy by artist characteristics according to one embodiment. FIG. 6 is a flowchart illustrating an AI-based artist timbre and style localization method according to one embodiment. FIG. 7 is a flowchart illustrating an evolution algorithm-based content optimization method according to one embodiment. Specific details for implementing the invention
[0020] The embodiments will be described in detail below based on the attached drawings. However, since various modifications are possible to these embodiments, the scope of the patent application is not limited by them. The specific structural or functional descriptions of the embodiments are presented merely for illustrative purposes and may be modified in various forms.
[0021] Unless otherwise specifically defined, all terms used herein have the meaning generally understood by those skilled in the art. Terms defined in commonly used dictionaries must be consistent in meaning within the context of the relevant technology, and where not explicitly defined, they are not interpreted in an overly formal sense. Additionally, when describing drawings, the same reference numeral is assigned to identical components regardless of drawing symbols, and redundant descriptions are omitted.
[0023] FIG. 1 is a drawing for explaining an artificial intelligence-based system according to one embodiment.
[0024] As shown in FIG. 1, an artificial intelligence-based system (100) may be composed of several user terminals (110-1 to 110-n), a server (120), and a database (130). In one embodiment, the database (130) is configured separately from the server (120), but the database (130) may be included within the server (120). For example, the server (120) may include several artificial intelligences for performing machine learning algorithms. According to another embodiment, the user terminals (110-1 to 110-n), the server (120), and the database (130) may be connected to and communicate with each other through a network (N).
[0025] The network (N) may support wireless or wired communication between user terminals (110-1 to 110-n), a server (120), and a database (130). For example, the network may perform wireless communication through methods such as LTE (long-term evolution), LTE-A (LTE Advanced), CDMA (code division multiple access), WCDMA (wideband CDMA), WiBro (Wireless BroadBand), WiFi (wireless fidelity), Bluetooth, NFC (near field communication), GPS (Global Positioning System), or GNSS (global navigation satellite system). The database (130) may store various data, and the stored data may include software (e.g., programs) as data collected, processed, and used by components of multiple user terminals (110-1 to 110-n) or the server (120). The database (130) may include volatile and / or non-volatile memory.
[0026] In the present invention, Artificial Intelligence (AI) refers to a technology that mimics human learning ability, reasoning ability, and perceptual ability and implements them in a computer, and may include concepts such as machine learning and symbolic logic. Machine Learning (ML) is an algorithmic technology that classifies or learns the characteristics of input data on its own. Artificial intelligence technology can perform judgments or predictions by analyzing input data through machine learning algorithms and learning the results. Furthermore, a technology that mimics the cognitive and judgment functions of the human brain by utilizing machine learning algorithms can also be understood as falling within the category of artificial intelligence.
[0027] Machine learning refers to the process of training neural network models based on experience in processing data, through which computer software can improve its own data processing capabilities. Neural network models are constructed by modeling the correlations between data, and these correlations can be expressed by various parameters. The core of machine learning lies in optimizing the model's parameters by extracting and analyzing features from given data to derive correlations between them, and repeating this process. For example, a neural network model can learn the mapping (correlation) between inputs and outputs for data provided as input-output pairs, and can also learn relationships by deriving regularities between data even when only input data is provided. Artificial intelligence learning models, or neural network models, are designed to implement the structure of the human brain on a computer and can include multiple network nodes with weights that mimic neurons in a neural network. These multiple network nodes can be interconnected, simulating the synaptic activity where neurons exchange signals through synapses. The multiple network nodes of an artificial intelligence learning model are located in layers of different depths and can exchange data according to convolutional connections. Examples of artificial intelligence learning models include artificial neural networks and convolutional neural networks (CNNs). In one embodiment, artificial intelligence learning models can be machine learned according to methods such as supervised learning, unsupervised learning, and reinforcement learning.Algorithms for machine learning that can be used include Decision Trees, Bayesian Networks, Support Vector Machines, Artificial Neural Networks, Ada-boost, Perceptrons, Genetic Programming, and Clustering.
[0028] CNNs are a type of multilayer perceptron designed to use minimal preprocessing. CNNs consist of one or more convolutional layers and general artificial neural network layers placed on top of them, and additionally utilize weights and pooling layers. Thanks to this structure, CNNs can effectively process two-dimensional input data. Compared to other deep learning structures, CNNs demonstrate superior performance in the fields of image and speech and can be trained using standard backpropagation. A convolutional network is a neural network composed of a set of nodes with bound parameters.
[0030] FIG. 2 is a diagram illustrating the learning of a neural network according to one embodiment.
[0031] As shown in FIG. 2, the learning device can train a neural network (123) to process review responses received from a plurality of user terminals (110-1 to 110-n) item by item. Additionally, the learning device can train a neural network (123) to extract user stay history from user movement path information. According to one embodiment, the learning device may be a separate entity from the server (120), but this is not limited thereto.
[0032] The neural network (123) includes an input layer (121) into which training samples are input and an output layer (125) into which training outputs are generated, and can be trained based on the difference between the training output and the label. Here, the label is defined according to the item corresponding to the review response and can be defined for the user's dwell history based on movement path information. The neural network (123) is connected by multiple groups of nodes and is defined by weights between the connected nodes and an activation function that activates the nodes.
[0033] The learning device can train the neural network (123) using the Gradient Descent (GD) technique or the Stochastic Gradient Descent (SGD) technique. The learning device can use a loss function designed based on the output and label of the neural network. The learning device can calculate the training error by utilizing a predefined loss function. The loss function can predefine the label, output, and parameters as input variables, where the parameters can be set as weights within the neural network (123). For example, the loss function can be designed in the form of Mean Square Error (MSE) or entropy, and various techniques or methods can be applied in designing the loss function.
[0034] The learning device can find weights that affect the training error through the backpropagation technique. Here, weights represent the relationships between nodes within the neural network (123). The learning device can use the SGD technique utilizing labels and outputs to optimize the weights found by the backpropagation technique. For example, the learning device can update the weights of the loss function defined according to the labels, outputs, and weights through the SGD technique.
[0036] FIG. 3 is a diagram illustrating the configuration of an artificial intelligence model according to one embodiment.
[0037] According to one embodiment, the artificial intelligence model may include an input layer, a hidden layer, and an output layer.
[0038] The input layer is the layer associated with the values input into the artificial intelligence model. In the hidden layer, feature maps can be output by performing multiply-accumulate (MAC) operations and activation operations on the input values. The MAC operation is a process of multiplying the input value by its corresponding weight and then summing the resulting values. The activation operation is a process of inputting the result of the MAC operation into an activation function to output a final result; activation functions can be of various types. For example, activation functions may include sigmoid functions, tangent functions, ReLU functions, Leaky ReLU functions, Max Out functions, and / or ELU functions, but their types are not limited. The hidden layer may consist of at least one layer. For example, if the hidden layer consists of a first hidden layer and a second hidden layer, the first hidden layer performs MAC operations and activation operations based on the input values of the input layer to output a feature map, and the feature map resulting from the first hidden layer can be used as the input value of the second hidden layer. The second hidden layer can perform MAC operations and activation operations according to the feature map resulting from the first hidden layer. The output layer may be a layer related to the result of the operations performed in the hidden layer.
[0039] In one embodiment, the learning model learns syllable (character) patterns that are frequently combined and used in a given corpus to automatically recognize the boundaries of compound words and named entities, and integrates object information from a first UI source with object information rendered in a browser to generate a learning object information file. By utilizing this learning object information file, data for training a deep learning network can be generated, and data can be received from various domains of a support system and standardized into an integrated format based on at least one standardization method corresponding to each domain. Data of a specific domain can be learned and inferred, information to be transmitted for standardization in that domain can be determined, and post-processing can be performed on data collected from various domains.
[0040] The first UI source includes an XML file, and the training object information file includes an input JSON file for feature learning and an output JSON file that serves as label data during training. This output JSON file is a file containing HTML DOM Tree information implemented in compliance with web standards. Various domains include at least one of a RAN (radio access network), a transport, or a core, and post-processing may include a correlation function.
[0042] Figure 4 is a block diagram showing the configuration of an AI-based system for generating overseas expansion strategies based on artist characteristics according to one embodiment.
[0043] A system (400) according to one embodiment may include a processor (420) and a memory (430), and some of the illustrated components may be omitted or substituted. A system (400) according to one embodiment may be a server or a terminal. According to one embodiment, the processor (420) is a component capable of performing operations or data processing regarding the control and / or communication of each component of the system (400), and may be composed of one or more processors. The memory (430) may store information related to the method described above or store a program in which the method described above is implemented. The memory (430) may be volatile memory or non-volatile memory. The memory (430) may store various file data, and the stored file data may be updated according to the operation of the processor (120).
[0044] According to one embodiment, the processor (420) can execute a program and control the device (400). The code of the program executed by the processor (120) can be stored in memory (430). Operations of the processor (420) can be performed by loading instructions stored in memory (430). The system (400) can be connected to an external device (e.g., a personal computer or a network) through an input / output device (not shown in the drawing) and exchange data.
[0045] A system (400) according to one embodiment includes a processor (420) and a memory (430). A system (400) according to one embodiment may be the server or terminal described above. The processor (420) may include at least one device described through FIGS. 1 to 3 or may perform at least one method described through FIGS. 1 to 3. The memory (430) may store information related to the method described above or store a program in which the method described above is implemented. The memory (430) may be volatile memory or non-volatile memory.
[0047] FIG. 5 is a flowchart illustrating a method for generating an AI-based overseas expansion strategy by artist characteristics according to one embodiment.
[0048] The operations described through FIG. 5 can be implemented based on instructions that can be stored in memory (e.g., memory (420) of FIG. 4). The order of each operation of FIG. 5 may be changed, some operations may be omitted, and some operations may be performed simultaneously.
[0049] In operation 510, the system (400) can receive content data of the artist from the user terminal.
[0050] According to one embodiment, the system (400) can receive music content data of an artist planning to expand overseas from a user terminal. The content data may include at least one of the artist's audio files, music videos, performance videos, or promotional images. The system (400) can perform a preprocessing process to extract the artist's musical characteristics from the received content data. The system (400) can standardize the audio sampling rate of the audio files, remove noise, and convert the data into a format suitable for analysis. The system (400) can also receive metadata such as the artist's genre, career history, and existing fan base information along with the content data to establish a foundation for comprehensive analysis. The system (400) can simultaneously receive content data for multiple artists and perform efficient analysis using a batch processing method.
[0051] In operation 520, the system (400) can derive preference characteristics by analyzing market data of the target country.
[0052] According to one embodiment, the system (400) can collect and analyze music market data of a target country to be entered and derive musical characteristics preferred in that country. The system (400) can collect data on music streaming platforms, digital music charts, and social media trends in the target country. From the collected data, the system (400) can statistically analyze musical elements such as tempo, rhythm patterns, chord progressions, melody structures, and instrumentation of music popular in that country. The system (400) can identify patterns of musical characteristics that show high preference in the target market by applying machine learning algorithms. The system (400) can provide more precise market analysis by deriving preference characteristics segmented by genre, age group, and time of day. The system (400) can even identify the fundamental causes of preference characteristics by considering the cultural background, linguistic characteristics, and historical music trends of the target country together.
[0053] In operation 530, the system (400) can generate virtual localized content that reflects preference characteristics in the content using an AI model and simulate expected reactions.
[0054] According to one embodiment, the system (400) can generate virtual localized content by utilizing an artificial intelligence model to reflect the preference characteristics of the target country in the artist's original content. The system (400) can synthesize a modified version by using a deep learning-based music generation model that applies the rhythm patterns, chord progressions, and instrument arrangements of the target market while maintaining the artist's unique timbre and style. The system (400) can calculate a market suitability index by comparing the generated virtual localized content with reference data of the target market. The system (400) can determine that immediate entry is possible if the market suitability index is above a set threshold, and determine that additional content correction is required if it is below the threshold. Based on the simulation results, the system (400) can predict and provide specific performance indicators such as the expected number of streams, the possibility of chart entry, and the size of the target fan base. The system (400) can simultaneously perform simulations for multiple target countries and propose a priority for market entry for each market.
[0055] According to one embodiment, the system (400) can receive content data of at least one artist from a user terminal. The system (400) can derive preference characteristics of a target country by analyzing market data of the target country to be entered. The system (400) can generate virtual localized content by using an artificial intelligence (AI) model to modify the content data to reflect the preference characteristics derived therefrom. The system (400) can control the simulation of expected reactions in the target country based on the generated virtual localized content.
[0056] According to one embodiment, the system (400) may receive content data of at least one artist from a user terminal. The content data may refer to a digital file of a musical work created or performed by the artist. The content data may include at least one of an audio waveform, a melody structure, lyric text, album artwork, or a music video. The system (400) may convert content data of various formats into a unified format to maintain consistency during subsequent analysis. The system (400) may also receive metadata such as the artist's activity history, major genres, and existing fan base distribution from the user terminal to improve the accuracy of content analysis.
[0057] The system (400) can derive preference characteristics of a target country by analyzing market data of the target country to be entered. Market data may refer to information indicating the consumption patterns and popularity rankings of music content distributed in the target country. Market data may include at least one of music charts, streaming platform statistics, social media reactions, or radio broadcast frequency. Preference characteristics may refer to a combination of musical elements that listeners in the target country consider important when selecting music. Preference characteristics may include at least one of tempo speed, rhythmic complexity, melody range, instrumentation density, or lyrical theme. The system (400) can extract statistically significant patterns from the market data by applying a machine learning algorithm. The system (400) can analyze the target market by deriving preference characteristics segmented by genre, age group, and time of day.
[0058] The system (400) can generate virtual localized content by using an artificial intelligence model to modify content data by reflecting preference characteristics derived from the data. The artificial intelligence model may refer to a neural network structure trained to change other attributes while maintaining specific attributes of the input music data. The artificial intelligence model may include at least one of a style transfer network or a conditional generation model. The virtual localized content may refer to music data generated algorithmically without actual recording. The system (400) can apply accompaniment compositions, rhythm patterns, or melody progressions corresponding to the preference characteristics of the target country while preserving the artist's unique timbre and vocal style. The system (400) may provide parameters to adjust the degree of modification, allowing the user to control the level of similarity with the original content. The system (400) may generate multiple modified versions to provide the user with options.
[0059] The system (400) can control the simulation of expected reactions in a target country based on generated virtual localized content. Expected reactions may refer to the preferences and acceptance levels estimated to be shown by listeners in the target country toward the localized content. Simulation may refer to the process of predicting performance through a computer model without actual market launch. The system (400) can calculate a similarity score by comparing the musical characteristics of the generated localized content with reference data from the target market. Based on the calculated similarity score, the system (400) can quantitatively predict at least one of the expected number of streams, the possibility of chart entry, or the possibility of viral spread. The system (400) can visualize the simulation results and provide them in a form that the user can intuitively understand. If the expected reaction exceeds a set threshold, the system (400) can generate a result recommending entry into the target country.
[0061] FIG. 6 is a flowchart illustrating an AI-based artist timbre and style localization method according to one embodiment.
[0062] The operations described through FIG. 6 can be implemented based on instructions that can be stored in memory (e.g., memory (420) of FIG. 4). The order of each operation of FIG. 6 may be changed, some operations may be omitted, and some operations may be performed simultaneously.
[0063] In operation 610, the system (400) can convert content data into STFT to separate vocals and accompaniment.
[0064] According to one embodiment, the system (400) can generate a spectrogram in the time-frequency domain by performing a short-time Fourier transform on the received artist's content data. The system (400) can input the generated spectrogram into a deep learning-based source separation network to estimate a mask that distinguishes between vocal components and accompaniment components. The system (400) can separate vocals and accompaniment with high accuracy by utilizing a modern source separation model such as the U-Net architecture or Demucs. The system (400) can independently extract vocal tracks and accompaniment tracks by applying the estimated binary mask or ratio mask to the original spectrogram. The system (400) can verify the separation performance by evaluating the quality of the separated vocal tracks and accompaniment tracks using an indicator such as the Signal-to-Distortion Ratio. The system (400) can apply post-processing filtering to minimize artifacts that may occur during the separation process.
[0065] In operation 620, the system (400) can generate an artist-specific characteristic vector by extracting MFCC and F0 from the vocal track.
[0066] According to one embodiment, the system (400) can quantify timbre characteristics and vocal techniques by analyzing separated vocal tracks on a frame-by-frame basis. The system (400) can obtain spectral envelope information representing the artist's unique timbre characteristics by extracting Mel frequency cepstrum coefficients. The system (400) can extract contour information of vocal techniques such as vocal pitch change, vibrato, and portamento by applying a fundamental frequency extraction algorithm. The system (400) can compress high-dimensional acoustic features into a low-dimensional latent space by inputting the extracted MFCC and F0 information into a pre-trained encoder model. The system (400) can define the embedding vector generated in the latent space as the artist's unique characteristic vector and utilize it in a subsequent localization process. The system (400) can generate a more stable and representative unique characteristic vector by averaging the characteristic vectors extracted from multiple songs by the same artist.
[0067] In operation 630, the system (400) can extract target style vectors from target country sound sources.
[0068] According to one embodiment, the system (400) can form a reference dataset of top-ranking popular chart tracks collected from market data of a target country. The system (400) can input the reference tracks into a style encoder network to extract musical style information unique to that country. The system (400) can compress macroscopic musical characteristics such as rhythm, groove patterns, textures based on instrument placement, and tempo distribution of the target country's music into a vector form through the style encoder. The system (400) can generate an integrated style vector representing the target market by averaging or clustering style vectors extracted from multiple reference tracks. The system (400) can extract different style vectors by genre and selectively apply the style most suitable for the artist's genre.
[0069] In operation 640, the system (400) can combine a unique characteristic vector and a style vector to synthesize localized content.
[0070] According to one embodiment, the system (400) can simultaneously input the artist's unique characteristic vector and the target country's style vector into a decoder or generator network. The system (400) can synthesize new music that reflects the accompaniment style and atmosphere of the target market while preserving the artist's timbre by utilizing a conditional generation model. The system (400) can control the intensity of localization by adjusting the weights of the style vectors during the generation process. The system (400) can maintain the characteristics of the original artist as much as possible in the vocal part of the synthesized localized content, and apply the instrumentation and rhythm patterns of the target market to the accompaniment part. The system (400) can convert the generated localized content into a time-domain audio waveform through inverse STFT and output it in a listenable form. The system (400) can provide the original content and the localized content simultaneously so that the user can directly compare and evaluate the degree of change.
[0071] According to one embodiment, the system (400) can generate a spectrogram in the frequency domain by performing a Short-Time Fourier Transform (STFT) on the received artist's content data. The system (400) can separate the vocal track and the accompaniment track by estimating a binary mask or ratio mask that distinguishes vocal components and accompaniment components on the spectrogram through a deep learning-based network and applying it. The system (400) can analyze the separated vocal track on a frame-by-frame basis to extract contour information of Mel-frequency cepstral coefficients (MFCC) representing timbre characteristics and Fundamental Frequency (F0) representing vocal technique and vibrato. The system (400) can input the extracted information into a pre-trained encoder model to generate an artist-specific characteristic vector in the form of an embedding vector in a multidimensional latent space. The system (400) can collect multiple reference audio sources belonging to popular charts or specific genres of a target country from the market data of that country. The system (400) can input the collected reference audio sources into a style encoder to extract a target style vector containing rhythmic sensations unique to that country, textures including instrument placement, and tempo information. The system (400) can input the artist-specific characteristic vector and the extracted target style vector into a decoder or generator to control the synthesis of virtual localized content that maintains the artist's timbre while reflecting the accompaniment style and atmosphere of the reference audio sources.
[0072] According to one embodiment, the system (400) can generate a frequency domain spectrogram by performing a short-time Fourier transform on the received artist's content data. A short-time Fourier transform may refer to a technique for dividing an audio signal that changes over time into short time intervals and analyzing the frequency components of each interval. A spectrogram may refer to visualization data that expresses the energy magnitude of each time-frequency point in color or brightness on a two-dimensional plane composed of a time axis and a frequency axis. The system (400) can simultaneously identify the temporal change and frequency distribution of the music signal through the spectrogram.
[0073] The system (400) can separate the vocal track and the accompaniment track by estimating a binary mask or ratio mask that distinguishes vocal components and accompaniment components on a spectrogram through a deep learning-based network and applying it. The deep learning-based network may refer to a neural network model that learns the frequency pattern differences between vocals and accompaniment from a large amount of audio data. The binary mask may refer to a filter composed of values of 0 and 1 that select either vocals or accompaniment for each time-frequency point of the spectrogram. The ratio mask may refer to a filter that enables smoother separation by expressing the energy ratio of vocals and accompaniment at each point as a real value between 0 and 1. The system (400) can generate a spectrogram containing only vocal components and a spectrogram containing only accompaniment components by multiplying the estimated mask by the original spectrogram. The system (400) can restore the audio signal in the time domain by applying an inverse Fourier transform to each separated spectrogram.
[0074] The system (400) can analyze a separated vocal track frame by frame to extract MFCCs representing timbre characteristics and contour information of the fundamental frequency representing vocal technique and vibrato. Frame-unit analysis may refer to a method of dividing a continuous audio signal into short time intervals and calculating independent characteristics for each interval. MFCCs are coefficients extracted through a Mel scale filter bank that mimics human auditory characteristics and can concisely express timbre characteristics. Fundamental frequency refers to the lowest frequency component generated by vocal cord vibration and can represent a key element in determining pitch. Contour information is a curve representing the pattern of how the fundamental frequency changes over time, which can quantify vocal technique such as vibrato and portamento. The system (400) can quantify the artist's unique timbre through MFCCs and identify the characteristics of the vocal style through F0 contour lines.
[0075] The system (400) can input the extracted information into a pre-trained encoder model to generate an artist-specific feature vector in the form of an embedding vector in a multidimensional latent space. The encoder model may refer to a neural network structure trained to convert high-dimensional input data into a low-dimensional compressed representation. The multidimensional latent space may refer to a reduced-dimensional vector space containing only the essential characteristics of the original data. An embedding vector may refer to a numerical array represented as a single point within the latent space, which implicitly contains the artist's musical identity. The system (400) can compress the complex pattern of MFCC and F0 contour lines into a single vector composed of hundreds of real values through the encoder.
[0076] The system (400) can collect multiple reference audio tracks belonging to popular charts or specific genres of a target country from market data of that country. Reference audio tracks may refer to music works that actually show high preference in the target market and represent samples that reflect the tastes of local listeners. The system (400) can automatically retrieve a list of top-ranking audio tracks on weekly or monthly charts through the API of a music streaming platform. The system (400) can increase the relevance of the analysis by prioritizing the selection of reference audio tracks belonging to categories similar to the artist's genre.
[0077] The system (400) can input collected reference audio sources into a style encoder to extract a target style vector containing rhythmic sensations unique to the country, textures including instrument placement, and tempo information. The style encoder may refer to a neural network model trained to extract macroscopic style characteristics of music in the form of a vector. Rhythm may refer to musical elements representing beat accent patterns and regularity of note placement. Texture may refer to the overall acoustic texture created by the timbres and density of multiple instruments played simultaneously. The target style vector may refer to a vector representation that quantifies the common style characteristics of music of the target country. The system (400) can generate a representative vector that reflects the overall market trend by averaging the style vectors extracted from multiple reference audio sources to remove the specificity of individual songs.
[0078] The system (400) can control the synthesis of virtual localized content that reflects the accompaniment style and atmosphere of reference audio sources while maintaining the artist's timbre by inputting the artist's unique characteristic vector and the extracted target style vector into a decoder or generator. The decoder may refer to a neural network structure that restores the original data form from a compressed vector in a latent space. The generator may refer to a neural network model that generates new data according to an input condition vector. The system (400) can constrain the generated music to maintain the artist's timbre by providing the artist's unique characteristic vector as a condition. The system (400) can induce the accompaniment composition and rhythm pattern of the generated music to match the preferences of the target market by providing the target style vector as an additional condition. The system (400) can control the balance between preserving the original identity and reflecting the local style by adjusting the weight ratio of the two vectors. The system (400) can convert the synthesized result into an audio waveform and provide it in a form that the user can listen to and evaluate directly.
[0080] According to one embodiment, the system (400) can generate a chromagram representing the energy distribution of the 12-note scale by performing a Constant-Q Transform (CQT) that transforms the received content data by frequency band. The system (400) can input the generated chromagram into a Hidden Markov Model (HMM) to extract a sequence of the original song's chord progression that changes over time. The system (400) can classify genres by analyzing metadata of top-ranking songs that are above a specified rank on the local charts of a target country. The system (400) can construct a Local Genre Chord Transition Matrix by accumulating and aggregating the number of transitions between all chord pairs occurring within the song data for each genre, and then normalizing the result to have a matrix value representing the conditional probability that a specific subsequent chord will appear after a specific preceding chord. The system (400) may refer to a local scale database that stores scale information used in the target country. The system (400) may identify a first code and a second code that are temporally consecutive within the chord progression sequence of the extracted original song as a transition pair. The system (400) may look up the transition probability value of the identified transition pair in the local genre code transition matrix, and if the probability value is less than a preset threshold probability value, determine the section where the second code is located as an anomaly section that does not match the local popular music style. The system (400) may replace the second code of the determined anomaly section with a substitute code that has the highest conditional probability when connected to the first code on the local genre code transition matrix.The system (400) can control the generation of virtual localized content by performing pitch shifting, which forcibly aligns the pitch of the melody corresponding to the dissonant section to the grid of the local scale database.
[0081] According to one embodiment, the system (400) can generate a chromagram representing the energy distribution of 12 scales by performing CQT, which converts received content data by frequency band. CQT may refer to a conversion technique that analyzes a signal into musically meaningful pitch units by dividing the frequency axis into a logarithmic scale. A chromagram may refer to a representation that visualizes a harmonic structure by condensing frequency energy into 12 semitone classes regardless of octave. Through the chromagram, the system (400) can observe the energy changes over time for each of the 12 scales, such as C, C#, and D. The system (400) can extract only the core information necessary for chord analysis by integrating frequencies with the same note name into one, even if they are in different octaves.
[0082] The system (400) can input the generated chromagram into a hidden Markov model to extract the chord progression sequence of the original song that changes over time. A hidden Markov model may refer to a probabilistic model that infers a sequence of hidden states from observable data. A chord progression sequence may refer to a sequence of chords that appear in chronological order throughout the song. The system (400) can apply an HMM by using the energy pattern of the chromagram as an observation and considering the chords of each time interval as hidden states. The system (400) can determine the most probable chord sequence using a decoding technique such as the Viterbi algorithm. The system (400) can store the extracted chord progression sequence in the form of symbols to be used for future analysis.
[0083] The system (400) can classify genres by analyzing the metadata of top-ranking audio tracks above a specified rank on the local charts of the target country. Metadata may refer to additional information such as genre, artist, and release year included in the audio files. The system (400) can collect metadata for each audio track along with chart data through the API of the audio streaming platform. The system (400) can group the collected audio tracks by genre category, such as pop, rock, hip-hop, and ballad. The types of genres are merely examples and are not limited thereto, and may vary depending on the characteristics of the music market in the target country.
[0084] The system (400) can construct a local genre code transition matrix having matrix values representing the conditional probability that a specific subsequent code will appear after a specific preceding code by accumulating and normalizing the number of transitions between all code pairs occurring within the audio data for each genre. A code pair may refer to a combination of two chords that appear sequentially in time. The number of transitions may refer to the total frequency at which a specific code pair appears in all audio sources of that genre. Normalization may refer to the process of converting absolute frequencies into a probability distribution by dividing them by the total sum. Conditional probability may refer to the probability that a specific subsequent code will follow given a preceding code. The local genre code transition matrix may refer to a two-dimensional array in which rows represent preceding codes and columns represent subsequent codes, with transition probability values stored in each cell. The system (400) can construct transition matrices independent of each genre to model the unique harmonic progression patterns of each genre. The system (400) can utilize the constructed matrix as a music grammar rule for the target market.
[0085] The system (400) may refer to a local scale database that stores scale information used in the target country. A scale may refer to a set of musically preferred pitches in a specific cultural sphere. The local scale database may refer to reference material that stores the structure of scales frequently used in the traditional music or popular music of the target country. The system (400) may include in the database not only Western major or minor scales but also specific regional pentatonic scales, heptatonic scales, or microtone systems. The system (400) may refer to the local scale database and use it as a standard for adjusting the melody during the localization process.
[0086] The system (400) can identify a first chord and a second chord that are temporally consecutive within the extracted original song's chord progression sequence as a transition pair. Temporal consecutiveness may mean that the two chords appear adjacently in the progression order of the song. The first chord may refer to a preceding chord that appears first in the transition pair. The second chord may refer to a succeeding chord that appears immediately after the first chord. The system (400) can sequentially scan the original song's chord progression sequence and extract all adjacent chord pairs as transition pair candidates.
[0087] The system (400) can look up the transition probability value of an identified transition pair in the local genre code transition matrix, and if the probability value is less than a preset threshold probability value, determine that the section where the second code is located is a dissonant section that does not match the local popular music style. Looking up the transition probability value may mean reading the value of the cell where the row corresponding to the first code and the column corresponding to the second code intersect in the matrix. The threshold probability value may mean the minimum probability standard for code transitions that are naturally accepted in the target market. A dissonant section may mean a section of harmonic progression that is likely to feel unfamiliar or alien to listeners in the target country. The system (400) may mark transition pairs whose probability value falls significantly below the threshold as priority targets for modification.
[0088] The system (400) can replace the second code of the identified dissonance section with a replacement code that has the highest conditional probability when connected to the first code on the local genre code transition matrix. The replacement code may represent a choice that conforms to the music grammar of the target market as a chord that can most naturally follow the first code. The system (400) can find the code with the maximum value by comparing all column values of the row corresponding to the first code in the local genre code transition matrix. The system (400) can perform localization of the chord progression by replacing the original second code with the found replacement code. The system (400) can maintain consistency by regenerating the accompaniment composition according to the replaced code.
[0089] The system (400) can control the generation of virtual localized content by performing pitch shifting, which forcibly aligns the pitch of a melody corresponding to a dissonant section to a grid in a local scale database. The grid may refer to a reference scale in which pitches included in the local scale are arranged at regular intervals. Forced alignment may refer to the operation of moving each note of the original melody to the nearest pitch on the grid. Pitch shifting may refer to a signal processing technique that changes the pitch of a note while preserving the timbre and temporal length. The system (400) can analyze the melody notes of the dissonant section to identify pitches not included in the local scale. The system (400) can move the identified notes by a semitone or a smaller unit to match the pitch of the local scale. The system (400) can combine chord substitution and pitch shifting to adjust both the harmony and the melody to conform to the musical conventions of the target market. The system (400) can synthesize the adjusted result into audio to complete the localized content.
[0091] FIG. 7 is a flowchart illustrating an evolution algorithm-based content optimization method according to one embodiment.
[0092] The operations described through FIG. 7 can be implemented based on instructions that can be stored in memory (e.g., memory (420) of FIG. 4). The order of each operation of FIG. 7 may be changed, some operations may be omitted, and some operations may be performed simultaneously.
[0093] In operation 710, the system (400) can assign random variation values to variation parameters such as tempo and pitch.
[0094] According to one embodiment, the system (400) may define a set of controllable transformation parameters for generating virtual localized content. The set of transformation parameters may include playback speed, pitch, energy ratio by frequency band, reverb depth, and compression ratio. The system (400) may generate and apply random variation values within a preset allowable range for each parameter. The system (400) may generate a first group of candidate content having different attribute values by utilizing an initial population generation method of a genetic algorithm. The system (400) may evenly explore the parameter space by setting the distribution of random variation values to a normal distribution or a uniform distribution.
[0095] In operation 720, the system (400) can calculate the market fit score of each candidate content.
[0096] According to one embodiment, the system (400) can extract musical feature vectors from each content included in a first candidate content group. The system (400) can quantify features such as rhythm density, tempo stability, frequency spectrum distribution, and harmonic complexity. The system (400) can calculate the similarity between the extracted feature vectors and reference feature vectors derived from market data of a target country. The system (400) can calculate a market fit score indicating how similar each candidate content is to local market data using the inverse of Euclidean distance or cosine similarity. The system (400) can improve the accuracy of the score calculation by applying weights based on importance to each dimension of the feature vectors. The system (400) can normalize the calculated scores to express them as values between 0 and 1, thereby enabling intuitive comparison.
[0097] In operation 730, the system (400) can set a new standard as the parameter average value of the top score contents.
[0098] According to one embodiment, the system (400) may select some content from a first group of candidate content that ranks high in market conformity scores as a superior candidate group. The system (400) may designate individuals corresponding to the top 20% to 30% of all candidates as a superior candidate group. The system (400) may calculate an arithmetic mean or a weighted average proportional to the score for the variation parameter values of the selected superior candidate group. The system (400) may set the calculated average value as a new reference parameter for generating next-generation candidates. The system (400) may transmit a combination of parameters with superior characteristics to the next generation by mimicking the selection and crossover operations of a genetic algorithm. The system (400) may gradually narrow the search range around the new reference parameter to converge to an optimal value.
[0099] In operation 740, the system (400) can repeat the process of generating next-ranking candidates by assigning a variable value to the new criteria.
[0100] According to one embodiment, the system (400) can generate a group of next-rank candidate content by assigning random variation values to newly set reference parameters. The system (400) can increase the precision of the search by gradually reducing the magnitude of the variation values compared to the previous generation. The system (400) can recalculate the market fit score of operation 720 for the generated next-rank candidates. The system (400) can repeatedly perform the process of selection, averaging, assigning variation values, and score calculation. The system (400) can monitor the optimization progress by recording the highest score and average score for each iteration cycle. The system (400) can prevent excessive calculation time by setting an upper limit on the number of iterations.
[0101] In operation 750, the system (400) can determine the optimal parameter combination when the score reaches the target value.
[0102] According to one embodiment, the system (400) can check at each iteration whether the maximum value of the market fit score calculated during the iteration process reaches a preset target threshold. The system (400) can determine that optimization is complete if the maximum value of the score exceeds the target threshold or if the amount of change in the score over several consecutive generations decreases within a preset convergence range. The system (400) can determine the variation parameter value at the time of optimization completion as the optimal parameter combination. The system (400) can apply the determined optimal parameter combination to the final content generation to output localized content most suitable for the target market. The system (400) can provide the user with optimization process information, such as the highest score achieved along with the optimal parameter combination, the number of iterations, and the convergence time. The system (400) can store the optimization results in a database and use them as initial values for future optimization for similar artists or target markets.
[0103] According to one embodiment, the system (400) may define a set of deformation parameters including tempo, pitch, and volume ratio, which are controllable variables for generating virtual localized content. The system (400) may generate a first group of candidate content having different attribute values by assigning a random offset within a preset error range to each item of the set of deformation parameters. The system (400) may extract a feature vector including rhythm density and frequency spectrum distribution from each content included in the first group of candidate content. The system (400) may calculate the inverse of the Euclidean distance or the cosine similarity between the extracted feature vector and a reference feature vector derived from market data of a target country, and calculate a Market Conformity Score indicating how similar the content is to the local market data. The system (400) can select some content items from the first candidate content group that have a high ranking in market compatibility score as an excellent candidate group. The system (400) can repeatedly perform the process of generating a next-rank candidate content group by calculating the average or weighted average value of the variation parameter values of the excellent candidate group, setting it as a new standard parameter, and then assigning random variation values again. When the maximum value of the market compatibility score calculated during the iteration process reaches a preset target threshold or the amount of change in the score decreases within a preset convergence range, the system (400) can control the system to determine the variation parameter value at that point in time as the optimal parameter combination and apply the determined optimal parameter combination to the final content generation.
[0104] According to one embodiment, the system (400) may define a set of transformation parameters including playback speed, pitch, and energy ratio by frequency band, which are controllable variables for generating virtual localized content. The set of transformation parameters may refer to a set of numerical variables that can be adjusted to change the characteristics of the music content. Playback speed may refer to the speed of a song expressed as the number of beats played per unit time. Pitch may refer to the degree of transposition that shifts the entire melody and chords up or down in semitone units. The energy ratio by frequency band may refer to the relative volume levels of the low, mid, and high frequencies, respectively. The system (400) may set an allowable minimum and maximum value for each parameter to limit the transformation so that it is performed within a musically meaningful range. The system (400) may represent the set of transformation parameters in the form of a vector and use it as input to an optimization algorithm.
[0105] The system (400) can generate a first candidate content group having different attribute values by assigning a random variation value within a preset error range to each item of the set of variation parameters. The error range may refer to the maximum deviation from a reference value that each parameter can deviate from. The random variation value may refer to a random number generated by a random number generator that follows a normal distribution or a uniform distribution. The first candidate content group may refer to a set of various variation versions generated during the initial exploration phase. The system (400) can generate dozens to hundreds of candidate contents by applying different parameter combinations to the original content. The system (400) may set the initial error range relatively large to explore the parameter space widely. The system (400) may render each generated candidate content into an audio file to prepare it as an evaluation target.
[0106] The system (400) can extract a feature vector including rhythm density and frequency spectrum distribution from each content included in the first candidate content group. Rhythm density may refer to the complexity of the rhythm, measured by the number of note events occurring per unit time. Frequency spectrum distribution may refer to the pattern of energy distributed in each frequency band. The feature vector may refer to a multidimensional vector that quantifies and expresses the characteristics of the music content. The system (400) can quantify the musical attributes of each candidate content by applying an automatic feature extraction algorithm. The system (400) can standardize the extracted feature vector to unify the scale of each dimension.
[0107] The system (400) can calculate the inverse of the Euclidean distance or the cosine similarity between the extracted feature vector and the reference feature vector derived from the market data of the target country, and calculate a market fit score indicating how similar the content is to the local market data. The reference feature vector may refer to the average value of feature vectors extracted from popular music tracks in the target country. The Euclidean distance is the straight-line distance between two vectors and may refer to the degree of divergence in the feature space. The inverse of the Euclidean distance may refer to a similarity indicator transformed to have a larger value as the distance is closer. Cosine similarity is the cosine value of the angle formed by two vectors and may refer to an indicator measuring directional similarity. The market fit score is normalized to a value between 0 and 1, and the closer it is to 1, the higher the fit with the target market. The system (400) can calculate the market fit score for each candidate content and rank the performance.
[0108] The system (400) can select some content from the first candidate content group that ranks high in market compatibility scores as a superior candidate group. The superior candidate group may refer to a certain percentage of entities among all candidates that have high suitability with the target market. The system (400) can set the size of the superior candidate group within the range of the top 10% to 30% of all candidates. The system (400) can form the superior candidate group by sorting in descending order based on scores and selecting the top n items. The system (400) can utilize the parameter combinations of the selected superior candidate group as the basis for generating the next generation.
[0109] The system (400) can repeatedly perform the process of calculating the average or weighted average of the variation parameter values of the excellent candidate group, setting them as new reference parameters, and then assigning random variation values to generate a next-rank candidate content group. The average value may refer to the result of arithmetic averaging all parameter values belonging to the excellent candidate group with equal weights. The weighted average value may refer to the average calculated by applying weights proportional to each candidate's market fit score. The new reference parameter may refer to a combination of parameters that serves as the center point for generating the next generation of candidates. The next-rank candidate content group may refer to the second and subsequent generations generated around the new reference parameter. As generations progress, the system (400) can gradually reduce the error range to precisely search around the optimal value. The system (400) can repeat the cyclical process of evaluation, selection, averaging, and generating new candidates for each generation. Through this iterative optimization process, the system (400) can converge to a parameter combination with high market fit.
[0110] The system (400) can control the process so that when the maximum value of the market fit score calculated during the iteration process reaches a preset target threshold or when the amount of change in the score decreases within a preset convergence range, the variation parameter value at that point is determined as the optimal parameter combination and the determined optimal parameter combination is applied to the final content creation. The target threshold is a score value indicating sufficiently high market fit and may represent the success criterion for optimization. The amount of change in the score may represent the difference between the highest scores in consecutive generations. The convergence range may represent a threshold value where further iteration is judged meaningless because the score improvement is insignificant. The optimal parameter combination may represent a set of parameter values that recorded the highest score at the point of goal achievement or convergence condition fulfillment. The system (400) may stop the optimization process if either of the two termination conditions is satisfied. The system (400) may create the final localized content by applying the determined optimal parameter combination to the original content. The system (400) can record and report to the user the number of iterations of the optimization process, the highest score reached, and the time taken.
[0112] According to one embodiment, the system (400) can convert the lyrics text of the content data into a multidimensional vector space to extract an Original Semantic Vector. The system (400) can select word combinations from a word database of the target country's language that have a cosine similarity with the Original Semantic Vector greater than or equal to a preset threshold as a candidate group. The system (400) can generate Adapted Lyrics by extracting information on the duration and accent location of notes from the melody line of the content data, and then selecting text from the selected candidate group in which the number of syllables matches the number of notes and the linguistic accent location is synchronized with the accent location of the melody. The system (400) can extract a Speaker Embedding Vector containing information on the vocal tract's resonance frequency (Formant) and vocal tone (Timbre) from a reference sound source of the target artist. The system (400) can execute an acoustic model that receives input of a sequence of modified text converted into phoneme units and pitch information of a melody, combines this with a speaker embedding vector, and generates a mel-spectrogram representing the energy distribution by frequency band. The system (400) can analyze a dataset of popular songs from a target country to construct a probability distribution model for the average pitch transition slope occurring in the transition interval between notes and the modulation amplitude and period of frequency fluctuations occurring in the note sustain interval.The system (400) can control the generation of virtual localized content by adjusting the numerical values of the pitch contour and energy envelope of the generated mell-spectrogram to converge to the mean value of the probability distribution model, and then converting them into an audio waveform through a vocoder.
[0113] According to one embodiment, the system (400) can convert the lyrics text of the content data into a multidimensional vector space to extract an original semantic vector. Lyric text may refer to data that expresses the linguistic content included in a song in characters. The multidimensional vector space may refer to a mathematical space that expresses the meaning of words or sentences using hundreds of dimensions of real coordinates. The original semantic vector may refer to a vector that compresses and expresses the overall meaning of the original lyrics. The system (400) can encode the lyrics text into a vector by utilizing a pre-trained language model. The system (400) can apply a sentence embedding technique to condense the meaning of the entire lyrics, rather than at the word level, into a single vector. The system (400) can use the extracted original semantic vector as a standard for translation and adaptation work into the local language.
[0114] The system (400) can select word combinations from a word database of the target country language as candidates, wherein the cosine similarity with the original text semantic vector is greater than or equal to a preset threshold. The word database may refer to reference materials that store the vocabulary of the target language and the vector representations of each word. A word combination may refer to multiple words combined to express the meaning of the lyrics. The threshold may refer to a minimum cosine similarity standard to ensure sufficient semantic similarity with the original text. The candidates may refer to a set of translatable texts that satisfy the semantic preservation condition. The system (400) can generate various vocabulary combinations of the target language and convert each into a vector to compare with the original text semantic vector. The system (400) can sort the candidates in order of highest cosine similarity to prioritize the semantically closest expressions.
[0115] The system (400) can generate rewritten text by extracting note duration and stress position information from the melody line of the content data, and then selecting from the selected candidate group texts in which the number of syllables matches the number of notes and the linguistic stress position is synchronized with the stress position of the melody. The melody line may refer to the main melody of a song that represents changes in pitch over time. The note duration may refer to the temporal length during which each note is maintained. The stress position may refer to a point in the melody that is emphasized with relatively large energy or a high pitch. A syllable may refer to a basic unit of pronunciation formed around a single vowel. Linguistic stress may refer to a part of a word or sentence that is naturally pronounced strongly. Synchronization may refer to matching the timing of the two elements. The rewritten text may refer to lyrics written to preserve the meaning of the original song, be translated into the target language, and harmonize syllably with the melody. The system (400) can calculate the number of syllables for each text in the candidate group and compare it with the number of notes of the melody. The system (400) can analyze the natural stress pattern of the text to verify whether it matches the strong beat position of the melody. The system (400) can select the optimal text that satisfies all three conditions of meaning, number of syllables, and stress pattern as the result of the rewriting.
[0116] The system (400) can extract a speaker embedding vector containing vocal tract resonance frequency and vocal tone information from a reference source of a target artist. The reference source may refer to existing recording data used to learn the timbre characteristics of the artist. The vocal tract resonance frequency may refer to the characteristic of the space extending from the human vocal cords to the throat, mouth, and nose emphasizing a specific frequency band. Vocal tone may refer to textures such as brightness, roughness, and softness as timbre characteristics of the voice. The speaker embedding vector may refer to a vector that quantifies and expresses the voice characteristics of a specific speaker. The system (400) can extract the artist's unique voice characteristics from the reference source using a speaker recognition model. The system (400) can utilize the extracted speaker embedding vector as a condition vector to control the timbre during the speech synthesis process.
[0117] The system (400) can execute an acoustic model that receives a sequence of converted text in phoneme units and pitch information of a melody, combines this with a speaker embedding vector, and generates a Mel-spectrogram representing the energy distribution by frequency band. A phoneme may refer to a sound that is the smallest unit for distinguishing meaning. A phoneme sequence may refer to text converted into a sequence of phonetic symbols. Pitch information may refer to the pitch value at each point in time of the melody. A Mel-spectrogram may refer to a spectrogram in which the frequency axis is converted to a Mel scale that reflects human auditory characteristics. The acoustic model may refer to a neural network that predicts acoustic features from linguistic information and prosodic information. The system (400) can specify the content to be pronounced through the phoneme sequence and control the melody through the pitch information. The system (400) can provide the speaker embedding vector as a condition to the acoustic model to induce the generated voice to reproduce the timbre of the target artist. The system (400) can obtain a Mel-spectrogram with energy distributions displayed in the time-frequency plane as the output of the acoustic model.
[0118] The system (400) can analyze a dataset of popular songs from a target country to construct a probability distribution model for the average slope of pitch fluctuation occurring in the transition interval between notes and the amplitude and period of frequency fluctuation occurring in the note sustain interval. The popular song dataset may refer to a collection of sound sources containing songs that are popular in the target country. The transition interval may refer to an intermediate point in time when moving from one note to another. The average slope of pitch fluctuation may refer to the average value of the speed at which the frequency changes in the transition interval. The note sustain interval may refer to the central part of the note where the pitch is maintained relatively constant. Frequency fluctuation may refer to a phenomenon in which the frequency fluctuates periodically, such as vibrato. Amplitude may refer to the maximum frequency deviation of the fluctuation. Period may refer to the time taken for the fluctuation to repeat once. The probability distribution model may refer to a model that mathematically expresses the probability of a specific parameter appearing. The system (400) can calculate a statistical distribution by extracting transition gradients and vibrato characteristics from all sound sources in the dataset. The system (400) can use the constructed probability distribution model as a standard for the natural vocal style of the target market.
[0119] The system (400) can control the generation of virtual localized content by correcting the numerical values of the pitch contours and energy envelopes of the generated Mel-spectrogram to converge to the mean value of the probability distribution model, and then converting them into an audio waveform through a vocoder. Pitch contours may refer to a curve representing the change in fundamental frequency over time. Energy envelopes may refer to a curve representing the change in volume over time. Correction may refer to the process of adjusting values to be closer to a target standard. A vocoder may refer to a neural network-based speech synthesizer that converts the Mel-spectrogram into a time-domain audio signal. An audio waveform may refer to data representing the change in air pressure over time as a digital signal. The system (400) can extract pitch contours from the Mel-spectrogram and analyze their slope and vibration characteristics. If the analyzed value deviates from the mean of the probability distribution model, the system (400) can adjust it toward the mean value. The system (400) can input the corrected Mel-spectrogram into a vocoder to generate an audible audio file. The system (400) can control the generated audio to reflect the vocal style of the target country while maintaining the artist's timbre.
[0121] According to one embodiment, the system (400) can access an open API (Application Programming Interface) of a local music streaming platform in a target country or a digital music chart database to collect audio data and metadata of top-ranked music tracks that have been ranked above a specified rank during a preset period. The system (400) can store the collected data as a Target Market Reference Dataset. The system (400) can calculate respective audio characteristic values by performing beat tracking to detect rhythmic speed, spectral centroid analysis to indicate timbre brightness, and chroma vector analysis to indicate harmonic characteristics on the generated virtual localized content and the music tracks within the Target Market Reference Dataset. The system (400) can calculate the Cosine Similarity between the audio characteristic values of virtual localized content and the mean distribution of audio characteristic values of a target market reference dataset to calculate a market fit index indicating the degree of similarity with local popular music trends. If the calculated market fit index is greater than or equal to a first threshold, the system (400) can recommend immediate entry into a target country and generate an active entry strategy that includes a short-form marketing guide tailored to the social media trends of that country. If the market fit index is less than the first threshold, the system (400) can control the generation of a content correction guide that analyzes the differences with the target market reference dataset and suggests tempo adjustment, volume adjustment of specific instrument sounds, or key changes.
[0122] According to one embodiment, the system (400) may access an open API of a local music streaming platform in a target country or a digital music chart database to collect audio data and metadata of top-ranked songs that have been ranked above a specified rank during a preset period. An open API may refer to a programming interface provided to allow external developers to access the platform's data. A digital music chart database may refer to a system that systematically stores information on the popularity rankings of songs. A preset period may refer to a time range appropriate for reflecting current trends, such as the last few weeks or months. A specified rank may refer to a baseline that guarantees sufficient popularity, such as the top 50 or top 100 of the chart. Audio data may refer to a file that stores the actual sound waveform of a song in a digital format. Metadata may refer to additional information such as song title, artist name, genre, release date, and playback count. The system (400) may download data by communicating with a platform server via an HTTP request.
[0123] The system (400) can store the collected data as a target market reference dataset. The target market reference dataset may refer to a set of sample audio sources representing the music market of a specific country. The system (400) can store the collected audio files and metadata in a structured format in a database. The system (400) can use the stored dataset as a criterion for evaluating the suitability of localized content. The system (400) can periodically update the dataset to reflect the latest trends.
[0124] The system (400) can calculate each audio characteristic value by performing beat tracking to detect the tempo of the rhythm, spectral centroids to indicate the brightness of the timbre, and chroma vector analysis to indicate harmonic features on the generated virtual localized content and the sound sources within the target market reference dataset. Beat tracking may refer to an algorithm that automatically detects beats that are periodically repeated in a music signal. Spectral centroids may refer to a value that quantifies whether the timbre is bright or dark as the center of gravity of the frequency spectrum. Chroma vectors may refer to a vector that represents the harmonic structure as having the energy of each of the 12 musical scales as an element. Audio characteristic values may refer to a set of numerical values that quantitatively represent the rhythm, timbre, and harmony of the music. The system (400) can apply the same feature extraction algorithm to all sound sources in the localized content and the reference dataset. The system (400) can calculate BPM values through beat tracking, quantify the brightness of the timbre through spectrum centroids, and determine the harmonic distribution through chroma vectors. The system (400) can organize the extracted characteristic values into vector forms and use them for similarity comparison.
[0125] The system (400) can calculate a market fit index representing the degree of similarity with local popular music trends by calculating the cosine similarity between the audio characteristic values of virtual localized content and the average distribution of audio characteristic values of a target market reference dataset. The average distribution may refer to a representative vector obtained by averaging the characteristic values of all sound sources included in the reference dataset. The market fit index may refer to an indicator that expresses how well the localized content matches the musical trends of the target market as a value between 0 and 1. The system (400) can generate a reference vector representing typical musical characteristics of the target market by calculating the average of characteristic vectors extracted from all sound sources in the reference dataset. The system (400) can measure directional similarity by calculating the cosine similarity between the characteristic vector of the localized content and the reference vector. The system (400) can define the calculated cosine similarity value as the market fit index. The system (400) can determine that the closer the index is to 1, the higher the fit with the target market.
[0126] If the calculated market suitability index is greater than or equal to a first threshold, the system (400) may recommend immediate entry into a target country and generate an aggressive entry strategy that includes a short-form marketing guide tailored to the social media trends of that country. The first threshold may be a baseline indicating sufficient market suitability, meaning a value that implies entry is possible without further modification. Immediate entry may mean a strategy of launching the content into the target market in its current state without further modification. The short-form marketing guide may mean a promotional plan utilizing short-form video platforms such as TikTok and Instagram Reels. The aggressive entry strategy may mean a plan that proposes aggressive marketing and a quick launch since the quality of the content has been verified. The system (400) may analyze popular content formats and hashtags on major social media platforms in the target country. Based on the analysis results, the system (400) may suggest a method of editing the highlight section of the localized content to a length of 15 to 30 seconds.
[0127] The system (400) can control the generation of a conditional entry strategy that includes a content correction guide suggesting tempo adjustment, volume adjustment of specific instrument sounds, or key change by analyzing the difference with the target market reference dataset when the market fit index is below a first threshold. Difference analysis may refer to the process of finding elements with large deviations by comparing the characteristic values of the localized content with the average values of the reference dataset item by item. Tempo adjustment may refer to modification work that speeds up or slows down the playback speed of the song. Volume adjustment may refer to work that changes the relative volume of a specific instrument or frequency band. Key change may refer to transposition that raises or lowers the pitch of the entire song by a semitone. The content correction guide may refer to specific modification guidelines to make the localized content more suitable for the target market. The conditional entry strategy may refer to a plan that recommends proceeding with market entry after reflecting the proposed modifications. The system (400) can calculate the difference between the localized content and the reference average for each characteristic item. The system (400) can identify major modification targets by sorting items with large differences by priority. The system (400) can provide specific instructions, for example, to increase the BPM by a specific amount if the tempo of the localized content is slower than the reference. The system (400) can recommend re-evaluating after applying the proposed modifications to check if the market fit index improves. The system (400) can include information such as a comparison before and after modification, expected improvement effects, and the time required for modification in the conditional entry strategy document.
[0129] According to one embodiment, the system (400) can construct a target visual dataset by collecting images of individuals that have recorded a specified number of views or more on a media platform in a target country. The system (400) can extract local preferred style attributes, including skin tone, makeup color tone, hair styling, and lighting tone, which are commonly found among the individuals in the dataset, through computer vision technology. The system (400) can re-synthesize images by using an image generation model trained by separating the structure and texture information of an image, while keeping the facial contours and features of the target artist fixed, and replacing the color tone, brightness, and texture of the image with the extracted values of the local preferred style attributes. The system (400) can generate a virtual localized profile image that predicts the appearance of the artist when local styling is applied. The system (400) can calculate the similarity between the feature vector of the generated virtual localization profile image and the average feature vector of the target visual dataset to produce a Visual Preference Score that quantifies how well the image matches the visual trends of the target country. The system (400) can control the generation of visual direction guide information, which is a dataset that can be referenced when producing album jackets or promotional materials, by extracting the main color value (Hex Code), lighting brightness value, and styling keyword from the image with the highest Visual Preference Score.
[0130] According to one embodiment, the system (400) can build a target visual dataset by collecting images of people that have recorded a specified number of views or more on a media platform in a target country. The media platform may refer to an online service where visual content is shared, such as YouTube, Instagram, or TikTok. The specified number may refer to a minimum number of views that guarantees sufficient public popularity. An image of a person may refer to a photo or video frame containing the face of an artist, influencer, actor, etc. The target visual dataset may refer to a set of images representing a visual style preferred in the target country. The system (400) can automatically download images from popular content on the platform using web crawling technology. The system (400) can select only images that clearly contain a person by applying a face detection algorithm. The system (400) can standardize the resolution and color profile of the collected images and store them as a dataset.
[0131] The system (400) can extract local preferred style attributes, including skin tone, makeup shade, hair styling, and lighting tone, which are commonly found among individuals within the dataset, through computer vision technology. Computer vision may refer to artificial intelligence technology that automatically extracts meaningful information from images or videos. Skin tone may refer to a value representing the average color of facial skin as a numerical value in the RGB or LAB color space. Makeup shade may refer to the trend of colors used in lipstick, eyeshadow, blush, etc. Hair styling may refer to characteristics of a hairstyle, such as hair length, shape, and color. Lighting tone may refer to the color temperature and brightness distribution of the entire image. Local preferred style attributes may refer to a combination of visual characteristics that are popularly preferred in the target country. The system (400) can calculate the average color by accurately separating the skin area through facial landmark detection. The system (400) can automatically segment the makeup area and analyze the color distribution of that part. The system (400) can extract shape features of the hair area and calculate a color histogram. The system (400) can quantify lighting characteristics by measuring the color temperature of the entire image and analyzing the brightness histogram. The system (400) can derive typical values of the local preferred style by calculating the mean and variance of attribute values extracted from all images in the dataset.
[0132] The system (400) can re-synthesize an image by using an image generation model that has been trained to separate and learn morphological information and textural information of an image, while keeping the face contour and facial feature shape information of the target artist fixed, and replacing the color tone, brightness, and texture of the image with the extracted values of local preferred style attributes. Morphological information may refer to geometric structures such as the contour of the face and the position and size of the eyes, nose, and mouth. Textural information may refer to visual appearances such as color, brightness, and fine patterns of the surface. The image generation model may refer to a deep learning model trained to independently control shape and texture, such as a style transfer network or a generative adversarial network. Face contour may refer to the shape of the outline of the face and the jawline. Facial feature shape may refer to the size and arrangement of the eyes, nose, mouth, and ears. Color tone may refer to the color distribution and saturation of the entire image. Brightness may refer to the degree of lightness and darkness of the image. Texture may refer to the texture of the skin or the detailed pattern of the hair. Re-synthesis may refer to the process of generating an image by combining separated elements into a new combination. The system (400) can input the artist's original image into an encoder to separate shape information and texture information into independent latent vectors. The system (400) can maintain the shape information vector as is and replace only the texture information vector with a local preferred style attribute value. The system (400) can input the modified vector combination into a decoder to generate a new image. Through this process, the system (400) can create an image in which only the color and style are localized, while the artist's facial structure remains unchanged.
[0133] The system (400) can generate a virtual localized profile image that predicts how an artist would look when local styling is applied. The virtual localized profile image is an image of the artist generated algorithmically without actual shooting, and may represent a result that reflects the visual preferences of the target country. The system (400) can output the image obtained through the preceding re-synthesis process as a profile image. The system (400) can provide options by generating multiple versions of the localized image with varying style intensities. The system (400) can evaluate the quality of the generated images and filter out unnatural results.
[0134] The system (400) can calculate a visual preference score that quantifies how well the image matches the visual trends of the target country by calculating the similarity between the feature vector of the generated virtual localization profile image and the average feature vector of the target visual dataset. A feature vector may refer to a representation of the visual characteristics of an image as a numerical array of hundreds of dimensions. An average feature vector may refer to a representative vector obtained by averaging the feature vectors extracted from all images in the target visual dataset. The visual preference score is normalized to a value between 0 and 1, and a value closer to 1 may indicate an indicator of high alignment with the visual trends of the target country. The system (400) can extract the feature vector of the localization profile image using a pre-trained image feature extraction network. The system (400) can also extract feature vectors for all images in the target visual dataset using the same network and calculate their average. The system (400) can define the visual preference score by calculating the cosine similarity or the inverse of the Euclidean distance between the feature vector of the localization image and the average feature vector. The system (400) can select the optimal version by comparing the scores of each version of the localized image when multiple versions of the image have been generated.
[0135] The system (400) can control the generation of visual direction guide information, which is a dataset that can be referenced when producing album jackets or promotional materials, by extracting major color values, lighting brightness values, and styling keywords from the image with the highest visual preference score. Major color values may refer to values expressed in hexadecimal codes of colors that appear predominantly in the image. Lighting brightness values may refer to the average brightness of the entire image measured as a value between 0 and 255. Styling keywords may refer to words that describe the characteristics of makeup, hair, and clothing in text. Visual direction guide information may refer to specific visual guidelines that can be referenced during actual shooting or design work. The system (400) can apply a color clustering algorithm to extract three to five major colors of the image and record their respective Hex Codes. The system (400) can quantify the average brightness and contrast by analyzing the brightness histogram of the image. The system (400) can use an image classification model to automatically label makeup styles, hairstyles, and clothing types and extract them as keywords. The system (400) can organize the extracted information into a structured document format and provide it to the user.
[0137] According to one embodiment, the system (400) can simultaneously analyze market data for a plurality of target countries to derive preference characteristics for each country, and then pre-evaluate the suitability between the artist's content data and the preference characteristics for each country to calculate an entry priority score. The system (400) can measure the basic suitability for each country by calculating the cosine similarity between the musical feature vector extracted from the artist's original content and the standard feature vector of each target country. The system (400) can calculate a final entry priority score by applying the music market size of each country, the acceptance of Korean music, and the digital music consumption rate as weights to the basic suitability. The system (400) can extract the top-ranked countries by sorting the calculated entry priority scores in descending order.
[0138] The system (400) can select the top three countries with the highest calculated entry priority scores as priority target markets and generate differentiated localization strategies for each market in parallel. The system (400) can simultaneously execute independent localization processes for each selected country. For the first country, the system (400) can generate first localization content by applying the preference characteristic vector of that country, and simultaneously generate second localization content by applying the preference characteristic vector of the second country for the second country, and generate third localization content by applying the preference characteristic vector of the third country for the third country. The system (400) can generate multiple localization versions optimized for each country by applying different tempo adjustment values, chord progression patterns, and instrument arrangement configurations for each country. For each localization version, the system (400) can individually calculate and evaluate a market fit index based on the market reference dataset of that country.
[0139] The system (400) can establish a localization strategy with minimal modification for a country where the artist's original content already shows a high degree of similarity to the preference characteristics of that country. The system (400) can classify a market as a market for uniqueness preservation if the similarity between the original content and the target country's preference characteristics is greater than or equal to a preset uniqueness preservation threshold. The uniqueness preservation threshold may refer to a minimum similarity standard that allows for market acceptance while maintaining the identity of the original. For markets for uniqueness preservation, the system (400) can limit the adjustment range of modification parameters to 50% or less compared to the general market. For example, the system (400) adjusts the tempo within a range of ±20% relative to the original in the general market, but only within a range of ±10% in markets for uniqueness preservation. The system (400) can adjust only superficial elements such as mixing balance or timbre correction without changing core musical elements such as chord progression or melody structure. The system (400) can provide content that is finely tuned to suit the listening environment of the target market while preserving the artist's unique musical identity as much as possible.
[0141] According to one embodiment, the system (400) can identify changes in music trends over the past six months by analyzing market data of a target country in a time series. For example, the system (400) can collect music chart data of the target country on a monthly basis and extract musical characteristic values of the top 100 songs at each point in time. The system (400) can calculate monthly average values for key parameters among the extracted characteristic values, including tempo, energy level, acoustic ratio, and danceability index. The system (400) can construct time series data by listing the monthly average values of each parameter in chronological order. The system (400) can calculate the slope of each parameter by applying linear regression analysis to the constructed time series data. The slope may represent the rate of change of the corresponding parameter value over time.
[0142] The system (400) determines whether the popularity of a specific musical characteristic is on an upward or downward trend, and can reflect the characteristic that is on an upward trend more strongly during the localization process. The system (400) can determine that the parameter is on an upward trend if the calculated slope is positive and its absolute value is greater than or equal to a preset trend threshold. The trend threshold may refer to the minimum slope value that indicates a statistically significant change; for example, an upward trend can be defined as a case where the monthly average growth rate is 5% or more. The system (400) can determine that the parameter is on a downward trend if the slope is negative and its absolute value is greater than or equal to the trend threshold. The system (400) can assign a trend reinforcement weight to the parameter determined to be on an upward trend. The trend reinforcement weight is a value greater than 1 and can be set within the range of 1.2 to 1.5 in proportion to the magnitude of the slope. When generating localization content, the system (400) can set the value obtained by multiplying the average characteristic value of the target market by the trend reinforcement weight as the target value. The system (400) can set the target tempo to 156 BPM, for example, if the current average tempo of the target country is 120 BPM and the tempo trend reinforcement weight is 1.3.
[0143] The system (400) can specifically process the tempo parameter as follows. The system (400) can obtain time series data such as, for example, 110 BPM in January, 112 BPM in February, 115 BPM in March, 118 BPM in April, 122 BPM in May, and 126 BPM in June by calculating the monthly average tempo of the chart tracks in the target country over the past six months. The system (400) can apply linear regression to this data to confirm a trend of increasing by approximately 3.2 BPM per month on average. The system (400) can determine that although this increase rate corresponds to approximately 2.8% compared to the previous month and falls short of the trend threshold of 5%, the cumulative increase rate over six months is approximately 14.5%, which is a significant increase. In this case, the system (400) can classify the tempo as an upward trend parameter and set the trend strengthening weight to 1.15. The system (400) can calculate a target tempo of approximately 145 BPM by applying a weight to the current average tempo of 126 BPM.
[0144] The system (400) can preemptively reflect future market preferences, rather than just the current state, through this trend prediction-based localization. The system (400) can predict future market characteristics if the current trend continues, considering that the release of localized content will take place approximately 2 to 3 months from now. The system (400) can calculate the expected average characteristic value at the scheduled release time by extrapolating the trend slope and use this as the localization target value. For example, if the average tempo in June is currently 126 BPM and is trending upward by an average of 3.2 BPM per month, the system (400) can estimate the expected average tempo in September, 3 months later, to be approximately 136 BPM. By applying a trend reinforcement weight to this expected value to set the final target value, the system (400) can generate content optimized for the market environment at the time of release.
[0146] According to one embodiment, the system (400) can identify the main fan age groups, gender distribution, and geographical locations by analyzing the artist's existing fan base data. The system (400) can collect fan base data from the artist's social media accounts, streaming platform listener statistics, and official fan club member information. The system (400) can extract information on each individual fan's age, gender, and country of residence from the collected data. The system (400) can classify age into age groups of the teens, 20s, 30s, and 40s and older, and calculate the proportion of fans belonging to each group. The system (400) can classify gender into male, female, and other, and calculate the proportion of each gender. The system (400) can aggregate geographical locations by country to identify the top 10 countries where fans are most distributed. The system (400) can identify a distribution such as, for example, that 45% of the total fans are in their 20s and 30% are in their teens, that by gender 68% are female and 32% are male, and by geography 60% are Korean, 15% are Japanese, and 10% are American. The system (400) can define the segment with the highest proportion as the core fan base based on the analysis results.
[0147] The system (400) can identify groups within a target country that have demographic characteristics similar to the artist's existing fan base and can prioritize reflecting the musical characteristics preferred by those groups. The system (400) can collect demographic segmentation data provided by music streaming platforms in the target country. The system (400) can obtain a list of songs preferred by each population group by segmenting music chart data in the target country by age group and gender. If the artist's core fan base is identified, for example, as women in their 20s, the system (400) can select the top 100 songs most streamed by female listeners in their 20s in the target country. The system (400) can organize the selected songs into a segment-specific reference dataset. The system (400) can extract musical characteristics, including tempo, timbre, chord progression, energy level, and acoustic ratio, from each song in the segment-specific reference dataset. The system (400) can model the distribution of musical characteristics preferred by the female segment in their 20s by calculating the mean and standard deviation of the extracted characteristic values. The system (400) can set the modeled segment preference characteristics to a higher priority than the general market preference characteristics.
[0148] The system (400) can determine the final target characteristic by weighting the segment preference characteristic and the general market preference characteristic when generating localized content. The system (400) can apply a weight of 0.7 to the segment preference characteristic and a weight of 0.3 to the general market preference characteristic. The weight values are merely examples and are not limited to this, and can be adjusted according to the artist's target strategy. For example, if the average preference tempo of the 20s female segment is 100 BPM and the average tempo of the general market is 80 BPM, the system (400) can calculate the final target tempo as (100 × 0.7) + (80 × 0.3) = 94 BPM. The system (400) can apply the calculated target characteristic values to the adjustment of each parameter in the localization process. If a specific instrument arrangement appears frequently in the segment-specific reference dataset, the system (400) can increase the proportion of that instrument in the accompaniment composition of the localized content. The system (400) can increase the volume of synthesizer tracks in the mixing of localized content by 15% compared to the default setting if, for example, it is analyzed that synthesizer sounds are prominently used in 80% of the sound sources preferred by women in their 20s.
[0150] According to one embodiment, the system (400) may conduct an online test on a panel of actual listeners in a target country regarding the generated virtual localized content. The system (400) may recruit listeners residing in the target country through an online survey platform or a specialized music test service. The system (400) may prioritize selecting participants from the recruited listeners who match the demographic characteristics of the previously identified core target segment. The system (400) may establish a test environment where the localized content can be streamed by providing the selected participants with a unique access link. The system (400) may set the number of test participants to at least 50 to ensure statistical significance. The number of participants is merely an example and is not limited to this, and may be adjusted according to budget and time constraints.
[0151] The system (400) can stream localized content to a small sample group and collect playback completion rates, repeat listening counts, and preference ratings. The system (400) can track each participant's playback session and calculate the playback completion rate as the ratio of the actual listened-to section to the total song length. The system (400) can record in seconds whether the participant listened to the song to the end, and if they stopped in the middle, at what point they stopped. The system (400) can count the number of times the same participant plays the same song and tally it as the repeat listening count. After listening is complete, the system (400) can ask the participant to rate their preference on a 5-point or 10-point scale and collect the response. The system (400) can additionally provide a timestamp-based feedback interface that allows users to mark a specific section as 'like' or 'dislike'. The system (400) can store all collected data in a database and manage it by structuring it for each participant.
[0152] The system (400) can recalibrate the market fit index based on collected actual listener response data and perform additional localization work if necessary. The system (400) can calculate the actual listening completion rate by averaging the playback completion rates of all participants. For example, if 38 out of 50 participants listened to the song to the end and 12 dropped out in the middle, the system (400) can calculate the playback completion rate as 76%. The system (400) can calculate the average repeat listening index per participant by calculating the average number of repeat listenings. The system (400) can derive the listener satisfaction score by calculating the mean and standard deviation of the preference ratings. The system (400) can calculate the empirical market fit index by applying a weight of 0.4 to the playback completion rate, 0.3 to the repeat listening index, and 0.3 to the preference rating. The weight values are merely examples and are not limited to this, and can be adjusted according to the evaluation criteria. The system (400) can compare the calculated empirical market fit index with the algorithm-based market fit index calculated in advance. If the difference between the two indices exceeds a preset threshold for correction, the system (400) may determine that the algorithm's prediction model needs to be readjusted. The threshold for correction may be set to a value indicating substantial discrepancy, such as 0.15. If the empirical index is lower than the algorithm index, the system (400) may adjust the method of calculating similarity between the target market reference dataset and the localized content or modify the weights of specific features. The system (400) may regenerate the localized content using the corrected model and output an improved version.
[0153] If panel test results show a high drop-off rate in a specific section, the system (400) can modify the melody or rhythm of that section to generate an improved version. The system (400) can collect the playback stop times of all participants in seconds and map the distribution on the timeline of the song. The system (400) can divide the entire song into segments of 10 seconds and count the number of times drop-off occurred for each segment. If the number of drop-offs in a specific segment is 20% or more of the total number of participants, the system (400) can identify that section as a high drop-off rate section. For example, if 12 out of 50 people stopped playback between 1 minute 30 seconds and 1 minute 40 seconds, the system (400) can mark this section as a problem section. The system (400) can analyze the audio data of the identified high drop-off rate section to extract the musical characteristics of that section. The system (400) can quantify the melody complexity, rhythm change rate, energy level, and harmonic progression of the corresponding section.
[0154] The system (400) may apply crossfade processing at the segment boundaries to naturally connect the modified segment with the rest of the original content. Crossfade may refer to a technique where one side fades out and the other side fades in to smoothly transition between two audio segments. The system (400) may implement a smooth transition by fading out the original from 2 seconds before the start of the modified segment and fading in the modified segment. The system (400) may generate an improved version of the modified content and test it again in a small panel test to verify the improvement effect. The system (400) may measure whether the dropout rate of the corresponding segment of the improved version has decreased compared to the initial version, and if the decrease rate is 30% or more, it may determine that the modification was effective.
[0155] The system (400) can improve the accuracy of the algorithm's prediction through this empirical feedback loop. The system (400) can accumulate actual listener response data obtained from each test cycle as training data. The system (400) can calculate the error between the market fit predicted by the algorithm and the actual measured listener response, and update the parameters of the prediction model in a direction that minimizes this error. The system (400) can learn patterns in which a specific combination of musical characteristics triggers a specific listener response by applying a supervised learning method. For example, the system (400) can learn a relationship in which the churn rate increases sharply when the melody complexity in the target market exceeds a threshold, and can apply this as a pre-constraint in future localization work. When the accumulated test data is secured to a certain scale or larger, the system (400) can retrain the deep learning model to gradually improve the prediction accuracy.
[0157] According to one embodiment, the system (400) can extract musical elements that receive positive evaluations in a target country by analyzing past review data of major music critics, influencers, and radio DJs in the target country using text mining techniques. The system (400) can collect music review texts from major music media websites, music-specialized blogs, YouTube channels, or podcast platforms in the target country. The system (400) can select the top 50 critics with high influence indices in the target country, influencers with more than 100,000 music-related followers, and DJs of radio programs in the top 20 listenership ratings. The selection criteria figures are merely examples and are not limited thereto, and can be adjusted according to the market size of the target country. The system (400) can automatically collect review texts written over the past two years using web crawling technology. The system (400) can automatically translate the collected text from the target country's language into English or use a natural language processing model capable of directly processing the language.
[0158] The system (400) can extract keywords referring to musical elements from the collected review text. Musical element keywords may refer to terms indicating instrument names, musical techniques, genre styles, and arrangement characteristics. The system (400) can detect the occurrence of the corresponding terms in the text by constructing a predefined dictionary of musical terms. The system (400) can automatically identify music-related noun phrases from the review text by applying a noun extraction algorithm. The system (400) can extract sentences surrounding the keywords to analyze the context in which each identified musical element keyword is mentioned within the review text. The system (400) can identify evaluation expressions regarding the corresponding elements by collecting adjectives and adverbs within a range of five words before and after the keywords.
[0159] The system (400) can score the extracted elements through sentiment analysis and reinforce the elements with high scores in the localized content. The system (400) can apply a sentiment analysis model to sentences associated with each musical element keyword. The sentiment analysis model may refer to a pre-trained natural language processing model that determines whether the emotional tone of the text is positive or negative and quantifies its intensity. The system (400) can calculate a sentiment score between -1 and +1 for each sentence through the sentiment analysis model. A sentiment score closer to +1 indicates a very positive evaluation, closer to -1 indicates a very negative evaluation, and closer to 0 indicates a neutral evaluation. The system (400) can collect sentiment scores for all sentences in which a specific musical element keyword is mentioned. For example, if the keyword 'acoustic guitar' is mentioned a total of 200 times in 150 reviews, the system (400) can calculate the sentiment score for each of those 200 sentences.
[0160] The system (400) can calculate the average sentiment score and mention frequency for each musical element to calculate the positive evaluation index for each element. The system (400) can classify elements with an average sentiment score of +0.5 or higher and a mention frequency of 10% or higher of the total reviews as high-rated music elements. The threshold is merely an example and is not limited to this, and can be adjusted according to the purpose of analysis. For example, the system (400) can identify an element called 'String Section' as a high-rated element if it records an average sentiment score of +0.72 and a mention frequency of 18%. The system (400) can select the top 10 high-rated music elements by sorting the list in order of highest positive evaluation index. The system (400) can designate the selected top elements as targets to be reflected first when creating localized content.
[0162] According to one embodiment, the system (400) can analyze playlist composition patterns on major streaming platforms in a target country. The system (400) can collect public playlist data through the open API of the streaming platform most widely used in the target country. The system (400) defines playlists with 100,000 or more followers as popular playlists and can select the top 100 of such playlists as targets for analysis. The selection criteria are merely an example and are not limited thereto, and can be adjusted according to the size of the target market. The system (400) can collect the list of songs included in each playlist and the order of arrangement of the songs. The system (400) can additionally obtain metadata and audio characteristic data of each song included in the playlist through the platform API. Audio characteristic data may refer to quantified musical characteristics including tempo, energy, danceability, balance, acousticness, and instrumentality.
[0163] The system (400) can build a playlist placement optimization model by learning musical similarity patterns between songs placed consecutively in a popular playlist. The system (400) can define two songs that are sequentially consecutive within each playlist as a transition pair. For example, the system (400) can extract the 5th and 6th songs of a playlist as a transition pair. The system (400) can form a dataset of thousands of transition pairs by extracting transition pairs from all selected playlists. For each transition pair, the system (400) can calculate the difference in audio characteristic values between the preceding song and the succeeding song. For example, if the tempo of the preceding song is 120 BPM and the tempo of the succeeding song is 125 BPM, the system (400) can record the tempo difference as +5 BPM.
[0164] The system (400) can perform statistical analysis on the characteristic difference values of the collected transition pair dataset. The system (400) can calculate the mean and standard deviation of the difference values for each characteristic to derive a typical range of variation allowed within the playlist. For example, if the system (400) analyzes that the mean of the tempo difference is +2 BPM and the standard deviation is 8 BPM, it can identify a pattern that, generally, consecutive songs in the playlist do not change their tempo significantly and mostly change within the range of ±16 BPM. If the system (400) analyzes that the mean of the energy level difference is -0.02 and the standard deviation is 0.15, it can confirm that when songs are placed, the energy tends to remain generally constant or decrease slightly. The system (400) can model the distribution of difference values for all audio characteristics as a normal distribution and store the allowable range of variation for each characteristic in the form of a probability distribution. The system (400) can define the characteristic difference distribution model constructed in this way as a playlist placement optimization model.
[0165] The system (400) can predict which popular playlist in the target country the generated localized content is suitable for and fine-tune the musical characteristics so that they blend naturally with other songs in the playlist. The system (400) can extract audio characteristic values of the generated localized content. The system (400) can calculate similarity by comparing the extracted characteristic values with the average characteristic values of previously collected popular playlists. The system (400) can define the average of the characteristic values of all songs included in each playlist as the representative characteristic vector of the playlist. The system (400) can calculate the Euclidean distance between the characteristic vector of the localized content and the representative characteristic vector of each playlist. The system (400) can select the top three playlists with the shortest distance as the target playlists most suitable for the localized content.
[0166] After all adjustment operations are completed, the system (400) can compare the modified localized content with the songs in the target playlist again. The system (400) can use the previously established playlist placement optimization model to verify whether the characteristic difference between the modified content and the preceding and succeeding songs is within an acceptable range when the modified content is inserted at a specific position in the playlist. For example, the system (400) can calculate the characteristic difference between the 14th song and the modified content, and the characteristic difference between the modified content and the 16th song, respectively, assuming that the modified content is inserted between the 15th and 16th songs in the playlist. The system (400) can determine that the adjustment is successful if the calculated difference values are within the acceptable distribution range of the playlist placement optimization model.
[0167] The system (400) can generate content optimized for the recommendation algorithm of a streaming platform through this. The system (400) can utilize the principle that the recommendation algorithm of a streaming platform recommends similar songs based on the musical consistency of the songs within a playlist. By adjusting the localized content to be similar to the average characteristics of a popular playlist, the system (400) can increase the likelihood of it being recommended to users who listen to the playlist. The system (400) can generate a playlist marketing guide that includes a list of recommended target playlists, the number of followers for each playlist, and the estimated number of listeners to reach, along with the adjusted localized content. The system (400) can verify the prediction accuracy and continuously improve the model by tracking which playlists the localized content was actually added to after its release.
[0169] According to one embodiment, the system (400) can calculate the expected return on investment along with market performance indicators expected from simulation results. The system (400) can estimate the expected market performance indicators based on the market fit index calculated for localized content. The market performance indicators may include the expected monthly streaming count, the expected chart entry ranking, and the expected viral spread index. The system (400) can construct a correlation model between the market fit index and actual market performance through the analysis of past data. The system (400) can inversely calculate the market fit index for music by Korean artists released in the target country over the past year and collect data on the actual streaming counts of said music. The system (400) can represent the collected data pairs as a scatter plot and apply regression analysis to derive a prediction function with the market fit index as the independent variable and the monthly streaming count as the dependent variable.
[0170] The system (400) can establish a relationship in which, for example, when the market fit index is 0.8, an average of 1.5 million streams are recorded per month. If the market fit index of the localized content is calculated to be 0.75, the system (400) can input this into a prediction function to estimate the expected monthly number of streams as approximately 1.2 million. The system (400) can investigate the payout per play rate for each streaming platform in the target country. The payout per play rate may refer to the average amount paid by the platform to the artist per stream. For example, the system (400) can determine that the average payout per play rate for major platforms in the target country is $0.004. The system (400) can calculate the expected monthly streaming revenue by multiplying the expected monthly number of streams by the payout per play rate. The system (400) can estimate a monthly streaming revenue of approximately $4,800 by multiplying 1.2 million by $0.004.
[0171] The system (400) can calculate the ROI for each target country by comprehensively considering the costs incurred for localization work, the marketing budget, the expected streaming revenue, and the potential for performance opportunities. The system (400) can define cost items incurred for localization work and set the unit price for each item. The costs of localization work may include costs for running an AI model, audio engineering costs, costs for translating and rewriting lyrics, and costs for recording vocals in the local language. For example, the system (400) can estimate the total localization cost as $4,300, with $500 for running an AI model, $1,000 for audio mixing and mastering, $800 for rewriting and translating lyrics, and $2,000 for re-recording vocals. The cost items and amounts are merely examples and are not limited to this, and may vary depending on the project scope.
[0172] The system (400) can estimate the budget required for marketing activities in a target country. The marketing budget may include social media advertising costs, influencer collaboration costs, playlist registration promotion costs, and promotional material production costs. The system (400) can calculate the budget by referring to advertising unit price data and standard marketing package information for the target country. For example, the system (400) can set the total marketing budget to $7,200 by allocating $3,000 for social media advertising, $2,000 for influencer collaboration, $1,500 for playlist promotion, and $700 for promotional material production. The system (400) can calculate the total investment cost by adding the localization costs and the marketing budget. The system (400) can calculate the total investment cost to be $11,500 by adding $4,300 and $7,200.
[0173] The system (400) can calculate the expected streaming revenue by extending it to the cumulative revenue over 12 months after launch. The system (400) can apply a monthly attenuation coefficient to reflect the tendency of the streaming pattern of the audio source to decrease over time. The system (400) can assume a pattern of gradual decrease, such as recording 100% streaming in the first month, 80% in the second month, and 65% in the third month. The system (400) can calculate the monthly revenue by multiplying the expected number of streams for each month by the price per play and sum the revenue for 12 months. The system (400) can estimate the cumulative streaming revenue for 12 months to be, for example, about $35,000.
[0174] The system (400) can quantify the potential for performance opportunities and include them as additional revenue items. The system (400) can estimate the probability of a performance invitation based on the market fit index and the size of the K-pop or Korean music performance market in the target country. For example, if the market fit index is 0.75 or higher and monthly streaming exceeds 1 million, the system (400) can set the probability of receiving a performance offer within 12 months to 60%. The system (400) can calculate expected performance revenue by investigating the average performance fee in the target country. If the average performance fee is $15,000 and the performance probability is 60%, the system (400) can calculate the expected performance revenue as $9,000. Expected performance revenue may mean the value obtained by multiplying the performance fee by the probability.
[0175] The system (400) can calculate total expected revenue by summing the expected streaming revenue and expected performance revenue. The system (400) can calculate total expected revenue as $44,000 by summing $35,000 and $9,000. The system (400) can calculate the expected return on investment (ROI) using the following formula. ROI can be calculated using the formula (Total expected revenue - Total investment cost) / Total investment cost × 100%. The system (400) can derive an ROI of approximately 283% by calculating (44,000 - 11,500) / 11,500 × 100%. The system (400) can generate country-specific comparison data by calculating ROI for all candidate target countries in the same way.
[0176] The system (400) can prioritize recommending countries with medium market fit but large market size and high profit potential over countries with high market fit but low profitability due to small market size. The system (400) can map the market fit index and calculated ROI for each target country onto a two-dimensional plane. The system (400) can compare, for example, cases where country A records a market fit of 0.85 and an ROI of 150%, and country B records a market fit of 0.70 and an ROI of 320%. The system (400) can calculate a comprehensive score by assigning weights to the market fit and ROI, respectively. The system (400) can calculate the overall score of Country A as 0.85 × 0.4 + 1.50 × 0.6 = 1.24 and the overall score of Country B as 0.70 × 0.4 + 3.20 × 0.6 = 2.20 by applying a weight of 40% to market fit and 60% to ROI. The weight ratios are merely examples and are not limited to this, and can be adjusted according to the artist's strategic direction. The system (400) can generate a list of priority recommended countries by sorting the countries in order of highest overall score.
[0177] The system (400) can provide decision support information that balances musical suitability and economic feasibility. The system (400) can generate a detailed report for each target country. The report may include a market suitability index, estimated monthly streaming counts, 12-month cumulative streaming revenue, probability of performance opportunities and expected revenue, details of total investment costs, calculated ROI, and an overall recommendation score. The system (400) can include visualized comparison charts in the report to enable an intuitive understanding of trade-offs between various countries. For example, the system (400) can provide information that a country with very high market suitability but low ROI is musically optimal but may have limited commercial revenue. Conversely, the system (400) can provide information that a country with medium market suitability but high ROI requires some musical adjustments but can expect a high return on investment. The system (400) can provide a weighting adjustment interface to allow the user to choose whether to place more weight on musical priority or economic priority.
[0179] Although the embodiments have been described above with reference to the limited drawings, those skilled in the art can apply various technical modifications and variations based on the above. For example, suitable results may be achieved even if the described techniques are performed in a different order than described, and / or if the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.
Claims
Claim 1 An AI-based system for generating overseas expansion strategies based on artist characteristics includes a memory and a processor for storing instructions, wherein, when the instructions are executed by the processor, the system receives content data of at least one artist from a user terminal, analyzes market data of a target country to be entered to derive preference characteristics of that country, generates virtual localized content by transforming the content data to reflect the derived preference characteristics using an AI model, controls the system to simulate the expected response in the target country based on the generated virtual localized content, performs a Constant-Q Transform (CQT) to transform the received content data by frequency band to generate a chromagram representing the energy distribution of the 12-note scale, inputs the generated chromagram into a Hidden Markov Model (HMM) to extract the chord progression sequence of the original song that changes over time, classifies genres by analyzing metadata of top-ranking songs above a specified rank on the local chart of the target country, and all occurrences within the music data for each genre After accumulating and normalizing the number of transitions between chord pairs, a Local Genre Chord Transition Matrix is constructed having matrix values representing the conditional probability of a specific subsequent chord appearing after a specific preceding chord; a local scale database storing scale information used in the target country is referenced; a first chord and a second chord that are temporally consecutive within the chord progression sequence of the extracted original song are identified as a single transition pair; and the transition probability value of the identified transition pair is queried from the Local Genre Chord Transition Matrix.A system that controls the generation of virtual localized content by determining, when the corresponding probability value is less than a preset threshold probability value, that the section where the second code is located is an anomaly section that does not conform to the local popular music style, and by replacing the second code of the determined anomaly section with an alternative code that has the highest conditional probability when connected to the first code on the local genre code transition matrix, or by performing pitch shifting to forcibly align the pitch of the melody corresponding to the anomaly section with the grid of the local scale database. Claim 2 In claim 1, when the instructions are executed by the processor, the system performs a Short-Time Fourier Transform (STFT) on the received artist's content data to generate a frequency domain spectrogram, estimates a binary mask or ratio mask that distinguishes vocal components and accompaniment components on the spectrogram through a deep learning-based network, and applies it to separate the vocal track and the accompaniment track, respectively; analyzes the separated vocal track on a frame-by-frame basis to extract contour information of Mel-frequency cepstral coefficients (MFCC) representing timbre characteristics and Fundamental Frequency (F0) representing vocal technique and vibrato, inputs the extracted information into a pre-trained encoder model to generate an artist-specific characteristic vector in the form of an embedding vector in a multidimensional latent space, and from the market data of the target country, the popular chart of that country or A system for controlling the synthesis of virtual localized content that maintains the artist's timbre while reflecting the accompaniment style and atmosphere of the reference audio, by inputting the collected reference audio into a style encoder to extract a target style vector containing texture and tempo information including rhythmic sense and instrument placement unique to the country, and inputting the artist's unique characteristic vector and the extracted target style vector into a decoder or generator. Claim 3 delete Claim 4 In claim 1, the instructions, when executed by the processor, define a set of transformation parameters including tempo, pitch, and volume ratio, which are controllable variables for the system to generate the virtual localized content; generate a first candidate content group having different attribute values by assigning a random offset within a preset error range to each item of the set of transformation parameters; extract a feature vector including rhythm density and frequency spectrum distribution from each content included in the first candidate content group; calculate the reciprocal of the Euclidean distance or cosine similarity between the extracted feature vector and a reference feature vector derived from the market data of the target country, and calculate a Market Conformity Score indicating how similar the content is to the local market data; select some content among the first candidate content group that ranks high in the Market Conformity Score as a superior candidate group; and the average value of the transformation parameter values of the superior candidate group Alternatively, a system that repeatedly performs the process of generating a next-ranked candidate content group by calculating a weighted average value and setting it as a new standard parameter, and then assigning a random variation value again, and controls the system to determine the variation parameter value at a given point in time as the optimal parameter combination and apply the determined optimal parameter combination to the final content generation when the maximum value of the market conformity score calculated during the above iteration process reaches a preset target threshold or when the amount of change in the score decreases within a preset convergence range. Claim 5 In claim 1, when the instructions are executed by the processor, the system converts the lyrics text of the content data into a multidimensional vector space to extract an Original Semantic Vector, selects word combinations from a word database of the target country's language that have a cosine similarity with the Original Semantic Vector greater than or equal to a preset threshold as a candidate group, extracts duration and accent position information of notes from the melody line of the content data, selects text from the selected candidate group in which the number of syllables matches the number of notes and the linguistic accent position is synchronized with the accent position of the melody to generate Adapted Lyrics, extracts a Speaker Embedding Vector including vocal tract resonance frequency (Formant) and timbre information from the target artist's reference audio source, and the sequence of the Adapted Lyrics converted into phoneme units and the melody An acoustic model is executed to receive pitch information as input and combine it with the speaker embedding vector to generate a Mel-spectrogram representing the energy distribution by frequency band; a dataset of popular songs from the target country is analyzed to construct a probability distribution model for the average pitch transition slope occurring during the transition between notes and the modulation amplitude / period of frequency fluctuations occurring during the note sustainment section; and after adjusting the numerical values of the pitch contour and energy envelope of the generated Mel-spectrogram to converge to the average value of the probability distribution model,A system that controls the generation of the aforementioned virtual localized content by converting this into an audio waveform through a vocoder.
Citation Information
Patent Citations
Method for providing analysis results of music related content and computer-readable recording medium
KR1020220043489A
Method, device and system for attracting overseas tourists based on regional content and providing platform services for exporting local products
KR102787515B1