AI-driven system and method for generating style fingerprints and compositions
Patent Information
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- MONARRCH INC
- Filing Date
- 2025-07-01
- Publication Date
- 2026-06-03
AI Technical Summary
The music industry faces challenges such as commoditization of music, declining revenue streams, ineffective rights management, and stifling of creativity due to traditional models being unsustainable in the digital era, necessitating a paradigm shift for equitable and transparent music ecosystems.
An AI-driven system that creates unique style fingerprints for artists, encapsulating their musical identity, and generates new compositions while preserving their signature style, utilizing a combination of AI models for style transfer and a centralized platform for managing and monetizing these rights.
Empowers artists with transparent rights management, opens new revenue streams, and fosters creativity by enabling accurate identification and monetization of their distinctive style, enhancing the music industry's capacity for innovation.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
AI-DRIVEN SYSTEM AND METHOD FOR GENERATING STYLEFINGERPRINTS AND COMPOSITIONSTECHNICAL FIELD
[0001] The present invention pertains to systems for managing and licensing style transfer rights within the music industry. More specifically, the present invention relates to a system and a method utilizing a combination of artificial intelligence (Al) models to create unique style fingerprints and generating new compositions / artistic outputs using trained style transfer Al model(s) while preserving the originality of the artist's style.BACKGROUND
[0002] The following references to and descriptions of prior proposals or products are not intended to be, and are not to be construed as, statements or admissions of common general knowledge in the art. In particular, the following prior arts discussion does not relate to what is commonly or well known by the person skilled in the art, but assists in the understanding of the inventive step of the present invention of which the identification of pertinent prior art proposals is but one part.
[0003] The music industry is still facing an unprecedented crisis. The rise of digital platforms and streaming services has fundamentally disrupted the traditional revenue models or systems, leaving artists and rights holders struggling to adapt.
[0004] The primary issue is the commoditization of music. The abundance of content and the ease of access through streaming platforms has led to the commoditization of music. The value of individual recordings has diminished, making it difficult for artists to monetize their work effectively.
[0005] Second issue is the declining revenue streams. With physical sales and downloads on a steady decline, streaming revenue has proven insufficient to offset these losses. The prevailing streaming model that is characterized by lower per-stream royalty rates, has left many artists and rights holders financially vulnerable.
[0006] Third issue grapples with ineffective rights management, navigating a complex maze of rights and licenses that increasingly hinders transparency and efficiency. This opacity contributes to lost revenue and undermines artists' control over their intellectual property.
[0007] Fourth issue is stifling of creativity, wherein the artists face pressure to conform to prevailing trends and commercially viable genres. This conformity often limits artistic innovation and expression, compelling musicians to operate within restrictive boundaries rather than exploring new creative territories.
[0008] Furthermore, the traditional music industry models are no longer sustainable in the face of these challenges. A paradigm shift is necessary to ensure the survival and prosperity of artists and rights holders. Effective solutions are needed to foster a more equitable and transparent music ecosystem that supports creative freedom, fair compensation, and sustainable career development for artists.
[0009] Style Transfer Fingerprinting is one such approach. It is the process of using advanced technologies to create digital profiles or fingerprints that captures the unique stylistic elements of an artist's work. This involves analysing various aspects of music, and encapsulate the artist's distinctive musical identity into a structured format.
[0010] Style fingerprints encompass various types that capture distinct creative aspects across different domains.
[0011] Musical style fingerprints include melody, harmony, rhythm, and instrumentation that encapsulates the unique sound and musical identity of an artist.
[0012] Visual style fingerprints capture nuances like colour palette choices, brushstroke techniques, composition, and overall artistic approach, reflecting the individuality of a painter or visual artist.
[0013] Writing style fingerprints focusses on aspects such as vocabulary usage, sentence structure, narrative voice, and thematic preferences, offering insights into an author's literary identity.
[0014] Fashion style fingerprints encapsulate preferences in colour schemes, fabric selections, garment designs, and overall fashion aesthetics, showcasing the distinctive fashion sense and creativity of designers.
[0015] Photography style fingerprints encompass techniques in lighting, framing, subject selection, editing style, and overall photographic vision, revealing the unique visual perspective and artistic expression of photographers.
[0016] Despite the potential benefits of style transfer fingerprinting methods / systems, there currently exists a gap in comprehensive Al-driven systems that effectively utilizes this technology to address the music industry's challenges.
[0017] To address the aforesaid challenges, there is a need for a groundbreaking Al-driven system that aims to redefine rights management in the music industry by harnessing the power of generative Al and style transfer.
[0018] This system may empower artists, revolutionizes revenue streams, and unlocks new realms of creativity.
[0019] Any discussion of the prior art throughout the specification should in no way be considered as an admission that such prior art is widely known or forms part of common general knowledge in the field.SUMMARY
[0020] PROBLEMS TO BE SOLVED
[0021] It may be an advantage to provide a system and a method utilizing a combination of artificial intelligence (Al) models to create unique style fingerprints and generating new compositions or artistic outputs.
[0022] It may be an advantage to provide a system that enables artists to create a distinctive style fingerprint by analyzing their existing work and encapsulating their unique musical identity.
[0023] It may be an advantage to provide a centralized platform for managing and monetizing style transfer rights for artists.
[0024] It may be an advantage to provide a system and method that opens up new avenues for music creation, licensing, and monetization, empowering artists and revolutionizing the music industry.
[0025] It may be an advantage to provide a system, wherein the style fingerprint is registered within a database and associated with the artist's profile for future reference and licensing.
[0026] It may be an advantage to provide a system that is able to generate new compositions and accordingly maintain the artist's signature style.
[0027] It may be an advantage to provide a system, wherein the style fingerprints encapsulate key characteristics of the artist's music, including but not limited to melodic patterns, harmonic progressions, and rhythmic structures.
[0028] It may be an advantage to provide a system that includes a Style Transfer Rights Agency (STRA) which acts as a collective licensing organization for style transfer rights.,maintaining the style transfer registry, managing artist memberships, and facilitating the integration of style transfer data with generative Al platforms through APIs and secure communication protocols.
[0029] It may be an advantage to provide a method that employs the style transfer Al model to create unique style fingerprints, once the artists have registered their music libraries with the STRA.
[0030] It may be an advantage to provide a method, wherein a generative Al platform processes user prompts, scans for references to registered artists or songs and matches them against the style transfer registry.
[0031] It is an object of the present invention to overcome or ameliorate at least one of the disadvantages of the prior art, or to provide a useful alternative.
[0032] MEANS FOR SOLVING THE PROBLEM
[0033] The present invention may be envisioned to a system and a method utilizing a combination of artificial intelligence (Al) models to create unique style fingerprints and generating new compositions or artistic outputs, thereby maintaining their signature style.
[0034] In a first aspect of the present invention, the Al-driven system modules comprise the library integration module for seamless connection to artists' music libraries and receipt of audio files; an audio indexing and preprocessing module that extracts metadata; a feature extraction module that employs advanced techniques to extract relevant features from the music to capture multiple levels of audio features from the artist's music.
[0035] In a second aspect of the present invention, the system uses various Al models, techniques and methods to represent and encode these musical features. The extracted features are encoded into a compressed representation of the artists style ‘fingerprint’ capturing a rich and nuanced stylistic element that closely replicate the artist's original style, whilst allowing for accurate style transfer across diverse forms. The generated stylefingerprint securely stores style fingerprints artefacts encapsulating artist's musical characteristics in a secure and encrypted database, linking to artists' profiles.
[0036] In a third aspect of the present invention, the system includes a composition generation module that utilizes the trained style transfer model to transfers between multiple musical styles or generate novel compositions. This module further synthesizes audio samples that authentically reflect the artist's unique musical traits, providing a comprehensive system for creating and managing artistic style in generative Al applications.
[0037] In another aspect of the present invention, the finger print is either a highdimensional vector or a set of learned model parameters, which serves to direct the generation of new musical content.
[0038] In another aspect of the present invention, the method for creating compositions in an artist's musical style involves a series of steps. First, the artist's music library connects to a system for high-quality audio analysis using a dedicated integration module. Metadata is extracted and processed using an audio indexing and preprocessing module. This step uses techniques like normalization and noise reduction to ensure the audio quality remains consistent. Once the audio is prepared, advanced feature extraction takes place on the music library to capture detailed musical elements. A deep neural network is then trained using these features, to encoded the information by adjusting model settings to minimize differences between the generated audio and the artist's original style. After training, a non- fungible style fingerprint is created to encapsulate the artist's unique musical characteristics.
[0039] In another aspect of the present invention this style fingerprint is securely stored in an encrypted form in a database linked to the artist's profile for future reference and licensing. Finally, the trained style transfer model is used to generate novel compositions that mirror the identified musical traits, enabling the creation of music in the artist's distinctive or signature style.
[0040] In the context of the present invention, the words “comprise”, “comprising” and the like are to be construed in their inclusive, as opposed to their exclusive, sense, that is in the sense of “including, but not limited to”.
[0041] The invention is to be interpreted with reference to the at least one of the technical problems described or affiliated with the background art of the invention. The present aims to solve or ameliorate at least one of the technical problems and this may result in one or more advantageous effects as defined by this specification and described in detail.BRIEF DESCRIPTION OF THE FIGURESFigure 1 illustrates a flow chart of the present method, according to an exemplary embodiment of the present invention;Figure 2 illustrates an architecture diagram of the Al system, according to an exemplary embodiment of the present invention;Figure 3 illustrates a block diagram of the Single Generative Model Fine-tuning approach;Figure 4 illustrates a block diagram of the Multi-modal Hybrid Model approach; andFigure 5 illustrates a flow chart relating to stylometric artefact creation.DESCRIPTION OF THE INVENTION
[0042] Preferred embodiments of the invention will now be described with reference to the accompanying drawings and non-limiting examples.
[0043] Although the invention has been described with reference to specific examples, it will be appreciated by those skilled in the art that the invention may be embodied in many other forms, in keeping with the broad principles and the spirit of the invention described herein.
[0044] The term “artist” specifically within the context of music and entertainment, refers to a person who creates and performs music, whether as a singer, musician, composer, or songwriter. They are recognized for their artistic expression, creativity, and ability toconnect with an audience through their work.
[0045] The term "signature style" refers to a distinctive and recognizable way of expressing oneself or creating art that sets an artist apart from others. It encompasses elements like tone, timbre, dynamics, and spatial characteristics that together create a recognizable auditory identity.
[0046] The term “hash output” refers to a unique identifier for the encoded content.
[0047] The term “model weights” refers to instructions for other generative Al engines to create content that resembles the artist's style and likeness.The term “data formats” refers to the relevant information for integration with existing rights organizations.
[0049] The present invention is directed to a system and a method utilizing a combination of artificial intelligence (Al) models to create unique style fingerprints and generating novel compositions or artistic outputs.
[0050] The system provides a specialized style fingerprint that captures the unique sonic characteristics of an artist. This fingerprint allows for precise representation and recognition of the artist's musical identity by integrating various audio features and metadata into a comprehensive profile.
[0051] This capability ensures accurate identification and differentiation from other artists, enhancing the system's ability to categorize, recommend, and analyze music based on its stylistic attributes.
[0052] The advantages of the system are recited in the below paragraphs.
[0053] Distinctive Style Fingerprint - The system offers a unique style fingerprint that encapsulates an artist's sonic signature, enabling precise representation and recognition of their musical identity.
[0054] New revenue streams: Using the system artists can monetize their style fingerprint whenever it's used in Al-generated music, diversifying income beyond traditional revenue sources like royalties and licensing fees.
[0055] Empowered rights management: The system provides artists with transparency and control over their creative assets, offering visibility into usage and ensuring fair compensation for their artistic contributions.
[0056] Expanded licensing opportunities: The system is able to open up a new licensing market where labels, publishers, and artists can capitalize on the commercial potential of an artist's distinctive style.
[0057] Supports Music Industry Growth: The system enhances the industry's capacity for innovation by encouraging the development of distinct musical identities and pushing creative boundaries.
[0058] The Al driven system comprises a library integration module that integrates the music library(s) and receives audio files; an audio indexing and preprocessing module extracts metadata and audio features from the musical library using preprocessing techniques; a feature extraction module that utilises advanced audio feature extraction techniques to capture plurality of level feature(s) of the artists music; a style transfer training model which harnesses the extracted musical features to train a deep neural network, optimizes parameters to closely match the artist's original style in the generated audio; a style fingerprint generation module that generates a style fingerprint that encapsulates the artist's musical characteristics, a database that stores generated style fingerprints securely and associates the fingerprints with artists' profiles; a composition generation module to create new compositions by applying the artist's style fingerprint from the style transfer trained model and further synthesizes new audio samples that capture the artist's unique musical characteristics.
[0059] Through the library integration module, the artists upload their music libraries for high-quality audio analysis and are compatible with formats including WAV, AIFF, and FL AC amongst others.
[0060] With the aid of the audio indexing and preprocessing module, the connected music library undergoes an indexing and feature extraction process. During this process, metadata and audio features are extracted.
[0061] This indexing involves digital signal processing allowing for detailed and granular analysis of the music content.
[0062] The Al system performs metadata extraction from audio files, utilizing available sources such as ID3 tags or embedded metadata. It parses and stores key metadata fields including title, artist, album, year, and genre.
[0063] In cases where metadata is absent or incomplete, the system employs techniques such as querying music databases or analysing file names to retrieve and supplement missing information.
[0064] Furthermore, preprocessing techniques like normalization and noise reduction are applied systematically to maintain consistent audio quality throughout the analysis process. These steps ensure that the system effectively processes and interprets the audio data, providing artists with reliable insights and enhanced capabilities for their creative endeavours.
[0065] Advanced and multiple techniques for extracting audio features are utilized to capture distinct elements of the artist's music within the system using the feature extraction module.
[0066] Low-level features such as spectral analysis, Mel-frequency cepstral coefficients (MFCCs), and chroma features are extracted to capture the tonal, timbral, and harmonic characteristics of the music. Moreover, the Mel-frequency cepstral coefficients (MFCCs)are computed using the short-time Fourier transform (STFT) to represent the spectral envelope of the audio.
[0067] Mid-level features like rhythm patterns, tempo, and key are identified using beat tracking and key detection algorithms to understand the timing and musical scale.
[0068] High-level features such as melody, harmony, and song structure are extracted using techniques like pitch tracking, chord recognition, and structural segmentation.
[0069] Melody identifies the main tune or sequence of notes that stands out.
[0070] Harmony examines he chord progressions and how different notes and chords are combined.
[0071] Rhythm analyses the beat patterns, tempo, and timing elements of the music.
[0072] Timbre studies the unique sound quality and texture of the instruments used.
[0073] Dynamics observes the changes in loudness and intensity throughout the piece.
[0074] With the help of style transfer training model, the features extracted from the artist's music catalog are inputted to train a deep neural network. Convolutional neural networks (CNNs) and recurrent neural networks (RNNs) are employed to understand both the temporal (time-related) and spatial (frequency-related) patterns in the audio data. The network process these features.
[0075] Techniques like autoencoders and generative adversarial networks (GANs) are utilized to capture the artist's unique style and generate new compositions that reflect this style.
[0076] Autoencoders aid in learning compressed representations of musical features, while GANs facilitate the generation of new compositions that closely match an artist's style.
[0077] During training, the model's parameters are adjusted to minimize the differences between the Al-generated music and the original music style of the artist. This process ensures that the Al accurately replicates and extends the artist's musical characteristics in the newly generated compositions.
[0078] Furthermore, once the style transfer model completes its training, it generates a concise representation known as the style fingerprint for the artist's music using the style fingerprint generating module.
[0079] This style fingerprint combines various elements extracted from the artist's music catalog, including features, metadata, stylistic components, and contextual details.
[0080] The style fingerprint is a high-dimensional vector or a set of learned model parameters. This representation is meticulously designed to encapsulate the distinctive characteristics and subtle nuances that distinguish the artist's musical style. Key components such as melodic patterns, harmonic progressions, and rhythmic structures are encoded within this fingerprint.
[0081] Once the style fingerprint is generated, it is securely stored within a dedicated database. It is closely associated with the artist's profile to ensure accurate referencing and efficient licensing processes. This repository serves as a resource for leveraging the artist's musical style in generating new compositions.
[0082] The Al system utilizes sophisticated algorithms to identify and categorize the cultural context, influences, and era related to the musical works. It applies techniques such as sentiment analysis and emotion recognition to analyze the audio files and determine the emotional range and expression conveyed in the music.
[0083] In order to generate new compositions via the composition generation module, the style transfer model utilizes either a set of input musical features or a seed composition asa starting point or input. The module processes these inputs to produce new audio samples that faithfully exhibit the characteristics encoded in the artist's style fingerprint.
[0084] In the system, TensorFlow and PyTorch are the deep learning frameworks used for building and training neural networks.
[0085] In fact, the system undergoes iterative training and optimization to improve the style fingerprint's quality and robustness.
[0086] The system employs Amazon S3 and / or Google Cloud Storage as the cloud storage solutions for storing and retrieving audio files securely. These platforms are chosen for their robustness, scalability, and security features, which are crucial for storing and accessing large volumes of audio files securely.
[0087] Through the system, the artists can monetize their style fingerprint every time it is used to generate new uses through Al. This creates a new revenue stream that complements traditional royalties and licensing fees.
[0088] Figure 1 refers to the flow chart of steps of the present method 100 . The flow chart depicts a detailed and methodical approach to develop and refine style fingerprints
[0089] Step 101- The process commences with data acquisition and preprocessing. During this initial step, the artist's music library integrates seamlessly with the system through the music library integration module. This integration ensures a comprehensive analysis of the artist's entire catalog.
[0090] Step 102- Once the data is pre-processed, the workflow splits into two concurrent processes the feature extraction and the metadata extraction
[0091] Step 102a - In the feature extraction step, advanced audio feature extraction techniques are employed to identify and extract specific attributes and characteristics fromthe pre-processed audio files. These features include tempo, rhythm, pitch, timbre, and other musical elements crucial for style analysis.
[0092] Step 102b - During the metadata extraction step, additional contextual information related to the audio files is gathered. This includes details such as genre, artist, release date, and other relevant metadata.
[0093] Step 103 -In Step 103, Steps 102a and 102b merge to form the stylistic analysis. Within this step, extracted features and metadata are collectively analyzed to identify the stylistic patterns and nuances within the music. Statistical and machine learning techniques are employed in this phase to pinpoint distinctive stylistic markers that characterize the artist's music.
[0094] Step 104- After the stylistic analysis step, the process moves to the context and emotion detection step. This step interprets the stylistic elements to identify the underlying context and emotional tone present in the music. Techniques such as sentiment analysis, mood detection, and contextual tagging are utilized to understand the emotional and contextual aspects of the music.
[0095] Step 105 - This step is called style fingerprint generation. A deep neural network is trained for style transfer using the extracted features and optimizing model parameters to minimize differences between Al-generated audio and the artist's original style. This training process involves the use of convolutional neural networks (CNNs) to capture spatial patterns, recurrent neural networks (RNNs) to model temporal dependencies, autoencoders for feature learning, and generative adversarial networks (GANs) to enhance the authenticity of the generated style. The output of this step is a unique style fingerprint that encapsulates the artist's musical characteristics.
[0096] Step 106- In the validation and refinement step, the fingerprint is evaluated to ensure its accuracy and reliability. This evaluation involves comparing the style fingerprint against a validation set of the artist's music and making iterative adjustments to address any inconsistencies. The validated style fingerprint is stored in a secured database andassociated with the artist's profile. This ensures that the fingerprint can be referenced in the future for licensing purposes, collaboration, and further analysis.
[0097] Step 107- The feedback loop is integral to this process, allowing for continuous improvement by feeding back new insights and adjustments into earlier stages of the workflow.
[0098] The final step of the method involves using the trained style transfer model to generate novel compositions in the artist's style. This step leverages the refined style fingerprint to create high-quality, stylistically consistent music that adheres to the unique characteristics of the artist's original works.
[0099] There includes a cloud-based platform on which the trained style transfer model is deployed to enable scalable and efficient generation of compositions.
[0100] The style fingerprint generated is a combination of extracted features, metadata, stylistic elements, and contextual information
[0101] The stylistic elements include vocal styles, instrumental techniques, production styles and lyrical themes.
[0102] Figure 2 illustrates a system architecture 200 of the present system. The details are recited in the below paragraphs. The system comprises of three panels namely style capture 201, style transfer 202 and style detector 203.
[0103] The artist music library 204 provides raw music data through an API to the library integration and pre-processing module 205.
[0104] The library integration and pre-processing module 205 manages this content and sends it to the feature extraction module 206.
[0105] Furthermore, the feature extraction module 206 performs multiple levels of feature extraction (low, mid, high), analyzes the dynamic range, and gathers metadata, these features are then forwarded to the encoding representations module 207.
[0106] This module 207 uses style encoders to convert the features into vector embeddings, which are then managed by the STRA management 208. The STRA management 208 takes these vector embeddings, creating encoded style fingerprints, and supplies them to the style transfer / generator 202.
[0107] Thereafter, the style transfer / generator 202 uses these fingerprints to synthesize or invert styles, decoding them into new music data, the generated music styles are then processed by style detection panel 203, which provides analysis and detection data to a combination of Al models.
[0108] The system starts with a cloud-based storage solution like Amazon S3 or Google Cloud Storage or a connection to artists content via linking data. Artists upload their music libraries comprising high-quality audio files in different musical formats. There is an admin panel with a user-friendly interface that allows artists to seamlessly connect their music libraries via a unique links on the admin panel.
[0109] Neural Style Encoding Agent: Within the system there is a deep learning model, that uses advanced frameworks like TensorFlow or PyTorch. This model is trained to encode musical attributes extracted from the audio content. The encoding process identifies and quantifies essential elements such as melody patterns, harmonic structures, rhythmic complexities, and nuanced production details. Thereafter, algorithms are designed to construct a distinctive fingerprint that captures the artist's unique style. With the assistance of the neural style encoding agent, several outputs are generated that include hash codes, model weights, and other relevant data formats essential for accurately representing and manipulating the artist's musical identity.
[0110] Neural Engine Admin Dashboard: The neural style encoding agent further includes an admin dashboard for artists to manage their encoded content. The dashboardoffers features for artists to continuously review, modify, augment and recalibrate their encoded styles.
[0111] This dashboard is integrated to a neural engine server ensuring seamless synchronization and management of encoded data, facilitating iterative refinements and updates to enhance the accuracy and fidelity of style representation over time.
[0112] Neural Engine Server: The neural engine server serves as the backend of the system. The primary function of this sever is to implement authentication and authorization mechanisms to ensure secure access to the encoded content.
[0113] The Style Transfer Rights Agency (STRA) safeguards rights of the artists within the realm of style transfer Al. STRA defines its organizational structure, governance model, and membership criteria, aiming to include artists and rights holders. Strategic partnerships with key stakeholders in the music industry, such as labels, publishers, and collecting societies, are crucial for fostering collaborative efforts and ensuring comprehensive protection of artists' intellectual property.
[0114] This agency includes a Style Transfer Registry, a centralized database designed to store and manage style transfer records for each registered artist. This entails building a secure and scalable infrastructure capable of handling large volumes of data while ensuring stringent access controls and authentication mechanisms to safeguard the confidentiality of artists' style transfer data.
[0115] Artists are facilitated through a user-friendly portal for registering their music libraries and creating style transfer fingerprints. Utilizing advanced style transfer Al models, developed in-house or in collaboration with Al platforms, the system analyses artists' music to generate unique style fingerprints. These records, including metadata such as artist names, song titles, and rights information, are securely stored within the registry.
[0116] Integration with a combination of Al models is a pivotal step, involving collaboration to seamlessly integrate the style transfer registry. This includes thedevelopment of APIs and secure communication protocols enabling real-time access to the registry during Al model processing. Additionally, a mechanism is implemented to scan user prompts, identifying references to registered artists or songs to initiate the style transfer matching process.
[0117] The establishment of standardized license agreements forms a critical component, outlining terms and conditions for utilizing artists' style transfers in generative Al outputs. Defining royalty structures, taking into account the prominence of the artist's style and commercial usage, ensures fair compensation. A robust system tracks and records style transfer usage, generates agreements, and facilitates royalty payments to the appropriate rights holders, ensuring transparency and accountability.
[0118] Monitoring and enforcement efforts involve developing sophisticated tools to detect unauthorized or unlicensed style transfer usage in generative Al outputs. A comprehensive legal framework and enforcement mechanisms are established to address infringements and protect artists' rights, supported by collaboration with legal experts and industry bodies to navigate evolving intellectual property landscapes within generative Al.
[0119] Education and outreach initiatives are integral to the organization's mission, encompassing campaigns to raise awareness among artists, rights holders, and the public about style transfer rights. Providing resources and support helps artists navigate registration processes, explore monetization opportunities, and adopt best practices for managing their style transfer rights. Engagement with music industry stakeholders, tech companies, and policymakers fosters dialogue on ethical and legal implications, ensuring continuous evolution and refinement of style transfer Al practices.
[0120] This integrated approach ensures that STRA effectively manages style transfer records, upholds artists' rights, and advances responsible use of Al technologies in the creative industry.
[0121] The system may employ a single generative model fine-tuning approach or multi-modal hybrid model approach for stylometrics engine development.
[0122] In the Single Generative Model Fine-Tuning approach the focus is on adapting a robust, pre-trained generative music model by training it exclusively on a specific artist's discography. The resulting fine-tuned model then functions as that artist's unique signature within a digital vault.
[0123] The Single Generative approach as illustrated in Figure 3 comprises Step 1, Step 2 and Step 3.
[0124] Step 1 - Base Music Model Selection
[0125] This step requires identifying suitable pre-trained generative music models (base model selection).
[0126] The promising candidates may include Transformer-based Autoregressive Models like Meta's MusicGen which is an open-source model and supports text-to-music and melody- conditioned generation and other open-source large-scale models, such as YuE, may be considered if they are fully open and fine-tunable. Additionally, diffusion models like Stable Audio, Riffusion (which leverages Stable Diffusion fine-tuned on spectrograms), Audio LDN may be used, though their fine-tuning capabilities and may depend on API access or open-source releases.
[0127] Step 2 - Per- Artist Data Preparation.
[0128] This step begins with audio collection and quality control, wherein the artist's complete works are gathered and converted to high-quality, consistent formats such as WAV or FLAC. Additionally, the lossy formats like MP3 are avoided to preserve maximum detail and volume levels are standardized through normalization.
[0129] Following the audio collection and quality control step, there includes a segmentation step that divides the full audio tracks into shorter segments (e.g., 10-30 seconds) to increase the dataset size and focus learning, with overlapping segments used to better capture transitions. The segments should be long enough to contain meaningful musical phrases or ideas.
[0130] A method for preparing the artist's audio data for fine-tuning is provided, wherein each audio track is systematically segmented. The method comprises dividing the full-length tracks into shorter segments, for instance, between 10 to 30 seconds in length. This segmentation serves the dual purpose of significantly increasing the size of the training dataset while also focusing the model's learning on discrete musical ideas.
[0131] For text-conditioned models, Metadata / Text Prompt Association is essential. In this particular step, each audio segment may need an associated text description, which may include existing lyrics, manual annotations (e.g., genre, mood, instrumentation, tempo), or automated captioning (requiring careful validation). These data pairs (audio file path, text description) are structured in a format suitable for training libraries.
[0132] Step 3 -Fine-tuning Workflow
[0133] In this step a high-performance Graphics Processing Unit (GPU) environment (e.g., NVIDIA Al 00, Hl 00) is selected with VRAM requirements depending on the chosen model size and fine-tuning technique. Libraries including PyTorch, Hugging Face (transformers, diffusers, accelerate, peft), and audiocraft (for MusicGen) are utilized for this purpose.
[0134] The pre-trained weights and architecture of the selected base model are loaded.
[0135] Pursuant to the loading step, Parameter-Efficient Fine-Tuning (PEFT), particularly methods like LoRA (Low-Rank Adaptation) or its variants (e.g., QLoRA for further memory saving), may be used for fine tuning of the model.
[0136] This approach is strongly recommended for efficiency and performance. This approach enables the adaptation of large, pre-trained models by inserting small, trainable "adapter" layers while keeping the vast majority of the base model's parameters frozen. This significantly reduces computational cost and helps prevent overfitting to an artist's specific dataset.
[0137] The decision of where to place these LoRA adapters is a critical part of the fine-tuning process. While common strategies involve adding adapters to the attention layers of a transformer, the optimal placement is an empirical question. The ideal configuration would be determined during the model development phase through experimentation, aiming to find the placement that best captures an artist's stylistic nuances without degrading the model's overall performance.
[0138] Varying adapter placement or fine-tuning strategy based on genre or audio signal characteristics is an advanced optimization that may be explored.
[0139] The fine tuning of the model reduces computational cost and GPU memory requirements, lowers the risk of "catastrophic forgetting" and overfitting to smaller artist datasets. The results are stored in very small adapter files (megabytes), making storage highly efficient. The artist's signature then becomes this small adapter file combined with a pointer to the common base model.
[0140] Within the fine-tuning step, there includes a training configuration mechanism that typically uses AdamW as the optimizer with a relatively small learning rate (e.g., le-4, 2e-4, 5e- 5), potentially with a warm-up phase and decay schedule, requiring tuning per model and dataset. The batch size is maximized to fit GPU memory, and training is usually limited to 1-5 epochs for LoRA fine-tuning to avoid over-training.
[0141] For evaluation purposes, it is crucial not to rely solely on loss metrics. Audio samples are generated frequently throughout training using relevant text prompts to subjectively assess if the style is being captured without sacrificing quality or becoming repetitive.
[0142] For execution purposes, tools like accelerate facilitate multi-GPU or mixed-precision training, and metrics and generated audio samples are logged using tools like Weights & Biases or MLflow.
[0143] The Stylometrics Engine Artefact as illustrated in Figure 5
[0144] In Step 4, if PEFT (LoRA) is used, the artefact is the set of trained adapter weights (a small file) specific to the artist and the identifier of the base model it was trained upon.
[0145] If full fine-tuning is employed (less recommended), the artefact is the entire set of modified model weights (a very large file).
[0146] Step 5 - Platform Integration
[0147] In the Platform Integration step (Vault and Usage), the artefact (adapter weights and base model ID) is securely stored and linked to the artist's profile in a centralised database. Functionality is developed to dynamically load the appropriate base model and subsequently apply the specific artist's adapter weights when their style is requested for music generation or analysis.
[0148] Thereafter, an API endpoint is exposed to accept requests (e.g., artist ID, text prompt) and return generated audio using the corresponding fine-tuned model.
[0149] The method for adapting a single generative model through LoRA (Low-Rank Adaptation) fine-tuning, focusing on efficiency and preserving artistic integrity includes the below steps.
[0150] The first key process involves music track segmentation for fine-tuning. This entails systematically dividing full-length audio tracks into shorter, typically 10 to 30-second segments. This strategic segmentation serves a dual purpose: it significantly expands the training dataset available to the model while simultaneously directing its learning towards discrete musical ideas.To ensure that the model also learns the nuances of transitions within an artist's music, these segments are created with overlaps. A crucial constraint throughout this process is ensuring each segment is long enough to encapsulate meaningful musical phrases or ideas, thereby safeguarding the artistic coherence of the source material.
[0151] The second core component is LoRA fine-tuning customization, which leverages Parameter-Efficient Fine-Tuning (PEFT). This technique involves freezing the original weights of the pre-trained generative model and injecting small, trainable "adapter" layers into its architecture. This approach is highly beneficial as it drastically reduces computational cost and GPU memory requirements, and importantly, mitigates the risk of "catastrophic forgetting," a common issue when fine-tuning large models. In this specific embodiment, an artist's unique musical "signature" is effectively encapsulated within these small adapter files, linked to the base model they modify.
[0152] Regarding the non-generic adaptations of this method, the precise placement of LoRA adapters within the base model architecture is a key configuration detail. The determination of which specific layers are augmented with these adapters is established during the model tuning and optimization phase. This empirical process ensures that the chosen placement optimally captures the artist's stylistic characteristics.
[0153] The concepts of dynamically selecting LoRA adapter rank based on metadata or creating custom tuning pipelines aligned to musical phrasing represent innovative, cutting-edge optimizations. These are precisely the kinds of advanced techniques that our flexible framework is designed to accommodate. However, specifying the exact mechanics of such novel adaptations would be premature at this stage. These are granular, implementation-level details that would be researched, developed, and validated during the build phase. The optimal approach would be determined through empirical testing to see which methods yield the most faithful and robust stylistic representations.
[0154] In the Single Generative Model Fine-Tuning approach there are several challenges and considerations. This includes scalability, which is significantly enhanced by PEFT for managing potentially thousands of fine-tuned adapters or models. Ensuring consistency in data preparation and fine-tuning across diverse artists is key.
[0155] GPU time for fine-tuning each artist represents a significant cost. This method implicitly defines "style" as "whatever the fine-tuned model generates," which might capture dominant traits well but could struggle with artists having very diverse styles within their own catalog.
[0156] While the fine-tuned model embodies the style, directly detecting that style in externalaudio is less straightforward than with a dedicated style encoder. Detection might involve techniques such as measuring the probability the fine-tuned model assigns to a piece of external music, comparing embeddings from internal layers of the model for known versus external tracks, or generating music from a suspected track's description using the artist's model and comparing outputs, all of which require further research and development. This focused fine-tuning approach offers a potentially faster path to integrating generative capabilities for specific artist styles into a platform, especially leveraging PEFT techniques.
[0157] The multi-modal hybrid model achieves the necessary depth and accuracy for representing an artist's style, moving beyond simply fine-tuning a single pre-existing model. This integrated system aims to deeply internalize the unique stylistic essence of musical creations and will form the technical foundation for tools that allow for creative extension of style and analytical understanding of its presence in the wider musical landscape.
[0158] The Multi-modal Hybrid Model approach as illustrated in Figure 4 comprises the following steps.
[0159] Step 1 - Data Foundation
[0160] The process begins with a robust data foundation which utilizes the artist's complete discography, converted to high-quality, lossless formats like WAV or FLAC, to preserve maximum detail. This also includes compiling and structuring comprehensive metadata for each track, such as precise genre / subgenre, instrumentation, tempo (BPM), key signature, time signature, mood descriptors, lyrical content, and potentially production notes or symbolic representations.
[0161] Step 2 - Advanced Audio Processing
[0162] In Step 2, Advanced Audio Processing is performed that involves converting the audio files into representations suitable for deep learning models, such as Mel spectrograms or Constant-Q Transforms (CQT). Specialized audio encoders (e.g., EnCodec, DAC, Audio Spectrogram Transformers - AST) are employed which are likely pre-trained models that are then fine-tuned specifically on the artist's music collection to learn the nuances of their sound palette and timbres.
[0163] For the Audio Representation Learners to function effectively, raw audio data must first be meticulously prepared. This involves converting it into suitable formats like spectrograms or Constant-Q Transforms (CQTs), followed by crucial preprocessing steps. These steps ensure the model receives clean, consistent, and meaningful input.
[0164] Firstly, the source audio files are analyzed for Noise Reduction and Quality Control. Systematic noise reduction techniques are applied to clean the audio and eliminate unwanted artifacts that could interfere with feature extraction. This ensures the model learns from the core musical content rather than incidental noise.
[0165] Audio is prioritized from high-quality, lossless formats like WAV or FLAC whenever possible to preserve maximum detail.
[0166] Secondly, normalization standardizes the volume levels across an artist's entire catalog, which may have been recorded and mastered at different times. This prevents the model from misinterpreting loudness as a stylistic feature, allowing it to focus on more nuanced characteristics of timbre, harmony, and rhythm.
[0167] Due to their length, full musical tracks are then subjected to Segmentation and Overlapping step. They are divided into shorter chunks, typically between 10 to 30 seconds. This technique serves two purposes: it significantly increases the size of the training dataset, and it helps the model focus its learning on discrete musical ideas.
[0168] To ensure the model also learns how an artist transitions between these ideas, overlapping segments are utilized. The chosen segment length is always long enough to contain a complete and meaningful musical phrase.
[0169] While this segmentation strategy aims to capture complete musical phrases, the highly specific task of Phrase Alignment. The necessity and specific methodology for such alignment, whether through automated beat-matching, structural analysis, or other techniques, would depend on the final model architecture and the artist's specific stylistic traits.
[0170] Given the evolving nature of audio processing in Al, committing to a single method now would be premature; the optimal approach would be determined during the build phase.
[0171] Step 3 -Multi-Component Model Architecture
[0172] Step 3 involves the multi-component model architecture, which is fine-tuned on the artist's data and comprises several specialized learners:
[0173] First learners are the Audio Representation Learners that utilize models like Transformers (potentially variants of MusicGen, YuE) or Diffusion Models (inspired by Riffusion, Stable Audio, or adapted AudioLDM). These are fine-tuned on the audio representations (spectrograms or embeddings) to deeply learn the characteristic sonic textures, harmonic language, and rhythmic patterns of the artist's music. Techniques such as LoRA (Low-Rank Adaptation) or PivotalParameters Tuning might be employed during fine-tuning to efficiently adapt these large models to the specific dataset without catastrophic forgetting or overfitting.
[0174] The Audio Representation Learners are tools are designed to deeply understand the characteristic sonic textures, distinctive harmonic language, and signature rhythmic patterns that define an artist's sound. They achieve this by being meticulously trained, or fine-tuned, on rich audio representations like Mel spectrograms or Constant-Q Transforms (CQTs), which are derived directly from the artist's high-quality audio files.
[0175] To facilitate this, pre-existing model architectures such as Transformers or Diffusion Models can be employed. Before these models can process raw audio, specialized audio encoders like EnCodec or Descript Audio Codec (DAC) convert the audio into a format they can understand. These encoders are crucial for generating high-fidelity, compressed representations (often called tokens) of the audio, which are then used in the fine-tuning process.
[0176] The fine-tuning itself is a carefully managed process that adapts these large models to an artist's specific body of work. Advanced techniques such as LoRA (Low-Rank Adaptation) are particularly effective here, as they efficiently adjust the models without issues like catastrophic forgetting or overfitting. The end result of this learning process is a sophisticated embedding that encapsulates the artist's essential timbral and textural palette.
[0177] Second learners are the sequence and structure learners that employ Transformer-based models to analyze the temporal evolution of the music, learning characteristic melodic contours, chord progressions, song structures, and rhythmic motifs over time. This component might leverage raw audio embeddings or symbolic data if available.
[0178] The Sequence & Structure Learners are all about understanding how an artist's music unfolds over time. Their core purpose is to analyze the temporal evolution of the music itself, zeroing in on an artist's signature melodic contours, common chord progressions, typical song structures, and recurring rhythmic motifs. By processing the music as a sequence, this component truly captures the architectural blueprint of an artist's unique compositional style.
[0179] These learners typically leverage Transformer-based models, which are incredibly adept at analyzing sequential data. The input for this component is quite versatile: it can process the embeddings created by the Audio Representation Learners to understand how sonic textures evolve, or it can directly use symbolic music formats like MIDI if they're available in the artist's dataset. Ultimately, the output is a powerful representation that encodes an artist's typical methods for building tension and release, smoothly transitioning between sections, and developing musicalideas across an entire composition.
[0180] Third leaners are the semantic and contextual learners that integrate Language Models (LLMs), fine-tuned on the artist's lyrics and associated metadata. This component aims to capture the thematic content, lyrical style, and conceptual underpinnings of the artist's work. It would also establish cross-modal links, perhaps using attention mechanisms or learned style tokens / textual inversion, to connect specific sonic events with descriptive tags or lyrical meanings.
[0181] These learners introduce a crucial layer of meaning and intent to the stylometric analysis, moving beyond purely sonic and structural elements to capture the conceptual and lyrical soul of an artist's work. Their function is to understand the thematic content, identify the unique lyrical style, and grasp the overall conceptual underpinnings that define the artist's message and narrative voice. This is achieved by integrating and fine-tuning Large Language Models (LLMs) on the artist's complete lyrical catalog and associated metadata, such as mood descriptors or production notes.
[0182] A key capability of this component is its ability to create bridges between different modes of expression. It establishes cross-modal links, using mechanisms like attention or learned style tokens to connect a specific lyrical phrase with a particular instrumental choice or sonic event identified by the other learners. This allows the system to learn not just that a certain guitar tone is used, but why — for instance, that it is often employed when the lyrical theme is one of melancholy or defiance. Consequently, this component produces a representation that captures the artist's "why" — the semantic and emotional context that imbues their music with its deeper meaning.
[0183] Step 4 - Fusion Mechanism
[0184] In Step 4, fusion mechanism is implemented. This intelligently integrates the learned representations from the audio, sequential, and semantic components into a unified, comprehensive "stylometric" vector or set of vectors from the disparate analyses.
[0185] The specific method for fusion may be embodied in several ways, including, but not limited to: Utilising cross-attention mechanisms to allow one modality to selectively weigh the importance of information from another; projecting the outputs of all learners into a shared embedding space where their relationships can be jointly analyzed; or concatenating the individual output vectors and processing them through further neural network layers to learn a final, fused representation.
[0186] Step 5 - Artefact Generation
[0187] The artefact of this approach is not a single file, but rather the entire integrated system. Thisincludes the specific model architectures chosen for each component, their fine-tuned weights learned from the artist's data, and the exact preprocessing steps required to feed new data into it. Essentially, it is a bespoke, state-of-the-art Al system grounded in multi-modal learning principles, designed to deeply internalize the unique stylistic essence of the musical creations.
[0188] The stylometric token serves as the digital asset representing an artist's unique style fingerprint. Its creation method varies based on the chosen model architecture. In one embodiment, employing a Single Generative Model, the token comprises the set of trained LoRA (Low-Rank Adaptation) adapter weights. This constitutes a small, portable file containing specific parameters designed to adapt a large base model to the artist's style.
[0189] Alternatively, in an embodiment utilizing a Multi-modal Hybrid Model, the stylometric token is a high-dimensional vector or a set of learned model parameters. This vector is the output of a Fusion Mechanism that intelligently integrates analyses of the artist's music across audio, structural, and semantic dimensions into a single, comprehensive representation. In both scenarios, the token is generated through training a deep neural network on the artist's music catalog, specifically designed to encapsulate key characteristics such as melodic patterns, harmonic progressions, and rhythmic structures.
[0190] A core component of the present invention is the ability for rights holders to specify limitations on the use of their stylometric token. This is managed through a combination of a userfacing dashboard and a robust backend enforcement system.
[0191] The Functional Aspects of the User Interface (UI) provide rights holders with access to a "Neural Engine Admin Dashboard" or a similar user-friendly portal. Within this UI, an artist or rights holder is presented with features to manage their encoded styles, including a dedicated section for defining the terms and conditions for the use of their stylometric token. The UI is configured to allow the specification of various parameters, such as defining royalty structures based on commercial usage, permitting or prohibiting use in certain contexts, or setting licensing terms for different platforms.
[0192] The Data Structure for the Token ensures that the stylometric token is more than just raw style data; it is part of a larger, secure data structure stored within the Style Transfer Registry. The method for structuring this data comprises a unique token identifier, a rights holder ID, the stylometric data itself (either the vector or a secure pointer to the LoRA file), and a "Policy" object. This Policy object contains machine-readable rules and limitations specified by the rights holder in the UI. The entire record is securely stored in an encrypted form and linked to the artist's profile,providing a robust framework for managing usage rights.
[0193] The enforcement of these limitations is handled by the system's backend, which implements stringent Technical Controls and One-Time Use mechanisms, including access controls and authentication.
[0194] Any request by an authorized system, such as a large language model (LLM), to use a token is managed via a secure API call to a central server. The method for a single, controlled use involves several steps: the authorized system makes an API call, authenticating itself with the server; the server then retrieves the token's data structure from the registry and checks the request against the rules defined in the Policy object. To enforce one-time use for that specific generative act, the system can be configured to issue a single-use cryptographic nonce or a time-limited access credential that must be presented to access the stylometric data. Once the generative act is completed, this credential becomes invalid. Crucially, the system meticulously tracks and records each transaction, logging the usage for auditing purposes and to facilitate royalty payments.
[0195] Furthermore, the method for the technical enforcement of stylometric use may be embodied in a system that employs blockchain-based tokenization as a foundational layer for trust and transparency. Each unique style fingerprint is represented as a secure and ownable cryptographic token on a distributed ledger. This Style Transfer Registry provides an immutable record of ownership and the associated usage rights defined by the artist. Access to use the style is then managed through secure, API-gated calls from authorized third-party generative platforms. When a platform requests to use a style, the system's authentication mechanisms validate the request against the on-chain record of the token. Each instance of use is then recorded as a new transaction on the blockchain, creating a transparent and automated method for tracking all usage, which is essential for facilitating royalty payments.
[0196] The system architecture provides a framework for implementing advanced and efficient fine-tuning techniques, such as LoRA (Low-Rank Adaptation). The potential embodiment of dynamically conditioning LoRA adapter placement and rank selection based on specific musical attributes is a sophisticated optimization that the framework is designed to accommodate. However, the specific method for implementing said dynamic conditioning represents a granular level of implementation detail. The determination of a definitive method for such a task — whether by genre, tempo, or other metadata — is a matter for further research and empirical validation during the development phase. Committing to a specific technique at this stage would be premature, and the present invention provides a flexible foundation for incorporating such future optimizations.
[0197] In another aspect of the present invention this style fingerprint is securely stored in an encrypted form in a database linked to the artist's profile for future reference and licensing. The system employs a comprehensive validation framework to ensure the fidelity and authenticity of the generated style fingerprints through quantitative assessment and subjective evaluation methods.
[0198] The validation process incorporates multiple assessment methodologies to ensure both technical accuracy and artistic authenticity. Subjective human evaluation forms a crucial component wherein audio samples generated throughout the training process are assessed by listeners to determine if the artist's style is captured authentically without sacrificing musical quality or becoming repetitive. This perceptual validation ensures the "feel" of the style is correctly represented, which purely mathematical metrics cannot always capture.
[0199] Latent space analysis provides quantitative assessment by comparing embeddings of newly generated compositions against embeddings from a validation set of the artist's original music. By measuring the distance or clustering of these points in the model's latent space, the system quantitatively scores stylistic similarity, wherein compositions that cluster closely with the artist's known works indicate high stylistic fidelity.
[0200] Probabilistic scoring utilizes the trained model as an evaluation tool, measuring the probability that the fine-tuned model assigns to a given piece of music. High probability scores indicate that the music closely conforms to the learned stylistic patterns of the artist, providing a direct quantitative measure of similarity.
[0201] The system architecture incorporates state-of-the-art Al components including transformers, diffusion models, and Low-Rank Adaptation (LoRA) fine-tuning techniques, each specifically configured and adapted for music domain applications.
[0202] Transformers serve as a cornerstone of the architecture, particularly adapted for understanding and generating complex musical sequences. For style fingerprinting, transformer encoder blocks process sequences of audio features from an artist's work, wherein the self-attention mechanism learns intricate, long-range dependencies characteristic of the artist's style, such as relationships between harmonic progressions and rhythmic motifs. The contextualized embedding produced by the encoder serves as a detailed style fingerprint capturing the essence of musical structure and language. For token generation, autoregressive transformers are fine-tuned on an artist's work to learn unique statistical patterns, enabling generation of new compositions token- by -token in the artist's signature style. For constrained generation, transformers are adapted toaccept external inputs through cross-attention layers, allowing text prompts or melody conditions to guide the generation process while maintaining the learned artistic style.
[0203] Diffusion models provide an alternative approach for music generation by progressively refining noise into coherent musical pieces. For style fingerprinting, the encoder component of the diffusion model's architecture creates meaningful latent representations of audio that serve as style fingerprints. For token generation, diffusion models operate on spectrograms and are trained on an artist's music to learn the denoising process, starting with random noise and meticulously refining it into spectrograms embodying the artist's timbral and textural qualities. For constrained generation, the denoising process is conditioned by external inputs such as text descriptions, guiding the refinement process to ensure outputs align with both the artist's style and specific prompt constraints.
[0204] LoRA fine-tuning provides efficient customization of large base models through injection of small, trainable adapter layers while freezing original model weights. For style fingerprinting, LoRA adapters are trained exclusively on a specific artist's discography, creating a compressed and portable representation of the artist's style that serves as the style fingerprint. For token generation, the LoRA fingerprint is loaded and merged with the general-purpose base model, creating a specialized model that generates music steered by the LoRA adapters to produce tokens in the artist's distinct style. For constrained generation, LoRA adapters modify how the base model interprets prompts, infusing outputs with the artist's unique stylistic traits while maintaining the constraint-handling capabilities of the underlying model.
[0205] Finally, the trained style transfer model is used to generate novel compositions that mirror the identified musical traits, enabling the creation of music in the artist's distinctive or signature style.
[0206] This approach is designed to deeply internalize the unique stylistic essence of the musical creations and will form the technical foundation for tools to creatively extend style and analytically understand its presence in the wider musical landscape.
[0207] The present invention and the described preferred embodiments specifically include at least one feature that is industrial applicable.
Claims
Claims:
1. An Al driven system for generating style transfer fingerprints and compositions in a computer environment, comprising;• a library integration module, wherein the module seamless integrates an artist sonic library(s) and receives audio files;• an audio indexing and preprocessing module operatively associated to the library integration module, wherein the module extracts metadata and audio features from the sonic library using preprocessing techniques;• a feature extraction module, wherein the module utilises advanced audio feature extraction techniques to capture a first level, a second level and a third level feature(s) of the artists sonics;• a style transfer training model coupled to the feature extraction module, wherein the style transfer model uses extracted artistic sonic features to train a deep neural network, optimizes parameters to closely match the artist's original style in the generated audio;• a style fingerprint generation module, wherein the style transfer model generates a style fingerprint that encapsulates the artist's sonic characteristics;• a database that stores generated style fingerprints securely and associates the fingerprints with artists' profiles; and• a composition generation module connected to the style fingerprint module, wherein the module creates new compositions by applying the artist's style fingerprint from the style transfer trained model and further synthesizes new audio samples that capture the artist's unique sonic characteristics.
2. The system of Claim 1 , wherein the sonic library and sonic characteristics may be the music library and musical characteristics.
3. The system of Claim 1, wherein first level features include spectral analysis, Mel- frequency cepstral coefficients (MFCCs), and chroma features.
4. The system of Claim 1, wherein second level features include rhythm patterns, tempo, and key using beat tracking and key detection algorithms.
5. The system of Claim 1, wherein third level features include melody, harmony, and song structures.
6. The system of Claim 1, wherein the audio files are in high-quality formats such as WAV, AIFF, or FLAC.
7. The system of Claim 1 , wherein the preprocessing techniques include normalization and noise reduction applied to ensure consistent audio quality.
8. The system of Claim 1, wherein the metadata fields include title, artist, album, year, and genre respectively.
9. The system of Claim 1 or Claim 2, wherein the characteristics of the artist's music encapsulated by the style fingerprint include melodic patterns, harmonic progressions, and rhythmic structures.
10. The system of Claim 1, wherein the system uses deep learning frameworks used for building and training neural networks.
11. The system of Claim 1, wherein the system undergoes iterative training and optimization to improve the style fingerprint's quality and robustness.
12. The system of Claim 1, wherein the system uses Amazon S3 and / or Google Cloud Storage as the cloud storage solutions for storing and retrieving audio files securely.
13. A method for generating compositions in an artist's sonic style, comprising the steps of;• connecting an artist's sonic library to the system for high-quality audio analysis via a library integration module;• extracting relevant metadata using an audio indexing and preprocessing module;• applying preprocessing techniques such as normalization and noise reduction to ensure consistent audio quality.• performing advanced feature extraction from the connected sonic library;• training a deep neural network for style transfer using the extracted features and optimizing model parameters to minimize differences between Al generated audio and the artist's original style;• generating a style fingerprint of the artist's sonic characteristics post-training through the style fingerprint generation module:• storing the style fingerprint in a secured database and associating it with the artist's profile for future reference and licensing purposes; and• generating new compositions in the artist's style by applying the trained style transfer model.
14. The method of Claim 13, wherein the sonic library and sonic characteristics may be the music library and musical characteristics.
15. The method of Claim 13 or Claim 14, wherein the generated style fingerprint is validated against a subset of the artist's musical works to assess its accuracy and representativeness.
16. The method of Claim 13, wherein the step of generating new compositions includes dynamically adjusting parameters of the trained style transfer model based on real-time user input or preferences.
17. The method of Claim 13 or Claim 14, wherein the step of performing advanced feature extraction includes calculating Mel-frequency cepstral coefficients (MFCCs), chroma features, spectral contrast, and tempo to capture harmonic, rhythmic, and timbral properties of the artist's music.
18. The method of Claim 13, further comprising deploying the trained style transfer model on a cloud-based platform to enable scalable and efficient generation of compositions.
19. The method of Claim 13, wherein the method records and manages transactions related to style transfer usage rights, including licensing agreements and royalty distributions.
20. The method of claim 13, wherein the style fingerprint generated is a combination of extracted features, metadata, stylistic elements, and contextual information.
21. The method of Claim 13 or Claim 20, wherein the stylistic elements include vocal styles, instrumental techniques, production styles, and lyrical themes.
22. A computer program product stored on a non-transitory computer-readable medium for generating style transfer fingerprints and compositions, the product performs:• integrating artist sonic libraries and receiving audio files;• extracting metadata, segmenting, and normalizing audio files;• extracting multiple levels of sonic features using advanced techniques and further training a deep neural network for style transfer and optimizing parameters; and• generating and storing style fingerprints in a secure database and developing new compositions using the stored style fingerprints.
23. The computer program product of claim 22, wherein the deep neural network is trained using a combination of convolutional and recurrent neural network architectures.
24. A method for managing style transfer rights in using the Al driven system, of Claim 1, comprising the steps of ;• establishing a Style Transfer Rights Agency (STRA) including defining organizational structure, governance, and membership criteria;• developing a style transfer registry by constructing a centralized database to manage and store style transfer records for registered artists and implementing access controls and authentication mechanisms to safeguard artists' style transfer data;• facilitating artist registration and style transfer creation via an intuitive portal for artists to register sonic libraries and create style transfer fingerprints;• deploying the style transfer Al model to analyze artists' sonics and generate unique style fingerprints;• storing generated style transfer records in the registry, including metadata;• integrating with a combination of Al models by collaborating with leading Al platforms to integrate the style transfer registry and further developing APIs and secure protocols for real-time access to the registry;• implementing a feature to scan user prompts for registered artists or songs, initiating style transfer matching;establishing license agreements and royalty distribution; and monitoring and enforcing style transfer rights