Intelligent music making system and method based on model context protocol

The intelligent music production system that combines a large language model with a digital audio workstation solves the problems of difficult traditional DAW operations and data barriers of AI systems, realizes efficient human-computer collaborative music creation, and improves the quality and efficiency of music production.

CN120599982APending Publication Date: 2025-09-05CHENGDU POTENTIAL ARTIFICIAL INTELLIGENCE TECH CO LTD

Patent Information

Application Number
CN202510843707.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing music automation generation technology has problems such as a steep learning curve for traditional DAW operations, data barriers between AI music generation systems and professional production software, and a lack of semantic understanding capabilities, which leads to a split between creativity and technical implementation.

Method used

By combining a large language model with a model context protocol and a digital audio workstation, and building a two-way data stream through the TCP/IP protocol stack, AI can be used as a virtual producer to analyze creative intent and accurately control the deep functions of the DAW, thereby reconstructing the human-computer collaborative music production method.

Benefits of technology

It realizes the full-process closed-loop control of the intelligent music production system, ensures the precise correspondence between artistic expression and technical implementation, and improves the quality and efficiency of music production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599982A_ABST
    Figure CN120599982A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of intelligent music generation, and relates to an intelligent music production system and method based on a model context protocol, and the system comprises a large language model, the model context protocol, and a digital audio workstation execution layer. The large language model is used for deconstructing a natural language instruction into structured operation logic and generating a semantic instruction based on the music knowledge base; the semantic instruction at least comprises a device parameter and a chord sequence; the model context protocol is used for compiling the semantic instruction into a JSON command and synchronizing working state data of the digital audio workstation in the digital audio workstation execution layer to a large language model so as to update a model context; the model context comprises track configuration, effect chain parameters and time sequence positioning information of the digital audio workstation; the digital audio workstation execution layer is used for converting the JSON command into music operation to obtain a music operation file; a man-machine cooperative music production mode is reconstructed, and the quality of music production is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent music generation, and specifically discloses an intelligent music production system and method based on a model context protocol. Background Art

[0002] Existing automated music generation technology identifies the user's target intent based on contextual information input by the user. Upon determining the target intent, a large language model is used to call a target tool and obtain the target tool's corresponding call protocol. Based on the parameter types corresponding to the tool call parameters to be obtained, a target prompt word is generated that matches the call protocol. The target prompt word is then input into the large language model, which then uses the target prompt word to obtain the tool call parameters from the contextual information based on the target prompt word and call the target tool using the tool call parameters. This approach allows for more accurate acquisition of the tool call parameters used to call the target tool from the user's input contextual information, enabling accurate call of the target tool. Existing technology suffers from the following drawbacks: First, traditional DAW operations rely on manual operation, resulting in a steep learning curve. Second, data barriers exist between AI music generation systems and professional production software, preventing integration. Finally, existing automated tools lack semantic understanding capabilities. This creates a disconnect between creativity and technical implementation in the music creation process.

[0003] In light of this, and to address these shortcomings, the present invention provides an intelligent music production system and method based on the Model Context Protocol. This system combines large language models, the Model Context Protocol, and audio workstation technology to materialize abstract concepts into multi-track mixing projects. Furthermore, these three elements establish a bidirectional data flow through the TCP / IP protocol stack, enabling AI to act as both a "virtual producer" to analyze creative intent and an "intelligent remote control" to precisely control the deeper functions of the DAW. This reconstructs the human-machine collaborative music production method and improves the quality of music production. Summary of the Invention

[0004] The purpose of the present invention is to provide an intelligent music production system based on a model context protocol, and the specific scheme is as follows:

[0005] Includes large language models, model context protocols, and digital audio workstation execution layers;

[0006] The large language model is used to deconstruct natural language instructions into structured operation logic and generate semantic instructions based on a music knowledge base; the semantic instructions include at least device parameters and chord sequences;

[0007] The model context protocol is used to compile the semantic instructions into JSON commands and synchronize the working status data of the digital audio workstation in the digital audio workstation execution layer to the large language model to update the model context; the model context includes the track configuration, effect chain parameters and timing positioning information of the digital audio workstation;

[0008] The digital audio workstation execution layer is used to convert the JSON command into a music operation to obtain a music operation file.

[0009] Furthermore, the large language model includes a music semantic parsing module, an instruction preprocessor, and a dynamic context manager;

[0010] The music semantic parsing module is used to perform style classification on the natural language instructions, extract style tags, and decompose the structural instructions into MIDI sequence generation tasks and / or paragraph structures;

[0011] The instruction preprocessor is used to dynamically map the abstract device requirements into a device parameter set to form a multi-dimensional parameter set;

[0012] The dynamic context manager is used to maintain the working status data of the digital audio workstation; the working status data includes track configuration, effect chain parameters and timing positioning information.

[0013] Furthermore, the model context protocol includes an MCP protocol encapsulation module, a command classification encoder and a heartbeat packet controller;

[0014] The MCP protocol encapsulation module uses a differential compression algorithm to encode the working status data of the digital audio workstation;

[0015] The command classification encoder is used to perform three-layer classification encoding on the semantic instructions to generate standardized instruction types; the three-layer classification encoding includes session control, track operation and segment editing;

[0016] The heartbeat packet controller is used to periodically send status verification packets.

[0017] Furthermore, the digital audio workstation execution layer includes a remote script interface module, a digital audio workstation and a reverse state collector;

[0018] The remote script interface module is used to convert JSON commands into atomic music operations of the audio workstation, including an API instruction execution engine; the API instruction set in the API instruction execution engine includes a track audio creation unit, a virtual instrument loading unit, and an automation configuration unit;

[0019] The track audio creation unit is used to create a MIDI track or an audio track;

[0020] The virtual instrument loading unit loads virtual instruments and effects using a URI standardization scheme;

[0021] The automation configuration unit verifies the parameter range through the sandbox environment and draws the automation curve;

[0022] The reverse state collector is used to capture the engineering state data of the digital audio workstation; the engineering state data includes track level peaks, device parameter snapshots, timeline positioning information and engineering metadata.

[0023] Furthermore, the digital audio workstation execution layer also includes an abnormal fuse mechanism module; the abnormal fuse mechanism module is used to terminate the current transaction chain and automatically roll back to the most recent valid state when detecting that the parameter is out of range.

[0024] Furthermore, the large language model, model context protocol and digital audio workstation execution layer construct a bidirectional data flow closed loop through the TCP / IP protocol stack, including a forward link module and a reverse link module;

[0025] The forward link module encapsulates the semantic instructions of the large language model into JSON commands through the model context protocol, and converts them into atomic operations by the API instruction execution engine of the digital audio workstation execution layer;

[0026] The reverse link module transmits the engineering status data of the digital audio workstation back to the dynamic context manager of the large language model via the reverse status collector and the differential compression module of the model context protocol to drive the optimization of the creative strategy.

[0027] The present invention also provides an intelligent music production method based on a model context protocol, comprising:

[0028] The large language model receives natural language instructions, deconstructs the natural language instructions into structured operation logic, and generates semantic instructions based on a music knowledge base; the semantic instructions include at least device parameters and chord sequences;

[0029] The model context protocol receives the semantic instruction and compiles it into a JSON command;

[0030] The model context protocol collects the working status data of the digital audio workstation and synchronizes it to the large language model to update the model context;

[0031] The digital audio workstation execution layer receives the JSON command, converts it into a music operation, and generates a music operation file.

[0032] Furthermore, the large language model receives natural language instructions, deconstructs the natural language instructions into structured operation logic, and generates semantic instructions based on the music knowledge base, including:

[0033] Performing intent recognition on natural language instructions and extracting intent parameters; the intent parameters include at least creative style intent, instrument requirement intent, and paragraph structure intent;

[0034] Based on the style template in the music knowledge base, mapping the composition style intention into preset chord progression and rhythm type parameters;

[0035] Performing parameter analysis on the instrument requirement intention to generate a loading instruction and initialization parameter set for the corresponding virtual instrument;

[0036] Decompose the paragraph structure intention and determine the segmentation logic of the prelude, verse and chorus and the duration configuration of each segment.

[0037] Furthermore, the model context protocol receives the semantic instruction and compiles it into a JSON command, including:

[0038] Parsing the action type in the semantic instruction to determine atomic music operations; the atomic music operations include track audio creation, virtual instrument loading, and automation configuration;

[0039] For the track audio creation, generate a first JSON command to create a MIDI track or an audio track;

[0040] For said virtual instrument loading, generating a second JSON command to load the virtual instrument and effector through a URI standardization scheme;

[0041] For the automation configuration, a third JSON command is generated, the parameter range is verified through the sandbox environment, and the automation curve is drawn.

[0042] Furthermore, it also includes

[0043] Check the tonal consistency of the chord sequence and the rhythmic rationality of the note arrangement; compare the style tags in the semantic instructions to verify whether the effect parameters and dynamic range match the preset style characteristics;

[0044] If all verifications fail, a feedback report of the mismatches is generated and sent back to the model context protocol to trigger instruction optimization.

[0045] The present invention has the following advantages and beneficial effects:

[0046] This patent adopts a three-layer collaborative architecture to achieve full-process closed-loop control of the intelligent music production system. The system uses the AI ​​interaction layer as the creative entrance. The natural language processing engine built based on the large language model (LLM) deeply integrates music field knowledge. The music semantic analysis module structures and deconstructs the vague creative intentions input by the user, accurately extracting core elements such as style tags, equipment requirements and chord structure. At the same time, it relies on the dynamic context manager to continuously maintain the project status snapshot of the digital audio workstation (DAW) at a refresh rate of no less than 5Hz, forming a real-time creative context that includes track configuration, effect chain parameters and timing positioning. In this process, the instruction preprocessor plays the role of a bridge between semantic and physical space, dynamically mapping abstract descriptions (such as "adding a sense of space") into a 42-dimensional parameter set such as reverberator decay time (Decay = 8.3s) and delay feedback amount (Feedback = 37%), ensuring the precise correspondence between artistic expression and technical implementation.

[0047] The protocol conversion layer establishes a two-way communication channel through the independently developed Model Context Protocol (MCP), using a hybrid data encapsulation format with a JSON body and nested binary attachments. This ensures command readability while supporting the efficient transmission of large amounts of data such as audio waveforms and spectral characteristics. The protocol defines a three-layer command classification and encoding system (0x01 Session Control / 0x02 Track Operation / 0x03 Clip Editing) and innovatively introduces a heartbeat packet mechanism to ensure the real-time reliability of the DAW connection through status verification every 200ms. At the data transmission level, the MCP protocol uses a differential compression algorithm to encode the DAW project status, reducing the transmission bandwidth of routine operation instructions to an average of 3.2KB / s, while supporting the synchronous transmission requirements of audio streams up to 18Mbps.

[0048] The DAW execution layer implements atomic operation control through a deeply encapsulated remote scripting interface. Its API instruction set covers basic operations such as creating audio tracks, loading virtual instruments, and configuring automation curves. Each instruction is verified for parameter range and dependency checks in a sandbox environment. The reverse state collector uses a multi-threaded architecture to capture track level peaks, device parameter snapshots, and timeline positioning information in parallel, building a real-time feedback loop from the physical layer to the cognitive layer. The specially designed exception fuse mechanism can immediately terminate the current transaction chain when a parameter out-of-bounds is detected (such as a MIDI velocity value > 127), and automatically roll back to the most recent valid state, ensuring the operational safety of the system in complex creative scenarios.

[0049] The three-layer architecture forms a closed-loop, bidirectional data flow through the TCP / IP protocol stack: The AI ​​layer semantically understands the user's natural language commands, generating a structured creative solution. The MCP protocol encapsulates these instructions into a standardized, machine-executable sequence, which is then translated into specific music engineering operations by the DAW execution layer. Simultaneously, the DAW's real-time status data is fed back into the AI ​​layer's dynamic context manager via a reverse channel, driving the large language model to dynamically optimize creative strategies based on the latest engineering environment. This collaborative mechanism of "creative expression - protocol conversion - physical execution - feedback optimization" achieves a deep integration of AI and professional music production tools at the semantic, protocol, and execution layers. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 An exemplary interaction diagram of an intelligent music production system based on a model context protocol provided by the present invention;

[0051] Figure 2 An exemplary flow chart of an intelligent music production method based on a model context protocol provided by the present invention. DETAILED DESCRIPTION

[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.

[0053] Figure 1 This is an exemplary interaction diagram of an intelligent music production system based on a model context protocol provided by the present invention. Figure 1 As shown in Figure 1, the intelligent music production system based on the Model Context Protocol consists of a large language model (AI interaction layer), a Model Context Protocol (protocol conversion layer), and a digital audio workstation execution layer (DAW execution layer). The AI ​​interaction layer builds a natural language processing engine based on the Large Language Model (LLM). The Model Context Protocol (MCP) enables two-way communication. The DAW execution layer implements atomic operations through remote scripting.

[0054] The large language model is used to deconstruct natural language instructions into structured operation logic and generate it based on a music knowledge base; the semantic instructions serve as an interaction carrier between the large language model and the model context protocol, and include at least device parameters and chord sequences.

[0055] Natural language instructions are text instructions entered by users that describe their music creation needs. For example, "Generate an electronic funk bassline with jazz chords." Structured operation logic breaks down natural language instructions into a set of executable music production tasks. For example, extracting style features, equipment requirements, and paragraph structure. A music knowledge base is a database that stores content related to music creation. For example, a music knowledge base can include style templates, chord combinations, and instrument parameters. Semantic instructions are structured instructions generated by a large language model. For example, an object containing a style label, chord sequence, and device parameters. Device parameters refer to the configuration parameters of a virtual instrument or effect. A chord sequence refers to a combination of chords arranged by key. For example, CG-Am-F. Based on the style templates in the music knowledge base, the creative style intention can be mapped to chord progressions and rhythmic pattern parameters.

[0056] The model context protocol is used to compile the semantic instructions into JSON commands and synchronize the working status data of the digital audio workstation to the large language model to update the model context.

[0057] JSON commands are standardized instruction frames used to drive DAWs to perform operations. JSON commands can include a header and a body. The header can include a transaction ID and command type. The body can include parameter payloads. For example, the "load compressor" instruction can be converted to

[0058] {"action":"load_instrument_or_effect","params":{"type":"effect","name":"Compressor"}}.

[0059] Model context is a real-time image of the DAW state, used for dynamic tuning of the large language model. Model context can include track configuration, effects chain parameters, and timing positioning. Track configuration can include track type and channel assignment. For example, creating a MIDI track and binding a piano sound. Effect chain parameters can include effect type and parameter settings. For example, a reverb decay time of 8.3s. Timing positioning can include timeline position and clip duration; a timeline position can be position="1.1.1". Synchronous state data captures the DAW state through a reverse state collector and transmits it back to the large language model after differential compression.

[0060] The digital audio workstation execution layer is used to convert the JSON command into a music operation to obtain a music project file.

[0061] A music operation is the smallest functional unit that can be executed by a DAW. A music operation includes at least one or more atomic operations, including basic operations such as creating audio tracks, loading virtual instruments, and configuring automation curves. Each instruction is verified for parameter ranges and dependencies in a sandbox environment. A music project file can refer to a complete music file, such as an Ableton Live project file. The DAW API is called through a remote scripting interface to convert the JSON command into a music operation, resulting in a music project file.

[0062] In some embodiments, the large language model includes a music semantic parsing module, an instruction preprocessor, and a dynamic context manager.

[0063] The music semantic parsing module is used to perform style classification on the natural language instructions, extract style tags, and decompose the structural instructions into MIDI sequence generation tasks and / or paragraph structures.

[0064] For example, the BERT model can be used to identify the "electro-funk" style keyword and extract the 16th-note rhythmic features. A style tag is a keyword that describes a musical style. A structural instruction is an instruction that describes the structure of a musical segment. For example, 8 bars of verse + 4 bars of chorus. A MIDI sequence generation task involves generating specific parameters for a note sequence. For example, pitch 60 (C4), velocity 80, and duration 1.0 second. A segment structure refers to the segmentation logic and duration of a musical segment. For example, an intro (4 bars) and a verse (8 bars).

[0065] The instruction preprocessor is used to dynamically map abstract device requirements into a device parameter set, obtain an LLM (Large Language Model) parsing result, and form a multi-dimensional parameter set.

[0066] An abstract device requirement is a user's vague description of the desired sound quality, such as warm bass. A device parameter set is a specific configuration for a virtual instrument, such as a low-pass filter cutoff frequency of 200Hz and tube saturation parameters. A multidimensional parameter set is composed of device parameter sets.

[0067] The dynamic context manager is used to maintain the working status data of the digital audio workstation.

[0068] Working status data can include track configuration, effects chain parameters, and timing positioning information. Track configuration refers to the properties and parameters of a track in a digital audio workstation (DAW), including track type, channel assignment, device loading, and parameter settings. Effect chain parameters refer to the specific configuration parameters of the virtual effects processor loaded for the track, which can include compressors, reverbs, and equalizers. Timing positioning information refers to data used to identify the position and timing attributes of musical elements on the timeline in a digital audio workstation (DAW).

[0069] In some embodiments, the Model Context Protocol (MCP) adopts a JSON-TCP hybrid architecture, including an MCP protocol encapsulation module, a command classification encoder, and a heartbeat packet controller.

[0070] The MCP protocol encapsulation module uses a differential compression algorithm to encode the working status data of the digital audio workstation. The command classification encoder is used to perform three-layer classification encoding on the semantic instructions to generate standardized instruction types. The three-layer classification encoding can refine the instruction type to adapt to DAW operations. The three-layer classification encoding can include session control, track operation and clip editing. Session control refers to managing the connection status and global session parameters of the digital audio workstation (DAW). Track operation refers to a set of operations for track creation, configuration and resource management. Clip editing refers to editing and control operations for MIDI clips (Clip) and notes. Standardized instruction types refer to instructions in standard form. For example, clip_creation (create Clip), set_tempo (set BPM), etc.

[0071] The heartbeat packet controller is used to send a status verification packet at a regular interval (eg, every 200 ms) to ensure the real-time reliability of the connection of the digital audio workstation.

[0072] The model context protocol also includes a state difference detection unit, which is used to compare project state changes, generate state difference reports and inject large language model context. Project state changes refer to state changes of a digital audio workstation (DAW) during the execution of music production operations. For example, real-time changes in project metadata, track status, device parameters, etc. The DAW state can be captured by the reverse state collector of the model context protocol (MCP), and the changes can be compared by the state difference detection unit. For example, device loading changes: when loading a new effector Ableton Reverb, the new_devices field in state_diff adds the device name Ableton Reverb. A state difference report refers to a collection of project state change data. For example, when the DAW executes the "load reverb effector" command, the state difference report contains the newly added device information. For example, to report a status difference for a device loading operation, the large language model sends the command "Load Ableton Reverb reverb effect on track 2." The DAW successfully loads the effect, and the device "Ableton Reverb" is added to track 2. The reverse state collector captures changes in the DAW browser and uses get_browser_items_at_path to query the specified path / Effects / Reverb to confirm the device's existence. The state difference detection unit compares the state before and after loading: Before loading, the effects chain for track 2 is empty; After loading, "Ableton Reverb" is added to the effects chain for track 2. A status difference report for the track 2 effects chain change is generated. To report a status difference for a global BPM adjustment, the large language model sends the command "Adjust global BPM from 120 to 128." The DAW updates the global BPM to 128, changing the project metadata. The reverse state collector calls get_session_info to obtain the project metadata and compares the BPM field. The state difference detection unit identifies the BPM change from 120 to 128 and generates a status difference report that includes the BPM change field and the project metadata update. For the status difference report of MIDIClip creation and note arrangement: the large language model sends the instruction "Create an 8-bar MIDI clip on track 3 and insert the note sequence [C2, D#2, F2]". The DAW creates a blank Clip at position "1.1.1" of track 3 with a length of 0:0:4.0 (8 bars) and inserts 3 notes. The reverse state collector obtains the Clip list and note data of track 3 through get_track_info. The state difference detection unit compares the original track 3 without Clips and the new track 3 with the specified note sequence after creation, and generates the corresponding state difference report.For the status difference report of the automated device parameter adjustment, the large language model sends the instruction "gradually change the filter cutoff frequency of the synthesizer on track 3 from 200Hz to 1kHz". The DAW draws the automation curve through set_device_parameter, and the parameter is linearly increased from 200Hz (0 seconds) to 1kHz (4 seconds). The reverse state collector obtains the device parameter change log, identifies the change in the cutoff_freq parameter of the "Retro Bass 3000" synthesizer, and the state difference detection unit extracts the keyframe data:

[0073] [0s:200Hz,4s:1000Hz], based on which a status difference report is generated.

[0074] In some embodiments, the model context protocol further includes an instruction frame construction unit, wherein the instruction frame construction unit is used to construct an instruction frame, including a generation module, an encapsulation module, and a response parsing module;

[0075] The generation module is used to generate a transaction_id in UUIDv4 format, define command_type and set a priority level of 1-5; the encapsulation module is used to encapsulate the target track index, MIDI note sequence and parameter constraints; the response parsing module is used to parse the status, execution_time and state_diff returned by the digital audio workstation.

[0076] In some embodiments, the digital audio workstation execution layer includes a remote script interface module, a digital audio workstation, and a reverse state collector.

[0077] The remote script interface module is used to convert JSON commands into atomic music operations of the audio workstation, and includes an API instruction execution engine; the API instruction set in the API instruction execution engine includes a track audio creation unit, a virtual instrument loading unit and an automation configuration unit.

[0078] Atomic music operations are used to implement specific music production tasks. Atomic music operations can include instrument loading, MIDI arrangement, effects chain configuration, and project parameter adjustment.

[0079] The API command set converts the JSON commands generated by the Model Context Protocol (MCP) into low-level commands executable by the DAW. For example, the API command set can be used for track management, device loading and control, MIDI and Clip operation management, parameter adjustment and synchronization status, and status acquisition and feedback.

[0080] The track audio creation unit is used to create a MIDI track or an audio track.

[0081] MIDI tracks are used to edit and process MIDI data. Audio tracks are used to record, play, and process audio wave files.

[0082] The virtual instrument loading unit loads virtual instruments and effects using a URI standardization scheme.

[0083] The automation control unit verifies the parameter range through the sandbox environment and draws the automation curve.

[0084] The Sandbox environment is an isolated parameter validation space used to verify parameter validity before executing DAW operations, preventing invalid parameters from causing system errors or project crashes. Parameter ranges refer to the legal range of values ​​for various DAW operation parameters, including minimum and maximum limits for device and project parameters. Automation curves are the trajectory of parameter changes over time. These curves define parameter values ​​at different points in time through keyframes or continuous interpolation, enabling dynamic effects in music production.

[0085] The reverse state collector is used to capture the engineering state data of the digital audio workstation.

[0086] Project status data can be used to reflect the project status of a digital audio workstation. Project status data can include track level peaks, device parameter snapshots, timeline positioning information, and project metadata. Track level peaks refer to the instantaneous maximum amplitude value of the audio signal in an audio track. Device parameter snapshots refer to the parameter configuration status of a virtual instrument or effects processor at a specific moment. Timeline positioning information identifies the position and timing attributes of musical elements on the DAW timeline. Project metadata is structured data that describes the overall properties of a music project and is used to record project status and configuration information.

[0087] In some embodiments, the digital audio workstation also includes an exception-fusing mechanism module, which is configured to terminate the current transaction chain and automatically roll back to the most recently valid state upon detecting a parameter out-of-bounds condition. A current transaction chain refers to an ordered sequence of related, atomic operations that collectively complete a task when a digital audio workstation (DAW) executes a music production operation. The most recently valid state refers to the project state after the last successful operation, as recorded by the DAW during the execution of the transaction chain. For example, when the large language model sends the instruction "set global BPM to 300," the range validator in the DAW execution layer detects that the BPM value of 300 exceeds the maximum allowed value of 250, triggering an exception-fusing mechanism. The "set_tempo" operation is immediately terminated to prevent the illegal parameter from being written to the project. The BPM is restored to its previous valid value of 1128.0 based on the state hash value of the most recently successful operation (e.g., "state_hash:a1 b2c3d4"), and the project metadata is synchronized via the reverse state collector.

[0088] In some embodiments, the large language model, model context protocol, and digital audio workstation establish a bidirectional data flow closed loop via the TCP / IP protocol stack, including a forward link module and a reverse link module. These components meet the following requirements: UTF-8 encoding is used for the transmission of natural language instructions to semantic instructions, ensuring the parsing of multilingual creative intent; a length prefix and CRC checksum are added to JSON command transmission to prevent data loss or parsing errors. Working status data synchronization uses an incremental transmission mode, uploading only changed parameters to reduce network bandwidth usage.

[0089] The forward link module encapsulates the semantic instructions of the large language model into JSON commands through the model context protocol, and converts them into atomic operations by the API instruction execution engine of the digital audio workstation;

[0090] The reverse link module transmits the engineering status data of the digital audio workstation back to the dynamic context manager of the large language model via the reverse status collector and the differential compression module of the model context protocol to drive the optimization of the creative strategy.

[0091] Figure 2 This is an exemplary flow chart of an intelligent music production method based on a model context protocol provided by the present invention. Figure 2 As shown, the intelligent music production method based on the model context protocol includes:

[0092] The large language model receives natural language instructions, deconstructs the natural language instructions into structured operation logic, and generates semantic instructions based on the music knowledge base; the semantic instructions serve as an interaction carrier between the large language model and the model context protocol, and include at least device parameters and chord sequences;

[0093] The model context protocol receives the semantic instruction and compiles it into a JSON command;

[0094] The model context protocol collects the working status data of the digital audio workstation and synchronizes it to the large language model to update the model context;

[0095] The digital audio workstation receives the JSON command, converts it into a music operation, and generates a music operation file.

[0096] In some embodiments, the large language model receives natural language instructions, deconstructs the natural language instructions into structured operation logic, and generates semantic instructions based on a music knowledge base, including:

[0097] Perform intent recognition on natural language instructions and extract intent parameters.

[0098] The intention parameters may at least include the creative style intention, the instrument requirement intention and the paragraph structure intention.

[0099] Based on the style template in the music knowledge base, the creative style intention is mapped into preset chord progression and rhythm type parameters.

[0100] A preset chord progression refers to a pre-stored chord progression combination in a music knowledge base that conforms to a specific musical style. Rhythmic pattern parameters are structured data describing the duration and dynamic patterns of notes. For example, to generate an electronic funk bassline: the natural language instruction "Generate an electronic funk bassline with jazz chords" extracts the creative style intent: electronic funk (16th-note rhythm), jazz chords (extended scale) and the equipment requirements: warm bass → matching the "Vintage Synth Bass" preset, applying a compressor (ratio 4:1). Through style template mapping, a preset chord progression is obtained: Based on the electronic funk style template, a I-IV-V chord progression is selected, and jazz elements are combined to adjust the chords to extended notes. Rhythmic pattern parameters: Extract the "16n / 0.8swing" rhythm pattern (16th-note swing rhythm, swing coefficient 0.8) from the style template, define the note duration as 16th notes, and the rhythm offset as 80% of the standard duration.

[0101] Parameters of the instrument requirement intention are parsed to generate a loading instruction and initialization parameter set of the corresponding virtual instrument.

[0102] A load command instructs a digital audio workstation (DAW) to call a specific virtual instrument plug-in and integrate it into the current music production project. An initialization parameter set is a set of configuration data related to the initial state of a virtual instrument. For example, given an electric guitar and a gentle style, a load command for the electric guitar is generated, and the initialization parameter set for the gentle style electric guitar is called to initialize the electric guitar.

[0103] Decompose the paragraph structure intention and determine the segmentation logic of the prelude, verse and chorus and the duration configuration of each segment.

[0104] The model context protocol receives the semantic instruction and compiles it into a JSON command, including:

[0105] Parse the action type in the semantic instruction and determine the atomic music operation; the atomic music operation includes track audio creation, virtual instrument loading and automation configuration; for the track audio creation, generate a first JSON command to create a MIDI track or audio track; for the virtual instrument loading, generate a second JSON command to load virtual instruments and effects through the URI standardization scheme; for the automation configuration, generate a third JSON command to verify the parameter range through the sandbox environment and draw the automation curve. The first JSON command refers to a JSON format instruction for creating a MIDI track or audio track. The second JSON command refers to a JSON instruction for loading a virtual instrument or effect through the URI standardization scheme. The third JSON command refers to a JSON instruction for generating a parameter automation curve after verifying the legitimacy of the parameters through the sandbox environment.

[0106] In some embodiments, the intelligent music production method based on the model context protocol also includes: checking the tonal consistency of the chord sequence and the rhythmic rationality of the note arrangement; comparing the style tags in the semantic instructions to verify whether the effect parameters and dynamic range match the preset style characteristics.

[0107] Tonality consistency means that the pitch relationship of each chord in the chord sequence must conform to the scale rules of a specific tonality (such as C major, A minor), avoiding pitch levels that conflict with the tonality and ensuring the auditory harmony and unity of the music. The rhythmic rationality of the note arrangement means that the duration, strength and weakness rules and arrangement of the notes must conform to the rhythmic characteristics of the target style. Effect parameters and dynamic range matching preset style characteristics means that the effect parameters (such as compressor ratio, reverb decay time) and dynamic range (volume fluctuation amplitude) must be consistent with the sound characteristics of the style. For example, warm bass needs to be matched with low-pass filtering and tube saturation effects, and aggressive rock style requires high compression ratio and distortion effects.

[0108] If all verifications fail, a feedback report of the mismatches is generated and sent back to the model context protocol to trigger instruction optimization.

[0109] Example 1

[0110] The Model Context Protocol (MCP) uses a JSON-TCP hybrid architecture:

[0111]

[0112]

[0113] Key field of the MCP protocol: command_type, technical function: an enumeration type field that defines the operation intention.

[0114] Track and device management: create_midi_track, function: dynamically create MIDI tracks, support index positioning and channel allocation, implementation: allocate MIDI track resources through DAW API (such as Ableton Live's Python API), default binding piano sound, support 16 MIDI channel selection.

[0115] create_audio_track, Purpose: Creates an audio track for recording or playing wave files; Features: Supports mono / stereo configurations and automatically handles sample rate conversion. load_instrument_or_effect, Purpose: Loads a virtual instrument or effect (e.g., "Operator / Drums / 808Core"); Implementation: Identifies cross-platform devices through standardized URI schemes and supports VST / AU plugin formats.

[0116] Clip and Note Operations: create_clip, Function: Creates a blank MIDI Clip at a specified track position; Parameters: Timeline position (e.g., position="1.1.1"), duration (duration="0:0:0.5"). add_notes_to_clip, Function: Adds MIDI note events to a Clip; Data structure: Contains fields such as pitch (pitch=60), velocity (velocity=80), and duration. fire_clip / stop_clip, Function: Triggers or stops Clip playback, supporting real-time synchronization; Underlying mechanism: Sends control signals to DAW ports (e.g., ports 11000 / 11001 in Ableton Live) via the OSC protocol.

[0117] Parameter Control: set_tempo, Function: Sets the global BPM value (e.g., 128.0); Extension: Supports dynamic BPM adjustment with Spread Spectrum Clocking (SSC) to prevent electromagnetic interference. set_device_parameter, Function: Adjusts device parameters (e.g., filter cutoff frequency); Encoding: Parameter IDs are mapped to DAW-internal identifiers (e.g., DeviceParameterID).

[0118] Playback control: start_playback / stop_playback, function: start / stop the global playback engine; mode: supports MODE_STATIC (preload buffer) and MODE_STREAM (real-time streaming).

[0119] Browser and Metadata: get_browser_tree, Purpose: Get the DAW browser tree structure (such as instrument / effect classification); Implementation: Recursively traverse the directory hierarchy and dynamically load device metadata.

[0120] get_browser_items_at_path, function: query the browser items under the specified path (such as

[0121] / Instruments / Synths); Optimization: Cache mechanism accelerates high-frequency access.

[0122] Status query: get_session_info, Purpose: Retrieves session metadata (BPM, track count, device chain); Output format: JSON structure containing hashed status value, used for incremental synchronization. get_track_info, Purpose: Retrieves track details (such as volume, pan, and inserted effects); Technical details: Fast location via Track.index.

[0123] Nested structure:

[0124]

[0125]

[0126] Technical Function: Implement incremental state comparison through state_hash to trigger model tuning.

[0127] Error handling mechanism: Code-level errors: Python exception stack capture (including thread-safe isolation).

[0128] Business-level errors: parameter out of bounds / resource not present, etc. Technical role: Establish an error classification system (connection level / protocol level / application level).

[0129] The JSON-TCP hybrid architecture achieves efficient coordination between data readability and transmission reliability by combining the core advantages of the two protocols. It is particularly suitable for scenarios that require mixed transmission of structured metadata and binary streams (such as audio streaming transmission).

[0130] Example 2

[0131] Remote script control module:

[0132] Ableton Live API deep package, including:

[0133]

[0134]

[0135] Example 3

[0136] Take the example of generating an electronic funk bassline with jazz chords:

[0137] Semantic analysis:

[0138] Extract key features: jazz chords (extended scales), electronic funk (16th note rhythm).

[0139] Device Matching: Select the "Vintage Synth Bass" preset and load the compressor (ratio 4:1).

[0140] Protocol Generation:

[0141]

[0142] DAW Execution:

[0143] Create a MIDI clip on track 3 and insert the note sequence [C2(0.0s),D#2(0.25s),...].

[0144] Load the "Retro Bass 3000" synthesizer and automate the filter cutoff frequency (from 200Hz to 1kHz).

[0145] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. An intelligent music production system based on model context protocol, characterized in that Includes large language models, model context protocols, and digital audio workstation execution layers; The large language model is used to deconstruct natural language instructions into structured operation logic and generate semantic instructions based on a music knowledge base; the semantic instructions include at least device parameters and chord sequences; The model context protocol is used to compile the semantic instructions into JSON commands and synchronize the working status data of the digital audio workstation in the digital audio workstation execution layer to the large language model to update the model context; the model context includes the track configuration, effect chain parameters and timing positioning information of the digital audio workstation; The digital audio workstation execution layer is used to convert the JSON command into a music operation to obtain a music operation file.

2. The intelligent music production system based on the model context protocol according to claim 1, characterized in that The large language model includes a music semantic parsing module, an instruction preprocessor and a dynamic context manager; The music semantic parsing module is used to perform style classification on the natural language instructions, extract style tags, and decompose the structural instructions into MIDI sequence generation tasks and / or paragraph structures; The instruction preprocessor is used to dynamically map the abstract device requirements into a device parameter set to form a multi-dimensional parameter set; The dynamic context manager is used to maintain the working status data of the digital audio workstation; the working status data includes track configuration, effect chain parameters and timing positioning information.

3. The intelligent music production system based on the model context protocol according to claim 1, characterized in that The model context protocol includes an MCP protocol encapsulation module, a command classification encoder and a heartbeat packet controller; The MCP protocol encapsulation module uses a differential compression algorithm to encode the working status data of the digital audio workstation; The command classification encoder is used to perform three-layer classification encoding on the semantic instructions to generate standardized instruction types; the three-layer classification encoding includes session control, track operation and segment editing; The heartbeat packet controller is used to periodically send status verification packets.

4. The intelligent music production system based on the model context protocol according to claim 1, characterized in that The digital audio workstation execution layer includes a remote script interface module, a digital audio workstation and a reverse state collector; The remote script interface module is used to convert JSON commands into atomic music operations of the audio workstation, including an API instruction execution engine; the API instruction set in the API instruction execution engine includes a track audio creation unit, a virtual instrument loading unit, and an automation configuration unit; The track audio creation unit is used to create a MIDI track or an audio track; The virtual instrument loading unit loads virtual instruments and effects using a URI standardization scheme; The automation configuration unit verifies the parameter range through the sandbox environment and draws the automation curve; The reverse state collector is used to capture the engineering state data of the digital audio workstation; the engineering state data includes track level peaks, device parameter snapshots, timeline positioning information and engineering metadata.

5. The intelligent music production system based on the model context protocol according to claim 1, characterized in that The digital audio workstation execution layer further includes an abnormal fusing mechanism module; the abnormal fusing mechanism module is used to terminate the current transaction chain and automatically roll back to the most recent valid state when a parameter out of bounds is detected.

6. The intelligent music production system based on the model context protocol according to claim 1, characterized in that The large language model, model context protocol and digital audio workstation execution layer construct a bidirectional data flow closed loop through the TCP / IP protocol stack, including a forward link module and a reverse link module; The forward link module encapsulates the semantic instructions of the large language model into JSON commands through the model context protocol, and converts them into atomic operations by the API instruction execution engine of the digital audio workstation execution layer; The reverse link module transmits the engineering status data of the digital audio workstation back to the dynamic context manager of the large language model via the reverse status collector and the differential compression module of the model context protocol to drive the optimization of the creative strategy.

7. An intelligent music production method based on model context protocol, characterized in that include: The large language model receives natural language instructions, deconstructs the natural language instructions into structured operation logic, and generates semantic instructions based on a music knowledge base; the semantic instructions include at least device parameters and chord sequences; The model context protocol receives the semantic instruction and compiles it into a JSON command; The model context protocol collects the working status data of the digital audio workstation and synchronizes it to the large language model to update the model context; The digital audio workstation execution layer receives the JSON command, converts it into a music operation, and generates a music operation file.

8. The intelligent music production method based on model context protocol according to claim 7, characterized in that: The large language model receives natural language instructions, deconstructs the natural language instructions into structured operation logic, and generates semantic instructions based on a music knowledge base, including: Performing intent recognition on natural language instructions and extracting intent parameters; the intent parameters include at least creative style intent, instrument requirement intent, and paragraph structure intent; Based on the style template in the music knowledge base, mapping the composition style intention into preset chord progression and rhythm type parameters; Performing parameter analysis on the instrument requirement intention to generate a loading instruction and initialization parameter set for the corresponding virtual instrument; Decompose the paragraph structure intention and determine the segmentation logic of the prelude, verse and chorus and the duration configuration of each segment.

9. The intelligent music production method based on model context protocol according to claim 7, characterized in that: The model context protocol receives the semantic instruction and compiles it into a JSON command, including: Parsing the action type in the semantic instruction to determine atomic music operations; the atomic music operations include track audio creation, virtual instrument loading, and automation configuration; For the track audio creation, generate a first JSON command to create a MIDI track or an audio track; For said virtual instrument loading, generating a second JSON command to load the virtual instrument and effector through a URI standardization scheme; For the automation configuration, a third JSON command is generated, the parameter range is verified through the sandbox environment, and the automation curve is drawn.

10. The intelligent music production method based on model context protocol according to claim 7, characterized in that: Also includes Check the tonal consistency of the chord sequence and the rhythmic rationality of the note arrangement; compare the style tags in the semantic instructions to verify whether the effect parameters and dynamic range match the preset style characteristics; If all verifications fail, a feedback report of the mismatches is generated and sent back to the model context protocol to trigger instruction optimization.

Citation Information

Patent Citations

  • Method, device, equipment and product for executing task

    CN120013493A

  • System

    JP2025050297A

  • A system for, and a method of, facilitating music composition and music performance

    WO2024132867A1

Cited By

  • Configuration method and device of storage cluster, electronic equipment and storage medium

    CN120762781A

  • Blind music creation method and electronic equipment

    CN120954363A

  • Blind music creation method and electronic device

    CN120954363B

  • Construction method and equipment of central air conditioner management and control system based on MCP protocol and medium

    CN121029781A

  • Method and device for constructing central air conditioning management and control system based on mcp protocol, and medium

    CN121029781B